Week 1 — Lecture Outline · Statistics, Data & Study Design
Course: Introduction to Statistics (18-week generic edition)
Objectives covered: Objective 1 — Distinguish populations from samples and identify appropriate sampling and study designs.
SLOs touched: A (reason quantitatively from data) · B (communicate results to a non-technical audience)
Meeting pattern: planned as 2 sessions × ~75 min = ~150 min. Segment minutes below total ~150; scale them to your own pattern.
Week at a Glance
| The week's big question | "Where do data come from, and when can the numbers be trusted to speak for more people than we actually measured?" |
| By the end of the week, students can… | (1) tell a population from a sample, and a parameter from a statistic; (2) classify a variable by its level of measurement (nominal / ordinal / interval / ratio); (3) name a sampling method and say whether it's likely to be biased; (4) tell an observational study from an experiment and explain why correlation isn't causation. |
| Key vocabulary | population, sample, parameter, statistic, census, variable, categorical vs. quantitative, nominal/ordinal/interval/ratio, simple random sample (SRS), stratified, cluster, systematic, convenience/voluntary-response, bias (undercoverage, nonresponse, response, voluntary-response), observational study, experiment, confounding variable |
| Materials | slides (Deck 1), the Week 1 chapter, the week's readings + video links, a spreadsheet (Google Sheets or Excel), the student's chatbot for the AI-critique moment and the tutorial |
| Timing note | 8 segments, ~150 min total. Session 1 = Segments 1–4 (~75). Session 2 = Segments 5–8 (~75). |
Segment 1 — Hook & the Promise (8 min) · Session 1 opens
Hook. "Think about the last 24 hours. Did you rate anything — a driver, a video, a product? Did an app count anything about you — steps, minutes watched, money spent?" Give them ten seconds. Then: every one of those numbers is now data in somebody's spreadsheet.
- Someone, somewhere, is using numbers like yours to make a claim about millions of people they never met: "riders are satisfied," "viewers prefer this," "customers spend more."
- "This course is about the machinery behind those claims — how a handful of measurements becomes a statement about everyone, and how to tell the honest versions from the garbage."
The promise (write it on the board): "By the end of this week you can look at any statistic in the wild — a poll, a star rating, a 'studies show' — and ask the three questions that decide whether it deserves your trust: Who was measured? How were they picked? What was actually recorded?"
Why it matters line (memory hook): "Statistics is not about the math. It's about trusting a number that describes people you didn't count."
Segment 2 — Population vs. Sample, Parameter vs. Statistic (20 min)
Plain language first.
- A population is everyone (or everything) we want to know about. A sample is the part we actually measured. We measure a sample because measuring the whole population is usually impossible, too slow, or too expensive.
- A census is the rare case where we do measure the whole population.
- Two more words, same split: a number that describes the population is a parameter; the matching number computed from the sample is a statistic.
Memory hook (put it on a slide):
Population → Parameter. Sample → Statistic. "The letters line up."
One fully worked example (do every step out loud).
Claim: "74% of a city's transit riders are satisfied with the new schedule."
The transit agency cares about all 24,000 weekday riders (the population). From the fare-card registry it randomly selects 800 riders to survey (the sample). Of those 800, 592 say they're satisfied.
- Sample statistic: 592 ÷ 800 = 0.74 → 74%. This is a statistic (it came from the sample).
- The true satisfaction rate among all 24,000 riders — what we'd get if every rider answered honestly — is the parameter. We never see it directly; 74% is our best estimate of it.
- Notation preview (notation comes after the idea): the population proportion is written p (parameter); the sample proportion is p̂ ("p-hat," statistic). "The hat means 'measured,' not 'true.'"
Land the key idea: the statistic (74%) is what we have; the parameter is what we want. The whole rest of the course is the bridge from one to the other.
Segment 3 — Levels of Measurement (25 min)
Plain language first. Before you can summarize a variable, you have to know what kind of thing it is. Two big families, four levels.
- Categorical (qualitative) — labels or groups. Two levels:
- Nominal — names with no order. (Major, blood type, delivery zone, yes/no, a driver's ID number.)
- Ordinal — ordered categories, but the gaps aren't equal/measurable. (Letter grade, "small/medium/large," a 1–5 star rating.)
- Quantitative (numerical) — actual amounts you can do arithmetic on. Two levels:
- Interval — ordered, equal gaps, but no true zero (zero doesn't mean "none"). (Temperature in °F/°C, the calendar year a phone model was released.)
- Ratio — ordered, equal gaps, and a true zero, so ratios make sense. (Height, weight, distance, income, the count of anything.)
Memory hook: N–O–I–R (like the French word for "black," noir) — Nominal, Ordinal, Interval, Ratio, in order of how much math they permit.
One fully worked example (classify and justify each).
A food-delivery app records, for each order: driver ID number, customer star rating (1–5), food temperature at drop-off (°F), delivery distance (miles), tip (dollars), delivery zone (North / East / South / West).
- Driver ID number → nominal (a number that names — averaging IDs is nonsense).
- Star rating → ordinal (5 beats 4, but the 4→5 jump isn't a measured, equal gap).
- Drop-off temperature (°F) → interval (0°F isn't "no heat").
- Delivery distance → ratio (0 miles means none; 4 miles is twice 2).
- Tip → ratio (a true zero, sadly for the driver).
- Delivery zone → nominal.
The "test" to give students: Does zero mean "none"? If yes → ratio. Equal gaps but zero is arbitrary? → interval. Ordered labels, fuzzy gaps? → ordinal. Just names? → nominal.
Segment 4 — Misconceptions + Quick Interaction (22 min) · Session 1 closes (~75)
Name the misconceptions out loud, then cure each:
- ❌ "If it's a number, it's quantitative."
✅ Cure: zip codes, phone numbers, bus route numbers, and ID numbers are numbers that label. The test isn't "is it a number," it's "does arithmetic mean anything." You can't average bus routes. - ❌ "A bigger sample is automatically a better sample."
✅ Cure: size doesn't fix bias. A huge sample drawn the wrong way is confidently wrong. (The most famous polling disaster in history — next session — had 2.4 million responses.) "How you pick beats how many you pick." - ❌ "Population means a lot of people."
✅ Cure: the population is whoever the question is about — it could be the 28 students in one section, or every order at one coffee shop. It's defined by the claim, not by being large. - ❌ "Sample and population are fixed labels."
✅ Cure: they're roles. This term's lab submissions can be a sample of all submissions ever, or the whole population of this term's submissions. Same data, different question.
Interaction — Think-Pair-Share (rapid-fire classification, ~12 min):
Put 6 variables on a slide; students decide the level of measurement solo (30 sec), compare with a neighbor (1 min), then the class votes by fingers (1 = nominal … 4 = ratio). Suggested items: a bus route number · a hotel's star rating · temperature in °C · monthly rent · blood type · the calendar year a phone model was released.
(Answers: nominal · ordinal · interval · ratio · nominal · interval.)
Debrief the two that always split the room: bus route number (nominal, not ratio — route 40 isn't "twice" route 20) and calendar year (interval — year zero is an arbitrary marker, not "no time").
Segment 5 — How We Pick: Sampling Methods (25 min) · Session 2 opens
Hook back in: "Last session: a statistic is only as good as the sample behind it. Today: how do you pick a sample you can actually trust?"
Plain language first — the gold standard.
- Simple Random Sample (SRS): every individual has an equal chance, and every group of that size is equally likely. Like names in a hat. This is the benchmark every other method is judged against.
The four probability methods (each with a one-line picture):
- Simple random — names in a hat. Fair, but can miss small groups by luck.
- Stratified — split the population into meaningful groups (strata) first, then random-sample within each. Use when you want every subgroup represented.
- Cluster — split into natural groups (clusters), randomly pick whole clusters, measure everyone in them. Cheaper when the population is spread out.
- Systematic — order the list, pick every k-th individual after a random start. Easy when you have a list or a line.
Memory hook: Stratified = sample within every group. Cluster = sample whole groups. (The classic mix-up; say it twice.)
The methods to distrust (name them as traps):
- Convenience sample — whoever's easy to reach (your friends, the front row). Cheap and almost always biased.
- Voluntary response — people opt in (online polls, "tap to rate"). The angry and the passionate over-reply.
Worked example (one scenario, four designs):
Goal: a public library system wants the average number of weekly visits across its 60,000 cardholders.
- Survey people walking into the main branch on one morning → convenience, and worse, biased toward frequent visitors (you sampled at the library).
- Email a random 500 pulled from the full cardholder database → SRS, the trustworthy move.
- Want every branch's community represented → random-sample within each branch's cardholders → stratified.
- Cheaper field option: randomly pick 2 of the 5 branches and survey every visitor there this week → cluster (and note what it risks: those 2 branches may not represent all 5).
Segment 6 — Bias, and a Famous Failure (18 min)
Plain language: bias is error baked into the method — it pushes results in the same wrong direction no matter how big the sample gets. Four kinds to recognize:
- Undercoverage — part of the population can't be reached or is left out of the frame.
- Nonresponse bias — the people who don't answer differ from those who do.
- Response bias — the question or setting pushes answers (leading wording, sensitive topics, who's asking).
- Voluntary-response bias — the opt-in crowd isn't typical.
The famous failure (the story that sticks):
The 1936 Literary Digest presidential poll. The magazine mailed 10 million ballots and got 2.4 million back — a colossal sample — and confidently predicted Landon would beat Roosevelt. Roosevelt won in a landslide.
What went wrong? The mailing list came from car registrations and telephone directories — in the Depression, that skewed wealthy (undercoverage) — and only motivated people mailed ballots back (nonresponse). Meanwhile George Gallup polled a few thousand people chosen well and called it correctly.
The lesson, in one line: 2.4 million badly-chosen people lost to a few thousand well-chosen ones. Method beats size.
Callback: this is the cure to Segment 4's "bigger is better" misconception. Point back to it explicitly.
Segment 7 — Observational Study vs. Experiment; Correlation ≠ Causation (20 min)
Plain language first.
- In an observational study, you watch and record — you don't change anything. (Survey people about their habits; track what happens.)
- In an experiment, you deliberately impose a treatment and compare. (Randomly assign some customers a new app layout, others the old one, compare usage.)
- Only an experiment with random assignment can support a cause-and-effect claim. Observational studies can show a link, never prove the arrow.
The reason — confounding (worked example):
Headline: "People who use a fitness tracker take more daily steps." (Observational — nobody was assigned a tracker.)
A confounding variable is a third thing tangled with both: people who already exercise are exactly the people who buy trackers. The data can't separate "tracker → steps" from "already-active → tracker AND steps."
Picture it: tracker — steps looks like a straight arrow, but a hidden "already active" pulls both strings.
Memory hook: "Correlation is a handshake, not a push." Two things moving together ≠ one shoving the other.
Misconception + cure:
- ❌ "They found a strong correlation, so X causes Y."
✅ Cure: ask "could a third variable explain both?" and "was anything randomly assigned?" If nothing was assigned, you have a link, not a cause.
Quick mini-debate (genuinely arguable, ~4 min): "A meal-kit company finds its subscribers cook at home more often than non-subscribers. Should it advertise 'our kits make you cook more'?" Have students argue both sides; surface the confounder (people who subscribe already like cooking) — and ask what experiment would settle it (randomly assign free kits).
Segment 8 — Technology Workflow + AI-Critique, Callback & Hand-off (12 min) · Session 2 closes (~75)
Technology workflow — draw an SRS in a spreadsheet (exact steps):
1. Put your list of names/IDs in column A (say A2:A101 for 100 people).
2. In B2 type =RAND() and fill down to B101 — a random number beside each person.
3. Select columns A:B → Data ▸ Sort range → sort by column B.
4. The top 10 rows are your simple random sample. (Re-sorting reshuffles — that's the randomness working.)
- Google Sheets: =RAND(); Excel identical. For a quick single pick, =RANDBETWEEN(1,100) also works.
AI-critique moment (students verify, not consume):
Paste this to your chatbot: "Classify these variables by level of measurement: a bus route number, a hotel's 1–5 star rating, temperature in °C, monthly rent."
Then check its work against NOIR. Chatbots often call a route number "ratio" (it's nominal) and stumble on °C ("ratio"? — it's interval, no true zero). Your job all term: the tool drafts, you judge. This is exactly how the weekly Lecture Tutorial and Data Lab work — you'll catch the model, not trust it.
Callback + tease:
- Callback: "Every trustworthy number this term rides on this week — who was measured and how they were picked."
- Tease next week: "Now that we can get good data, Week 2 is the first thing we do with it: turn a pile of numbers into a picture — histograms, shapes, and the graphs that lie."
Hand-off (the week's work):
- Chapter 1 (the primary reading) if they haven't read it yet — then Lecture Tutorial 1 (AI tutor, share-link submission) — population/sample, NOIR, sampling & bias.
- Data Lab 1 (real dataset: penguins!) · Quiz 1 (end of week) · Discussion 1 (the star-rating question) · Assignment 1 (AI-coached).
Instructor FAQ — Common Stumbles
| Student says / does | Quick cure |
|---|---|
| "Is a driver ID quantitative? It's a number." | Apply the test — does averaging it mean anything? No. It's a nominal label that happens to look numeric. |
| "What's the difference between interval and ratio again?" | One question: does zero mean "none"? Yes → ratio (rent, distance, counts). No, zero is just a mark → interval (°F, °C, calendar year). |
| Confuses stratified and cluster. | Stratified = sample within every group; cluster = randomly pick whole groups and measure all of them. Stratified buys representation; cluster buys cheapness. |
| "The poll had 2 million responses, so it's reliable." | Size never fixes bias. The 1936 Literary Digest poll: 2.4 million responses, wrong winner. Method beats size. |
| Calls a strong correlation "proof" of cause. | Ask: was anything randomly assigned? If no, it's observational → a link, not a cause. Hunt the confounder. |
| "Population = a big group." | The population is whoever the question is about — could be 25 people. It's a role, not a size. |
| Thinks a census is a kind of sample. | A census measures the whole population — the opposite of sampling. Rare because it's usually too costly or slow. |
| Picks a convenience sample because it's "random enough." | Convenience ≠ random. "Whoever's nearby" has a hidden pattern (people at the library visit the library a lot). Only a chance-based method earns the word random. |
Scope flag
This outline stays within Objective 1. The °C/calendar-year interval-vs-ratio nuance and the 1936 Literary Digest case are added context (not strictly required by the objective) — kept because they make the misconceptions stick; cut them for a leaner session.