Back to the Introduction to Statistics outline The Course Maker
Introduction to Statistics outline
Week 1 · Lecture outline

Week 1 — Lecture Outline · Statistics, Data & Study Design

Introduction to Statistics Generic evergreen edition

Course: Introduction to Statistics (18-week generic edition)
Objectives covered: Objective 1 — Distinguish populations from samples and identify appropriate sampling and study designs.
SLOs touched: A (reason quantitatively from data) · B (communicate results to a non-technical audience)
Meeting pattern: planned as 2 sessions × ~75 min = ~150 min. Segment minutes below total ~150; scale them to your own pattern.


Week at a Glance

The week's big question "Where do data come from, and when can the numbers be trusted to speak for more people than we actually measured?"
By the end of the week, students can… (1) tell a population from a sample, and a parameter from a statistic; (2) classify a variable by its level of measurement (nominal / ordinal / interval / ratio); (3) name a sampling method and say whether it's likely to be biased; (4) tell an observational study from an experiment and explain why correlation isn't causation.
Key vocabulary population, sample, parameter, statistic, census, variable, categorical vs. quantitative, nominal/ordinal/interval/ratio, simple random sample (SRS), stratified, cluster, systematic, convenience/voluntary-response, bias (undercoverage, nonresponse, response, voluntary-response), observational study, experiment, confounding variable
Materials slides (Deck 1), the Week 1 chapter, the week's readings + video links, a spreadsheet (Google Sheets or Excel), the student's chatbot for the AI-critique moment and the tutorial
Timing note 8 segments, ~150 min total. Session 1 = Segments 1–4 (~75). Session 2 = Segments 5–8 (~75).

Segment 1 — Hook & the Promise (8 min) · Session 1 opens

Hook. "Think about the last 24 hours. Did you rate anything — a driver, a video, a product? Did an app count anything about you — steps, minutes watched, money spent?" Give them ten seconds. Then: every one of those numbers is now data in somebody's spreadsheet.

  • Someone, somewhere, is using numbers like yours to make a claim about millions of people they never met: "riders are satisfied," "viewers prefer this," "customers spend more."
  • "This course is about the machinery behind those claims — how a handful of measurements becomes a statement about everyone, and how to tell the honest versions from the garbage."

The promise (write it on the board): "By the end of this week you can look at any statistic in the wild — a poll, a star rating, a 'studies show' — and ask the three questions that decide whether it deserves your trust: Who was measured? How were they picked? What was actually recorded?"

Why it matters line (memory hook): "Statistics is not about the math. It's about trusting a number that describes people you didn't count."


Segment 2 — Population vs. Sample, Parameter vs. Statistic (20 min)

Plain language first.
- A population is everyone (or everything) we want to know about. A sample is the part we actually measured. We measure a sample because measuring the whole population is usually impossible, too slow, or too expensive.
- A census is the rare case where we do measure the whole population.
- Two more words, same split: a number that describes the population is a parameter; the matching number computed from the sample is a statistic.

Memory hook (put it on a slide):

Population → Parameter. Sample → Statistic. "The letters line up."

One fully worked example (do every step out loud).

Claim: "74% of a city's transit riders are satisfied with the new schedule."
The transit agency cares about all 24,000 weekday riders (the population). From the fare-card registry it randomly selects 800 riders to survey (the sample). Of those 800, 592 say they're satisfied.
- Sample statistic: 592 ÷ 800 = 0.74 → 74%. This is a statistic (it came from the sample).
- The true satisfaction rate among all 24,000 riders — what we'd get if every rider answered honestly — is the parameter. We never see it directly; 74% is our best estimate of it.
- Notation preview (notation comes after the idea): the population proportion is written p (parameter); the sample proportion is ("p-hat," statistic). "The hat means 'measured,' not 'true.'"

Land the key idea: the statistic (74%) is what we have; the parameter is what we want. The whole rest of the course is the bridge from one to the other.


Segment 3 — Levels of Measurement (25 min)

Plain language first. Before you can summarize a variable, you have to know what kind of thing it is. Two big families, four levels.

  • Categorical (qualitative) — labels or groups. Two levels:
  • Nominal — names with no order. (Major, blood type, delivery zone, yes/no, a driver's ID number.)
  • Ordinal — ordered categories, but the gaps aren't equal/measurable. (Letter grade, "small/medium/large," a 1–5 star rating.)
  • Quantitative (numerical) — actual amounts you can do arithmetic on. Two levels:
  • Interval — ordered, equal gaps, but no true zero (zero doesn't mean "none"). (Temperature in °F/°C, the calendar year a phone model was released.)
  • Ratio — ordered, equal gaps, and a true zero, so ratios make sense. (Height, weight, distance, income, the count of anything.)

Memory hook: N–O–I–R (like the French word for "black," noir) — Nominal, Ordinal, Interval, Ratio, in order of how much math they permit.

One fully worked example (classify and justify each).

A food-delivery app records, for each order: driver ID number, customer star rating (1–5), food temperature at drop-off (°F), delivery distance (miles), tip (dollars), delivery zone (North / East / South / West).
- Driver ID number → nominal (a number that names — averaging IDs is nonsense).
- Star rating → ordinal (5 beats 4, but the 4→5 jump isn't a measured, equal gap).
- Drop-off temperature (°F) → interval (0°F isn't "no heat").
- Delivery distance → ratio (0 miles means none; 4 miles is twice 2).
- Tip → ratio (a true zero, sadly for the driver).
- Delivery zone → nominal.

The "test" to give students: Does zero mean "none"? If yes → ratio. Equal gaps but zero is arbitrary? → interval. Ordered labels, fuzzy gaps? → ordinal. Just names? → nominal.


Segment 4 — Misconceptions + Quick Interaction (22 min) · Session 1 closes (~75)

Name the misconceptions out loud, then cure each:

  • "If it's a number, it's quantitative."
    Cure: zip codes, phone numbers, bus route numbers, and ID numbers are numbers that label. The test isn't "is it a number," it's "does arithmetic mean anything." You can't average bus routes.
  • "A bigger sample is automatically a better sample."
    Cure: size doesn't fix bias. A huge sample drawn the wrong way is confidently wrong. (The most famous polling disaster in history — next session — had 2.4 million responses.) "How you pick beats how many you pick."
  • "Population means a lot of people."
    Cure: the population is whoever the question is about — it could be the 28 students in one section, or every order at one coffee shop. It's defined by the claim, not by being large.
  • "Sample and population are fixed labels."
    Cure: they're roles. This term's lab submissions can be a sample of all submissions ever, or the whole population of this term's submissions. Same data, different question.

Interaction — Think-Pair-Share (rapid-fire classification, ~12 min):
Put 6 variables on a slide; students decide the level of measurement solo (30 sec), compare with a neighbor (1 min), then the class votes by fingers (1 = nominal … 4 = ratio). Suggested items: a bus route number · a hotel's star rating · temperature in °C · monthly rent · blood type · the calendar year a phone model was released.
(Answers: nominal · ordinal · interval · ratio · nominal · interval.)
Debrief the two that always split the room: bus route number (nominal, not ratio — route 40 isn't "twice" route 20) and calendar year (interval — year zero is an arbitrary marker, not "no time").


Segment 5 — How We Pick: Sampling Methods (25 min) · Session 2 opens

Hook back in: "Last session: a statistic is only as good as the sample behind it. Today: how do you pick a sample you can actually trust?"

Plain language first — the gold standard.
- Simple Random Sample (SRS): every individual has an equal chance, and every group of that size is equally likely. Like names in a hat. This is the benchmark every other method is judged against.

The four probability methods (each with a one-line picture):
- Simple random — names in a hat. Fair, but can miss small groups by luck.
- Stratified — split the population into meaningful groups (strata) first, then random-sample within each. Use when you want every subgroup represented.
- Cluster — split into natural groups (clusters), randomly pick whole clusters, measure everyone in them. Cheaper when the population is spread out.
- Systematic — order the list, pick every k-th individual after a random start. Easy when you have a list or a line.

Memory hook: Stratified = sample within every group. Cluster = sample whole groups. (The classic mix-up; say it twice.)

The methods to distrust (name them as traps):
- Convenience sample — whoever's easy to reach (your friends, the front row). Cheap and almost always biased.
- Voluntary response — people opt in (online polls, "tap to rate"). The angry and the passionate over-reply.

Worked example (one scenario, four designs):

Goal: a public library system wants the average number of weekly visits across its 60,000 cardholders.
- Survey people walking into the main branch on one morning → convenience, and worse, biased toward frequent visitors (you sampled at the library).
- Email a random 500 pulled from the full cardholder database → SRS, the trustworthy move.
- Want every branch's community represented → random-sample within each branch's cardholders → stratified.
- Cheaper field option: randomly pick 2 of the 5 branches and survey every visitor there this week → cluster (and note what it risks: those 2 branches may not represent all 5).


Segment 6 — Bias, and a Famous Failure (18 min)

Plain language: bias is error baked into the method — it pushes results in the same wrong direction no matter how big the sample gets. Four kinds to recognize:
- Undercoverage — part of the population can't be reached or is left out of the frame.
- Nonresponse bias — the people who don't answer differ from those who do.
- Response bias — the question or setting pushes answers (leading wording, sensitive topics, who's asking).
- Voluntary-response bias — the opt-in crowd isn't typical.

The famous failure (the story that sticks):

The 1936 Literary Digest presidential poll. The magazine mailed 10 million ballots and got 2.4 million back — a colossal sample — and confidently predicted Landon would beat Roosevelt. Roosevelt won in a landslide.
What went wrong? The mailing list came from car registrations and telephone directories — in the Depression, that skewed wealthy (undercoverage) — and only motivated people mailed ballots back (nonresponse). Meanwhile George Gallup polled a few thousand people chosen well and called it correctly.
The lesson, in one line: 2.4 million badly-chosen people lost to a few thousand well-chosen ones. Method beats size.

Callback: this is the cure to Segment 4's "bigger is better" misconception. Point back to it explicitly.


Segment 7 — Observational Study vs. Experiment; Correlation ≠ Causation (20 min)

Plain language first.
- In an observational study, you watch and record — you don't change anything. (Survey people about their habits; track what happens.)
- In an experiment, you deliberately impose a treatment and compare. (Randomly assign some customers a new app layout, others the old one, compare usage.)
- Only an experiment with random assignment can support a cause-and-effect claim. Observational studies can show a link, never prove the arrow.

The reason — confounding (worked example):

Headline: "People who use a fitness tracker take more daily steps." (Observational — nobody was assigned a tracker.)
A confounding variable is a third thing tangled with both: people who already exercise are exactly the people who buy trackers. The data can't separate "tracker → steps" from "already-active → tracker AND steps."
Picture it: tracker — steps looks like a straight arrow, but a hidden "already active" pulls both strings.

Memory hook: "Correlation is a handshake, not a push." Two things moving together ≠ one shoving the other.

Misconception + cure:
- ❌ "They found a strong correlation, so X causes Y."
Cure: ask "could a third variable explain both?" and "was anything randomly assigned?" If nothing was assigned, you have a link, not a cause.

Quick mini-debate (genuinely arguable, ~4 min): "A meal-kit company finds its subscribers cook at home more often than non-subscribers. Should it advertise 'our kits make you cook more'?" Have students argue both sides; surface the confounder (people who subscribe already like cooking) — and ask what experiment would settle it (randomly assign free kits).


Segment 8 — Technology Workflow + AI-Critique, Callback & Hand-off (12 min) · Session 2 closes (~75)

Technology workflow — draw an SRS in a spreadsheet (exact steps):
1. Put your list of names/IDs in column A (say A2:A101 for 100 people).
2. In B2 type =RAND() and fill down to B101 — a random number beside each person.
3. Select columns A:B → Data ▸ Sort range → sort by column B.
4. The top 10 rows are your simple random sample. (Re-sorting reshuffles — that's the randomness working.)
- Google Sheets: =RAND(); Excel identical. For a quick single pick, =RANDBETWEEN(1,100) also works.

AI-critique moment (students verify, not consume):

Paste this to your chatbot: "Classify these variables by level of measurement: a bus route number, a hotel's 1–5 star rating, temperature in °C, monthly rent."
Then check its work against NOIR. Chatbots often call a route number "ratio" (it's nominal) and stumble on °C ("ratio"? — it's interval, no true zero). Your job all term: the tool drafts, you judge. This is exactly how the weekly Lecture Tutorial and Data Lab work — you'll catch the model, not trust it.

Callback + tease:
- Callback: "Every trustworthy number this term rides on this week — who was measured and how they were picked."
- Tease next week: "Now that we can get good data, Week 2 is the first thing we do with it: turn a pile of numbers into a picture — histograms, shapes, and the graphs that lie."

Hand-off (the week's work):
- Chapter 1 (the primary reading) if they haven't read it yet — then Lecture Tutorial 1 (AI tutor, share-link submission) — population/sample, NOIR, sampling & bias.
- Data Lab 1 (real dataset: penguins!) · Quiz 1 (end of week) · Discussion 1 (the star-rating question) · Assignment 1 (AI-coached).


Instructor FAQ — Common Stumbles

Student says / does Quick cure
"Is a driver ID quantitative? It's a number." Apply the test — does averaging it mean anything? No. It's a nominal label that happens to look numeric.
"What's the difference between interval and ratio again?" One question: does zero mean "none"? Yes → ratio (rent, distance, counts). No, zero is just a mark → interval (°F, °C, calendar year).
Confuses stratified and cluster. Stratified = sample within every group; cluster = randomly pick whole groups and measure all of them. Stratified buys representation; cluster buys cheapness.
"The poll had 2 million responses, so it's reliable." Size never fixes bias. The 1936 Literary Digest poll: 2.4 million responses, wrong winner. Method beats size.
Calls a strong correlation "proof" of cause. Ask: was anything randomly assigned? If no, it's observational → a link, not a cause. Hunt the confounder.
"Population = a big group." The population is whoever the question is about — could be 25 people. It's a role, not a size.
Thinks a census is a kind of sample. A census measures the whole population — the opposite of sampling. Rare because it's usually too costly or slow.
Picks a convenience sample because it's "random enough." Convenience ≠ random. "Whoever's nearby" has a hidden pattern (people at the library visit the library a lot). Only a chance-based method earns the word random.

Scope flag

This outline stays within Objective 1. The °C/calendar-year interval-vs-ratio nuance and the 1936 Literary Digest case are added context (not strictly required by the objective) — kept because they make the misconceptions stick; cut them for a leaner session.