Week 10 — Lecture Outline · Sampling Distributions & the Central Limit Theorem
Course: Introduction to Statistics (18-week generic edition)
Objectives covered: Objective 5 — Use normal and sampling distributions to describe how sample statistics behave (this week completes the objective the Week 8 normal model began).
SLOs touched: A (reason quantitatively from data) · B (communicate results to a non-technical audience)
Meeting pattern: planned as 2 sessions × ~75 min = ~150 min. Segment minutes below total ~150; scale them to your own pattern.
Week at a Glance
| The week's big question | "Every honest sample gives a slightly different answer. So how can any one sample's answer be trusted — and by exactly how much does an average wobble?" |
| By the end of the week, students can… | (1) explain sampling variability and describe a sampling distribution (the pile of a statistic's values across all possible samples); (2) give the center and spread of the sampling distribution of x̄ — mean μ and standard error σ/√n — and keep SD vs. SE straight; (3) state and apply the Central Limit Theorem, including when it kicks in; (4) compute probabilities about a sample mean (and, via totals ⟺ means, about a sum) using the Week 8 z-table; (5) do the same for a sample proportion p̂, checking np ≥ 10 and n(1 − p) ≥ 10. |
| Key vocabulary | sampling variability, sampling distribution, unbiased, standard error (SE), SD vs. SE, Central Limit Theorem (CLT), n ≥ 30 rule of thumb, sampling distribution of p̂, success/failure condition (np ≥ 10, n(1 − p) ≥ 10) |
| Materials | slides (Deck 10), the Week 10 chapter (the friendly z-table rides again in its Section 4), the week's readings + video links, a spreadsheet (Google Sheets or Excel), a Desmos-class graphing/statistics tool, the student's chatbot for the AI-critique moment and the tutorial |
| Timing note | 8 segments, ~150 min total. Session 1 = Segments 1–4 (~75). Session 2 = Segments 5–8 (~75). |
Segment 1 — Hook & the Promise (8 min) · Session 1 opens
Welcome to the second half. The midterm is behind you — Weeks 1–8 taught you to get data honestly, describe it, quantify chance, and model it with the normal curve. Everything from here on answers one question: what can a sample really tell us about the world it came from? This week is the engine room.
Hook. Hold up (or describe) a snack-size candy bag: "The label says 200 g. Has anyone ever weighed one? You'd get 197.2. Weigh another: 203.1. A third: 199.4. Nobody's bag is exactly 200. So what is the label actually promising — and how could the company possibly keep a promise about an average when every single bag misses it?"
- Same puzzle everywhere: a shipping company promises an average delivery time, a call center reports an average hold time, an elevator placard trusts an average passenger weight. Individuals bounce all over; somehow the averages are steady enough to promise.
- "This week you learn the single most useful fact in statistics: averages behave. Individual values wobble a lot; averages of n values wobble exactly √n times less — and their wobble follows the bell curve you mastered in Week 8, almost no matter what the individuals do."
The promise (write it on the board): "By the end of this week you can predict the behavior of a sample average you haven't taken yet — its center, its spread, and its shape — and compute the probability it lands anywhere."
Why it matters line (memory hook): "One value is a guess. An average is a promise with a wobble you can compute."
Segment 2 — Sampling Variability & the Sampling Distribution Idea (20 min)
Plain language first. Draw a random sample and compute its mean: you get one number. Draw a different random sample the same honest way: you get a different number. That bounce is sampling variability — Week 1's lab showed it to you (two students' penguin samples never matched), and it is not an error. It's what sampling does.
Now the week's big mental move. Imagine taking every possible sample of size n from the population and computing x̄ for each one. Pile up all those x̄'s into one distribution. That pile is the sampling distribution of the sample mean — the complete catalog of every answer your sampling method could give, and how often.
- Three distributions to keep straight (this is the week's map — slide it, say it, repeat it):
1. The population distribution — every individual value. Any shape at all.
2. The distribution of one sample — your n measured values. It resembles the population (bumpy, same shape-ish).
3. The sampling distribution of x̄ — one entry per sample, not per individual. New shape, new (smaller) spread. - A statistic is a random variable (Week 6 callback): before you draw the sample, x̄ has a distribution, an expected value, and an SD — and we can find all three.
One fully worked example (do every step out loud — the four-bag toy build).
A tiny population: 4 candy bags weighing 196, 198, 202, 204 g. Population mean μ = (196 + 198 + 202 + 204) ÷ 4 = 800 ÷ 4 = 200 g.
List every possible sample of n = 2 (order doesn't matter — there are 6) and each sample's mean:
- {196, 198} → x̄ = 197 · {196, 202} → x̄ = 199 · {196, 204} → x̄ = 200
- {198, 202} → x̄ = 200 · {198, 204} → x̄ = 201 · {202, 204} → x̄ = 203
The sampling distribution of x̄ is the pile {197, 199, 200, 200, 201, 203}. Two things to notice, both huge:
1. Its average is (197 + 199 + 200 + 200 + 201 + 203) ÷ 6 = 1200 ÷ 6 = 200 — exactly μ. Sample means don't systematically miss high or low: x̄ is an unbiased estimator of μ.
2. It huddles. Individuals span 196 to 204 (a range of 8 g); the means span 197 to 203 (a range of 6 g). Averaging pulled the extremes in — a lucky-high bag tends to get paired with an ordinary one.
Land the key idea: any one sample mean can miss μ. But the collection of possible sample means is centered exactly on μ and is tighter than the individuals. Next segment: how much tighter — the exact law.
Segment 3 — Center & Spread of x̄: the Standard Error (25 min)
Plain language first. For a random sample of size n from any population with mean μ and standard deviation σ, the sampling distribution of x̄ has:
- Center: mean = μ (unbiased — the toy build just showed it).
- Spread: standard deviation = σ ⁄ √n — so important it gets its own name: the standard error (SE) of the sample mean.
The two rulers (say this until it's a reflex): σ is the ruler for individuals — how far ONE value typically strays from μ. σ/√n is the ruler for averages — how far a sample MEAN of n values typically strays from μ. Same units, totally different jobs.
One fully worked example (do every step out loud).
A candy company's filling machine drops pieces into snack bags with individual-bag weights averaging μ = 200 g with σ = 8 g (fills are approximately bell-shaped — machines are like that). Bags ship in cases of n = 25.
- SE of the case average: σ/√n = 8 ⁄ √25 = 8 ⁄ 5 = 1.6 g.
- In words: one bag typically sits about 8 g from 200; a case's average typically sits only about 1.6 g from 200.
- That's how the label keeps its promise: the company can't control your bag, but averages across many bags barely move.
The √n economics (the part that runs every study-design budget):
- n = 25 → SE = 8/5 = 1.6. To get SE = 0.8 — half the wobble — you need √n = 10, i.e., n = 100. Quadruple the sample to halve the noise. Precision gets expensive fast; that trade-off is why real studies agonize over sample size.
Misconception + cure (the week's #1):
- ❌ "SD and SE are basically the same thing."
✅ Cure: ask "one value, or an average?" before touching a formula. σ answers questions about ONE bag; σ/√n answers questions about a CASE average. Using σ where SE belongs is the wrong-ruler error — the single most common mistake of the entire second half.
Honesty note (scope): the four-bag build sampled without replacement, so the ÷√n law doesn't apply to it exactly — we read only "centered at μ" and "narrower" from the toy. The σ/√n formula is exact for independent draws and excellent whenever the population dwarfs the sample (our standing situation).
Segment 4 — Misconceptions + Think-Pair-Share (22 min) · Session 1 closes (~75)
Name the misconceptions out loud, then cure each:
- ❌ "n is the number of samples."
✅ Cure: n is the size of EACH sample. How many samples you take is a different number entirely. In this week's lab you take 30 samples of n = 10 penguins: n = 10, number of samples = 30. Blur those and every formula breaks. - ❌ "A big sample makes the data normal."
✅ Cure: your one sample's histogram looks like the population — skew, lumps, and all — no matter how big n gets. What turns bell-shaped is the pile of sample means. (Full story next segment.) - ❌ "Doubling the sample halves the standard error."
✅ Cure: the √n is merciless — doubling n divides SE by √2 ≈ 1.41 only. Quadrupling halves it. (25 → 100, not 25 → 50.) - ❌ "The sample mean equals the population mean."
✅ Cure: unbiased means the sampling distribution is centered on μ — no systematic lean — not that any one x̄ hits it. Your x̄ almost certainly misses μ; the SE tells you by about how much.
Interaction — Think-Pair-Share (rapid-fire, ~12 min):
Put 6 quick items on a slide; students answer solo (30 sec), compare with a neighbor (1 min), then vote. Suggested items:
1. σ measures the typical wobble of __. (individual values)
2. σ/√n measures the typical wobble of _. (sample means)
3. A process has σ = 30; samples of n = 36. SE = . (30/√36 = 30/6 = 5)
4. To cut an SE in half, multiply n by . (4)
5. True or false: with n = 400, the histogram of the 400 individual values is approximately normal no matter the population's shape. (False — it mirrors the population)
6. The sampling distribution of x̄ is centered at ___. (μ)
Debrief items 3 and 5 — the compute-and-classify pair that predicts quiz performance.
Segment 5 — The Central Limit Theorem (25 min) · Session 2 opens
Hook back in: "We know the center (μ) and the spread (σ/√n) of the sample mean's distribution. One question left — the shape. Week 8 you learned what a normal shape buys you: exact percentages. Here's the astonishing part: for averages, the bell shows up whether or not the population has one."
Plain language first — the theorem.
The Central Limit Theorem (CLT): take random samples of size n from any population with mean μ and SD σ. As n grows, the sampling distribution of x̄ becomes approximately normal — no matter what shape the population has — with mean μ and SE σ/√n.
When the bell kicks in (the fine print, taught as three lines):
- Population already normal → x̄ is normal at every n (even n = 2).
- Population any shape at all — skewed, lumpy, bimodal → n ≥ 30 is the working rule of thumb (heavier skew wants more).
- Always required: a genuinely random sample. The CLT fixes shape, never bias — a biased sample's mean is confidently wrong at any n (Week 1's lesson, still undefeated).
Callback that makes it land (Week 8's lab): the penguins' pooled body-mass histogram was not one bell — right-skewed with a second bump (three species mixed). The CLT's claim: draw samples of penguins and histogram the sample MEANS, and the bell appears anyway. This week's data lab has students build exactly that with their own hands.
One fully worked example (do every step out loud — the case-average contrast).
The candy machine again: individual bags ≈ normal, μ = 200 g, σ = 8 g; cases of n = 25; SE = 1.6 g. Question: how often is a weight over 204 g? — asked two ways.
- One random bag over 204 g: the ruler is σ = 8. z = (204 − 200) ⁄ 8 = 0.5 → area above = 1 − 0.6915 = 0.3085 ≈ 31%. Common.
- A case's average over 204 g: the ruler is SE = 1.6. z = (204 − 200) ⁄ 1.6 = 2.5 → area above = 1 − 0.9938 = 0.0062 = 0.62%. Rare.
Same threshold, two rulers: 31% vs. 0.62% — a 50-fold drop. Averages don't just wobble less; extreme averages become genuinely rare. That rarity is what all of inference will lean on.
Misconception + cure:
- ❌ "The CLT says the population becomes normal when n is large."
✅ Cure: the population never changes — it doesn't know you're sampling it. Only the distribution of x̄ goes normal. Say the theorem's subject out loud every time: the sampling distribution of the sample mean.
Segment 6 — The Elevator Problem: the CLT at Work (18 min)
Set up the drama. "Every elevator has a placard. Ours says: capacity 9 people or 1,800 lb. Nine adults step in. Should you be nervous?"
Worked example (do every step out loud).
Suppose adults in this building have weights that are approximately normal with μ = 190 lb and σ = 24 lb, and the 9 riders are (for now) like a random sample.
- Totals ⟺ means bridge: the total exceeds 1,800 lb exactly when the average rider exceeds 1,800 ⁄ 9 = 200 lb. One divide converts a sum question into a mean question — our machinery applies.
- SE of the 9-rider average: 24 ⁄ √9 = 24 ⁄ 3 = 8 lb.
- z = (200 − 190) ⁄ 8 = 1.25 → area above = 1 − 0.8944 = 0.1056 ≈ 10.6% — about one full ride in nine.
That is not a comfortable number — which is exactly why real placards carry big engineering safety margins beyond the printed limit, and why engineers do this calculation with averages, not one rider at a time.
The assumption check (honesty moment — say it plainly): real riders are not a random sample. People board in groups — a visiting basketball team, a family with kids — so weights arrive correlated, and the true risk can sit above or below 10.6%. The formula is only as good as the random-sampling assumption feeding it. (This "check the conditions before trusting the output" reflex is the second half's professional habit.)
Quick mini-debate (~4 min, genuinely arguable): "Should the placard print the statistical reality — 'with 9 average-population adults, about a 1-in-9 chance over the rated load' — or just '9 persons max'?" Surface the tension: transparency vs. usable rules; and where the safety margin actually lives (the cable is rated far above the placard).
Segment 7 — The Sample Proportion p̂ (20 min)
Plain language first. Weeks of this course run on percentages — and a percentage is secretly an average. Code every individual as 1 (success) or 0 (not), and the proportion p̂ is just the mean of the 0s and 1s. So everything today already applies:
- Center: the sampling distribution of p̂ has mean p (unbiased again).
- Spread: SE = √( p(1 − p) ⁄ n ).
- Shape: approximately normal when the sample expects at least 10 of each kind: np ≥ 10 and n(1 − p) ≥ 10 (the success/failure condition — the binomial-to-bell bridge you saw building in Week 7).
One fully worked example (do every step out loud).
A parcel service's long-run on-time rate is p = 0.80. Each day, quality control pulls a random sample of n = 100 tracking scans. How should the daily sample proportion p̂ behave?
- Center: 0.80. SE = √(0.80 × 0.20 ⁄ 100) = √0.0016 = 0.04.
- Shape check: np = 80 ✓ and n(1 − p) = 20 ✓ — both ≥ 10, so the bell applies.
- "How often would a day show 75% or worse purely by sampling luck?" z = (0.75 − 0.80) ⁄ 0.04 = −1.25 → area below = 0.1056 ≈ 10.6%. About one day in nine looks that bad with nothing wrong — a manager who panics at every 75% day is chasing sampling noise.
Misconception + cure:
- ❌ "SE of p̂ = p(1 − p)/n" (the missing square root).
✅ Cure: p(1 − p)/n is the variance; the SE wears the √. Sanity-check the units of wobble: for p = 0.80, n = 100, an SE of 0.0016 would mean essentially no wobble — too good to be true is the tell. (σ vs. σ² strikes again.)
Segment 8 — Technology Workflow + AI-Critique, Callback & Hand-off (12 min) · Session 2 closes (~75)
Technology workflow — hand-crank a sampling distribution, then let the software confirm it (exact steps, the lab's engine):
1. Data in a sheet (the lab uses the penguin body masses). In the first empty column type =RAND() beside every row and fill down.
2. Select the data range → Data ▸ Sort range ▸ Advanced (sort by the RAND column). The top 10 rows are a fresh random sample.
3. In a spare cell, =AVERAGE( the 10 sampled values ) — then copy → Paste special ▸ Values only into a results column. That's one sample mean, frozen.
4. Re-sort (the =RAND() values reshuffle automatically) and repeat until you've banked 30 sample means. =AVERAGE() and =STDEV() of that results column, plus Insert ▸ Chart ▸ Histogram, reveal the sampling distribution you built: centered near the full-data mean, SD near σ/√10, mound-shaped.
5. Theory-side shortcut: =NORM.DIST(204, 200, 1.6, TRUE) returns P(case average < 204) = 0.9938 for the candy example — the third argument must be the SE (1.6), never σ (8). That argument box is where the wrong-ruler error hides in software. (Desmos-class tools: normaldist(200, 1.6) and shade — same rule, the SD slot takes the SE.)
AI-critique moment (students verify, not consume):
Paste this to your chatbot: "A population is strongly right-skewed. I take ONE random sample of 100 values. Will my histogram of those 100 values be approximately normal thanks to the Central Limit Theorem?"
Chatbots regularly answer yes — or hedge into a muddle. The correct answer is no: your 100 values' histogram resembles the population (skewed); the CLT speaks only about the distribution of sample means across samples. If the bot got it right, push once — "so a big sample never looks normal?" — and watch for the over-correction. The tool drafts, you judge — and this week's lab makes the chatbot explain YOUR two SDs (802 vs. ~254) and catches it conflating them.
Callback + tease:
- Callback: "Week 1 told you method beats size; Week 8 gave you the bell. This week glued them: an honest sample's average follows a bell centered on the truth, with a wobble of σ/√n we can compute."
- Tease next week: "One more flip and inference begins. This week: 'given the truth, how do sample means behave?' Next week: 'given ONE sample mean, where is the truth probably hiding?' — the confidence interval, built directly on today's SE."
Hand-off (the week's work):
- Chapter 10 (the friendly z-table rides again) — then Lecture Tutorial 10 (AI tutor; share link + summary).
- Data Lab 10 (the signature build: 30 penguin samples, the bell from nowhere) · Quiz 10 (end of week, closed to AI) · Discussion 10 (the fine print on "average") · Assignment 10 (AI-coached, delivery-time problems).
Instructor FAQ — Common Stumbles
| Student says / does | Quick cure |
|---|---|
| "Standard error, standard deviation — same thing, right?" | Two rulers. σ = typical stray of ONE value; σ/√n = typical stray of an AVERAGE of n. Ask "one value, or an average?" before every formula. |
Uses σ where the SE belongs (or feeds σ to =NORM.DIST for a mean question). |
The wrong-ruler error. Have them write the ruler before the z: "ruler = 8/√25 = 1.6" on its own line. In software, the SD argument takes the SE for any sample-mean question. |
| "So n is 30 because we took 30 samples?" | n = the size of EACH sample (10 penguins). The 30 is how many samples — a different number with no formula role. Lab language: "30 samples of n = 10." |
| "With 500 data points my histogram should look normal." | One sample's histogram mirrors the population. Only the pile of sample means earns the bell. If their histogram is skewed, the population probably is too — that's information, not failure. |
| "We doubled n, so the SE halved." | √n: doubling divides by 1.41. Halving the SE costs 4× the sample. (8/√25 = 1.6 → need n = 100 for 0.8.) |
| "The sampling distribution is centered at μ, so my x̄ = μ." | Unbiased = no systematic lean, not a bullseye. Any single x̄ misses; the SE says by about how much (and Week 11 turns that into a margin of error). |
| "Why 30? What's magic about it?" | Nothing — it's a rule of thumb that works for moderate skew. Normal-ish population: any n. Heavy skew/outliers: want more than 30. The magic word is approximately. |
| Forgets the √ in the p̂ standard error. | p(1 − p)/n is the variance; the SE wears the √. For p = 0.8, n = 100: √0.0016 = 0.04, not 0.0016 (σ vs. σ² in new clothes). |
| "The elevator math says 10.6%, so that's the real risk." | Only if riders were a random sample — they aren't (people board in groups). Conditions in, trust out: the number is as good as the assumptions. |
Scope flag
This outline stays within Objective 5. Deliberate boundaries: the four-bag toy build samples without replacement, so it demonstrates "centered at μ" and "narrower" only — the finite population correction is out of scope for this course and never invoked; the law of large numbers appears informally (averages settle as n grows) without formal statement; n ≥ 30 is taught as a rule of thumb, not a theorem; the elevator example's not-an-SRS caveat is taught as an assumption check, not pursued into dependence math. The candy-case contrast (31% vs. 0.62%) is added drama the objective doesn't strictly require — kept because it makes the two-rulers idea unforgettable; cut it for a leaner session.