Back to the Introduction to Statistics outline The Course Maker
Introduction to Statistics outline
Week 12 · Lecture outline

Week 12 — Lecture Outline · Confidence Intervals for a Proportion

Introduction to Statistics Generic evergreen edition

Course: Introduction to Statistics (18-week generic edition)
Objectives covered: Objective 6 — Construct and interpret confidence intervals (this week: for a population proportion), and determine the sample size a target margin of error requires.
SLOs touched: A (reason quantitatively from data) · B (communicate results to a non-technical audience)
Meeting pattern: planned as 2 sessions × ~75 min = ~150 min. Segment minutes below total ~150; scale them to your own pattern.


Week at a Glance

The week's big question "Every poll in your feed ends with the same fine print — 'margin of error ±3 points.' Where does that number come from, what does it promise, and what does it quietly leave out?"
By the end of the week, students can… (1) compute a sample proportion p̂ and its standard error; (2) check the conditions for a one-proportion z-interval (random sample; at least 10 successes and 10 failures; population at least 10× the sample); (3) build the interval p̂ ± z·SE at 90/95/99% and state the confidence–width trade-off; (4) interpret it correctly — and spot the two classic misreadings; (5) choose the sample size a target margin of error requires (and always round up); (6) audit a media poll's margin of error* from its own reported numbers.
Key vocabulary point estimate, sample proportion p̂, standard error of p̂, critical value z, margin of error (ME), confidence interval, confidence level, one-proportion z-interval, large-counts (success/failure) condition, 10% condition, conservative planning value (p = 0.5), sample-size formula
Materials slides (Deck 12), the Week 12 chapter (with the z* multiplier table), the week's readings + video links, a spreadsheet (Google Sheets or Excel), a Desmos-class statistics tool, the student's chatbot for the AI-critique moment and the tutorial
Timing note 8 segments, ~150 min total. Session 1 = Segments 1–4 (~75). Session 2 = Segments 5–8 (~75).

Segment 1 — Hook & the Promise (8 min) · Session 1 opens

Hook. "Sometime this month, a headline in your feed ended with the same eight words of fine print: '…with a margin of error of ±3 percentage points.' Show of hands — who has ever actually used that ±3 for anything?" (Few hands.) "By the end of this week, that fine print is yours. You'll know where the 3 comes from, you could recompute it from the article itself, and you'll know the two things it does not protect you from."

  • Callback to Week 11: "Last week we built our first confidence intervals — for a mean, with the t-multiplier. Same machine this week, new fuel: the statistic being estimated is now a percent — the most public statistic in the world."
  • Everything runs on Week 10's engine: p̂ varies from sample to sample in a predictable, bell-shaped way. Predictable variation is what lets us attach a margin to one sample's answer.

The promise (write it on the board): "By the end of this week you can take any yes/no question about a big group — will they renew? do they recycle? did it get returned? — and turn one random sample into an honest range for the truth, plus a price tag: exactly how many people a ±3 (or ±2, or ±1) costs."

Why it matters line (memory hook): "A percent without a margin of error is a number with its fine print torn off."


Segment 2 — From p̂ to an Interval: The Anatomy (20 min)

Plain language first — the Week 10 recap that powers everything. Take a random sample and compute the proportion of "yes" — that's , our point estimate of the population proportion p. Different random samples give different p̂'s, but Week 10 told us how they differ: p̂'s pile up in a bell shape centered at p, with a typical distance from p of

SE = √( p̂(1 − p̂) ⁄ n ) — the standard error of p̂ ("how far a sample's percent typically sits from the truth").

  • If p̂ usually lands within a couple of SEs of p, then reaching a couple of SEs out from our p̂ usually reaches back and captures p. That reach is the whole idea of a confidence interval.
  • The anatomy (say it as a chant): estimate ± multiplier × standard errorp̂ ± z*·SE. The margin of error (ME) is the reach: ME = z·SE. Memory hook: "center ± reach."*

The multiplier table (embedded in the chapter, tutorial, and assignment coach — the only three values this week uses):

The table below lists the confidence multiplier z* for each confidence level used in this course.

Confidence level z*
90% 1.645
95% 1.96
99% 2.576

Why z this week, when last week used t? For a mean we had to estimate the spread with a second, separate statistic (s), and the t-multiplier pays for that extra uncertainty. For a proportion there is no separate spread number — p̂ itself feeds both the center and the SE — so the normal (z) multiplier is the standard tool. One-line version for students: means → t, proportions → z. (And they've met these z's before: Week 11's t-table columns head toward exactly these values as df grows.)

Conditions before construction (the password, three checks):
1. Random. The data are a random sample from the population (bias, Week 1, is not fixable by any formula).
2. Large counts. At least 10 successes and 10 failures in the sample — that's what makes p̂'s distribution normal enough for z.
3.
10% condition. When sampling without replacement, the population should be at least 10 × n* (keeps the draws effectively independent).


Segment 3 — One Full Build, Every Step Out Loud (22 min)

One fully worked example (the week's anchor — do every step).

A campus tech office wants the proportion of the campus's 18,000 students who have installed the campus app. It draws a random sample of 100 students from the enrollment database; 60 have the app.
- Point estimate: p̂ = 60 ÷ 100 = 0.60.
- Conditions: random ✓ (drawn from the full enrollment list) · successes 60 ≥ 10 and failures 40 ≥ 10 ✓ · 18,000 ≥ 10 × 100 ✓.
- Standard error: SE = √(0.60 × 0.40 ⁄ 100) = √0.0024 ≈ 0.049.
- Margin of error (95%): ME = 1.96 × 0.049 ≈ 0.096.
- Interval: 0.60 ± 0.096 → (0.504, 0.696) — "somewhere between about 50.4% and 69.6%."
- Interpretation (the sentence to memorize the shape of): "We are 95% confident that the interval from 0.504 to 0.696 captures the true proportion of all 18,000 students who have installed the app."

Now turn the confidence dial (same p̂, same SE — only z* changes):
- 90%: 1.645 × 0.049 ≈ 0.081 → (0.519, 0.681) — narrower, but the method misses more often.
- 99%: 2.576 × 0.049 ≈ 0.126 → (0.474, 0.726) — surer, but wider.

Land the trade-off: surer means wider; narrower means less sure. At a fixed sample size you cannot have both — the only way to be both surer and narrower is more data (Segment 5 prices that out).

Say the units out loud: the interval is about the proportion — never "95% of students." 0.504 to 0.696 is a range of possible truths about p, not a range of people.


Segment 4 — Misconceptions + Quick Interaction (25 min) · Session 1 closes (~75)

Name the misconceptions out loud, then cure each:

  • "There's a 95% chance p is between 0.504 and 0.696."
    Cure: p is a fixed (unknown) number — it doesn't hop around, so it gets no probability. What's random is the interval: a different sample would produce a different one. "95% confident" describes the method: intervals built this way capture the truth in about 19 of every 20 random samples. Once yours is built, it either captured p or missed — you just don't know which. The confidence is in the recipe, not in one batch.
  • "95% of the students are inside the interval."
    Cure: the interval brackets a proportion, not people. Individuals are 0s and 1s (have the app / don't); no student is "0.53."
  • "p̂ IS the true proportion — the survey found 60%."
    Cure: Week 1's hat is still on duty: p̂ is measured, p is true. The entire reason we build an interval is that p̂ ≠ p, sample after sample.
  • "The margin of error accounts for everything that could be wrong with the poll."
    Cure: ME covers random sampling variation only — the luck of who landed in the sample. Leading wording, nonresponse, undercoverage (Week 1's whole bias catalog) are not in the ±, and they can dwarf it. (Segment 6 makes this the headline.)

Interaction — Think-Pair-Share (rapid-fire, ~10 min). Six quick items on a slide; solo 30 sec, pair 1 min, fingers vote:
1. A random sample of 100 has 50 successes — what is p̂? → 0.50
2. Same sample, same data: which is wider, the 90% or the 99% interval? → 99%
3. SE for p̂ = 0.50, n = 100? → √(0.25 ⁄ 100) = 0.05
4. True or false: "(0.504, 0.696) at 95% means p has a 95% chance of being in there." → False (the 95% belongs to the method)
5. A sample of 40 has 8 successes — may we build the z-interval? → No (8 < 10; large-counts condition fails)
6. Quadruple the sample size: the SE becomes… → half as big (√4 = 2)
Debrief items 4 and 6 — they're the week's two load-bearing ideas.


Segment 5 — Choosing the Sample Size: Buying Precision (20 min) · Session 2 opens

Hook back in: "Last session we read the margin of error off a sample. Real studies run the other direction: someone says 'I need ±3' — and you have to answer 'that will cost you this many people.'"

Plain language first. Before the data exist, pick the target margin ME and confidence level, then solve the margin's formula backwards for n:

n = p(1 − p) · (z ⁄ ME)² — where p is a planning value for the proportion.

  • No prior idea of p? Use the conservative value p* = 0.5 — it makes p(1 − p) as big as it can be (0.25), so the n it demands is always enough, whatever p turns out to be.
  • Have a pilot estimate? Use it — a p* far from 0.5 shrinks the bill.

One fully worked example (do every step).

A city sustainability office wants to estimate the proportion of households that set out curbside recycling, with a 95% interval and a margin of error of at most ±0.03. No prior estimate.
- n = 0.25 × (1.96 ⁄ 0.03)² = 0.25 × 4268.44 = 1067.11 → survey 1,068 households.
- Round UP, always: 1,067 would leave the margin a hair over target. The ceiling is a guarantee, not an approximation.
- With a pilot estimate p = 0.8 (a nearby city saw ~80%): n = 0.8 × 0.2 × 4268.44 = 0.16 × 4268.44 = 682.95 → 683 households* — prior knowledge cut the bill by about a third.

The cost law (say it twice): ME shrinks like 1 ⁄ √n. Half the margin costs four times the sample; a tenth of the margin costs a hundred times. Precision is bought on a square-law budget — which is why ±3, not ±1, is the industry standard.

Misconception + cure:
- ❌ Rounding n to the nearest whole number (1067.11 → 1067).
Cure: n must meet or beat the target margin. Round the formula's answer up, every time.


Segment 6 — Polls & the Margin of Error in the Media (20 min)

Plain language first — the anatomy of a poll report. A well-reported poll tells you: who was sampled (the population it claims to speak for), n, how they were reached, the margin of error, and the confidence level (almost always 95%, often unstated). This segment: recompute a poll's ± from its own numbers.

One fully worked example (the audit — every step).

A national outlet reports: "52% of 1,067 randomly sampled adults support expanding passenger rail service; margin of error ±3 percentage points."
- SE = √(0.52 × 0.48 ⁄ 1067) ≈ 0.0153.
- ME = 1.96 × 0.0153 ≈ 0.030±3.0 points. The fine print checks out.
- The napkin rule: at 95% with the conservative p = 0.5, ME = 0.98 ⁄ √n ≈ 1 ⁄ √n. For n = 1,067: 1 ⁄ √1067 ≈ 0.031 → "about ±3." (That's why so many professional polls have n near 1,000 — it's the ±3 price point.)
-
Read the interval, not the point: 52 ± 3 → 49% to 55%. The interval dips below 50%, so "majority support" is plausible but not established*. A headline that says "majority backs rail" is over-reading its own data.

What the ±3 does NOT cover (the segment's headline): the margin of error prices sampling luck only. Question wording, nonresponse, undercoverage — Week 1's entire bias catalog — ride outside the ±, invisible and often bigger. Callback to the forever-example: the 1936 Literary Digest poll's 2.4 million responses came with a tiny computed margin — and still called the wrong winner, because the method was biased. A ± on a biased sample is a precise measurement of the wrong thing.

Media-reading habit (give them the 3-question checklist): ① Who was actually sampled, and how? ② What's the interval (point estimate ± ME) — and does the story's claim survive the whole interval? ③ What could bias this that the ± doesn't cover?


Segment 7 — Interpretation Clinic + Mini-Debate (15 min)

The three-sentence game (run it as a vote). A valid 95% interval for a proportion is (0.42, 0.52). Which interpretation survives?
- A. "95% of the individuals fall between 0.42 and 0.52." → ✗ (people aren't proportions)
- B. "There's a 95% probability that p is between 0.42 and 0.52." → ✗ (the 95% belongs to the method, not to this one interval)
- C. "We are 95% confident the interval from 0.42 to 0.52 captures the true population proportion." → ✓
Then have pairs write their own sentence C for the campus-app interval and read two aloud. Police the three ingredients: confidence level · the interval · the population's proportion.

Mini-debate (genuinely arguable, ~6 min): "A sample of 1,000 people can't possibly speak for tens of millions — agree or disagree?" Let both sides argue, then reveal the resolution hiding in the formula: n is in the SE formula; the population size N is not. A well-stirred pot of soup needs only one spoonful to taste, whether it's a saucepan or a vat — if the stirring (random sampling) was real. Size of the pot: nearly irrelevant (the 10% condition is the only cameo). Stirring: everything.


Segment 8 — Technology Workflow + AI-Critique, Callback & Hand-off (20 min) · Session 2 closes (~75)

Technology workflow — the whole interval in six cells (exact steps):
1. A1: successes 60 · A2: sample size 100.
2. A3 — p̂: =A1/A20.6.
3. A4 — SE: =SQRT(A3*(1-A3)/A2)0.04899.
4. A5 — ME at 95%: =1.96*A40.09602.
5. A6 and A7 — the interval: =A3-A50.50398 · =A3+A50.69602. (Rounded: 0.504 to 0.696 — matching the hand build.)
6. Change A1/A2 and the whole interval re-computes — that's the audit tool for any poll. (Google Sheets and Excel are identical here; a Desmos-class statistics tool or graphing calculator has this packaged as 1-PropZInt.)

AI-critique moment (students verify, not consume):

Paste this to your chatbot: "What sample size do I need to estimate a proportion to within ±3 percentage points at 95% confidence, and what's the formula?"
Then check its work against the formula: n = 0.25 × (1.96 ⁄ 0.03)² = 1067.11 → 1,068. Chatbots frequently answer "1,067" — quoting the famous unrounded figure from memory — or round 1067.11 down. The round-UP rule is exactly the kind of fine print they fumble. Your job all term: the tool drafts, you judge. This same audit is the heart of this week's Data Lab.

Callback + tease:
- Callback: "Week 10 gave us the engine (sampling distributions); Week 11 built the machine for means (t); this week the same anatomy — estimate ± multiplier × SE — handled percents with z. One machine, two fuels."
- Tease next week: "A confidence interval lists every plausible value of the truth. Next week someone walks in with a specific claim — 'our support is 60%,' 'the coin is fair' — and we test whether that particular value is plausible. Same engine, sharper question: hypothesis testing."

Hand-off (the week's work):
- Chapter 12 (primary reading, with the z table) — then Lecture Tutorial 12 (AI tutor; share link + summary).
-
Data Lab 12 ("Anatomy of a Poll" — audit a real or provided poll's margin of error) · Quiz 12 (end of week) · Discussion 12 ("The 400 vs. 18,000 Problem") · Assignment 12* (AI-coached).


Instructor FAQ — Common Stumbles

Student says / does Quick cure
"Why z this week when last week was t?" For a mean, spread needed its own second estimate (s) — t pays for that. p̂ feeds both center and SE, so no extra tax: means → t, proportions → z.
Plugs the count into the SE formula: √(60 × 40 ⁄ 100). The formula eats proportions, not counts: √(0.60 × 0.40 ⁄ 100). Quick tell: an SE bigger than 1 for a proportion is impossible.
"So there's a 95% chance p is in my interval." The 95% belongs to the method — about 19 of every 20 samples produce a capturing interval. p is fixed; your one interval either captured it or didn't.
"95% of customers are in the interval." The interval brackets the proportion, never individuals. People are yes/no data points, not values of p.
Rounds n = 1067.11 to 1067. Round up, always — 1067 misses the promised margin. The ceiling is a guarantee, not an estimate.
"The interval includes 50%, so it's a tie." Not a tie — undecided. The data can't rule out values on either side of 50; that's weaker than "equal" and honest reporting says so.
"To poll a bigger country you need a bigger sample." N isn't in the SE formula — one well-stirred spoonful tastes any size pot. What matters is n and random stirring (plus the 10% condition).
"The ±3 covers biased wording and nonresponse too." ME prices sampling luck only. Bias rides outside the ± — a precise interval from a biased sample is precisely wrong (Literary Digest, forever-example).
Checks 10-successes/10-failures with p instead of the data. Use the observed counts: successes = 60, failures = 40, both ≥ 10. If either count is single-digit, the z-interval is off the table.

Scope flag

This outline stays within Objective 6's proportion portion. The napkin rule (ME ≈ 1 ⁄ √n), the why-z-not-t contrast, and the Literary Digest callback are added context (not strictly required by the objective) — kept because they make the media segment and the conditions stick; cut them for a leaner session. Plus-four / Wilson adjusted intervals and exact binomial intervals are out of scope; two-proportion comparisons wait for Week 15.