Back to the Introduction to Statistics outline The Course Maker
Introduction to Statistics outline
Week 9 · Study guide

Week 9 — Midterm Study Guide · Weeks 1–8

Introduction to Statistics Generic evergreen edition

Course: Introduction to Statistics (18-week generic edition)
What it covers: Weeks 1–8 — everything from "where do data come from?" through the normal model. Objectives 1–4 in full, plus Objective 5's normal-distribution portion.
The exam it prepares you for: the Week 9 midterm — 50 questions × 2 points = 100, Midterm group (5% of the grade), closed to AI, sitting mid-week after the review session. Calculator fine; one page of notes if your instructor permits; any z-table values an item needs are printed inside that item.


How to use this guide (20 minutes now saves hours later)

  1. Diagnose first, reread last. Skim each week's "Can you…?" checklist below and mark every box you can't check out loud, from memory. Those boxes — not the whole guide — are your study list. Rereading what you already know feels productive and isn't.
  2. Study in the order: idea → trap → mini-example. For each weak box: read that week's big idea, then its traps (the exam's wrong answers are built from exactly these), then work the mini-example before looking at its solution.
  3. Then rehearse, actively. Take the practice exam closed-book (unlimited attempts, ungraded, zero items shared with the real exam), and run the exam-prep tutorial — it re-teaches whatever the practice exam catches. The plan at the end of this page sequences it all.
  4. Memorize almost nothing. The only numbers to know by heart: 68 – 95 – 99.7. Every z-table value the exam needs is printed in the item that needs it.

Week 1 — Statistics, Data & Study Design

The big idea. Statistics turns a measured sample into a claim about an unmeasured population: the statistic (the value we have) estimates the parameter (the value we want). Whether that leap deserves trust depends entirely on how the sample was chosen and what was actually recorded — method beats size.

Must-know terms: population · sample · census · parameter · statistic · categorical (nominal, ordinal) vs. quantitative (interval, ratio) · simple random sample · stratified · cluster · systematic · convenience · voluntary response · bias (undercoverage, nonresponse, response, voluntary-response) · observational study · experiment · confounding variable.

Can you…
- [ ] name the population, sample, parameter, and statistic in any survey scenario? (P→P, S→S — the letters line up.)
- [ ] classify any variable with the NOIR tests — is there order? are the gaps equal? does zero mean "none"?
- [ ] tell stratified (sample within every group) from cluster (take whole groups)?
- [ ] spot the bias in an opt-in or convenience sample — and explain why a bigger sample doesn't fix it?
- [ ] say why only a randomized experiment supports a cause-and-effect claim?

Classic traps. A number that labels (route number, ID, zip code) is nominal, not quantitative — ask whether arithmetic means anything. "Two million responses" never repairs a biased method (the 1936 Literary Digest poll is the forever-example). And a strong link is a handshake, not a push — hunt the confounder before believing any arrow.

Worked mini-example. A trail association records, for each volunteer shift: the crew name, the shift's difficulty tag (easy / moderate / strenuous), and the hours worked. Classify: crew name — nominal (a label); difficulty tag — ordinal (ordered, gaps unmeasurable); hours — ratio (zero hours means none). If the association then averages the hours of the 40 shifts that happened to file reports early to describe all 500 shifts, that average is a statistic from a voluntary-response-flavored sample — early filers may differ from everyone else.


Week 2 — Tables & Graphs

The big idea. Count, then share: a frequency table counts each category or class, and relative frequencies (count ÷ total) make shares comparable — they must sum to 1. Match the display to the variable — categories get separated bars; quantitative data get a number line (dot plot, histogram, stem plot) — then read the shape and stay alert for graphs that lie about honest numbers.

Must-know terms: frequency · relative frequency · class · bar chart · pie chart · histogram · dot plot · stem plot · distribution · shape (symmetric, skewed right/left, uniform, bimodal) · outlier (informal) · misleading graph.

Can you…
- [ ] compute a relative frequency and check that the shares total 1?
- [ ] pick the right display — and say why a pie chart needs parts of one whole totaling 100%?
- [ ] read a histogram's classes ("how many below…?") by adding bars?
- [ ] name a shape from the tail, not the peak?
- [ ] spot a truncated axis and say what it exaggerates?

Classic traps. Skew is named for the tail — a histogram with tall bars on the left and a long right tail is skewed right (the tail tells the tale). Bars apart = categories, bars touching = a number line — you can sort a bar chart, never a histogram. Outliers get investigated, never silently deleted. Bars start at zero.

Worked mini-example. A snack cart sorts one day's 40 sales: 18 pretzels, 12 lemonades, 10 fruit cups. Relative frequencies: 18 ÷ 40 = 0.45, 12 ÷ 40 = 0.30, 10 ÷ 40 = 0.25 — and 0.45 + 0.30 + 0.25 = 1.00 ✓ (your built-in error check). Right display: a bar chart (three categories, separated bars, axis starting at zero). A pie chart would also be legal here — the three counts are non-overlapping parts of one whole.


Week 3 — Center & Spread

The big idea. Three centers (mean = balance point, median = sorted middle, mode = most frequent) and two spread kits (SD with the mean; IQR with the median). Resistance decides which pair to report: extreme values drag the mean and SD but barely budge the median and IQR — the mean chases the tail. z-scores turn any value into "how many SDs from the mean," a ruler that works across different scales.

Must-know terms: mean (x̄, μ) · median · mode · resistant measure · range · deviation · variance (s², σ²) · standard deviation (s, σ) · five-number summary · quartiles · IQR · 1.5×IQR rule · boxplot · z-score.

Can you…
- [ ] compute a small dataset's mean, median, and sample SD (deviations → square → add → ÷(n−1) → root)?
- [ ] choose the honest summary pair for skewed data — and defend the choice?
- [ ] build fences at Q1 − 1.5·IQR and Q3 + 1.5·IQR and flag formal outliers?
- [ ] compute and interpret a z-score — sign = direction, size = distance?

Classic traps. Variance is in squared units — it's scratch work; the SD is the number you say out loud (s² = 9 means s = 3). The boxplot's middle line is the median, never the mean. A negative z isn't "bad" — it's below the mean, which for times and bills is often good. And z is a distance in SDs, not a percentile — turning z into a percentage takes the Week 8 normal model.

Worked mini-example A (SD by hand). A bike-repair bench logs five tune-up times: 12, 12, 15, 18, 18 minutes. Mean = 75 ÷ 5 = 15. Deviations: −3, −3, 0, 3, 3 → squares 9, 9, 0, 9, 9 → sum 36 → 36 ÷ 4 = 9 (variance) → s = 3 minutes: tune-ups typically sit about 3 minutes from the mean.

Worked mini-example B (five-number summary). A pop-up market logs seven stall setup times, sorted: 20, 24, 26, 30, 34, 38, 48 minutes. Median = 30; Q1 = median of the lower half (20, 24, 26) = 24; Q3 = median of (34, 38, 48) = 38; IQR = 38 − 24 = 14. Upper fence = 38 + 1.5(14) = 59 — so the 48 is a long shift, but not a formal outlier. The rule decides, not the feeling.


Week 4 — Relationships Between Two Variables

The big idea. Two quantitative variables → a scatterplot (x explains, y responds), read as Direction, Form, Strength — then Stragglers, with r compressing the linear pattern into one unitless number in [−1, +1]: sign = direction, size = strength. Two categorical variables → a two-way table, where association means the conditional distributions differ across groups. Either way: association is never, by itself, causation.

Must-know terms: explanatory vs. response · scatterplot · direction/form/strength · correlation r · two-way table · marginal distribution · conditional distribution · association · lurking variable · confounding.

Can you…
- [ ] assign explanatory (x) and response (y) roles and describe a scatterplot in a sentence?
- [ ] rank strengths by |r| — and say what r deliberately ignores (units, axis-swap)?
- [ ] compute a conditional percent with the right denominator ("among ___" → that group's total)?
- [ ] decide association by comparing conditionals — and then refuse the causal leap, naming a lurking variable?

Classic traps. r sees straight lines only — a strong arch can produce r ≈ 0, so look at the plot first. r = −0.9 is stronger than r = +0.5 (strength is distance from 0). The classic table error is the wrong denominator — the joint percent (out of everyone) answers a different question than the conditional (out of the group). And precision about the handshake is not evidence of a push.

Worked mini-example. A community-garden sign-up sheet: of 60 plot-holders with compost training, 45 renewed their plots (45 ÷ 60 = 75%); of 40 without training, 10 renewed (10 ÷ 40 = 25%). The conditionals differ (75% vs. 25%) → renewal is associated with training. Causal? Not on this evidence — nobody assigned training; keener gardeners may both seek training and renew (a lurking variable). An experiment would randomly assign the training.


Week 5 — Probability Foundations

The big idea. Probability is a long-run promise, not a short-run guarantee: P(A) is the proportion of times A would occur over endless repetitions. Four word-cues run every problem — NOT (complement), OR (add, minus the overlap), AND (multiply, if independent), GIVEN (shrink the world and re-count).

Must-know terms: random phenomenon · law of large numbers · sample space · event · complement · disjoint · addition rules · independence · multiplication rule · conditional probability P(A | B) · gambler's fallacy.

Can you…
- [ ] interpret "P = 0.15" correctly ("over many repetitions, about 15%…") without promising any short-run count?
- [ ] apply P(not A) = 1 − P(A), and treat any answer outside 0-to-1 as a smoke alarm?
- [ ] use P(A or B) = P(A) + P(B) − P(A and B), subtracting the overlap?
- [ ] compute P(A | B) from a table by finding the "given" group's total first?
- [ ] explain why streaks create no debts (the die has no memory)?

Classic traps. "Two outcomes, so 50/50" — equal probabilities must be earned, never assumed. Disjoint and independent are different facts (disjoint events are never independent). P(A | B) ≠ P(B | A) — say the "given" world out loud before dividing. And the long run fixes proportions by swamping, not compensating — no outcome is ever "due."

Worked mini-example. A ferry snack bar finds P(a passenger buys anything) = 0.85, so P(buys nothing) = 1 − 0.85 = 0.15. Separately, P(buys a drink) = 0.5, P(buys a snack) = 0.3, and P(buys both) = 0.2 — so P(drink or snack) = 0.5 + 0.3 − 0.2 = 0.6. (Adding without subtracting would double-count the both-buyers.)


Week 6 — Random Variables

The big idea. A random variable attaches numbers to chance — discrete you count, continuous you measure — and its distribution table (probabilities in [0,1] summing to exactly 1) is its whole personality. E(X) = Σ x·P(x) is the long-run average; SD(X) is the typical distance from it. Under Y = a + bX: adding shifts the center; multiplying stretches both.

Must-know terms: random variable · discrete vs. continuous · probability distribution · legitimate distribution · expected value E(X) · variance and SD of X · linear transformation · density curve · uniform density.

Can you…
- [ ] find a missing probability using "the table totals 1"?
- [ ] compute E(X) — weighting each value by its probability — and interpret it as a long-run average?
- [ ] transform: E(a + bX) = a + b·E(X), SD(a + bX) = |b|·SD(X)?
- [ ] explain why P(X = one exact value) = 0 for a continuous variable (no width, no area, no probability)?

Classic traps. Expected value is what you'd average, not what you'd expect — E(X) needn't be a possible value. Averaging the listed values (ignoring the probabilities) is the classic E(X) error. Variance vs. SD: one square root apart, never interchangeable. And adding a constant moves the whole number line without changing the spacing — SD untouched.

Worked mini-example. A weekend kayak rental returns X extra paddles requested per booking: P(0) = 0.3, P(1) = 0.4, P(2) = 0.2, P(5) = 0.1 (group bookings). Legitimacy: 0.3 + 0.4 + 0.2 + 0.1 = 1.00 ✓. E(X) = 0(0.3) + 1(0.4) + 2(0.2) + 5(0.1) = 0.4 + 0.4 + 0.5 = 1.3 paddles — the long-run average per booking, even though no single booking can request 1.3 paddles.


Week 7 — The Binomial Distribution

The big idea. When a situation passes B·I·N·S — Binary outcomes, Independent trials, Number of trials fixed, Same p — the count of successes X is binomial, and everything about it comes from n and p: P(X = k) = C(n, k) p^k (1−p)^(n−k) (ways × wins × losses), μ = np, σ = √(np(1−p)).

Must-know terms: binomial setting (B·I·N·S) · trial · success/failure · n, p · binomial coefficient C(n, k) · binomial formula · "at least one" complement shortcut · binomial mean and SD · cumulative probability · =BINOM.DIST(…, TRUE).

Can you…
- [ ] check B·I·N·S and name the impostors (no fixed n; without-replacement draws)?
- [ ] compute P(exactly k) with the ways factor included?
- [ ] use at least one = 1 − P(none)?
- [ ] read "exactly / at least / at most" like a lawyer, and know that TRUE in =BINOM.DIST means cumulative?
- [ ] compute and interpret np and √(np(1−p)) — expected, not guaranteed?

Classic traps. Forgetting C(n, k) makes answers absurdly small — did I count the ways? p doesn't have to be 0.5 — it just has to be the same every trial. μ = np is a long-run average, not a promise. And independent trials don't remember — no "due" successes.

Worked mini-example. A pottery studio's kiln cracks each mug independently with probability 0.2. Load 5 mugs: P(exactly 1 cracks) = C(5, 1)(0.2)¹(0.8)⁴ = 5 × 0.2 × 0.4096 = 0.4096. Across 25 mugs fired in a week: expected cracks = 25 × 0.2 = 5, give or take σ = √(25 × 0.2 × 0.8) = √4 = 2. Nine cracked mugs in a week (z-thinking: two SDs above 5) would be a real signal, not ordinary wobble.


Week 8 — The Normal Distribution

The big idea. A density curve carries proportion as area (total area 1; height is never proportion). The normal model N(μ, σ) is drawn entirely by two numbers, obeys 68–95–99.7 (IF bell-shaped — that password is mandatory), and converts values to percentages in both directions: forward (x → z = (x−μ)/σ → left-tail area) and inverse (percent → z → x = μ + z·σ). Before trusting any of it, check that the bell is actually there.

Must-know terms: density curve · normal model N(μ, σ) · empirical rule · z-score · standardizing · standard normal N(0, 1) · left-tail area · percentile · inverse normal · assessing normality.

The course z-table (the exam prints any of these values inside the items that need them):

z area to the left z area to the left
−2.5 0.0062 0.5 0.6915
−2.0 0.0228 1.0 0.8413
−1.5 0.0668 1.25 0.8944
−1.25 0.1056 1.5 0.9332
−1.0 0.1587 2.0 0.9772
−0.5 0.3085 2.5 0.9938

(And z = 0 → 0.5000 by symmetry.)

Can you…
- [ ] apply 68–95–99.7 to any N(μ, σ) — and refuse to apply it to a skewed pile?
- [ ] sketch-and-shade before every lookup, and sanity-check which tail should be big?
- [ ] run a forward calculation (below / above = 1 − area / between = difference of areas)?
- [ ] run an inverse calculation (area → z → x = μ + z·σ), walking the correct direction from the mean?
- [ ] compare two values from different distributions with z-scores?

Classic traps. Height is not proportion — area is. The wrong tail (reporting 0.9772 when the question asked "above") — the sketch catches it every time. Divide by σ, never σ². And one passed 68% check doesn't prove normality — judge the histogram, the empirical-rule check, and the skew/outlier hunt together, then defend the call.

Worked mini-example (both directions on one model). Delivery-window arrivals at a bakery counter are approximately N(80, 8) seconds. Forward: P(under 92 s): z = (92 − 80)/8 = 1.5 → area 0.9332 — about 93% of arrivals beat 92 seconds. Inverse: the slowest 15.87% start where 84.13% lie below → z = 1.0 → 80 + 1(8) = 88 seconds. Sketch both, shade both, and check: 92 is above the mean, so "below 92" should be a big number ✓.


The cross-week connections map (how the midterm thinks)

The first half is one pipeline — the synthesis items simply run more than one segment of it:

  • W1 → everything: if the sample was biased, every later computation is confidently wrong. (Expect at least one item that hands you a beautiful statistic from an ugly sample.)
  • W2 ↔ W3: shape chooses the summary — skew spotted in the histogram is why you report the median/IQR pair.
  • W3 → W8: the z-score you learned as relative standing becomes the normal model's input; the SD you computed by hand becomes the σ in N(μ, σ).
  • W4 ↔ W5: the two-way table's conditional percent is a conditional probability wearing Week 4 clothes — same denominator discipline.
  • W5 → W6 → W7: probability rules price single events; random variables average whole distributions (E(X)); the binomial packages n independent yes/no events into one formula with μ = np.
  • W7 → W8: the binomial's histogram reaches for a bell as n grows — so np and √(np(1−p)) plus 68–95–99.7 can price a whole season of trials (that's a synthesis item shape).

Your prep sequence (fits in two or three sittings)

  1. Sitting 1 — diagnose (≈45 min). Work every "Can you…?" box above from memory; list the misses. Read only those sections' ideas and traps; work their mini-examples on paper.
  2. Sitting 2 — rehearse (≈60 min). Take the practice exam closed-book, one honest attempt, no AI, notes away — it mirrors the real blueprint and shares zero items with it. Sort your misses by week; re-drill with the matching mini-examples. Retake until clean (attempts are unlimited and ungraded).
  3. Sitting 3 — drill and seal (≈60–90 min). Run Exam-Prep Tutorial 9 (this week's graded tutorial — submit the share link + completion summary). It diagnoses across all eight weeks, re-teaches your weak spots, and ends with a 10-question mixed mock round. Finish with one pass through the eight classic traps in the review outline — the exam's distractors are built from them.

Honest test-taking notes for this exam: answer everything (no penalty for guessing); "at least / exactly / at most" are different questions — read like a lawyer; find the "among ___" denominator before touching a table item; sketch-and-shade before any z item; and if a probability comes out above 1 or below 0, that's the smoke alarm, not a rounding issue. The midterm is 5% — a checkpoint that tells you what consolidated. The weekly work remains the grade engine.

Placement note (for the instructor): publish as an ungraded Page in the Week 9 module, available from the module's start so it posts ahead of the exam window.