Back to the Introduction to Statistics outline The Course Maker
Introduction to Statistics outline
Week 13 · Quiz

Week 13 — Quiz (auto-graded) · Hypothesis Testing: Foundations

Introduction to Statistics Generic evergreen edition

Course: Introduction to Statistics (18-week generic edition)
Objective tested: Objective 7 — the logic of significance testing: hypotheses, p-values, α, decisions, and error types.
Points: 10 (1 each) · Assignment group: Quizzes (15% of grade) · Due: end of Week 13 · Closed to AI.

This is the human-readable quiz with its vetted answer key and feedback. The import-ready Classic QTI is in F-quiz-week-13-qti.xml; the reusable item-bank entries and the Canvas placement block are at the bottom of this file.


Blueprint

# Type Concept Objective
1 Multiple choice Writing H₀ and Hₐ (direction; parameter not statistic) 7
2 Multiple choice p-value meaning in context 7
3 Multiple choice Decision at α + conclusion language 7
4 Multiple choice Type I error in context 7
5 Multiple choice Type II error in context 7
6 Matching Core terms: H₀ / Hₐ / p-value / α 7
7 Multiple answer Properties of the significance level α 7
8 True / False "Fail to reject = proven true" misconception 7
9 Multiple choice Statistical vs. practical significance 7
10 Multiple choice Significance depends on the chosen α 7

No trick questions; every p-value is supplied (no computation this week); distractors target the Week 13 misconceptions named in the lecture outline.


Questions, key, and feedback

Q1 (MC). A cereal maker prints "500 grams of cereal per box, on average" on its packaging. A consumer group suspects the true mean fill is less than 500 grams and plans a test. Which hypotheses should the group use?
- A. H₀: μ < 500 vs. Hₐ: μ = 500
- B. H₀: μ = 500 vs. Hₐ: μ < 500
- C. H₀: x̄ = 500 vs. Hₐ: x̄ < 500
- D. H₀: μ = 500 vs. Hₐ: μ > 500
Feedback: H₀ holds the claim as stated with the equals sign; the suspicion ("less") points the alternative left. Hypotheses describe the population μ — never the sample's x̄. (A = the roles swapped; C = the statistic-on-trial trap; D = the wrong tail.)

Q2 (MC). The group weighs a random sample of boxes, and its software reports a p-value of 0.03. Which statement is what "0.03" actually means?
- A. There is only about a 3% probability that the company's 500-gram average-fill claim is actually true
- B. About 3% of the individual boxes the company fills contain less than 500 grams of cereal
- C. If the true mean really were 500 g, samples this light or lighter would occur about 3% of the time
- D. About 3% of the measurements in the consumer group's sample must have contained weighing errors
Feedback: The p-value lives inside the what-if world where H₀ is true: it reports how often chance alone produces data at least this extreme. It is never the probability that the claim is true. (A = the classic misread.)

Q3 (MC). A commuter-rail operator advertises that 90% of its trains arrive on time. A transit watchdog tests H₀: p = 0.90 vs. Hₐ: p < 0.90, and its software reports a p-value of 0.21. At α = 0.05, what is the correct decision?
- A. Reject H₀ — there is convincing evidence the true on-time rate is below 90%
- B. Accept H₀ — the test confirms the operator's 90% on-time claim is true
- C. Reject Hₐ and conclude that the true on-time rate is exactly 90 percent
- D. Fail to reject H₀ — there is not convincing evidence the rate is below 90%
Feedback: 0.21 > 0.05, so the evidence doesn't clear the bar: fail to reject, stated in context. "Accept H₀" and "the claim is confirmed" are banned readings — a big p-value means compatible, not proven.

Q4 (MC). In the watchdog's test of the rail operator (H₀: p = 0.90 on time), which outcome would be a Type I error?
- A. Concluding the on-time rate is below 90% when the true rate really is 90%
- B. Concluding there's not enough evidence when the true rate really is below 90%
- C. Choosing a significance level that is too small before collecting the sample
- D. Recording several trains as late when they actually arrived at the station on time
Feedback: Type I = rejecting a true H₀ — the false alarm (its probability is α). Option B describes the Type II miss; D is a data-recording error, not an inference error.

Q5 (MC). Back to the cereal test (H₀: μ = 500 vs. Hₐ: μ < 500). Suppose the truth is that boxes really do average only 493 grams — but the group's sample yields p = 0.14, so the test fails to reject H₀. What has happened?
- A. A Type I error — the test raised a false alarm about the fill claim
- B. No error — failing to reject is always the correct call when p exceeds α
- C. A Type II error — the test missed a claim that really is false
- D. A sampling error — with p = 0.14 the sample must have been biased
Feedback: The procedure was followed correctly, yet the verdict is factually wrong: a false H₀ escaped rejection. That miss is precisely the Type II error — following the rule doesn't make the conclusion true.

Q6 (Matching). Match each term to its description.
| Term | Correct description |
|---|---|
| Null hypothesis (H₀) | The "nothing's going on" explanation — holds the equals sign and gets the benefit of the doubt |
| Alternative hypothesis (Hₐ) | The suspicion — the conclusion that requires convincing evidence before we may reach it |
| p-value | The probability, computed assuming the null is true, of data at least as extreme as observed |
| Significance level (α) | The false-alarm risk chosen before the data are examined — the line the p-value must beat |
Feedback: The courtroom in four rows: the defendant, the accusation, the strength of the evidence, and the standard of proof.

Q7 (Multiple answer — select all that apply). Which statements about the significance level α are true?
- A. It is chosen before the data are examined
- B. It equals the probability of a Type I error when H₀ is true
- C. It is calculated from the sample after the data are collected
- D. Making α smaller raises the risk of a Type II error, everything else equal
- E. It gives the probability that the null hypothesis is true
Feedback: α is a pre-set standard of proof: your false-alarm budget (A, B), with the see-saw cost that a stricter bar produces more misses (D). It is never computed from the data (C) and never a probability about H₀ itself (E).

Q8 (True / False). "If a test's p-value is 0.47, the correct conclusion is that the null hypothesis has been proven true."
- True
- False
Feedback: False. A large p-value means the data are compatible with H₀ — the evidence didn't clear the bar. "Fail to reject" is never "accept": not guilty ≠ innocent, and absence of evidence is not evidence of absence.

Q9 (MC). A cereal company tests a tweaked filling machine on 250,000 boxes and finds the mean fill rose by 0.4 gram, with a p-value of 0.001. What is the best reading of this result?
- A. The rise must be both large and practically important, because the p-value is so extremely small
- B. Statistically significant, but 0.4 g may be too small to matter — huge samples detect tiny effects
- C. The result must be a computational mistake, because an effect of 0.4 gram is too trivial to be real
- D. The tiny p-value means there is only a 0.1% chance that the filling machine was actually unchanged
Feedback: The p-value measures surprise, not size: with n = 250,000, even a trivial true difference is nearly impossible to blame on chance. Ask "is it significant?" and then always "how big is it?" (D repeats the P(H₀ | data) misread at a new address.)

Q10 (MC). A test produces a p-value of 0.03. Using the standard decision rule, at which significance levels would the result be declared statistically significant?
- A. At α = 0.05 but not at α = 0.01
- B. At α = 0.01 but not at α = 0.05
- C. At both α = 0.05 and α = 0.01
- D. At neither α = 0.05 nor α = 0.01
Feedback: 0.03 ≤ 0.05 (clears the everyday bar) but 0.03 > 0.01 (fails the stricter one). "Statistically significant" is always relative to the α set before the data arrived.


Answer key (quick reference)

Q Answer
1 B
2 C
3 D
4 A
5 C
6 H₀→nothing's going on, equals sign / Hₐ→the suspicion needing evidence / p-value→P(data at least this extreme, assuming null) / α→pre-chosen false-alarm risk
7 A, B, D
8 False
9 B
10 A

Quality gate (self-checked): each single-answer item has exactly one correct option; the multiple-answer item's three true statements are the only true options listed; no positional pattern in the key (B C D A C · B A) and no length giveaway (options within each item are comparable lengths); every p-value is supplied in the stem — no computation to mis-key; all decisions re-verified in the week's math script (0.03 ≤ 0.05, 0.21 > 0.05, 0.14 > 0.05, 0.001 ≤ 0.05, 0.03 > 0.01); no item asserts a fact outside the Week 13 course definitions; no scenario reuses the tutorial, practice, chapter, lab, or assignment surfaces (quiz surfaces: cereal-box fill claims and rail on-time claims only).


Item-bank entries (for variants + the final)

All ten items are tagged week=13 · objective=7 · topic=hypothesis-testing-foundations and deposited in Item Bank: Week 13 — Hypothesis Testing: Foundations with idents w13q1w13q10. The final (Week 18) and per-term variant updates draw fresh variants from this bank's concepts — never these live stems. (Tags: w13q1 hypothesis-setup, w13q2 p-value-meaning, w13q3 decision-conclusion, w13q4 type-i, w13q5 type-ii, w13q6 core-terms, w13q7 alpha-properties, w13q8 fail-to-reject, w13q9 practical-significance, w13q10 alpha-dependence.)

Canvas placement block

canvas_object    = Quizzes::Quiz
title            = "Week 13 Quiz — Hypothesis Testing: Foundations"
assignment_group = "Quizzes"
points_possible  = 10
grading_type     = points
due_offset_days  = 6        # end of the module's week
published        = true
shuffle_answers  = true
This is the human-readable quiz with its vetted answer key and rationale. The import-ready Classic-QTI version (F-quiz-week-13-qti.xml) ships inside the course's .imscc package — it lands in the Canvas gradebook on import.