The Interior — MathGrades 11–12

Unit 26 · Statistics: From a Sample to a Claim

A unit of the course: the story, then chapter by chapter — sections, numbered lessons, a source or the numbers to read, three checks each — a review per chapter, and the wrap-up at the end.

← Math, the whole course

Drawn scene: a polling place at dusk with a line of people waiting at a lit door, a bell-curve-shaped hill behind, a clipboard and a ballot box on the table
26Unit

Statistics: From a Sample to a Claim

Statistics and Probability

A nurse says a child is at the 90th percentile. A news site says a candidate leads 48% to 46%. A commercial says a supplement cut risk by 30%. Each of these is a claim built from data, and each can be read well or read badly. This unit is about reading them well. You will learn how a single bell-shaped curve describes heights, test scores and coin flips; how a few hundred randomly chosen people can stand in for millions; and how to decide whether a surprising result is real or just the luck of the draw.

The tools are few and they build on each other. Center, spread and shape describe a pile of data. Z-scores and percentiles locate one value inside the pile. Probability rules say how likely combinations of events are, and expected value turns those probabilities into decisions. Then the direction flips: instead of using a known population to predict a sample, you use one sample to make a claim about an unknown population, and you attach an honest margin of error to that claim.

By the end you will be able to compute a percentile from a z-score, test whether a two-way table shows independence, find the expected value of a game, explain why a poll of 1,000 people has a margin of error near 3 points, run the logic of a hypothesis test, and take apart a headline that confuses a 20% relative change with a real difference. Those are skills for a statistics course, and they are also skills for being a citizen who cannot be fooled by a number.

How we figured it out
1662

John Graunt studies London's death records and shows that patterns emerge from counts of many people

1713

Jacob Bernoulli's Ars Conjectandi proves that sample proportions settle toward the true probability as trials increase

1733

Abraham de Moivre shows that binomial counts pile up in a bell-shaped curve

1809

Carl Friedrich Gauss uses the normal curve to describe errors in astronomical measurements

1835

Adolphe Quetelet applies the bell curve to human heights and other social measurements

1900

Karl Pearson introduces the chi-square test for comparing observed counts with expected ones

1908

William Gosset, writing as Student, publishes a method for drawing conclusions from small samples

1925

Ronald Fisher's Statistical Methods for Research Workers spreads significance testing and randomized design

1936

The Literary Digest poll fails despite two million responses; Gallup's smaller representative sample succeeds

1948

Polls that stopped too early help produce the famous wrong headline 'Dewey Defeats Truman'

1954

The Salk polio vaccine field trial uses random assignment, placebos and blinding on hundreds of thousands of children

Today

Simulation on computers lets students build sampling distributions and p-values without formulas

Chapter

Distributions and the Normal Curve

Statistics
Big questionHow can one curve describe heights, test scores and coin flips, and when should we refuse to trust it?
The story

The Chart on the Wall of Exam Room 3

A little sister, a nurse with a pencil, and a number on a chart that sounds like a grade but is not one.

Marcus takes his six-year-old sister Nia to her checkup at a clinic on Chicago's West Side because their mother cannot leave work. The nurse measures Nia against the wall, writes 48 inches on a clipboard, and finds the spot on a chart taped beside the door. She makes a small dot, looks up and says, 'She is right around the 90th percentile for height.' Nia grins. Marcus nods, but on the bus home he admits he is not sure what that meant. Was 90 a score? Did she pass?

A *percentile* is not a grade. It is a position in a crowd. Saying Nia is at the 90th percentile means that if you lined up 100 six-year-old girls by height, about 90 of them would be shorter than Nia and about 10 would be taller. The chart on the wall was built by measuring many thousands of children and recording where each height fell. The curves printed on it, one for the 5th percentile, one for the 50th, one for the 95th, are simply lines drawn through those positions at every age.

Here is the part that surprised Marcus when he looked it up. Heights at a given age pile up in the middle and thin out at both ends in a very regular way. Most six-year-old girls are close to the average, a few are quite short, a few are quite tall, and the drop-off on each side is nearly the same. That bell-shaped pile has a name, the normal curve, and once you know its center and its spread you can predict what percent of children fall in any range without measuring anyone new.

That is the promise of this chapter. One curve, described by two numbers, turns a height into a percentile, a test score into a rank, and a run of coin flips into a probability. But the same tools also tell you when the curve is the wrong model, when the pile of data leans hard to one side, or when a count cannot possibly go below zero. Marcus wanted to know what 90 meant. By the end of the chapter you will be able to explain it to him, and to say when the chart would lie.

Talk about itIf Nia is at the 90th percentile and a classmate is at the 50th, does that mean Nia is almost twice as tall? Explain what the two numbers actually compare.
Section 1

Describing a Pile of Data

59.1

Center, Spread and Shape

Main ideaTo describe a set of numbers, report its center, how far the numbers spread from that center, and the shape of the pile.

Eight students report how many hours they slept last night: 2, 4, 4, 4, 5, 5, 7, 9. The is the total divided by the count: 2 + 4 + 4 + 4 + 5 + 5 + 7 + 9 = 40, and 40 ÷ 8 = 5 hours. The is the middle value when the numbers are in order. With eight values the middle is between the 4th and 5th, which are 4 and 5, so the median is 4.5 hours. Mean and median are both measures of center, but they answer slightly different questions: the mean balances the total, the median splits the group in half.

Center alone is not enough. Two classes can both average 5 hours while one class is tightly packed near 5 and the other ranges from 1 to 10. The measures spread as a typical distance from the mean. For the sleep data: subtract the mean from each value to get −3, −1, −1, −1, 0, 0, 2, 4. Square each: 9, 1, 1, 1, 0, 0, 4, 16, which sum to 32. Divide by the count, 32 ÷ 8 = 4, and take the square root: √4 = 2. The standard deviation is 2 hours, so a typical student is about 2 hours from the mean. A calculator’s sample standard deviation divides by one less than the count and gives a slightly larger number; either way, bigger means more spread out.

Shape is the third piece. Draw a , a bar for each range of values, and look at the pile. If the left and right halves are mirror images, the shape is , and the mean and median match. If one tail stretches far out on the right, as with household incomes where a few very large numbers pull the mean up, the shape is right and the mean sits above the median. In the sleep data the 9 pulls the mean (5) above the median (4.5), a small right skew. A common mistake is to report the mean for skewed data as if it were typical; for skewed data the median is usually the more honest center.

Words to know
mean
the total of the values divided by how many values there are
median
the middle value when the numbers are placed in order; half the values sit below it
standard deviation
a typical distance of the values from the mean; larger means more spread out
histogram
a bar graph where each bar counts how many values fall in a range
symmetric
a shape whose left and right halves mirror each other
skewed
a shape with one tail stretched much farther out than the other
Check yourself

1. The values 4, 8, 9 and 11 have what mean?

2. For the sleep data 2, 4, 4, 4, 5, 5, 7, 9 (mean 5), what is the standard deviation when you divide the squared deviations by the count?

3. A histogram of home prices in a city has a long tail stretching to the right. Which statement is true?

59.2

The Normal Curve and the 68-95-99.7 Rule

Main ideaWhen data follow a normal curve, about 68% fall within 1 standard deviation of the mean, 95% within 2, and 99.7% within 3.

Many measurements, such as heights of adults of the same sex, pile up in a smooth bell shape: highest at the mean, symmetric, tails thinning out on both sides. This bell is the , and it is completely described by two numbers, its mean and its standard deviation. Suppose a model for adult women’s heights uses a mean of 64 inches and a standard deviation of 2.5 inches. The peak of the bell is at 64. One standard deviation on each side reaches from 64 − 2.5 = 61.5 to 64 + 2.5 = 66.5 inches.

The says what fraction of a normal pile sits in those bands. About 68% of values lie within 1 standard deviation of the mean, about 95% within 2, and about 99.7% within 3. For the height model: about 68% of women are between 61.5 and 66.5 inches; about 95% are between 64 − 5 = 59 and 64 + 5 = 69 inches; about 99.7% are between 56.5 and 71.5 inches. Because the curve is symmetric, each half of a band holds half its percent. So 34% are between 64 and 66.5, and 34% are between 61.5 and 64.

Now the ends. If 95% are within 2 standard deviations, the other 5% are outside, split evenly: 2.5% below 59 inches and 2.5% above 69 inches. If 68% are within 1 standard deviation, 32% are outside, so 16% are above 66.5 inches. A common error is to forget the split and say 32% are above 66.5. Another is to add standard deviations without multiplying: 2 standard deviations above 64 is 64 + 2 × 2.5 = 69, not 64 + 2 = 66. Write the band endpoints first, then read off the percents.

Words to know
normal curve
the symmetric bell-shaped curve, described entirely by its mean and standard deviation
empirical rule
for a normal curve, about 68%, 95% and 99.7% of values lie within 1, 2 and 3 standard deviations of the mean
tail
the thin far ends of a distribution, where few values fall
Check yourself

1. Scores on a test are normal with mean 100 and standard deviation 15. Between which two scores do about 95% of people fall?

2. With the same mean of 100 and standard deviation of 15, about what percent of people score above 115?

3. Using the 68-95-99.7 rule, about what percent of values fall between the mean and 1 standard deviation above the mean?

59.3

Z-Scores Measure Distance in Standard Deviations

Main ideaA z-score tells how many standard deviations a value sits above or below the mean, so scores from different scales can be compared.

A history test has mean 70 and standard deviation 8. Jada scores 82. How unusual is that? Measure her distance from the mean in units of standard deviation: 82 − 70 = 12 points above the mean, and 12 ÷ 8 = 1.5 standard deviations. That number, 1.5, is her . The formula is z = (value − mean) ÷ standard deviation. A score of 58 gives z = (58 − 70) ÷ 8 = −12 ÷ 8 = −1.5, the same distance but below the mean. The sign matters: positive means above, negative means below.

Z-scores let you compare things measured on different scales. Suppose Test A has mean 500 and standard deviation 100, and Test B has mean 20 and standard deviation 5. Luis scores 650 on Test A and 28 on Test B. Which is the stronger result? On A: z = (650 − 500) ÷ 100 = 150 ÷ 100 = 1.5. On B: z = (28 − 20) ÷ 5 = 8 ÷ 5 = 1.6. The 28 is slightly farther above its mean, so it is the stronger performance, even though 650 is a much bigger number.

You can also go backwards. If a value has z = 2 on the 100-and-15 scale, it is 2 × 15 = 30 points above 100, so the value is 130. In general, value = mean + z × standard deviation. Two mistakes are common: dividing the raw value by the standard deviation without subtracting the mean first (82 ÷ 8 ≈ 10.25, which means nothing), and dropping the negative sign on values below the mean. Always subtract, then divide, and keep the sign.

Words to know
z-score
how many standard deviations a value is above (positive) or below (negative) the mean
standardize
convert a value into a z-score so it can be compared with values on other scales
raw score
the original measured value before it is turned into a z-score
Check yourself

1. A distribution has mean 50 and standard deviation 10. What is the z-score of the value 35?

2. On a scale with mean 100 and standard deviation 15, which raw score has a z-score of 2?

3. Test A: mean 80, SD 4, Priya scores 86. Test B: mean 60, SD 10, Priya scores 72. Which result is relatively stronger?

Section 2

Percentiles and the Limits of the Model

59.4

From a Z-Score to a Percentile

Main ideaA percentile is the percent of values below a given value; on a normal curve, the z-score tells you the percentile through a standard table.

Back to Nia’s height chart. A is the percent of the group that falls below a value. On a normal curve, each z-score matches one percentile, and a short table gives the match. A value exactly at the mean has z = 0 and is the 50th percentile: half the group is below. A value with z = 1 is at about the 84th percentile, because 50% are below the mean and another 34% are between the mean and 1 standard deviation above, and 50 + 34 = 84. A value with z = −1 is at about the 16th percentile, since 50 − 34 = 16.

The table fills in the values between. The z-score for the 90th percentile is about 1.28. If six-year-old girls’ heights are modeled with mean 45 inches and standard deviation 2 inches, then the 90th percentile height is 45 + 1.28 × 2 = 45 + 2.56 ≈ 47.6 inches. That is roughly where the nurse’s dot for Nia fell. Going the other way, a score of 82 on the test with mean 70 and standard deviation 8 has z = 1.5, and the table gives about 93% below, so 82 is at the 93rd percentile.

To find the percent between two values, find each percentile and subtract. Between z = −1 (about 15.9%) and z = 0.5 (about 69.1%): 69.1 − 15.9 = 53.2% of values. Two cautions. First, a percentile is a rank, not an amount: the 90th percentile is not twice the 45th. Second, the table only works if the data are actually normal, which is the subject of the next lesson. A student who plugs a skewed data set into a z-table will get a tidy percent that is simply wrong.

Words to know
percentile
the percent of the group whose values fall below a given value; the 90th percentile is above 90% of the group
z-table
a table that gives, for each z-score, the percent of a normal curve below it
rank
a position in an ordered group, as opposed to an amount
Check yourself

1. On a normal curve, a value with z-score 1 is at about which percentile?

2. A test has mean 70 and standard deviation 8. Using the table (z = 1.5 gives about 93.3% below), at what percentile is a score of 82?

3. Using the table, about what percent of values lie between z = −1 (15.9% below) and z = 0.5 (69.1% below)?

59.5

When the Normal Model Fits

Main ideaUse the normal model only when the data are roughly symmetric and bell-shaped; skewed, two-peaked or bounded data break it.

The normal curve is a model, and a model can be wrong. Before using z-scores or the empirical rule, look at a histogram. Heights of adult men, measurement errors, and averages of many small effects tend to be close to normal. Household income is not: most households are near the middle, but a few earn enormous amounts, and the tail on the right stretches for miles. Time waiting in line is not either: it cannot go below zero, but it can run very long. A normal model applied to either would predict percents that never happen.

Here is a quick test. Suppose a data set of wait times has mean 10 minutes and standard deviation 8 minutes. If the data were normal, the empirical rule would put 2.5% of waits below 10 − 2 × 8 = −6 minutes. A negative wait is impossible, so the model cannot fit. Any time the mean is less than about two standard deviations from a hard boundary such as zero, the normal curve is a poor choice. A second test is to count: if far fewer than 68% of the values are within 1 standard deviation of the mean, the shape is not a bell.

data, which have two separate peaks, also break the model. Ages at a middle-school open house are bimodal: one hump near 13 for students and one near 42 for parents. The mean, somewhere around 28, describes almost nobody, and a z-score built from it is meaningless. When the model does not fit, use percentiles computed directly from the sorted data instead of from a z-table, and report the median and the middle 50% rather than the mean and standard deviation.

Words to know
model
a simplified mathematical description of real data that may fit well or badly
bimodal
a distribution with two distinct peaks, often from two different groups mixed together
boundary
a hard limit that values cannot cross, such as zero for a wait time
right-skewed
having a long tail of large values on the right, so the mean exceeds the median
Check yourself

1. Which data set is most likely to be well described by a normal curve?

2. A data set of delivery times has mean 20 minutes and standard deviation 15 minutes. Why does the normal model fit poorly?

3. A histogram shows two clear peaks. What is the best description?

Section 3

Rules of Probability

59.6

Independence and the Multiplication Rule

Main ideaWhen two events are independent, the probability that both happen is the product of their separate probabilities.

Flip a coin twice. The chance of heads on the first flip is 1/2, and the coin has no memory, so the chance of heads on the second flip is still 1/2 no matter what happened first. Two events are when one happening does not change the probability of the other. For independent events, the gives the chance that both happen: P(A and B) = P(A) × P(B). So P(heads, then heads) = 1/2 × 1/2 = 1/4. You can check by listing the four equally likely outcomes, HH, HT, TH, TT: exactly one of the four is HH.

A basketball player makes 80% of her free throws. If her shots are independent, the chance she makes two in a row is 0.8 × 0.8 = 0.64. The chance she misses at least one of the two is everything else: 1 − 0.64 = 0.36. That subtraction uses the : P(not A) = 1 − P(A). The chance of three makes in a row is 0.8 × 0.8 × 0.8 = 0.512, so even a good shooter misses at least one of three shots almost half the time.

Drawing without replacement is the classic case where independence fails. Pull one card from a standard 52-card deck: P(ace) = 4/52. Keep it out and pull a second: if the first was an ace, only 3 aces remain among 51 cards, so P(second ace) = 3/51. The chance of two aces is 4/52 × 3/51 = 12/2652 = 1/221. A common mistake is to multiply 4/52 × 4/52 as if the deck had been reset; that overstates the chance. Before multiplying, ask whether the first event changes the odds for the second.

Words to know
independent
two events are independent when one happening does not change the probability of the other
multiplication rule
for independent events, P(A and B) = P(A) × P(B)
complement
the event that A does not happen; P(not A) = 1 − P(A)
without replacement
drawing items and not putting them back, so later draws have changed probabilities
Check yourself

1. A fair die is rolled twice. What is the probability of rolling a 6 both times?

2. Three independent parts each work with probability 0.9. What is the probability all three work?

3. Which pair of events is NOT independent?

59.7

Conditional Probability and Two-Way Tables

Main ideaA conditional probability restricts attention to one row or column of a table; if it matches the overall probability, the events are independent.

A survey of 200 students asks two yes-or-no questions: do you play a school sport, and do you have a part-time job? The results fit in a . Of the 120 who play a sport, 40 have a job and 80 do not. Of the 80 who do not play a sport, 30 have a job and 50 do not. Column totals: 70 have a job, 130 do not. The overall chance that a randomly chosen student has a job is 70/200 = 0.35.

A asks about a smaller group. Given that a student plays a sport, what is the chance they have a job? Now only the 120 sport players matter, and 40 of them have jobs: P(job | sport) = 40/120 = 1/3 ≈ 0.333. The vertical bar reads ’given.’ Flip the condition: P(sport | job) = 40/70 = 4/7 ≈ 0.571, because now the 70 job holders are the whole world and 40 of them play a sport. Notice these two are different numbers; mixing them up is the most common error in this topic.

Conditional probability gives a clean test for independence: two events are independent exactly when P(A | B) = P(A). Here P(job | sport) ≈ 0.333 but P(job) = 0.35, and P(job | not sport) = 30/80 = 0.375. The numbers are close but not equal, so job and sport are slightly linked: sport players are a little less likely to hold a job. When a table has raw counts, always divide by the total of the row or column you are conditioning on, never by the grand total.

Words to know
two-way table
a table that counts how many items fall into each combination of two categories
conditional probability
the probability of A given that B is known to have happened, written P(A | B)
given
the word that signals a condition: 'the chance of rain given that it is cloudy'
grand total
the total count of everything in a two-way table
Check yourself

1. In the student survey (120 play a sport, 40 of them have a job), what is P(job | sport)?

2. Same survey: 70 students have a job and 40 of those play a sport. What is P(sport | job)?

3. Two events A and B are independent exactly when which statement is true?

59.8

The Addition Rule

Main ideaTo find the chance that A or B happens, add their probabilities and subtract the overlap so nothing is counted twice.

Draw one card from a standard deck. What is the chance it is an ace or a heart? There are 4 aces and 13 hearts, but the ace of hearts is in both groups. Adding 4 + 13 = 17 counts that card twice. The fixes this: P(A or B) = P(A) + P(B) − P(A and B). So P(ace or heart) = 4/52 + 13/52 − 1/52 = 16/52 = 4/13. Check by counting: 4 aces plus the 12 hearts that are not aces gives 16 favorable cards out of 52.

Some events cannot happen together. Rolling a 1 and rolling a 2 on the same die are : the overlap is empty and P(A and B) = 0. Then the rule simplifies to P(A or B) = P(A) + P(B) = 1/6 + 1/6 = 2/6 = 1/3. Now try ’even or greater than 4.’ Even is {2, 4, 6}, probability 3/6. Greater than 4 is {5, 6}, probability 2/6. Both at once is {6}, probability 1/6. So P = 3/6 + 2/6 − 1/6 = 4/6 = 2/3. Listing the winning faces, {2, 4, 5, 6}, confirms 4 out of 6.

Two mistakes come up again and again. The first is adding probabilities for events that overlap without subtracting: a student says P(even or greater than 4) = 5/6, which counts the 6 twice. The second is confusing ’or’ with ’and’: ’and’ asks for both, which is the overlap and is found by multiplying (when independent) or by reading the table cell, never by adding. Word the question to yourself first: is it ’both,’ or is it ’at least one’?

Words to know
addition rule
P(A or B) = P(A) + P(B) − P(A and B); subtract the overlap so it is not counted twice
mutually exclusive
two events that cannot both happen at once, so their overlap has probability 0
overlap
the outcomes that belong to both events, written A and B
Check yourself

1. P(A) = 0.5, P(B) = 0.3 and P(A and B) = 0.1. What is P(A or B)?

2. One roll of a fair die: what is the probability of an even number or a number greater than 4?

3. What does it mean for two events to be mutually exclusive?

Section 4

Expected Value and Repeated Trials

59.9

Expected Value and Decisions

Main ideaExpected value is the long-run average result per trial: multiply each outcome by its probability and add.

A club sells 100 raffle tickets at $2 each, and one ticket wins a $100 prize. If you buy one ticket, what do you get back on average? You win $100 with probability 1/100 and $0 with probability 99/100. The of the prize is 100 × 1/100 + 0 × 99/100 = $1. You paid $2, so your expected net result is 1 − 2 = −$1 per ticket. That does not mean you lose exactly $1; you lose $2 or gain $98. It means that over many tickets your average result is a loss of $1 each, which is exactly how the club makes money.

The general recipe: list every outcome, multiply each by its probability, and add the products. A game pays the number of dollars showing on one roll of a fair die. The expected payout is 1 × 1/6 + 2 × 1/6 + 3 × 1/6 + 4 × 1/6 + 5 × 1/6 + 6 × 1/6 = 21/6 = $3.50. A fair price to play would be $3.50; at $4 the house wins 50 cents per game on average. A common error is to average the outcomes without their probabilities, which only works when all outcomes are equally likely, as they happen to be here.

Expected value guides decisions, but it is not the only consideration. A phone-repair plan costs $60 a year. Suppose a phone breaks with probability 0.1 in a year and the repair costs $300. Expected repair cost without the plan: 0.1 × 300 + 0.9 × 0 = $30. The plan costs $60, twice the expected loss, so on average it is a bad deal. But a person who could not possibly pay $300 at once might reasonably buy it anyway: expected value tells you the average, not how much a bad outcome would hurt. Report the expected value, then talk about the risk.

Words to know
expected value
the long-run average result per trial, found by adding each outcome times its probability
fair game
a game whose expected net result is zero for the player
risk
how far the actual outcomes can swing from the expected value, and how much a bad swing would hurt
net result
what you gain or lose after subtracting what you paid
Check yourself

1. A spinner pays $10 with probability 0.2 and costs you $2 with probability 0.8. What is the expected net result per spin?

2. A raffle sells 1,000 tickets at $5 each and gives one $2,000 prize. What is the expected net result of buying one ticket?

3. Option A pays a sure $50. Option B pays $120 with probability 0.5 and $0 otherwise. Which statement is correct?

59.10

The Binomial Setting

Main ideaWhen you repeat the same independent yes-or-no trial a fixed number of times, the count of successes follows a binomial distribution.

A free-throw shooter who makes half her shots takes 4 shots. How likely is exactly 2 makes? This is a : a fixed number of trials (4), each with two outcomes (make or miss), the same probability of success on each (0.5), and independent trials. Write all 16 equally likely sequences of M and X. Exactly 2 makes happens in 6 of them: MMXX, MXMX, MXXM, XMMX, XMXM, XXMM. So P(exactly 2) = 6/16 = 0.375.

Counting sequences by hand gets slow, so use the pattern. The number of ways to choose which k of the n trials are successes is written C(n, k). For n = 4: C(4, 0) = 1, C(4, 1) = 4, C(4, 2) = 6, C(4, 3) = 4, C(4, 4) = 1. Each specific sequence with k successes has probability p^k × (1 − p)^(n − k). Multiply: P(k successes) = C(n, k) × p^k × (1 − p)^(n − k). For an 80% shooter taking 3 shots, P(exactly 2 makes) = C(3, 2) × 0.8^2 × 0.2^1 = 3 × 0.64 × 0.2 = 0.384.

Check the whole distribution for that shooter: P(0) = 1 × 0.2^3 = 0.008; P(1) = 3 × 0.8 × 0.04 = 0.096; P(2) = 0.384; P(3) = 0.8^3 = 0.512. The four add to 1.000, which is a good sign that nothing was dropped. Two things break the binomial setting: a probability that changes from trial to trial (drawing cards without replacement), and a number of trials that is not fixed in advance (shooting until the first miss). Check the four conditions before using the formula.

Words to know
binomial setting
a fixed number of independent trials, each a success or failure with the same probability
trial
one repetition of the chance process, such as one shot or one flip
success
the outcome being counted, whatever it is; a 'success' can even be a defect
C(n, k)
the number of ways to choose which k of n trials are successes; C(4, 2) = 6
Check yourself

1. A fair coin is flipped 3 times. What is the probability of exactly 1 head?

2. A player makes 80% of free throws. In 3 independent shots, what is the probability of exactly 2 makes?

3. Which situation is a binomial setting?

59.11

Mean and Spread of a Binomial Count

Main ideaIn a binomial setting the expected count is n × p and the standard deviation is √(n × p × (1 − p)), which tells you which counts are unusual.

If a 70% free-throw shooter takes 100 shots, how many makes should we expect? Common sense says about 70, and the formula agrees: the of a binomial count is n × p = 100 × 0.7 = 70. The count varies from night to night, and its is √(n × p × (1 − p)) = √(100 × 0.7 × 0.3) = √21 ≈ 4.6 makes. For a large n the binomial pile is close to a normal curve, so the empirical rule applies: on about 95% of 100-shot nights she makes between 70 − 2 × 4.6 ≈ 61 and 70 + 2 × 4.6 ≈ 79.

That spread lets you judge a claim. A student guesses randomly on a 20-question multiple-choice test with 4 choices each, so p = 0.25. Mean = 20 × 0.25 = 5 correct. Standard deviation = √(20 × 0.25 × 0.75) = √3.75 ≈ 1.94. Suppose the student gets 12 right. The z-score is (12 − 5) ÷ 1.94 ≈ 3.6, far beyond 3 standard deviations. Pure guessing almost never produces that, so the honest conclusion is that the student knew something (or copied). The formula turned a hunch into a number.

Flip a coin 100 times and get 65 heads. Is the coin fair? For a fair coin, mean = 50 and standard deviation = √(100 × 0.5 × 0.5) = √25 = 5. Then 65 heads is (65 − 50) ÷ 5 = 3 standard deviations above the mean, a result the empirical rule puts at about 1 in 1,000 for the upper tail alone. That is strong evidence against fairness, and it is the first step toward the hypothesis tests of the next chapter. A common slip is to take √(n × p) and forget the (1 − p) factor: √50 ≈ 7.1 would wrongly make 65 heads look ordinary.

Words to know
binomial mean
the expected number of successes, n × p
binomial standard deviation
the spread of the success count, √(n × p × (1 − p))
unusual
a result more than about 2 standard deviations from what the model expects
Check yourself

1. A fair coin is flipped 400 times. What are the mean and standard deviation of the number of heads?

2. A process succeeds 20% of the time. In 50 independent tries, what is the expected number of successes?

3. A coin flipped 100 times gives 65 heads. Using mean 50 and standard deviation 5, what is the z-score, and what does it suggest?

Chapter review

Distributions and the Normal Curve

0 / 8

1. For the values 3, 5, 5, 7 (mean 5), what is the standard deviation when you divide by the count?

2. Heights are modeled as normal with mean 64 inches and standard deviation 2.5 inches. About what percent are taller than 69 inches?

3. Which z-score corresponds to a raw score of 58 when the mean is 70 and the standard deviation is 8?

4. A test score is at the 84th percentile of a normal distribution. About what is its z-score?

5. A two-way table shows 60 phones with a case, 10 of them cracked, and 40 without a case, 20 of them cracked. What is P(cracked | no case)?

6. P(A) = 0.4, P(B) = 0.5, and A and B are independent. What is P(A or B)?

7. A game pays $6 with probability 1/3 and $0 otherwise, and costs $3 to play. What is the expected net result?

8. A shooter makes 60% of free throws and takes 100 independent shots. What is the standard deviation of the number of makes?

Chapter

Sampling, Inference and Claims

Statistics
Big questionHow can a few hundred people tell us about millions, and when should we refuse to believe what a number claims?
The story

Two Million Ballots and the Wrong Answer

In 1936 a famous magazine collected more than two million responses and still predicted the wrong president. The problem was not the count.

In the fall of 1936 the Literary Digest, a popular American magazine, ran the biggest straw poll anyone had ever attempted. It mailed roughly ten million postcard ballots asking who the reader would vote for: President Franklin Roosevelt or the Republican challenger, Alf Landon. More than two million cards came back. The magazine had called the previous several elections correctly with the same method, and it announced its forecast with confidence: Landon would win comfortably, with a clear majority of the vote.

Roosevelt won in a landslide, carrying all but two states and roughly 61% of the popular vote. The magazine had missed by nearly 20 percentage points with a sample a thousand times larger than most polls use today. What went wrong? The Digest had built its mailing list from telephone directories, automobile registrations and its own subscribers. In the middle of the Great Depression, people who owned a phone or a car were, on average, wealthier than people who did not, and wealthier voters leaned toward Landon. The list itself was tilted before a single card was mailed.

There was a second problem. Only about one card in four came back. People who bother to return a mail-in ballot are not a random slice of the people who received one; they tend to be angrier, more engaged, or more eager for change. That year, the voters most motivated to respond were the ones who wanted Roosevelt out. Two tilts, one in who was asked and one in who answered, pointed the same way, and no amount of extra postcards could fix them. A bigger biased sample is just a bigger mistake.

Meanwhile a young researcher named George Gallup predicted Roosevelt's win using a sample of only tens of thousands of people, chosen to look like the whole electorate. His question was not 'how many people can we ask?' but 'who are we asking, and who are we missing?' That question is the heart of this chapter. You will learn how a small random sample can stand in for a huge population, how much it can be trusted, how to test a claim against data, and how to read a headline built on numbers without being fooled.

Talk about itSuppose the Literary Digest had mailed twenty million cards instead of ten million, using the same lists. Would its prediction have gotten better, worse, or stayed about the same? Explain.
Section 1

Collecting Data That Can Be Trusted

60.1

Populations, Samples and Statistics

Main ideaA population is the whole group you want to know about; a sample is the part you actually measure, and a statistic from it estimates a parameter of the population.

Chicago Public Schools serves more than 300,000 students. A student newspaper wants to know what fraction of them ride a CTA bus or train to school. Asking every student is a , which is slow and expensive. Instead the paper asks 500 students. The whole group of interest, every CPS student, is the . The 500 who are asked are the . If 310 of the 500 say yes, the sample proportion is 310 ÷ 500 = 0.62, or 62%.

That 62% is a : a number computed from the sample. The true fraction of all CPS students who ride the CTA, whatever it is, is a : a fixed number about the population that we usually never learn exactly. Statistics estimate parameters. A second sample of 500 might give 295 ÷ 500 = 0.59. Neither is wrong; each is an estimate, and the rest of the chapter is about how close such estimates tend to be. Keep the vocabulary straight: parameter goes with population, statistic goes with sample.

The same words apply to averages. If the mean commute time in the sample of 500 is 34 minutes, that sample mean is a statistic estimating the parameter, the mean commute of all 300,000 students. Two common confusions: calling the sample the population (the 500 students are not ’everyone’), and treating the statistic as if it were exact (writing ’exactly 62% of CPS students ride the CTA’). A careful sentence says ’about 62% of the students surveyed,’ and then asks how far off that might be.

Words to know
population
the entire group you want to draw conclusions about
sample
the part of the population you actually measure
census
an attempt to measure every member of the population, like the U.S. Census every ten years
parameter
a number that describes the whole population, usually unknown
statistic
a number computed from a sample, used to estimate a parameter
Check yourself

1. A newspaper surveys 500 of the more than 300,000 CPS students. Which is the population?

2. In the survey, 310 of 500 students ride the CTA. The value 62% is best called what?

3. Why do two random samples of 500 from the same population usually give slightly different percentages?

60.2

Surveys, Observational Studies and Experiments

Main ideaA survey asks, an observational study watches, and an experiment assigns treatments; only an experiment can show that one thing causes another.

Three ways to collect data answer three different kinds of question. A asks people to report something: who they will vote for, how many hours they sleep, whether they ride the CTA. An records what people already do and looks for patterns, such as comparing the grades of students who eat breakfast with the grades of those who skip it. An actively assigns a : the researcher decides which students get breakfast and which do not, then compares.

The difference matters because of . Suppose the observational study finds that breakfast eaters average 8 points higher. Did breakfast cause the gain? Maybe. But students who eat breakfast may also sleep more, have more organized mornings, or come from homes with more support, and any of those could be the real reason. A hidden variable tangled up with the one you are studying is a confounder, and observation alone cannot untangle it. The study shows an association, not a cause.

An experiment breaks the tangle by random assignment. Flip a coin for each of 200 volunteers: heads eats a provided breakfast for a month, tails does not. Chance, not personal habits, decides who is in which group, so sleep, family support and everything else are spread roughly evenly across both groups. If the breakfast group then scores higher, breakfast is the most reasonable explanation. Whenever you read that something ’causes’ or ’leads to’ an outcome, ask: did the researchers assign it, or just observe it?

Words to know
survey
a study that collects data by asking people questions
observational study
a study that records what happens without controlling who gets what
experiment
a study in which the researcher assigns treatments to subjects, ideally at random
treatment
the condition given to a group in an experiment, such as a new drug or a new lesson plan
confounding
when a hidden variable is mixed up with the one being studied, so you cannot tell which caused the result
Check yourself

1. Researchers compare test scores of students who chose to attend tutoring with scores of students who did not. What type of study is this?

2. What is the main advantage of randomly assigning subjects to treatment groups?

3. A study finds that people who own more books have higher incomes. Which is the most careful conclusion?

60.3

Random Samples and Sources of Bias

Main ideaA simple random sample gives every group of the same size the same chance of being chosen; anything that tilts who is asked or who answers is bias.

A is chosen so that every possible group of the chosen size has the same chance of being picked. To draw 60 students at random from a school of 1,200, number the students 1 to 1,200, and let a random number generator pick 60 different numbers. No student can lobby to be included or excluded, and no teacher’s opinion of who is ’typical’ enters in. That is the point: chance does the choosing, so the sample’s tilt, if any, is only the luck of the draw, which the next lessons show how to measure.

is a systematic tilt, a reason the sample tends to differ from the population in a predictable direction. The Literary Digest showed two kinds. means parts of the population never had a chance to be picked: no phone, no car, no subscription meant no ballot. bias arises when the people who answer differ from those who do not. A third kind, , is worse still: a website poll that anyone can click answers only the question ’who felt strongly enough to click?’ A survey of sports fans who volunteer to rate the team’s coach will not describe the average fan.

Convenience samples, such as asking the first 50 people at the cafeteria door, share the same flaw: the people at the door at 11:15 are not a random slice of the school. Wording matters too. ’Do you support the plan to cut wasteful spending?’ and ’Do you support the plan to cut school programs?’ can describe the same plan and get opposite answers. When you judge a sample, ask three questions: who was left out, who did not answer, and what exactly was asked?

Words to know
simple random sample
a sample chosen by chance so every group of the same size is equally likely to be picked
bias
a systematic tilt that makes a sample differ from the population in a predictable direction
undercoverage
when part of the population has no chance of being selected
nonresponse
when the people who refuse or fail to answer differ from those who do
voluntary response
a sample made of people who chose to participate, usually those with strong opinions
Check yourself

1. A TV station asks viewers to text in whether they support a new stadium, and 78% say yes. What is the main problem?

2. A poll of Illinois voters is drawn only from people with landline phones. What kind of bias is this?

3. Which method produces a simple random sample of 60 from a school of 1,200 students?

Section 2

How Much Do Samples Vary?

60.4

Why Samples Differ From Each Other

Main ideaEven perfect random samples give different results each time; the pattern of those results, the sampling distribution, is predictable.

Imagine a jar of 10,000 beads, exactly half red and half white, so the true proportion of red is p = 0.5. Scoop out 100 beads at random and count the red ones. You might get 47. Put them back, stir, scoop again: 53. Again: 49, 55, 44. None of these is wrong. This bead-to-bead wobble is , and it happens even when the sampling is done perfectly. The parameter never moved; only the sample did.

Now imagine doing 1,000 scoops of 100 and making a histogram of the 1,000 sample proportions. That histogram is the . For random samples it has a beautiful regularity: it is centered at the true value 0.5, it is roughly bell-shaped, and its spread depends on the sample size. Most scoops of 100 fall between 0.40 and 0.60. If instead each scoop held 400 beads, the histogram would be much narrower, with most scoops between 0.45 and 0.55.

Two lessons follow. First, a single sample’s proportion is a draw from this distribution, so it is ’off’ by a typical amount you can estimate. Second, bigger samples shrink that typical amount, but only for random samples. If the jar were stirred badly so that red beads sat on top, every scoop would run red, and scooping 400 instead of 100 would not help. Sampling variability is chance and can be reduced by size; bias is a tilt and cannot.

Words to know
sampling variability
the natural chance difference between one random sample's result and the next
sampling distribution
the pattern of values a statistic takes over many repeated random samples
sample size
the number of individuals in the sample, written n
Check yourself

1. A jar is exactly half red. A random scoop of 100 beads has 44 red. What is the best explanation?

2. Which change makes the sampling distribution of a sample proportion narrower?

3. A sampling distribution built from many random samples of the same size is centered where?

60.5

Simulating a Sampling Distribution

Main ideaSimulate many random samples with coins, dice or a computer to see how far a statistic typically strays; the spread shrinks like 1 divided by the square root of n.

You do not need a jar of beads. To sampling 100 people from a population where half say yes, flip a coin 100 times and count heads. Each run of 100 flips is one simulated sample. Do it 50 times and you have a rough sampling distribution. A computer can do 10,000 runs in a second, but the idea is the same: rebuild the chance process, run it many times, and look at what the results usually do.

The simulated histogram has a spread you can predict with the binomial formula from the last chapter. For a proportion, the is √(p × (1 − p) ÷ n). With p = 0.5 and n = 100: √(0.5 × 0.5 ÷ 100) = √0.0025 = 0.05. So about 95% of simulated sample proportions land within 2 × 0.05 = 0.10 of 0.5, that is, between 0.40 and 0.60, matching the bead scoops. With n = 400: √(0.25 ÷ 400) = √0.000625 = 0.025, so 95% land between 0.45 and 0.55.

Notice the pattern: quadrupling n from 100 to 400 cut the spread in half, from 0.05 to 0.025. The spread shrinks with the square root of n, not with n itself. To cut it in half again you would need 1,600, and to reach a spread of 0.01 you would need 2,500. Doubling the sample only makes it about 1.4 times more precise. That is why polls of about 1,000 people are so common: they give a standard error near 0.016, a margin of about 3 points, and going to 10,000 people costs ten times as much just to bring the margin down to about 1 point.

Words to know
simulate
imitate a chance process many times with coins, dice, cards or a computer to see its typical results
standard error
the standard deviation of a statistic across repeated samples; for a proportion it is √(p(1 − p) ÷ n)
square root rule
the spread of a sample statistic shrinks in proportion to 1 ÷ √n, so four times the sample halves the spread
Check yourself

1. For a population proportion of 0.5, what is the standard error of a sample proportion when n = 100?

2. A sample proportion has standard error 0.05 with n = 100. To make the standard error 0.025, what sample size is needed?

3. In a simulation of 100 fair-coin flips repeated many times, about what range holds 95% of the sample proportions of heads?

60.6

Margin of Error and Confidence Intervals

Main ideaA confidence interval reports a sample estimate plus and minus a margin of error, a range that captures the true value in about 95% of samples.

A poll of 1,000 Illinois adults finds 54% in favor of a new transit plan. The true percent for all adults is unknown, but the sampling distribution tells us how far a sample of 1,000 typically strays. Standard error: √(0.54 × 0.46 ÷ 1,000) = √0.0002484 ≈ 0.0158. Two standard errors is about 0.032, or 3.2 percentage points. That is the . The is 54% ± 3.2%, from 50.8% to 57.2%. A pollster would write: ’54%, with a margin of error of about 3 points.’

What does ’95% confidence’ mean? Not that there is a 95% chance the truth is in this one interval; the truth is a fixed number and is either in it or not. It means the method works 95% of the time: if you took many samples of 1,000 and built an interval from each, about 95 of every 100 intervals would capture the true value and about 5 would miss. You do not know whether yours is one of the 5, which is why a result at 50.8% to 57.2% supports ’more than half’ only cautiously.

A quick rule for a 95% margin of error when p is near 0.5 is 1 ÷ √n. For n = 1,000: 1 ÷ √1,000 ≈ 1 ÷ 31.6 ≈ 0.032, matching the careful calculation. For n = 400 it is 1 ÷ 20 = 0.05, or 5 points; for n = 2,500 it is 1 ÷ 50 = 0.02, or 2 points. Two cautions: the margin of error covers only random sampling variability, not bias, and a poll with 60% ± 5% does not mean ’60% is right and the rest is error.’ Every value from 55% to 65% is consistent with the data.

Words to know
margin of error
the plus-or-minus amount around a sample estimate, about 2 standard errors for 95% confidence
confidence interval
the range from estimate minus margin of error to estimate plus margin of error
95% confidence
the method produces an interval that captures the true value in about 95 out of 100 samples
percentage point
the difference between two percents; from 54% to 57% is 3 percentage points
Check yourself

1. A random sample of 400 people finds 60% support a proposal. Using the 1 ÷ √n rule, what is the approximate 95% confidence interval?

2. A pollster wants to cut a margin of error from 4 points to 2 points. Roughly how must the sample size change?

3. A 95% confidence interval for a proportion is 50.8% to 57.2%. Which interpretation is correct?

Section 3

Testing a Claim

60.7

The Logic of a Hypothesis Test

Main ideaA hypothesis test assumes the claim is true, asks how surprising the data would then be, and rejects the claim only if the data are very unlikely under it.

A friend hands you a coin and says it is fair. You flip it 100 times and get 60 heads. Fair coins usually give around 50, but 60 is not impossible. How do you decide? A works like a courtroom. The claim on trial, called the , is ’the coin is fair, p = 0.5.’ It gets the benefit of the doubt. The is what you would conclude if the evidence is strong enough: ’the coin favors heads.’ The data are the evidence, and the question is whether they are consistent with the null.

Assume the null is true and measure how unusual the result is. For a fair coin in 100 flips, the expected count is 50 with standard deviation √(100 × 0.5 × 0.5) = 5. Then 60 heads is (60 − 50) ÷ 5 = 2 standard deviations above expected. From the normal table, a fair coin produces 60 or more heads only about 2.3% of the time. That is rare. If we would see this in fewer than 5 of 100 tries, the usual convention says the result is , and we reject the null: the evidence points to a coin that favors heads.

The logic has two sides that students often flip. Rejecting the null does not prove the alternative; it says the data would be surprising if the null were true, so the null is doubtful. And failing to reject does not prove the null; 54 heads is consistent with a fair coin, but also with a coin that has p = 0.52. A test can be wrong in two ways: rejecting a true null (convicting an innocent coin, which happens about 5% of the time at the 5% cutoff), or failing to reject a false null (letting a guilty coin go). Small samples make the second error common.

Words to know
hypothesis test
a procedure that judges a claim by asking how unlikely the observed data would be if the claim were true
null hypothesis
the claim being tested, usually 'no effect' or 'no difference,' which gets the benefit of the doubt
alternative hypothesis
the conclusion you adopt if the data make the null hypothesis hard to believe
statistically significant
so unlikely under the null hypothesis (usually under 5%) that the null is rejected
Check yourself

1. A company claims its battery lasts 10 hours on average. In a test of the claim, what is the null hypothesis?

2. A fair coin flipped 100 times gives 60 heads, which is 2 standard deviations above 50. About how often does a fair coin do at least this well?

3. A test fails to reject the null hypothesis that a coin is fair. What can you conclude?

60.8

How Surprising Is Surprising? The P-Value

Main ideaThe p-value is the chance of a result at least as extreme as yours if the null hypothesis were true; a small p-value is evidence against the null.

The number that measures surprise is the : the probability, assuming the null hypothesis is true, of getting a result at least as extreme as the one observed. For the coin with 60 heads in 100 flips, the p-value is about 0.023, the fair-coin chance of 60 or more heads. A p-value is not the chance the null is true; it is the chance of data like yours if the null is true. Those are different sentences, and swapping them is the most common misreading of statistics in news stories.

You can find a p-value by instead of a table. Take the cafeteria claim that 80% of students like the menu, with a sample of 50 that found only 32. Program a computer to draw 50 ’students’ at random from a population where 80% say yes, count the yeses, and repeat 10,000 times. Then count how many of the 10,000 runs gave 32 or fewer. If about 60 runs did, the p-value is roughly 60 ÷ 10,000 = 0.006. The claim would produce a sample this low only about 6 times in 1,000, so we reject it.

The usual cutoff, called the , is 0.05: reject the null when p is below it. But the cutoff is a convention, not a law of nature. A p-value of 0.06 is barely different from 0.04, and a p-value of 0.30 means the data are ordinary under the null, not that the null is true. Try one: a candidate claims 50% support, a random sample of 200 shows 55%. Standard error √(0.5 × 0.5 ÷ 200) ≈ 0.0354, so 0.55 is about 1.41 standard errors above 0.5, and the upper-tail p-value is about 0.08. Above 0.05, so not significant, even though 55 sounds like more than 50.

Words to know
p-value
the probability, if the null hypothesis were true, of a result at least as extreme as the one observed
significance level
the cutoff, usually 0.05, below which a p-value leads to rejecting the null hypothesis
simulation
generating many fake samples from the null hypothesis to see how often results like yours occur
extreme
farther from what the null hypothesis expects than the observed result
Check yourself

1. Which sentence correctly describes a p-value of 0.02?

2. A candidate claims 50% support. A random sample of 200 shows 55%, about 1.41 standard errors above 0.5, with p-value about 0.08. At the 0.05 level, what is the conclusion?

3. In 10,000 simulated samples under the null hypothesis, 30 gave a result as extreme as the real one. What is the estimated p-value?

60.9

Statistical Versus Practical Significance

Main ideaWith a huge sample, even a tiny, unimportant difference becomes statistically significant; always ask how big the effect is, not just whether p is small.

A tutoring company tests a new app on 50,000 students, split at random into two groups. The app group averages 70.5 on a final exam and the control group averages 70.0. Because the samples are enormous, the standard error of the difference is tiny, about 0.2 points, so a 0.5-point gap is 2.5 standard errors and the p-value is below 0.05. The result is statistically significant. It is also nearly worthless: half a point on a 100-point exam changes nobody’s grade. The gap is real, but it is not important.

That is the difference between and . Statistical significance answers: could this gap be a fluke of sampling? With 50,000 students, almost no gap is a fluke. Practical significance answers: is the gap big enough to matter to a student, a doctor or a city? Only judgment, not a formula, answers that. A report should always give the , the actual difference or ratio, alongside the p-value, so a reader can see whether 0.5 points is worth an app subscription.

The mirror-image mistake happens with small samples. Suppose a new reading program is tried on 12 students and their scores rise 8 points on average, but with only 12 students the standard error is about 5 points, so the gain is 1.6 standard errors and the p-value is about 0.11, not significant. That does not mean the program failed; it means 12 students cannot tell an 8-point gain from noise. An honest summary says ’promising, but the study was too small to be sure.’ Sample size decides what a study can detect; it does not decide what matters.

Words to know
statistical significance
a result unlikely to be due to sampling variability alone, usually meaning p is below 0.05
practical significance
an effect large enough to matter in real life, judged by its size, not by its p-value
effect size
the actual size of a difference or relationship, such as 0.5 points or a 2% drop in risk
power
a study's ability to detect a real effect; larger samples give more power
Check yourself

1. A study of 50,000 students finds a 0.5-point exam gain with a p-value below 0.05. What is the best description?

2. Why can a very large sample make a tiny difference statistically significant?

3. A program tested on 12 students shows an 8-point gain with p = 0.11. What is the fairest conclusion?

Section 4

Reading Claims in the Wild

60.10

Headlines Built on Numbers

Main ideaBefore believing a data claim, ask who was sampled, how big the sample was, whether the effect is relative or absolute, and who paid for the study.

’Daily coffee linked to 20% lower risk!’ Before sharing that, read it like a statistician. First, linked to is observational language: coffee drinkers may differ in exercise, income or sleep, so this is an association until an experiment says otherwise. Second, 20% lower than what? If the risk falls from 5 in 1,000 to 4 in 1,000, that is a 20% drop but only 1 in 1,000 in terms. Relative percents make small changes sound enormous. Ask for the two actual rates every time.

Third, look for the sample. ’Nine out of ten dentists’ means little without knowing how many dentists were asked, how they were chosen, and who asked them. A survey of 20 dentists selected by a toothpaste company is not the same as a random sample of 2,000. Also look for the margin of error. If a poll says a candidate leads 48% to 46% with a margin of error of 3 points, the intervals overlap, and the honest headline is ’too close to call,’ not ’candidate takes the lead.’ A 2-point gap inside a 3-point margin is noise.

Fourth, ask who benefits. A study funded by the company whose product it praises is not automatically wrong, but it deserves a closer look, and so does a result that has never been repeated by anyone else. Finally, watch the axes on any graph: a bar chart whose vertical axis starts at 90 instead of 0 can make a 2% difference look like a cliff. None of these checks requires a formula. They require the habit of asking ’compared to what, measured how, and says who?’

Words to know
relative change
a change expressed as a percent of the starting value; 5 to 4 per 1,000 is a 20% relative drop
absolute change
the plain difference between two rates; 5 to 4 per 1,000 is a drop of 1 per 1,000
association
two things tending to occur together, which does not by itself show one causes the other
replication
repeating a study with new data to see whether the result holds up
Check yourself

1. A risk falls from 8 in 1,000 to 6 in 1,000. What are the relative and absolute changes?

2. A poll shows candidate A at 48% and candidate B at 46% with a margin of error of 3 points. Which headline is most accurate?

3. A bar chart shows sales of 96 and 98 units, but the vertical axis starts at 95 so the second bar looks twice as tall. What is the problem?

60.11

Experiments That Earn a Cause

Main ideaRandom assignment, a control group, blinding and replication are what let an experiment say 'this caused that.'

To show that a treatment causes an outcome, an experiment needs four ingredients. First, a that does not get the treatment, so there is something to compare against. Second, of subjects to treatment or control, so the groups are alike except for the treatment. Third, when people’s expectations could affect the result, a , a fake treatment that looks real, and , so subjects (and ideally the people measuring them) do not know who got which. Fourth, enough subjects that a real effect is not lost in sampling variability.

The 1954 field trial of the Salk polio vaccine is the classic example. Hundreds of thousands of American children were randomly assigned to receive the vaccine or a placebo injection, and neither the children nor the doctors diagnosing polio knew which was which. Polio cases were much rarer in the vaccinated group, and because assignment was random and diagnosis was blind, that difference could not be explained by which families volunteered or by doctors expecting vaccinated children to be healthier. The vaccine was declared effective and released to the public.

Compare a weaker design: give the vaccine to children whose parents ask for it and compare them with everyone else. Parents who ask for a vaccine tend to differ in income, education and health habits, so any difference in polio rates would be tangled with those. The lesson for reading any claim of cause: find the control group, find the word random, and ask whether anyone could have known who got what. When an experiment is impossible or unethical, as with smoking, researchers rely on many observational studies that point the same way, plus a plausible mechanism, and they say so.

Words to know
control group
the subjects who do not receive the treatment, used for comparison
random assignment
using chance to decide which subjects get which treatment, balancing hidden differences
placebo
a fake treatment that looks like the real one, used so expectations affect both groups equally
blinding
keeping subjects, and often researchers, from knowing who received which treatment
Check yourself

1. What is the purpose of a placebo in a medical experiment?

2. In the 1954 polio vaccine trial, why was random assignment essential?

3. A company gives its new app to employees who volunteer and compares their productivity with everyone else's. What is the flaw?

Chapter review

Sampling, Inference and Claims

0 / 8

1. A city surveys 800 randomly chosen residents and finds a mean commute of 34 minutes. The number 34 minutes is best described as a

2. The 1936 Literary Digest poll used lists of phone and car owners. Which type of bias does this illustrate?

3. For a proportion near 0.5, the standard error with n = 400 is 0.025. What is it with n = 1,600?

4. A random sample of 2,500 voters gives 52% for a measure. Using 1 ÷ √n, what is the approximate 95% confidence interval?

5. A coin flipped 100 times shows 58 heads. With mean 50 and standard deviation 5, what is the z-score, and is it significant at the 5% level (one-sided cutoff about 1.65)?

6. Which statement about a p-value of 0.40 is correct?

7. A study of 1,000,000 shoppers finds that a new checkout layout saves an average of 1.2 seconds per visit, p below 0.001. Which is the best description?

8. Which study design could justify the claim 'the new fertilizer causes taller corn'?

Unit wrap-up

Statistics: From a Sample to a Claim

Twelve words, twelve meanings

0 / 12

Tap a word, then tap its meaning. A right pair locks in green.

Words
Meanings
Unit test

Fifteen questions across the unit

0 / 15

1. The values 6, 7, 8, 9, 10 have mean 8. What is their standard deviation when you divide the squared deviations by the count?

2. Scores are normal with mean 500 and standard deviation 100. About what percent of scores are between 400 and 600?

3. With mean 500 and standard deviation 100, what is the z-score of a score of 350?

4. On a normal curve, a z-score of 2 corresponds to about which percentile?

5. Which data set would a normal model describe poorly, and why?

6. A fair coin is flipped 3 times. What is the probability of getting heads all 3 times?

7. A table shows 120 students play a sport; 40 of those have a job. Overall 70 of 200 students have a job. Which comparison tests whether sport and job are independent?

8. P(A) = 0.6, P(B) = 0.3 and P(A and B) = 0.2. What is P(A or B)?

9. A game pays $12 with probability 1/4 and nothing otherwise, and costs $4 to play. What is the expected net result per game?

10. A 25% free-throw shooter takes 48 independent shots. What are the mean and standard deviation of the number of makes?

11. A magazine polls its own subscribers about a statewide election. What is the main weakness?

12. For a proportion near 0.5, which sample size gives a 95% margin of error of about 2 percentage points using 1 ÷ √n?

13. A poll of 1,000 people gives 52% support with a margin of error of 3 points. What can be concluded?

14. A company claims a defect rate of 2%. A random sample of 1,000 items finds 35 defects. Under the claim, the expected count is 20 with standard deviation about 4.4. What is the z-score, and what does it suggest?

15. A headline says a treatment cut risk by 50%. The rates were 2 in 10,000 and 1 in 10,000. What is the absolute change?

Spiral review

Five questions from earlier units

0 / 5

1. (Unit 25) Which rule is recursive?

2. (Unit 24) A substance loses 20% of its mass each year. Its half-life, to the nearest tenth, is

3. (Unit 23) What are the excluded values of (x − 1)/(x^2 − 9)?

4. (Unit 25) A loan follows B_n = 1.01 × B_(n−1) − 200 with B_0 = 5,000. What is B_1?

5. (Unit 24) The solutions of cos x = −√3/2 on [0, 2π) are

Write it

A random sample of 400 students at a large high school finds 232 who say they get less than 7 hours of sleep. Estimate the fraction of all students who sleep less than 7 hours, give a margin of error and interval, and decide whether the school can claim that more than half its students are short on sleep. Then name one bias the margin of error would not catch.

  • State the sample proportion first: 232 ÷ 400, as a decimal and a percent.
  • Compute the margin of error with 1 ÷ √n and show the arithmetic, then write the interval.
  • Say what 95% confidence means using the idea of many repeated samples, not a chance about this one interval.
  • Compare the whole interval with 50% before deciding whether 'more than half' is supported.
  • Name a specific bias (who was left out, who did not answer, or how the question was worded) and say why a bigger sample would not fix it.
0 wordsSaved on this device as you type.

Practice rooms

Rooms already on the site that belong to this unit — cards, quizzes, a lab.

For the teacher

Every lesson keeps its own three checks; a lesson is ticked when all three are right. Chapter reviews, the unit test and its spiral review (five questions from earlier units in this band) score on the page. When the site is connected to your sheet, or the link carries ?dest=, each one also has a Send box: the first-try score, the standards, the supports used, the attempt number and the minutes go to your sheet as an IEP data point.

Print this page for a paper copy of the readings, the sources, the words and the questions; the answers print as dashed boxes under each question.

Fact-check notes for this course live in the handoff: quotes marked (paraphrased) were set that way on purpose.