The Interior — MathGrades 9–10

Unit 22 · Statistics in Algebra I

A unit of the course: the story, then chapter by chapter — sections, numbered lessons, a source or the numbers to read, three checks each — a review per chapter, and the wrap-up at the end.

← Math, the whole course

Drawn scene: a baseball diamond under stadium lights at night with a scoreboard and a scatter of fans in the stands that drifts upward like a trend
22Unit

Statistics in Algebra I

Statistics

Every day you meet numbers that claim to summarize something bigger: an average test score, a poll result, a headline that says one thing is linked to another. This unit is about looking past the single number to the whole set of data behind it. You will draw dot plots, histograms and box plots, measure where data are centered and how far they spread, and decide what to do when one value sits far from the rest.

Then you will put two variables on one graph and ask whether they move together. You will draw a line through the cloud of points, read its slope in plain words, check it with residuals, and measure the strength of the pattern with the correlation coefficient. You will also learn when a curve fits better than a line and how far a prediction can safely reach.

The unit ends with the most important warning in statistics: two things rising together does not mean one causes the other. By the end you will be able to describe a data set honestly, compare two groups fairly, fit and check a model, and read a claim about data with the right mix of trust and suspicion.

How we figured it out
1662

John Graunt tallies London death records into tables, an early step toward statistics as a way of reading data

1790

The first United States Census counts the population, a data set that repeats every ten years to this day

1805

Adrien-Marie Legendre publishes the method of least squares for fitting a line to observations

1809

Carl Friedrich Gauss publishes his own account of least squares and the bell-shaped error curve

1858

Florence Nightingale uses graphs of hospital deaths to argue for better sanitation in army hospitals

1886

Francis Galton describes regression toward the mean while studying the heights of parents and children

1890s

Karl Pearson develops the correlation coefficient r into the form still used today

1925

Ronald Fisher's Statistical Methods for Research Workers spreads modern data analysis to scientists

1935

Fisher's The Design of Experiments explains random assignment, the tool that separates cause from correlation

1973

Francis Anscombe publishes four data sets with the same line and r but wildly different scatter plots

1977

John Tukey's Exploratory Data Analysis popularizes the box plot and the habit of graphing data first

Chapter

Describing and Comparing Data

Statistics
Big questionHow can two data sets share the same average and still tell completely different stories?
The story

Two Pitchers, One Average

A coach on the Northwest Side of Chicago has two pitchers with the same average, and she still has to pick one for the playoff game.

Coach Alvarez runs the varsity softball team at a public high school on the Northwest Side of Chicago. In late May the team qualifies for the playoffs, and she has to choose a starting pitcher for the first game. Two pitchers have each started eight games this season. Her assistant prints one number for each of them: runs allowed per start. Dana averages 3.0 runs. Marcus averages 3.0 runs. The assistant shrugs and says it is a coin flip.

Coach Alvarez does not trust one number, so she asks for the game-by-game list. Dana's eight starts read 3, 3, 2, 4, 3, 3, 4, 2. Marcus's eight starts read 0, 1, 0, 7, 0, 8, 1, 7. Both lists add up to 24 runs, so both averages are 24 ÷ 8 = 3. But the lists do not look alike at all. Dana is never brilliant and never terrible. Marcus is either untouchable or he gets hammered.

She draws each list as a row of dots on a number line. Dana's dots pile up between 2 and 4, a tight little hill. Marcus's dots sit in two clumps, one near 0 and one near 7 or 8, with nothing in between. The middle value, the median, is 3 for Dana and only 1 for Marcus. The distance from lowest to highest is 2 for Dana and 8 for Marcus. Same average, completely different seasons.

Now the choice depends on the question. In a game where the team's hitters are hot and can score plenty, Dana's steady 3 runs is probably enough to win. In a game against a great opposing pitcher, where the team might score only 1 run, Dana almost surely loses, while Marcus has a real chance to throw a shutout. The average could not tell Coach Alvarez any of this. The shape and the spread of the data could.

This chapter is about seeing a whole data set, not just its average. You will draw dot plots, histograms and box plots, measure center and spread, decide what to do about a strange value, compare two groups honestly, and read tables that sort people by two questions at once. Every tool answers the coach's question: what does the whole list say, not just its middle?

Talk about itIf you were Coach Alvarez, which pitcher would you start against a team with a great pitcher of its own, and what number from the data supports your choice?
Section 1

Pictures of a Data Set

51.1

Dot Plots and Shape

Main ideaA dot plot stacks one dot per value on a number line, so the shape of the whole data set is visible at a glance.

A teacher asks 12 students how many siblings they have and writes down the answers: 0, 1, 1, 1, 2, 2, 2, 2, 3, 3, 4, 6. A list like this is hard to read. A fixes that. Draw a number line from 0 to 6. For each answer, put one dot above that number, stacking dots when a number repeats. You get 1 dot above 0, 3 dots above 1, 4 dots above 2, 2 dots above 3, 1 dot above 4, none above 5, and 1 dot above 6. The 12 dots make a picture of the whole , which is the full set of values and how often each one occurs.

Once the dots are stacked, describe the shape. Ask three questions. Where is the tallest stack? Here it is at 2, so 2 siblings is the most common answer. Is the picture balanced? No. Most dots crowd the left side, and a thin tail stretches to the right toward the lone 6. A distribution shaped like this is , because the long tail points right. If the tail pointed left, it would be skewed left. A distribution with two matching halves, like a hill with the same slope on each side, is .

Students often mix up which way a skew points. The skew is named for the tail, not the peak. Picture a slide: the pile of dots is the platform, and the tail is the slide going down. In the sibling data, the slide goes to the right, so it is skewed right. Another common slip is forgetting a value that appears zero times. There are no students with 5 siblings, so 5 gets no dot, but 5 still stays on the number line. Leaving it out would squeeze the 6 next to the 4 and hide the gap.

Words to know
dot plot
a graph that puts one dot above a number line for each value in a data set, stacking repeats
distribution
the whole set of values in a data set together with how often each value occurs
skewed right
a shape with most values on the left and a long thin tail stretching to the right
symmetric
a shape whose left and right halves are mirror images of each other
Check yourself

1. In the sibling dot plot, how many dots sit above the number 2?

2. A dot plot has most of its dots on the left and a long thin tail reaching to the right. What is its shape?

3. A dot plot of 20 quiz scores shows 5 dots above the score 8. What fraction of the scores are 8?

51.2

Histograms and Bins

Main ideaA histogram groups numbers into equal-width bins and draws a bar for each bin, so it works for large data sets and messy decimals.

A dot plot works for 12 answers. It falls apart for 300 commute times measured to the minute, because almost every value is different and the dots never stack. A solves this by grouping. Suppose 30 students report their morning commute in minutes. Split the number line into equal-width : 0 to 9 minutes, 10 to 19, 20 to 29, 30 to 39, and 40 to 49. Count how many commutes land in each bin. Say the counts are 4, 9, 10, 5 and 2. Those five counts add to 30, which is a quick check that no student was lost or counted twice.

Now draw. The horizontal axis is minutes, split at 0, 10, 20, 30, 40 and 50. Above each bin draw a bar whose height is the , the count for that bin. The bars touch each other, with no gaps, because the bins cover the whole number line without skipping any minutes. That is the visual difference between a histogram and a bar graph of categories like favorite pizza topping, where the bars are separated because the categories are not numbers on a line.

Reading a histogram means adding bins. How many students commute less than 20 minutes? Add the first two bins: 4 + 9 = 13 students. What share commute 30 minutes or more? Add the last two bins: 5 + 2 = 7, and 7/30 is about 0.233, or 23%. The most common bin is 20 to 29 minutes with 10 students, and the shape has a slight tail to the right, since the long commutes trail off.

The width of the bins changes the picture. Bins of width 5 would show more detail but might look jagged. Bins of width 25 would smooth everything into two fat bars and hide the peak. A frequent mistake is making bins that overlap, like 0 to 10 and 10 to 20. A commute of exactly 10 minutes would then belong to both bins. Use bins like 0 to 9 and 10 to 19, or agree on a rule such as ’the left edge is included and the right edge is not.’

Words to know
histogram
a graph that splits a number line into equal-width bins and draws a touching bar for the count in each bin
bin
one interval of values on a histogram, such as 10 to 19 minutes
frequency
how many data values fall in a bin or at a value; the height of a histogram bar
Check yourself

1. Using the commute histogram, how many students commute less than 20 minutes?

2. About what percent of the 30 students commute 30 minutes or more?

3. Why do the bars of a histogram touch each other?

51.3

Box Plots and Quartiles

Main ideaA box plot draws five numbers, the minimum, first quartile, median, third quartile and maximum, and shows where the middle half of the data lives.

Eleven students take a test and score 62, 68, 70, 74, 75, 78, 80, 84, 88, 90 and 95, already in order. The is the middle value. With 11 scores, the middle is the 6th one, which is 78. Five scores sit below it and five above. Now split each half again. The lower half is 62, 68, 70, 74, 75, and its middle is 70. That is the , written Q1. The upper half is 80, 84, 88, 90, 95, and its middle is 88. That is the third quartile, Q3. Together with the minimum 62 and the maximum 95, these make the : 62, 70, 78, 88, 95.

A draws exactly those five numbers above a number line. Draw a box from Q1 = 70 to Q3 = 88. Draw a line inside the box at the median, 78. From each end of the box draw a whisker, a thin line out to the minimum 62 on the left and the maximum 95 on the right. Roughly one quarter of the scores lie in each of the four pieces: below the box, in the left part of the box, in the right part, and beyond the box. The box itself holds the middle half of the class.

When the count is even, the median is the average of the two middle values. Take 3, 5, 6, 8, 10, 12, 15, 20. The two middle values are 8 and 10, so the median is (8 + 10) ÷ 2 = 9. The lower half is 3, 5, 6, 8 and Q1 is (5 + 6) ÷ 2 = 5.5. The upper half is 10, 12, 15, 20 and Q3 is (12 + 15) ÷ 2 = 13.5. A common mistake is finding the quartiles from the unsorted list. Always sort first. Another is including the median itself in both halves when the count is odd; in this course, leave the middle value out of both halves.

Words to know
median
the middle value of a sorted data set, or the average of the two middle values
first quartile
the median of the lower half of a sorted data set, written Q1
five-number summary
the minimum, first quartile, median, third quartile and maximum of a data set
box plot
a graph that draws a box from Q1 to Q3 with a line at the median and whiskers to the minimum and maximum
Check yourself

1. For the test scores 62, 68, 70, 74, 75, 78, 80, 84, 88, 90, 95, what is the first quartile Q1?

2. About what fraction of a data set lies inside the box of a box plot?

3. What is the median of the data set 3, 5, 6, 8, 10, 12, 15, 20?

Section 2

Measuring Center and Spread

51.4

Mean Versus Median

Main ideaThe mean shares the total equally and the median finds the middle; one very large value pulls the mean but barely moves the median.

Five people work at a small café in Pilsen. Four of them earn $12, $13, $13 and $14 an hour. The manager earns $48 an hour. The is the total shared equally: add the wages, 12 + 13 + 13 + 14 + 48 = 100, and divide by the 5 workers to get $20 an hour. The median is the middle wage when the list is sorted: 12, 13, 13, 14, 48, so the median is $13. If a job ad said ’our workers average $20 an hour,’ it would be telling the truth and still misleading four of the five people who applied.

The mean is sensitive to every value. Raise the manager’s wage from $48 to $98 and the total becomes 150, so the mean jumps to $30. The median stays $13, because the middle of the sorted list has not moved. A measure that barely changes when one extreme value changes is called . The median is resistant. The mean is not. That is why news reports about house prices or incomes usually give the median: a few mansions or billionaires would drag the mean far above what a typical family sees.

Neither measure is wrong. They answer different questions. The mean answers this: if the total were shared equally, how much would each get? That matters for a budget. The café’s payroll is 5 × 20 = $100 an hour no matter how the wages are split. The median answers a different question: what does a typical member of this group look like? For a symmetric distribution the two are close together. For a skewed distribution they pull apart. The mean gets dragged toward the tail. Here is a quick check on any data set. If the mean is far above the median, expect a right tail of large values.

Words to know
mean
the sum of all values divided by how many values there are; the average
resistant
describes a measure that changes very little when one extreme value in the data changes
center
a single number, such as the mean or median, that stands for a typical value in a data set
Check yourself

1. What is the mean of 4, 7, 9, 12 and 18?

2. What is the median of 15, 3, 8, 21, 10 and 6?

3. A neighborhood has 40 homes worth about $200,000 and one worth $5,000,000. Which statement is true?

51.5

Range and Interquartile Range

Main ideaThe range measures the full stretch of the data and the interquartile range measures the stretch of the middle half, which extreme values cannot touch.

One week in June, a student records the daily high temperature outside her apartment in Chicago: 71, 75, 68, 80, 74, 73 and 90 degrees Fahrenheit. Sort them: 68, 71, 73, 74, 75, 80, 90. The is the biggest value minus the smallest: 90 − 68 = 22 degrees. The range is quick and easy, but it depends entirely on the two most extreme days. That single 90-degree Saturday sets the whole range by itself.

The , or IQR, ignores the extremes. Find the median first. With 7 values the middle is the 4th, 74. The lower half is 68, 71, 73, so Q1 = 71. The upper half is 75, 80, 90, so Q3 = 80. The IQR is Q3 − Q1 = 80 − 71 = 9 degrees. That is the width of the box in a box plot, the stretch of the middle half of the week. Notice that the 90 does not appear anywhere in that subtraction.

Now suppose the Saturday reading was a typo and it should have been 120, a number that would break Chicago’s all-time record. The range becomes 120 − 68 = 52, more than doubling. The IQR is still 80 − 71 = 9, because Q3 is still the middle of the upper half, and 120 is still just the biggest value in that half. Like the median, the IQR is resistant. Range and IQR are both measures of : how far apart the values are. The range is honest about the extremes; the IQR is honest about the crowd in the middle.

Words to know
range
the maximum value minus the minimum value of a data set
interquartile range
the third quartile minus the first quartile; the width of the middle half of the data, written IQR
spread
how far apart the values in a data set are from each other; also called variability
Check yourself

1. What is the range of 5, 9, 14, 22 and 30?

2. A data set has Q1 = 15 and Q3 = 27. What is its interquartile range?

3. If the largest value in a data set of 20 numbers is increased by 100, which measure stays exactly the same?

51.6

Standard Deviation, Introduced

Main ideaStandard deviation measures the typical distance of the values from the mean, using every value in the data set.

The IQR uses only two numbers from the data. The uses all of them. Start with five quiz scores: 6, 7, 8, 9, 10. The mean is 40 ÷ 5 = 8. Now find each score’s , its distance from the mean with a sign: 6 − 8 = −2, 7 − 8 = −1, 8 − 8 = 0, 9 − 8 = 1, and 10 − 8 = 2. The deviations always add to zero, −2 − 1 + 0 + 1 + 2 = 0, so averaging them tells you nothing. That is the reason for the next step.

Square each deviation to erase the signs: 4, 1, 0, 1, 4. Add them: 4 + 1 + 0 + 1 + 4 = 10. Divide by the number of values, 5, to get 2. That number is the . Its units are squared points, which is awkward, so take the square root: √2 ≈ 1.41. The standard deviation of these scores is about 1.41 points. Read it this way: a typical quiz score in this set sits about 1.4 points from the mean of 8.

Compare a second set with the same mean: 8, 8, 8, 8, 8. Every deviation is 0, so the variance is 0 and the standard deviation is 0. A third set, 4, 6, 8, 10, 12, has deviations −4, −2, 0, 2, 4, squares 16, 4, 0, 4, 16 that add to 40, variance 40 ÷ 5 = 8, and standard deviation √8 ≈ 2.83. Same mean, twice the spread, and the standard deviation doubles from 1.41 to 2.83. Larger standard deviation means the values wander farther from the mean.

Two warnings. First, many calculators offer two versions, one dividing by the number of values and one dividing by one less than that. The second version gives a slightly bigger answer, √(10 ÷ 4) ≈ 1.58 for the quiz scores, and is used when the data are a sample from a bigger group. In this course, divide by the number of values unless you are told otherwise. Second, do not skip the square root. A variance of 2 and a standard deviation of 1.41 are different numbers, and only the second is in the same units as the data.

Words to know
standard deviation
the square root of the variance; the typical distance of the values from the mean
deviation
a value minus the mean; negative for values below the mean and positive for values above it
variance
the average of the squared deviations from the mean
Check yourself

1. The data set 2, 4, 6 has mean 4. What is the sum of its squared deviations?

2. Which data set has a standard deviation of 0?

3. The data set 1, 5, 5, 9 has mean 5. Dividing by the number of values, what is its standard deviation?

Section 3

Outliers and Comparisons

51.7

Outliers and Their Pull

Main ideaAn outlier is a value far outside the pattern of the rest, and it pulls the mean and range much harder than the median and IQR.

Eight friends count the text messages they sent in one hour: 12, 14, 15, 15, 16, 17, 18 and 40. Seven of the numbers are between 12 and 18. The 40 is far from the rest. A value like that is an . Sometimes it is a mistake, like a typo. Sometimes it is real: one friend was planning a party. Either way, you need a rule for deciding when a value is far enough away to count, and you need to know what it does to your summaries.

The standard rule uses the IQR. The median of the eight values is (15 + 16) ÷ 2 = 15.5. The lower half, 12, 14, 15, 15, has median (14 + 15) ÷ 2 = 14.5, so Q1 = 14.5. The upper half, 16, 17, 18, 40, has median (17 + 18) ÷ 2 = 17.5, so Q3 = 17.5. The IQR is 17.5 − 14.5 = 3. Multiply it by 1.5 to get 4.5. Any value more than 4.5 below Q1 or more than 4.5 above Q3 is an outlier. The are 14.5 − 4.5 = 10 and 17.5 + 4.5 = 22. The 40 is above 22, so it is an outlier. The 12 is not below 10, so it is fine.

Now see the pull. With the 40, the total is 147 and the mean is 147 ÷ 8 ≈ 18.4. Without it, the total is 107 and the mean is 107 ÷ 7 ≈ 15.3. One value moved the mean by about 3 messages. The median moves from 15.5 to 15, half a message. The range collapses from 40 − 12 = 28 to 18 − 12 = 6. Without the 40, Q1 is 14 and Q3 is 17, so the IQR is still 3, unchanged. This is the whole reason the median and IQR are called resistant. The mean and range are not.

Do not just delete an outlier because it is inconvenient. First ask whether it is an error. A height of 5.9 meters is a typo for 1.59 meters; fix it. A student who really did send 40 texts belongs in the data, and the honest move is to report the summary both ways and say why. On a box plot, outliers are drawn as separate dots beyond the whiskers, and the whisker stops at the last value inside the fence, here at 18.

Words to know
outlier
a data value far outside the pattern of the rest; by the usual rule, more than 1.5 IQRs beyond a quartile
fences
the cutoffs Q1 − 1.5 × IQR and Q3 + 1.5 × IQR; values past a fence are outliers
pull
the effect an extreme value has on a summary such as the mean or range
Check yourself

1. A data set has Q1 = 20 and Q3 = 30. What is the upper fence for outliers?

2. In the sorted data 3, 4, 4, 5, 5, 6, 20, which value is an outlier by the 1.5 × IQR rule?

3. When an outlier is removed from a data set, which summary usually changes the most?

51.8

Comparing Two Distributions

Main ideaTo compare two data sets fairly, compare their shape, their center and their spread, not just one number.

Back to Coach Alvarez. Dana’s runs allowed in eight starts: 3, 3, 2, 4, 3, 3, 4, 2. Marcus’s: 0, 1, 0, 7, 0, 8, 1, 7. Both totals are 24, so both means are 3. A one-number comparison ends there in a tie. A three-part comparison, shape, center and spread, does not. Sort Dana: 2, 2, 3, 3, 3, 3, 4, 4. Sort Marcus: 0, 0, 0, 1, 1, 7, 7, 8. Dana’s dot plot is a single symmetric hill. Marcus’s has two clumps with a gap in the middle, one at 0 and 1 and one at 7 and 8.

Center: Dana’s median is (3 + 3) ÷ 2 = 3. Marcus’s median is (1 + 1) ÷ 2 = 1. Spread: Dana’s range is 4 − 2 = 2 and Marcus’s is 8 − 0 = 8. For the IQR, Dana’s Q1 is the middle of 2, 2, 3, 3, which is 2.5, and Q3 is the middle of 3, 3, 4, 4, which is 3.5, so the IQR is 1. Marcus’s Q1 is the middle of 0, 0, 0, 1, which is 0, and Q3 is the middle of 1, 7, 7, 8, which is 7, so the IQR is 7. Dana is far more : her spread is a fraction of his by every measure.

Write the comparison in sentences, with the numbers. ’Both pitchers allow a mean of 3 runs per start, but their distributions differ. Dana’s runs are symmetric and tightly clustered, with an IQR of 1. Marcus’s are split into two clumps, with a median of only 1 but an IQR of 7.’ When two groups are drawn as box plots on the same number line, look for : if one box sits entirely to the right of the other, the groups differ clearly. If the boxes overlap heavily, the difference between the groups is smaller than the inside each group.

One trap is comparing groups of different sizes by raw counts. Suppose 40 students in one class and 20 in another take a survey. Then 10 ’yes’ answers mean 25% in the first class and 50% in the second. When the groups are not the same size, compare percents, medians or means, never raw counts. And always use the same scale on both plots. A box plot drawn from 0 to 10 next to one drawn from 0 to 100 makes the second group look more crowded than it is.

Words to know
consistent
describes data with small spread; the values stay close to one another
overlap
the part of the number line where two distributions both have values, visible when two box plots are stacked
variability
another word for spread; how much the values in a group differ from each other
Check yourself

1. What is the median of Marcus's runs allowed: 0, 1, 0, 7, 0, 8, 1, 7?

2. Two classes have the same median test score of 78. Class A has an IQR of 6 and class B has an IQR of 20. Which statement is true?

3. A survey gets 12 'yes' answers from a class of 24 and 12 'yes' answers from a class of 48. Which comparison is fair?

51.9

Choosing the Right Summary

Main ideaReport the mean and standard deviation for symmetric data with no outliers; report the median and IQR when the data are skewed or have outliers.

Every data set can be summarized two ways. The mean and standard deviation form one pair: both use every value, and both get pulled by extremes. The median and IQR form the other pair: both depend on position in the sorted list, and both resist extremes. Choosing between the pairs is a judgment about the shape of the data, so look at a dot plot, histogram or box plot before you pick.

Take the café wages: 12, 13, 13, 14, 48. The distribution is strongly skewed right, with the manager’s 48 far out in the tail. The mean of 20 describes nobody. The median of 13 and the IQR describe the four ordinary workers well. For skewed data or data with outliers, report the median and IQR. Now take Dana’s runs: 2, 2, 3, 3, 3, 3, 4, 4. Symmetric, no outliers, a tidy hill. Here the mean of 3 and the standard deviation tell the story completely, and they use all eight games instead of just positions in a list. For symmetric data without outliers, report the mean and standard deviation.

Compute Dana’s standard deviation to see the pair in action. The mean is 3. Deviations are −1, −1, 0, 0, 0, 0, 1, 1. Squares: 1, 1, 0, 0, 0, 0, 1, 1, summing to 4. Divide by 8 to get 0.5, and take the square root: √0.5 ≈ 0.71. A typical Dana start is about 0.7 runs from her mean. For Marcus, deviations from 3 are −3, −3, −3, −2, −2, 4, 4, 5, squares 9, 9, 9, 4, 4, 16, 16, 25 summing to 92, variance 92 ÷ 8 = 11.5, standard deviation √11.5 ≈ 3.39. About five times Dana’s. But notice Marcus’s two-clump shape: no single center describes him well, so any summary should come with the picture.

The most common mistake is reporting a mean because it is familiar, without checking the shape. The second is reporting a spread measure from the wrong pair, like a median with a standard deviation. Keep the pairs together: mean with standard deviation, median with IQR. And when in doubt, report both pairs and show the plot. A reader can always ignore an extra number; a reader cannot recover the one you left out.

Words to know
summary statistic
one number, such as a mean or an IQR, that describes some feature of a whole data set
shape
the overall pattern of a distribution: symmetric, skewed, one peak or two, with or without gaps
skewed
describes a distribution with a longer tail on one side than the other
Check yourself

1. A data set is strongly skewed right with two large outliers. Which pair of summaries should you report?

2. Dana's runs allowed, 2, 2, 3, 3, 3, 3, 4, 4, have mean 3. Dividing by 8, what is the standard deviation, to two decimals?

3. Which pairing of center and spread is correctly matched?

Section 4

Two Questions at Once

51.10

Two-Way Frequency Tables

Main ideaA two-way table counts people by two categories at once, with the cells giving joint counts and the margins giving totals for each category.

A school surveys 100 students with two questions: what grade are you in, and how do you get to school? Every student is in grade 9 or grade 10, and every student either rides the bus or walks. Two questions with two answers each make four combinations, and a has one cell for each. Say 35 ninth graders ride the bus and 25 walk, while 20 tenth graders ride the bus and 20 walk. Rows are grades, columns are how they travel. The four cells hold 35, 25, 20 and 20, which add to 100.

The numbers inside the four cells are joint frequencies. Each one counts students who fit two categories at the same time, like ’grade 9 and bus.’ Now add across each row and down each column. Row totals: grade 9 has 35 + 25 = 60 students, and grade 10 has 20 + 20 = 40. Column totals: bus riders number 35 + 20 = 55, and walkers number 25 + 20 = 45. These totals are written in the margins of the table. They are the marginal frequencies. Each margin counts one category by itself, ignoring the other question. Both sets of margins add to the same grand total: 60 + 40 = 100 and 55 + 45 = 100. That is a built-in check on your arithmetic.

Reading a two-way table is a matter of finding the right cell. How many tenth graders walk? Go to the grade 10 row and the walk column: 20. How many students ride the bus in total? Read the bus column margin: 55. How many students are in grade 9? Read the grade 9 row margin: 60. A common mistake is answering a marginal question with a joint cell, like saying 35 ninth graders when the question asked for all ninth graders. Read the question twice: does it name one category or two?

Words to know
two-way table
a table that sorts each person by two categories at once, with rows for one category and columns for the other
joint frequency
a count in one cell of a two-way table; the number of people in two categories at the same time
marginal frequency
a row total or column total of a two-way table; the count for one category alone
Check yourself

1. In the survey table, how many tenth graders walk to school?

2. How many of the 100 students ride the bus?

3. Which number in the table is a marginal frequency?

51.11

Conditional Relative Frequencies

Main ideaA conditional relative frequency divides a cell by its row or column total, which lets you compare groups of different sizes and look for an association.

The survey table says 35 ninth graders and 20 tenth graders ride the bus. Does that mean ninth graders like the bus more? Not yet. There are 60 ninth graders and only 40 tenth graders, so of course the ninth grade has more of everything. To compare fairly, turn counts into fractions of the right group. A is a cell count divided by some total. Divide by the grand total, 100, and you get the joint relative frequency: 35 ÷ 100 = 0.35, so 35% of all students are ninth graders who ride the bus.

The more useful move is to divide by a row or column total. That gives a , because you are asking a question on the condition that someone belongs to a certain group. Among ninth graders, what fraction ride the bus? Divide the cell by the row total: 35 ÷ 60 ≈ 0.583, about 58%. Among tenth graders: 20 ÷ 40 = 0.50, exactly 50%. Now the comparison is fair. Ninth graders are somewhat more likely to ride the bus than tenth graders, 58% versus 50%.

You can condition on the columns instead. Among bus riders, what fraction are in grade 9? Divide by the column total: 35 ÷ 55 ≈ 0.636, about 64%. Among walkers: 25 ÷ 45 ≈ 0.556, about 56%. Notice that ’the fraction of ninth graders who ride the bus’ (58%) and ’the fraction of bus riders who are ninth graders’ (64%) are different questions with different answers. The most common mistake in this whole topic is dividing by the wrong total. Say the condition out loud: ’of the ninth graders’ means divide by 60; ’of the bus riders’ means divide by 55.

When the conditional relative frequencies differ from group to group, the two categories have an : knowing a student’s grade tells you something about how they travel. If every row gave the same percents, the categories would be independent of each other. Here the gap is 58% versus 50%, a modest association. Whether a gap that size matters is a question for later courses; for now, the skill is computing the right fraction and stating clearly which group it describes.

Words to know
relative frequency
a count divided by a total, written as a fraction, decimal or percent
conditional relative frequency
a cell count divided by its row total or column total; the fraction of one group that falls in a category
association
a relationship between two categories, shown when the conditional relative frequencies differ from row to row
Check yourself

1. Of the 40 tenth graders, 20 walk to school. What percent of tenth graders walk?

2. Of the 45 walkers, 25 are in grade 9. About what percent of walkers are ninth graders?

3. What is the joint relative frequency of students who are in grade 9 and ride the bus?

51.12

Reading Data Honestly

Main ideaCheck the axis, the sample and the choice of summary before you believe a graph or a statistic, because each one can be bent to mislead.

A poster shows two bars: last year’s test average, 50, and this year’s, 52. The bar for 52 is twice as tall as the bar for 50. How? The vertical axis does not start at 0. It starts at 48. Measured from 48, the first bar is 2 units tall and the second is 4 units tall, so it looks like a doubling. The real change is 52 ÷ 50 = 1.04, a 4% rise. A , one that cuts off the bottom, is the most common trick in graphs. Always find the number where the axis begins before you compare bar heights.

The second thing to check is who was asked. A gym hands out a survey about exercise habits to its own members and reports that 90% of people exercise three times a week. The , the group that was actually measured, is gym members, and gym members are not typical of everyone. The sample is : it leans in one direction before a single question is asked. A fair sample gives every person in the group you care about a chance to be picked. Small samples are a related problem: a 3-out-of-4 result from four people is not a 75% result for a school of 1,600.

The third check is the summary itself. A company with 40 workers earning $40,000 and one owner earning $2,000,000 can honestly say its mean salary is over $87,000: the total is 1,600,000 + 2,000,000 = 3,600,000, divided by 41 is about $87,800. The median is $40,000. Both numbers are true. Only one of them describes a typical worker. When you see an average, ask which kind, and ask what the shape of the data looks like. When you see a percent, ask what the total was. When you see a graph, find zero on the axis.

None of this makes statistics untrustworthy. It makes you a careful reader. The same tools that can mislead, axes, samples and summaries, are the tools that tell the truth when used well. A histogram with an honest axis, a sample chosen fairly and a summary that fits the shape of the data is the strongest kind of argument there is. Your job as a reader is to check that all three are in place, and your job as a writer is to make sure they are.

Words to know
truncated axis
a graph axis that does not start at zero, which exaggerates small differences between bars
sample
the group of people or things that were actually measured, chosen from a larger population
biased
describes a sample or method that leans toward one answer before the data are even collected
Check yourself

1. A bar graph's axis starts at 48. One bar shows 50 and another shows 52. How many times taller does the 52 bar look compared with the 50 bar?

2. A company has 40 workers earning $40,000 and one owner earning $2,000,000. Which summary better describes a typical worker's pay?

3. A gym surveys its own members and reports that 90% of people exercise three times a week. What is the main problem?

Chapter review

Describing and Comparing Data

0 / 8

1. A dot plot has a tall stack of dots on the right and a long thin tail stretching left. Its shape is best described as:

2. A histogram of 40 test scores has bins 50 to 59, 60 to 69, 70 to 79, 80 to 89 and 90 to 99 with counts 3, 7, 14, 11 and 5. How many scores are 80 or higher?

3. What is the median of 7, 2, 9, 4, 11, 6, 3?

4. The data set 10, 12, 15, 18, 25 has mean 16. Which value has a deviation of −4?

5. A data set has Q1 = 12 and Q3 = 20. Which value would be flagged as an outlier by the 1.5 × IQR rule?

6. Two runners have the same mean 5K time of 24 minutes. Runner A's times have a standard deviation of 0.5 minutes; Runner B's have 4 minutes. What can you conclude?

7. In a two-way table, 30 of 50 freshmen and 21 of 30 sophomores own a bike. What percent of sophomores own a bike?

8. A graph's vertical axis starts at 95 instead of 0, and bars of 96 and 99 are drawn. What is the effect?

Chapter

Lines of Fit and Correlation

Statistics
Big questionWhen two things rise and fall together, how do we measure the pattern, and how do we know whether one is causing the other?
The story

The Number That Fooled a Town

A lakeside town council finds a striking pattern in its records and nearly passes a law based on it.

Picture a small town on the shore of Lake Michigan with a public beach and a boardwalk lined with ice cream stands. Every year the town clerk records two numbers for each month: total ice cream sales on the boardwalk and the number of drownings and near-drownings reported by the lifeguards. One autumn a new council member plots the twelve months on a graph, sales on the horizontal axis and water rescues on the vertical axis, and gasps. The dots climb from lower left to upper right in a nearly straight line.

He brings the graph to the next meeting. In January, sales are close to zero and so are rescues. In July, sales are at their peak and so are rescues. He calculates the correlation coefficient, a number that measures how tightly points hug a line, and gets 0.93, about as strong as real-world data ever gets. His proposal follows: close the ice cream stands, or at least move them off the beach, and drownings will fall. Several council members nod. The pattern is right there in the data.

The lifeguard captain asks for the floor. She agrees that the graph is accurate. She asks the council to add one more column to the table: the average air temperature for each month. Then she asks two questions. When is it hot? When do people buy ice cream, and when do people swim? The answers are the same word three times: summer. Heat drives people to buy ice cream, and heat drives people into the water. Ice cream and rescues rise together because a third thing pushes both of them.

The proposal dies, and the town instead hires two more lifeguards for July and August. This story is a classic example that statistics teachers tell, and the two variables change from telling to telling, but the lesson does not. A strong correlation is real information: the two numbers really do move together, and one really can be used to predict the other. What a correlation cannot do, by itself, is say why. That is the central warning of this chapter, and it comes after you learn how to find the line, measure the fit and use it to predict.

By the end you will be able to draw a line through a cloud of points and say what its slope means, check the line with residuals, read the correlation coefficient without overreading it, recognize when a curve fits better than a line, and know how far from your data a prediction can safely reach. Each skill is a tool. The council member had the tools; he was missing the question the lifeguard asked.

Talk about itThe council member's graph was accurate and his correlation was strong. What exactly was his mistake, and what single question would have caught it?
Section 1

Points and Lines

52.1

Reading a Scatter Plot

Main ideaA scatter plot puts one dot per pair of measurements, and its overall drift up or down shows whether the two variables move together.

Five students record how many hours they studied for a quiz and the score they got: 1 hour and 60 points, 2 hours and 65, 3 hours and 75, 4 hours and 80, 5 hours and 90. Each student gives one pair of numbers, and each pair becomes one dot. A draws those dots on a grid: hours on the horizontal axis, score on the vertical axis. The student who studied 3 hours and scored 75 is the dot at (3, 75). Five students, five dots.

Which variable goes on which axis is not random. The is the one you think does the explaining or comes first, and it goes on the horizontal axis. Hours of studying come before the quiz, so hours are explanatory. The is the outcome you want to predict, and it goes on the vertical axis. The score responds to the studying. In an experiment the explanatory variable is the one you control; in a survey it is the one you would use to predict the other.

Now look at the drift. As hours go up, scores go up: the dots climb from lower left to upper right. That is a . If the dots fell from upper left to lower right, like hours of TV against quiz score, that would be a negative association. If the dots formed a shapeless cloud with no drift, there would be no association. Also note the form, whether the dots follow a straight line or a curve, and the strength, whether they hug the pattern tightly or scatter widely. These five dots follow a nearly straight line, tightly, going up.

A common mistake is reading a scatter plot as a time graph and connecting the dots in order. Do not connect them. Each dot is a separate student, not a step in a sequence, and the pattern is in the cloud as a whole. Another is describing a single dot as ’high’ without saying which axis: the student at (5, 90) is high in hours and high in score, and both facts matter.

Words to know
scatter plot
a graph that draws one dot for each pair of measurements, one variable on each axis
explanatory variable
the variable placed on the horizontal axis, used to explain or predict the other
response variable
the variable placed on the vertical axis, the outcome being predicted
positive association
a pattern in which larger values of one variable go with larger values of the other
Check yourself

1. In the study-time data, which variable belongs on the horizontal axis?

2. The dots in a scatter plot drift from upper left down to lower right. What does this show?

3. What does the dot at (3, 75) represent?

52.2

Drawing a Line of Fit

Main ideaA line of fit runs through the middle of the cloud of dots, and its equation lets you predict the response from the explanatory variable.

The five study dots, (1, 60), (2, 65), (3, 75), (4, 80) and (5, 90), nearly form a line. A is a straight line drawn through the cloud so that the dots are balanced around it, about as many above as below and none very far away. To find an equation by hand, pick two points that seem to sit on the trend, not necessarily data points. Here the first and last dots look right. Use (1, 60) and (5, 90).

The is rise over run: (90 − 60) ÷ (5 − 1) = 30 ÷ 4 = 7.5. Each extra hour of studying goes with 7.5 more points. Now find the using one point. If y = 7.5x + b and the point (1, 60) is on the line, then 60 = 7.5 × 1 + b, so b = 52.5. The line is y = 7.5x + 52.5. Check it with a point you did not use: at x = 3, y = 7.5 × 3 + 52.5 = 22.5 + 52.5 = 75, which is exactly the middle dot. At x = 4, y = 30 + 52.5 = 82.5, close to the actual 80.

A calculator or spreadsheet finds the . That is the one line that makes the total of the squared vertical gaps as small as possible. For these data it gives y = 7.5x + 51.5, just 1 point below the hand-drawn line. Two people drawing by eye will get slightly different lines. That is fine, as long as each line runs through the middle of the cloud. The least-squares line is the standard answer everyone can agree on. This chapter uses it from here on.

The most common error is picking the two points badly, such as the highest dot and the lowest dot even when they are off the trend, or forcing the line through the origin. The line does not have to pass through (0, 0); it has to pass through the middle of the data. Another slip is mixing up rise and run. Slope is change in y divided by change in x, vertical over horizontal, always.

Words to know
line of fit
a straight line drawn through a scatter plot so that the dots are balanced around it
slope
change in y divided by change in x between two points on a line; how much y changes per one unit of x
y-intercept
the value of y where a line crosses the vertical axis, at x = 0
least-squares line
the line of fit that makes the sum of the squared vertical gaps from the dots as small as possible
Check yourself

1. What is the slope of the line through (2, 10) and (6, 22)?

2. Using the line y = 7.5x + 52.5, what score does it predict for 4 hours of studying?

3. Which describes a good line of fit?

52.3

Slope and Intercept in Context

Main ideaIn a line of fit, the slope is the predicted change in the response for each one-unit increase in the explanatory variable, and the intercept is the prediction at zero.

A used-car website fits a line to the ages and prices of a certain model. With age a in years and value v in thousands of dollars, the line is v = 24 − 2.5a. Numbers in an equation are only useful once you can say them in words. The slope is −2.5. In context: for each additional year of age, the predicted value drops by 2.5 thousand dollars, that is, $2,500 per year. The is 24. In context: a car of age 0, brand new, is predicted to be worth $24,000.

Use the line to . A 4-year-old car: v = 24 − 2.5 × 4 = 24 − 10 = 14, so about $14,000. A 6-year-old car: v = 24 − 15 = 9, about $9,000. Notice the word ’predicted’ or ’about’ in every sentence. The line describes the trend across many cars; any one car can sit above or below it because of mileage, accidents or paint color. A line of fit predicts the typical value, not the exact value.

Always attach units to slope and intercept. ’The slope is −2.5’ is incomplete. ’Value drops by $2,500 per year’ is a sentence a car buyer can use. The slope’s units are always the response units per one explanatory unit: dollars per year, points per hour, centimeters per month. The intercept has the response’s units alone.

Sometimes the intercept has no real meaning. A line predicting a child’s height from age might have an intercept of 50 cm, which is sensible. But a line predicting shoe size from height might have an intercept of −20. Nobody has a negative shoe size. The intercept is just where the line crosses the axis. It only describes reality if x = 0 is a situation that actually occurs in the data. A common mistake is to explain a slope backwards, as ’years per dollar.’ Read it as y per x. The response changes this much when the explanatory variable rises by one.

Words to know
intercept
the predicted value of the response when the explanatory variable is 0
predict
to use a line of fit to estimate the response for a given value of the explanatory variable
rate of change
the slope of a line read as a rate: how many response units per one explanatory unit
Check yourself

1. In the line v = 24 − 2.5a, what does the slope mean?

2. Using v = 24 − 2.5a, what is the predicted value of a 4-year-old car?

3. A line predicting a person's weight from height has an intercept of −90 kg. What is the best interpretation?

Section 2

Checking the Line

52.4

Residuals Measure the Miss

Main ideaA residual is the actual value minus the predicted value; positive means the dot is above the line and negative means below.

A line of fit rarely passes through every dot, so measure the misses. Use the least-squares line for the study data, y = 7.5x + 51.5. For the student at (2, 65), the line predicts y = 7.5 × 2 + 51.5 = 15 + 51.5 = 66.5. The student actually scored 65. The is actual minus predicted: 65 − 66.5 = −1.5. The dot sits 1.5 points below the line. Negative residual, dot below.

Do all five. At x = 1 the is 7.5 + 51.5 = 59 and the actual is 60, so the residual is 60 − 59 = 1. At x = 3 the prediction is 22.5 + 51.5 = 74 and the actual is 75, residual 1. At x = 4 the prediction is 30 + 51.5 = 81.5 and the actual is 80, residual −1.5. At x = 5 the prediction is 37.5 + 51.5 = 89 and the actual is 90, residual 1. The residuals are 1, −1.5, 1, −1.5, 1. Add them: 1 − 1.5 + 1 − 1.5 + 1 = 0. That is not a coincidence. The least-squares line always makes the residuals sum to zero; the dots above balance the dots below exactly.

Residuals are how the least-squares line earns its name. Square each residual: 1, 2.25, 1, 2.25, 1, which add to 7.5. Any other line through these dots gives a bigger total. Try the hand-drawn line y = 7.5x + 52.5 from the last lesson: its residuals are 0, −2.5, 0, −2.5, 0, and their squares add to 12.5, larger than 7.5. Smaller total means better fit, and the least-squares line is the champion by definition.

Keep the order straight: actual minus predicted, never the reverse. Students who flip it get the right size and the wrong sign, and then describe a dot as above the line when it is below. A second slip is reading the residual off the horizontal axis. A residual is a vertical distance, measured straight up or down from the dot to the line, because it is a miss in the predicted response, not in the explanatory variable.

Words to know
residual
the actual value minus the value the line predicts; the vertical gap from a dot to the line
predicted value
the y-value the line of fit gives for a particular x
actual value
the y-value that was really measured for a data point
Check yourself

1. Using y = 7.5x + 51.5, what is the residual for the point (2, 65)?

2. A point has a negative residual. Where is it?

3. A line predicts 44 for a point whose actual value is 50. What is the residual?

52.5

Residual Plots and Patterns

Main ideaA residual plot with no pattern says a line fits well; a residual plot with a curve or a fan says the line is the wrong model.

A draws the residuals themselves: x on the horizontal axis and each residual on the vertical axis, with a horizontal line at 0 marking the fit line. For the study data the residuals are 1, −1.5, 1, −1.5, 1. They bounce above and below zero with no trend, all within 1.5 points. That is what a good fit looks like. If a line is the right model, the leftover misses should look like random noise, small and patternless.

Now try data that grow faster and faster: (1, 2), (2, 5), (3, 10), (4, 17), (5, 26). The least-squares line is y = 6x − 6. Its predictions are 0, 6, 12, 18 and 24. The residuals are 2 − 0 = 2, 5 − 6 = −1, 10 − 12 = −2, 17 − 18 = −1, and 26 − 24 = 2. In order they read 2, −1, −2, −1, 2: positive at both ends, negative in the middle. Plotted, they form a U. The line is too high in the middle and too low at both ends, because the data bend upward and a straight line cannot bend.

A U or an upside-down U in the residual plot means a curve would fit better than a line. A fan shape is different: residuals are tiny on the left and huge on the right. It means the line’s predictions get less reliable as x grows. Either is a warning. These patterns are easy to miss on the scatter plot, where the dots may look close to the line. They are impossible to miss once the line is flattened out to zero. That is the point of the residual plot.

Do not judge a fit by the size of the residuals alone. The U-shaped residuals above are only 2 units off on data that reach 26, which sounds fine, yet the pattern says the is wrong and will be badly wrong beyond x = 5. Small residuals with a pattern are a worse sign than larger residuals without one. Also remember that a residual plot uses the same x-values as the scatter plot; only the vertical axis changes.

Words to know
residual plot
a graph of residuals against x, with a horizontal line at zero standing for the line of fit
pattern
a visible shape in a residual plot, such as a U or a fan, that shows the model is missing something
linear model
a straight-line equation used to describe and predict a relationship
Check yourself

1. A residual plot reads 2, −1, −2, −1, 2 from left to right. What does this suggest?

2. A residual plot shows dots scattered randomly above and below zero with no shape. What does this tell you?

3. For the point (5, 26), the line y = 6x − 6 predicts 24. What is the residual?

Section 3

Measuring the Relationship

52.6

The Correlation Coefficient

Main ideaThe correlation coefficient r, between −1 and 1, gives the direction and strength of a straight-line relationship in a single number.

Words like ’strong’ and ’weak’ are vague. The , written r, replaces them with a number. It is always between −1 and 1. The sign gives the : positive r means the dots drift upward, negative r means they drift downward. The size gives the : r near 1 or −1 means the dots hug a straight line tightly, and r near 0 means the dots scatter with little straight-line pattern. If every dot sits exactly on an upward line, r = 1. On a downward line, r = −1.

A calculator computes r from the data. The formula multiplies how far each x is from its mean by how far the matching y is from its mean. It adds those products and scales the total so the result lands between −1 and 1. For the study data, (1, 60) through (5, 90), r ≈ 0.99. That is a very strong positive correlation, and it matches the picture of five dots almost on a line. Here are rough guides for reading r. Above 0.8 in size is strong. From 0.5 to 0.8 is moderate. Below 0.5 is weak. The cutoffs shift by field, though. A medical study may call 0.3 meaningful; a physics lab would call it noise.

Two facts make r easy to work with. First, r has no units. Change hours to minutes or points to percents and r does not move, because it measures pattern, not scale. Second, r is the same whichever variable you put on which axis, since it treats x and y symmetrically. The slope of the line of fit has neither property; it changes with the units and flips when you swap the axes.

Students often confuse the strength of r with the steepness of the line. A slope of 0.001 can have r = 1 if the dots sit exactly on that gentle line, and a steep slope of 50 can have r = 0.2 if the dots scatter widely around it. Steepness is slope. Tightness is r. Another mistake is reading r = −0.9 as weaker than r = 0.6 because it is ’negative.’ Compare sizes: 0.9 is larger than 0.6, so the negative relationship is the stronger one.

Words to know
correlation coefficient
a number r between −1 and 1 measuring the direction and strength of a straight-line relationship
direction
whether a relationship is positive (dots drift up) or negative (dots drift down), shown by the sign of r
strength
how tightly the dots follow a straight line, shown by how close r is to 1 or −1
Check yourself

1. A data set has r = −0.92. Which description fits?

2. Which correlation coefficient shows the strongest straight-line relationship?

3. Data measured in inches has r = 0.75. The same data are converted to centimeters. What is r now?

52.7

What r Cannot See

Main ideaThe correlation coefficient only measures straight-line pattern, so a curve, an outlier or a cluster can make r badly misleading.

Take five dots that sit perfectly on a U: (−2, 4), (−1, 1), (0, 0), (1, 1) and (2, 4). Every y equals x squared, so the relationship is as exact as a relationship can be. Compute r and you get exactly 0. Why? The left half of the U falls and the right half rises, and the two halves cancel in the formula. A student who looks only at r would say ’no relationship’ and be completely wrong. The value r = 0 means no relationship, nothing more. Always look at the scatter plot before you trust r.

Now the reverse trap. The curving data from the residual lesson, (1, 2), (2, 5), (3, 10), (4, 17), (5, 26), give r ≈ 0.98. That sounds like a perfect line. But the residual plot was a clear U, and the line predicts 6 × 10 − 6 = 54 at x = 10 when the curve y = x^2 + 1 gives 101. A high r does not prove that a line is the right model; it only says the dots rise together. The residual plot catches what r misses.

A single can also swing r. Picture nine dots forming a shapeless blob with r near 0, and then add one dot far away in the upper right. The formula counts that dot’s huge distances from both means, and r can jump to 0.7 or more, from a single point. The opposite happens too: one dot far off a tight line can drag r from 0.95 down to 0.5. When r changes a lot after removing one point, report both values and say so.

Finally, r says nothing about how much y changes per unit of x. That is the slope’s job. And r says nothing about why the variables move together, which is the next lesson. The correlation coefficient is one of one feature of a scatter plot. It is a good summary of that one feature. Everything else, form, outliers, clusters and cause, needs the picture and your judgment.

Words to know
straight-line relationship
a pattern in which the dots follow a line; the only kind of pattern r measures
outlier
a dot far from the rest of the scatter plot, which can pull the line and change r by a large amount
summary
one number that captures a single feature of a data set, leaving other features out
Check yourself

1. A scatter plot shows a clear U shape and r = 0.05. What is the right conclusion?

2. The curving data (1, 2) through (5, 26) have r ≈ 0.98, yet their residual plot is a U. What does this show?

3. Removing a single dot from a scatter plot changes r from 0.95 to 0.50. What should you report?

52.8

Correlation Is Not Causation

Main ideaA strong correlation shows that two variables move together; it does not show that one causes the other, because a third variable may drive both.

In the chapter story, ice cream sales and water rescues had a correlation near 0.93, and a council member concluded that ice cream causes drowning. His graph was right and his conclusion was wrong. Heat, the , drives both: hot days send people to the ice cream stand and into the lake. A lurking variable is one that was left out of the plot but affects both variables in it. Whenever two things rise and fall together, ask what else changes at the same time.

Here is another. Among elementary school children, shoe size and reading level are strongly correlated. Bigger feet, better readers. Does buying larger shoes improve reading? Of course not; the lurking variable is age. Older children have bigger feet and have had more years of reading lessons. The correlation is real and could even be used to predict: tell me a child’s shoe size and I can guess their reading level better than chance. Prediction works. , one thing making another happen, was never shown.

Correlation can also run backwards. Cities with more police officers tend to have more reported crime. Do officers cause crime? More likely, crime leads cities to hire officers. The direction of cause is not in the scatter plot; both variables are just numbers on axes. And sometimes a correlation is pure coincidence: with thousands of variables being tracked, some pairs will move together by chance, like the number of movies an actor makes and the price of a vegetable.

How do scientists ever show cause? With an : they change the explanatory variable on purpose, assign subjects to groups at random so lurking variables spread evenly, and then compare the responses. A survey or a record book, where nobody controlled anything, is , and it can show correlation only. Every headline that says ’X linked to Y’ is reporting a correlation. The careful reader asks: was this an experiment, or could a lurking variable explain it, or could the cause run the other way?

Words to know
lurking variable
a variable left out of the analysis that affects both plotted variables and creates the correlation between them
causation
a relationship in which changing one variable directly makes the other change
experiment
a study in which the researcher changes the explanatory variable on purpose and assigns subjects to groups at random
observational
describes a study that records what happens without controlling anything; it can show correlation but not cause
Check yourself

1. Ice cream sales and water rescues rise and fall together across the year. What is the most likely explanation?

2. Among children, shoe size and reading level are strongly correlated. What is the lurking variable?

3. Which kind of study can show that one variable causes another?

Section 4

Beyond the Straight Line

52.9

When a Line Fails: Exponential Growth

Main ideaWhen each step multiplies the response by the same factor instead of adding the same amount, the data follow an exponential model, not a line.

A biology class counts bacteria in a dish every hour: 100 at hour 0, then 200, 400, 800 and 1,600. Try a line. The changes from hour to hour are +100, +200, +400 and +800. A line needs the same change every step, and these changes double. A line fit to these points would be too high at hours 1 and 3 and far too low at hour 4, and the residual plot would show a U. The line fails because the pattern is not additive. It is multiplicative.

Look at instead of differences. 200 ÷ 100 = 2, 400 ÷ 200 = 2, 800 ÷ 400 = 2, 1,600 ÷ 800 = 2. Every hour the count multiplies by 2. A pattern with a constant ratio is . Its model is y = a × b^x, where a is the starting value and b is the , the constant ratio. Here y = 100 × 2^x. Check: at x = 3, 100 × 2^3 = 100 × 8 = 800. To predict hour 6, compute 100 × 2^6 = 100 × 64 = 6,400.

Real data are never this clean, so the test is whether the ratios are roughly constant. Say a savings account shows $1,000, $1,050, $1,102, $1,158 over four years. Ratios: 1.05, 1.0495, 1.0508. Close enough to call it exponential with growth factor about 1.05, or 5% per year. The differences, 50, 52 and 56, creep upward, which is the fingerprint of exponential growth: each year’s increase is a little bigger than the last because it is a percent of a bigger amount.

The test for a line is constant differences; the test for an exponential is constant ratios. Students often check only differences, see them change, and give up. Check ratios next. A second common mistake is writing the growth factor as the percent: 5% growth is a factor of 1.05, not 0.05 or 5. And a ratio less than 1, like 0.9, still gives an exponential model, one that decays, such as a medication leaving the bloodstream by 10% each hour.

Words to know
exponential model
an equation of the form y = a × b^x, where each step in x multiplies y by the same factor b
growth factor
the constant ratio b between one value and the next in an exponential pattern
ratio
one value divided by the previous one; constant ratios point to an exponential model
Check yourself

1. A quantity starts at 5 and triples every step. What is its value after 4 steps?

2. Which table of y-values, for x = 0, 1, 2, 3, shows an exponential pattern?

3. Using y = 100 × 2^x, what is the predicted count at hour 6?

52.10

Quadratic Models and Second Differences

Main ideaWhen first differences change but second differences are constant, the data follow a quadratic model with an x-squared term.

A ball is dropped from a tall building, and a camera records how far it has fallen after each second, in feet: 0, 16, 64, 144, 256. First differences: 16, 48, 80, 112. Not constant, so not a line. Ratios: 64 ÷ 16 = 4, 144 ÷ 64 = 2.25, 256 ÷ 144 ≈ 1.78. Not constant either, so not exponential. Now take differences of the differences: 48 − 16 = 32, 80 − 48 = 32, 112 − 80 = 32. The are constant. That is the fingerprint of a , an equation with an x^2 term.

Here the model is d = 16t^2, distance in feet after t seconds. Check: at t = 3, 16 × 9 = 144. Predict t = 5: 16 × 25 = 400 feet. A simpler example: y-values 0, 5, 20, 45, 80 for x = 0 through 4 have first differences 5, 15, 25, 35 and second differences 10, 10, 10. The model is y = 5x^2. At x = 6 it predicts 5 × 36 = 180.

Quadratic patterns show up whenever something speeds up steadily or turns around. A falling object is one example. So is the height of a thrown basketball on its way up and back down. So is the area of a square as its side grows. A calculator can fit a quadratic to messy data just as it fits a line. The same tools apply: look at the residual plot. If a line leaves a U in the residuals and a quadratic leaves patternless noise, the quadratic is the better model.

Keep the three tests in one place. Constant first differences: linear, y = mx + b. Constant ratios: exponential, y = a × b^x. Constant second differences: quadratic, with an x^2 term. Always check the tests in that order, and only on equally spaced x-values; differences between x = 1 and x = 2 and between x = 2 and x = 5 are not comparable. A common slip is to compute second differences from the original values instead of from the first differences.

Words to know
quadratic model
an equation with an x^2 term, such as y = 5x^2, that fits data which curve up or down and turn around
first differences
the changes between one y-value and the next for equally spaced x-values
second differences
the changes between one first difference and the next; constant for a quadratic pattern
Check yourself

1. A table of y-values for x = 0, 1, 2, 3, 4 has second differences that are all equal and not zero. What kind of model fits?

2. Using y = 5x^2, what is the predicted value at x = 6?

3. For x = 0, 1, 2, 3, the y-values are 3, 7, 11, 15. Which model fits?

52.11

How Far a Prediction Can Reach

Main ideaPredicting inside the range of the data is interpolation and is usually safe; predicting outside it is extrapolation and can fail badly.

The study-time line, y = 7.5x + 51.5, was built from students who studied 1 to 5 hours. Ask it about 3.5 hours: y = 7.5 × 3.5 + 51.5 = 26.25 + 51.5 = 77.75, about 78 points. That is , a prediction between data points you actually have. Nobody in the data studied exactly 3.5 hours, but people studied 3 and 4, and the line between them is a reasonable guess.

Now ask about 10 hours: y = 7.5 × 10 + 51.5 = 75 + 51.5 = 126.5. A score of 126.5 on a 100-point quiz is impossible. The line did nothing wrong; it kept going straight, as lines do. The mistake was asking it about a place with no data. That is , predicting beyond the range that was measured. Real relationships bend, level off or stop, and the line cannot know that. Ask about 0 hours and the line says 51.5, which might be reasonable, or might not; no student in the data studied 0 hours either.

Extrapolation is not forbidden, but it needs a warning label, and the farther you go, the louder the label. Predicting 5.5 hours from data that reach 5 is a small stretch. Predicting 20 hours is a leap into the dark. Exponential models extrapolate especially badly: the bacteria model y = 100 × 2^x predicts 100 × 2^24, over 1.6 billion, after one day, and the dish would have run out of food long before. Every model has a range where it is trustworthy, and that range is roughly the range of the data that built it.

When you report a prediction, state the range of the data next to it. For example: ’Based on students who studied 1 to 5 hours.’ That one clause tells the reader whether your prediction is interpolation or extrapolation. A frequent mistake is reading a prediction off a line that has been extended on the graph to fill the page. The drawn line may extend to the edges. The data do not. Only the data part deserves trust.

Words to know
interpolation
using a model to predict a value inside the range of the data it was built from
extrapolation
using a model to predict a value outside the range of the data it was built from
range of the data
the span from the smallest to the largest x-value that was actually measured
Check yourself

1. The line y = 7.5x + 51.5 was built from 1 to 5 hours of study. Predicting a score for 10 hours gives 126.5. What is the main problem?

2. Using y = 7.5x + 51.5, what score is predicted for 2.5 hours of study?

3. A model was built from data with x from 20 to 60. Which prediction is interpolation?

Chapter review

Lines of Fit and Correlation

0 / 8

1. A scatter plot of car weight against gas mileage drifts from upper left to lower right. Which statement is correct?

2. What is the slope of the line through (1, 8) and (5, 20)?

3. A line predicts a plant's height as h = 4 + 1.5w, with w in weeks and h in cm. What does 1.5 mean?

4. A line predicts 32 for a point whose actual value is 28. What is the residual, and where is the point?

5. Which correlation coefficient describes the weakest straight-line relationship?

6. A study finds that towns with more churches have more crime. What is the most likely lurking variable?

7. For x = 0, 1, 2, 3, the y-values are 4, 12, 36, 108. Which model fits?

8. A line was fit to data with x from 5 to 25. Which prediction is an extrapolation?

Unit wrap-up

Statistics in Algebra I

Twelve words, twelve meanings

0 / 12

Tap a word, then tap its meaning. A right pair locks in green.

Words
Meanings
Unit test

Fifteen questions across the unit

0 / 15

1. Which set of summaries should you report for a data set that is strongly skewed with outliers?

2. What is the mean of 6, 9, 11, 14 and 20?

3. What is the median of 18, 4, 9, 13, 6, 11?

4. A sorted data set is 10, 12, 14, 16, 18, 20, 22. What is its interquartile range?

5. The data set 3, 5, 7 has mean 5. Dividing by the number of values, what is its standard deviation?

6. A histogram of 50 commute times has bins 0 to 9, 10 to 19, 20 to 29 and 30 to 39 with counts 8, 17, 15 and 10. What percent of commutes are under 20 minutes?

7. A data set has Q1 = 40 and Q3 = 60. Which value is an outlier by the 1.5 × IQR rule?

8. In a two-way table, 24 of 60 juniors and 18 of 30 seniors have a part-time job. Which statement is true?

9. A poll about a new stadium surveys only people leaving a home game. What is the main flaw?

10. What is the slope of the line through (2, 5) and (6, 17)?

11. A line predicts a taxi fare as f = 3 + 2.5m, with m in miles. What does the 3 mean?

12. A line predicts 70 for a point whose actual value is 64. What is the residual?

13. A residual plot shows a clear U shape. What should you conclude?

14. Which statement about the correlation coefficient r is true?

15. Cities with more ice cream shops have more sunburn cases. What is the best explanation?

Spiral review

Five questions from earlier units

0 / 5

1. (Unit 21) A 45-45-90 triangle has legs of 4. How long is the hypotenuse?

2. (Unit 20) Two triangles have two pairs of congruent angles and a pair of congruent sides that are not between those angles. Which criterion applies?

3. (Unit 19) Solve x^2 + 2x − 15 = 0.

4. (Unit 18) How many solutions does y = −x + 4 and y = −x − 2 have?

5. (Unit 21) What are the center and radius of (x − 1)² + (y + 6)² = 16?

Write it

Two classes took the same test. Class A scores: 70, 72, 75, 78, 80, 82, 85. Class B scores: 55, 60, 78, 80, 82, 95, 98. Compute the median, range and IQR for each class, then write a paragraph comparing the two classes. Explain which class did better, which was more consistent, and why the mean alone would not tell the whole story.

  • Sort each list first, then find the median, Q1 and Q3 before computing anything else.
  • Show every subtraction: range is max minus min, IQR is Q3 minus Q1.
  • Use shape, center and spread as three separate sentences in your comparison.
  • Say what each number means for a student in that class, not just what the number is.
  • Check your work by confirming that each class has the same number of scores on each side of its median.
0 wordsSaved on this device as you type.

Practice rooms

Rooms already on the site that belong to this unit — cards, quizzes, a lab.

For the teacher

Every lesson keeps its own three checks; a lesson is ticked when all three are right. Chapter reviews, the unit test and its spiral review (five questions from earlier units in this band) score on the page. When the site is connected to your sheet, or the link carries ?dest=, each one also has a Send box: the first-try score, the standards, the supports used, the attempt number and the minutes go to your sheet as an IEP data point.

Print this page for a paper copy of the readings, the sources, the words and the questions; the answers print as dashed boxes under each question.

Fact-check notes for this course live in the handoff: quotes marked (paraphrased) were set that way on purpose.