The Interior — MathGrades 6–8

Unit 17 · Statistics and Probability

A unit of the course: the story, then chapter by chapter — sections, numbered lessons, a source or the numbers to read, three checks each — a review per chapter, and the wrap-up at the end.

← Math, the whole course

Drawn scene: a classroom at evening with a dot plot growing on the whiteboard, a jar of coins, two dice and a spinner on the teacher's desk, dusk in the window
17Unit

Statistics and Probability

Statistics and Probability

How tall is a seventh grader? Is this coin fair? Does more sleep really go with better grades? None of these questions has a single answer you can look up. Each one asks about a whole group, or about chance, and the honest answer comes from collecting data and reading it carefully. That is what statistics and probability do. They turn a pile of numbers into a picture, a few summary values and a claim you can defend.

In this unit you will measure a class and describe it with dot plots, histograms and box plots. You will use the mean and median for the center, and the range, IQR and MAD for the spread. You will put a number on chance, from 0 for impossible to 1 for certain. You will list every outcome of a compound event and run simulations. You will learn why a random sample of 50 students can tell you something true about 800. Finally you will plot two measurements at once, fit a line by eye, and read its slope and intercept. You will also sort categories into two-way tables.

Along the way you will learn what data can and cannot prove. A bigger sample beats a smaller one. But a biased sample of a million is worse than a random sample of a hundred. Two things that rise together are associated, but association is not cause. By the end you will be able to answer a statistical question from start to finish. You will explain, with numbers, why your answer deserves to be believed.

How we figured it out
1654

Pascal and Fermat exchange letters about games of chance, starting the mathematics of probability

1662

John Graunt studies London's death records and publishes some of the first statistical tables

1713

Jacob Bernoulli's Ars Conjectandi is published, proving that long runs of trials settle near the true probability

1763

Thomas Bayes's essay on updating probabilities with evidence is published after his death

1786

William Playfair prints some of the first bar charts and line graphs of economic data

1790

The first United States Census counts about 3.9 million people

1805–1809

Legendre and Gauss publish the method of least squares for fitting a line to scattered data

1858

Florence Nightingale uses diagrams of hospital deaths to argue for better sanitation

1880s

Francis Galton describes regression and correlation from measurements of parents and children

1900

Karl Pearson introduces the chi-square test for comparing counts in tables

1977

John Tukey's Exploratory Data Analysis popularizes the box plot

Chapter

Describing Data

Statistics
Big questionHow can a few numbers and one picture describe a whole group fairly?
The story

Twenty-Eight Heights, One Class

A seventh-grade class in Chicago tries to answer a simple question and finds that one number is never enough.

Ms. Ortiz wrote one question on the board: How tall is a seventh grader? Her class of 28 laughed. Jamal stood up and said, "I am 158 centimeters, so that is the answer." Priya shook her head. "You are one seventh grader. The question asks about all of us." So they taped a measuring tape to the wall and took turns. By the end of the period they had 28 numbers, from 141 centimeters up to 176.

The list on the board was a mess of numbers in the order students had been measured. Nobody could see a pattern in it. Ms. Ortiz asked, "Can one number describe this class?" Jamal said the tallest, 176. Priya said "most of us are between 152 and 162." Devon added all 28 heights and divided by 28 and got 157.5. Three students, three different answers, and every one of them was telling the truth about something.

Then Priya drew a number line from 140 to 180 and put one dot above each height. Suddenly the class could see itself. The dots piled up between 152 and 162. A few sat lower. One dot, Marcus, stood alone at 176. Ms. Ortiz showed them the middle dot when the 28 were put in order: 157. The mean, 157.5, was a little higher, and the class figured out why. Marcus's 176 pulled it up, while the middle dot did not care how tall the tallest student was.

Next door, the eighth-grade class had done the same thing. When the two dot plots were drawn on the same line, the eighth-grade dots sat a bit to the right, but the two piles overlapped a lot. Was the eighth grade really taller, or just a little different by chance? Answering that takes more than a mean. It takes a way to measure how spread out each pile is. That is what this chapter builds: pictures of data, numbers for the center, numbers for the spread, and a way to compare two groups.

Talk about itIf a new student who was 190 centimeters tall joined the class, which would change more, the mean or the middle value? Why?
Section 1

Questions and Pictures of Data

37.1

Statistical Questions

Main ideaA statistical question expects answers that vary, and you answer it by collecting and describing data from a group.

Ask one friend, "How many pets do you have?" You get one answer, maybe 2. Ask the same question to everyone in your class and you get a whole set of answers: 0, 0, 1, 2, 3, 0, 1, and so on. That set of answers is . A is a question you expect to have many different answers, so you need data to answer it. "How many pets does Ana have?" is not statistical. "How many pets do students in our class have?" is.

The word for the differences among the answers is . Heights vary, bus wait times vary, and the number of pets varies. If there is no variability, you do not need statistics. "How many days are in a week?" has one answer, so it is not a statistical question. "How many minutes did students in Illinois spend on homework last night?" has thousands of answers, so it is.

To answer a statistical question, follow four steps. First, ask the question clearly. Second, collect the data from a group. Third, organize and picture the data with a plot. Fourth, describe the data with numbers such as the center and the spread. A common mistake is to stop after one answer, like "My cousin has 3 pets, so seventh graders have 3 pets." One answer is a fact about one person, not a description of a group.

Words to know
data
a set of answers or measurements collected from a group
statistical question
a question whose answers vary from person to person, so it takes data to answer
variability
the differences among the values in a set of data
Check yourself

1. Which question is a statistical question?

2. Why is "How many pets does Marcus have?" not a statistical question?

3. A class collects 28 heights. What is the next step in answering the statistical question?

37.2

Dot Plots

Main ideaA dot plot stacks one dot per data value above a number line, so you can see every value and where they pile up.

Ten students report their number of pets: 1, 0, 2, 1, 3, 0, 1, 2, 0, 1. A shows this on a number line from 0 to 3. Put one dot above the number for each answer and stack repeats. Above 0 you stack 3 dots, above 1 you stack 4 dots, above 2 you stack 2 dots, and above 3 you place 1 dot. Count the dots: 3 + 4 + 2 + 1 = 10, which matches the ten students. Always check that total.

The dot plot answers questions at a glance. The tallest stack shows the most common value, called the ; here it is 1 pet. The dots spread from 0 to 3, so the smallest and largest values are easy to spot. You can also count how many students have at least 2 pets: 2 + 1 = 3 students, which is 3 out of 10. The height of each stack is that value’s .

A common mistake is to put the dot in the wrong place, for example stacking a dot at 2 when the student said 3 pets. Another is skipping numbers on the line. The line must include every whole number from the smallest value to the largest, even if a number has no dots. A gap with no dots is information too. Dot plots work best for small sets of whole numbers, like pets or siblings. For hundreds of values, or measurements like 157.3 cm, a histogram works better.

Words to know
dot plot
a graph with one dot above the number line for each data value, stacking repeats
mode
the value that appears most often in a data set
frequency
how many times a value appears in the data
Check yourself

1. In the pets dot plot, how many students have fewer than 2 pets?

2. A dot plot of shoe sizes has 2 dots above 5, 5 dots above 6, 4 dots above 7 and 1 dot above 8. How many students were surveyed?

3. Which value is the mode of that shoe-size dot plot?

37.3

Histograms

Main ideaA histogram groups data into equal-width intervals and shows how many values fall in each one, with bars that touch.

Twenty students record last night’s homework time in minutes; the values range from 5 to 75. A dot plot with a dot at 5, 12, 18, 23, 25, 27 and so on would be hard to read. Instead, group the values into equal : 0–19, 20–39, 40–59 and 60–79 minutes. Count how many values land in each: 3, 8, 6 and 3. Check: 3 + 8 + 6 + 3 = 20. A draws a bar for each interval, with height equal to the count. The bars touch because the intervals sit next to each other on the number line.

Reading the histogram: the tallest bar is 20–39 minutes, so most students spent between 20 and 39 minutes. Nine students (6 + 3) spent 40 minutes or more. Nine out of twenty is 9/20 = 0.45, or 45% of the class. You cannot see exact values, only the interval each value belongs to. That is the trade-off: a histogram shows the shape clearly but hides the details.

Two common mistakes: using intervals of different widths, which makes bars unfair to compare, and putting a value on the wrong side of a boundary. Decide ahead of time that 20 minutes belongs to 20–39, not 0–19, and stick with it. Changing the interval width changes the picture. Intervals of 10 minutes would give eight bars and more detail; intervals of 40 would give two bars and almost no detail.

Words to know
histogram
a bar graph of counts in equal intervals, with bars that touch
interval
a range of values grouped together, like 20 to 39 minutes
Check yourself

1. In the homework histogram, how many students spent fewer than 40 minutes?

2. What percent of the 20 students spent 40 minutes or more?

3. Why do the bars of a histogram touch?

Section 2

Finding the Center

37.4

The Mean

Main ideaThe mean is the fair-share value: add all the values and divide by how many there are.

Five friends bring cookies to share: 3, 7, 4, 6 and 5 cookies. If they pool the cookies and share equally, how many does each get? Add: 3 + 7 + 4 + 6 + 5 = 25. Divide by the 5 friends: 25 ÷ 5 = 5. Each friend gets 5 cookies. That fair-share number is the , also called the average. The mean is the value every data point would have if the total were spread out evenly.

Try it with four test scores: 80, 90, 70 and 100. The is 80 + 90 + 70 + 100 = 340. Divide by 4 scores: 340 ÷ 4 = 85. The mean score is 85. Notice that the mean does not have to be one of the values, and it can be a decimal. The mean of 2, 3 and 5 is 10 ÷ 3, which is about 3.33.

You can also work backward. Suppose four numbers have a mean of 10, and three of them are 7, 9 and 12. The total must be 4 × 10 = 40. The three known numbers add to 28, so the missing number is 40 − 28 = 12. A common mistake is dividing by the wrong count, such as dividing 340 by 3 because you skipped a score. Another is forgetting a zero: if one friend brought 0 cookies, that friend still counts in the division.

Words to know
mean
the sum of the values divided by the number of values; the fair-share value
sum
the total you get by adding all the values
Check yourself

1. What is the mean of 12, 15, 9 and 20?

2. The mean of five numbers is 8. What is their sum?

3. Four numbers have a mean of 10. Three of them are 7, 9 and 12. What is the fourth?

37.5

The Median

Main ideaThe median is the middle value of the ordered data; half the values are at or below it and half at or above.

Five students’ heights, in centimeters, are 155, 150, 160, 152 and 158. To find the , first put the values in order: 150, 152, 155, 158, 160. The middle value is 155, with two values below and two above. The median height is 155 cm. The ordering step matters. If you take the middle of the unordered list you get 160, which is wrong.

When the count is even there are two middle values, and the median is their mean. Six values: 2, 3, 5, 8, 9, 12. The two middle values are 5 and 8. Their mean is (5 + 8) ÷ 2 = 13 ÷ 2 = 6.5. So the median is 6.5, even though 6.5 is not in the list. A quick way to find the middle: cross off the smallest and largest together, again and again, until one or two values remain.

The median is a position, not a total. Change the largest height from 160 to 200 cm. The median does not move; it is still 155. That makes the median a good choice when a few values are far from the rest. Watch out for repeated values: in 4, 4, 4, 9, 10, the median is 4, and each 4 counts as its own value in the .

Words to know
median
the middle value when the data are in order, or the mean of the two middle values
ordered list
the data written from smallest to largest
Check yourself

1. What is the median of 7, 2, 9, 4, 5?

2. What is the median of 10, 20, 30, 40?

3. The data set 3, 3, 4, 8, 9, 10, 12, 15 has median 8.5. If the 15 is changed to 50, what is the new median?

37.6

Mean or Median?

Main ideaUse the mean when values are fairly even; use the median when a few extreme values would drag the mean away from the typical value.

Six workers at a small shop earn these yearly amounts, in thousands of dollars: 30, 32, 35, 38, 40 and 200. The 200 is the owner. The mean is (30 + 32 + 35 + 38 + 40 + 200) ÷ 6 = 375 ÷ 6 = 62.5, so "the average worker earns $62,500." But five of the six earn under $41,000. The median is the mean of the two middle values, (35 + 38) ÷ 2 = 36.5, or $36,500. The median describes a typical worker far better.

A value far from the rest is called an . Outliers pull the mean toward them because the mean uses every value in its sum. The median ignores how far away an outlier is; it only counts it as one value. When the data have outliers or a long tail on one side, the median is usually the better . When the data are roughly symmetric, the mean and median are close, and either works.

Ask two questions before choosing. Are there outliers? Does the total matter? If you want the total, like the total cookies to buy, the mean is the right tool, because mean × count = total. If you want a typical value, like a typical home price in a Chicago neighborhood, the median is safer, because one mansion sale would distort the mean. A common mistake is to call the mean "more accurate." Both are exact; they just answer different questions.

Words to know
outlier
a value that sits far from the rest of the data
measure of center
a single number, like the mean or median, that describes a typical value
Check yourself

1. Data: 5, 6, 6, 7, 8, 40. Which statement is true?

2. For which purpose is the mean the better choice?

3. A set has mean 20 and median 20. A value of 100 is added. What happens?

Section 3

Measuring Spread

37.7

Range and Interquartile Range

Main ideaRange measures the full spread; interquartile range measures the spread of the middle half and is not thrown off by outliers.

Eight students count their books at home: 4, 6, 7, 9, 10, 12, 15, 18. The is the largest value minus the smallest: 18 − 4 = 14. The range is quick, but it depends only on two values. One student with 200 books would push the range to 196, and it would say nothing about the other seven.

The , or IQR, measures the spread of the middle half of the data. First find the median: the two middle values are 9 and 10, so the median is 9.5. The lower half is 4, 6, 7, 9; its median is (6 + 7) ÷ 2 = 6.5. That is the first , Q1. The upper half is 10, 12, 15, 18; its median is (12 + 15) ÷ 2 = 13.5. That is the third quartile, Q3. IQR = Q3 − Q1 = 13.5 − 6.5 = 7. The middle half of the students have roughly 6.5 to 13.5 books.

With an odd count, leave the median out of both halves. For 2, 4, 5, 7, 8, 10, 11, the median is 7. The lower half is 2, 4, 5, so Q1 = 4. The upper half is 8, 10, 11, so Q3 = 10. IQR = 10 − 4 = 6. Common mistakes: forgetting to order the data first, and subtracting the wrong pair. Range uses max and min; IQR uses Q3 and Q1. If one value becomes an outlier, the range jumps but the IQR barely moves.

Words to know
range
the largest value minus the smallest value
interquartile range
Q3 minus Q1; the spread of the middle half of the data
quartile
a value that splits the ordered data into quarters; Q1 is the median of the lower half, Q3 of the upper half
Check yourself

1. What is the range of 3, 11, 8, 15, 6?

2. Data: 2, 4, 5, 7, 8, 10, 11. What is the IQR?

3. Which measure of spread changes the least when one outlier is added?

37.8

Box Plots

Main ideaA box plot draws the five-number summary, so each of its four pieces holds about a quarter of the data.

Ten quiz scores, in order: 5, 6, 6, 7, 8, 8, 9, 9, 10, 10. Find five numbers. Minimum: 5. Maximum: 10. Median: the middle two are 8 and 8, so 8. The lower half 5, 6, 6, 7, 8 has median 6, so Q1 = 6. The upper half 8, 9, 9, 10, 10 has median 9, so Q3 = 9. The is 5, 6, 8, 9, 10.

A draws these on a number line. Draw a box from Q1 = 6 to Q3 = 9. Draw a line inside the box at the median, 8. Draw lines called from the box out to the minimum 5 and the maximum 10. The box covers the middle half of the scores. Its width is the IQR, 9 − 6 = 3. Each of the four pieces (whisker, half-box, half-box, whisker) holds about 25% of the data.

Reading a box plot takes care. A long whisker means the values in that quarter are spread out. It does not mean there are more of them. A common mistake is to think a longer box or whisker holds more data. It holds the same quarter, spread across more space. Box plots hide the exact values and the count. They are best for comparing groups side by side, like two classes’ quiz scores on one number line.

Words to know
box plot
a graph showing the minimum, Q1, median, Q3 and maximum as a box with whiskers
five-number summary
minimum, first quartile, median, third quartile, maximum
whisker
the line from the box out to the minimum or the maximum
Check yourself

1. A box plot has Q1 = 14 and Q3 = 22. What is the IQR?

2. Data: 5, 6, 6, 7, 8, 8, 9, 9, 10, 10. What is the median shown by the line inside the box?

3. About what fraction of the data lies inside the box of a box plot?

37.9

Mean Absolute Deviation

Main ideaMean absolute deviation (MAD) is the average distance of the values from the mean; a bigger MAD means more spread.

Four friends scored 2, 4, 6 and 8 in a game. The mean is (2 + 4 + 6 + 8) ÷ 4 = 20 ÷ 4 = 5. How far is each score from the mean? 2 is 3 away, 4 is 1 away, 6 is 1 away, 8 is 3 away. These distances are ; distance is never negative, so ignore the sign. Average the distances: (3 + 1 + 1 + 3) ÷ 4 = 8 ÷ 4 = 2. The , or MAD, is 2. On average, a score sits 2 points from the mean.

Try 10, 12, 14, 20. Mean: 56 ÷ 4 = 14. Distances from 14: 4, 2, 0, 6. MAD: (4 + 2 + 0 + 6) ÷ 4 = 12 ÷ 4 = 3. A data set where every value is the same, like 6, 6, 6, 6, has MAD 0, because nothing deviates. The larger the MAD, the more spread out the values are around their mean.

Common mistakes: using signed differences, which always add to zero and give a MAD of 0 by accident; dividing by the wrong count; and measuring from the median instead of the mean. MAD goes with the mean, the way IQR goes with the median. When you report a mean as the center, report MAD as the spread. When you report a median, report the IQR.

Words to know
mean absolute deviation
the average distance of the data values from the mean
absolute deviation
the distance of one value from the mean, with no sign
Check yourself

1. What is the MAD of 1, 3, 5, 7, 9?

2. What is the MAD of the two values 3 and 7?

3. Two classes have the same mean score of 80. Class A has MAD 2 and Class B has MAD 12. What does this mean?

Section 4

Shape and Comparison

37.10

The Shape of a Distribution

Main ideaDescribe a distribution by its shape, symmetric or skewed, its peaks, clusters, gaps and outliers, before you pick a measure of center.

A is the whole pattern of the data: which values appear and how often. When you look at a dot plot or histogram, describe its shape first. If the left and right sides are close to mirror images, the shape is . Heights of students in one grade are usually close to symmetric, with a hump in the middle and fewer very short or very tall students.

If one side stretches out much farther than the other, the shape is . Number of pets is skewed to the right: most students have 0, 1 or 2, and a few have 5 or 6, so the tail points right toward the large values. In a right-skewed set, the mean is usually larger than the median, because the tail pulls the mean. In a left-skewed set, the mean is usually smaller. Also look for peaks. Two peaks might mean two different groups are mixed together, like the heights of sixth graders and eighth graders on one plot.

Finally, look for clusters, gaps and outliers. A cluster is a bunch of values close together. A gap is a stretch of the number line with no values. An outlier sits far from the rest. A complete description of homework minutes might be: "The data are skewed right, with a cluster from 20 to 40 minutes. There is a gap from 60 to 74 and one outlier at 120 minutes." A common mistake is to describe the bars instead of the data, like calling a histogram "tall." Say where the values pile up and where they thin out.

Words to know
distribution
the pattern of which values appear in the data and how often
symmetric
a shape whose left and right sides are close to mirror images
skewed
a shape with one tail stretched much farther than the other
Check yourself

1. A dot plot of family sizes has most dots at 3 and 4, a few at 5, 6 and 7, and one at 12. How would you describe the shape?

2. In a distribution that is skewed right, which is usually true?

3. A histogram of ages at a family reunion has one peak near 10 and another near 40. What is the most likely explanation?

37.11

Comparing Two Groups

Main ideaTo compare two groups, compare their centers, then judge the difference against the spread: a gap of about 2 MADs or more matters; a gap much smaller than the MAD may not.

Two seventh-grade classes measure their heights. Class A has a mean of 152 cm with a MAD of 4 cm. Class B has a mean of 160 cm with a MAD of 4 cm. The is 160 − 152 = 8 cm. How big is that? Compare it to the spread: 8 ÷ 4 = 2, so the difference is 2 MADs. When the gap between centers is about 2 or more MADs, the groups clearly sit in different places; most students in B are taller than most in A.

Now compare two classes’ quiz scores. Class C has mean 78 with MAD 8; Class D has mean 80 with MAD 8. The difference is 2 points, only 2 ÷ 8 = 0.25 of a MAD. The dot plots would almost completely. You should not conclude that Class D does better. A small difference in the center is not meaningful when the spread is large.

You can do the same with box plots, using medians and IQRs. Draw both box plots on one number line. If the boxes barely overlap, the groups differ. If one box sits mostly inside the other, they are similar. A common mistake is to compare only the maximum values: "Class A has the tallest student, so Class A is taller." One student is not a group. Compare centers, then spreads, then shapes.

Words to know
difference in means
one group's mean minus the other's; judge it by how many MADs it is
overlap
how much of the number line two groups' values share
Check yourself

1. Group X: mean 40, MAD 5. Group Y: mean 55, MAD 5. How many MADs apart are the means?

2. Two groups have means of 62 and 64, and both have a MAD of 10. What should you conclude?

3. Which is the best evidence that Class B is generally taller than Class A?

Chapter review

Describing Data

0 / 8

1. What is the mean of 6, 9, 12 and 13?

2. What is the median of 15, 3, 8, 11, 20, 4?

3. What is the range of 21, 14, 30, 9, 17?

4. A histogram uses intervals 0–9, 10–19 and 20–29 with counts 4, 10 and 6. How many values are 10 or more?

5. Which is a statistical question?

6. What is the IQR of 1, 3, 5, 7, 9, 11?

7. What is the MAD of 5, 7, 9, 11?

8. Data: 20, 22, 23, 24, 25, 90. Which is the best measure of a typical value, and why?

Chapter

Probability and Sampling

Probability
Big questionHow can we make reliable claims about chance and about a whole group when we can only see part of it?
The story

One Hundred Flips

A student flips a coin 100 times, gets 56 heads, and has to decide whether the coin is cheating.

Devon's little brother lost four coin flips in a row and announced that the quarter was rigged. Devon did not think so, but saying "it's fair" was not going to settle anything. So Devon sat at the kitchen table with the quarter and a sheet of paper and flipped it 100 times, making a tally mark for each result. The final count was 56 heads and 44 tails. His brother pointed at the paper. "See? More heads. It's not fair."

Devon thought about it. A fair coin has a probability of 1/2 for heads, so in 100 flips you expect about 50 heads. But expect is not the same as guarantee. Devon looked back at the tally marks. The first 10 flips had 7 heads. The next 10 had only 4. The counts wandered up and down the whole time; 56 was just where they happened to land at flip number 100.

The next day Devon brought the question to math class, and the teacher turned it into an experiment. Twenty-five students each flipped a coin 100 times. Their head counts ranged from 41 to 59, and most sat between 45 and 55. Nobody claimed those coins were rigged. Devon's 56 was on the high side but well inside what fair coins had produced. When the class added everything up, 2,500 flips had produced 1,261 heads: a fraction of 0.5044, remarkably close to one half.

That is the heart of this chapter. Probability is a long-run number. Any short run can wander, and the wandering shrinks as the number of trials grows. The same idea explains why a random sample of 25 students can tell you something true about a whole school, and why a bigger sample tells you more. Devon's brother was not convinced, so Devon offered a deal: 1,000 more flips, and the loser does the dishes.

Talk about itIf Devon had gotten 80 heads out of 100 flips, would you still believe the coin is fair? What would you do next to find out?
Section 1

What Probability Means

38.1

A Number From 0 to 1

Main ideaProbability is a number from 0 to 1 that tells how likely an event is: 0 means impossible, 1 means certain, and 1/2 means as likely as not.

Will it snow in Chicago in January? Very likely. Will it snow in July? Almost impossible. turns words like "likely" into numbers. The probability of an is a number from 0 to 1. A probability of 0 means the event cannot happen. A probability of 1 means it is certain. A probability of 1/2, or 0.5, means the event is as likely to happen as not, like a fair coin landing heads.

Probabilities can be written as fractions, decimals or percents. 1/4 = 0.25 = 25% all describe the same chance. Events near 0, like 0.05, are unlikely. Events near 1, like 0.9, are . Rolling a 7 on a regular six-sided die has probability 0, since no face shows 7. Rolling a number less than 7 has probability 1, since every face does.

A common mistake is to write a probability bigger than 1 or less than 0. If you get 7/6 or 120%, something went wrong, usually a count on top that is too large. Another mistake is to think "unlikely" means "will not happen." An event with probability 0.1 happens about 1 time in 10 in the long run. It still happens.

Words to know
probability
a number from 0 to 1 that tells how likely an event is
event
something that may or may not happen, like rolling an even number
likely
having a probability closer to 1 than to 0
Check yourself

1. Which probability describes an event that is unlikely but possible?

2. A student says the probability of rain tomorrow is 1.3. What is wrong?

3. What is the probability of rolling a number less than 7 on a six-sided die?

38.2

Theoretical Probability

Main ideaWhen all outcomes are equally likely, probability = number of favorable outcomes ÷ total number of outcomes.

A spinner has 8 equal sections; 3 are red and 5 are blue. Each section is an , and because the sections are the same size, the outcomes are . The of red is the number of red sections divided by the total: 3/8 = 0.375. You do not need to spin at all. This is what should happen in the long run.

A regular die has six equally likely outcomes: 1, 2, 3, 4, 5, 6. The probability of an even number is 3/6 = 1/2, because 2, 4 and 6 are the three favorable outcomes. A bag holds 4 red and 6 blue marbles. P(red) = 4/10 = 2/5 = 0.4. P(blue) = 6/10 = 3/5 = 0.6. The two add to 1, because a marble must be red or blue. The probability that an event does not happen is 1 minus the probability that it does: P(not red) = 1 − 0.4 = 0.6.

You can predict counts too. If you roll a die 60 times, you expect a 3 about 1/6 of the time: 60 × 1/6 = 10 times. Expect, not guarantee. Common mistakes: treating outcomes as equally likely when they are not (a spinner with one big section and three small ones), and putting the wrong number on top. P(red) uses the red count on top and the total, 10, on the bottom, not the blue count.

Words to know
outcome
one possible result, like one section of a spinner
theoretical probability
favorable outcomes divided by total outcomes, when all outcomes are equally likely
equally likely
outcomes that each have the same chance of happening
Check yourself

1. What is the probability of rolling a number greater than 4 on a fair die?

2. A bag holds 5 green, 3 yellow and 2 white marbles. What is P(not green)?

3. A spinner has 4 equal sections numbered 1 to 4. In 80 spins, about how many times would you expect it to land on 2?

38.3

Experimental Probability

Main ideaExperimental probability = times the event happened ÷ number of trials; with more trials it usually settles near the theoretical probability.

Devon’s class spins the 3-red, 5-blue spinner 50 times and records 18 reds. The of red is the number of reds divided by the number of : 18/50 = 0.36. The theoretical probability was 3/8 = 0.375. They are close but not identical. That is normal. Fifty spins are not enough to land exactly on 0.375; in fact, 0.375 × 50 = 18.75 is not even a whole number.

The more trials you run, the closer the experimental probability tends to get to the theoretical one. Ten flips of a coin might give 7 heads, an experimental probability of 0.7. A hundred flips gave 56 heads, or 0.56. Devon’s class combined 2,500 flips and got 1,261 heads, which is 1,261 ÷ 2,500 = 0.5044. Each bigger set of trials sits closer to 0.5.

Experimental probability is the only choice when you cannot count equally likely outcomes. What is the chance a thumbtack lands point up? A tack is not symmetric, so you drop it 200 times and count. If it lands point up 130 times, the experimental probability is 130/200 = 0.65. Common mistakes: dividing by the number of successes instead of the number of trials, and trusting a tiny experiment. Three flips that all land heads do not mean P(heads) = 1.

Words to know
experimental probability
times an event happened divided by the number of trials
trial
one run of an experiment, like one spin or one flip
Check yourself

1. A coin is flipped 40 times and lands heads 24 times. What is the experimental probability of heads?

2. A marble is drawn from a bag and put back 200 times. Red comes up 30 times. What is the experimental probability of red?

3. Which experiment gives the best estimate of a spinner's true probability of landing on green?

Section 2

Compound Events

38.4

Listing Every Outcome

Main ideaFor a compound event, list the sample space in an organized way; then probability is favorable outcomes over total outcomes.

Flip two coins. What are the possible results? Students often say three: two heads, two tails, or one of each. But list them carefully, first coin then second: HH, HT, TH, TT. There are four outcomes, and they are equally likely. The full list is the . "One head and one tail" happens in two of the four, HT and TH, so its probability is 2/4 = 1/2, not 1/3. This is why the organized list matters.

Now flip a coin and roll a die. For each of the 2 coin results there are 6 die results, so the sample space has 2 × 6 = 12 outcomes. They are H1, H2, H3, H4, H5, H6, T1, T2, T3, T4, T5, T6. P(heads and an even number) = 3/12 = 1/4, since H2, H4 and H6 are favorable. A quick check: multiply the number of outcomes for each part. That product is the size of the sample space for a .

A table works well for two dice. The rows are the first die (1 to 6) and the columns are the second (1 to 6), giving 6 × 6 = 36 cells. Count the cells whose sum is 7: (1,6), (2,5), (3,4), (4,3), (5,2), (6,1). That is 6, so P(sum of 7) = 6/36 = 1/6. A sum of 2 appears in only one cell, (1,1), so P(sum of 2) = 1/36. Common mistake: treating (2,5) and (5,2) as one outcome. They are different rolls.

Words to know
sample space
the organized list of every possible outcome
compound event
an event made of two or more parts, like a coin flip and a die roll together
Check yourself

1. A lunch has 3 sandwich choices and 2 drink choices. How many different lunches are possible?

2. Two fair coins are flipped. What is the probability of two heads?

3. Two dice are rolled. What is P(sum of 2)?

38.5

Tree Diagrams

Main ideaA tree diagram branches once for each choice; each path from start to tip is one outcome, and the number of tips is the product of the branch counts.

A cafe sells sandwiches with 2 breads (white, wheat) and 3 fillings (turkey, cheese, veggie). A starts with 2 branches for bread. From each bread, draw 3 branches for filling. Count the ends: 2 × 3 = 6 sandwiches: white-turkey, white-cheese, white-veggie, wheat-turkey, wheat-cheese, wheat-veggie. Each path from the start to a tip is one outcome.

Add a third choice, 2 drinks (water, juice). From each of the 6 sandwich tips draw 2 more branches. Now there are 2 × 3 × 2 = 12 complete meals. If a customer picks each choice at random, every path is equally likely, so the probability of any one meal, such as wheat-cheese-juice, is 1/12. The probability of "a wheat sandwich with juice" counts 3 paths, one for each filling, so it is 3/12 = 1/4.

Tree diagrams shine when the are not the same at every step. Suppose a game: flip a coin; if heads, you roll a die; if tails, you spin a 4-section spinner. The tree has 6 tips under heads and 4 under tails, 10 total, but the tips are not equally likely, because heads and tails each get half the probability. Common mistake: multiplying counts when the choices depend on each other. Draw the tree and count paths instead of assuming.

Words to know
tree diagram
a drawing that branches once for each choice, so every path is one outcome
branch
one line in a tree diagram, standing for one option at that step
Check yourself

1. A tree diagram has 4 branches for color, then 3 for size, then 2 for sleeve length. How many outcomes are there?

2. In the cafe with 2 breads, 3 fillings and 2 drinks, what is the probability of choosing wheat-cheese-juice at random?

3. Why draw a tree diagram instead of only multiplying?

38.6

Probability of Compound Events

Main ideaFind the probability of a compound event by counting favorable outcomes in the sample space, or by multiplying the probabilities of independent parts.

Spin a 4-section spinner (1, 2, 3, 4) and flip a coin. The sample space has 4 × 2 = 8 equally likely outcomes. What is P(3 and heads)? Only one outcome, 3H, so 1/8. What is P(even number and tails)? Two outcomes, 2T and 4T, so 2/8 = 1/4. Counting always works when outcomes are equally likely.

There is a shortcut when the parts do not affect each other, which we call . The spinner does not change the coin. So P(3 and heads) = P(3) × P(heads) = 1/4 × 1/2 = 1/8, the same answer. For two dice, P(both show 6) = 1/6 × 1/6 = 1/36. For three coin flips, P(all heads) = 1/2 × 1/2 × 1/2 = 1/8.

"At least one" events are easier through the opposite, called the . P(at least one head in two flips) = 1 − P(no heads) = 1 − P(TT) = 1 − 1/4 = 3/4. Check by listing: HH, HT, TH are three of the four outcomes. One common mistake is adding when you should multiply. For "3 and heads," 1/4 + 1/2 = 3/4 is far too large. Another is multiplying when the events are not independent, like drawing two marbles without putting the first back.

Words to know
independent events
events where one result does not change the chances of the other
complement
the event not happening; its probability is 1 minus the event's probability
Check yourself

1. Three fair coins are flipped. What is P(all three land heads)?

2. A spinner with 3 equal colors is spun twice. What is P(red both times)?

3. Two fair coins are flipped. What is P(at least one head)?

Section 3

Simulations and Samples

38.7

Simulating Chance

Main ideaA simulation uses a random tool matched to the real probabilities, run many times, to estimate a probability that is hard to compute.

The forecast says a 30% chance of rain each day this week. What is the probability of rain on at least 2 of the 5 school days? Computing this by hand is messy. Instead, run a . Use 0 to 9. Let 0, 1 and 2 stand for rain; that is 3 out of 10 digits, matching 30%. Digits 3 to 9 mean no rain. Take 5 random digits for one week: 8 1 4 2 7 means rain on Tuesday and Thursday, 2 rainy days.

One week is one trial. Run 50 trials, count how many have 2 or more rainy days, and divide by 50. If 24 trials qualify, the estimate is 24/50 = 0.48. The key rule: the tool must match the real probability. For a basketball player who makes 80% of free throws, let digits 0 to 7 mean make and 8, 9 mean miss. For a 1/6 chance, roll a die and pick one face. For 1/4, use a spinner with four equal sections, or two coin flips where HH counts.

Common mistakes: matching the wrong number of digits, since 0 to 3 gives 40%, not 30%. Reusing the same digits every trial. Stopping after a handful of trials. A simulation is an experiment, so its answer is an experimental probability. It gets more trustworthy with more trials, but it is still an estimate.

Words to know
simulation
an experiment with a random tool that stands in for a real situation
random digits
digits 0 to 9 produced so that each is equally likely, used to model chance
Check yourself

1. You want to simulate a 40% chance with random digits 0–9. Which assignment works?

2. A simulation of 5-day weeks runs 40 trials; 14 of them have at least 2 rainy days. What is the estimated probability?

3. Which random tool best simulates a 1/6 chance?

38.8

Random Samples

Main ideaA random sample gives every member of the population an equal chance of being chosen, so the sample is likely to look like the whole.

You want to know how many hours students at your school sleep, but the school has 800 students. Asking all 800 is a census. Asking a smaller group is taking a . The whole group you care about is the . A sample only helps if it looks like the population, and the best way to get that is to choose at random: every student has the same chance of being picked.

Random does not mean careless. Put every student’s name or ID number in a list, then use a random number generator, or draw numbers from a hat, to pick 50. That is a . Compare that to asking the 50 students in the cafeteria line at 7:30 a.m., who are early risers, or asking the basketball team, who are taller than average. Those samples are : some students had a much better chance of being chosen than others, so the sample leans one way.

Random samples still vary. Two random samples of 50 students might give mean sleep of 7.1 and 7.4 hours. Larger samples vary less. The U.S. Census counts every person once every ten years; between counts, the government estimates by sampling. Common mistake: thinking a big biased sample beats a small random one. A million responses to an online poll that only night owls saw still tell you about night owls.

Words to know
population
the whole group you want to learn about
sample
a smaller group taken from the population
random sample
a sample in which every member of the population had an equal chance of being chosen
biased
leaning one way because some members were more likely to be chosen than others
Check yourself

1. Which sample of a school's 800 students is a random sample?

2. A survey about lunch is given only to students who buy lunch. What is the problem?

3. Which will vary least from the true population value?

Section 4

Making Inferences

38.9

From Sample to Population

Main ideaFind the proportion in a random sample, then multiply by the population size to estimate the count in the whole population.

A random sample of 50 students at an 800-student school is asked their favorite lunch. 20 pick pizza. The sample is 20/50 = 0.4, or 40%. If the sample looks like the school, about 40% of all 800 students prefer pizza: 0.4 × 800 = 320 students. This is an : a conclusion about the population drawn from the sample.

Another: a random sample of 40 students shows 12 walk to school, so 12/40 = 0.3. In a district of 1,200 students, estimate 0.3 × 1,200 = 360 walkers. The estimate is not exact. A second sample might show 10 or 14 walkers, giving 300 or 420. Report it as "about 360," and remember that larger samples give tighter estimates.

The same idea works for quality checks. A factory pulls a random sample of 100 light bulbs and finds 3 that fail. The proportion is 3/100 = 0.03. In a shipment of 20,000 bulbs, expect about 0.03 × 20,000 = 600 failures. Common mistakes: multiplying the sample count by the population instead of using the proportion (20 × 800 is nonsense), and making inferences from a biased sample. A sample of the pizza line will always say pizza.

Words to know
proportion
the part divided by the whole, like 20/50 = 0.4
inference
a conclusion about a population based on a sample
Check yourself

1. In a random sample of 50 students, 20 prefer pizza. About how many of the school's 800 students prefer pizza?

2. A random sample of 40 students shows 12 walk to school. In a district of 1,200 students, about how many walk?

3. A scoop of 40 gumballs from a jar of 500 has 14 blue. About how many blue gumballs are in the jar?

38.10

Comparing Two Populations

Main ideaCompare random samples from two populations the way you compare two groups: look at the difference in centers relative to the spread.

Do seventh graders and eighth graders spend different amounts of time on homework? You cannot ask everyone, so take a random sample of 30 from each grade. Seventh-grade sample: mean 45 minutes, MAD 10. Eighth-grade sample: mean 65 minutes, MAD 10. The difference in means is 65 − 45 = 20 minutes, which is 20 ÷ 10 = 2 MADs. That is a large, visible separation, so you can infer that eighth graders in this school generally spend more time.

Change the numbers: seventh-grade mean 52, eighth-grade mean 55, both MAD 12. The difference is 3 minutes, or 3 ÷ 12 = 0.25 MAD. Two random samples from the same population could easily differ by that much. Do not claim a real difference. Say the samples overlap too much to tell. A is one that is large compared with the spread.

Repeating the sampling shows why. If you drew many pairs of samples of 30, the difference in sample means would bounce around by a few minutes even when the populations were identical. That bounce is . A difference must be big compared to that bounce, and compared to the spread within each group, to mean something. Common mistakes: comparing two biased samples, and comparing samples of very different sizes without noticing. Sizes of 30 and 3 do not carry equal weight.

Words to know
meaningful difference
a gap between two centers that is large compared with the spread, usually 2 or more MADs
sampling variability
the natural bouncing around of results from one random sample to the next
Check yourself

1. Sample A: mean 45, MAD 10. Sample B: mean 65, MAD 10. How many MADs apart are the means?

2. Two random samples have means of 52 and 55, each with a MAD of 12. What is the best conclusion?

3. Which change would make a difference between two sample means more convincing?

Chapter review

Probability and Sampling

0 / 8

1. A bag has 3 red and 9 blue marbles. What is P(red)?

2. A fair die is rolled 120 times. About how many times should it show a 5?

3. A coin is flipped 50 times and lands tails 22 times. What is the experimental probability of tails?

4. How many outcomes are in the sample space when you flip a coin and roll a die?

5. Two dice are rolled. What is P(sum of 7)?

6. Which sample of a school is random?

7. In a random sample of 60 students, 15 ride bikes to school. The school has 900 students. About how many ride bikes?

8. A player makes 70% of free throws. Which digit assignment simulates one shot?

Chapter

Two-Variable Data

Statistics
Big questionWhen two things are measured together, how can we tell whether they move together, and what a pattern between them does and does not prove?
The story

Sleep and Scores

Twenty-four students track their sleep for a week, and a scatter plot reveals a pattern no single number could show.

It started as an argument in Ms. Bell's class. Nadia said students who sleep more get better grades. Leo said that was nonsense; he knew someone who slept ten hours a night and failed everything. Ms. Bell did not take sides. She handed out a sheet and asked each of the 24 students to record their hours of sleep every night for a week and then write down their score on Friday's quiz. By Monday the class had 24 pairs of numbers: an average sleep time and a quiz score for every student.

The table was hard to read. Twenty-four rows of two numbers each did not say much. Then Nadia drew two number lines at a right angle, hours of sleep going across and quiz score going up, and put one dot for each student. Ms. Bell called it a scatter plot. The moment the last dot went on, the class could see it: the dots formed a loose cloud that rose from the lower left to the upper right. Less sleep, lower scores. More sleep, higher scores. Not perfectly, but clearly.

Leo pointed at one dot sitting all by itself: nine hours of sleep and a score of 55. "That's me," he said. "I was sick on Friday." The class agreed that his dot did not follow the pattern, and that there was a good reason. Most of the other dots crowded between six and a half and eight hours. Nadia laid a ruler across the middle of the cloud and drew a line. The line rose about 5 points for every extra hour of sleep.

Then Ms. Bell asked the hard question. Does more sleep cause higher scores? Or do students who sleep more also finish their homework earlier, or worry less, or have quieter homes? The dots showed that the two things moved together. They could not, by themselves, show why. The class also realized some questions do not have number answers at all. Did you study with a group, yes or no? Did you pass, yes or no? For questions like those, they would need a different tool: a table with two ways in.

Talk about itWhich is more convincing: one student who slept 10 hours and scored 98, or 24 dots that mostly rise together? Why?
Section 1

Scatter Plots

39.1

Two Measurements Per Point

Main ideaA scatter plot shows two measurements for each person or thing as one point, with one variable on each axis.

Five students report hours of sleep and quiz score: (5, 68), (6, 72), (7, 80), (8, 85), (9, 88). Each pair is one student. A puts sleep on the horizontal axis and score on the vertical axis and marks one point per student. The student who slept 7 hours and scored 80 is the point 7 units to the right and 80 units up. Five students, five points. The data now have two , sleep and score, and the plot lets you see them together.

Choose axes carefully. The variable you think might explain the other usually goes across (sleep), and the one being explained goes up (score). Pick scales that fit the data: sleep from 4 to 10 hours, score from 50 to 100. A scale that starts at 0 on both axes would squeeze the points into one corner. A plot with 24 students has 24 points, one for each pair of measurements.

Watch for three mistakes. Drawing two separate dot plots loses the pairing. Swapping the coordinates puts (5, 68) at 68 across and 5 up. Connecting the points with lines suggests a path. Points in a scatter plot are separate people, not a path. Leave them as dots and look for a pattern in the cloud.

Words to know
scatter plot
a graph with one point for each pair of measurements, one variable on each axis
variable
something measured that can take different values, like hours of sleep
Check yourself

1. A student slept 6 hours and scored 72. Where is that point on a scatter plot with sleep across and score up?

2. How many points does a scatter plot of 24 students' sleep and scores have?

3. Why should you not connect the points in a scatter plot?

39.2

Reading a Scatter Plot

Main ideaRead a scatter plot by asking: as one variable increases, what happens to the other, and how tightly do the points follow that trend?

Look at the sleep-and-score plot. Start at the left: students with little sleep have lower scores. Move right: scores climb. The cloud of points rises from lower left to upper right. This is an between sleep and score: knowing one tells you something about the other. Now ask how tight the pattern is. If the points hug a line, the association is strong. If they spread widely, it is weak.

You can read individual facts too. Find the student with 8 hours of sleep: the point at 8 across sits at 85 up, so that student scored 85. Find the highest scorer: the top point, at 88, belongs to the 9-hour sleeper. Find how many students slept under 7 hours: count the points left of 7. In the five-student data, that is 2.

Be careful about what an association does not prove. More sleep goes with higher scores, but sleep might not be the cause. Students who sleep more may also finish homework earlier, or worry less. A scatter plot shows that two variables follow a together; it cannot by itself show that one causes the other. Common mistake: reading the plot as if the points were in time order. The leftmost point is the least sleep, not the first day.

Words to know
association
a pattern where knowing one variable tells you something about the other
trend
the general direction the points follow across the plot
Check yourself

1. On a scatter plot of sleep (across) and score (up), the point at 8 across is 85 up. What does it mean?

2. Points on a scatter plot rise to the right but are widely spread. Which describes it?

3. Cities with more ice cream sales also have more sunburns. What can you conclude?

Section 2

Patterns in the Points

39.3

Positive and Negative Association

Main ideaA positive association means both variables rise together; a negative association means one falls as the other rises; a shapeless cloud shows no association.

Sleep and score rise together: that is a . Now picture a plot of a car’s age (across) and its price (up). Older cars generally cost less, so the points fall from upper left to lower right. That is a . Both can be strong or weak. A third possibility: a plot of shoe size and quiz score shows a shapeless cloud, because the two have nothing to do with each other. That is .

A quick test: slide your hand across the plot from left to right. If the points under your hand move up, the association is positive. If they move down, it is negative. If they stay level or jump around, there is none. Words matter: negative does not mean bad. A negative association between minutes of exercise and resting heart rate is good news.

Examples to sort. Hours of TV and hours of homework: likely negative. Height and arm span: positive and strong. Day of the month and high temperature in Chicago: no clear association within one month. Common mistake: calling a plot negative because the values are low, or positive because the values are high. Direction comes from the slope of the cloud, not from where it sits.

Words to know
positive association
as one variable increases, the other tends to increase
negative association
as one variable increases, the other tends to decrease
no association
the points show no trend; one variable tells you nothing about the other
Check yourself

1. As outdoor temperature rises, heating bills fall. Which kind of association is this?

2. A scatter plot of height and arm span shows points in a tight band rising to the right. Which describes it?

3. Which pair would most likely show no association?

39.4

Clusters, Outliers and Curves

Main ideaBesides direction, look for clusters, outliers and whether the pattern is a line or a curve.

On the sleep plot, most points bunch between 6.5 and 8 hours. That bunch is a . A cluster can hint at a group: maybe most students have a similar bedtime. Two clusters might mean two kinds of students, like those with after-school jobs and those without. Name clusters by where they sit on both axes: "a cluster around 7 hours and 80 points."

One student slept 9 hours but scored 55. That point sits far below the rising band. It is an in the scatter plot: it does not follow the pattern the others follow. An outlier can be a data-entry error, a special case (Leo was sick), or a real surprise. Do not delete it silently. Note it, find out why if you can, and describe the pattern "with one outlier at (9, 55)."

Not every pattern is a line. Plot the height of a thrown ball against time and the points rise, then fall, in an arch. That is a association. Plot a bacteria count against hours and the points bend upward faster and faster. When the points bend, a straight line will fit badly, and you should say the association is nonlinear. Common mistake: forcing a line through a curved pattern and reporting its slope as if it held for all values.

Words to know
cluster
a group of points bunched close together on the plot
outlier
a point that sits far from the pattern the other points follow
nonlinear
a pattern that bends instead of following a straight line
Check yourself

1. A point at (9, 55) sits far below a rising band of points. What is it?

2. Points rise, then fall, in an arch. Which describes the pattern?

3. What should you do with an outlier in a scatter plot?

Section 3

Fitting a Line

39.5

A Line by Eye

Main ideaWhen points show a linear pattern, draw a straight line through the middle of the cloud with about as many points above as below, then use it to predict.

The five sleep points (5, 68), (6, 72), (7, 80), (8, 85), (9, 88) rise in a nearly straight band. Lay a ruler on the plot and draw a . It should follow the direction of the cloud and pass close to as many points as you can. A good line has about as many points above it as below it. The points above should not all sit at one end. The line does not have to touch any point.

One easy line goes through (5, 68) and (9, 88). Its steepness is (88 − 68) ÷ (9 − 5) = 20 ÷ 4 = 5. Check the middle: at 7 hours, this line gives 68 + 2 × 5 = 78, close to the actual 80. At 8 hours it gives 83, close to 85. It sits a little low in the middle, so a slightly better line is score = 5 × hours + 43, which gives 68, 73, 78, 83 and 88 at 5 through 9 hours. It is above the data twice, below twice, and exact twice.

Use the line to make a . For 6.5 hours, the line gives 5 × 6.5 + 43 = 32.5 + 43 = 75.5, so predict about 75 or 76. Predictions inside the range of the data are reasonable. Predictions far outside, like 20 hours of sleep, are not; the line would say 143, which is not even a possible score. Common mistakes: drawing the line from the lowest point to the highest point no matter where the rest sit, and running it through two outliers.

Words to know
line of fit
a straight line drawn through the middle of a linear cloud of points to summarize the trend
prediction
a value read from the line of fit for an input you did not measure
Check yourself

1. Which is a sign of a good line of fit?

2. Using score = 5 × hours + 43, what score does the line predict for 6.5 hours of sleep?

3. Why is predicting the score for 20 hours of sleep from this line unreasonable?

39.6

Slope of a Fitted Line

Main ideaThe slope of a fitted line tells how much the vertical variable changes, on average, for each 1-unit increase in the horizontal variable.

A line of fit through plant-growth data passes through (1, 4) and (4, 13): after 1 week the plant was 4 cm, after 4 weeks it was 13 cm. is rise over run: (13 − 4) ÷ (4 − 1) = 9 ÷ 3 = 3. The slope is 3 cm per week. It means: for each extra week, the line predicts about 3 cm more height. Always attach both units, the vertical unit per one horizontal unit.

Slope can be negative. A line through used-car data passes through (1, 20) and (5, 12), with age in years and price in thousands of dollars. Slope = (12 − 20) ÷ (5 − 1) = −8 ÷ 4 = −2. Each extra year of age goes with about $2,000 less in price. The sign tells the direction of the association, and the size tells how steep it is.

In the sleep line, score = 5 × hours + 43, the slope is 5 points per hour. Each extra hour of sleep goes with about 5 more quiz points, on average. "On average" matters; individual students sit above or below the line. Common mistakes: dividing run by rise, which gives 1/3 instead of 3. Subtracting in different orders on top and bottom, which gives −3. Forgetting that slope is a , not a total.

Words to know
slope
rise divided by run: how much the line goes up for each 1 unit it goes right
rate of change
how much one quantity changes for each unit of another, like 3 cm per week
Check yourself

1. A fitted line passes through (0, 5) and (4, 17). What is its slope?

2. A line passes through (1, 20) and (5, 12). What is its slope?

3. A line of fit for car age (years) and price (thousands of dollars) has slope −2. What does it mean?

39.7

Intercept and Predictions

Main ideaThe y-intercept is the line's value when the horizontal variable is 0; it means something only if 0 is inside or near the data.

In the line score = 5 × hours + 43, set hours to 0 and the score is 43. That number is the , where the line crosses the vertical axis. Does it mean a student with no sleep would score 43? No. No student in the data slept 0 hours; the least was 5. The intercept here is a mathematical anchor, not a fact about zero-sleep students.

Sometimes the intercept does mean something. A line for taxi fares, fare = 2.5 × miles + 3, has intercept 3. At 0 miles the fare is $3, which is the flat fee for getting in the cab. Here 0 miles is a real starting point. To predict, plug in: 8 miles gives 2.5 × 8 + 3 = 20 + 3 = 23 dollars. A line with slope −2 and intercept 50, y = −2x + 50, gives y = −2 × 15 + 50 = −30 + 50 = 20 when x = 15.

Always ask two questions about a prediction. Is the input inside the ? Does the answer make sense? A line for ice cream sales, cones = 4 × temp − 100, predicts 4 × 80 − 100 = 220 cones at 80°F, which is sensible. At 20°F it predicts 4 × 20 − 100 = −20 cones, which is impossible; that temperature is far outside the summer data. Common mistake: reading the intercept as "the starting value" in every situation.

Words to know
y-intercept
the value of the line where the horizontal variable is 0
range of the data
the stretch from the smallest to the largest measured input; predictions are safest inside it
Check yourself

1. Using y = 2.5x + 10, what is y when x = 8?

2. A line has slope −2 and y-intercept 50. What is y when x = 15?

3. For the taxi line fare = 2.5 × miles + 3, what does the intercept 3 mean?

Section 4

Categories in Tables

39.8

Two-Way Tables

Main ideaA two-way table counts people by two categories at once; row and column totals let you check the count and answer how-many questions.

Some data are categories, not numbers: grade level, cat or dog, yes or no. This is . A organizes counts by two categories at once. Survey 100 students: do you prefer cats or dogs? Rows are grade (7th, 8th), columns are the pet. Seventh grade: 20 cats, 30 dogs. Eighth grade: 15 cats, 35 dogs. Add each row: 7th total 50, 8th total 50. Add each column: cats 35, dogs 65. The grand total is 100 either way, which confirms the table.

Each answers one how-many question. How many eighth graders prefer dogs? 35. How many students prefer cats? Read the column total: 35. How many seventh graders were surveyed? Read the row total: 50. Given some cells, you can find missing ones by subtraction. If 80 students are surveyed, 50 play a sport, and 18 of those also play an instrument, then 50 − 18 = 32 play a sport but no instrument.

Common mistakes: mixing up row and column when reading a cell, and double-counting when categories overlap. Each person belongs in exactly one cell. Also, do not use a two-way table when a variable is numerical, like exact height; use a scatter plot for two numerical variables. Two categories, two-way table; two numbers, scatter plot.

Words to know
categorical data
data sorted into groups with names, like yes/no or cat/dog, instead of numbers
two-way table
a table that counts how many fall into each pair of categories
cell
one box in the table, holding the count for one pair of categories
Check yourself

1. In the pet table, how many students in all prefer cats?

2. 80 students are surveyed: 50 play a sport, and 18 of those also play an instrument. How many play a sport but no instrument?

3. Which pair of variables belongs in a two-way table?

39.9

Relative Frequencies

Main ideaDivide a cell by its row total (or column total) to get a relative frequency, which lets you compare groups of different sizes.

Counts alone can mislead when groups differ in size. Use : a count divided by a total. In the pet table, 30 of the 50 seventh graders prefer dogs: 30/50 = 0.6, or 60%. For eighth graders, 35/50 = 0.7, or 70%. Dividing by the gives row relative frequencies. Because both grades had 50 students, the counts and the percents tell the same story here.

Now a table with unequal groups. 80 students: 50 play a sport and 30 do not. Of the sport players, 18 play an instrument. Of the non-players, 12 do. Raw counts, 18 versus 12, make it look like sport players play instruments more. But 18/50 = 0.36, or 36%, and 12/30 = 0.4, or 40%. The non-players actually have the higher share. Compare percents, not counts, when the groups are different sizes.

You can also divide by the grand total to get the share of everyone: 18/80 = 0.225, so 22.5% of all students play both a sport and an instrument. Choose the total that matches the question. "What percent of sport players" means divide by the sport row. "What percent of all students" means divide by 80. Common mistake: dividing by the wrong total, or mixing a row percent with a column percent in one comparison.

Words to know
relative frequency
a count divided by a total, written as a fraction, decimal or percent
row total
the sum of all the cells in one row of the table
Check yourself

1. Of 50 sport players, 18 play an instrument. What is the row relative frequency?

2. 12 of the 30 non-players play an instrument. What percent is that?

3. Why compare relative frequencies instead of counts?

39.10

Association in a Table

Main ideaTwo categories are associated when the row percents differ noticeably between rows; similar percents mean little or no association.

Is there an association between playing a sport and playing an instrument? Compare the row percents: 36% of sport players play an instrument, and 40% of non-players do. Those are close, only 4 apart. Knowing whether a student plays a sport tells you little about instruments. That is weak or no association.

Now a different survey of 200 students: studied with a group or alone, and passed or did not. Group: 60 passed, 20 did not, total 80. Alone: 60 passed, 60 did not, total 120. Row percents: group 60/80 = 0.75, or 75% passed. Alone 60/120 = 0.5, or 50% passed. Notice the raw pass counts are the same, 60 and 60, but the percents differ by 25 percentage points. That is a clear : students who studied in groups passed at a higher rate.

Association again is not proof of cause. Group studiers might already be stronger students. Also, a difference of a couple of percentage points in small groups is not convincing; 5 of 10 versus 6 of 10 is one student. Common mistakes: comparing counts across rows of different size, and comparing a row percent to a column percent. Pick one direction, compute both rows, then compare.

Words to know
association
in a table, a difference in the row percents that shows the two categories are linked
percentage point
the difference between two percents, like 75% minus 50% = 25 percentage points
Check yourself

1. Group: 60 of 80 passed. Alone: 60 of 120 passed. What percent of each group passed?

2. Two row percents are 36% and 40%. What does this suggest?

3. Which is the fairest way to look for an association in a two-way table?

Chapter review

Two-Variable Data

0 / 8

1. Points on a scatter plot fall from upper left to lower right. Which association is this?

2. What is the slope of a line through (2, 10) and (6, 22)?

3. Using y = 3x + 5, what is y when x = 10?

4. A point sits far from a rising band of points. What is it called?

5. In a two-way table of 100 students, 35 prefer cats and the rest prefer dogs. How many prefer dogs?

6. 24 of 40 seventh graders bring lunch; 30 of 60 eighth graders do. Which grade has the higher share?

7. What is the y-intercept of y = −2x + 50?

8. A line of fit was made from students who slept 5 to 9 hours. Why not use it to predict a score for 20 hours?

Unit wrap-up

Statistics and Probability

Twelve words, twelve meanings

0 / 12

Tap a word, then tap its meaning. A right pair locks in green.

Words
Meanings
Unit test

Fifteen questions across the unit

0 / 15

1. What is the mean of 4, 8, 9 and 15?

2. What is the median of 9, 2, 6, 11, 5?

3. What is the IQR of 3, 5, 7, 9, 11, 13, 15, 17?

4. What is the MAD of 4, 8, 12, 16?

5. Data: 12, 14, 15, 15, 16, 80. Which measure of center best describes a typical value?

6. A bag has 2 red and 6 blue marbles. What is P(red)?

7. A spinner is spun 25 times and lands on green 10 times. What is the experimental probability of green?

8. You flip two coins and roll one die. How many outcomes are in the sample space?

9. Two dice are rolled. What is the probability that both show an even number?

10. Which is a random sample of a school's students?

11. In a random sample of 75 students, 30 prefer soccer. The school has 1,500 students. About how many prefer soccer?

12. What is the slope of a line through (3, 8) and (7, 20)?

13. Using y = 4x − 6, what is y when x = 5?

14. 18 of 60 girls and 12 of 30 boys ride the bus. Which group has the higher share of bus riders?

15. A scatter plot of hours of sleep and quiz score rises to the right. Which describes it?

Spiral review

Five questions from earlier units

0 / 5

1. (Unit 16) What is the distance between (0, 3) and (6, 11)?

2. (Unit 15) What is the equation of the line through (0, −2) and (3, 7)?

3. (Unit 14) Factor 10y + 25 completely.

4. (Unit 13) What is −45 ÷ 5?

5. (Unit 12) How much simple interest does $1,500 earn at 4% per year for 2 years?

Write it

Two classes took the same quiz. Class A: mean 78, MAD 4. Class B: mean 84, MAD 4. Is Class B really doing better, or could the difference be chance? Find how many MADs apart the means are, decide, and explain each step. Then say what else you would want to know before trusting the conclusion.

  • State your decision in the first sentence.
  • Show the subtraction and the division that give the number of MADs.
  • Explain what a difference of that many MADs means about how much the two groups overlap.
  • Say how the classes were chosen and whether that could bias the comparison.
  • Check your arithmetic a second time before you finish.
0 wordsSaved on this device as you type.

Practice rooms

Rooms already on the site that belong to this unit — cards, quizzes, a lab.

For the teacher

Every lesson keeps its own three checks; a lesson is ticked when all three are right. Chapter reviews, the unit test and its spiral review (five questions from earlier units in this band) score on the page. When the site is connected to your sheet, or the link carries ?dest=, each one also has a Send box: the first-try score, the standards, the supports used, the attempt number and the minutes go to your sheet as an IEP data point.

Print this page for a paper copy of the readings, the sources, the words and the questions; the answers print as dashed boxes under each question.

Fact-check notes for this course live in the handoff: quotes marked (paraphrased) were set that way on purpose.