A unit of the course: the story, then chapter by chapter — sections, numbered lessons, a source or the numbers to read, three checks each — a review per chapter, and the wrap-up at the end.
Drawn scene: a lecture hall board with a bar graph and trend line, John Snow's pump and cholera map in silhouette, a magnifying glass and a notebook on the desk
27Unit
Capstone: Argue From Evidence
Science Practices
You have spent years learning what science has found: how cells divide, why the seasons turn, what holds an atom together. This last unit is about something different: how anyone knows any of it, and how you can tell when a claim deserves belief. A doctor in 1854 drew dots on a map and found the cause of cholera before anyone had seen the germ. A scientist in a cornfield saw genes move when the textbooks said they could not. A tiny study of twelve children started a fear that studies of a million children could not confirm. In each case the difference between truth and error was not who was more confident, but who had the better evidence and the more careful reasoning.
The first chapter is about reading evidence: telling correlation from cause, spotting a biased sample, reading a graph whose axis has been trimmed, knowing what a plus-or-minus means, understanding why one paper is a lead and a replicated result is knowledge. It ends with three cases, smoking and cancer, the ozone layer, and vaccines and autism, where careful evidence settled questions that arguments alone never could. The second chapter is about making evidence: asking a question sharp enough to test, designing a fair comparison, collecting data you could defend, building and testing a model or a prototype, and writing and standing behind a claim.
By the end you will plan and carry out an investigation of your own, from a testable question to a public defense. The tools you will use are the same ones that John Snow and Barbara McClintock used, and the same ones that a citizen needs to judge a headline, a product label or a proposal at a city council meeting. Science is not a pile of facts. It is a set of habits for finding out what is true and admitting when you were wrong, and those habits are what this unit is for.
How we figured it out
1620
Francis Bacon argues that knowledge should be built up from careful observation rather than from authority
1660
The Royal Society of London is founded; its motto, Nullius in verba, means take nobody's word for it
1747
James Lind tests six remedies for scurvy on twelve sailors kept under the same conditions
1854
John Snow maps cholera deaths around the Broad Street pump and argues the disease travels in water
1935
Ronald Fisher publishes The Design of Experiments, laying out randomization and control
1948
A British trial of streptomycin for tuberculosis assigns patients to groups at random
1950
Richard Doll and Austin Bradford Hill report that lung cancer patients are almost all smokers
1951
Barbara McClintock presents evidence that genes can move within a chromosome; few believe her
1964
The United States Surgeon General's committee concludes that smoking causes lung cancer
1974
Mario Molina and Sherwood Rowland predict that CFCs will destroy ozone high in the atmosphere
1985
British Antarctic Survey scientists report the ozone hole; nations sign the Montreal Protocol two years later
2010
The Lancet retracts the 1998 paper linking vaccines and autism after its data are shown to be false
59
Chapter
Reading and Judging Evidence
Science Practices
Big questionHow can you tell whether a scientific claim deserves your belief?
The story
The Map and the Pump Handle
In the summer of 1854 a London doctor drew dots on a street map and found the source of a killer that everyone else blamed on bad air.
At the end of August 1854, cholera exploded in Soho, a crowded neighborhood of London. Cholera kills by violent diarrhea and vomiting that drain the body of water; a healthy adult could sicken at breakfast and be dead by night. Within ten days more than five hundred people in a few streets were dead. Families fled. The accepted explanation, held by most doctors and the government, was miasma: a poisonous vapor rising from filth and sewage and drifting through the air.
John Snow, a physician who lived nearby, did not believe it. He had studied earlier outbreaks and suspected that cholera was swallowed, not inhaled, and that it spread through water fouled by the waste of the sick. But suspecting is not showing. So Snow went door to door with the registrar's list of deaths and marked each one on a map of the streets. The dots piled up around one spot: the public water pump on Broad Street.
A pattern on a map is a correlation, and Snow knew a skeptic would say the pump was simply in the middle of a crowded, dirty district. So he hunted for the cases that did not fit. A brewery a few doors from the pump had lost no workers; the men drank the beer they made and never touched the pump. A workhouse with its own private well, housing hundreds of people, was nearly untouched. And a widow who had moved to Hampstead, miles away, died of cholera after having a bottle of Broad Street water delivered because she liked its taste.
On September 7 Snow presented his evidence to the parish officials. The next day they removed the pump handle. The outbreak was already fading, and Snow himself admitted that the handle probably did not end it. What mattered was the argument: a claim, evidence gathered on foot, reasoning that explained both the pattern and the exceptions, and a prediction that could be checked. A local clergyman, Henry Whitehead, set out to prove Snow wrong, gathered his own data, and ended up finding the first case: a sick baby whose washing water had drained into a cesspool beside the well.
The cholera bacterium itself was not confirmed under the microscope until the 1880s, three decades later. Snow never saw the cause. He found it anyway, because he treated a popular belief as a claim to be tested rather than a fact to be accepted. This chapter is about doing what Snow did: reading evidence honestly, asking what else could explain it, and knowing when a claim has earned belief.
Talk about itSnow's map alone could not prove the pump was the cause. Which of his other pieces of evidence do you think did the most to rule out the miasma explanation, and why?
Section 1
Claims, Evidence and Reasoning
59.1
What a Claim Needs
Main ideaA scientific claim is an answer to a question that is backed by evidence and linked to that evidence by clear reasoning.
A fish kill at a park pond gets people talking. One neighbor says a factory dumped poison; another says the heat wave did it. Each of these is a : a statement that answers a question about what happened or why. A claim is not right or wrong just because someone says it with confidence. It becomes worth believing only when it is tied to , meaning observations and measurements that anyone could check.
Evidence by itself does not settle anything. A dissolved-oxygen reading of 1.5 milligrams per liter is just a number until someone explains why it matters. That link is the : the scientific principle that connects the evidence to the claim. Warm water holds less oxygen, decaying algae use up what is left, and most fish suffocate below about 2 milligrams per liter. Now the number supports the heat-wave claim, and the poison claim needs its own evidence.
Good scientific claims share one more feature. They are : there is some observation that would show them to be wrong. ’The fish died because of low oxygen’ can be checked against oxygen readings, fish species and the timing of the deaths. ’The fish died because the pond was unlucky’ cannot be checked at all. A claim that no possible evidence could ever count against is not a scientific claim, however attractive it sounds.
Throughout this unit you will use this claim, evidence and reasoning pattern to take arguments apart and to build your own. The questions to ask are always the same. What exactly is being claimed? What evidence is offered, and could I look at it myself? What reasoning links the two, and does that reasoning hold up? And what would it take to show the claim is wrong?
Words to know
claim
a statement that answers a scientific question about what happens or why it happens
evidence
observations, measurements and data that can be checked and that bear on a claim
reasoning
the scientific principle or logic that explains why the evidence supports the claim
falsifiable
able to be shown wrong by some possible observation or test
Check yourself
1. In the fish-kill example, which of these is the reasoning rather than the claim or the evidence?
Why: Reasoning is the principle that links the measurement to the claim; the oxygen reading is evidence and the heat-wave statement is the claim.
2. Why is 'the pond was unlucky' not a scientific claim?
Why: A scientific claim must be falsifiable; luck explains everything and therefore can be tested by nothing.
3. A student says, 'My claim must be true because I am very sure of it.' What is missing?
Why: Confidence is not evidence; a claim earns belief through checkable observations and sound reasoning, not through certainty or authority.
59.2
Correlation Is Not Causation
Main ideaTwo things that rise and fall together may be linked by a hidden third factor, by chance, or by cause running the other way.
In most American cities, ice cream sales and drowning deaths rise and fall together across the year. Nobody believes that ice cream causes drowning. Both rise in summer, when people swim more and eat more cold treats. Summer is a : a hidden factor that drives both things and creates a between them. A correlation means two quantities move together. means one actually produces the other.
There are three main ways a correlation can fool you. First, a confounder like the season may cause both. Second, the cause may run backward: hospitals with more patients in intensive care have more deaths, but the care is not killing them; the sickest patients are sent there. Third, with enough pairs of numbers, some will line up by pure chance. If you compare a hundred unrelated things, a few will match closely by accident, and a careless researcher will report those few.
So how did scientists ever conclude that smoking causes cancer, when no one can run an experiment forcing people to smoke? In 1965 the British statistician Austin Bradford Hill laid out a set of questions. Is the association strong? Does more of the cause bring more of the effect? Does the cause come before the effect in time? Does the link show up in many different studies and populations? Is there a known mechanism? No single answer proves causation, but together they build a case that a chance correlation cannot match.
The strongest tool is still the controlled experiment, where you change one factor on purpose and hold the rest steady. When that is impossible or unethical, scientists lean on the Bradford Hill questions, on natural experiments, and on looking hard for confounders. When you read that ’a study links’ one thing to another, your first question should be: what else could explain this?
Words to know
correlation
a pattern in which two quantities tend to rise or fall together
causation
a relationship in which one thing actually produces a change in another
confounding variable
a hidden factor that affects both quantities and creates a misleading correlation between them
Check yourself
1. Ice cream sales and drownings rise together each summer. What best explains the correlation?
Why: Season is a confounding variable that drives both quantities, so they move together without either causing the other.
2. Which of these is an example of the cause running backward?
Why: Being near death causes a patient to be in intensive care, not the other way around; that is reverse causation.
3. Why did Bradford Hill propose a set of questions instead of a single test for causation?
Why: When you cannot run a controlled experiment, strength, dose-response, timing, consistency and mechanism together make a causal case that a coincidence could not produce.
59.3
Samples, Bias and Controls
Main ideaA conclusion is only as good as the sample it rests on, and a fair sample needs a comparison group chosen the same way.
In 1936 the magazine Literary Digest mailed about ten million ballots asking Americans how they would vote for president. More than two million came back, an enormous , and the magazine confidently predicted that Alf Landon would beat Franklin Roosevelt. Roosevelt won in one of the biggest landslides in history. The mailing list had been built from telephone directories and car registrations, and in the middle of the Depression the people who owned phones and cars were richer than average and leaned against Roosevelt. A young pollster named George Gallup, using a far smaller but carefully chosen sample, got it right.
The lesson is that size does not cure . Bias is any systematic way that a sample differs from the population it is supposed to represent. Surveying only people who answer the phone, only patients who came back to the clinic, or only the mice that survived the first week all build bias into the result. The cure is random selection, where every member of the population has an equal chance of being picked, so that the sample’s differences from the population are only due to chance, and chance shrinks as the sample grows.
A second kind of fairness is the . In 1747 the ship’s surgeon James Lind took twelve sailors sick with scurvy, kept them on the same diet in the same quarters, and gave each pair a different treatment. The pair who got oranges and lemons recovered within days; the others did not. Without the comparison, a recovery proves nothing, because people often recover on their own. Modern trials go further: patients are assigned to groups at random, the control group gets a that looks identical, and neither patients nor doctors know who got which until the end.
When you read a study, ask three questions. How were the subjects chosen, and who was left out? Was there a control group, and was it treated exactly the same except for the one thing being tested? And how large was the sample, remembering that a small random sample beats a huge biased one every time?
Words to know
sample
the part of a population that is actually measured or surveyed
bias
a systematic difference between a sample and the population it is meant to represent
control group
a comparison group treated the same as the test group except for the one factor being studied
placebo
a fake treatment that looks like the real one, used so that the control group's experience matches the test group's
Check yourself
1. Why did the Literary Digest poll fail despite more than two million responses?
Why: The sample was biased toward phone and car owners, and no amount of size fixes a systematic bias.
2. What made Lind's scurvy test convincing rather than a lucky story?
Why: Comparison groups under matched conditions show that the recovery came from the treatment, not from time or chance.
3. Why do modern drug trials give the control group a placebo?
Why: Expecting a treatment can change how people feel and report; a matching placebo keeps expectations equal so only the drug differs.
Section 2
Numbers That Tell the Truth
59.4
Reading a Graph Honestly
Main ideaA graph can tell the truth or bend it depending on its axes, its scale and the window of data it chooses to show.
The same data can be drawn to look like a crisis or like nothing at all. Suppose a city’s water use fell from 100 million gallons a day to 97 million over a decade. Plotted on an that runs from 0 to 100, the line is nearly flat. Plotted on an axis that runs from 96 to 100, the same line plunges toward the floor. Neither graph lies about the numbers, but one of them lies about their importance. Before you read the line, read the axes: where they start, what units they use and whether the is evenly spaced.
The second trick is the window. A stock that has fallen for five years can look like a winner if the chart shows only the last three months. Choosing the start and end points that make your case is called , and it is the most common way honest-looking graphs mislead. Ask what happened just before the graph begins and whether a longer view would change the story. A real holds up when you widen the window.
Watch also for two lines on two different vertical axes, which can be stretched to make anything appear to match anything; for bars whose width or picture size grows along with their height, so the eye sees a bigger change than the numbers show; and for lines with no error bars at all, which hide how uncertain each point is. A logarithmic scale, where each step up multiplies by ten, is not a trick, but it does make explosive growth look like a gentle slope, so check the labels.
None of this means graphs are untrustworthy. A well-made graph shows in one glance what would take a page of numbers to say. The habit to build is simply to read a graph the way you read a contract: the labels, the units, the range, the source and the sample size first, and the dramatic line last.
Words to know
axis
one of the two number lines a graph is drawn on, showing the range and units of a variable
scale
how much each step along an axis represents, which can be even, stretched or logarithmic
cherry-picking
choosing only the data points or time window that support your case
trend
the overall direction of change in data over time, seen when the window is wide enough
Check yourself
1. A graph's vertical axis starts at 96 instead of 0. What effect does this have?
Why: Truncating the axis stretches a small range across the whole height of the graph, exaggerating small differences.
2. What is the best way to check whether a rising line is a real trend or a cherry-picked window?
Why: A genuine trend survives a longer view; a cherry-picked window falls apart when you look before the chosen start date.
3. On a logarithmic vertical axis, what does each equal step upward represent?
Why: On a log scale each equal step multiplies by a fixed factor, usually ten, so rapid growth looks like a gentle slope.
59.5
Error and Uncertainty
Main ideaEvery measurement carries uncertainty, and honest science reports how large that uncertainty is instead of hiding it.
Measure a table with a meter stick and you might get 152.3 centimeters. Measure again and get 152.1. Neither is a mistake. In science the word error does not mean a blunder; it means the unavoidable gap between a measurement and the true value. Every instrument has a finest mark, every hand wobbles, every room changes temperature. The size of that gap is the , and a measurement reported without it is only half reported.
Errors come in two kinds. scatters readings above and below the true value with no pattern, like the wobble of a hand-held ruler. Repeating the measurement and averaging shrinks random error, because the highs and lows cancel. pushes every reading the same way, like a scale that reads 2 grams heavy with nothing on it. Averaging does nothing for it; you find it only by checking the instrument against a known standard. A tight cluster of readings tells you the is high; it does not tell you their .
The measured speed of light shows how uncertainty gets smaller over time. Hippolyte Fizeau in 1849 used a spinning toothed wheel and got a value about 5 percent too high. Albert Michelson, using rotating mirrors and a long baseline, narrowed the range to a few kilometers per second by the 1920s. Each later scientist did not simply announce a number; they announced a number with a plus-or-minus, and the next experiment had to land inside or outside that range.
When you read a result, look for the plus-or-minus, the error bars on a graph or the words ’confidence interval.’ Two measurements that differ by less than their uncertainties agree; nothing has been discovered. A difference that is many times larger than the uncertainty is a real signal. Learning to tell those apart is most of what it means to read data.
Words to know
uncertainty
the range within which the true value of a measurement probably lies, often written as plus-or-minus
random error
unpredictable scatter in repeated measurements that averaging can reduce
systematic error
a consistent shift in every measurement in the same direction, caused by the instrument or method
precision
how closely repeated measurements agree with one another
accuracy
how close a measurement is to the true value
Check yourself
1. A kitchen scale reads 2 grams with nothing on it, so every reading is 2 grams too high. What kind of error is this?
Why: The error pushes every reading the same direction by the same amount, which is the definition of systematic error.
2. Which technique reduces random error but not systematic error?
Why: Averaging cancels random highs and lows, but a consistent shift is present in every reading and survives the average.
3. Five readings of a length cluster tightly around 30.2 cm, but the true length is 31.0 cm. How should the readings be described?
Why: Precision is how well readings agree with each other; accuracy is closeness to the true value. Tight agreement around the wrong value is precise but inaccurate.
59.6
Significant Figures and False Precision
Main ideaThe digits you report should match what your instrument can actually tell you; extra digits claim a certainty you do not have.
A ruler marked in millimeters lets you read a pencil as 14.7 centimeters, with the last digit an honest estimate between marks. Writing 14.7000 centimeters would claim you know the length to a ten-thousandth of a centimeter, which the ruler cannot support. The digits that carry real information are the . In 14.7 there are three. The rule of the game is simple: report every digit you can defend and not one more.
Calculators are the main source of trouble. Divide a 14.7 centimeter pencil into 3 equal pieces and the screen says 4.9 centimeters; divide 14.7 by 7 and it shows 2.1, but the screen may first show 2.1000000. Those extra zeros mean nothing. A result cannot be more precise than the least precise measurement that went into it. When multiplying or dividing, keep the same number of significant figures as the measurement with the fewest. When adding or subtracting, keep the same number of decimal places as the roughest measurement. Counting numbers, like 3 pieces, are exact and do not limit anything.
The same idea guards you against outside the lab. A headline that says the average adult walks 3,812 steps a day sounds exact, but a survey of a few hundred people with step counters of unknown quality cannot pin the number to the last step. A report that a hurricane has a 47.3 percent chance of landfall is claiming more than any forecast model can deliver. When a number carries more digits than its source could possibly justify, it is a sign that someone wants it to sound more certain than it is.
is therefore not laziness; it is honesty. Say about 3,800 steps, roughly a 50 percent chance, 4.9 centimeters. Keep the extra digits in your calculator during the work so rounding errors do not pile up, and round once at the end to the precision your data can actually support.
Words to know
significant figures
the digits in a measurement that carry real information, from the first nonzero digit to the estimated last digit
false precision
reporting more digits than the measurement or method can justify
rounding
trimming a number to the digits that the underlying data can support
Check yourself
1. A length of 14.7 cm is divided by an exact count of 3. How should the result be reported?
Why: The count of 3 is exact, so the measurement 14.7 with three significant figures sets the limit, and 4.9 keeps that precision without adding false digits.
2. Why is a headline claiming '3,812 steps a day' an example of false precision?
Why: The number carries more digits than the sample size and instruments could support, so it sounds more certain than it is.
3. When should you round during a multi-step calculation?
Why: Rounding at every step lets small errors pile up; keeping extra digits during work and rounding once at the end is both accurate and honest.
Section 3
How Science Checks Itself
59.7
Peer Review and Replication
Main ideaScience trusts a result not because an expert approved it but because others can repeat it and get the same answer.
Before a scientific paper appears in a journal, editors send it to other scientists in the field for . The reviewers look for weak methods, missing controls, overreaching conclusions and math errors, and they can demand changes or recommend rejection. Peer review catches a great deal, but it is not proof. Reviewers rarely see the raw data, cannot repeat the experiment, and are volunteers working for free. A peer-reviewed paper is a serious claim worth taking seriously, not a settled fact.
The deeper test is : can a different lab, following the published methods, get the same result? In March 1989 two chemists at the University of Utah announced at a press conference that they had produced nuclear fusion in a jar of water at room temperature. Labs around the world rushed to repeat it. Within months most had failed, the few early successes evaporated, and the claim collapsed. That is the system working. In 2015 a team of psychologists tried to repeat 100 published experiments and found that fewer than half held up, which has pushed many fields to raise their standards.
When a published paper turns out to be wrong through fraud or serious error, the journal issues a , formally withdrawing it from the record. Retractions are rare and painful, but a field without them would be a field that never admits mistakes. Many scientists now also post , drafts shared online before review, so that colleagues can criticize the work early. A preprint has not passed review at all, and news stories built on preprints deserve extra caution.
The lesson for a reader is that no single study is the final word. Ask whether the result has been repeated by independent groups, whether the raw data and methods are available, and whether the finding fits with the rest of what is known. A surprising result from one lab is a lead. A result that many labs have reproduced is knowledge.
Words to know
peer review
the checking of a scientific paper by other experts before a journal publishes it
replication
repeating a study independently to see whether the same result appears
retraction
the formal withdrawal of a published paper because of serious error or misconduct
preprint
a draft scientific paper posted publicly before it has been peer reviewed
Check yourself
1. What does passing peer review guarantee about a paper?
Why: Reviewers judge the paper's methods and logic, but they usually cannot see the raw data or repeat the work, so peer review is a filter, not proof.
2. What happened to the 1989 claim of room-temperature fusion?
Why: Replication is the real test; when independent labs could not reproduce the result, the scientific community set it aside.
3. A news story reports an exciting result from a preprint. Why should a reader be extra cautious?
Why: A preprint is a draft that has skipped even the first filter of peer review, so its claims are the least tested kind of published science.
59.8
Judging a Science News Story
Main ideaA science headline is a claim about a claim; check who did the study, on whom, how big the effect really is and who paid for it.
Headlines compress. A study that found a small effect in 40 mice over eight weeks becomes ’Coffee Cures Memory Loss’ by the time it reaches your feed. The first questions to ask of any science story are the ones a scientist would ask of the paper behind it. Who did the study, and was it published in a peer-reviewed journal? What was studied, humans or animals or cells in a dish? How many subjects, and for how long? Was there a control group? Most stories that fall apart fall apart here.
The second trap is how risk is described. In 2015 a World Health Organization agency reported that eating about 50 grams of processed meat a day raised the of colorectal cancer by about 18 percent. Headlines compared bacon to cigarettes. But the lifetime of colorectal cancer is around 5 or 6 percent, so an 18 percent increase moves it by about one percentage point. Both numbers are true. Only the absolute number tells you how much it matters to you, and stories almost always give you the relative one because it is bigger.
Third, follow the money and the megaphone. A university is written to attract attention, and a study of the topic found that exaggeration in news stories often begins in the release itself. A , such as a study of a sweetener funded by a sweetener company, does not prove the result is wrong, but it raises the bar for how carefully you should check it. Good journals require authors to declare funding; good stories mention it.
Finally, one study is one study. Science moves by the weight of many results, and a single paper that overturns decades of evidence is far more likely to be wrong than to be a revolution. The most reliable science stories are the least exciting ones: they report a body of work, quote scientists who were not involved, give absolute numbers and say plainly what is still unknown.
Words to know
relative risk
how much a factor multiplies the chance of an outcome compared with people without that factor
absolute risk
the actual chance of an outcome, such as 5 in 100, before or after a factor is added
press release
a promotional summary of a study written by an institution to attract news coverage
conflict of interest
a situation in which a researcher or funder stands to gain from a particular result
Check yourself
1. A story says a food raises cancer risk by 18 percent. What number do you need to know how much that matters?
Why: Relative risk multiplies a baseline; only the absolute risk tells you the actual change in your chances, which was about one percentage point in the processed meat case.
2. Why is a study funded by the maker of the product being tested a concern?
Why: A conflict of interest does not prove a result wrong, but it raises the chance of bias and so raises the standard of checking required.
3. A single new study claims to overturn decades of consistent evidence. What is the most reasonable first reaction?
Why: Science rests on the weight of many results; one surprising paper is more likely to be an error than a revolution until others reproduce it.
Section 4
Three Cases Where Evidence Decided
59.9
Smoking and Lung Cancer
Main ideaWithout any experiment on humans, many kinds of evidence together proved that smoking causes lung cancer.
In the first half of the twentieth century, lung cancer went from a rare disease to a common one, and doctors argued about why. Cars, paved roads, factory smoke and cigarettes had all spread at the same time. In 1950 Richard Doll and Austin Bradford Hill published a from London hospitals. They compared hundreds of patients with lung cancer to similar patients with other diseases and asked about their habits. Almost every lung cancer patient was a smoker, and heavy smokers were far more common among the cases than the controls.
A case-control study looks backward and can be fooled by memory and by how patients are chosen, so Doll and Hill launched a in 1951. They wrote to every doctor in Britain and recorded the smoking habits of about 40,000 who replied. Then they simply waited to see who died of what. Within a few years the pattern was unmistakable: the heavier the smoking, the higher the lung cancer death rate, a clear relationship. Doctors who quit saw their risk fall over time. Similar studies in the United States found the same thing.
No one could run a controlled experiment forcing people to smoke, so the case had to be built from the Bradford Hill questions. The association was extremely strong, it showed a dose-response, smoking came before the cancer, it appeared in every country studied, and tar from cigarettes caused tumors when painted on the skin of mice, giving a mechanism. In January 1964 the United States Surgeon General’s committee reviewed the evidence and concluded that cigarette smoking causes lung cancer in men.
The tobacco industry responded not by producing better evidence but by manufacturing doubt. It funded studies designed to muddy the picture and insisted the question was still open. That strategy is worth remembering, because it has been copied for other products since. The demand for more evidence is a good scientific instinct; the demand for evidence that could never be enough is not.
Words to know
case-control study
a study that compares people who have a condition with similar people who do not, looking backward for differences
cohort study
a study that records a group's habits or exposures first and then follows them forward to see what happens
dose-response
a pattern in which more of a suspected cause produces more of the effect
Check yourself
1. What is the main weakness of a case-control study that led Doll and Hill to start a cohort study?
Why: Case-control studies rely on recalled history and on the choice of cases and controls; a cohort study records exposure first and then watches, avoiding those biases.
2. Which finding is an example of dose-response in the doctors study?
Why: Dose-response means more of the cause produces more of the effect, which is exactly what rising death rates with heavier smoking show.
3. How did the tobacco industry respond to the mounting evidence?
Why: The industry's strategy was to keep the question looking open, an approach that has been reused for other products since.
59.10
The Ozone Layer
Main ideaA prediction made from laboratory chemistry was confirmed by measurements a decade later, and the world acted on the evidence.
High in the , roughly 15 to 35 kilometers up, a thin layer of , a molecule of three oxygen atoms, absorbs most of the sun’s harmful ultraviolet light. In 1974 the chemists Mario Molina and Sherwood Rowland published a warning. , the chlorofluorocarbon gases then used in spray cans, refrigerators and foam, were so stable that nothing at ground level destroyed them. They would drift upward for decades until ultraviolet light broke them apart, releasing chlorine atoms. Each chlorine atom, they calculated, could act as a and destroy thousands of ozone molecules before being removed.
This was a claim built from laboratory rates and a model of the atmosphere, not from any observed loss. Critics, including the companies that made CFCs, said the effect was speculative. Then in 1985 three scientists from the British Antarctic Survey, Joseph Farman, Brian Gardiner and Jonathan Shanklin, reported that the ozone above their Halley Bay station every October had fallen sharply since the 1970s, to far below anything the models had predicted. A satellite had been measuring the same region, but its software had been flagging the extremely low values as instrument errors.
The Antarctic ozone hole was a stronger effect than Molina and Rowland had predicted, and the reason turned out to be ice clouds in the polar winter that speed up the chlorine chemistry. Within two years, in 1987, nations signed the Montreal Protocol to phase out CFCs, and the treaty has been tightened several times since. Molina and Rowland shared the 1995 Nobel Prize in Chemistry with Paul Crutzen. Measurements show the amount of ozone-destroying chlorine in the stratosphere peaking around the turn of the century and slowly falling, and the ozone layer is now recovering, with the Antarctic hole expected to heal around the 2060s.
The case shows the full arc of an argument from evidence. A mechanism and a prediction came first. Independent measurements confirmed the prediction, and the exceptions, the values too low to believe, turned out to be the most important data. The claim survived attack, the world acted, and later measurements tested whether the action worked.
Words to know
stratosphere
the layer of the atmosphere above the weather, roughly 10 to 50 kilometers up, where the ozone layer sits
ozone
a molecule of three oxygen atoms that absorbs ultraviolet light high in the atmosphere
CFC
chlorofluorocarbon, a stable human-made gas once used in spray cans and refrigerators that releases chlorine high in the atmosphere
catalyst
a substance that speeds up a chemical reaction without being used up, so it can act again and again
Check yourself
1. What was the basis of Molina and Rowland's 1974 warning?
Why: Their claim was a prediction from mechanism, not an observation; the observed loss came a decade later.
2. Why can a single chlorine atom destroy thousands of ozone molecules?
Why: A catalyst is regenerated after each reaction cycle, so one chlorine atom can break ozone apart again and again.
3. What role did the 1985 Halley Bay measurements play in the argument?
Why: The measurements were the observational test of a prediction made from chemistry, and the size of the loss pushed nations to act.
59.11
Vaccines and Autism
Main ideaA tiny, flawed study created a fear that enormous, careful studies have repeatedly failed to support.
In February 1998 the medical journal The Lancet published a paper describing twelve children with developmental problems, most of them autism, whose parents linked the symptoms to the measles, mumps and rubella vaccine. Twelve children, no control group, no comparison to unvaccinated children, and a raised at a press conference by the lead author, Andrew Wakefield, who suggested the combined vaccine be split up. Vaccination rates in Britain fell. Measles, which had nearly vanished, came back.
The proper way to test the hypothesis is to compare large groups of vaccinated and unvaccinated children and see whether autism is more common in one. Researchers in Denmark, where national health records track every child, did exactly that. A 2002 study of more than 500,000 children found no difference in autism rates. A 2019 study of more than 650,000 children, including children considered at higher risk, found none either. Studies in the United States, Japan and elsewhere agreed. Autism diagnoses have risen, but the rise tracks broader definitions and better screening, and it appears in vaccinated and unvaccinated children alike.
Meanwhile an investigation found that the 1998 paper’s data had been altered, that the children had been recruited through lawyers preparing a lawsuit against vaccine makers, and that the lead author had been paid by those lawyers. In 2010 The Lancet the paper, and Britain’s medical council struck Wakefield from the register of doctors. The paper’s central claim had failed every test that evidence can apply: it could not be replicated, its data were unreliable and its author had an undisclosed conflict of interest.
The case matters not because scientists were wrong for a while, which happens, but because of how the two kinds of evidence compare. On one side is a story about twelve children told by a person with a financial stake. On the other are studies of more than a million children with no link found. Learning to weigh those two against each other, rather than treating them as two equal sides, is the whole skill this chapter has been about.
Words to know
hypothesis
a proposed explanation that can be tested by gathering evidence
retracted
formally withdrawn from the scientific record by the journal that published it
cohort
a large group of people followed over time to see what happens to them
Check yourself
1. Which weakness of the 1998 paper would a reader of this chapter spot first?
Why: Twelve cases with no comparison group cannot show whether autism is any more common among vaccinated children than among others.
2. What did the large Danish studies find?
Why: Cohort studies of hundreds of thousands of children found no difference in autism rates, and other countries' studies agreed.
3. Why should the 1998 paper and the later cohort studies not be treated as two equal sides of a debate?
Why: Evidence is weighed by its quality and quantity, and the two sides differ enormously in both.
Chapter review
Reading and Judging Evidence
0 / 8
1. John Snow's map showed deaths clustered around the Broad Street pump. Why was the brewery evidence so important?
Why: Brewery workers breathed the same air as their neighbors but drank beer instead of pump water and stayed healthy, which fits the water explanation and not the bad-air one.
2. A study finds that students who eat breakfast get higher grades. Which is the best confounding variable to consider?
Why: A hidden factor that drives both breakfast and grades, such as household resources, could create the correlation without breakfast causing better grades.
3. Which sample would give the most trustworthy estimate of a city's opinion?
Why: A smaller random sample avoids the systematic bias built into surveys of car owners, volunteers or callers.
4. A measurement is reported as 52.3 plus or minus 0.4 grams. A second measurement reads 52.6 plus or minus 0.4 grams. Do they disagree?
Why: Two results whose uncertainty ranges overlap are consistent with each other; nothing has been detected.
5. Why did most laboratories abandon the 1989 cold fusion claim?
Why: Replication by independent labs is the decisive test, and the claim failed it.
6. A news story reports that a habit doubles the risk of a rare disease affecting 1 in 10,000 people. What is the absolute risk after doubling?
Why: Doubling a relative risk from a baseline of 1 in 10,000 gives about 2 in 10,000, still a very small absolute risk.
7. Which piece of evidence gave the smoking and cancer case its mechanism?
Why: Laboratory evidence that a component of smoke causes tumors explained how the association seen in people could be causal.
8. What do the ozone and vaccine cases have in common about how evidence was used?
Why: In the ozone case the first claim was confirmed by later data; in the vaccine case the first claim was refuted. Either way the deciding factor was the weight of independent evidence.
Send it to your teacher
60
Chapter
Designing and Making the Case
Science Practices
Big questionHow do you turn a question you care about into an investigation whose answer others will believe?
The story
The Genes That Would Not Sit Still
A scientist studying spotted corn kernels saw something that the textbooks said was impossible, and then waited thirty years for the rest of biology to catch up.
Barbara McClintock spent her working life in cornfields. Corn, or maize, is a wonderful organism for a geneticist, because every kernel on a cob is a separate offspring, and a single ear can display hundreds of results at once. By the 1930s McClintock had already helped show, with her colleague Harriet Creighton, that when chromosomes physically swap pieces, traits get swapped along with them. Colleagues considered her one of the best geneticists alive.
At Cold Spring Harbor on Long Island in the 1940s she turned to a puzzle. Some kernels that should have been a solid color came out streaked and spotted, and the spots did not follow the tidy rules of inheritance. Genes at that time were pictured as beads fixed in order on a string. McClintock's crosses, tracked kernel by kernel across years of growing seasons, pointed to something stranger. A piece of genetic material was jumping from one place on a chromosome to another. When it landed inside a pigment gene, it switched the gene off; when it jumped out again in a growing cell, the color came back in that cell's descendants, painting a spot.
She called these pieces controlling elements and presented the work at a Cold Spring Harbor symposium in 1951. The response was silence, and then polite doubt. The idea did not fit the model everyone was using, the evidence was a mountain of corn crosses that few others had the patience to follow, and the person presenting it was a woman in a field run by men. Within a few years McClintock largely stopped publishing on the subject, though she never stopped the work.
In the late 1960s and 1970s, researchers working on bacteria found stretches of DNA that inserted themselves into genes and moved around. Then they were found in yeast, in fruit flies and in people. Today we know that nearly half of the human genome is made of such transposable elements, or transposons, and that they play a part in disease and in how genes are regulated. In 1983, at the age of 81, McClintock received the Nobel Prize in Physiology or Medicine, alone, for a discovery made three decades earlier.
Her story is not simply one of a genius ignored. It is a story about what it takes to make a case: a question sharp enough to test, years of carefully designed crosses, records that could survive scrutiny, and the willingness to present a result that the audience did not want. This chapter is about building that kind of case yourself, from the first question to the moment you stand up and defend the answer.
Talk about itMcClintock's evidence was correct in 1951, but it was not accepted until other methods found the same thing. What does this suggest about the difference between being right and making a convincing case?
Section 1
From Curiosity to a Testable Question
60.1
Asking a Testable Question
Main ideaA good investigation begins with a question that names what you will change, what you will measure and how you would know the answer.
Why does bread rise? Are plants happier in the sun? Does music help you study? Each of these is a real curiosity, but none of them is yet a , one that an investigation could actually answer. ’Happier’ is not something a plant shows on a dial. ’Help you study’ could mean a dozen things. The first job in any investigation is to sharpen the curiosity until it points at something you could measure and something you could change.
The sharpening step is to write an for every fuzzy word. Instead of ’happier,’ choose ’height in centimeters after three weeks’ or ’number of leaves.’ Instead of ’help you study,’ choose ’score on a 20-word recall test after ten minutes.’ Now the question can be rewritten: how does the number of hours of daily light affect the height of bean seedlings after 21 days? It names the thing you will change, the thing you will measure and the time frame. Someone else could run it.
From the question comes a , a proposed answer with a reason behind it: seedlings with more daily light will grow taller, because light supplies the energy for photosynthesis. And from the hypothesis comes a that could fail: if we give one group 4 hours and another 12 hours, the 12-hour group will be taller at 21 days. If they are not, the hypothesis is in trouble. A question whose every possible answer would leave you believing the same thing is not worth the time.
McClintock’s question was of this kind. Not ’why are these kernels strange?’ but ’where on the chromosome is the element that switches the color gene off, and does it stay put across generations?’ Each cross she planned was a prediction that the next season’s ears could confirm or wreck. Your capstone question needs the same shape, whatever its subject.
Words to know
testable question
a question that an investigation could answer by measuring something under conditions you can set
operational definition
a precise statement of how a vague term will actually be measured in an investigation
hypothesis
a proposed answer to a testable question, with a reason behind it
prediction
a specific result that should happen if the hypothesis is right, and that could fail if it is wrong
Check yourself
1. Why is 'Are plants happier in the sun?' not yet a testable question?
Why: Without an operational definition of the outcome and a variable to change, no investigation could answer it.
2. Which of these is an operational definition?
Why: An operational definition states exactly how a term will be measured, including the unit and the conditions.
3. What must be true of a good prediction?
Why: A prediction is useful only if it can fail; that is what lets the investigation actually test the hypothesis.
60.2
Variables and Controls
Main ideaA fair test changes one factor on purpose, measures one outcome and holds everything else the same across a comparison group.
Every experiment has three kinds of quantities. The is the one you deliberately change, such as the hours of light a seedling gets. The is the one you measure to see the effect, such as height. The are everything else that could matter and that you hold the same for every group: the kind of seed, the soil, the pot size, the water, the temperature. A careless experiment lets a controlled variable drift, and then you cannot tell whether the light or the drift caused the difference.
The comparison is what makes a test fair. If every seedling gets 12 hours of light, you learn what 12-hour seedlings look like but nothing about the effect of light. You need at least two levels of the independent variable, and it is usually best to have several, so that you can see whether the effect grows steadily or levels off. A control group, given the ordinary or zero condition, anchors the comparison. In a medical trial the control gets a placebo; in a plant experiment the control might be the light level seedlings normally get.
One variable at a time is the beginner’s rule, and it is a good one. It is not the only design, though. Statisticians led by Ronald Fisher in the 1920s and 1930s worked out how to vary several factors at once in carefully arranged patterns and then untangle their separate effects, which is how modern agricultural and industrial experiments are run. What never changes is the principle: the only difference between the groups should be the one you put there, and chance should decide which unit goes into which group.
That last point, random assignment, deserves emphasis. If you put the healthiest-looking seedlings in the 12-hour group because you expect them to do well, you have built your expectation into the result. Number the pots, flip a coin or use a random number list, and let chance sort them. Then a difference at the end can only come from the treatment or from chance, and you can calculate how likely chance is.
Words to know
independent variable
the factor an experimenter changes on purpose
dependent variable
the outcome an experimenter measures to see the effect of the change
controlled variable
a factor held the same for every group so that it cannot cause the difference
random assignment
letting chance decide which subjects go into which group, so that expectations do not bias the result
Check yourself
1. In an experiment on how fertilizer amount affects tomato yield, what is the dependent variable?
Why: The dependent variable is the outcome measured; fertilizer amount is the independent variable and plant type and pot size are controlled.
2. Why does an experiment need more than one level of the independent variable?
Why: Without a comparison there is nothing to attribute the outcome to; the effect of a variable is a difference between levels.
3. A student puts the strongest-looking seedlings in the treatment group. What has gone wrong?
Why: Assigning by expectation builds a hidden difference into the groups; random assignment prevents it.
60.3
Planning the Investigation
Main ideaA written plan with a procedure, sample size, repeated trials, safety steps and a pilot run turns an idea into an investigation others could repeat.
A is the written procedure for an investigation, detailed enough that a stranger could follow it and get the same kind of data. It lists materials, the exact steps, the order of steps, the timing, what will be measured and how. Writing it forces decisions you would otherwise make on the fly, and those on-the-fly decisions are where bias creeps in. It also lets a reader judge the work later, which is why published papers include a methods section.
Two numbers matter most in the plan: how many subjects and how many . One seedling per group tells you almost nothing, because seedlings differ from one another for reasons that have nothing to do with light. Ten per group lets the individual differences average out. Repeating each measurement, and repeating the whole experiment if you can, shows how much scatter to expect, and without knowing the scatter you cannot say whether a difference is real. As a rule, plan for more subjects than you think you need, because some will be lost.
Before committing weeks to the full experiment, run a : a small, quick version to find out what goes wrong. The pots may dry out faster than expected; the recall test may be too easy; the microscope slide may need a different stain. Pilots are cheap and save you from discovering a broken procedure at the end. Many scientists now also write down their hypothesis and analysis plan before collecting data, a practice called preregistration, so that they cannot quietly change the question to match the result.
Finally, plan for safety and ethics. Chemicals, heat, sharp tools and electricity need specific precautions in writing. Any investigation involving people, even a survey of classmates, needs their informed consent and a way to keep their answers private. Research on animals and people at universities must be approved by a review board before it starts. A plan that skips these is not ready.
Words to know
protocol
the written, step-by-step procedure for an investigation, detailed enough for someone else to follow
trial
one run of a measurement or experiment; repeated trials show how much results scatter
pilot study
a small, quick version of an experiment run first to find problems in the procedure
preregistration
writing down the hypothesis and analysis plan before collecting data, so the question cannot be changed to fit the result
Check yourself
1. Why should a protocol be detailed enough for a stranger to follow?
Why: Replication and evaluation both depend on knowing exactly what was done; that is the purpose of a methods section.
2. What is the main reason to use ten seedlings per group instead of one?
Why: Subjects vary for many reasons; a larger group lets that random variation cancel so the treatment effect can be seen.
3. What is the purpose of a pilot study?
Why: A quick small run exposes practical problems, such as pots drying out or a test that is too easy, while they are still cheap to fix.
Section 2
Data and Models
60.4
Collecting Data You Can Trust
Main ideaGood data come from a checked instrument, a table planned in advance, units on every number and a record of everything odd that happened.
Data collection begins before the first measurement, with the instrument. means checking the instrument against a known standard: a balance against a certified mass, a thermometer in an ice-water bath that should read 0 degrees Celsius. An uncalibrated instrument can be perfectly precise and perfectly wrong. Know its too, the smallest change it can show. A ruler marked in millimeters cannot report tenths of a millimeter, no matter how carefully you squint.
Design the data table before you collect anything. Each column gets a label and a unit; each row is one trial or one subject. Leave room for the date, the time, the conditions and a notes column. Record , the numbers as the instrument gave them, and never overwrite them. Do calculations in separate columns later. A smudged, corrected or ’improved’ raw number is worthless, because no one, including you, can tell what was actually observed. Scientists keep bound notebooks with numbered pages for exactly this reason.
Write down everything unusual. A pot got knocked over; the room was unusually warm on day 9; the microscope lamp flickered. When a strange value, an , appears in the data, the notes tell you whether it came from a real event or from a known accident. Throwing out an outlier because it spoils the pattern is not allowed. Throwing it out because the notes say the sample was contaminated, and saying so in the report, is good practice.
If you use this site’s microscope for a project, the same rules apply: note the magnification, the stain, the time since the slide was made and the field of view you counted. Two counts of cells in a drop of pond water made under different conditions are not the same measurement, and the notes are what let you tell the difference.
Words to know
calibration
checking an instrument against a known standard so its readings can be trusted
resolution
the smallest change in a quantity that an instrument can show
raw data
measurements recorded exactly as the instrument gave them, before any calculation or correction
outlier
a data value far from the rest, which may be a real event or a known error
Check yourself
1. A thermometer placed in an ice-water bath reads 2 degrees Celsius. What has the student learned?
Why: An ice-water bath is a known standard at 0 degrees Celsius; a reading of 2 degrees shows a calibration error.
2. Why should raw data never be overwritten with corrected values?
Why: Raw data are the only record of the observation itself; corrections belong in separate columns where they can be checked.
3. When is it acceptable to exclude an outlier from analysis?
Why: Exclusion is justified only by a documented problem with the measurement, and it must be stated in the report.
60.5
Finding the Pattern
Main ideaAnalysis means summarizing the data, plotting them and asking whether the pattern is bigger than the scatter.
With the data in hand, the first step is to summarize each group. The , the sum divided by the count, is the everyday average. The , the middle value when the data are sorted, is better when a few extreme values would drag the mean around. The , or a more careful measure of spread like the standard deviation, tells you how scattered the values are. Report both a center and a spread; a mean without a spread is a claim without an uncertainty.
Next, plot the data. Put the independent variable on the horizontal axis and the dependent variable on the vertical, with every point shown, not just the averages. A picture reveals what a table hides: whether the relationship is a straight line, a curve that flattens out, or no relationship at all. If the points fall roughly along a line, a drawn through them gives a rate, and its slope has units, such as centimeters of growth per hour of light per day.
The key question is whether the pattern is larger than the noise. If the 12-hour seedlings average 14 centimeters and the 4-hour seedlings average 12, but each group’s values range from 9 to 17, the difference is inside the scatter and you have not shown anything. If the ranges barely overlap, the difference is real. Statistical tests put a number on this: how often would chance alone produce a difference this big? A result that chance would produce fewer than 1 time in 20 is usually called significant, but that is a convention, not a law of nature, and a tiny significant difference may not matter.
Beware of the pattern that fits too well. With enough freedom you can draw a wiggly curve through any set of points, and it will predict nothing. Prefer the simplest relationship that fits the data within its uncertainty, and check whether it fits new data you did not use to draw it.
Words to know
mean
the sum of the values divided by how many there are; the ordinary average
median
the middle value when the data are sorted; less affected by extreme values than the mean
range
the difference between the largest and smallest values, a simple measure of spread
best-fit line
a straight line drawn through plotted data so that it passes as close as possible to all the points
Check yourself
1. Reaction times for a class are mostly near 0.25 seconds, but two students scored 1.5 seconds. Which summary best represents a typical student?
Why: The median ignores how extreme the outliers are, while the two slow times would pull the mean well above what most students scored.
2. Two groups have means of 14 cm and 12 cm, but values in each group range from 9 to 17 cm. What can be concluded?
Why: A difference smaller than the spread within groups could easily arise by chance; more data or a larger effect would be needed.
3. Why prefer the simplest curve that fits the data over a wiggly one that passes through every point?
Why: Overfitting captures random scatter rather than the real relationship; a simpler fit generalizes better and can be tested on new data.
60.6
Models and Their Limits
Main ideaA model is a deliberate simplification that lets you predict and explain, and it is only as trustworthy as its record against real data.
A is a stand-in for something too big, too small, too slow or too complicated to work with directly. A globe is a model of Earth. An equation relating a pendulum’s length to its period is a mathematical model. A computer of the atmosphere, dividing the air into millions of boxes and stepping their temperature and wind forward in time, is a model too. Every one of them leaves things out on purpose. The globe has no weather; the pendulum equation ignores air resistance; the climate model cannot track every cloud.
The leaving-out is the point, and it is also the danger. Each model rests on , and it is only reliable where those assumptions hold. The pendulum equation works beautifully for small swings and fails for wide ones. The bead-on-a-string picture of genes that McClintock’s colleagues held was a useful model that made good predictions for decades, right up until her corn showed genes moving. A model that has never been wrong may just never have been tested outside its comfortable range.
So the test of a model is its record. Does it predict measurements it was not built from? Where does it start to fail, and does the failure make sense? Weather models are checked every single day against what actually happened, which is why a five-day forecast today is about as good as a three-day forecast was decades ago. When a model and a careful measurement disagree, the measurement usually wins, and the disagreement is where the next discovery hides.
For your own project you will build at least a simple model: a line through data points, a diagram of a food web, a formula relating your variables. State its assumptions. Say where you expect it to hold and where you do not. A reader who sees that you know the limits of your model will trust the parts inside them.
Words to know
model
a simplified representation of a system, physical, mathematical or computer-based, used to explain and predict
simulation
a computer model that calculates how a system changes step by step over time
assumption
something a model takes as true in order to simplify, which limits where the model applies
Check yourself
1. The pendulum equation ignores air resistance. Why is that acceptable?
Why: Models leave things out on purpose; the omission is fine as long as you know where it makes the model fail.
2. What is the best test of whether a model can be trusted?
Why: A model earns trust by its record against new data; the more detailed a model is, the harder it is to check, not the more correct.
3. McClintock's colleagues used a model of genes as fixed beads on a string. What does her story show about models?
Why: The bead model made good predictions until data outside its range appeared; that is how models are corrected.
Section 3
Engineering a Solution
60.7
Criteria and Constraints
Main ideaAn engineering problem is defined by measurable criteria for success and firm constraints that any solution must obey.
Science asks what is true; engineering asks what will work. A school wants to cut the flooding on its playing field after storms. Before anyone sketches a solution, the problem needs a sharp definition. The are what a successful design must do, written so that they can be measured: standing water gone within 12 hours of a 2-centimeter rain; no damage to the running track; safe for students. Vague criteria like ’better drainage’ cannot be tested and cannot settle an argument between two designs.
The are the limits every design must stay inside: a budget of a certain amount, a completion date before the fall season, no digging near the buried gas line, compliance with the city’s stormwater rules. A design that fails a constraint is out, no matter how well it scores on the criteria. Listing constraints early saves everyone from falling in love with a design that could never be built.
Almost no design wins on every criterion, so engineers face . A deep gravel trench drains fastest but costs the most and takes longest to build. A raised field is cheap but changes the playing surface. A makes the trade-offs visible: list the designs as rows and the criteria as columns, score each cell, weight the criteria by how much they matter, and add up. The matrix does not decide for you, but it forces you to say out loud which criteria you value most and why.
This is the same claim, evidence and reasoning habit in a new dress. The claim is that a design meets the criteria within the constraints; the evidence is test data and cost figures; the reasoning shows how the trade-offs were weighed. Notice that criteria and constraints usually come from people, not from nature, and that different stakeholders, the coach, the neighbors, the budget office, will weight them differently.
Words to know
criteria
the measurable requirements a successful design must meet
constraints
the limits, such as cost, time, materials, safety or law, that any design must stay within
trade-off
giving up some of one desirable feature to gain more of another
decision matrix
a table that scores each design against each criterion to make comparisons and trade-offs visible
Check yourself
1. Which of these is a constraint rather than a criterion?
Why: A constraint is a hard limit any design must obey; the others describe what a successful design should accomplish.
2. Why must criteria be written so they can be measured?
Why: Only measurable criteria allow a test to show whether a design succeeded, which is how evidence rather than opinion decides.
3. What does a decision matrix actually do?
Why: The matrix organizes judgments and forces the weighting of criteria into the open; people still make the decision.
60.8
Build, Test, Fail, Improve
Main ideaDesigns improve through cycles of building a prototype, testing it against the criteria, studying the failures and trying again.
Engineering rarely gets a design right the first time, and it does not expect to. The design cycle runs: define the problem, brainstorm, build a , test it, study what went wrong, revise, and test again. A prototype is a working version made quickly and cheaply enough that you are willing to break it. A cardboard model of the drainage trench, a scale bridge made of craft sticks, a spreadsheet of the costs: each is a prototype of a different part of the design.
The test must be against the written criteria, not against a feeling that it looks good. Load the craft-stick bridge with weights until it breaks and record the mass. Pour a measured amount of water on the model field and time how long it stands. Comparing prototypes on the same test is what turns opinion into evidence. Numbers from a fair test also settle arguments between team members far more peacefully than volume does.
The most valuable step is : figuring out why a prototype failed rather than just that it did. Did the bridge break at a joint or in the middle of a stick? Did the water pool because the slope was wrong or because the gravel clogged? In 1901 the Wright brothers found that the published tables of lift they had trusted were wrong, built their own wind tunnel, and tested more than a hundred wing shapes before the 1903 flight. In 1940 the Tacoma Narrows Bridge twisted itself apart in a moderate wind, and the study of that failure changed how every long bridge since has been designed.
Each cycle is an , and the number of iterations is not a sign of poor work. Commercial products often go through hundreds or thousands of prototypes. The record of what was tried, what failed and why is as important as the final design, because the next engineer, or you next year, will need it.
Words to know
prototype
an early, quick version of a design built to be tested and probably broken
failure analysis
the study of exactly how and why a design failed, to guide the next version
iteration
one complete cycle of building, testing and revising a design
Check yourself
1. What is the purpose of a prototype?
Why: A prototype is built to be tested and often broken, so that failures are found while they are still cheap.
2. Why did the Wright brothers build their own wind tunnel?
Why: When their gliders underperformed, they traced the failure to bad published data and generated their own through systematic testing.
3. Which is the best response to a prototype bridge breaking at a glued joint?
Why: Failure analysis targets the actual weak point; a blanket fix wastes effort and may not address the cause.
Section 4
Making the Case
60.9
Writing a Scientific Explanation
Main ideaA written explanation states the claim, presents the evidence with numbers and uncertainty, reasons from principle, and admits what it cannot rule out.
The claim comes first, in one sentence, answering the question you asked. ’Bean seedlings given 12 hours of daily light grew taller over 21 days than seedlings given 4 hours.’ Not ’light affects plants,’ which is too vague to test, and not ’light is essential to all life,’ which your data cannot support. State it in the past tense about what you found, and keep it inside the range you tested.
The evidence follows, with numbers. Means with their spread, the number of subjects, and a graph or table the reader can inspect. ’The 12-hour group averaged 14.2 centimeters, ranging 12.5 to 16.0, across 10 seedlings; the 4-hour group averaged 9.8 centimeters, ranging 8.1 to 11.4.’ Then the reasoning: the scientific principle that explains why this evidence supports the claim, such as light supplying the energy for photosynthesis, which builds the sugars used for growth. Reasoning is where you connect your small experiment to what is already known.
The strongest explanations then turn on themselves. Name the you considered and say what evidence rules them out: could the 12-hour lamp also have warmed the pots? If you measured temperature and it was the same, say so; if you did not, say that too. State the : one bean variety, three weeks, indoor conditions. A reader who sees you searching for your own weaknesses is a reader who begins to trust you.
Finally, give credit. A for every fact, value or method that came from someone else tells the reader where to check and shows that you built on real work rather than invention. McClintock’s papers were dense with cross-references to decades of maize genetics, which is part of why they held up when they were finally read carefully.
Words to know
alternative explanation
a different cause that could also account for the evidence, which a good explanation tries to rule out
limitation
a boundary on what a study can show, such as its sample, duration or conditions
citation
a reference that tells the reader where a fact, method or value came from
Check yourself
1. Which claim stays properly inside the range of the seedling experiment?
Why: A claim should state what was found under the tested conditions, neither too vague nor broader than the data.
2. Why should a written explanation name alternative explanations?
Why: Considering and testing alternatives is how an explanation earns trust; leaving them out invites the reader to find them.
3. What does a citation accomplish?
Why: Citations give credit and, more importantly, let readers trace and verify what the explanation rests on.
60.10
Presenting and Defending a Claim
Main ideaDefending a claim means answering objections with evidence, separating questions of fact from questions of value, and changing your mind when the evidence demands it.
A scientific presentation, whether a talk, a poster or a written report, is an argument, and arguments draw fire. Prepare for it. Before you present, list the five hardest questions someone could ask and work out your answers. Most objections come in a few forms: your sample was too small, you did not control something, there is a simpler explanation, or your numbers do not show what you say they show. Each of these is a , and each deserves a that points at evidence, not at your effort or your confidence.
Some objections will be right. The honest response to a good objection is to say so, and to say what you would do about it: ’That is a fair point; we did not measure temperature, and the lamp could have warmed the pots. The next run should include a thermometer in each pot.’ Scientists who cannot say ’you are right’ in public are not trusted in private either. The goal of a defense is not to win but to find out, together with your critics, what the evidence actually supports.
Learn to separate two kinds of disagreement. Whether CFCs destroy ozone is a question of fact, settled by measurement. Whether the cost of phasing them out was worth it is a question of values, on which measurement informs but does not decide. In a or a public meeting, people mix the two constantly. Naming which kind of question is on the table is often the most useful thing a presenter can do.
Two cautionary tales frame the skill. The cold fusion scientists of 1989 were confident, presented dramatically and were wrong. McClintock was quiet, thorough and right, but her audience was not ready and she stopped pushing. The lesson is not that confidence is bad or that being ignored proves you correct. It is that a claim stands or falls on the evidence and on how carefully it was gathered, and the presenter’s job is to make that evidence as clear and as checkable as possible.
Words to know
counterclaim
an opposing claim or objection raised against an argument
rebuttal
a response to an objection that uses evidence and reasoning to answer it
poster session
a scientific meeting format in which researchers stand beside a printed summary of their work and answer questions
Check yourself
1. A critic says your control group was not treated identically. What is the best response?
Why: A rebuttal rests on evidence; where the objection is valid, conceding it and proposing a fix is the honest and trusted response.
2. Which is a question of values rather than of fact?
Why: Whether to keep the wetland depends on what people value; the other questions can be settled by measurement.
3. What lesson do the cold fusion and McClintock stories teach together?
Why: One confident claim failed and one quiet claim succeeded because of the evidence behind each, not the presentation style.
60.11
Your Capstone Plan
Main ideaA capstone plan lays out the question, the design, the data, the analysis, the safety steps and the timeline before any measurement is made.
Your project is one investigation, chosen by you, carried through every stage this unit has covered. It can be an experiment, a field study, an analysis of existing data or an engineering design, but it must produce evidence and end with a defended claim. Start with what you are curious about near you: the temperature of Chicago’s lakefront versus an inland neighborhood, the number of organisms in a drop of pond water under this site’s microscope, the craters visible on the moon with its telescope on different nights, the drainage on your own school’s field. Local questions are easier to measure and harder to fake.
The written plan has a fixed set of parts. The testable question and hypothesis. Background: what is already known, with citations. The variables, identifying which is independent, which dependent and which controlled, or the criteria and constraints if it is a design. The protocol, step by step, with materials. The data table, empty, with units. The analysis plan: what you will calculate and what graph you will draw. Safety and ethics. And a line stating what result would make you abandon the hypothesis.
Then the . Work backward from the presentation date and set : plan approved, pilot run done, data collection finished, analysis done, draft written, practice defense. Plants take weeks to grow and pond water changes with the seasons, so the calendar is a real constraint. Build in slack, because a pilot will reveal something that needs changing.
Everything in the plan is a promise to your future self and to your readers about how the work will be done, made before the results can tempt you to bend it. That is preregistration in miniature, and it is the single best protection against fooling yourself. When the data come in, you will analyze them the way you said you would, report what you find whether or not it is what you hoped, and stand up to defend it. That is the whole practice of science in one project.
Words to know
capstone
a final project that brings together every skill of a course in one complete investigation
timeline
a schedule that sets when each stage of a project will be done
milestone
a checkpoint in a project where a specific piece of work must be complete
Check yourself
1. Why does the plan include a statement of what result would make you abandon the hypothesis?
Why: Deciding in advance what would count against the hypothesis is preregistration in miniature and protects against fooling yourself.
2. Which of these is a milestone?
Why: A milestone is a dated checkpoint for a piece of work; the other items are parts of the plan, not points on the timeline.
3. Why does the reading recommend local questions for a capstone?
Why: Local phenomena can be measured firsthand with available tools, which makes the evidence real and checkable.
Chapter review
Designing and Making the Case
0 / 8
1. Which question is testable as written?
Why: It names what is changed, what is measured and the conditions; the others have no measurable terms.
2. In a test of whether music affects recall, students in the music group are tested in the morning and the silent group after lunch. What is wrong?
Why: The groups differ in two ways at once, so any difference in recall cannot be attributed to the music alone.
3. What is the main purpose of a pilot study?
Why: A quick small run exposes practical problems while they are cheap to fix.
4. A student's data table has numbers but no units and several erased, rewritten entries. Why is this a problem?
Why: Raw data must be recorded as observed and labeled with units, or the record cannot be checked or trusted.
5. Two groups differ in their means, but the ranges of values in each group overlap almost completely. What follows?
Why: A difference smaller than the spread within groups could arise by chance, so more evidence is needed.
6. What does the McClintock story show about scientific models?
Why: The fixed-gene model worked for decades and then failed when McClintock's crosses tested it in a new situation.
7. Which of these is a criterion rather than a constraint for a footbridge?
Why: A criterion describes measurable success; budget, deadline and protected creek bed are limits any design must obey.
8. A written explanation lists its own limitations and the alternative explanations it could not rule out. What effect should this have on a careful reader?
Why: An explanation that honestly bounds itself shows the author is reporting what the evidence supports rather than selling a result.
Send it to your teacher
★
Unit wrap-up
Capstone: Argue From Evidence
Twelve words, twelve meanings
0 / 12
Tap a word, then tap its meaning. A right pair locks in green.
Words
Meanings
Unit test
Fifteen questions across the unit
0 / 15
1. John Snow's map showed deaths clustered around one pump. What turned that pattern into a strong argument?
Why: A cluster could have many causes; cases that the rival theory could not explain but the water theory could are what made the case.
2. Counties with more churches also have more crime. What is the most likely explanation?
Why: Population size is a confounding variable that drives both counts, creating a correlation without causation.
3. Which sample is most likely to represent the opinions of a whole city?
Why: Random selection avoids the systematic bias built into volunteer or car-owner samples, regardless of their size.
4. A bar chart's vertical axis runs from 9.8 to 10.2 hours, and the bars show 10.1 and 9.9 hours. How does this affect a viewer?
Why: A truncated axis stretches a tiny range across the full height, exaggerating a small difference.
5. A scale reads 3 grams with nothing on it. Repeating a measurement ten times and averaging will do what?
Why: Averaging reduces random scatter but cannot remove a consistent offset; only calibration catches systematic error.
6. A student's calculator shows a density of 0.944700461 g per cubic centimeter from measurements with two significant figures. How should it be reported?
Why: A result cannot be more precise than its least precise measurement, so two significant figures is the honest report.
7. What was decisive in rejecting the 1989 cold fusion claim?
Why: Replication is the real test; the claim failed it worldwide within months.
8. A headline says a food raises the relative risk of a disease by 20 percent. The absolute risk was 5 percent. What is the new absolute risk?
Why: Twenty percent of 5 percent is 1 percentage point, so the risk rises from about 5 to about 6 percent.
9. How did scientists establish that smoking causes lung cancer without an experiment on people?
Why: The Bradford Hill questions, answered by many kinds of evidence together, built a causal case that chance could not match.
10. Why do the ozone and vaccine cases both count as science working well?
Why: One prediction was confirmed and one was refuted, but in both the weight of independent evidence settled the matter.
11. Which of these is a testable question?
Why: It names the variable changed, the outcome measured and the range tested; the others have no measurable terms or are questions of value.
12. A student tests fertilizer on tomatoes but gives the fertilized plants a sunnier spot. What kind of error is this?
Why: Sunlight now differs along with fertilizer, so any difference in yield cannot be attributed to fertilizer alone.
13. A pendulum model ignores air resistance and works well for small swings but fails for very wide ones. What does this show?
Why: Every model simplifies; its usefulness is bounded by the range where its assumptions are close enough to true.
14. A craft-stick bridge must span 40 centimeters and hold 5 kilograms within a materials budget. Which is a constraint?
Why: A constraint is a firm limit any design must stay within; the others are criteria for judging how good a design is.
15. Why should a capstone plan state, before data collection, what result would count against the hypothesis?
Why: Deciding in advance what would falsify the claim protects against fooling yourself, the easiest person to fool.
Send it to your teacher
Spiral review
Five questions from earlier units
0 / 5
1. (Unit 26) Why was the ozone problem easier to solve than climate change?
Why: The Montreal Protocol succeeded because replacements existed and the change was manageable; carbon dioxide comes from the core of the energy system.
2. (Unit 25) Dim ultraviolet light releases electrons from a metal but bright red light does not. This supports the idea that:
Why: Each electron takes energy from one photon. Red photons are below the threshold, and more of them do not help a single electron.
3. (Unit 24) A student walks 3 m east, then 3 m north, then 3 m west. What is the magnitude of her displacement?
Why: East and west cancel, leaving only the 3 m north. Distance is 9 m, but displacement is 3 m.
4. (Unit 23) In 2 NaN3 gives 2 Na + 3 N2, how many moles of N2 form from 1.0 mole of NaN3?
Why: The ratio of NaN3 to N2 is 2 to 3, so 1.0 mole of azide gives 1.5 moles of nitrogen.
5. (Unit 22) Which explains why water has a far higher boiling point than hydrogen sulfide?
Why: Hydrogen bonds between water molecules are much stronger than the forces between H2S molecules.
Send it to your teacher
Write it
Choose one of the unit's three cases: smoking and lung cancer, the ozone layer, or vaccines and autism. Write a claim about what the evidence shows, support it with at least three specific pieces of evidence from the unit, explain the reasoning that links them, and describe the strongest objection someone could raise and how the evidence answers it.
Claim: one sentence stating what the evidence shows, kept inside what the studies actually tested.
Evidence: use specific studies, numbers and dates from the chapter, such as sample sizes, dose-response patterns or measurements that confirmed a prediction.
Reasoning: explain why each piece of evidence supports the claim, using ideas like control groups, replication, mechanism or the Bradford Hill questions.
The other side: state the best counterclaim honestly, including who made it and why it seemed plausible, then show which evidence answers it.
Uncertainty: say what remained unknown at the time and what later evidence added.
0 wordsSaved on this device as you type.
Practice rooms
Rooms already on the site that belong to this unit — cards, quizzes, a lab.
Every lesson keeps its own three checks; a lesson is ticked when all three are right. Chapter reviews, the unit test and its spiral review (five questions from earlier units in this band) score on the page. When the site is connected to your sheet, or the link carries ?dest=, each one also has a Send box: the first-try score, the standards, the supports used, the attempt number and the minutes go to your sheet as an IEP data point.
Print this page for a paper copy of the readings, the sources, the words and the questions; the answers print as dashed boxes under each question.
Fact-check notes for this course live in the handoff: quotes marked (paraphrased) were set that way on purpose.