Scientific testability
A Socratic walk-through of scientific testability — reasoned out one step at a time, not lectured.
The question we started with
THE QUESTION #What makes a claim genuinely testable rather than merely compatible with any possible result?
We praise a theory for how much it explains. Notice the assumption: that fitting the facts is evidence of truth. But consider a claim that fits every fact — markets rose because sentiment was bullish, markets fell because sentiment turned. Whatever happened, the claim was waiting for it. Did it ever risk anything? And if it risked nothing, what did the world's cooperation tell us?
So the interesting property may not be how much a claim explains, but how much it forbids.
Reasoning it through
REASONING #Try measuring a claim by its prohibitions. "Something will happen tomorrow" forbids nothing and is worthless. "The star's apparent position will shift by about 1.75 arcseconds during the eclipse" forbids nearly every other number — and had the plates shown no shift, the claim would have been in serious trouble. That asymmetry is Popper's insight: a theory earns its keep by sticking its neck out, and confirmations only count when they were the sort of thing that could have gone badly wrong.
Now push on it, because this is where the tidy version breaks. Suppose the prediction fails. Have you refuted the hypothesis? Ask what you actually deployed to generate the prediction. Not the hypothesis alone — also that the instrument reads true, that the sample was uncontaminated, that no unaccounted body or field intervened, that the mathematics was applied correctly. The prediction came from the bundle. So a failure tells you the bundle contains an error. Does it tell you where? It cannot. This is the Duhem-Quine problem, and it is not a technicality; it is the ordinary condition of every experiment ever run.
Which means the falsifying blow can always be absorbed somewhere else. Is that cheating? Consider two cases from the same century. Uranus wandered from its predicted orbit; rather than abandon Newtonian gravitation, astronomers blamed an auxiliary — there must be an unseen planet. Neptune was found in 1846, roughly where the rescue said to look. Encouraged, the same move was made for Mercury's stubborn perihelion: another unseen planet, Vulcan. Nobody ever found it, and the anomaly was eventually settled by general relativity instead. Identical logical move, opposite verdicts. So what distinguished them? Not the form of the rescue — the content. One rescue made a fresh, risky, independently checkable prediction. The other, as the searches failed, had to keep shrinking and hiding its planet, buying survival with no new commitments.
The analogy
THE ANALOGY #Think of a court where the defendant is not a person but a whole partnership. The alibi collapses, so you know the partnership is guilty of something — but the verdict names no member. You may always convict the most junior clerk and let the senior partners walk. Do that once with good reason and justice is served; do it every single time, always sacrificing whoever is nearest the door, and observers will start to suspect the firm itself.
A court eventually returns a verdict and the case ends, whereas science has no judge and no closing date — the "firm" is only discredited when working scientists collectively lose patience, which is a sociological event, not a logical one.
Clarifying the model
THE MODEL #Two corrections follow. First, naive falsificationism — one failed prediction, one dead theory — is not the settled view in philosophy of science, and it never described real practice; scientists live with anomalies for decades and are usually right to. Second, the criterion is not thereby empty. The surviving line is not "can this be refuted by a single result?" but something closer to Lakatos's question: when the programme repairs itself, does it generate new testable content, or merely re-describe what already happened?
So testability is best treated as a matter of degree, and as a property of a research programme over time rather than of an isolated sentence. That weakens Popper's sharp demarcation line, and it is contested — some think it dissolves the criterion, others that it makes it usable for the first time.
A picture of it
THE PICTURE #How to readStart at the top box and follow the two arrows into the slanted node — the forbidden result comes from the hypothesis and the hexagon of auxiliary assumptions together, never the hypothesis alone. The first diamond is the experiment; the second is the Duhem-Quine question, and nothing in the diagram tells you which branch to take. The loop back from "revise an auxiliary" is legitimate science — that is how Neptune was found. The rounded node beneath it is what the same loop becomes when taken every time and never yielding a new prediction.
What became clearer
WHAT CLEARED #A claim is testable to the extent that it forbids things — but no experiment tests a claim by itself, only the claim bundled with everything else assumed to get a prediction out. A failure therefore never names its own culprit, and blaming an auxiliary is always available. What separates honest science from unfalsifiable dogma is not whether that rescue is used, but whether the rescue itself sticks its neck out.
Where to go next
ONWARD #- Whether Kuhn's paradigms make demarcation the wrong question to ask.
- How Bayesian confirmation distributes blame across the bundle by prior probability.
- Why pre-registration is, in effect, an institutional device for restoring risk.
Key terms
TERMS #| Term | What it means |
|---|---|
| Falsifiability | Popper's criterion: a claim is scientific insofar as it forbids observable outcomes. |
| Auxiliary assumptions | the background claims about instruments, samples and conditions needed to derive any prediction. |
| Duhem-Quine thesis | the argument that hypotheses face the evidence only in bundles, so a failed test cannot single out the guilty member. |
| Degenerating research programme | Lakatos's term for a programme whose repairs only accommodate known results rather than predicting new ones. |
Every term the collection defines is gathered in the glossary.