THIS EXPLANATION
THE ROOM
ENG·04 Engineering & Technology 6 MIN · 8 STATIONS

Bathtub failure curve

A Socratic walk-through of the bathtub failure curve — reasoned out one step at a time, not lectured.

abcdefgh
a

The question we started with

THE QUESTION #

Why is a brand-new machine more likely to break this week than a six-month-old one?

Take a pump out of its crate and a pump that has run since January. Which is more likely to fail this week? The new one, and often by a wide margin. Reliability engineers have a name for that early stretch of elevated, falling failure rate — infant mortality — and it is the left-hand wall of the bathtub curve.

The explanation everyone reaches for is that a machine needs to bed in: it starts fragile and toughens up with use. That is worth testing rather than repeating, because it makes a strong claim. It says a particular machine becomes less likely to fail as it runs. What, physically, would have got better?

b

Reasoning it through

REASONING #

Run through what use does to hardware. Surfaces abrade, fasteners loosen, seals harden, insulation embrittles, metal accumulates fatigue cycles. Nearly every process you can name runs one way. A few genuinely do improve early — bearing surfaces conforming to each other, a seal seating properly, electrical contacts wiping clean — and it would be wrong to say bedding-in is a myth. But it is a minority of mechanisms, and it cannot carry a curve that falls this steeply across nearly everything manufactured, including electronics with no moving parts at all.

So try assuming the opposite: that no individual unit improves. Can the failure rate still fall?

Here is the move that unlocks it. The falling rate is not a measurement of a machine. It is a measurement of a batch. And a batch is not uniform — it is a mixture. Most units left the factory built correctly. A small fraction left with a latent defect: a cold solder joint, an O-ring nicked on assembly, a bolt at the wrong torque, a contaminant trapped in a casting. Those two groups have completely different prospects, and only one of them is fragile.

Now watch the mixture over time. The defective units fail quickly, because their defects are the sort that show up under the first load cycles. The sound ones do not. Week by week, the survivors contain a smaller and smaller share of defectives — not because anything healed, but because the weak members have been removed from the population being counted.

The arithmetic is sharper than the picture suggests. Suppose you have two subpopulations, each with a constant failure rate — no ageing whatsoever in either group, so no individual item's prospects change by a hair. Mix them and the observed failure rate of the mixture declines over time regardless, because the high-rate members are eliminated first and the average drifts towards the low-rate ones. Decreasing risk, built entirely out of components with no decreasing risk in them.

That reframes the original question. The six-month-old machine is not tougher than the new one. It is evidence. It has sat an examination the new one has not sat, and passed. "It has run since January" is data about which subpopulation it was drawn from.

Three things follow that the bedding-in story cannot explain. Burn-in — running units hard before shipping them — works, and what it does is cull rather than strengthen; the manufacturer is deliberately spending some of the good units' life to buy out the defective ones before they reach a customer. Replacing a working part with a fresh one moves you from a proven unit back into the unfiltered population, which is why "we serviced it and then it broke" is a real pattern rather than bad luck. And early failures should be concentrated in a minority of units and traceable to specific defects rather than spread evenly across the fleet — which is what failure analysis actually finds when it opens them up.

One honest complication: a large share of early failures are not the machine's fault at all. They are installation errors, wrong commissioning settings, and an operator who has not yet learned the thing. That produces the same falling shape — a fixed stock of mistakes surfacing early and then running out — from a completely different cause.

c

The analogy

THE ANALOGY #
THE FIGURE

Think of a hundred people hired on the same day. Over the first year, departures are frequent and then taper off. It is tempting to say the job gets easier or people grow loyal, but most of it is simpler: the badly matched leave early, so the group that remains is increasingly made of people who fit. Nobody changed; the composition did.

WHERE IT BREAKS DOWN

People genuinely do adapt and build ties, so there is real individual change mixed into that cohort's numbers — whereas for most components there is no equivalent improvement at all, which makes the selection story the whole of the explanation rather than part of it.

d

Clarifying the model

THE MODEL #

This cuts against the way the bathtub is usually taught. Drawn as a single curve against age, it invites you to read it as one machine's biography: fragile youth, steady middle age, worn-out old age. It is nothing of the kind. It is a composite — the left wall belongs to a population containing defectives, the flat floor to failures triggered by external events rather than by age, and the right wall to genuine accumulating wear. Different mechanisms, different subsets of items, superimposed on one axis.

And most items never show all three. In the study that founded reliability-centred maintenance, Nowlan and Heap's 1978 analysis of United Airlines component data, failure modes sorted into six patterns, and the full bathtub was the rarest but one. The dominant pattern, at about 68 percent, was early failures followed by a flat line that never rises — items for which there is no wear-out region to schedule against at all. Only around 11 percent showed any age-related wear-out zone. Those figures are much quoted and much over-extrapolated: they describe aircraft components under one airline's maintenance regime, not machinery in general.

e

A picture of it

THE PICTURE #
Bathtub failure curve
Bathtub failure curve Each slice is the share of failure modes whose failure rate behaved a given way as the item aged -- not the share of failures, which is a different quantity. The largest slice by far is the one that answers the question: for most items the rate starts high and then flattens permanently, so a new unit really is the risky one and age never makes it risky again. Add the bottom two slices plus the slowly-rising one and you have the only items for which "it is old, replace it" means anything -- roughly a ninth of the total. {"generator":"mermaid-svg-renderer@3.2.1","source":"../Socrates/.diagram-cache/_src/bathtub-failure-curve.md","sourceIndex":1,"sourceLine":4,"sourceHash":"46030668468d66bfc8e939737834257b804be708693f520b2adfdb329351f355","diagramType":"pie","layoutVariant":"source","repairedDuplicateIds":[],"motion":"entrance-with-reduced-motion-fallback","presentation":"editorial","attempt":1,"viewBox":{"x":0,"y":0,"width":920,"height":545},"qa":{"passed":true,"findings":[]}} 68% 14% 7% 5% 4% 2% TOTAL 100 Failure patterns in the United Airlines data, Nowlan and Heap 1978 Early failures then constant 68 Constant, unrelated to age 14 Low at first, then constant 7 Slowly rising with age 5 Full bathtub 4 Constant then wear-out 2

How to readEach slice is the share of failure modes whose failure rate behaved a given way as the item aged — not the share of failures, which is a different quantity. The largest slice by far is the one that answers the question: for most items the rate starts high and then flattens permanently, so a new unit really is the risky one and age never makes it risky again. Add the bottom two slices plus the slowly-rising one and you have the only items for which "it is old, replace it" means anything — roughly a ninth of the total.

f

What became clearer

WHAT CLEARED #
WHAT CLEARED

Infant mortality is a selection effect, not a maturation effect. New machines fail more because every batch contains a few that were born broken, and time removes them; the survivors look sturdier only because the weak ones have already gone. That is why burn-in is culling rather than toughening, why an unrepaired machine that has run for months is safer than a freshly serviced one, and why the bathtub curve should be read as a picture of a population rather than the life story of a thing.

g

Where to go next

ONWARD #
  • How burn-in duration is chosen, given that it consumes life from the good units too.
  • Why some software shows the same early-failure shape despite nothing physical wearing out.
  • What changes in the reasoning when a fleet is repaired rather than replaced, so units re-enter the population part-aged.
h

Key terms

TERMS #
TermWhat it means
Failure rate (hazard rate)the chance an item fails in the next interval given that it has survived until now.
Infant mortalitythe early period in which the failure rate of a population falls as defective units are eliminated.
Burn-indeliberately operating units before delivery to precipitate the failure of defective ones.
Mixture of populationsa group made of subgroups with different failure rates, whose combined rate falls over time even when no subgroup's does.

Every term the collection defines is gathered in the glossary.

Nearby on the shelf

4