THIS EXPLANATION
THE ROOM
WRK·06 Work, Careers & Skilled Trades 7 MIN · 8 STATIONS

Blameless review

A Socratic walk-through of blameless review — reasoned out one step at a time, not lectured.

abcdefgh
a

The question we started with

THE QUESTION #

Why does holding someone accountable for an outage make the next one more likely?

An engineer runs a command with the wrong flag and takes down payments for forty minutes. The organisation does the thing that feels like seriousness: it names the person, notes it in their file, and requires a sign-off from a manager before anyone runs that command again.

Everything about that is defensible in isolation. And yet organisations that respond this way tend to have more serious incidents afterwards, not fewer. So there is something wrong with the reasoning, and it is worth finding out exactly where — because the case for blame is not stupid, it is just built on an assumption that does not hold.

b

Reasoning it through

REASONING #

The assumption is this: that the outage happened because that person was insufficiently careful, and that raising the personal cost of carelessness will raise care. Let us test each half.

Take the first half. Ask what the engineer knew at the moment they pressed enter. Not what we know now, but what was visible then. Almost always the command looked ordinary, the flag looked right, the environment looked like the one they meant, and nothing on the screen suggested otherwise. This is what safety researchers call the local rationality principle: people's actions make sense given their goals, their attention, and the information actually in front of them. The action looks reckless only from the far side of the outcome, once we know which of the many ordinary-looking moments turned out to matter.

That knowledge changes what we can see. Hindsight bias makes the path to the failure look obvious and inevitable after the fact — the missed signal now glows, though it sat in a stream of identical-looking signals at the time. Outcome bias does something further: it makes us judge the decision by how it turned out. Notice that the same command with the same flag, run on a quieter day, would have caused nothing and been reviewed by nobody. If the judgement of the act depends on the outcome rather than on the act, we are not really assessing carefulness at all.

Now the second half — does raising the personal cost raise care? Consider what the engineer's colleagues learn from watching. They learn that reporting a near miss is dangerous, because a near miss is an outage that merely happened to land well and will be judged once someone knows about it. So they stop volunteering the small stuff: the confusing tool, the mistake caught at the last second, the workaround everyone quietly uses. Do you see what the organisation has just lost? Not the errors — those continue — but the reports of them. The rate of visible incidents falls, which reads as improvement, while the reservoir of unexamined hazard grows.

There is a second loss during the incident itself. In a blaming culture, the first person to notice a fault has a private incentive to check whether they caused it before saying anything. That is minutes, sometimes hours. And in the review afterwards, the person who understands the failure best — the one with their hands on it — is precisely the one who must now be careful about what they say, so the account you get is defensible rather than complete.

Set against that, what does blame buy? Usually only the belief that the cause has been addressed. And here the deeper point arrives: the flag was accepted by a tool that offered no confirmation, in an environment indistinguishable from production, without a dry-run mode, at a time when it was possible for one command to remove payments entirely. All of those conditions survive the individual's departure. Removing the person removes the last link in a chain and leaves the rest of the chain in place, wearing a new name.

This is not a purely theoretical position. Aviation acted on it deliberately: NASA has run the Aviation Safety Reporting System since 1976 as a confidential channel, with limited enforcement immunity for reporters, on the explicit reasoning that near-miss information is worth more than punishing the people best placed to supply it. The evidence base is largely from aviation, healthcare and case studies rather than controlled trials — worth saying plainly — but it points consistently in the same direction.

One caution, because "blameless" is routinely misheard. It does not mean nothing has consequences. The just-culture framing separates honest error, at-risk behaviour where a hazard was underestimated, and genuine recklessness where a known risk was knowingly disregarded. The first is a systems problem; the second calls for coaching and guardrails; the third can carry sanction. What blameless review removes is punishment for the ordinary error any competent person would have made in that position.

c

The analogy

THE ANALOGY #
THE FIGURE

Think of a smoke alarm you can only trigger by admitting you left the pan on. If every activation means a fine, the household's alarms all go quiet — and the quiet is not evidence of fewer near-fires. It is evidence that the reporting channel has been closed by the very people whose reports you needed.

WHERE IT BREAKS DOWN

a smoke alarm reports a hazard that already exists whether or not anyone speaks, whereas an incident review is generative — the account itself is what turns a confusing hour into knowledge — so what blame destroys is not just detection but the organisation's only means of learning.

d

Clarifying the model

THE MODEL #

Three refinements.

Blamelessness is not the same as no accountability. The engineer is entirely accountable — for narrating what happened honestly, and for helping fix the conditions that made it possible. That is a heavier obligation than being written up, and one the person can only meet if it is safe to.

The claim is not that individuals never matter. It is that the individual is the least durable of the causes. A review that ends at "human error" has stopped at the first plausible explanation rather than the most useful one, and it will produce the same class of incident again from a different pair of hands.

Finally, culture here is not a poster but a set of observable behaviours: what happened to the last person who admitted a mistake determines what any future review can discover, and no stated policy overrides it.

e

A picture of it

THE PICTURE #
Blameless review
Blameless review Both branches start from the same honest account, so the difference is entirely in what the review does with it. Trace the upper branch and watch the arrows back to the engineer close the reporting channel, leaving the conditions untouched for the next incident. Trace the lower branch and the same information ends up in organisational memory as a changed tool, with the loop back to the engineer keeping the channel open. The messages after the branch, not the ones before it, are where the two futures diverge. {"generator":"mermaid-svg-renderer@3.2.1","source":"../Socrates/.diagram-cache/_src/blameless-review.md","sourceIndex":1,"sourceLine":4,"sourceHash":"c6c5a01b28978ccfd5a1449df9b0b9b8e1f16e65be226b01c521883cda3f6887","diagramType":"sequence","layoutVariant":"source","repairedDuplicateIds":[],"motion":"entrance-with-reduced-motion-fallback","presentation":"editorial","attempt":1,"viewBox":{"x":0,"y":0,"width":1339,"height":876},"qa":{"passed":true,"findings":[]}} "The next incident" 01 "Organisational memory" 02 "Incident review" 03 alt [review names a culprit] [review names the conditions] E Engineer full account of what was on the screen 1 sanction recorded 2 withholds near misses from now on 3 the enabling conditions are still there 4 same failure, different hands, found later 5 confirmation prompt, dry run, blast radius limit 6 reporting stays cheap and safe 7 near misses arrive before they become outages 8 this class of failure is harder to reach 9
KINDSlifelineparticipantalternativemessage

How to readBoth branches start from the same honest account, so the difference is entirely in what the review does with it. Trace the upper branch and watch the arrows back to the engineer close the reporting channel, leaving the conditions untouched for the next incident. Trace the lower branch and the same information ends up in organisational memory as a changed tool, with the loop back to the engineer keeping the channel open. The messages after the branch, not the ones before it, are where the two futures diverge.

f

What became clearer

WHAT CLEARED #
WHAT CLEARED

Blame feels like accountability because it produces a consequence, but the consequence lands on the one part of the system that will not be there next time — while removing the information flow you needed to fix the parts that will. A review's yield is measured in what people are willing to tell you afterwards, and blame is the most reliable way to reduce it to nothing.

g

Where to go next

ONWARD #
  • Just culture, and how it draws the line between honest error and knowing recklessness.
  • Why "human error" is best treated as the starting point of an investigation rather than its finding.
  • Counterfactual language in postmortems ("should have noticed"), and why it quietly reintroduces hindsight.
h

Key terms

TERMS #
TermWhat it means
Blameless postmorteman incident review conducted so that participants can describe their actions and reasoning without fear of punishment, in order to surface the conditions that made the failure possible.
Local rationalitythe principle that people's actions were sensible given their goals, attention, and the information available to them at that moment.
Hindsight biasthe tendency, once an outcome is known, to see the path to it as obvious and the warning signs as conspicuous.
Just culturea framework distinguishing honest error, at-risk behaviour, and reckless behaviour, applying different responses to each.

Every term the collection defines is gathered in the glossary.

Nearby on the shelf

4