THIS EXPLANATION
THE ROOM
SPT·25 Sports, Exercise & Recreation 6 MIN · 8 STATIONS

Judging order effects

A Socratic walk-through of judging order effects — reasoned out one step at a time, not lectured.

abcdefgh
a

The question we started with

THE QUESTION #

Why does the same gymnastics routine score higher when it is performed late in the running order than early?

Take a routine, filmed once, and imagine it performed identically first in the running order and then eighth. Nothing about the gymnast has changed, and the judges are the same people applying the same code. Yet across judged sports — gymnastics, diving, figure skating, dressage, ski jumping's style marks, and music competitions besides — the later performance tends to draw the higher mark.

The lazy reading is that judges are biased or corrupt. But the effect turns up in panels with no stake in the outcome and no acquaintance with the competitors, which suggests something structural rather than something venal. What is it about going later that a scrupulous judge cannot escape?

b

Reasoning it through

REASONING #

Begin with what the judge is actually asked to do. Not to rank — ranking could wait until everyone had performed — but to map a single performance onto a number, immediately, and then not touch it again. Three conditions are buried in that sentence, and each does work.

First, the scale is bounded. A judge working an old-style ten-point code has a ceiling and cannot exceed it. Second, the mark is committed before the field is known: performance one is scored without any idea whether performance twelve will be sublime or a disaster. Third, marks are irreversible — a judge who later realises the first routine was the best of the day cannot go back and raise it.

Now ask what a careful judge should do under those constraints. Suppose you award a 9.7 to the opening routine and then watch four better ones. You have painted yourself into the top of the scale and can no longer separate them. The rational hedge is to hold something back — to score the early field slightly below what you privately believe, keeping headroom for whatever may come. That is score reservation, and it produces exactly the observed gradient without a shred of bad faith. The judge is not favouring late competitors. The judge is rationing a scarce resource, and early competitors are the ones who pay for the rationing.

There is a second mechanism worth separating out, because it makes different predictions. As the session runs, the judge's sense of what a routine of this class looks like sharpens against the routines already seen. Early marks are made against an imagined standard, later ones against an observed one — and the observed standard, in a field where the strong tend to be there, tends to sit higher than the imagined one. Call that reference drift. It would raise late marks even on an unbounded scale.

And a third factor is not a judging effect at all: in most sports the running order is not random. Seeding, qualification rank, or reverse-standings ordering put the better athletes later. So a raw correlation between start position and score is confounded from the outset, and much of the popular version of this claim rests on exactly that confound.

Which is why the cleanest evidence comes from settings where the order really is drawn by lot. Studies of the Queen Elisabeth music competition, whose performance order is randomised, found that later performers won more often than chance allows — with no seeding to explain it. That is the observation the account has to answer to.

c

The analogy

THE ANALOGY #
THE FIGURE

Think of a teacher marking a stack of essays one at a time, in ink, with no rereading and a hard cap of 100. The first essay is genuinely excellent — but excellent compared with what? Give it 96 and you have nowhere to put anything better. So you write 91, telling yourself you can always come back. You cannot come back. Twenty essays later you know precisely what the top of this stack looks like, and the last excellent essay gets its 96.

WHERE IT BREAKS DOWN

The teacher's caution is a private habit and can be trained away, whereas a judging panel's is enforced by rules — the marks are published, or averaged, or fed into a live standings board within seconds, so the option to revise does not merely go unused, it does not exist.

d

Clarifying the model

THE MODEL #

It is tempting to read this as a claim about judges being poor at their job. It is closer to the opposite: reservation is the correct response to being asked for an irreversible absolute mark under uncertainty about the remaining field. The defect is in the scoring architecture, not in the person.

The two mechanisms should also be kept apart, because they can be pulled apart empirically. Reservation depends on the ceiling; reference drift does not. So the natural test is to compare a bounded scale against an open-ended one. Figure skating's move away from the closed 6.0 system toward an open-ended points code is precisely that natural experiment, and if reservation is doing the work, the order gradient should have shrunk there while surviving in sports that kept a capped scale.

The sharper test is on irreversibility. Have judges score the same routines from video, in shuffled order, with the option to revise every mark after seeing all of them. If reservation is the mechanism, the gradient should collapse. If a randomised, fully revisable scoring session still shows late performances marked higher, score reservation is not the explanation and the effect belongs to reference drift or to plain contrast with the immediately preceding competitor.

I am deliberately quoting no effect size. The published estimates vary a great deal by sport, panel, and method, and the ones that look largest are usually the ones least protected against the seeding confound. What is reasonably firm is the direction and its appearance in randomised-order settings. How big it is, and how much of it survives modern judging rules, is genuinely unsettled.

e

A picture of it

THE PICTURE #
Judging order effects
Judging order effects Read downward as elapsed time in one session. The two solid arrows into the judge are two physically identical routines. What differs is only what the judge exchanges with the scale in between: early on, the scale reports uncommitted headroom that must be protected, so the mark comes back cautious. By the eighth performance the field's standard is known, the headroom no longer needs guarding, and the same work converts into a higher number. Nothing in the picture requires the judge to prefer the later gymnast. {"generator":"mermaid-svg-renderer@3.2.1","source":"../Socrates/.diagram-cache/_src/judging-order-effects.md","sourceIndex":1,"sourceLine":4,"sourceHash":"b9b02c24dca388a1f42f3d4758650bd01baa99aa241ecc80082d524bc7f92b27","diagramType":"sequence","layoutVariant":"source","repairedDuplicateIds":[],"motion":"entrance-with-reduced-motion-fallback","presentation":"editorial","attempt":1,"viewBox":{"x":0,"y":0,"width":1237,"height":684},"qa":{"passed":true,"findings":[]}} Late gymnast 01 The bounded scale 02 Judge 03 Early gymnast 04 Performs, first in the order Must commit a mark now, field unknown Only one band is left for the best of the day Cautious mark, top of the range held back Performs the same routine, eighth in the order Standard of the field now observed Held-back band can safely be spent Higher mark for identical work
KINDSlifelineparticipantmessage

How to readRead downward as elapsed time in one session. The two solid arrows into the judge are two physically identical routines. What differs is only what the judge exchanges with the scale in between: early on, the scale reports uncommitted headroom that must be protected, so the mark comes back cautious. By the eighth performance the field's standard is known, the headroom no longer needs guarding, and the same work converts into a higher number. Nothing in the picture requires the judge to prefer the later gymnast.

f

What became clearer

WHAT CLEARED #
WHAT CLEARED

The gradient is not a bias inside the judge but a cost imposed by the format: an absolute mark, on a capped scale, committed in ignorance of the rest of the field and never revisable. Early competitors subsidise the panel's uncertainty. That reframing also tells you where to intervene — not by lecturing judges about fairness, but by removing the ceiling, allowing revision, or scoring from video after the fact.

g

Where to go next

ONWARD #
  • Why open-ended scoring codes introduce their own pathology: the difficulty-inflation race that a capped scale suppressed.
  • How dropping the highest and lowest marks changes panel behaviour, and why it does not touch an effect that moves the whole panel in one direction.
  • Whether the same reservation logic explains grade inflation across a marking season, where the field is even less knowable.
h

Key terms

TERMS #
TermWhat it means
Score reservationholding back the top of a bounded scale early in a session, to preserve the ability to separate better performances later.
Reference driftthe sharpening and rising of a judge's internal standard as the session supplies real examples of the class being judged.
Running orderthe sequence in which competitors perform, often set by seeding or qualification rank rather than by lot.

Every term the collection defines is gathered in the glossary.

Nearby on the shelf

4