THIS EXPLANATION
THE ROOM
EDU·31 Education & Learning 7 MIN · 8 STATIONS

Teacher expectations

A Socratic walk-through of teacher expectations — reasoned out one step at a time, not lectured.

abcdefgh
a

The question we started with

THE QUESTION #

Why can a teacher's private guess about a pupil end up matching that pupil's later results?

In September a teacher writes a private note: this one will struggle, that one will fly. In June the results arrive and the note was right. The story everyone reaches for is that the note made it right — that expectation shapes reality, and a child written off in September is finished by June.

But hold on. What we have observed is a match between a guess and an outcome, and a match is only a correlation. A correlation can be produced by the guess causing the outcome, by the outcome's real causes producing the guess, or by something producing both. The popular story picks one of the three and never mentions the others. Which one is actually doing the work?

b

Reasoning it through

REASONING #

Take the accuracy possibility seriously first, because it is the one that gets skipped. What does a teacher have in September? Several weeks of work, how a child speaks, whether homework arrives, how quickly a new idea lands, often the previous school's record. Those are not superstitions — they are the same variables that will drive the June result, and teacher judgements do correlate quite well with later externally marked tests. So a substantial part of the match needs no causal story at all.

Now the causal possibility. If a private thought is to move a public result, there must be a physical channel, so ask what a teacher actually does differently. Rosenthal's summary is four-fold: a warmer climate toward the pupil expected to do well, more and harder material, more turns to answer and longer patience waiting for the answer, and more detailed feedback rather than a bare mark. None of that is mysterious — it is more instruction, more practice, more correction, delivered unevenly.

Then the famous study. Rosenthal and Jacobson in 1968 told teachers that certain randomly chosen pupils were about to bloom, and reported that those pupils gained more in IQ. It is the origin of the whole idea, and it is weaker than its fame: the gains were concentrated in the youngest two year groups and were slight or absent higher up, and the measure has been much criticised. It shows something happened, not that expectations are powerful.

The result that actually explains the mechanism came later, from Raudenbush's meta-analysis of experiments that planted false expectations in teachers. The effect turned out to depend on when the planting happened. Told before meeting the class, or in roughly the first fortnight, teachers' behaviour shifted and pupil outcomes moved. Told after they had already taught the children for a while, the false label did almost nothing.

That is the hinge of the whole topic, so let us read it properly. Why should a fortnight matter? Because an expectation is a belief, and beliefs compete with evidence. In week one the label is the only information there is. By week six the teacher has watched the child solve fifty problems, and a stranger's claim cannot outweigh that. The self-fulfilling effect is therefore not a general property of teaching — it is what happens in the gap before real evidence arrives.

Which lets us assemble the answer. The match in June is mostly accuracy, with a real but modest causal contribution — reviews of the naturally occurring case, Jussim and Harber's especially, put typical expectancy effects at the small end, the size implied by correlations of around a tenth to two-tenths, and find that they tend to fade rather than accumulate across years. That last point matters: the runaway-loop picture predicts growth over time, and growth over time is largely not observed.

Let me try to break my own account. If accuracy were the whole story, planting a false expectation should do nothing at all — and it does something, in the early window. So the pure-accuracy story fails too. Both channels are real. The question was only ever their relative size.

c

The analogy

THE ANALOGY #
THE FIGURE

A doctor glances at a patient in the waiting room and privately predicts a poor recovery. In June the prediction is right. Almost all of that is diagnosis — the pallor and the gait were evidence. A little of it may be treatment: a pessimistic doctor may push less hard on rehabilitation. Both are present, and confusing them makes the doctor's skill look like a curse.

WHERE IT BREAKS DOWN

a doctor's prognosis is checked against measurements they did not produce, whereas much of what a teacher uses to check an expectation — classwork, effort grades, the mark on an essay — is generated by the teacher themselves, which is precisely the condition under which a loop can run rather than correct.

d

Clarifying the model

THE MODEL #

Three refinements connect the pieces.

First, the loop closes only when the feedback is contaminated. If a pupil's evidence comes back through an externally marked assessment, the teacher's belief is disciplined by something they did not create. If it comes back as their own grading of their own tasks, the belief marks its own homework. That is where I would expect self-fulfilling effects to be strongest, and it is a claim about assessment design rather than about teacher virtue.

Second, "small on average" is not "small everywhere". The reviews report larger effects for pupils who are low-achieving or from stigmatised groups — plausibly because the real signal is more ambiguous there, leaving more room for the prior to steer. An average effect can conceal a serious one in a subgroup.

Third, a methodological caution shared with any expectation research: children are often noticed at their extremes, and a pupil measured at a low point will tend to move toward their own average regardless of anyone's belief. Regression to the mean flatters and flatters alike, so uncontrolled classroom observations of expectation effects cannot be trusted.

Softest to firmest: the share of the September-to-June match attributable to causation is not cleanly quantified and I decline to give a figure. The correlation range implied above is a summary of reviews and should be held loosely. The dependence on prior contact is firm in direction and less firm in magnitude. That teacher judgements predict external tests reasonably well, and that the original Pygmalion result was concentrated in the youngest grades, are both solid.

e

A picture of it

THE PICTURE #
Teacher expectations
Teacher expectations Read downward as time. The first arrow is the accuracy path -- real evidence flowing from pupil to teacher before any expectation exists. The boxed loop is the causal path, and note that every arrow in it is an ordinary teaching behaviour rather than anything transmitted invisibly. The branch at the bottom is the finding that decides the whole question: the upper branch is the ordinary case, where an external result the teacher did not produce corrects the belief, and the lower branch is the narrow early window in which a planted label actually steers outcomes. {"generator":"mermaid-svg-renderer@3.2.1","source":"../Socrates/.diagram-cache/_src/teacher-expectations.md","sourceIndex":1,"sourceLine":4,"sourceHash":"34150a3239cccaa646c74c79932d7a9241c153ef3fda9916544b786f7370a07f","diagramType":"sequence","layoutVariant":"source","repairedDuplicateIds":[],"motion":"entrance-with-reduced-motion-fallback","presentation":"editorial","attempt":1,"viewBox":{"x":0,"y":0,"width":872,"height":996},"qa":{"passed":true,"findings":[]}} External test 01 Teacher 02 Pupil 03 forms a private expectation loop [every week of term] belief corrected, accuracy path dominates alt [teacher already has independent evidence] [first fortnight, evidence still thin] prior work, speech, homework 1 warmer climate, harder material 2 more turns to answer, longer wait 3 fuller answers, more practice done 4 sits an externally marked paper 5 a result the teacher did not set 6 the label alone steers the treatment 7
KINDSlifelineparticipantalternativemessage

How to readRead downward as time. The first arrow is the accuracy path — real evidence flowing from pupil to teacher before any expectation exists. The boxed loop is the causal path, and note that every arrow in it is an ordinary teaching behaviour rather than anything transmitted invisibly. The branch at the bottom is the finding that decides the whole question: the upper branch is the ordinary case, where an external result the teacher did not produce corrects the belief, and the lower branch is the narrow early window in which a planted label actually steers outcomes.

f

What became clearer

WHAT CLEARED #
WHAT CLEARED

The September note usually matches June because it was a decent reading of real evidence, not because it cast a spell. A genuine causal channel exists and runs through unequal teaching — climate, material, turns, feedback — but it is small on average, largest exactly where the teacher's evidence is thinnest, and it tends to fade rather than compound. The load-bearing claim is that most of the correlation is accuracy, and the self-fulfilling part is conditional on the absence of better information — which is why the practical lever is not asking teachers to think positively but making sure their beliefs keep meeting evidence they did not generate.

g

Where to go next

ONWARD #
  • Why expectancy effects are reported as larger for stigmatised groups, and whether ambiguity of signal explains it.
  • How externally marked assessment changes the strength of the loop compared with teacher-set grading.
h

Key terms

TERMS #
TermWhat it means
Self-fulfilling prophecya belief that causes the behaviour which makes it true, as distinct from a belief that merely predicts accurately.
Pygmalion effectthe specific claim, from Rosenthal and Jacobson's 1968 study, that induced teacher expectations raise pupil attainment.
Regression to the meanthe tendency of an extreme measurement to be followed by one nearer the average, with no cause involved.

Every term the collection defines is gathered in the glossary.

Nearby on the shelf

4