THIS EXPLANATION
THE ROOM
LAN·31 Language, Media & Communication 6 MIN · 8 STATIONS

Transcribed speech

A Socratic walk-through of transcribed speech — reasoned out one step at a time, not lectured.

abcdefgh
a

The question we started with

THE QUESTION #

Why does an ordinary conversation look ungrammatical the moment it is written down word for word?

Read a genuinely verbatim transcript of a conversation you took part in and enjoyed. On the page it is a wreck: sentences started twice, clauses abandoned in mid-air, filled pauses everywhere, a phrase begun and swapped for another.

Yet nobody in the room noticed anything wrong, and you would testify that the speaker was articulate. Both records describe the same minutes. So which is mistaken — the page, or your memory of the room?

b

Reasoning it through

REASONING #

First, settle whether the mess is real or an artefact of a hostile transcriber. It is real. Speech is planned while it is being produced, under a clock nobody controls, and the planning routinely lags the talking. Counts of ordinary conversation commonly land near six disfluencies per hundred words, which is roughly one every couple of sentences. Repairs even have a regular shape, described by Levelt: the stretch to be replaced, then an editing signal — uh, sorry, I mean — then the correction.

So the page is accurate. Why then did nobody hear it?

Because hearing speech is not transcription. Consider Warren's demonstration: splice a cough over a speech sound in a recorded sentence, and listeners report hearing the missing sound perfectly clearly, and cannot say where the cough fell. The system is not passing sound through to consciousness for inspection. It is producing the most probable utterance and handing you that.

Now add a second fact. Sachs showed that memory for exact wording decays within seconds of hearing it, while the meaning stays. So even if you had noticed a false start, you could not report it a minute later; the surface form is already gone and only the reconstructed content remains.

Put the two together and the question dissolves. The repair happens during comprehension, before anything reaches awareness, and the record you keep afterwards is of the output rather than the input. Your memory of an articulate speaker is a memory of your own parser's finished product.

There is a further turn, and it changes the moral. The disfluencies are not noise the listener heroically filters out; they carry information. Clark and Fox Tree argued that uh and um are not accidents but signals of how long a delay to expect, and listeners use a filled pause to anticipate something new or hard to name. An editing signal announces that a repair is coming, so you know to discard what preceded it.

That reframes the transcript. It is not a faithful record with warts. It is a record made in a medium with no place for a whole layer of what was there: timing, pitch, loudness, overlap, gaze, gesture. Strip the channel that made the false starts interpretable and only the false starts remain.

And there is the yardstick problem. Speech is not written prose delivered aloud. It is built in short intonation units, chained, extended and revised as it goes, with constructions that are entirely ordinary in speech and simply not available in writing. Judging a transcript by the sentence is judging speech against a standard it was never aiming at.

Which is why almost every institution that produces transcripts cleans them. Court reporters work to conventions that drop most of it, and journalists tidy quotes as a matter of routine. That cleaning is defensible — it recovers the utterance the listener actually received — but it is an editorial act, not a neutral one. Someone decides which start was abandoned and which was meant, and the decision is not always obvious. It reached the United States Supreme Court in Masson v. New Yorker Magazine, which had to rule on when altering a quotation crosses into falsity; the answer turned on whether the change materially altered the meaning, which is a judgement call, exactly as the tidying is.

One asymmetry is worth naming, because it does damage. Disfluency rises with planning difficulty, unfamiliarity and stress. So a nervous witness, a second-language speaker, or anyone taken by surprise transcribes worse than a rehearsed professional saying much less — and a reader with no access to the room reads that difference as intelligence.

c

The analogy

THE ANALOGY #
THE FIGURE

Watching someone speak is watching a building go up. You see scaffolding, a beam hoisted and rejected, a wall started in the wrong place and taken down — but what you take away is the building, because that is what you were watching it become. A verbatim transcript is a photograph of the site, scaffolding and rubble included, handed to someone who never saw the building.

WHERE IT BREAKS DOWN

scaffolding is removed after the fact, whereas the disfluency is never removed from the record at all — it simply was never perceived as separate from the structure; and unlike scaffolding, the false starts and filled pauses were load-bearing for the listener, telling them what was coming.

d

Clarifying the model

THE MODEL #

The temptation is to say the listener "ignores" or "forgives" the errors. Neither is right, and the difference matters. Nothing is experienced as an error in the first place. There is no moment of noticing followed by charity; the repaired version is the only version that arrives.

Nor does this make transcripts wrong. A verbatim transcript answers a real question — what sounds were produced — with high fidelity. It just is not the question a reader thinks they are asking, which is what the speaker said. Two different questions, two different correct answers, and confusion only when one is served up as the other.

A calibration, finally. Not all disfluency is invisible. A speaker who stalls constantly, or repairs the same phrase four times, does register, and listeners draw conclusions from it. The claim is about the ordinary rate in fluent speech, not a promise that the parser hides everything.

e

A picture of it

THE PICTURE #
Transcribed speech
Transcribed speech Read the top half first: the speaker sends three things, and the two self-arrows on the listener are the repair, with the note recording why nothing survives to be reported later. Then compare the bottom half, where the transcript receives the same syllables minus the channel that made them interpretable, so the reader is handed the abandoned starts with none of the cues that told the listener to drop them. Nothing was added to the transcript -- something was subtracted. {"generator":"mermaid-svg-renderer@3.2.1","source":"../Socrates/.diagram-cache/_src/transcribed-speech.md","sourceIndex":1,"sourceLine":4,"sourceHash":"6a1e9c1fe7f316907593ce2fe3149fd827c876ea9f514b608da1b40d89bfacba","diagramType":"sequence","layoutVariant":"source","repairedDuplicateIds":[],"motion":"entrance-with-reduced-motion-fallback","presentation":"editorial","attempt":1,"viewBox":{"x":0,"y":0,"width":1458,"height":842},"qa":{"passed":true,"findings":[]}} Reader 01 Transcript 02 Listener 03 Speaker 04 what is stored is the meaning, not the wording starts a clause and abandons it 1 editing signal, uh or I mean 2 the repaired clause, with timing and pitch 3 discards the abandoned start 4 reads the filled pause as a cue, not a fault 5 every syllable, but no timing and no pitch 6 the abandoned starts, with no cue about what to drop 7 no repair is possible, so the speaker reads as incoherent 8
KINDSlifelineparticipantmessage

How to readRead the top half first: the speaker sends three things, and the two self-arrows on the listener are the repair, with the note recording why nothing survives to be reported later. Then compare the bottom half, where the transcript receives the same syllables minus the channel that made them interpretable, so the reader is handed the abandoned starts with none of the cues that told the listener to drop them. Nothing was added to the transcript — something was subtracted.

f

What became clearer

WHAT CLEARED #
WHAT CLEARED

Fluent speech is not clean, and the cleanliness we remember is our own work. The parser repairs in real time, using signals the disfluencies themselves provide, and stores the meaning rather than the wording — so the repair leaves no trace to compare against.

A verbatim transcript therefore looks broken not because the speech was, but because it renders speech into a medium missing half the original channel and judges it against a grammar the speaker never targeted. The cleaning that fixes this is both genuinely necessary and genuinely editorial, which is an uncomfortable pair of facts for anyone who quotes people for a living.

g

Where to go next

ONWARD #
  • How conversation-analytic transcription preserves timing and overlap, and what it trades away.
  • Why automatic speech recognition distorts differently again from a human transcriber.
h

Key terms

TERMS #
TermWhat it means
Disfluencyan interruption in the flow of speech: a filled pause, repetition, false start or self-repair.
Filled pauseuh, um and their equivalents, which signal a delay and its expected length rather than merely marking hesitation.
Phoneme restorationthe effect in which a speech sound masked by noise is heard as present, and the noise cannot be located.
Verbatim transcripta record of what was uttered including disfluencies, as distinct from the cleaned transcripts used in courts and journalism.

Every term the collection defines is gathered in the glossary.

Nearby on the shelf

4