THIS EXPLANATION
THE ROOM
ART·41 Arts, Design & Culture 6 MIN · 8 STATIONS

Voice recording mismatch

A Socratic walk-through of the voice recording mismatch — reasoned out one step at a time, not lectured.

abcdefgh
a

The question we started with

THE QUESTION #

Why does your own recorded voice sound like a stranger's while everyone else says it sounds normal?

Play someone a recording of themselves and the reaction is remarkably consistent: a wince, and that is not what I sound like. The people in the room, who have heard both, are unimpressed — the recording sounds exactly like them. Two claims that cannot both be about the same thing, and yet each party is reporting honestly.

So which ear is wrong? Neither, it turns out, and noticing why forces a question you may never have asked: how does the sound of your own voice actually get to you?

b

Reasoning it through

REASONING #

Assume for a moment it travels the way every other sound does — out of your mouth, through the air, round to your ear canal, down to the eardrum. If that were the whole story, you and the microphone would be sampling the same signal from slightly different positions, and playback should be unremarkable.

There is a cheap experiment against that assumption. Block both ears with your fingers and hum. The hum does not vanish; it gets louder and considerably deeper. Something is reaching the inner ear without going through the outside air at all. That is the second route: your vocal folds vibrate the tissue and bone of the head directly, and the cochlea, which is embedded in the skull, is driven by that vibration as well as by the airborne sound. Plugging the ear canal traps the airborne portion and makes the bone-conducted portion relatively more prominent — the occlusion effect.

Now the crucial detail. The two routes do not carry the same thing. Bone conduction through the skull favours the lower frequencies; the mass and elasticity of tissue pass slow vibrations far more readily than fast ones. So the voice you have always heard is the airborne one plus a low-frequency reinforcement nobody else receives. It is fuller and deeper than what leaves your mouth.

Which makes the recording predictable. A microphone stands in the air and captures the airborne route alone. Play it back and the low-frequency supplement is simply absent — so your voice sounds thinner, reedier, higher, more nasal than the one in your head. And note that this is not a distortion introduced by the recording. The recording is the honest one; your own version is the augmented one.

Is that the whole answer? Probably not, and the second part is less often cited. Ask what else changes when you hear yourself from outside. While speaking, you are not really listening — you are monitoring, using your own voice as feedback to steer articulation, attending to what you are saying rather than to how it sounds. In playback you are a listener for the first time, and things you never monitored become audible: hesitations, the shape of your accent, tics of rhythm and emphasis. Some of the shock is not spectral at all. It is meeting a speaking manner you have owned for decades and never observed.

Layer familiarity on top. Mere exposure — the tendency for repeated stimuli to feel right — has been shown across a broad range of material, and no version of your voice has been repeated to you more than the bone-conducted one. It is not merely different from the recording; it is the reference against which the recording registers as wrong. That framing is an inference from a general effect rather than something demonstrated for voices specifically, and it is worth holding loosely; there is work reporting that people rate their own recorded voice more favourably when they are not told whose it is, which suggests recognition itself is doing part of the damage.

c

The analogy

THE ANALOGY #
THE FIGURE

Suppose you had only ever heard your favourite album through headphones with a bass boost you did not know was switched on. Play it through a friend's speakers and it sounds thin, brittle, wrong — though that is the mix the engineer made and everyone else has always heard. Your objection is not to the recording. It is to the missing boost.

WHERE IT BREAKS DOWN

the headphone boost is a filter applied to an external signal, whereas bone conduction is a second delivery route from a source inside your own head, which also carries feedback you use to control your speech in real time — something no listener, and no album, ever supplies.

d

Clarifying the model

THE MODEL #

A few refinements hold the picture together.

First, the recording is closer to what others hear, not identical to it. A microphone has its own frequency response, sits at its own distance, and the room adds reflections that were not in anyone's ear; a phone's small speaker then rolls off the bass on playback. Cheap equipment therefore exaggerates the effect, but it does not create it — the mismatch persists on excellent gear.

Second, resist the tempting summary that "your voice is really higher than you think". The pitch you produce is what it is, and it is unchanged by any of this. What differs is the balance of frequencies — the relative weight of the low end — which is why the recorded voice reads as thinner rather than as transposed upward.

Third, the two explanations are not rivals; they sit at different levels. The bone-conduction account says why the signal differs. The familiarity account says why a difference of that size is experienced as alienating rather than merely noticed. A purely acoustic story would predict mild surprise, not the reliable flinch.

e

A picture of it

THE PICTURE #
Voice recording mismatch
Voice recording mismatch Read downward as time. The single source at the left sends its vibration two ways at once -- out through the air, and directly through the bone and tissue of the head -- and both arrive at your cochlea, so the voice you know is their sum, with the bone route adding weight at the bottom of the range. The last three messages are the recording: the microphone can only stand in the air, so it captures one of the two routes, and playback delivers that one back to you. Nothing has been altered; a contribution has been dropped. {"generator":"mermaid-svg-renderer@3.2.1","source":"../Socrates/.diagram-cache/_src/recorded-voice.md","sourceIndex":1,"sourceLine":4,"sourceHash":"9f8941e7dbacddb778d6fc3e84255766c9b44c83c2b834180b86c4ba214b092c","diagramType":"sequence","layoutVariant":"source","repairedDuplicateIds":[],"motion":"entrance-with-reduced-motion-fallback","presentation":"editorial","attempt":1,"viewBox":{"x":0,"y":0,"width":1388,"height":776},"qa":{"passed":true,"findings":[]}} Microphone and playback 01 Your cochlea 02 Skull and soft tissue 03 Air outside the head 04 Vocal folds 05 one source, two routes live self-hearing equals air plus bone same voice, without the low-end reinforcement sound radiates from the mouth 1 vibration conducts through the head 2 full spectrum, from outside 3 low frequencies favoured 4 only the airborne route is captured 5 playback arrives by air alone 6
KINDSlifelineparticipantmessage

How to readRead downward as time. The single source at the left sends its vibration two ways at once — out through the air, and directly through the bone and tissue of the head — and both arrive at your cochlea, so the voice you know is their sum, with the bone route adding weight at the bottom of the range. The last three messages are the recording: the microphone can only stand in the air, so it captures one of the two routes, and playback delivers that one back to you. Nothing has been altered; a contribution has been dropped.

f

What became clearer

WHAT CLEARED #
WHAT CLEARED

You have never heard your voice the way a microphone does, because you have always heard it twice — once through the air and once through your own skull, with the bone route adding a low-frequency reinforcement that no one else receives and no recording can capture. The playback is not a distortion of your voice; it is your voice minus something. And the discomfort exceeds the acoustic difference, because that augmented version is also the one a lifetime of exposure has made feel correct.

g

Where to go next

ONWARD #
  • Why bone-conduction headphones work at all, and what they trade away.
  • How singers and broadcasters learn to trust an external monitor over their internal one.
h

Key terms

TERMS #
TermWhat it means
Bone conductiontransmission of sound to the inner ear through vibration of the skull and soft tissue rather than through the ear canal.
Occlusion effectthe boost in perceived loudness and depth of one's own voice when the ear canals are blocked.
Frequency responsehow a system's output level varies across the range of pitches it passes.
Mere exposure effectthe tendency for repeated exposure to a stimulus to increase liking or the sense that it is right.

Every term the collection defines is gathered in the glossary.

Nearby on the shelf

4