Massed string sound
A Socratic walk-through of massed string sound — reasoned out one step at a time, not lectured.
The question we started with
THE QUESTION #Why does doubling an orchestra's violin section add so little loudness compared with doubling one player's effort?
An orchestra fields sixteen first violins and one flute. If sixteen players were sixteen times one player, the flute would be inaudible and the section deafening. Neither is true. And a conductor wanting a louder string sound does not ask for more players — they ask the players they have to dig in.
Sixteen instruments really are producing sixteen instruments' worth of sound energy. So where does it go, and why does an orchestra keep so many players anyway?
Reasoning it through
REASONING #Start by being careful about what adds. Sound arrives as a pressure fluctuation, and pressure is signed: at any instant a wave is pushing or pulling. Two contributions arriving together may reinforce or partly cancel, depending where each is in its cycle. So the question is not "how much sound is there" but "how do these sixteen waves line up at the ear".
Take the extremes. Suppose all sixteen were perfectly in step — identical waveform, identical phase. The pressures add straight: sixteen times one player's amplitude, and since intensity goes as the square of pressure, 256 times one player's power. Now suppose instead they are unrelated in phase. The cross terms — the products of one player's wave with another's — are as often negative as positive and average away over time. What survives is each player's own squared contribution, so total power is sixteen times one player's, and amplitude is the square root of sixteen: four times, not sixteen.
Which case is a violin section in? Their pitches differ by a few cents, their bow changes fall at slightly different instants, and each has their own vibrato at its own rate and phase. Nothing locks them at the sample level — they are locked at the level of the beat, thousands of times coarser than a cycle of the note being played. So the section is very close to the incoherent case: its power adds, its coherence does not.
Now put decibels on it. Level in decibels is ten times the base-ten logarithm of the power ratio. Incoherent addition gives power proportional to the number of players, so level above one player is ten log ten of N, and doubling N gives 3.01 dB. In the coherent case power goes as N squared, so doubling would give 6.02 dB. A doubled section buys three decibels, not six, and nothing like the sixteen-fold impression the count suggests.
Three decibels is a real increase and a small one. To judge how small we need a rule about hearing rather than physics, and the usual figure is that a ten-decibel increase is experienced as a doubling of loudness. Ten decibels is a factor of ten in power, so ten times the players: sixteen violins would need to become a hundred and sixty. No hall and no payroll accommodates that, which is why the loudness lever an orchestra pulls is the one on each individual player.
So if sections are a bad way to buy volume, what are they for? Two things, the first being the same incoherence that cost us the loudness. Sixteen slightly detuned, independently bowed instruments produce a sound whose fine structure constantly fluctuates — a wash no single instrument can imitate, unlocatable in space, without the grain of one violin. That is the sound composers write for, and it follows from the players not matching. The second is headroom: sixteen players at a sustainable dynamic hold a line one player at full effort could not.
The analogy
THE ANALOGY #Imagine a hundred people on the same spot, each taking one long stride in a direction chosen at random. They do not end up a hundred paces from the start; their displacements partly cancel, and the crowd ends up on average about ten paces out — the square root of a hundred. Had they agreed on a direction, they would be a hundred paces out. Ten versus a hundred is the difference between adding things lined up and things not.
Each walker takes one step and stops, so the result is a single draw from a distribution that could come out unusually large or small, whereas players sound continuously and their phase relationships are re-scrambled thousands of times a second — which is why a section's level is a dependable average, not a lucky outcome.
Clarifying the model
THE MODEL #The misconception to dislodge is that the missing sound has been destroyed. It has not. Total acoustic power is genuinely sixteen times one player's; nothing is cancelled on average. What fails to happen is the amplitude piling up, and since our loudness impression tracks something far closer to power than to a count of instruments, sixteen times the power is only twelve decibels above one.
Second, this is why studio doubling behaves as it does. Mix a track with a bit-identical copy and you get 6 dB, the two being perfectly coherent; mix it with a second take of the same part and you get about 3 dB, plus the characteristic thickening. The waveform is all that changed.
Third, a boundary against a neighbour. Concert hall reverberation describes a room handing back delayed copies of the same sound, and the mechanisms are cousins: the room's returns arrive at random phase relative to the direct sound, so they too add in power rather than amplitude. The difference is what does the adding — there, one source returning many times; here, many sources arriving once.
Fourth, honest limits. The ten-decibels-per-doubling-of-loudness figure is a recalled psychoacoustic convention, not a derived quantity: it depends on level, bandwidth and how the question is asked, and I would not defend the exact number. Perfect incoherence is an idealisation too — players partially synchronise their attacks. And the account is about level, not the impression of size, which the section's spatial spread affects in ways no decibel captures.
The falsification test is direct. In a dead room, at a fixed dynamic, record a section as players are added one at a time and plot the level. Incoherent addition predicts ten log ten of N: 3 dB from one to two, about 12 dB by sixteen. The refuting observation is twenty log ten of N — 6 dB per doubling, 24 dB by sixteen — which would mean the sources are phase-locked and my account is wrong.
A picture of it
THE PICTURE #How to readThe horizontal axis doubles at every step, so equal spacing means equal doublings. The bars are the real case, incoherent addition, computed as ten times the base-ten logarithm of the player count — each doubling adds a flat 3 dB, and thirty-two players sit only 15 dB above one. The line is the hypothetical coherent case, twenty times the logarithm, rising twice as fast to 30 dB. The widening gap is the loudness phase disagreement costs — and, read the other way, the ensemble texture it buys.
What became clearer
WHAT CLEARED #Sixteen violinists are not one violinist multiplied by sixteen, because sound adds as a signed pressure and their waves are not lined up. Their powers add, so level rises as the logarithm of the count: 3 dB per doubling, about 12 dB for a full section, a factor of ten in players for a subjective doubling of loudness. Orchestras keep the players anyway, because the same failure to line up produces the shimmering, unlocatable sound of a string section. The section is not a volume control but a texture, whose price is the loudness it fails to deliver.
Where to go next
ONWARD #- Why a solo violin stays audible over an orchestra while being quieter than the section it competes with.
- Whether the same square-root logic explains massed voices, where pitch spread and vibrato behave differently.
Key terms
TERMS #| Term | What it means |
|---|---|
| Coherent addition | summation of waves holding a fixed phase relationship, so amplitudes add and power rises as the square of the count. |
| Incoherent addition | summation of waves with unrelated phases; cross terms average away and power rises in proportion to the count. |
| Decibel | ten times the base-ten logarithm of a power ratio; doubling power is 3.01 dB. |
Every term the collection defines is gathered in the glossary.