THIS EXPLANATION
THE ROOM
GOV·04 Government, Law & Civics 6 MIN · 8 STATIONS

Census undercount

A Socratic walk-through of census undercount — reasoned out one step at a time, not lectured.

abcdefgh
a

The question we started with

THE QUESTION #

What changes when a government miscounts its own population?

Every census misses people. Households do not answer, addresses do not exist on any list, someone is between homes on the reference night. Nobody expects a perfect count, and everyone involved knows it is imperfect.

So here is the question worth asking: if the count is known to be wrong, why does it matter so much how it is wrong? Suppose a census missed exactly two percent of every group, everywhere, uniformly. What would actually break?

b

Reasoning it through

REASONING #

Follow the number to where it is used. In a system like the United States', the census determines how many legislative seats each state receives, where district lines are drawn, and how large sums of formula funding are distributed. Notice what those uses have in common: they are almost all about shares. Seats are apportioned among states; money is allocated by each area's proportion of a national total.

Now apply the uniform error. If every state is two percent low, every state's share of the total is unchanged, and the apportionment comes out identical. A uniform error, for these purposes, is remarkably harmless. That is the first surprising conclusion, and it inverts the intuitive worry: the accuracy of the national total is not what is at stake.

What is at stake is differential undercount — error that is concentrated. Miss five percent of one group and none of another, and every share moves. Representation shifts from the places where the missed people live to the places where they do not. And here the pattern is not random. Coverage measurement in successive censuses has consistently found the same shape: people who rent rather than own, young children, people in irregular or shared housing, some racial and ethnic minority groups, and remote rural populations are undercounted, while other groups are simultaneously *over*counted, through duplication — a student recorded both at university and at home, a family with two residences. The 2020 post-enumeration survey in the United States found exactly this structure: a national net coverage error not statistically distinguishable from zero, alongside statistically significant undercounts of some groups and overcounts of others. A count can be excellent in aggregate and badly wrong everywhere that matters.

There is a second harm, quieter than the first. The census is not consulted once. It becomes the denominator for a decade — the base for annual population estimates, the frame from which other surveys draw their samples, the divisor in every per-capita rate. If a population is undercounted, its crime rate, vaccination rate, and unemployment rate are all computed against a base that is too small. The original error propagates into statistics whose users have no idea a census is involved.

And the error is not neutral in direction, which is the uncomfortable part. Those hardest to enumerate are broadly those with the greatest claim on the programmes the count funds, so the mechanism systematically transfers representation and resources away from them.

c

The analogy

THE ANALOGY #
THE FIGURE

Think of a shared restaurant bill split between tables by how many people each table reports. If every table undercounts by one, the split is unchanged and no one is harmed. If only one table undercounts, everyone else pays a little more and that table's error is invisible in the arithmetic.

WHERE IT BREAKS DOWN

The bill is settled once and then forgotten, whereas a census figure stays in service for ten years as the denominator of other statistics — so unlike the diners, the miscounted party keeps paying in places nobody is looking, long after the original count is out of anyone's mind.

d

Clarifying the model

THE MODEL #

The natural response is: measure the error and correct for it. That is technically possible. An independent post-enumeration survey re-canvasses a sample of addresses; comparing who appears in both sources and who in only one yields an estimate of who appears in neither. This is dual-system estimation, the same capture-recapture logic used to estimate wildlife populations, and the statistical theory is mature.

Why is it not simply done? Two reasons, and honesty requires giving both weight.

The technical objection is real. The correction has sampling error of its own, and it suffers correlation bias: people the census missed are disproportionately the people the follow-up survey also misses, so the adjustment systematically understates the gap. Adjustment can improve national and state figures while making some small-area figures worse — and small areas are what funding formulas often use.

The political objection is that adjustment predictably moves seats and money in a direction that is known in advance, which makes it a contested act rather than a technical one. In the United States the courts held that statistical sampling may not be used to produce the apportionment count, and the Bureau has declined to adjust the published figures even when its own coverage measurement found the differential clearly. The result is a standing gap between the best available estimate of the population and the number the law uses — maintained deliberately, on grounds that are partly statistical and partly not.

e

A picture of it

THE PICTURE #
Census undercount
Census undercount Start at the top parallelogram, a single household, and follow the branches by whether anyone was reached; each labelled edge is a different quality of data feeding the same store. The two shaded risk nodes on the right are the whole problem -- what leaves the flow there does not leave at random, which is why the harm is labelled differential. At the bottom, the second parallelogram brings in the independent survey that could measure the gap, and the final decision shows the split outcome: the measured correction is permitted to inform estimates but not the count that assigns seats and money. {"generator":"mermaid-svg-renderer@3.2.1","source":"../Socrates/.diagram-cache/_src/census-undercount.md","sourceIndex":1,"sourceLine":4,"sourceHash":"18d6a2bb2032f4e36677b80b76c5a48e50665afea2defc0fe077d1901c95282f","diagramType":"flowchart-v2","layoutVariant":"source","repairedDuplicateIds":[],"motion":"entrance-with-reduced-motion-fallback","presentation":"editorial","attempt":1,"viewBox":{"x":0,"y":0,"width":1201,"height":1231},"qa":{"passed":true,"findings":[]}} yes, self-response no reply an occupant or neighbouranswers nobody reached unit never located barred for apportionment permitted for estimates A household is asked to respond Does the household reply? The enumerated count Does a field visit reach anyone? Proxy response, less accurate Impute from records and nearbyunits Person missed entirely Differential undercount,concentrated by group andplace Post-enumeration surveyre-canvasses a sample May the count be adjusted? Seats, districts, formula funding Population estimates and surveyframes
KINDSsourcedecisionprocessriskoutcomeconnector

How to readStart at the top parallelogram, a single household, and follow the branches by whether anyone was reached; each labelled edge is a different quality of data feeding the same store. The two shaded risk nodes on the right are the whole problem — what leaves the flow there does not leave at random, which is why the harm is labelled differential. At the bottom, the second parallelogram brings in the independent survey that could measure the gap, and the final decision shows the split outcome: the measured correction is permitted to inform estimates but not the count that assigns seats and money.

f

What became clearer

WHAT CLEARED #
WHAT CLEARED

The damage from a miscount is not proportional to its size but to its concentration. Because the count is used to divide fixed things into shares, an error spread evenly cancels out, while an error concentrated in particular groups and places moves representation and money away from exactly those groups — and then persists for a decade inside every rate computed on top of it. That the correction is technically available and legally unusable is the part that no amount of better fieldwork resolves.

g

Where to go next

ONWARD #
  • How the choice of a residence rule — where someone is deemed to live on one night — creates its own systematic errors.
  • Why administrative registers, used instead of a headcount in several countries, trade one set of biases for another.
h

Key terms

TERMS #
TermWhat it means
Differential undercountcoverage error that varies systematically by group or place, rather than being spread evenly across the population.
Apportionmentthe allocation of legislative seats among regions in proportion to their counted populations.
Post-enumeration surveyan independent re-canvass of sampled addresses, used to estimate how many people the census missed.
Dual-system estimationthe capture-recapture method that infers those missed by both the census and the survey from the overlap between them.
Imputationfilling in a missing record using administrative data or nearby units when no response can be obtained.

Every term the collection defines is gathered in the glossary.

Nearby on the shelf

4