Measurement Epistemology 5 min read

The Prism

The Prism

How many Americans have Long COVID?

There are five major instruments trying to answer this question in the United States right now. They return five different numbers. The gap between the lowest and highest is nearly threefold. None of them are wrong.

Five Instruments, Five Diseases

0% 5% 10% 15% 20% ~6% Household Pulse 6.4% BRFSS 8.3% NHIS 14.3% NYC CHS 16.3% Tian EHR/AI nearly 3× range

These numbers span from roughly 6% to 16% — and they are not measuring the same thing.

Instrument Question Asked Method Result
Tian et al. (JAMA Net Open, 2026) Does this patient's EHR match a PASC phenotype? AI algorithm on 58-hospital records (n=457,950) 16.3%
NYC CHS (medRxiv, 2025) Have you ever had Long COVID? City-level phone survey (2021–2023) 14.3%
NHIS (Jia et al., 2026) Have you ever had Long COVID? National household interview (n≈88,000/yr) 8.3%
BRFSS (MMWR, 2024) Do you currently have Long COVID? State-level phone survey (n≈400,000) 6.4%
Household Pulse (Census) Do you currently have Long COVID? Online panel survey (biweekly) ~6%

Some of this divergence is definitional: "ever" and "currently" are different constructs. But even within the same construct, the instruments disagree. NHIS says 8.3% "ever." NYC CHS says 14.3% "ever." Same question, different sample frames, nearly double the rate. For "current" Long COVID: NHIS reports 3.3%, BRFSS reports 6.4% — again, a twofold gap.

And Tian et al. found that diagnostic codes — the data the healthcare system actually uses — capture fewer than 7% of algorithmically detected PASC cases. More than 10 million Americans with Long COVID may be invisible to the instruments that count them.

The Shape Changes Too

The instruments don't just disagree on how many. They disagree on what the disease is.

O'Mahoney et al. (Nature Communications, 2025) conducted the largest controlled meta-analysis of Long COVID symptoms to date: 50 studies, 14.66 million participants, all with uninfected comparator groups. When you compare infected patients to uninfected controls rather than to a pre-COVID baseline, the symptom hierarchy reshuffles:

Symptom risk relative to uninfected controls

Loss of smell — RR 4.31

Loss of taste — RR 3.71

Poor concentration — RR 2.68

Impaired memory — RR 2.50

Fatigue — drops in the hierarchy

Without controls, fatigue dominates every Long COVID study. With controls, it recedes — because uninfected people also report fatigue at high rates. The disease looks different depending on whether you subtract the background.

This is not a technical footnote. It is a shape-change. Uncontrolled studies describe a fatigue-dominant disease. Controlled studies describe a chemosensory-cognitive disease. Physicians, patients, and policymakers reading these two versions would make different decisions about treatment, research funding, and disability accommodation.

And Panagiotopoulos and Ioannidis (medRxiv, 2026) found that 84.5% of published Long COVID prevalence studies lack any control group at all. Nearly two-thirds rely exclusively on self-report. The version of the disease that most studies describe — the uncontrolled version — may be the less accurate one.

The Pattern That Survives

If every instrument shows a different disease, is there anything left? Yes — but you have to look at the structure, not the numbers.

Two independent instruments measuring Long COVID prevalence over time show the same paradox:

NYC Community Health Survey

Population "ever LC": 5.9% → 14.3% ↑ rising

Per-infection risk: 34.6% → 26.3% ↓ falling

National Health Interview Survey

Population "ever LC": 7.0% → 8.3% ↑ rising

Per-infection risk: 17.7% → 13.7% ↓ falling

Two different instruments, two different populations (New York City vs. national), different survey methods — and they converge on the same directional structure: the total number of people who have ever had Long COVID keeps rising even as each new infection is less likely to cause it.

The specific numbers disagree wildly. NYC CHS says per-infection risk started at 34.6%; NHIS says 17.7%. But both show the same arrow: down. Both show population prevalence going the other way: up. The paradox resolves simply — reinfections accumulate. Each individual infection is less dangerous than earlier-variant infections, but there are more total infections per person, so the population sum keeps climbing.

This is what a prism does. White light enters. The prism doesn't break the light — it separates it into components that were always there. Each instrument refracts Long COVID differently. They disagree on the specific numbers. They disagree on the symptom hierarchy. But some structural properties of the disease — the rising-while-falling paradox, the cumulative ratchet — survive every refraction.

The Missing Update

The CDC promised updated Long COVID prevalence estimates by June 30, 2026. The deadline passed. The tracking page still shows data from March 9, 2026. No explanation was given.

This matters because the CDC's estimates are what policymakers use. The five instruments above exist in academic papers. What exists in congressional briefings, disability determinations, and insurance coverage debates is whatever number the CDC last published. When the update doesn't come, the policy world operates on stale data while the disease accumulates.

What I Don't Know

The NYC Community Health Survey paper is a preprint. Its per-infection estimates (34.6% falling to 26.3%) are substantially higher than the NHIS figures, which may reflect New York City's early, severe epidemic waves biasing upward or may reflect genuine geographic variation. The directional convergence with NHIS is informative; the absolute magnitudes are not directly comparable.

The Tian et al. AI algorithm was validated against clinician review (PPV 79.5%) but operates on EHR data, which overrepresents people who seek care and underrepresents those who don't. Its 16.3% may overcapture in healthcare-utilizing populations while undercapturing in populations that avoid or lack access to the healthcare system.

I've framed this as "the instruments refract differently" rather than "the instruments are broken." But the less comfortable interpretation is also possible: some of these instruments may simply be wrong, and the pattern-convergence I identify may be coincidental rather than structural. Two instruments showing the same paradox is suggestive. It is not proof.

The most important thing I don't know is what happens next. The NHIS data shows "ever LC" plateauing between 2023 (8.4%) and 2024 (8.3%). If each infection carries lower risk, and infection rates stabilize, the population burden may be leveling off. Or the plateau may be an artifact of the specific instrument — a ceiling on what self-report surveys can capture. We will need the CDC update to tell them apart. We are still waiting.