What happens in these experiments
For three decades, researchers have studied false confessions by accusing innocent people of minor transgressions in the laboratory. Four paradigms dominate the literature. In each, the participant truly did not do the thing — the question is whether they will sign a statement saying they did.
Alt-Key
Participants type in a speeded computer task and are warned not to touch a forbidden key. The computer crashes by design, and the participant is accused of pressing it — in many versions, a confederate claims to have seen it happen. Signing an admission typically costs nothing beyond agreeing.
Cheating
Participants solve problems alongside a partner (secretly a confederate) and are later accused of sharing answers on a problem they were told to solve alone. The act feels more like real wrongdoing and ground truth is verifiable — but there is still no arrest, charge, or court.
Fake video
Participants play a computerized gambling task and are then shown doctored video “evidence” of themselves taking money on a losing round. Every condition in this paradigm confronts the participant with fabricated evidence.
Theft accusation
Participants are accused of taking money or an item — a missing $20 bill, for example — during the session. The allegation is closer to a real one, but it still unfolds inside a university lab with no real consequences attached.
The headline number — and its enormous spread
Pooling all 99 conditions, the meta-analytic false confession rate is 41%. But the more revealing number is the spread. A new experiment drawn from this literature could credibly produce a rate anywhere from 3% to 93% — near zero to near universal — depending on task details.
Each dot is one experimental condition (dot size = number of participants; hover for details). The horizontal line marks the pooled 41% estimate; the shaded band is the 95% prediction interval.
If these tasks measured a stable psychological construct, different implementations should produce reasonably consistent results. Instead, rates scatter across nearly the entire 0–100% range. That pattern looks less like a stable trait being measured and more like an outcome that is highly sensitive to how the task is built.
37% confess before interrogation even starts
In 30 control conditions, the experimenter simply stated the accusation — no false evidence, no minimization, no pressure tactics of any kind. 37% of innocent participants signed anyway. The most coercive tactic in the literature, the false evidence ploy, averages 62% — which means the no-tactic baseline already accounts for 60% of it.
Unweighted mean false confession rate by interrogation method, ordered from least to most coercive (computed directly from the 99 conditions; k = number of conditions). Dashed line: control baseline.
Tactics do matter at the margin — rates rise about 0.23 logit units per step up the coerciveness scale (Spearman ρ = .33). But in a formal dose-response model, coerciveness explains less than 1% of the variation across conditions. Most of the measured behavior is already there when an authority figure merely states an accusation in a low-stakes setting. That is the signature of compliance, not of interrogation-induced false confession.
Small details move the needle more than tactics do
If tactics explain almost none of the variation, what does? Task structure and context. Explore the three views below: the same “false confession rate” swings dramatically with the paradigm used, minor procedural details like which key the participant is accused of pressing, and the country where the study was run.
Model-based (posterior) estimates with 90% credible intervals, for observed paradigm–method combinations. Open circles (†) are cells with a single condition — interpret those with caution. Note the false evidence ploy and mixed accusatorial tactics tie at ≈69% in the Alt-Key paradigm.
Posterior estimates with 90% credible intervals across Alt-Key task variants. Change nothing but the forbidden key — accuse participants of pressing Shift instead of Esc — and the rate roughly triples (23% → 60%). A construct this sensitive to trivial procedural details is not behaving like a stable measure of interrogation vulnerability.
Posterior estimates with 90% credible intervals, holding paradigm (Alt-Key), variant (classic), and method (control) constant. Same task, same absence of tactics: 45% in the United States, 6% in Japan. Non-U.S. estimates rest on fewer studies and are less precise.
Two ways to read this literature
The validity question can be framed as a contest between two interpretations, each making observable predictions. The interrogation mechanism account says lab paradigms tap the processes that produce real false confessions. The compliance account says they mostly measure willingness to go along with an authority figure’s accusation when nothing is at stake.
| Observable prediction | Interrogation mechanism account | Compliance account | What the data show | Verdict |
|---|---|---|---|---|
| Confessions when no tactic is applied | Rare | Common | 37% across 30 control conditions | ✗ Mechanism ✓ Compliance |
| Consistency across implementations | Reasonably stable | Swings with task details | 3–93% prediction interval; rate triples across key variants | ✗ Mechanism ✓ Compliance |
| What moves the rate | Coercive tactics | Procedural & contextual details | Tactics: <1% of variation paradigm, variant & country do far more work | ✗ Mechanism ✓ Compliance |
| Magnitudes commensurable with the real world | At least roughly | Far above any defensible rate | See the calibration test below compatibility requires true rates near 41% | ✗ Mechanism ✓ Compliance |
These accounts are not mutually exclusive — laboratory tasks may capture mechanisms genuinely relevant to interrogation while still being poorly calibrated to the real-world phenomenon. The question is which account better describes the absolute rates and patterns of variation the literature actually produces.
Could these rates come from the real world?
Here is the paper’s central exercise. Suppose the true real-world false confession rate were some value — say 10%. If we then ran this exact literature (99 conditions, the same sample sizes) on that world, how often would it produce the observed pooled mean of 41%? Drag the slider.
Probability that a synthetic literature (99 conditions, the actual sample sizes, binomial sampling) yields a mean at or above the observed 0.41, by hypothesized true rate — the normal approximation to the paper’s 10,000-draw simulation. Amber band: plausible real-world range (5–15%).
At true rates of 5%, 10%, or 15%, the probability is essentially zero — not small, zero to hundreds of decimal places. The most forgiving version of this test (allowing the corpus’s own heterogeneity into the hypothetical real world) lowers the edge of compatibility from about 40% to about 32% overall, and to 23% for control conditions — still double the field’s upper-bound survey estimates. The laboratory numbers are only compatible with a world in which false confession is routine.
If lab rates were real-world rates
Another way to see the calibration gap: take the laboratory rate literally and see what it implies. The FBI reports roughly 7.5 million arrests per year in the United States. Choose what share you think involves custodial interrogation of an innocent person.
To be clear about the logic: this is a reductio, not an accusation. We do not claim researchers believe the field produces hundreds of thousands of false confessions a year. The point is the converse — because that implication is not reasonable, a laboratory rate that produces it cannot be measuring the same thing as real-world false confession. Even the control conditions alone (37%, no tactics) imply roughly 278,000 per year.
Even generous assumptions leave an enormous gap
Effect sizes are the standard language for how big a difference is. Express the gap between the pooled laboratory rate (41%) and conservative real-world rates (5–15%) in those terms, and it dwarfs the conventional benchmarks for a “large” effect — at every baseline.
Odds ratio comparing the pooled laboratory rate (41%) with each assumed real-world rate. Dashed line: the conventional “large effect” benchmark (OR = 2). Hover a bar for the other metrics.
| Assumed real-world rate | Risk ratio | Odds ratio | Cohen’s h | Cohen’s d | Correlation r |
|---|---|---|---|---|---|
| 5% | 8.18 | 13.1 | 0.94 | 1.42 | .58 |
| 10% | 4.09 | 6.23 | 0.75 | 1.01 | .45 |
| 15% | 2.73 | 3.92 | 0.59 | 0.75 | .35 |
Large-effect benchmarks: RR = 1.5, OR = 2.0, h = 0.5, d = 0.8, r = .30. Every value exceeds its benchmark except Cohen’s d at the 15% baseline (0.75, just below the 0.8 cutoff).
At a 5% baseline, the gap corresponds to an odds ratio of 13.1 and a Cohen’s d of 1.42. Even granting 15% — above the field’s own upper-bound survey estimates — nearly every metric still clears its large-effect threshold, and the one exception (d = 0.75) misses by a hair. Gaps of this size are rare for complex social behavior: associations this strong usually involve direct mechanistic relationships, and behavioral effects routinely shrink once confounding and measurement error are addressed. These numbers are best read as the size of the calibration gap itself — not as the size of any real-world interrogation effect.
High baselines are a feature — and the problem
Why would paradigms evolve toward tasks where a third of innocent people comply immediately? Statistical power. Detecting tactic effects at realistic base rates requires enormous samples; at laboratory base rates, small samples suffice.
Participants needed per condition for 80% power (two-sided α = .05) to detect the labeled increase, versus the average condition actually run in this literature (n = 37).
A four-arm study at realistic base rates would need nearly 5,900 participants; the average condition in this corpus enrolled 37. A paradigm that produces a 30–40% baseline makes tactic comparisons feasible at ordinary lab sizes — that is precisely why these tasks are useful experimental tools. But the same feature is the validity problem: a behavior common enough to appear in over a third of innocent control participants is unlikely to be the same quantity as the rare, high-cost event of false confession in a criminal case.
What this means in practice
For courts & attorneys
Laboratory rates, odds ratios, and effect sizes from this literature are not field-calibrated risk estimates. When lab numbers are offered to characterize real-world interrogation risk, Daubert/Frye “fit” questions apply: the behavior measured (low-stakes compliance with an accusation) is not the phenomenon at issue (custodial false confession).
For police & trainers
The experiments do show that false evidence ploys and minimization elicit more admissions than a bare accusation — in the lab. Treat that as evidence about compliance pressure worth considering in context. Do not treat 41%, 69%, or “13× the odds” as the real-world risk attached to any tactic.
For researchers & reviewers
Report absolute baselines alongside contrasts, and flag calibration limits when citing pooled rates. A 3–93% prediction interval means “the” laboratory false confession rate does not exist; a 37% no-tactic baseline means the outcome needs a construct label chosen with care.
What the literature does offer
None of this means laboratory research on false confessions lacks value. The paradigms isolate real psychological mechanisms — minimization cues, false evidence, authority signals — and can rank tactics by how much compliance they elicit under controlled conditions. Read as model systems for theory-building, they are informative. What they cannot responsibly supply is the number: the prevalence, magnitude, or field risk of interrogation-induced false confession.
We tried to break the result — it held
- Publication bias? Adjusting for small-study effects raises the estimate rather than lowering it (PET 64%, PEESE 49%; Egger z = −4.61). Selective reporting does not explain the elevated rates.
- One influential study? Dropping each of the 41 studies in turn moves the pooled rate only between 37.2% and 41.3% — no single paper drives the result.
- Threshold choice? Five alternative calibration thresholds (posterior mean, median, 25th/75th percentiles, frequentist estimate) all yield the same verdict: essentially zero probability at real-world rates ≤ 15%.
- Heterogeneity? Injecting the corpus’s own between-condition heterogeneity (τ ≈ 1.5) into the hypothetical real world lowers the compatibility edge from ≈40% to ≈32% (23% for controls) — still far above any defensible real-world rate.