What Do Laboratory False Confession Paradigms Actually Measure?
Our new paper, “What Do Laboratory False Confession Paradigms Measure? A Calibration Meta-Analysis”, is now out in the Journal of Quantitative Criminology.
The paper started with a fairly simple question.
When an innocent college student signs a statement saying that they pressed a computer key they did not press, cheated on a laboratory task, or committed some other minor transgression invented by an experimenter, what exactly has been measured?
The experimental literature calls that outcome a false confession. But the label does a lot of work.
A false confession in a criminal case is a high-cost act of self-incrimination. It occurs in a legal setting, usually under custodial pressure, with the possibility of arrest, prosecution, conviction, imprisonment, public stigma, and very real consequences for the person making the statement.
Those features are largely impossible to reproduce ethically in a laboratory.
So laboratory studies necessarily strip most of them away. Participants generally face a minor accusation involving an artificial task, know they are participating in research, face little meaningful consequence for signing, and can leave the experiment without being prosecuted or sent to prison.
That does not make the experiments bad experiments.
It does raise a basic measurement question: does signing the laboratory statement measure the same psychological construct as falsely confessing to a crime?
Our results give us considerable reason to doubt that it does.
Start with the control conditions
We analyzed 99 experimental conditions from 41 studies, covering 3,682 participants.
The pooled laboratory false confession rate was about 41%.
But the number I find more important is 37%.
That is the pooled rate in control conditions, where no substantive interrogation tactic was applied. In most of these conditions, the experimenter essentially accused the participant and asked them to sign.
More than one-third did.
That should matter for how these studies are interpreted.
If 37% of innocent participants produce the outcome before false evidence, minimization, maximization, or another focal interrogation tactic is introduced, then a very large part of what the experiment calls a false confession cannot be attributed to interrogation pressure.
The most straightforward explanation is that these tasks are very good at eliciting compliance with an authority figure in a low-stakes setting.
That behavior is psychologically interesting. It may even tell us something useful about mechanisms that also operate during interrogations.
But compliance with an experimenter and falsely incriminating yourself for a serious crime are not interchangeable outcomes.
The experiments behave more like tasks than like a stable measure of false confession
Another result reinforced that concern.
Across the literature, the expected false confession rate for a new experimental condition ranges from roughly 3% to 93%.
Think about that range for a moment.
A researcher can run one of these paradigms and, depending on relatively small differences in implementation, plausibly obtain almost no false confessions or nearly universal false confession.
Even within the classic Alt-Key paradigm, changing which keyboard key participants were accused of pressing produced large changes in the resulting rate. Rates also varied substantially across countries.
That is not what we would expect from a well-calibrated measure of a reasonably stable underlying construct.
It looks much more like a task-sensitive behavioral outcome.
In other words, the experimental setup itself appears to matter enormously.
That is important because the literature has generally been interested in the effect of interrogation tactics. And those tactics do matter. More coercive approaches tend to produce more admissions.
But they explain very little of why confession rates vary so dramatically from one experimental condition to another.
The control baseline alone is about 60% of the rate produced under false evidence ploys.
So the picture is not one in which an otherwise reluctant innocent person is exposed to an interrogation tactic and suddenly becomes likely to confess. The experiment starts with an unusually high probability of signing, and the tactic moves that already-high probability upward.
That is a different phenomenon.
Then there is the calibration problem
We also asked a deliberately uncomfortable question: what would these laboratory rates imply if they were even roughly calibrated to false confession in the field?
The answer is difficult to take seriously.
The FBI reported roughly 7.5 million arrests in 2024. Suppose, for illustration, that only 10% involved the custodial interrogation of an innocent suspect. That would be 750,000 interrogations.
Apply the laboratory rate of 41%, and the implication is roughly 307,000 false confessions every year.
Even the 37% control rate, with essentially no interrogation tactic, implies about 278,000 per year.
The National Registry of Exonerations has documented roughly 375 false confession cases across its entire history.
Obviously, we are not claiming that every false confession results in an exoneration, that the Registry captures every case, or that 10% is the true rate at which innocent people are interrogated.
We also are not claiming that false confession researchers believe there are 300,000 false confessions each year. Most clearly do not.
The calculation is useful because the implied quantity is so implausible.
If a laboratory outcome occurs in 37% to 41% of innocent participants, but nobody thinks the analogous real-world behavior occurs at anything remotely approaching that frequency, then we should ask whether the laboratory and field outcomes are actually the same quantity.
Our calibration analyses formalize that problem. Across a wide range of assumptions, the observed laboratory rates are incompatible with plausible real-world false confession rates.
That is not a small generalizability problem. It goes directly to what the dependent variable means.
Why this matters even if nobody calls 41% a prevalence estimate
One response to this argument is that laboratory researchers do not claim 41% of innocent suspects will falsely confess in the real world.
That is generally correct.
But it does not solve the problem.
Laboratory rates are used to estimate odds ratios, risk ratios, and differences among interrogation tactics. Those findings are then used to say that particular interrogation practices increase false confession risk. They appear in expert testimony, amicus briefs, policy arguments, scientific reviews, and discussions of interrogation reform.
Those quantitative claims still depend on what the experiment’s outcome represents.
Suppose a false evidence ploy increases signing from 37% to 62% in a laboratory task. That tells us something real about the effect of false evidence on behavior in that experiment.
But if the underlying 37% is primarily low-stakes compliance, the experiment has shown that false evidence increases that behavior. It has not automatically shown that the same magnitude of effect applies to the high-cost decision to falsely confess during custodial interrogation.
Relative comparisons do not escape the construct-validity problem.
An effect is only as interpretable as the outcome being affected.
There is a reason these paradigms work so well in the laboratory
There is also an interesting statistical tension here.
False confessions in the field are rare. Rare outcomes are expensive to study experimentally because very large samples are needed to detect differences between conditions.
The average experimental condition in this literature had only about 37 participants.
Those sample sizes become much more workable when the laboratory task produces confession rates of 30%, 40%, 50%, or more.
For example, detecting an increase from 40% to 60% requires vastly fewer participants than detecting an increase from 5% to 7.5%.
So the feature that makes these paradigms experimentally useful is also part of the problem.
They generate enough admissions to make tactic comparisons possible with ordinary laboratory samples because the measured behavior is common.
But real-world false confession is not a common, low-cost behavior.
The laboratory paradigm gains statistical leverage by creating a situation in which innocent people are unusually willing to sign. We should not then forget that fact when interpreting what the resulting numbers mean.
What I think these studies do tell us
None of this means the experimental false confession literature should be thrown out.
These studies demonstrate that people will sometimes comply with accusations they know are false. They show that authority, false evidence, minimization, and other social influences can change behavior. They allow clean experimental comparisons that field research often cannot.
Those are worthwhile contributions.
But I think the evidence supports a narrower interpretation of them.
These paradigms are useful model systems for studying compliance, social influence, and responses to accusation. They can tell us whether one experimental manipulation produces more admissions than another under controlled conditions.
What they do not currently provide is a calibrated quantitative model of false confession under custodial interrogation.
That distinction becomes especially important when the findings leave the laboratory.
If a result is being used in court, policy, training, or debates over interrogation practice, it is not enough to say that the underlying experiment was randomized and internally valid. The scientific evidence also has to fit the real-world phenomenon to which it is being applied.
A beautifully controlled experiment can answer the wrong question with great precision.
The larger problem is construct validity
Ecological validity is the obvious concern. College students are not criminal suspects. Pressing a computer key is not homicide. Signing a laboratory form is not risking years in prison.
But I think the deeper issue is construct validity.
Those differences in stakes may not simply make the laboratory a weaker version of the real world. They may change the behavior itself.
The decision to sign a harmless experimental statement may involve acquiescence, trust in the experimenter, demand characteristics, inconvenience avoidance, conflict avoidance, or a belief that the statement does not really matter.
A criminal suspect considering a confession faces a fundamentally different decision environment.
The motivational, legal, and personal consequences that define that environment are precisely the features researchers cannot ethically reproduce.
That creates a hard problem for the field. It may not be possible to build a laboratory false confession paradigm with both experimental control and genuine fidelity to the phenomenon.
But the ethical difficulty of measuring something does not permit us to assume that a convenient proxy measures it.
That is ultimately the argument of the paper.
Across decades of experiments, the pattern we observe looks more like low-stakes compliance with an accusation than a calibrated laboratory version of interrogation-induced self-incrimination.
That does not make the experiments meaningless.
It does mean we should be much more careful about what we say they mean.
Explore the results
We built an interactive walkthrough of the paper, including the different paradigms, control-condition rates, variation across experimental designs and countries, calibration analyses, interrogation-tactic comparisons, and statistical power calculations.
The paper
Mourtgos, S. M., & Adams, I. T. (2026). What do laboratory false confession paradigms measure? A calibration meta-analysis. Journal of Quantitative Criminology. Advance online publication. https://doi.org/10.1007/s10940-026-09687-1
The article is open access. Replication data and code are available at github.com/smourtgos/false-confession-meta-analysis.