Interactive Analysis

Seeing Confessions Clearly

A method that gets fewer confessions is not the same thing as a method that gets better ones — and the score the field uses to rank interrogation methods cannot tell the two apart.

Scott M. Mourtgos, J. Pete Blair & Ian T. Adams · CrimRxiv preprint  |  Data and code

Four interrogation approaches, one study, two different answers.

Researchers score interrogation methods with a number called diagnosticity — how much more often guilty suspects confess than innocent ones. By that score the winner is the study’s No tactic condition: no technique at all, just asking.

But ask a different question — which approach is actually best at telling a guilty suspect from an innocent one? Call that the method’s separation. On separation, No tactic drops to third, and Minimization, which diagnosticity ranks third, comes first.

The two answers disagree because of a third quantity the score quietly absorbs: the threshold — how hard it was to get a confession out of anyone at all, guilty or innocent.

All of this comes from the same four numbers. This dashboard is about why they disagree, and which one you should care about.

The same four conditions, sorted two ways

Data from Russano and colleagues (2005), the study that introduced this scoring method. Click the buttons to re-sort. Watch where No tactic lands.

Separation
How well the approach tells guilty from innocent. The method’s actual accuracy. Higher is better.
Threshold
How hard it is to get a confession from anyone, guilty or innocent. “Tight” means few confessions of any kind. A setting, not a skill.
Interrogation approach Of 100 guilty,
how many confessed
Of 100 innocent,
how many confessed
Diagnosticity Separation Threshold
Read the last two columns together. The diagnosticity order and the threshold order are identical, method for method. The diagnosticity order and the separation order are not. That is the whole finding: the score is tracking how hard it was to get anyone to talk, not how well the method sorted the guilty from the innocent.
The courthouse metal detector

Every courthouse has one, and it has two completely separate things going on. First, how good the machine is — whether it can tell a handgun from a belt buckle. Second, where the sensitivity is set — how much metal it takes before it beeps.

Turn the sensitivity down and the machine almost never beeps. When it does, it is usually something real, so the proportion of alarms that are genuine goes way up. You have not bought a better machine. You have bought a quieter one — and you are now walking guns past it.

Diagnosticity is that proportion. It goes up when you turn the sensitivity down.

On the source study. Before this paper was circulated, the authors shared it with Dr. Melissa B. Russano and Dr. Christian A. Meissner, two of the authors of the 2005 study reanalyzed here. They reviewed it and confirmed the paradigm is accurately described. They also noted that, to the extent the manuscript implies they present diagnosticity as a measure of posterior probability, they disagree. That is accepted: Russano and colleagues did not present diagnosticity as a probability. The concern here is with how the ratio gets interpreted downstream.

Any interrogation method has two separate properties. Diagnosticity blends them into one number.

Separation is how far apart guilty and innocent suspects are — the method's actual ability to tell them apart. Threshold is how hard it is to get a confession out of anyone at all, guilty or innocent. Separation is a skill. Threshold is a setting.

Move the two sliders below. Watch what happens to diagnosticity when you change only the threshold.

The two dials
Drag either one. Everything on the right updates live.
Separation is locked. The method's ability to tell guilty from innocent literally cannot change while this is on. Now move the threshold and watch diagnosticity anyway.
Separation 1.45
How far apart are guilty and innocent suspects? This is the method's real ability to tell them apart.
0 — coin flip4 — near perfect
Confession threshold Tight (+0.83)
How hard is it to get a confession out of anyone at all? This is a setting, not a skill.
Loose — almost anyone talksTight — almost nobody talks
Try this
Jump to a real condition
Diagnosticity
7.67
Separation 1.45 · Threshold +0.83
46
of 100 guilty confess
6
of 100 innocent confess
Move a slider to see what changed.
The two groups, and where the line falls
Everything to the right of the dashed line confessed. Blue is innocent suspects, orange is guilty. Separation moves the humps apart; threshold slides the line.
The map
The dark curve is every threshold setting for a method of this separation. The dashed gold line is diagnosticity — its steepness is the score. Slide the threshold and the dot travels along the curve while the gold line pivots.
Flip it around: the two rates are not independent

On the companion dashboard, The Risk of False Confession Wrongful Convictions, you set sensitivity and specificity with two separate sliders. Those are the same two rates shown here — and they are not free to move independently. Any pair you can name is produced by exactly one combination of separation and threshold. The dials above are just a different set of coordinates for the same point.

Drag the rates instead
Now you set the two confession rates, and the two dials move on their own.
Of 100 guilty, how many confess 46
Of 100 innocent, how many confess 6
Separation
1.45
implied by these two rates
Threshold
+0.83
implied by these two rates
Two sliders you can move independently quietly assume you can gain on one without paying on the other. For a fixed method, you cannot.

Diagnosticity is an odds multiplier, not a risk.

A diagnosticity of 7.67 means guilty suspects confess about 7.7 times as often as innocent ones. It does not mean a confession is 7.7 times more likely to be true. Getting from one to the other requires something the score does not contain: how many of the people being interrogated actually did it.

Laboratory studies fix that number at 50 out of 100 by design, because they assign equal numbers of guilty and innocent participants. Real interrogation populations are not laboratories.

The interrogated population
The diagnosticity below carries over from The Two Dials.
Of every 100 people interrogated, how many actually did it 50
50 is the laboratory's number, built in by design. It is not a finding about the field.
5 — wide net99 — near certainty going in
Diagnosticity in use
Chance this confession is false
11.5%
approximately 1 in 9
Same score, different room
The curve is one fixed diagnosticity. Only the room changes. The grey diamond marks the laboratory's built-in 50.
Same score, five times the innocent people

Because diagnosticity is a ratio, it throws away the count. These two techniques score identically.

Low-pressure technique
D = 5
2 of 100 innocent suspects confess · 20 false confessions per 1,000 innocent people interrogated
High-pressure technique
D = 5
10 of 100 innocent suspects confess · 100 false confessions per 1,000 innocent people interrogated
A ratio throws away the count. Two methods can post identical scores while one of them puts five times as many innocent people in the position of having confessed. For policy, that difference is the whole question.

Diagnosticity divides by the false confession rate — a small number measured on a small group.

That makes it jumpy. Run the same study twice and you can get very different scores. Run it and have no innocent participant confess, and the score does not exist at all: you cannot divide by zero.

Press the button. Each press runs the same experiment again — 37 guilty and 37 innocent suspects, at the real No tactic rates.

Suspects per group:
The true diagnosticity for this condition is 7.67. Press “Run the study again” and see what a single study would have reported.
0
studies run
0
score undefined
median score
middle 95% of scores
Every study you have run
Each dot is one run of the same experiment. The black line is the true value, 7.67. Red marks at the bottom are runs where no innocent participant confessed and the score could not be computed at all.
At 37 per group — the size of the original study — the score is undefined in about 10% of runs. When it can be computed at all, it averages 9.21 against a true value of 7.67. Separation, computed from exactly the same simulated data, is defined every single time.
A question worth asking

One published study reported a diagnosticity of 87.0. Not one innocent participant in that condition confessed — so the true value of the ratio was undefined. The 87 appears to come from substituting 1% for the observed 0%.

The general point: when a diagnosticity figure is reported, ask what the false confession rate was. If it was zero, the ratio could not be computed at all — so whatever number appears in its place came from a decision about how to handle the zero, not from the experiment.

Everything below is the technical machinery behind the rest of the dashboard. None of it is needed to use the tool — it is here so the work can be checked.

The formulas

Write TCR for the true confession rate (the share of guilty suspects who confess) and FCR for the false confession rate (the share of innocent suspects who confess). Write z( ) for the standard normal quantile — the z-score that cuts off a given share of a bell curve.

Diagnosticity  D = TCR / FCR
Separation  d′ = z(TCR) − z(FCR)
Threshold  c = −½ [ z(TCR) + z(FCR) ]

and running it backwards: TCR = Φ(d′/2 − c), FCR = Φ(−d′/2 − c)

Diagnosticity is the slope of the line from the origin to the point (FCR, TCR). Moving the threshold slides that point along a fixed accuracy curve, which changes the slope. That is the entire mechanism this dashboard demonstrates.

What we assume, and why it matters

Each condition in the source study gives exactly one operating point — a single pair of rates. Reading that point as a separation and a threshold requires a measurement model. We use the standard equal-variance Gaussian signal detection model: both the guilty and innocent groups are treated as bell curves of the same width.

With one point per condition, the data cannot trace the shape of the curve, and cannot test that assumption. So separation and threshold here are model-implied quantities, not directly observed ones.

That is a real limitation and we state it plainly. The honest comparison, though, is this: using diagnosticity as if it were a criterion-free measure of a method's accuracy requires a stronger assumption — that the slope of the line from the origin does not depend on the threshold — and that assumption is almost never stated at all. Here it is at least visible.

When nobody innocent confesses

If no innocent participant confesses, the observed false confession rate is zero and diagnosticity divides by zero. It is undefined — not infinitely good. This dashboard prints “undefined” rather than an infinity symbol for exactly that reason.

Separation has a standard, principled repair for the same situation: add one half to every cell before computing the rates (the log-linear correction). Applied to the published condition with a zero false confession rate, that yields a finite, high separation of roughly 2.5 to 2.9, depending on which standard correction you use. The ranking survives all of them.

A similar correction could be applied to diagnosticity to make it finite. It would not fix the real problem: the corrected ratio would still be tied to a single operating point, and would still rank methods by a mixture of separation and threshold.

How the simulation works

The “Run the study again” button draws a fresh sample of guilty and innocent suspects from the real No tactic rates (46% and 6%) and recomputes both scores from the simulated counts. It is the same procedure as the paper's Monte Carlo, which used 50,000 replications.

At 37 per group the ratio is undefined in 10.3% of replications; at 20 per group, 28.5%; at 80, 0.8%; at 200, effectively never. Among the replications where it can be computed at 37 per group, the mean is 9.21 against a true value of 7.67, and the upper 2.5% tail reaches 21.

Does the reversal survive sampling error?

The four conditions are small, so the reversal could in principle be a fluke of four point estimates. The paper checks this by drawing each condition's rates from the posterior distributions implied by its cell counts and re-ranking 50,000 times.

The diagnosticity ordering differs from the separation ordering in 87% of draws. The probability that Minimization genuinely out-separates No tactic is 0.70. No tactic ranks first on diagnosticity in 65% of draws, but first on separation in only 17%.

So the disagreement is not an artifact of treating four rounded percentages as exact. It persists once sampling uncertainty is accounted for.

Where the spread in diagnosticity comes from

Across the four conditions, observed diagnosticity spans a 3.79-fold range from lowest to highest. Two counterfactuals ask where that spread comes from:

  • Hold separation fixed at its average and let only the threshold vary: the spread is still 3.76-fold.
  • Hold the threshold fixed at its average and let only separation vary: the spread collapses to 1.60-fold.

Nearly all of what diagnosticity registers as variation across these conditions travels with the threshold rather than with the ability to tell guilty from innocent.

Glossary
Diagnosticity (D)
The true confession rate divided by the false confession rate. Formally a likelihood ratio — the same kind of number a DNA analyst reports. It multiplies odds; it is not itself a probability.
Separation (d′)
How far apart the guilty and innocent groups sit on the underlying scale. Larger means the method distinguishes them better. Called discriminability or d-prime in the literature.
Threshold (c)
Where the line for confessing falls. Positive is tight (few confessions of any kind); negative is loose. Called the response criterion.
True confession rate (TCR)
Of guilty suspects, the share who confess. A hit rate.
False confession rate (FCR)
Of innocent suspects, the share who confess. A false alarm rate.
The map
What the literature calls ROC space: false confession rate across the bottom, true confession rate up the side. Every threshold setting for one method traces a curve.
Cite
Mourtgos, S. M., Blair, J. P., & Adams, I. T. (2026). Seeing Confessions Clearly: Diagnosticity, Risk, and Signal Detection: A Signal-Detection Reappraisal of Interrogation Methods and Confession Evidence. CrimRxiv. https://www.crimrxiv.com/pub/emake4tt Underlying study reanalyzed: Russano, M. B., Meissner, C. A., Narchet, F. M., & Kassin, S. M. (2005). Investigating true and false confessions within a novel experimental paradigm. Psychological Science, 16(6), 481–486.  |  Data and code