Four interrogation approaches, one study, two different answers.
Researchers score interrogation methods with a number called diagnosticity — how much more often guilty suspects confess than innocent ones. By that score the winner is the study’s No tactic condition: no technique at all, just asking.
But ask a different question — which approach is actually best at telling a guilty suspect from an innocent one? Call that the method’s separation. On separation, No tactic drops to third, and Minimization, which diagnosticity ranks third, comes first.
The two answers disagree because of a third quantity the score quietly absorbs: the threshold — how hard it was to get a confession out of anyone at all, guilty or innocent.
All of this comes from the same four numbers. This dashboard is about why they disagree, and which one you should care about.
Data from Russano and colleagues (2005), the study that introduced this scoring method. Click the buttons to re-sort. Watch where No tactic lands.
| Interrogation approach | Of 100 guilty, how many confessed |
Of 100 innocent, how many confessed |
Diagnosticity | Separation | Threshold |
|---|
Every courthouse has one, and it has two completely separate things going on. First, how good the machine is — whether it can tell a handgun from a belt buckle. Second, where the sensitivity is set — how much metal it takes before it beeps.
Turn the sensitivity down and the machine almost never beeps. When it does, it is usually something real, so the proportion of alarms that are genuine goes way up. You have not bought a better machine. You have bought a quieter one — and you are now walking guns past it.
Diagnosticity is that proportion. It goes up when you turn the sensitivity down.
Any interrogation method has two separate properties. Diagnosticity blends them into one number.
Separation is how far apart guilty and innocent suspects are — the method's actual ability to tell them apart. Threshold is how hard it is to get a confession out of anyone at all, guilty or innocent. Separation is a skill. Threshold is a setting.
Move the two sliders below. Watch what happens to diagnosticity when you change only the threshold.
On the companion dashboard, The Risk of False Confession Wrongful Convictions, you set sensitivity and specificity with two separate sliders. Those are the same two rates shown here — and they are not free to move independently. Any pair you can name is produced by exactly one combination of separation and threshold. The dials above are just a different set of coordinates for the same point.
Diagnosticity is an odds multiplier, not a risk.
A diagnosticity of 7.67 means guilty suspects confess about 7.7 times as often as innocent ones. It does not mean a confession is 7.7 times more likely to be true. Getting from one to the other requires something the score does not contain: how many of the people being interrogated actually did it.
Laboratory studies fix that number at 50 out of 100 by design, because they assign equal numbers of guilty and innocent participants. Real interrogation populations are not laboratories.
Because diagnosticity is a ratio, it throws away the count. These two techniques score identically.
Setting the share of guilty suspects is the hardest and most consequential input on this page — and most published false-confession statistics get its role backwards. That is the entire subject of the companion dashboard.
Diagnosticity divides by the false confession rate — a small number measured on a small group.
That makes it jumpy. Run the same study twice and you can get very different scores. Run it and have no innocent participant confess, and the score does not exist at all: you cannot divide by zero.
Press the button. Each press runs the same experiment again — 37 guilty and 37 innocent suspects, at the real No tactic rates.
One published study reported a diagnosticity of 87.0. Not one innocent participant in that condition confessed — so the true value of the ratio was undefined. The 87 appears to come from substituting 1% for the observed 0%.
The general point: when a diagnosticity figure is reported, ask what the false confession rate was. If it was zero, the ratio could not be computed at all — so whatever number appears in its place came from a decision about how to handle the zero, not from the experiment.
Everything below is the technical machinery behind the rest of the dashboard. None of it is needed to use the tool — it is here so the work can be checked.
Write TCR for the true confession rate (the share of guilty suspects who confess) and FCR for the false confession rate (the share of innocent suspects who confess). Write z( ) for the standard normal quantile — the z-score that cuts off a given share of a bell curve.
Separation d′ = z(TCR) − z(FCR)
Threshold c = −½ [ z(TCR) + z(FCR) ]
and running it backwards: TCR = Φ(d′/2 − c), FCR = Φ(−d′/2 − c)
Diagnosticity is the slope of the line from the origin to the point (FCR, TCR). Moving the threshold slides that point along a fixed accuracy curve, which changes the slope. That is the entire mechanism this dashboard demonstrates.
Each condition in the source study gives exactly one operating point — a single pair of rates. Reading that point as a separation and a threshold requires a measurement model. We use the standard equal-variance Gaussian signal detection model: both the guilty and innocent groups are treated as bell curves of the same width.
With one point per condition, the data cannot trace the shape of the curve, and cannot test that assumption. So separation and threshold here are model-implied quantities, not directly observed ones.
That is a real limitation and we state it plainly. The honest comparison, though, is this: using diagnosticity as if it were a criterion-free measure of a method's accuracy requires a stronger assumption — that the slope of the line from the origin does not depend on the threshold — and that assumption is almost never stated at all. Here it is at least visible.
If no innocent participant confesses, the observed false confession rate is zero and diagnosticity divides by zero. It is undefined — not infinitely good. This dashboard prints “undefined” rather than an infinity symbol for exactly that reason.
Separation has a standard, principled repair for the same situation: add one half to every cell before computing the rates (the log-linear correction). Applied to the published condition with a zero false confession rate, that yields a finite, high separation of roughly 2.5 to 2.9, depending on which standard correction you use. The ranking survives all of them.
A similar correction could be applied to diagnosticity to make it finite. It would not fix the real problem: the corrected ratio would still be tied to a single operating point, and would still rank methods by a mixture of separation and threshold.
The “Run the study again” button draws a fresh sample of guilty and innocent suspects from the real No tactic rates (46% and 6%) and recomputes both scores from the simulated counts. It is the same procedure as the paper's Monte Carlo, which used 50,000 replications.
At 37 per group the ratio is undefined in 10.3% of replications; at 20 per group, 28.5%; at 80, 0.8%; at 200, effectively never. Among the replications where it can be computed at 37 per group, the mean is 9.21 against a true value of 7.67, and the upper 2.5% tail reaches 21.
The four conditions are small, so the reversal could in principle be a fluke of four point estimates. The paper checks this by drawing each condition's rates from the posterior distributions implied by its cell counts and re-ranking 50,000 times.
The diagnosticity ordering differs from the separation ordering in 87% of draws. The probability that Minimization genuinely out-separates No tactic is 0.70. No tactic ranks first on diagnosticity in 65% of draws, but first on separation in only 17%.
So the disagreement is not an artifact of treating four rounded percentages as exact. It persists once sampling uncertainty is accounted for.
Across the four conditions, observed diagnosticity spans a 3.79-fold range from lowest to highest. Two counterfactuals ask where that spread comes from:
- Hold separation fixed at its average and let only the threshold vary: the spread is still 3.76-fold.
- Hold the threshold fixed at its average and let only separation vary: the spread collapses to 1.60-fold.
Nearly all of what diagnosticity registers as variation across these conditions travels with the threshold rather than with the ability to tell guilty from innocent.
- Diagnosticity (D)
- The true confession rate divided by the false confession rate. Formally a likelihood ratio — the same kind of number a DNA analyst reports. It multiplies odds; it is not itself a probability.
- Separation (d′)
- How far apart the guilty and innocent groups sit on the underlying scale. Larger means the method distinguishes them better. Called discriminability or d-prime in the literature.
- Threshold (c)
- Where the line for confessing falls. Positive is tight (few confessions of any kind); negative is loose. Called the response criterion.
- True confession rate (TCR)
- Of guilty suspects, the share who confess. A hit rate.
- False confession rate (FCR)
- Of innocent suspects, the share who confess. A false alarm rate.
- The map
- What the literature calls ROC space: false confession rate across the bottom, true confession rate up the side. Every threshold setting for one method traces a curve.
- This is a measurement argument, not a policy verdict. Nothing here says which interrogation method anyone should use. It says the number used to rank them cannot support the ranking.
- Laboratory paradigms are not interrogations. Whether a cheating-paradigm confession indexes real interrogation-induced false confession, or something closer to compliance with an accusation, is a separate and unsettled question.
- The four conditions are small. Roughly 37 suspects per group. Their estimates carry real sampling error — which is itself one of the paper's points.