Scott M. Mourtgos, Ph.D. ← Back to Dashboards

100% of Shark Attack Victims Were Wet.

A short, friendly tour of why case files can describe the people something happened to, but can't tell you how risky anything is. Two silly stories, one switch, one slider.

tl;dr
If the only people in your data are the people something happened to, you can describe them, but you can't say how much more likely that something was for anyone. A risk comparison is a ratio. A ratio needs two groups. Outcome-selected data has one.
Example one

The wedding buffet

Forty guests got sick after the reception. Somebody with a clipboard interviewed all forty. Thirty-six of them had eaten the potato salad. Case closed?

Example two

Shark attacks

A lifeguard reviews every shark attack on record. All sixteen victims were wet (obviously). Fifteen were wearing a swimsuit. Should you swim in a tuxedo?

The part that matters

Same case file. Three different worlds.

Three different wedding receptions. In every one, forty guests got sick and thirty-six of them ate the potato salad. The case files are identical. The worlds are not.

Case files only. Can you tell which world is which?
Show everyone
Okay, but seriously

This is the same problem in my own field.

I study false confessions and wrongful convictions, and a lot of what we know about them comes from case files: databases of exonerations, collections of proven false confessions, and so on. Those collections are enormously valuable. They tell us what wrongful convictions look like, demonstrate possible mechanisms behind them, and show what tends to appear in them. I rely on them constantly.

What they can't do is what the potato salad can't do. "X% of exonerees falsely confessed" is a fact about the people in the file. It is the sticky note. It does not, by itself, tell you how much more likely a wrongful conviction becomes when someone confesses, or when a particular interrogation tactic is used. That is the gauge, and the gauge needs the people who confessed, or were interrogated the same way, and were not wrongfully convicted. They aren't in the file, so we simply do not know that number.

🥗 The potato salad

  • In the file: 90% of the sick guests ate the salad.
  • Not in the file: the guests who ate it and were fine.
  • The question people want answered: "Is the salad dangerous?" Needs both.
  • Without both: the slider can put the truth anywhere from "protective" to "very dangerous," and the file looks identical.

⚖️ The research

  • In the file: a large share of exonerees falsely confessed, or were questioned with tactic T.
  • Not in the file: the people who confessed, or faced tactic T, and were never wrongfully convicted. (Also missing: the wrongfully convicted who were never exonerated.)
  • The question people want answered: "How much does tactic T raise the risk of a wrongful conviction?" Needs both.
  • Without both: the true risk could be large or small and the database looks the same.
The rule If your sample was chosen because the outcome happened, you can describe the cases. You cannot say how much anything changed the odds of the outcome. That takes the people it didn't happen to.

None of this is controversial in research methods. It is why the textbooks warn against "selecting on the dependent variable." I made this page because the idea is genuinely simple, and these examples make it easier to understand than the equations do.

If you'd like the grown-up version, with the actual studies and the actual numbers, the interrogation duration and false confession risk dashboards walk through what it would take to estimate the denominator for real.

My work on this problem.

Mourtgos, S. M., & Adams, I. T. (2026). Recalibrating the risk of false confession wrongful convictions: Interrogation tactics and inverse probability. Journal of Criminal Justice, 103, 102600. https://doi.org/10.1016/j.jcrimjus.2026.102600

Mourtgos, S. M., & Adams, I. T. (2026). Interrogation duration and the estimation of false confession wrongful conviction risk: A reply to Smith and colleagues. Journal of Criminal Justice, 107, 102747. https://doi.org/10.1016/j.jcrimjus.2026.102747

Mourtgos, S. M., & Adams, I. T. (2026). What do laboratory false confession paradigms measure? A calibration meta-analysis. Journal of Quantitative Criminology. https://doi.org/10.1007/s10940-026-09687-1

About the numbers. Every population on this page is invented. Each has 200 people. The number of cases, and the share of cases with the feature, is fixed for each story; the slider only changes how common the feature is among the people the outcome didn't happen to. The risk ratio is P(outcome | feature) ÷ P(outcome | no feature), computed from the 2×2 table. Nothing here estimates anything about real weddings or sharks.