Policy Evaluation

Did Banning Deceptive Interrogation Tactics Affect Juvenile Case Outcomes?

Between 2022 and 2024, nine states banned police from using deception when interrogating juveniles, raising concerns among some that the tradeoff would be fewer cases solved. Across 510,582 serious juvenile-involved incidents in NIBRS (2021–2024), arrest clearances did not decline after adoption, and estimates are null to modestly positive across every method. Prosecution-declined clearances fell in adopting states, but the same decline appears for adult suspects, and triple-difference analyses indicate the pattern is not juvenile-specific.

0
Treated states
0
Control states (incl. DC)
0
Juvenile-involved incidents
+2.25 pp
Arrest-clearance change after adoption
0.5%
Posterior probability arrests declined
Step 1 · The policy

What nine states banned — and the question it raised

American police may legally lie to suspects during interrogation (within constitutional and statutory limits). Because juveniles are widely considered more susceptible than adults to suggestive questioning and false confession, nine states passed laws between 2021 and 2024 restricting deceptive interrogation of minors — typically by making statements obtained through deception inadmissible in court.

These statutes sit between two research traditions. One treats deception as a contributor to unreliable admissions, with juveniles of special concern. The other treats restrictions on interrogation as potential constraints on police fact-development — suggesting a possible tradeoff: fewer confessions, fewer cases solved, more cases that prosecutors decline. A competing possibility is substitution: agencies adapt to legal constraints, and investigators shift toward other lawful techniques. Which pattern the data show is an empirical question. This page walks through it.

How to read this page: it is a descriptive companion to a peer-reviewed study, scoped to what administrative data can show for 2021–2024. “No decline” is the claim being tested; where estimates come out positive, we report them as associations, not causal benefits — no mechanism linking deception bans to more cleared cases is claimed. Two states (Nevada and California) adopted in July 2024 and contribute only about six post-ban months.
Step 2 · The natural experiment

Nine states adopted; 42 stayed put

Staggered adoption — Illinois and Oregon first (January 2022), California and Nevada last (July 2024) — creates a natural comparison: how did case outcomes move in ban states, relative to the 42 state units (including Washington, DC) that never adopted? The data are FBI NIBRS incident records: 510,582 serious juvenile-involved incidents reported by 11,333 law-enforcement agencies from 2021 through 2024.

Treated states are shaded by statute strength (darker = stronger law; see Step 3). Hover a state for its effective date and statute score. Gray states never adopted a ban and serve as controls.

Step 3 · The laws

These nine laws are not the same law

The statutes differ on five dimensions: how strong the remedy is (presumptive exclusion vs. weaker), how broadly “deception” is defined, how general the prohibition is, and whether good-faith or public-safety exceptions soften it. Summing the first three and subtracting the exceptions gives a composite strength score from 2 to 8. Oregon and Connecticut have the strongest laws; Nevada the weakest. Step 9 uses this variation.

StateEffectiveRemedyBreadthScope Good-faith exc.Public-safety exc.CompositeStrength

Remedy, breadth, and scope are each coded 1–3 (higher = stricter/broader); exceptions subtract 1 each. Composite = remedy + breadth + scope − good-faith − public-safety. Utah and Colorado additionally require recording; Illinois and Connecticut also address deception more broadly than the juvenile provisions scored here.

Step 4 · The yardsticks

Two ways a juvenile case gets resolved — and two hypothesized changes

NIBRS records how each incident is cleared. We track the two outcomes the hypothesized tradeoff speaks to directly.

Arrest clearance

The concern: could fall after a ban

The case is closed because someone was arrested, charged, and turned over for prosecution — the everyday sense of a case being “solved.” Across the full panel, roughly 41% of these juvenile-involved incidents end in an arrest clearance.

≈ 41% of incidentsconcern: decline

Prosecution-declined exceptional clearance

The concern: could rise after a ban

Police identified the offender and presented the case, but the prosecutor declined to pursue it — the incident is closed “exceptionally.” If bans weakened the evidence in juvenile cases, these refusals could become more common. They are rare to begin with: about 3.3% of incidents.

≈ 3.3% of incidentsconcern: increase

Throughout this page, probabilities are reported in the direction of the hypothesized cost: for arrest clearance, the posterior probability of a decline; for prosecution-declined, the probability of an increase. Small values mean the data point away from the hypothesized tradeoff.

Step 5 · The headline

Arrest clearances did not fall

The primary model is a hierarchical Bayesian difference-in-differences fit on the agency-month panel (11,333 agencies, roughly 153,000 agency-months), which lets the treatment effect vary by state and agency and weights by caseload. It compares ban states, after adoption, to what the never-treated states say we should have expected.

+2.25 pp
arrest clearance after adoption vs. the no-ban expectation (95% CrI +0.99 to +3.60)
0.5%
posterior probability that arrest clearances declined at all
−0.26 pp
prosecution-declined change (95% CrI −0.76 to +0.25); P(increase) = .13

Dots are posterior mean effects in percentage points; bars are 95% credible intervals. The dashed line at zero is “no change.” A tradeoff would appear below zero for arrest clearance and above zero for prosecution declined.

The posterior probability of any decline in arrest clearances within the study period is 0.5%. We treat the positive point estimate as descriptive, not causal: concurrent reforms, agency adaptation, or compositional shifts could produce it. On the available evidence, and over the study period examined, adoption is not associated with an aggregate decline in arrest clearances.

Step 6 · The test

The hypothesized tradeoff vs. what the data show

If restricting deception reduced case resolution, several patterns should be observable. Each row is one of them, next to what the data actually show for 2021–2024.

If bans cost cases… What the data show Verdict
Arrest clearances should fall in ban states +2.25 pp [+0.99, +3.60] P(decline) = .005; the point estimate runs modestly positive ✗ Not supported
The strictest laws should hurt the most Strongest statutes (OR, CT) show the largest positive annual change: +2.28 pp/yr [+1.00, +3.51] the gradient runs opposite to the hypothesized direction, though strength cells are thin (Step 9) ✗ Not supported
Each ban state should show its own decline 9 of 9 state-specific synthetic-control estimates are positive +1.33 to +5.20 pp (Step 8) ✗ Not supported
Prosecutors should refuse more juvenile cases Declinations fell (−0.26 pp), and fell for adults in the same states too juvenile-minus-adult difference is null (Step 10) ✗ Not supported
Any null should be the artifact of one model Twelve estimation strategies, one direction every arrest estimate is null-to-positive (Step 7) ✗ Not supported

To be precise about scope: this scorecard tests the hypothesis within the first one to three years of each ban, in states that report to NIBRS. It does not rule out longer-run effects, and it does not establish that bans improve clearances — only that the hypothesized deterioration is not observed in any of these tests.

Step 7 · Many methods

Twelve estimates, one direction

Any single model can mislead, so the paper re-asks the question with different estimators, samples, and windows: frequentist two-way fixed effects, randomization tests, synthetic controls, restricted samples, an extended 2016–2024 panel, and placebo contrasts. If the aggregate result were an artifact of one specification, the alternatives would disagree.

Dots are point estimates in percentage points with 95% intervals (credible or confidence, per method); rows without a dot are ranges across states or leave-one-out samples. Colors group the methods. The randomization-inference p-value (.47 for arrest) is the paper’s primary frequentist significance measure, because nine treated states are too few clusters for reliable clustered standard errors — by that standard the pooled arrest effect is not statistically significant, merely directionally positive. The augmented synthetic control is reported as directional only (its interval is wide and spans zero).

The methods disagree about magnitude — from +1.05 to +3.53 percentage points for arrest clearance — and about statistical significance. They do not disagree about direction: no method, sample, or window produces evidence of the hypothesized decline.

Step 8 · State by state

All nine states point the same way

Pooled averages can hide a state that really was hurt. So each state gets its own Bayesian synthetic control: a weighted blend of never-treated states tuned to track that state’s pre-ban history (typical tracking error 1–4 percentage points), then projected forward as the no-ban counterfactual.

Average post-ban arrest-clearance effect per state, in percentage points, with 95% credible intervals; hover for each state’s pre/post months and P(decline). Open circles mark Nevada and California, which have only ~6 post-ban months. Indiana is the only state whose interval excludes zero.

Every point estimate is positive — and at two-decimal rounding, four of nine states have P(decline) ≤ .05. The honest summary is not “bans raised clearances in nine states”; single-state intervals are wide, and only Indiana’s (+3.71, CrI +0.61 to +5.86) excludes zero. The consistent finding is that no state shows evidence of a post-adoption decline in arrest clearances.

Step 9 · Statute strength

Stronger statutes do not show larger declines

If statutory restrictiveness reduced case resolution, more restrictive statutes should show more negative post-adoption trajectories. Using the Step 3 composite, the model estimates how each outcome changes per year after adoption at each observed strength level — from Nevada’s 2 to Oregon and Connecticut’s 8.

Estimated post-adoption trend (percentage points per year) at each observed composite score, with 95% credible intervals; blue = arrest clearance, amber = prosecution declined. Hover for the states at each level and the cost-direction probability.

For arrest clearance, the estimated annual change becomes more positive as restrictiveness increases — the opposite of the hypothesized gradient — with the strongest statutes (OR, CT) at +2.28 pp/yr (CrI +1.00 to +3.51). The pattern is monotonic, but the strength cells are thin (one state at scores 2 and 3; two at scores 4 and 8), so the gradient reflects differences among a small number of states as much as a general relationship with statutory strength. For prosecution declined, estimated annual changes are negative below the strongest level; at OR/CT the estimate is +0.14 pp/yr with an interval straddling zero. No gradient consistent with the claim that restrictiveness reduces case resolution appears. As throughout, these are associations, not mechanisms.

Step 10 · The prosecution story

Declinations fell — but not just for juveniles

One outcome did move: prosecution-declined clearances drifted down in ban states — the opposite of the predicted surge. Before crediting the bans, run the placebo: adult suspects in the same states are untouched by juvenile interrogation statutes, so if the decline were the ban’s doing, it should appear only for juveniles.

Estimates in percentage points with 95% credible intervals. “Juv − adult” is the triple-difference: the juvenile-specific residual after subtracting whatever happened to adults in the same states. Zero line = no effect.

Juvenile declinations fell (−1.03 pp), but so did adult declinations in the same states (−0.51 pp, credibly negative) — and the juvenile-minus-adult difference (−0.27 pp, CrI −1.11 to +0.49) is null. The decline is consistent with broader changes in those states’ case processing rather than a juvenile-specific policy effect. The same placebo sharpens the arrest result in the other direction: adult arrest clearances did not rise (−0.74 pp, interval spanning zero), so the juvenile-specific arrest difference is +2.10 pp (CrI +0.54 to +3.68). One footnote: Connecticut is omitted from state-level declination summaries — its base rate is 0.33% and 14 of its 15 post-ban months had zero declined cases, too little signal to summarize responsibly.

Step 11 · Who and what

Where the arrest estimates concentrate

Splitting the arrest-clearance result by offender age and race, and by offense type, shows where the positive drift sits — and where there is none.

Arrest-clearance estimates in percentage points with 95% credible intervals; hover for the prosecution-declined counterpart. Groups are modeled separately; intervals overlap, so differences between groups are suggestive, not established.

Same convention, by offense type (n = incidents). † The prosecution-declined model for murder is not estimable — too few declined murder cases to identify the parameter.

By age and race, the positive arrest-clearance drift is largest for the youngest juveniles (≤13: +3.23 pp) and absent for 16–17-year-olds; both Black (+1.99) and non-Black (+1.11) youth show credibly positive estimates. By offense, results are mixed but broadly consistent with the aggregate pattern: property offenses — burglary, motor vehicle theft — show positive and relatively precise estimates, while several violent and more complex offenses have point estimates near zero or negative with wide intervals spanning both directions. The offense with the most credible evidence of a possible reduction is sex crimes (P(decline) = .89), though the estimate is modest (−0.63 pp) and its interval includes zero (−1.61 to +0.39). The most negative point estimate is murder, but it is highly imprecise (−1.41 pp, CrI −8.35 to +5.71). These patterns are consistent with the possibility that clearance outcomes in serious, complex cases — those that may depend more heavily on statement evidence — respond differently to interrogation constraints; the data do not allow that possibility to be distinguished confidently from noise. These splits are descriptive; we claim no mechanism for any of them.

Step 12 · Caveats

Caveats

A fair reading of these results requires being clear about what they can and cannot settle.

What to keep in mind

  • Aggregate, not individual cases. The analysis is of aggregate clearance rates. It should not be read to imply that no individual case was affected — it remains possible that some cases that would otherwise have been cleared were not, even if such effects are not detectable at the aggregate level.
  • Juveniles only. These statutes govern the interrogation of juveniles, and every estimate here is for juvenile-involved incidents. The results say nothing about what restricting deception in adult interrogations would do — adult outcomes appear on this page only as a placebo comparison, and results could differ for adults.
  • Long-run effects. Nevada and California contribute only ~6 post-ban months; the longest-treated states contribute three years. Costs that emerge later would not be visible here.
  • Pre-existing trend differences. Formal tests reject equal pre-ban slopes (arrest p = .008; prosecution declined p = .025), driven substantially by California’s long pre-period. For arrest clearance the treated states were trending more negative before adoption — a bias running against the positive finding, which makes it conservative. For prosecution declined the pre-trend runs the same direction as the post-ban decline, so that result warrants genuine caution.
  • Mechanisms. We measure case outcomes, not interrogation practice. Whether investigators changed tactics, recorded more, or obtained fewer confessions is not observed — and the positive arrest estimates are associations, not demonstrated effects of the bans. NIBRS also does not flag which incidents involved interrogation at all, so the denominator includes many cases where the restricted tactic could not have mattered, diluting any true effect toward zero.
  • Justice quality. A cleared case is not necessarily a justly cleared case; administrative clearance rates cannot speak to wrongful convictions averted or caused.
  • Coverage. NIBRS reporting grew over the window; results describe reporting agencies, and the consistent-reporter check (see robustness) addresses but cannot eliminate this.
Step 13 · Takeaways

What this means in practice

For legislators & advocates

Over the first one to three years of adoption, the data show no aggregate decline in juvenile case resolution in any adopting state, including those with the strictest statutes. On current evidence, a clearance penalty should not be assumed in debates over these laws — though longer-run effects remain an open question.

For police executives

Aggregate arrest clearances in adopting states held or drifted upward after adoption, including for the youngest suspects, and the decline in prosecutorial declinations appears for adult cases too. One plausible reading is substitution — investigators adapting toward other lawful techniques rather than absorbing the constraint as lost capacity.

For researchers

The absence of an aggregate decline is consistent across methods; the modest positive drift is unexplained. Mechanism data — recording rates, confession rates, interrogation practice — are the next step, along with longer post-periods for the 2024 adopters and court data separating dismissal from declination.

We tried to break the result — it held
  • Shuffle the treatment. Randomization inference (reassigning the 9-state treatment structure across states, 500 permutations) gives p = .47 for the pooled arrest effect and p = .63 for prosecution declined — the pooled effects are not significant by the paper’s primary frequentist standard, consistent with “no decline” rather than “credible increase.” A rank-based randomization test asking whether treated states rank unusually high lands at mean-rank p = .01 for arrest (p = .83 for prosecution declined).
  • Drop California. The largest treated state (and a short post-period) could dominate. Refitting the primary model without it: arrest +2.59 pp [+0.92, +4.08] — the result strengthens slightly.
  • Drop any state. All nine leave-one-state-out TWFE estimates stay positive (+0.95 to +1.54 pp); dropping California gives the largest and only conventionally significant one (+1.54, p = .02).
  • Extend the window. Refit on 2016–2024 (265,412 agency-months), letting treated states carry up to eight pre-years: arrest +1.05 pp [+0.43, +1.68].
  • Hold the sample fixed. Restricting to agencies present in every year 2021–2024 (guarding against NIBRS rollout churn): arrest +3.53 pp [+1.62, +5.48].
  • Check the pre-trends. Treated and control pre-ban slopes differ (arrest p = .008, prosecution declined p = .025). The arrest divergence is negative — it biases against the positive finding. The prosecution-declined divergence runs with its post-ban decline and is flagged as a genuine caution in Step 12.
  • Ask adults. The adult-offender placebo and triple-difference (Step 10) are themselves robustness checks: the arrest effect is juvenile-specific (+2.10 pp DDD); the declination drop is not.