PIJ Benchmark for LLM Criminal Profiling and Investigation
September 18, 2026
The Profiling, Investigation, and Judgment (PIJ) benchmark uses 2,500 real homicide cases to evaluate LLM performance on criminal profiling, crime reconstruction, and sentence prediction. Testing across nine models shows significant performance degradation when models move from explicit fact extraction to abductive reasoning tasks.
HOW THIS AFFECTS YOU
●
researcherYou can use this benchmark to test the reasoning limits of LLMs in complex, abductive legal scenarios.
●
policyYou should be aware of the systematic reasoning failures in models when applied to critical criminal justice tasks.