LLM Reasoning Performance is Substantively Shallow in European Court of Human Rights Cases
August 19, 2026
A study of GPT 5.4 on ECtHR legal cases reveals that while models produce structurally complete responses, their legal reasoning remains substantively shallow. Findings show that LLM-as-a-Judge evaluators align weakly with human legal experts in these domains.
HOW THIS AFFECTS YOU
●
researcherThe results suggest current reasoning benchmarks may not capture domain-specific depth.
●
policyBe cautious when using LLMs for automated legal analysis or case forecasting.