Legal LLM Benchmarks Fail to Correctly Couple Answer and Authority
August 5, 2026
Testing on Taiwan bar-examination items shows that model answer correctness and legal authority grounding are decoupled. Up to 42.4% of models provided correct answers without the required legal citations, making answer-only scoring an unreliable proxy for legal reasoning.
HOW THIS AFFECTS YOU
●
builderDo not rely on accuracy scores alone when building legal-tech applications.
●
policyYou must require citation-based evaluation for legal AI to ensure actual grounding.