Same Formulas, Different Semantics: Do Language Models Follow Modal Logic Specifications?
August 6, 2026
Four of five recent LLMs perform below baseline when evaluated on modal logic tasks where identical premises require different semantic conclusions. Enabling reasoning mode increases DeepSeek V4 Flash performance from 4.4% to 88.1% on these tasks. This suggests semantic adherence depends more on inference modality than model scale.