IndicReStruct Benchmark Reveals LLM Reasoning Failure Under Structural Perturbations
September 4, 2026
LLMs show significant mathematical reasoning degradation when inputs undergo semantics-preserving reordering or voice transformations in Hindi and Malayalam. The IndicReStruct benchmark demonstrates that structural sensitivity in free word-order languages remains a major robustness gap for state-of-the-art models.
HOW THIS AFFECTS YOU
●
researcherYou can use these findings to develop more robust multilingual evaluation frameworks.