Multilingual GSM-Symbolic Quantifies Cross-Lingual Mathematical Capability Transfer
October 5, 2026
This benchmark uses 30,000 symbolic templates across 15 languages to evaluate mathematical reasoning while preventing overfitting. Findings indicate model size (beta = 1.77) and language resource level (beta = 0.77) are the primary determinants of capability transfer.
HOW THIS AFFECTS YOU
●
researcherYou can use these symbolic templates to evaluate if your model's reasoning is truly transferable across low-resource languages.