●builderBe aware that model performance in non-English languages may be constrained by API or generation limits rather than intelligence.
●researcherEnsure your multilingual evaluations use length normalization or variable token budgets to avoid reporting artifacts.