Cultural and Linguistic Failures of LLMs in Urdu Story Generation
September 11, 2026
Evaluation of GPT-5.1, Qwen-3-Max, and DeepSeek-3.1 on Urdu-Stories reveals significant grammatical, semantic, and cultural errors. The study finds that even few-shot prompting fails to resolve pervasive cultural shallowness and lack of coherence in low-resource language generation.
HOW THIS AFFECTS YOU
●
researcherCurrent multilingual LLMs remain unreliable for high-fidelity generation in low-resource languages like Urdu.
●
designerAvoid relying on current LLMs for culturally nuanced creative content in non-Western contexts.