TurEngMix Benchmark Reveals High Error Rates in Turkish-English Code-Mixing
September 9, 2026
A new 5.5K post corpus and benchmark for Turkish-English code-mixing shows that current LLMs struggle with morphologically integrated tokens. GPT-4o and Qwen exhibit NER error rates over 5x higher on mixed-language tokens compared to monolingual ones.
HOW THIS AFFECTS YOU
●
builderExpect significant performance degradation if your product relies on Turkish-English code-mixed social media data.
●
researcherThis benchmark provides a new target for improving NER and LID in morphologically complex code-switching scenarios.