CNeo-Bench introduces a 4,759-item benchmark testing how LLMs handle Chinese neologisms involving phonetic substitution and visual character decomposition. Testing 18 LLMs shows most score below 40% on definition generation and struggle to restore source forms from semantic paraphrases.
HOW THIS AFFECTS YOU
●
builderThis informs the need for better handling of slang and visual-based text in Chinese localization.
●
researcherYou can use this to evaluate linguistic robustness in non-English training data.