Transformers Internalize Linguistic Features for Multilingual Readability Assessment
September 11, 2026
An analysis using SHAP and TCAV shows that multilingual Transformer models recover surface-length, syntactic, and lexical-diversity signals used by traditional readability classifiers. The study confirms that XLM-R and language-specific encoders reflect the ordinal CEFR structure across five languages.
HOW THIS AFFECTS YOU
●
researcherYou can leverage these findings to probe how Transformer models internalize specific linguistic properties compared to feature-based models.