LLMs Outperform BERT and TF-IDF in Out-of-Distribution Bloom Level Classification
September 24, 2026
Evaluation of Bloom-level pedagogical classifiers shows significant performance degradation on AI-generated datasets. While TF-IDF models drop to a 0.48 Macro F1-score on out-of-distribution data, LLMs and BERT demonstrate superior robustness for assessing educational content quality.
HOW THIS AFFECTS YOU
●
researcherYou should account for dataset shift when using automated classifiers to evaluate AI-generated educational materials.