Data-driven hypothesis generation for natural language question difficulty
October 2, 2026
This method uses Item Response Theory from LLM responses to estimate question difficulty and then prompts LLMs to generate natural-language explanations for those differences. The approach moves beyond single-score difficulty metrics to provide qualitative reasoning for why specific questions are harder.
HOW THIS AFFECTS YOU
●
researcherThis provides a way to move from scalar difficulty metrics to interpretable, linguistic explanations of model failure modes.