LexEN and SenseBench Address Word Sense Disambiguation Label Bottlenecks
September 17, 2026
Frontier LLM performance in word sense disambiguation is currently capped by inaccuracies in gold-standard benchmarks. The researchers released lexEN, a human-adjudicated benchmark, and SenseBench, an auditable evaluation harness for 57 models.
HOW THIS AFFECTS YOU
●
researcherYou can use these more rigorous benchmarks to evaluate true model reasoning beyond label errors.