FreqBLiMP evaluates LLM grammatical robustness under lexical rarity
September 9, 2026
FreqBLiMP extends the BLiMP benchmark by regenerating 67 grammatical paradigms under explicit Zipf-frequency regimes. Results show that while sentence likelihood decreases monotonically with lexical rarity, overall contrastive acceptability accuracy remains relatively stable across model scales.
HOW THIS AFFECTS YOU
●
researcherYou can better assess if your model's linguistic capabilities are driven by frequency biases rather than grammar.