LMEnt suite for analyzing knowledge acquisition in language models
September 18, 2026
LMEnt provides a knowledge-rich pretraining corpus and an entity-based retrieval method that outperforms existing tools by 80.4%. It includes 12 pretrained models up to 1B parameters with 4K checkpoints to track how entity mentions translate to downstream knowledge.
HOW THIS AFFECTS YOU
●
researcherYou can use these checkpoints and annotated data to study the precise mechanics of knowledge acquisition during training.