SA-Bench Reveals High Semantic Drift in AI-Generated Research Code
August 26, 2026
SA-Bench evaluates LLM agents on their ability to reproduce machine learning papers using 1,491 Semantic Alignment Units (SAUs). Even top configurations like Claude+PaperCoder achieved a mean SAU score of only 0.301, showing significant divergence from original paper specifications.
HOW THIS AFFECTS YOU
●
builderYou should be cautious when using LLM agents to translate theoretical specifications into functional code.
●
researcherThis highlights the current unreliability of using LLM agents for automated literature implementation.