LLMs Outperform Symbolic Baselines in PDDL Domain Repair
August 19, 2026
Open-weight LLMs can repair errors in Planning Domain Definition Language (PDDL) models, reaching an F1 score of 0.87 compared to a symbolic baseline of 0.49. While highly effective, performance drops significantly in complex domains like Thoughtful, where the test pass rate fell to 0.06.
HOW THIS AFFECTS YOU
●
builderYou can use LLMs to automate the correction of symbolic world models in AI planning workflows.
●
researcherThe gap between F1 and pass rates suggests significant reliability issues in complex reasoning tasks.