Pac-Bench Evaluates One-Shot Code Generation for Pac-Man Games
September 28, 2026
Pac-Bench measures a model's ability to generate a fully functional Pac-Man game in a single HTML file from one prompt. The benchmark requires the model to complete the task without any follow-up instructions or iterative debugging steps.
HOW THIS AFFECTS YOU
●
builderYou can use this to benchmark how reliably your agents can generate complete, single-file software assets.
●
researcherIt provides a strict one-shot evaluation metric for zero-shot code generation capabilities.