Decision model performance benchmarked via Pac-Man gameplay
October 8, 2026
A comparative benchmark of six decision models, including GPT-6 Luna and Clef Flash, was conducted using Pac-Man gameplay. GPT-6 Luna achieved the highest average score of 2,568 with 179 ms latency, while Laya recorded the lowest latency at 104 ms.
HOW THIS AFFECTS YOU
●
builderYou can use this open-source leaderboard to evaluate the real-time responsiveness of your fine-tuned decision models.
●
researcherThis provides a real-time latency and reasoning benchmark for decision-making endpoints.