MetroLLM-Bench Evaluates LLMs as Transit Kiosk Runtimes
September 8, 2026
MetroLLM-Bench is a 955-case benchmark designed to test LLMs acting as policy layers for transit kiosks. It evaluates routing, fare calculation, and tool-calling capabilities across six metro systems using 14 deterministic and 8 semantic scoring components.
HOW THIS AFFECTS YOU
●
builderYou can use this to benchmark the reliability of LLMs in high-stakes, structured agentic environments.