Safety Benchmarking for Vehicle Voice Command Authorization
September 18, 2026
A new 202-scenario benchmark evaluates LLM decision-making in vehicle voice assistants across seven safety classes. Results show decision alignment ranging from 40.1% for Llama 3.2 3B to 89.1% for Gemini 3.1 Pro Preview.
HOW THIS AFFECTS YOU
●
builderYou must account for significant safety alignment gaps when deploying small local models in critical hardware interfaces.
●
policyThis highlights the urgent need for standardized safety taxonomies in automotive LLM integration.