LOGIC Benchmark for Intent-Grounded Aerospace Electrical System Changes
October 7, 2026
LOGIC introduces a benchmark for evaluating how language models ground engineering requests within electrical traceability graphs. The framework separates candidate-selection errors from downstream propagation errors across 168 scenarios, demonstrating that structured-evidence methods achieve a 1.0000 candidate F1 on anchored selection cases.
HOW THIS AFFECTS YOU
●
builderThis provides a specialized evaluation path for applying LLMs to high-stakes aerospace design automation.
●
researcherYou can use this framework to isolate specific failure modes in model reasoning within complex engineering domains.