FramingQA Benchmark Measures LLM Sensitivity to Compositional Question Framing
September 9, 2026
FramingQA evaluates how subtle rephrasings in law, medicine, finance, and robotics influence LLM responses. The benchmark tests three levels of bias injection to quantify how much model advice is tainted by user-implied stances.
HOW THIS AFFECTS YOU
●
researcherThis exposes a critical vulnerability in how models handle compositional reasoning under biased premises.
●
policyThis highlights the need for safety guardrails in high-stakes domains where framing can lead to incorrect expert advice.