Wiggle Framework Reveals Epistemic Instability in LLM Judges
August 14, 2026
The Wiggle Framework tests LLM judge robustness across mechanical consistency, single-turn conviction, and multi-turn persistence. Research shows frontier models frequently flip verdicts by 25-91% when subjected to static or adaptive pressure.
HOW THIS AFFECTS YOU
●
researcherYou should account for significant judge instability when using LLMs for automated evaluation.
●
policyThis highlights the unreliability of using single LLMs for safety or political alignment grading.