HARDEN Method Reduces LLM Accuracy via Constrained Evolutionary Search
September 28, 2026
HARDEN uses constrained evolutionary search to generate more difficult evaluation variants for language models while preserving task semantics and expected outputs. Testing on Qwen3.5 models shows it reduces task-model accuracy by an average of 22.7% and up to 49.9% compared to single-pass baselines.
HOW THIS AFFECTS YOU
●
researcherYou can use this to identify hidden performance gaps in models using existing benchmarks.
●
policyThis provides a way to stress-test model robustness against more complex, real-world enterprise scenarios.