Research shows that including a completed audit-repair episode in an LLM's context reduces false alarms by 2.8 to 11.5 percentage points. This context shift makes the model's verification behavior more lenient, contradicting standard accumulated-message predictions.
HOW THIS AFFECTS YOU
●
builderThis affects the reliability of your automated testing and verification pipelines.
●
researcherYou should account for context-induced leniency when designing automated checker-fixer loops.