[arXiv]score: 0.18
JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language Models
July 23, 2026
JailMeter introduces a dual-feedback optimization framework based on Information Bottleneck theory to filter noise from model responses during jailbreak evaluation. By isolating content relevant to malicious intent, the framework achieves 97.27% accuracy on the JailMeter-Eva benchmark, which contains 330 human-labeled, non-rejected attack instances.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy