RiskChainBench: Evaluating Web Investigation and Message Restoration
September 14, 2026
RiskChainBench evaluates an agent's ability to restore obfuscated platform messages (using emojis or homophones) and conduct web-based evidence investigation. It pairs 3,600 restoration tasks with 600 human-labeled web environments to track how recovery affects downstream tasks.
HOW THIS AFFECTS YOU
●
researcherThis provides a more holistic benchmark for VLM-driven web agents that must perform both text restoration and visual investigation.
●
policyThis helps in developing tools to combat platform abuse and fraudulent redirection campaigns.