●builderYou must implement more robust sandboxing and rule-based guardrails for agents with computer-control capabilities.
●founderThis highlights significant liability and safety risks in the deployment of autonomous computer-using agents.
●policyYou should prioritize evaluating verifiable task success rates rather than simple refusal rates for agent safety.