Adaptive Tversky Policy Optimization for Multi-label Video Safety Detection
October 2, 2026
Adaptive Tversky Policy Optimization (ATPO) uses reinforcement learning to enable controllable precision-recall trade-offs in multi-label video safety detection. The framework introduces an Adaptive Tversky Reward to dynamically adjust penalties for false positives and false negatives across different unsafe categories.
HOW THIS AFFECTS YOU
●
researcherYou can use this RL framework to optimize models for specific moderation requirements rather than binary classification.
●
policyThis allows for more nuanced safety enforcement by tuning sensitivity for different types of harmful content.