THV-UCB Algorithm for Multi-Objective Pareto Bandits
July 30, 2026
The THV-UCB algorithm optimizes multi-objective slate selection by maximizing dominated hypervolume regret. It achieves a gap-free regret bound of O(d*sqrt(nkT)) when selecting subsets of arms that jointly approximate the Pareto frontier.
HOW THIS AFFECTS YOU
●
researcherThis provides a theoretical framework for selecting optimal subsets of arms in multi-objective optimization settings.