QDOS: Offline-to-Online Skill Learning via Advantage-Weighted Quality-Diversity
August 21, 2026
QDOS introduces a unified pipeline for hierarchical skill policies using an Advantage-Weighted Quality-Diversity pretraining objective. This method weights skill extraction and diversity by estimated advantage, reducing the dependency of low-level policies on raw dataset quality.
HOW THIS AFFECTS YOU
●
researcherThis approach provides a more robust way to extract diverse skills from imperfect offline datasets for reinforcement learning.