Expert Behavior Prior Reinforcement Learning via Q-CVAE
July 24, 2026
The Expert Behavior Prior (EBP) algorithm improves online reinforcement learning sample efficiency by using a Q-guided conditional variational autoencoder (Q-CVAE). This method generates high-value expert policy priors directly from online replay buffers to overcome the limitations of static, low-diversity offline datasets.
HOW THIS AFFECTS YOU
●
researcherThis method offers a way to mitigate suboptimal trajectory quality in online RL training.