Scaling network depth to 1024 layers enables significant performance gains in self-supervised, goal-conditioned reinforcement learning. In unsupervised locomotion and manipulation tasks, this depth increased performance of contrastive RL algorithms by 2x to 50x compared to standard shallow architectures.
HOW THIS AFFECTS YOU
●
builderIf building high-capacity RL agents, you should consider much deeper architectures than the typical 2-5 layer standard.
●
researcherScaling depth in RL may be the key to matching the progress seen in self-supervised language and vision models.