A paper posted to arXiv on 8 October argues that the bottleneck in very large-scale reinforcement learning for robot control is not compute but which starting configurations the simulator hands the policy. Train across up to a million parallel environments with resets sampled uniformly, and a growing share of that data is spent on configurations the policy has already mastered or cannot yet attempt at all.
The authors, from the University of Washington with one also at NVIDIA, propose Success-Guided Sampling: an adaptive sampler that concentrates training on configurations sitting at the edge of what the policy can currently do. Their reported results are stark at scale. On quadruped locomotion across all terrains at one million environments, the method reaches a 0.73 success rate against 0.54 for prioritised level replay and 0.00 for both uniform and linear curriculum sampling. A nut-and-bolt assembly task on a Franka arm goes to 0.70 against 0.06 and 0.05.
They then distil the manipulation policies into versions that work from camera images alone and report zero-shot transfer to a real UR5e assembling parts on a physical task board, with no demonstrations and a single reward function across tasks.
The paper is listed for CoRL 2026. These are the authors’ own figures on their own benchmarks, not independent replication, and the locomotion and assembly tasks were chosen by the people proposing the method. Read the preprint or the project page with the videos.
