A group from Vesoma, ETH Zurich, the Max Planck Institute for Intelligent Systems, the University of Tuebingen and the ELLIS Institute has published VioLA, a humanoid control policy that gets around the shortage of robot training data by learning mainly from recordings of people.
The trick is what the policy predicts. Instead of joint commands, VioLA outputs body and hand motion latents, which pretrained controllers execute on the robot. Because motion encoders map human and robot movement into one latent space, a video of a person becomes a labelled example in the policy’s action space. The pool runs to 140.6 million frames, 93.2% of them human.
The reported results are strong: 100% success on locomotion instructions on a real Unitree G1 with no task-specific fine-tuning, against 16.7% and 0% for two existing generalist policies, and 88.6% on manipulation.
The numbers need their denominator, which the paper gives. The main evaluation is thirteen tasks on a single robot with five trials each, so 100% means five out of five, and the authors say broader testing across tasks, environments and robot bodies is needed. They also report that robot-only training scores higher on manipulation than the human-led recipe, so the central claim does not hold uniformly. The hand decoder has no contact-force feedback, and code and checkpoints are promised rather than posted.
