Sakana AI and the University of Tokyo have published SAIL, a method for getting a robot to copy a task from a handful of demonstrations without retraining anything. The work is due to be presented at IROS2026.
The setup uses two vision-language models. One acts as a trajectory generator, conditioned on a few successful demonstrations. The other watches the resulting video from a simulator and identifies where the robot’s progress stalled. Monte Carlo tree search uses that feedback to decide which candidate trajectory to pursue next.
The headline result is about search budget rather than model size. Across six manipulation tasks in simulation, raising the number of candidates from one to 45 lifted the average rate of finding a successful trajectory from 25% to 73%. That is a data point for an argument currently live in robotics: how much of the gap between a plan and a working action can be closed by spending compute at inference time, rather than by collecting more demonstration data.
The caveat is the important one. These are simulation results; Sakana’s post reports no results on physical hardware, and a trajectory that works in a simulator is not the same as one that works on a real arm. The write-up is on Sakana’s site.
