A page from Sakana AI and the University of Tokyo introduces SAIL. The work was done during Makoto Sato's internship at Sakana, with Yusuke Iwasawa of the University of Tokyo and Yujin Tang and So Kuroki of Sakana. The paper was accepted to IROS 2026.
The question is whether a vision-language model that has already learned from images, video, and language can generate robot trajectories without a weight update. The page places GPT-6 Astra in earlier context: Astra has done several manipulation tasks on a physical robot. SAIL's own experiments use Gemini Robotics-ER 1.5 as both the model that generates trajectories and the model that scores task progress, and they do not update its weights.
[1]
From a few successful demonstrations, the policy model generates a full trajectory of end-effector poses and gripper states. SAIL runs that trajectory in simulation. An evaluation model watches the video and estimates progress through predefined subtasks, then returns scores at selected waypoints so Monte Carlo tree search can revise the trajectory. Only the chosen trajectory is sent to the physical robot.
The simulation is ALOHA, with six manipulation tasks and 20 initial configurations each. Raising the search budget from one node to 45 nodes raised the average rate of finding a successful trajectory from 25% to 73%. The averages at 6, 15, 30, and 45 nodes are 55%, 65%, 71%, and 73%. In the same table, breadth-first search at 15 nodes averages 51, depth-first search averages 37, and a single generation stays at 25. This success rate is how often the search finds a trajectory that passes the simulator's ground-truth check inside the budget. The scores that guide the search are the model's estimates of completion. Those are not the same number.
[1]The physical robot is a LeRobot SO-101 arm, and the task is placing a block in a bowl. The scene is reconstructed from color and depth. With randomized object positions and a budget of 15 candidate trajectories, the method succeeded in five of six trials. They attribute the remaining failure to pose-estimation error and to differences in contact between simulation and the real setup. An imitation policy trained on the successful search trajectories also succeeded in five of six trials and reduced the reported execution time relative to the tree search. The page does not say by how much.
The page lists three limits. Trajectories run open-loop, with no visual feedback during the motion. Searching and revising in simulation costs extra computation. The real-world evaluation is one task and six trials per method, so wider testing is still needed. The 73% figure is the average rate of finding a successful trajectory on these six simulated tasks at a budget of 45 nodes. It is not a success rate for an arbitrary new kitchen.
[1]要点
- SAIL does not update Gemini Robotics-ER 1.5. It searches and revises trajectories in simulation.
- On six simulated tasks with 20 starts each, one node averages 25% and 45 nodes average 73%.
- A LeRobot SO-101 placed a block in a bowl, five of six trials, at a budget of 15 candidates.
- Trajectories are open-loop, and the real test is one task. Astra did not run these experiments.