SAIL: Scaling In-Context Imitation Learning
Sakana AI
Sakana AI says it found a way to make robot plans more reliable without retraining the model. It lets the model test its own moves in simulation and revise them before the real robot acts.
Based on reporting by Sakana AI — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Sakana AI and the University of Tokyo are set to present SAIL, short for Scaling In-Context Imitation Learning, at IROS 2026. The pitch is simple: instead of teaching a robot by changing the model, ask a foundation model to use what it already knows, then pressure-test its plan before anything touches hardware.
That matters because a robot plan can look fine on first try and still collapse on a tiny mistake. The source says earlier work has shown large language and vision models can produce whole action sequences from a few demonstrations, but those models do not always land on a usable trajectory in one shot. Context helps, and the wrong movement target can break the task.
SAIL tries to fix that with test-time scaling. A policy VLM generates a candidate trajectory from a few successful demonstrations. The trajectory gets checked in a simulator, then an evaluation VLM watches the resulting video and points out where progress got stuck. The policy model uses that feedback to revise the plan, while Monte Carlo tree search explores alternatives and keeps refining the better candidates. Only the chosen trajectory is passed to the physical robot.
The results are strongest in simulation. Across six manipulation tasks, raising the search budget from one candidate to 45 increased the average success rate for finding a workable trajectory from 25% to 73%. Sakana AI also tested SAIL on a physical robot, and says the results suggest extra computation can help a model test and polish its own actions before execution.
There’s an obvious subtext here: maybe the next step for robotics is not bigger training runs, but better ways to interrogate the model at runtime. That’s a more boring story than magic general intelligence, which is usually how you know it might actually work.
My take — AI-written commentary, not fact-checked reporting
This is the right kind of robotics progress: less grandstanding, more trying the plan twice before the arm smashes into the table. Test-time compute is becoming the adult in the room for AI, especially where a small error has a physical bill attached. The industry loves pretending every model failure needs a bigger model; sometimes it just needs a simulator and a bit of humility.
Read more about this at: Sakana AI