Signal BATON: Transition-Aware Memory for Long-Horizon Robot Manipulation
Summary
A study addresses two failure modes that arise when LLM agents direct vision-language-action systems on long-horizon robot manipulation tasks. One is that exploration cost grows multiplicatively as the number of stages increases, and the other is that there is no way to represent transitions between subtasks. To address this, the authors present a framework called BATON, which treats the subtask itself as the unit of exploration so cost becomes additive rather than multiplicative, storing each subtask's solution in memory. It pairs this with a transition-aware memory, in which a verifier agent governs when the vision-language-action model is invoked within a subtask, a handoff mechanism restores entry states disturbed by the prior subtask, and a lookahead mechanism picks strategies whose outcomes help the next subtask, all without updating any model parameters. On the long-horizon benchmark RoboMemArena, BATON is reported to raise task success by 11 point 6 percentage points and cumulative success by 14 point 9 percentage points over the prior best-performing method.
Classification
Evidence 1
- Don't Drop the BATON: Long-Horizon Robot Manipulation via Agentic Subtask Exploration and Transition-aware Memory arXiv (cs.AI) 2026-08-17 accessed 2026-08-20T05:08:18+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-755f65ae50d4
