One Word, Different Action: A Real-Robot Benchmark for Language-Conditioned Embodied Reasoning
Abstract
Natural-language instruction changes can directly alter robot behavior. A reliable embodied system should preserve its action when the task is unchanged and update it correctly when the task itself changes. We introduce One Word, Different Action, a real-robot benchmark built on physical decision states and executable actions, using task-preserving and task-changing instruction pairs to jointly evaluate Decision Invariance and Decision Sensitivity, with further evaluation under multi-constraint ...
Description / Details
Natural-language instruction changes can directly alter robot behavior. A reliable embodied system should preserve its action when the task is unchanged and update it correctly when the task itself changes. We introduce One Word, Different Action, a real-robot benchmark built on physical decision states and executable actions, using task-preserving and task-changing instruction pairs to jointly evaluate Decision Invariance and Decision Sensitivity, with further evaluation under multi-constraint reasoning and real-RGB grounding. Experiments show that modern models are near saturation on single-constraint instruction changes, yet several models degrade noticeably when multiple task constraints must be integrated into one executable decision. These results suggest that the more salient remaining challenge is no longer recognizing an isolated instruction change, but reliably composing multiple task requirements into a correct robot action decision.
Source: arXiv:2609.05260v1 - http://arxiv.org/abs/2609.05260v1 PDF: https://arxiv.org/pdf/2609.05260v1 Original Link: http://arxiv.org/abs/2609.05260v1
Please sign in to join the discussion.
No comments yet. Be the first to share your thoughts!
Sep 7, 2026
Robotics
Robotics
0