Explorerβ€ΊData Scienceβ€ΊMachine Learning
Research PaperResearchia:202610.08005

Rephrase Before You Act: Characterizing and Mitigating Language Sensitivity in Vision-Language-Action Models

Mikey Watts

Abstract

Vision-language-action models (VLAs) are strikingly sensitive to instruction phrasing and do not inherit the language robustness of the vision-language models they are built on. A one-word edit can move success by tens of points: $Ο€_{0.5}$ turns on a LIBERO stove 100% of the time for "switch on the stove" and 2% for "switch on the hot plate", and a $Ο€_0$ checkpoint finetuned with rephrase augmentation still shows swings of up to 61 points. We characterize this sensitivity with statistically test...

Submitted: October 8, 2026Subjects: Machine Learning; Data Science

Description / Details

Vision-language-action models (VLAs) are strikingly sensitive to instruction phrasing and do not inherit the language robustness of the vision-language models they are built on. A one-word edit can move success by tens of points: Ο€0.5Ο€_{0.5} turns on a LIBERO stove 100% of the time for "switch on the stove" and 2% for "switch on the hot plate", and a Ο€0Ο€_0 checkpoint finetuned with rephrase augmentation still shows swings of up to 61 points. We characterize this sensitivity with statistically tested single-edit swings and an oracle phrase search, which shows that phrasing alone nearly closes the 21-point gap between in-distribution and out-of-distribution tasks. We then reduce it without modifying the policy. Because the sensitivity is systematic, it can be expressed as explicit rules: we score many phrasings of a few training tasks, have a large language model distill the evidence into ten to twenty rephrasing rules, and at deployment rewrite each incoming instruction once under these rules. The rules improve the frozen Ο€0Ο€_0 by 16 to 27% relative on twelve held-out tasks across adversarial, VLM-generated, and human-generated phrasings, with gains concentrated on out-of-distribution tasks. The pipeline replicates on Ο€0.5Ο€_{0.5} and LIBERO, lifting in-finetune success from 93.6% to 97.8%. The method requires no retraining and no per-step verification, and applies zero-shot to unseen tasks and instructions. Project website: https://sttawm.github.io/rephrase-before-you-act


Source: arXiv:2610.10526v1 - http://arxiv.org/abs/2610.10526v1 PDF: https://arxiv.org/pdf/2610.10526v1 Original Link: http://arxiv.org/abs/2610.10526v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Oct 8, 2026
Topic:
Data Science
Area:
Machine Learning
Comments:
0
Bookmark
Rephrase Before You Act: Characterizing and Mitigating Language Sensitivity in Vision-Language-Action Models | Researchia