ExplorerRoboticsRobotics
Research PaperResearchia:202608.11088

SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation

Jingkai Wang

Abstract

Vision-language-action policies rely on large multimodal backbones to jointly perform perception, language conditioning, and action generation at every control step. Much of this capacity supports open-domain semantics, whereas continuous robot manipulation primarily requires compact representations of observations, actions, and the transitions induced by actions. Pixel-level world models provide another route, but predicting visual details irrelevant to control can be unnecessarily expensive. W...

Submitted: August 11, 2026Subjects: Robotics; Robotics

Description / Details

Vision-language-action policies rely on large multimodal backbones to jointly perform perception, language conditioning, and action generation at every control step. Much of this capacity supports open-domain semantics, whereas continuous robot manipulation primarily requires compact representations of observations, actions, and the transitions induced by actions. Pixel-level world models provide another route, but predicting visual details irrelevant to control can be unnecessarily expensive. We propose SLIM (Self-supervised Latent Interaction Model), a compact 0.5B-parameter latent interaction policy. SLIM learns action-grounded predictive latents that capture both action-conditioned future transitions and the actions that explain observed changes. SLIM learns these representations through self-supervised masked trajectory prediction, combining action reconstruction with future-latent prediction. A compact Mixture-of-Transformers (MoT) backbone models interactions between observation latents and action tokens. The resulting policy is trained with flow matching for language-conditioned action generation. Across simulation benchmarks and real-world evaluation, SLIM matches or exceeds representative large-scale VLA and world-action-model baselines with fewer parameters, no additional embodied pretraining, lower inference latency, and substantially lower GPU memory usage.


Source: arXiv:2608.09771v1 - http://arxiv.org/abs/2608.09771v1 PDF: https://arxiv.org/pdf/2608.09771v1 Original Link: http://arxiv.org/abs/2608.09771v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Aug 11, 2026
Topic:
Robotics
Area:
Robotics
Comments:
0
Bookmark
SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation | Researchia