Explorerโ€บRoboticsโ€บRobotics
Research PaperResearchia:202607.29075

SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models

Zonghe Liu

Abstract

Vision-Language-Action (VLA) models have shown strong potential for general robot manipulation, but most existing models rely on 2D visual-language backbones and lack fine-grained 3D understanding of target objects, especially under occlusion, pose variation, scale changes, and precise spatial interaction. We propose an object-centric 3D representation alignment framework built upon $ฯ€_0$, using SAM3D as a frozen 3D teacher to provide target-object 3D priors during training. Specifically, we loc...

Submitted: July 29, 2026Subjects: Robotics; Robotics

Description / Details

Vision-Language-Action (VLA) models have shown strong potential for general robot manipulation, but most existing models rely on 2D visual-language backbones and lack fine-grained 3D understanding of target objects, especially under occlusion, pose variation, scale changes, and precise spatial interaction. We propose an object-centric 3D representation alignment framework built upon ฯ€0ฯ€_0, using SAM3D as a frozen 3D teacher to provide target-object 3D priors during training. Specifically, we localize task-relevant objects with object recognition models, generate corresponding object masks, and use SAM3D to extract dense object-level 3D representations, which are aligned with intermediate visual features of ฯ€0ฯ€_0. This enables the policy to internalize target-object 3D information while preserving the original RGB-language-to-action inference pipeline without requiring depth, point clouds, masks, SAM3D, or additional 3D modules at test time. Simulation experiments show consistent improvements, achieving 99.1% on LIBERO and an average length of 4.11 on CALVIN. Real-world experiments further demonstrate that our method is particularly effective in long-horizon manipulation scenarios where the robot must focus on different target objects across multiple subtasks.


Source: arXiv:2607.25912v1 - http://arxiv.org/abs/2607.25912v1 PDF: https://arxiv.org/pdf/2607.25912v1 Original Link: http://arxiv.org/abs/2607.25912v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Jul 29, 2026
Topic:
Robotics
Area:
Robotics
Comments:
0
Bookmark
SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models | Researchia