Explorerโ€บRoboticsโ€บRobotics
Research PaperResearchia:202609.22010

MIGU: Multimodal Instruction Grounding under Uncertainty for Manipulation Planning

Mingke Lu

Abstract

Understanding natural human instructions is crucial for deploying robots in human-centric environments. We study multimodal instruction grounding, where language and gesture provide complementary but uncertain cues. We present MIGU, a modular framework that combines semantic and geometric evidence into a unified grounding belief and connects it to manipulation planning. MIGU constructs a 3D geometric likelihood by propagating viewing-direction and depth uncertainty through eye-finger geometry wh...

Submitted: September 22, 2026Subjects: Robotics; Robotics

Description / Details

Understanding natural human instructions is crucial for deploying robots in human-centric environments. We study multimodal instruction grounding, where language and gesture provide complementary but uncertain cues. We present MIGU, a modular framework that combines semantic and geometric evidence into a unified grounding belief and connects it to manipulation planning. MIGU constructs a 3D geometric likelihood by propagating viewing-direction and depth uncertainty through eye-finger geometry while accounting for hand-direction estimation error. A vision-language model (VLM) provides semantic priors over candidate objects and regions, which are combined with the geometric likelihood through Bayes-inspired fusion. The resulting belief supports behavior planning to either proceed directly to downstream planning or request clarification. Grounded targets then define goals for mobile manipulation and tabletop task-and-motion planning. On a real-world benchmark, MIGU outperforms all evaluated baselines, while ablations support the benefit of explicit multimodal uncertainty modeling. Project website: multimodal-instruction.github.io


Source: arXiv:2609.24995v1 - http://arxiv.org/abs/2609.24995v1 PDF: https://arxiv.org/pdf/2609.24995v1 Original Link: http://arxiv.org/abs/2609.24995v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Sep 22, 2026
Topic:
Robotics
Area:
Robotics
Comments:
0
Bookmark
MIGU: Multimodal Instruction Grounding under Uncertainty for Manipulation Planning | Researchia