ExplorerArtificial IntelligenceAI
Research PaperResearchia:202607.29002

$π\mathbf{R}^2$: Reactive Real-time Flow Policies

Sungjae Park

Abstract

Generalist manipulation policies increasingly take the form of action-chunking flow policies built on large pretrained backbones. Such chunks run open-loop, so the policy cannot react to sensory input arriving mid-execution, sacrificing \emph{reactivity}. Replanning more often would restore it, but the perception-to-action pipeline (a large backbone plus multiple denoising steps) is too slow: this \emph{latency} forbids frequent replanning and leaves committed actions stale, making such policies...

Submitted: July 29, 2026Subjects: AI; Artificial Intelligence

Description / Details

Generalist manipulation policies increasingly take the form of action-chunking flow policies built on large pretrained backbones. Such chunks run open-loop, so the policy cannot react to sensory input arriving mid-execution, sacrificing \emph{reactivity}. Replanning more often would restore it, but the perception-to-action pipeline (a large backbone plus multiple denoising steps) is too slow: this \emph{latency} forbids frequent replanning and leaves committed actions stale, making such policies ill-suited for dynamic, closed-loop control. We present πR2π\mathbf{R}^2, which makes these policies reactive and real-time while retaining large backbones, expressive multi-modal policies, and multi-action prediction. Built on the per-position noise schedule of diffusion forcing, πR2π\mathbf{R}^2 contributes two ideas. First, it splits conditioning into a fast channel (proprioception, fresh every tick) and an asynchronously updated slow channel (vision-language features), so the policy reacts to proprioception within a chunk while tolerating stale vision. Second, a latency-adaptive flow schedule treats in-flight actions as inpainting conditioning and emits actions in one denoising step per call, letting one trained model adapt to varying hardware latency. Requiring minimal modification to existing architectures, πR2π\mathbf{R}^2 can be finetuned from a pretrained policy: applied to GR00T-N1.7 on a real xArm6+XHand platform, it replans closed-loop roughly 4×4\times faster than the base policy (~2525Hz on an A5000 GPU), acting on a fresh observation every 4040ms. Across simulation and real-world manipulation tasks, πR2π\mathbf{R}^2 improves the success rate by up to 23%23\% in simulation and 30%30\% in the real world over the strongest baseline. Project page: https://pi-r2-flow.github.io/


Source: arXiv:2607.26055v1 - http://arxiv.org/abs/2607.26055v1 PDF: https://arxiv.org/pdf/2607.26055v1 Original Link: http://arxiv.org/abs/2607.26055v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Jul 29, 2026
Topic:
Artificial Intelligence
Area:
AI
Comments:
0
Bookmark