ExplorerRoboticsRobotics
Research PaperResearchia:202609.01086

SleepWalking: Privileged Representation Shaping for End-to-End Blind Locomotion in Legged Robots

Zheng Pan

Abstract

Partially observable locomotion requires a policy to act when task-relevant properties of the robot--environment state are not fully specified by instantaneous observations. Existing approaches often address this challenge by explicitly estimating missing physical variables or processing extended observation histories through structured architectures. We take a different view: partial observability is fundamentally an information-retention problem. The decisive question is not how task-relevant ...

Submitted: September 1, 2026Subjects: Robotics; Robotics

Description / Details

Partially observable locomotion requires a policy to act when task-relevant properties of the robot--environment state are not fully specified by instantaneous observations. Existing approaches often address this challenge by explicitly estimating missing physical variables or processing extended observation histories through structured architectures. We take a different view: partial observability is fundamentally an information-retention problem. The decisive question is not how task-relevant information enters the network, but whether the policy's internal state retains it. Guided by this perspective, we propose SleepWalking for Robot Locomotion (SWAQ), a one-stage end-to-end framework that uses next-step privileged physical reconstruction to shape what a recurrent history representation retains during policy learning, while the deployed actor uses only a direct history-to-action pathway. Under aligned training settings, SWAQ achieves a 15.0% higher peak mean terrain level than DWAQ, the strongest non-exteroceptive baseline, while using 44.4% fewer inference MACs per control step. Layerwise probes further show that information associated with the reconstructed physical variables remains linearly decodable through the policy head up to the layer preceding the action output. Complementary theoretical analysis relates privileged-variable recoverability to the achievable-return gap between history-based and privileged-information policy classes. These results suggest that semantic objectives can structure learning without requiring a corresponding architectural decomposition of the deployed controller.


Source: arXiv:2608.30883v1 - http://arxiv.org/abs/2608.30883v1 PDF: https://arxiv.org/pdf/2608.30883v1 Original Link: http://arxiv.org/abs/2608.30883v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Sep 1, 2026
Topic:
Robotics
Area:
Robotics
Comments:
0
Bookmark