ExplorerRoboticsRobotics
Research PaperResearchia:202608.06086

Toward Integrating Adaptive Experience Replay and Online Uncertainty Estimation in Safe Actor-Critic Optimal Control

Mahshad Rastegarmoghaddam

Abstract

Safe actor-critic control often treats barrier filtering, uncertainty estimation, and experience replay as separate modules, even though each changes the data used for learning and control. We develop an integrated architecture in which the uncertainty estimate updates the obstacle geometry used by a control barrier function, filter interventions and estimation residuals determine replay priority, and the critic learns from the executed rather than nominal action. We instantiate the architecture...

Submitted: August 6, 2026Subjects: Robotics; Robotics

Description / Details

Safe actor-critic control often treats barrier filtering, uncertainty estimation, and experience replay as separate modules, even though each changes the data used for learning and control. We develop an integrated architecture in which the uncertainty estimate updates the obstacle geometry used by a control barrier function, filter interventions and estimation residuals determine replay priority, and the critic learns from the executed rather than nominal action. We instantiate the architecture on a two-dimensional robot-navigation task with corrupted obstacle measurements and compare six component-matched configurations under common training budgets, random seeds, sensor streams, exploration, and disturbances. Evaluation includes a moderate post-training test, an eleven-level perception-noise sweep, and an exploratory extreme-stress test at multiplier 6.06.0. In the extreme test, the integrated configuration recorded no contacts and reached the goal in all five evaluation seeds. Its mean cost was 7.63±0.447.63\pm0.44 and its obstacle-belief root-mean-square error was 3.52±0.553.52\pm0.55 cm. The uncertainty-estimation ablation also recorded no contacts but reached the goal in four of five seeds, with mean cost 8.96±2.088.96\pm2.08 and belief error 11.08±1.2311.08\pm1.23 cm. A finite-training bound clarifies replay exposure, and a robust barrier condition states the required estimation-error and feasibility assumptions. The results support coupling estimation, safety filtering, and replay on this benchmark; broader safety and convergence claims require further study.


Source: arXiv:2608.04732v1 - http://arxiv.org/abs/2608.04732v1 PDF: https://arxiv.org/pdf/2608.04732v1 Original Link: http://arxiv.org/abs/2608.04732v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Aug 6, 2026
Topic:
Robotics
Area:
Robotics
Comments:
0
Bookmark
Toward Integrating Adaptive Experience Replay and Online Uncertainty Estimation in Safe Actor-Critic Optimal Control | Researchia