When Robots Mishear Us: Mapping the Safety Risks of Voice-Controlled Embodied AI
Abstract
We investigate whether automatic speech recognition (ASR) errors in user input can lead to unsafe outputs from Embodied AI (EAI) models. We find that ASR errors can lead to harmful instructions being accepted and executed by EAI models, thereby reducing safety. We simulate ASR errors and combine them with existing safety benchmarks (SafeAgentBench and POEX) to evaluate how different errors affect embodied AI safety. We find that some of them preserve semantic structure but increase harmful ambig...
Description / Details
We investigate whether automatic speech recognition (ASR) errors in user input can lead to unsafe outputs from Embodied AI (EAI) models. We find that ASR errors can lead to harmful instructions being accepted and executed by EAI models, thereby reducing safety. We simulate ASR errors and combine them with existing safety benchmarks (SafeAgentBench and POEX) to evaluate how different errors affect embodied AI safety. We find that some of them preserve semantic structure but increase harmful ambiguity, while others weaken the model refusal behaviour and allow unsafe plans to be generated and executed. We show that in some cases automatic correction of ASR errors can reduce the risk, but this is not always effective. Overall, we show that ASR errors lead to significant safety risks for embodied AI.
Source: arXiv:2608.28518v1 - http://arxiv.org/abs/2608.28518v1 PDF: https://arxiv.org/pdf/2608.28518v1 Original Link: http://arxiv.org/abs/2608.28518v1
Please sign in to join the discussion.
No comments yet. Be the first to share your thoughts!
Aug 31, 2026
Computational Linguistics
NLP
0