ExplorerComputational LinguisticsNLP
Research PaperResearchia:202608.18010

Model Hypnosis: Strong control of AI via additive subliminal effects

Enric Boix-Adsera

Abstract

We demonstrate that AI models are broadly susceptible to a phenomenon we call model hypnosis, in which individually weak and seemingly irrelevant cues in the prompt can be systematically combined to strongly control model behavior. Model hypnosis occurs across model families and scales, including in frontier reasoning models, and hypnotic prompts can transfer between models. Because the model is controlled by inconspicuous textual choices, such as paraphrases and typos, model hypnosis presents n...

Submitted: August 18, 2026Subjects: NLP; Computational Linguistics

Description / Details

We demonstrate that AI models are broadly susceptible to a phenomenon we call model hypnosis, in which individually weak and seemingly irrelevant cues in the prompt can be systematically combined to strongly control model behavior. Model hypnosis occurs across model families and scales, including in frontier reasoning models, and hypnotic prompts can transfer between models. Because the model is controlled by inconspicuous textual choices, such as paraphrases and typos, model hypnosis presents new challenges and avenues for AI safety, and is a major hurdle for AI interpretability.


Source: arXiv:2608.16834v1 - http://arxiv.org/abs/2608.16834v1 PDF: https://arxiv.org/pdf/2608.16834v1 Original Link: http://arxiv.org/abs/2608.16834v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Aug 18, 2026
Topic:
Computational Linguistics
Area:
NLP
Comments:
0
Bookmark
Model Hypnosis: Strong control of AI via additive subliminal effects | Researchia