ExplorerArtificial IntelligenceAI
Research PaperResearchia:202609.04054

From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research

Yakov Pyotr Shkolnikov

Abstract

Research and news coverage of language-model deception increasingly attributes human-like mental-state concepts to language models. Such claims can blur the distinction between behavior that looks deceptive and a mechanism that is actually deceptive. We introduce a causal taxonomy separating prior commitment from retrospective report, model preference from realized output, false preference from sensitivity to the utility of misleading a recipient, and deceptive behavior from the provenance of ...

Submitted: September 4, 2026Subjects: AI; Artificial Intelligence

Description / Details

Research and news coverage of language-model deception increasingly attributes human-like mental-state concepts to language models. Such claims can blur the distinction between behavior that looks deceptive and a mechanism that is actually deceptive. We introduce a causal taxonomy separating prior commitment from retrospective report, model preference from realized output, false preference from sensitivity to the utility of misleading a recipient, and deceptive behavior from the provenance of the objective or strategy producing it. We test these distinctions in two open-weight model families. Across controlled guessing-game and stock-trading experiments, we find that deceptive-looking behavior can arise without the corresponding proposed mechanism, while other interventions provide direct evidence that recipient information state can causally affect deceptive preference. These results show that deceptive behavior can provide evidence for a deceptive mechanism. But even evidence for such a mechanism does not establish model agency in the deception.


Source: arXiv:2609.04166v1 - http://arxiv.org/abs/2609.04166v1 PDF: https://arxiv.org/pdf/2609.04166v1 Original Link: http://arxiv.org/abs/2609.04166v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Sep 4, 2026
Topic:
Artificial Intelligence
Area:
AI
Comments:
0
Bookmark
From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research | Researchia