ExplorerComputational LinguisticsNLP
Research PaperResearchia:202608.18008

Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text

Benjamin Belay

Abstract

A language model's output does not by itself provide verifiable evidence about the internal computation that produced it. We study computational provenance: whether generated text can carry detectable evidence of which causally relevant internal state occurred. We test a bounded form of this idea in two controlled architectures: a modular feed-forward neural network and a transformer-based model. Both architectures are trained on the same arithmetic task with a mandatory pathway through two disc...

Submitted: August 18, 2026Subjects: NLP; Computational Linguistics

Description / Details

A language model's output does not by itself provide verifiable evidence about the internal computation that produced it. We study computational provenance: whether generated text can carry detectable evidence of which causally relevant internal state occurred. We test a bounded form of this idea in two controlled architectures: a modular feed-forward neural network and a transformer-based model. Both architectures are trained on the same arithmetic task with a mandatory pathway through two discrete intermediate states, allowing different internal paths to produce the same answer. We deliberately switch between these paths, authenticate the state actually used, and let that verified state determine a subtle statistical pattern in the generated text that can later be detected. The feed-forward and transformer systems each passed all 128 matched pairs in both their public and separately sealed protected end-to-end evaluations, with the detector recovering the signal associated with the authenticated internal state. The required causal computation also reproduced across five independently trained feed-forward models and three independently trained transformers. In a separate answer-only transformer experiment, our linear probes did not recover a naturally learned intermediate state. These results provide a controlled proof of concept that information about a verified, causally relevant internal state can be preserved in generated text even when the answer is unchanged.


Source: arXiv:2608.16868v1 - http://arxiv.org/abs/2608.16868v1 PDF: https://arxiv.org/pdf/2608.16868v1 Original Link: http://arxiv.org/abs/2608.16868v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Aug 18, 2026
Topic:
Computational Linguistics
Area:
NLP
Comments:
0
Bookmark
Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text | Researchia