Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
Abstract
We investigate whether natural-language documentation helps coding agents resolve software issues, and we build the tools to construct and evaluate it. We introduce a roundtrip benchmark that scores code descriptions by whether code regenerated from them passes the original tests, and show that completeness, not length, drives a description's fidelity. Using the benchmark as an optimization signal, we discover a description-writing prompt that reaches full fidelity and generalizes to unseen file...
Description / Details
We investigate whether natural-language documentation helps coding agents resolve software issues, and we build the tools to construct and evaluate it. We introduce a roundtrip benchmark that scores code descriptions by whether code regenerated from them passes the original tests, and show that completeness, not length, drives a description's fidelity. Using the benchmark as an optimization signal, we discover a description-writing prompt that reaches full fidelity and generalizes to unseen files. We then test the hypothesis that motivated the work: that better documentation helps an agent resolve real repository issues. Across two model families and ten repositories, and against a positive control confirming that our evaluation can detect a genuine improvement, we find that it does not. When the source is present, neither static compact documentation nor retrieved context beats the issue alone. We report this negative result together with the benchmark and the optimizer, and we characterize the boundary at which documentation helps.
Source: arXiv:2609.31587v1 - http://arxiv.org/abs/2609.31587v1 PDF: https://arxiv.org/pdf/2609.31587v1 Original Link: http://arxiv.org/abs/2609.31587v1
Please sign in to join the discussion.
No comments yet. Be the first to share your thoughts!
Sep 28, 2026
Artificial Intelligence
AI
0