Explorerโ€บData Scienceโ€บMachine Learning
Research PaperResearchia:202609.09086

Do Reasoning Representations Help Humans Evaluate LLM Outputs?

Jaewoo Lim

Abstract

Reasoning representations are increasingly used as explanations for large language model outputs. Yet they are typically evaluated with model-centric criteria, such as answer accuracy and faithfulness, leaving it unclear whether they help people evaluate model responses. In this work, we study reasoning representations as human-facing interfaces rather than proxies for model reasoning ability. We conduct a controlled human study of six reasoning formats across tasks of varying complexity, suppor...

Submitted: September 9, 2026Subjects: Machine Learning; Data Science

Description / Details

Reasoning representations are increasingly used as explanations for large language model outputs. Yet they are typically evaluated with model-centric criteria, such as answer accuracy and faithfulness, leaving it unclear whether they help people evaluate model responses. In this work, we study reasoning representations as human-facing interfaces rather than proxies for model reasoning ability. We conduct a controlled human study of six reasoning formats across tasks of varying complexity, supported by a web-based framework that randomizes task domains, problem instances, and representation order. The study collects fine-grained judgments of structural understanding, error detection and localization, and trust calibration. Our study shows a mismatch between perceived preference and support for human evaluation. Participants prefer planning- and decomposition-based representations, but simpler chain-of-thought traces better support verification, trust, and interpretability. Preferred representations also introduce calibration risks, with more false alarms on correct traces and high trust despite low willingness to verify.


Source: arXiv:2609.09038v1 - http://arxiv.org/abs/2609.09038v1 PDF: https://arxiv.org/pdf/2609.09038v1 Original Link: http://arxiv.org/abs/2609.09038v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Sep 9, 2026
Topic:
Data Science
Area:
Machine Learning
Comments:
0
Bookmark
Do Reasoning Representations Help Humans Evaluate LLM Outputs? | Researchia