ExplorerData ScienceStatistics
Research PaperResearchia:202608.13031

Coarsening Latent-Class Probabilities: Directional Distortion and Coverage Loss

Marcell T. Kurbucz

Abstract

Outcomes are increasingly regressed on a calibrated probability vector for unobserved class membership, and that vector is often coarsened to a hard label first. Under a constant-coefficient structural mean and conditional calibration, the observed-data problem is a partially linear regression of the outcome on the probability vector; we take this reduction as the starting point and ask what coarsening costs. For any coarsening, the plug-in estimator converges to $\mathcal{A}τ$, where the coarse...

Submitted: August 13, 2026Subjects: Statistics; Data Science

Description / Details

Outcomes are increasingly regressed on a calibrated probability vector for unobserved class membership, and that vector is often coarsened to a hard label first. Under a constant-coefficient structural mean and conditional calibration, the observed-data problem is a partially linear regression of the outcome on the probability vector; we take this reduction as the starting point and ask what coarsening costs. For any coarsening, the plug-in estimator converges to Aτ\mathcal{A}τ, where the coarsening operator satisfies A=I+D1E[ahu]\mathcal{A}=I+D^{-1}\mathbb{E}[a_{h}u^{\top}] with uu the discarded signal. Coarsening is therefore free exactly when what is discarded is uncorrelated with what is kept, and is otherwise anisotropic: it distorts some contrasts far more than others. The same operator governs inference. The Wald interval built from coarsened labels has limiting coverage Φ(zλ)Φ(zλ)Φ(z-λ)-Φ(-z-λ), with λλ the ratio of the coarsening bias to the reported standard error; because A\mathcal{A} and that standard error depend on observables alone, the coverage implied by the estimated index can be approximated before the interval is reported. Simulations show severe coverage loss after argmax coarsening, and three real-data audits exhibit the direction-specific distortion that hard labels induce.


Source: arXiv:2608.11784v1 - http://arxiv.org/abs/2608.11784v1 PDF: https://arxiv.org/pdf/2608.11784v1 Original Link: http://arxiv.org/abs/2608.11784v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Aug 13, 2026
Topic:
Data Science
Area:
Statistics
Comments:
0
Bookmark