An Entropy-based Coefficient of Determination with Adjustment of Optimization Bias
Abstract
Classical likelihood-ratio tests and $Ξ$AIC exacerbate the statistical significance crisis by scaling with sample size, often flagging negligible improvements as highly significant. While causal estimands like the average treatment effect (ATE) quantify practical magnitude, their reliance on the expectation operator ties them to the data's original coordinate scale. Furthermore, existing pseudo-$R^2$ metrics are inadequate: variance-based measures ignore higher-order distributional changes, and ...
Description / Details
Classical likelihood-ratio tests and AIC exacerbate the statistical significance crisis by scaling with sample size, often flagging negligible improvements as highly significant. While causal estimands like the average treatment effect (ATE) quantify practical magnitude, their reliance on the expectation operator ties them to the data's original coordinate scale. Furthermore, existing pseudo- metrics are inadequate: variance-based measures ignore higher-order distributional changes, and current formulations lack invariance to monotone transformations. We resolve these limitations by introducing Entropic Variance (EV) as a rigorous, scale-independent generalization of error variance in ordinary least squares. We define the population EV-based parameter, , which projects unbounded cross-entropy onto a standardized scale, and establish that the EV-based statistic asymptotically follows an -distribution. Building on these distributional properties, we propose two estimators: the empirical population and the out-of-sample predictive . Both are derived by exponentiating per-observation cross-entropy and incorporate a degrees-of-freedom correction for training optimism. Leveraging the -distribution, we derive refined -values and confidence intervals for without requiring intractable Fisher information matrices. Simulation studies and a Parkinson's disease microbiome application demonstrate the superiority of variable selection via these EV- metrics. Notably, evaluating the of a LASSO path via data-splitting reduced false discovery rates from 80% to 6% in simulations while fully preserving signal recall.
Source: arXiv:2608.06624v1 - http://arxiv.org/abs/2608.06624v1 PDF: https://arxiv.org/pdf/2608.06624v1 Original Link: http://arxiv.org/abs/2608.06624v1
Please sign in to join the discussion.
No comments yet. Be the first to share your thoughts!
Aug 10, 2026
Biotechnology
Biology
0