Selective Inference for CART with Binary Outcomes
Abstract
Binary classification trees select subgroups using the same outcomes later used to assess their differences. We develop finite-sample conditional tests of a common success probability within a parent selected by deterministic Gini CART. The construction retains all eligible cutpoints and conditions on the selected split, its ancestor path, the parent success total, and outside outcomes. The resulting uniform label fiber gives an exact count distribution, while a reversible parallel Monte Carlo c...
Description / Details
Binary classification trees select subgroups using the same outcomes later used to assess their differences. We develop finite-sample conditional tests of a common success probability within a parent selected by deterministic Gini CART. The construction retains all eligible cutpoints and conditions on the selected split, its ancestor path, the parent success total, and outside outcomes. The resulting uniform label fiber gives an exact count distribution, while a reversible parallel Monte Carlo construction yields super-uniform inclusive and exactly uniform tie-randomized p-values for any prespecified finite run budget. Within a two-child constant-risk model, the selected count law is an exponential family with information equal to its conditional count variance. We separate information loss from computational limitations and exhibit a selected fiber disconnected under single-label swaps. Simulations with 200 or 400 observations, ten independent or correlated predictors, and trees of depth three show conservative inclusive tests and nontrivial power for large risk differences. At 400 observations and a generating risk difference of 0.4, randomized rejection conditional on reaching the prespecified third-level target is 51--71%, while selection followed by rejection occurs in 11--21% of datasets. Smaller signals remain difficult to detect, and a representative fivefold increase in computation gives little power improvement. The guarantee concerns parent homogeneity, or equality of two constant child risks, and does not cover equality of heterogeneous regional averages.
Source: arXiv:2609.24949v1 - http://arxiv.org/abs/2609.24949v1 PDF: https://arxiv.org/pdf/2609.24949v1 Original Link: http://arxiv.org/abs/2609.24949v1
Please sign in to join the discussion.
No comments yet. Be the first to share your thoughts!
Sep 22, 2026
Data Science
Statistics
0