ExplorerData ScienceStatistics
Research PaperResearchia:202608.31032

Learning between the peaks: sharp asymptotics for kernel ridge regression under power-law anisotropy

Lorenzo Rizzi

Abstract

We study kernel ridge regression under anisotropic Gaussian data, where the input covariance decays as a power law with exponent $α\geq 0$ for polynomial inner-product kernels. We derive asymptotically sharp expressions for the kernel spectrum and the generalization error in the polynomial high-dimensional regime $n=Θ(d^κ)$, revealing how anisotropy reshapes the learning curves. For weak anisotropy ($0<α<1$), the problem remains effectively high-dimensional and retains some features of the isotr...

Submitted: August 31, 2026Subjects: Statistics; Data Science

Description / Details

We study kernel ridge regression under anisotropic Gaussian data, where the input covariance decays as a power law with exponent α0α\geq 0 for polynomial inner-product kernels. We derive asymptotically sharp expressions for the kernel spectrum and the generalization error in the polynomial high-dimensional regime n=Θ(dκ)n=Θ(d^κ), revealing how anisotropy reshapes the learning curves. For weak anisotropy (0<α<10<α<1), the problem remains effectively high-dimensional and retains some features of the isotropic case, while departing from it in others: the variance still peaks at integer sample complexities κNκ\in\mathbb{N}, but these peaks are progressively damped as αα grows; meanwhile, for targets strongly aligned with the data's principal directions, the bias drops at fractional sample complexities, decoupling the bias transitions from the interpolation peaks. For strong anisotropy (α>1α> 1), the effective dimension of the problem is constant, and the variance stops depending on sample size altogether, plateauing under ridgeless interpolation or vanishing at an explicit rate under fixed ridge penalty. The bias undergoes a sharp transition governed by the target's decay rate: below a threshold, learning is abrupt rather than gradual; above it, the bias decays as a power law that recovers the classical source and capacity rates. We finally specialize these results to single-index targets, showing how the alignment of the index with the data's principal directions determines the effect of anisotropy on learning. Together, our results clarify how the input geometry shapes the kernel features and fundamentally impacts its generalization properties.


Source: arXiv:2608.28564v1 - http://arxiv.org/abs/2608.28564v1 PDF: https://arxiv.org/pdf/2608.28564v1 Original Link: http://arxiv.org/abs/2608.28564v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Aug 31, 2026
Topic:
Data Science
Area:
Statistics
Comments:
0
Bookmark