High-Dimensional Learning Dynamics of Attention-Indexed Models
Abstract
Attention mechanisms are central to modern foundation models, yet their training dynamics remain poorly understood, especially when the attention matrices have extensive rank. In this work, we study attention-indexed models, a broad framework that can represent multi-layer and multi-head attention architectures. First, we show that, in a suitable high-dimensional limit, the population-loss landscape is characterized by a finite set of trace order parameters. In contrast, online stochastic gradie...
Description / Details
Attention mechanisms are central to modern foundation models, yet their training dynamics remain poorly understood, especially when the attention matrices have extensive rank. In this work, we study attention-indexed models, a broad framework that can represent multi-layer and multi-head attention architectures. First, we show that, in a suitable high-dimensional limit, the population-loss landscape is characterized by a finite set of trace order parameters. In contrast, online stochastic gradient descent (SGD) is governed by an infinite hierarchy of matrix moments, which we show can be exponentially well-approximated by a finite truncated system. Second, this framework reveals that attention parameterization itself can act as an architectural implicit bias. Direct optimization of an attention matrix can remain trapped in an uninformative state. Tied attention () induces an automatic symmetry-breaking mechanism and yields weak recovery in samples. For untied attention, , we uncover a fast-slow mechanism: the pre-activation mean first evolves on a fast timescale, while the overlaps evolve on a slower one. Weak recovery on the scale occurs when the state selected by the fast dynamics breaks the initial symmetry.
Source: arXiv:2609.03858v1 - http://arxiv.org/abs/2609.03858v1 PDF: https://arxiv.org/pdf/2609.03858v1 Original Link: http://arxiv.org/abs/2609.03858v1
Please sign in to join the discussion.
No comments yet. Be the first to share your thoughts!
Sep 4, 2026
Data Science
Statistics
0