Explorerβ€ΊArtificial Intelligenceβ€ΊAI
Research PaperResearchia:202608.27051

How Much Rank Does LoRA Need? Rank-Error Bounds for Transformer Attention

Gerard Conangla Planes

Abstract

Choosing the rank of a low-rank adaptation (LoRA) update is usually an empirical task. In this paper, we provide a task-dependent theory of the approximation error achievable at each LoRA rank for Transformer attention. We fix a pretrained attention head, a target attention function, and a distribution over inputs from the downstream task, and bound the smallest expected Kullback--Leibler (KL) error achievable by a rank-$r$ query LoRA update. When target attention probabilities are bounded away ...

Submitted: August 27, 2026Subjects: AI; Artificial Intelligence

Description / Details

Choosing the rank of a low-rank adaptation (LoRA) update is usually an empirical task. In this paper, we provide a task-dependent theory of the approximation error achievable at each LoRA rank for Transformer attention. We fix a pretrained attention head, a target attention function, and a distribution over inputs from the downstream task, and bound the smallest expected Kullback--Leibler (KL) error achievable by a rank-rr query LoRA update. When target attention probabilities are bounded away from zero, we prove a lower bound of the error proportional to ψ(βˆ₯dβˆ₯2)ψ(\|d\|_2), where dd is the difference between candidate and target attention scores and ψ(t)=min⁑{t2,t}ψ(t)=\min\{t^2,t\}. We also prove an unconditional upper bound min⁑{βˆ₯dβˆ₯22/4,2βˆ₯dβˆ₯2}\min\{\|d\|_2^2/4,\sqrt2\|d\|_2\}. Under explicit realizability, geometry, and moment conditions, we then bound the best rank-rr error between an explicit multiple of ψ(Tr)ψ(\sqrt{T_r}) and min⁑{Tr/4,2Tr}\min\{T_r/4,\sqrt{2T_r}\}, where TrT_r is the downstream-weighted tail energy of the target update. We also provide target-Fisher bounds when candidate scores remain within a fixed range of the target scores, and an unrestricted lower bound when a subset of tokens carries most of the probability mass. These spectral bounds describe finite-score approximation. We then construct explicit families in which softmax saturation makes the rank required to match the attention function strictly smaller than the rank required to match the finite logits. Finally, we extend the analysis to fused multi-head LoRA and joint query/key updates, exposing the effects of rank sharing and query/key factorization constraints.


Source: arXiv:2608.26052v1 - http://arxiv.org/abs/2608.26052v1 PDF: https://arxiv.org/pdf/2608.26052v1 Original Link: http://arxiv.org/abs/2608.26052v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Aug 27, 2026
Topic:
Artificial Intelligence
Area:
AI
Comments:
0
Bookmark