ExplorerData ScienceMachine Learning
Research PaperResearchia:202608.05080

Muon Meets Mamba: Spectral Optimization for State Space Models

Arslan Battalov

Abstract

Muon is a recent optimizer that orthogonalizes the update to each weight matrix with a Newton-Schulz iteration, which performs steepest descent under the spectral norm. Almost all the evidence for it comes from Transformer models, and its behavior on state-space models is largely unreported. We compare Muon with AdamW on Mamba-2 130M under a controlled protocol that varies only which weight groups are trained with Muon. The benefit is localized. Muon on the output projection alone beats Muon on ...

Submitted: August 5, 2026Subjects: Machine Learning; Data Science

Description / Details

Muon is a recent optimizer that orthogonalizes the update to each weight matrix with a Newton-Schulz iteration, which performs steepest descent under the spectral norm. Almost all the evidence for it comes from Transformer models, and its behavior on state-space models is largely unreported. We compare Muon with AdamW on Mamba-2 130M under a controlled protocol that varies only which weight groups are trained with Muon. The benefit is localized. Muon on the output projection alone beats Muon on the input projection or on both. The advantage is mainly one of token efficiency. It holds on two corpora and two token budgets, and persists when training continues well past the compute-optimal point. Conditioning does not explain the gain. Muon lowers the condition number of whichever projection it trains, but the better-conditioned input projection is not the one that helps.


Source: arXiv:2608.03941v1 - http://arxiv.org/abs/2608.03941v1 PDF: https://arxiv.org/pdf/2608.03941v1 Original Link: http://arxiv.org/abs/2608.03941v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Aug 5, 2026
Topic:
Data Science
Area:
Machine Learning
Comments:
0
Bookmark