LESSER: Post-Training Data Selection with Output-Layer Gradients
Abstract
The choice of post-training data for large language models substantially affects downstream performance. Gradient-based data selection is a popular approach that ranks training data by how well their gradients align with those of a small validation set. However, ranking with full-parameter gradients requires an expensive backward pass on every sample, making computation intractable for large candidate pools. This raises a natural question: can we approximate full-gradient features at a fraction ...
Description / Details
The choice of post-training data for large language models substantially affects downstream performance. Gradient-based data selection is a popular approach that ranks training data by how well their gradients align with those of a small validation set. However, ranking with full-parameter gradients requires an expensive backward pass on every sample, making computation intractable for large candidate pools. This raises a natural question: can we approximate full-gradient features at a fraction of the cost? Conveniently, we find that output-layer gradients suffice for effective data selection, yet require only the cheaper forward pass. We implement this as LESSER, a drop-in wrapper for selection methods that reduces the feature-extraction FLOP cost by for SFT and for RL benchmarks, while tracking full-gradient performance on downstream tasks. Empirically, we find that even when output-layer and full gradients rank individual samples differently, they select batches with aligned gradients.
Source: arXiv:2610.03702v1 - http://arxiv.org/abs/2610.03702v1 PDF: https://arxiv.org/pdf/2610.03702v1 Original Link: http://arxiv.org/abs/2610.03702v1
Please sign in to join the discussion.
No comments yet. Be the first to share your thoughts!
Oct 5, 2026
Data Science
Machine Learning
0