ExplorerData ScienceMachine Learning
Research PaperResearchia:202608.12004

Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders

Nikolai Bolik

Abstract

Shani et al. (2026) show that LLM representations broadly recover human category boundaries, while failing to reflect fine-grained typicality structure. Their analysis uses cosine similarity over dense model representations. We revisit their approach using overlap over active sparse autoencoder (SAE) latent sets as a more interpretable similarity measure. We first verify that this set-level measure is meaningful: SAE latent sets can recover union-like compositional structure in controlled toy mo...

Submitted: August 12, 2026Subjects: Machine Learning; Data Science

Description / Details

Shani et al. (2026) show that LLM representations broadly recover human category boundaries, while failing to reflect fine-grained typicality structure. Their analysis uses cosine similarity over dense model representations. We revisit their approach using overlap over active sparse autoencoder (SAE) latent sets as a more interpretable similarity measure. We first verify that this set-level measure is meaningful: SAE latent sets can recover union-like compositional structure in controlled toy models and induce semantically coherent neighborhoods in natural text. Extending the human-concepts analysis to SAE set similarities, we find that SAE activation sets do not recover human category boundaries or within-category typicality more faithfully than dense embeddings or residual-stream states, but instead track model-internal similarity structure. To probe this gap further, we study active latent sets under well-controlled semantic modifications, revealing a substantial mismatch between human judgements of conceptual change and change in the SAE active set. We interpret this as evidence that, outside idealised settings, SAE features do not compose via simple bag-of-features semantics.


Source: arXiv:2608.11197v1 - http://arxiv.org/abs/2608.11197v1 PDF: https://arxiv.org/pdf/2608.11197v1 Original Link: http://arxiv.org/abs/2608.11197v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Aug 12, 2026
Topic:
Data Science
Area:
Machine Learning
Comments:
0
Bookmark
Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders | Researchia