ExplorerArtificial IntelligenceAI
Research PaperResearchia:202607.28069

Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects

Phu Gia Hoang

Abstract

The wide-scale use of sparse autoencoders (SAEs) as interpretability tools is limited by inconsistent links between SAE features and model behavior. Features with clear activation descriptions may have weak or unexpected causal effects; steering can vary across prompts or oppose the intended direction; and activation-based feature selection can miss features that produce the desired output change. Prior work has studied feature geometry inside the model, where features are computed. We instead s...

Submitted: July 28, 2026Subjects: AI; Artificial Intelligence

Description / Details

The wide-scale use of sparse autoencoders (SAEs) as interpretability tools is limited by inconsistent links between SAE features and model behavior. Features with clear activation descriptions may have weak or unexpected causal effects; steering can vary across prompts or oppose the intended direction; and activation-based feature selection can miss features that produce the desired output change. Prior work has studied feature geometry inside the model, where features are computed. We instead study the geometry of changes in model logits caused by feature interventions. We introduce Feature-Effect Geometry Analysis (FEGA), an unsupervised framework that removes the same active SAE feature across contexts and analyzes the resulting cloud of logit changes. Across SAE variants, consistent one-dimensional effects are rare: few features behave like reusable directions. To interpret this variation, we distinguish value-like features, tied to static information such as factual attributes, from pointer-like features, associated with context-dependent operations. Value-like features more often exhibit structured, low-dimensional effects, although these effects typically span several directions. Pointer-like features, by contrast, predominantly exhibit diffuse effects. Our results show that a feature can be interpretable and causally relevant without providing a stable direction for steering.


Source: arXiv:2607.24645v1 - http://arxiv.org/abs/2607.24645v1 PDF: https://arxiv.org/pdf/2607.24645v1 Original Link: http://arxiv.org/abs/2607.24645v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Jul 28, 2026
Topic:
Artificial Intelligence
Area:
AI
Comments:
0
Bookmark
Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects | Researchia