Study Reveals Sparse Autoencoders Learn Reproducible Subspaces Despite Unstable Individual Features
Researchers have found that sparse autoencoders (SAEs), tools used to interpret neural network representations, produce features that often fail to reproduce consistently across different training runs. However, while individual unstable features are non-reproducible, they tend to cluster within reproducible lower-dimensional subspaces, suggesting the instability reflects basis ambiguity rather than pure noise. The findings have implications for the reliability of mechanistic interpretability methods that depend on SAE-learned features.
A new study posted to arXiv investigates 'feature stability' in sparse autoencoders — a class of models widely used in mechanistic interpretability research to decompose neural network activations into human-interpretable features. The researchers developed a per-feature metric estimating the probability that a given SAE feature reappears when the model is retrained from a different random seed, enabling large-scale comparisons across seeds, model types, layers, dictionary sizes, and SAE variants. Their central finding is a functional asymmetry: stable features carry the bulk of reconstruction and prediction-relevant signal, while unstable features have weak marginal impact and appear driven by low-frequency surface-form patterns. Geometrically, unstable features are individually non-reproducible but concentrate within reproducible lower-rank subspaces, indicating that different training runs are carving up the same region of activation space in different but equally valid ways. A controlled synthetic experiment confirmed this mechanism, showing that low-rank ground-truth structure can be recovered at the subspace level even when individual latents remain non-identifiable across seeds. The authors also demonstrate that pooling features across multiple training runs can yield more stable SAEs without sacrificing explained variance. The work suggests that instability in SAE features should not be dismissed as noise, but understood as a structural ambiguity with meaningful geometric underpinnings.
What's missing
The study does not address how subspace-level reproducibility translates into practical interpretability workflows — it remains unclear whether human-readable explanations derived from pooled or subspace-aligned features are qualitatively more informative than those from standard SAEs. The paper also does not evaluate whether the proposed cross-seed pooling method scales efficiently to very large language models or production-scale SAE dictionaries.
What different sources said
- arXiv cs.AICenter
Unstable Features, Reproducible Subspaces: Understanding Seed Dependence in Sparse Autoencoders
Related
Gut Bacteria Enzyme Found to Break Down Heat-Processed Food Compounds, Producing Novel Biogenic Amines
Researchers have discovered that an enzyme in common gut bacteria can degrade N-epsilon-carboxymethyllysine (CML), a compound formed during thermal food processing, producing previously unknown biogenic amines. The enzyme, ornithine decarboxylase SpeC from enterobacteria, acts on CML and related modified lysine derivatives through a low-level 'underground' catalytic activity. This finding suggests a previously unrecognized communication axis between thermally processed dietary compounds and gut microbial physiology, with potential implications for host health.
Full-Length Gene Sequencing Reveals Two Distinct Bacterial Communities in Black-Legged Ticks Expanding Into Canada
Researchers used Oxford Nanopore full-length 16S rRNA gene sequencing to characterize the microbiome of Ixodes scapularis black-legged ticks collected in Nova Scotia, Canada, distinguishing between tick-adapted bacteria and environmentally acquired bacteria. The study comes as I. scapularis — the primary vector of Lyme disease — is rapidly expanding northward into Canada due to climate change. The findings suggest that environmentally derived bacteria in tick microbiomes are not mere contamination, which has implications for how tick microbiome data is collected and interpreted across surveillance studies.
Study Identifies Metabolic Link Between Cell Envelope Stress and Biofilm Formation in Bacteria
Researchers have discovered that the metabolite acetyl-CoA directly inhibits enzymes that degrade the bacterial signaling molecule c-di-GMP, connecting cell envelope biosynthesis stress to biofilm formation in Pseudomonas aeruginosa. The study found that sub-inhibitory concentrations of antibiotics targeting early peptidoglycan biosynthesis — but not other antibiotic classes — elevate c-di-GMP levels by reducing phosphodiesterase activity, with acetyl-CoA competing for the enzyme active site. Because the relevant enzyme domain is broadly conserved across bacterial species, this checkpoint mechanism may be widespread and could have implications for understanding antibiotic-induced biofilm responses.