jh

Home | Projects | Blog | Quotes | Connect

Scaling Monosemanticity

In October 2023, Anthropic’s interpretability team demonstrated that dictionary learning with sparse autoencoders (SAEs) can find interpretable features in a toy, one-layer transformer language model (LM) in Towards Monosemanticity: Decomposing Language Models With Dictionary Learning. Now, they have scaled up this approach to an actual, production-grade language model in Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet.

Setup

A sparse autoencoder is a component that allows the neural network to expand its representation by giving it an intermediate layer with more neurons. This is the inverse of what happens with regular autoencoders, which force the network to compress its representation by shrinking the number of neurons in the intermediate layer.

Definitions:

Like regular autoencoders, SAEs are trained to minimize a) reconstruction error, i.e. the activations after passing the autoencoder should match the original activation, and b) the feature activations. This forces the model to use as few features as possible for the activation reconstruction, nudging it to use features representing individual concepts (monosemanticity).

This additional layer is placed in the middle layer of the network on a residual stream, as opposed to placement on an MLP block in Towards Monosemanticity, and the new weights are trained with the explained loss on a dataset, which is from a similar distribution as the text from the pretraining dataset.

SAE placement in a Transformer.

They train 3 SAEs of three sizes: ~1M, ~4M, and ~33M features. Scaling laws apply, saving a bunch of computation during training.

After the training, the reconstructed activations explain a 65% of model activations variance. Features often fire together, with an average of less than 300 active features per token (in all sizes). Some features end up “dead”, i.e. ~never active during inference — up to 65% for the biggest version.

Results

So, how interpretable and clean were the found features?

Unfortunately, there is no objective metric to answer this. In the paper, they define two assessment methods, specificity and influence on behaviour, and demonstrate them on four example features.

Specificity: does the feature fire2 only on inputs containing the concept it should represent? For example, the "The Golden Gate Bridge" feature should have a high value on input related to the Golden Gate Bridge and a low value on unrelated inputs.

In the studied examples, the large activations indeed correspond to the hypothesized concepts behind the features. Moreover, this transfers to images as well, despite the SAEs trained only on text3.

Influence on behaviour: if the feature is increased or decreased4, does this influence the generated way in expected ways? So increasing "The Golden Gate Bridge" should ... make the model talk about the bridge a lot?

The authors did not observe any feature where the resulting behaviour diverged from what they expected based on activating inputs5.

Like in Towards Monosemanticity, the SAE features are more interpretable than the original neurons. This is demonstrated with an “automated interpretability pipeline” from Language models can explain neurons in language models, where (1) LM generates an explanation of neurons based on their activation given different inputs, (2) another LM  instance predicts the neuron's activations on a new text, given the generated explanation in the previous step, and (3) the predicted activations are compared to the actual activations.

We can also measure the distance between the feature vectors based on cosine similarity and inspect related clusters, even between different SAEs. For example, the Immunology feature is close to features related to common diseases. Close by is also a "Vaccines and immunizations" cluster, and a little bit further are legal and social immunity concepts. These visualizations also show how the features from the three sizes of SAEs are related, where a feature from 1M often splits into multiple features in bigger models. Sometimes, the bigger models have features with no close counterparts from the 1M model.

Finally, they present a handful of safety-related features, such as “Backdoor”, “Secrecy or discreetness”, or “Treacherous turns”. Once again, the downstream effect of increasing these activations has expected results: increase the “Unsafe Code” feature activation, and the model completes the code with buffer overflow; decrease the “AI Assistant” feature activation, and the model responds with “I am a person who is here to help you.” instead of “I am an artificial intelligence created by Anthropic…”.

Comments

This was an interesting and fun read!

My main takeaway is that perhaps we can scale up our interpretability tools to the largest models. This is not a one-layer transformer nor a GPT-2. We even got interpretability scaling laws!

The transfer to images was cool, though probably not entirely unexpected.

I was more intrigued by the claim that the authors always observed the expected influence on behaviour with feature activation modifications. That looks like a strong claim, but it would be great if it were true!

I agree with the authors that it’s not surprising to find the safety features present.

  1. So it has the same size for differently-sized SAEs. ↩

  2. The activations are calculated after each token; that’s why multiple tokens ina a given prompt have different activations. If the paper mentions only a single activation given a prompt, I believe it’s the activation after the last token. ↩

  3. The underlying model, Claude 3 Sonnet, was trained on images as well. ↩

  4. Modifying the features to produce the desired changes has to be quite severe: from -10x to 10x the observed maximum value of the activation during training. The authors hypothesize that it is because they modify only one feature, which often co-fires with others. But if it is changed even more, say 100x instead of 10x, the model starts producing non-sensical values, like repeating the same token. ↩

  5. “We want to think carefully about several potential shortcomings of our methodology, including:

    • Illusions from suboptimal dictionary learning, […]

    • Cases where the downstream effects of features diverge from what we might expect given their activation patterns.

    We have not seen evidence of either of these potential failure modes, but these are just a few examples, and in general we want to keep an open mind as to the possible ways we could be misled.” (emphasis mine) ↩