Interpretability Study Session
Details
Tracing Attention Computation Through Feature Interactions
Jonathan deWerd continues steamrolling through Anthropic's Transformer Circuits Thread (https://transformer-circuits.pub/) from oldest to now. Starting again with a super high level recap of previous work but the main focus will be the material in Tracing Attention.
This week's goal:
* Tracing Attention Computation Through Feature Interactions
Previous sessions:
* Towards Monosemanticity
* Scaling Monosemanticity
* Circuit Tracing
* On the Biology of a Large Language Model (Video overview (1 hour youtube))
Future sessions:
* Emergent Introspective Awareness in Large Language Models
* Emotion Concepts and their Function in a Large Language Model
* Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations
* Verbalizable Representations Form a Global Workspace in Language Models
...and then we can finally take this meetup to J-space.
The setting is informal and open to everyone!
