LLM Interpretability and Evaluation

Move from a model view to compatible captures, interventions and bounded comparisons.

Read the map first

Start with the model map article to separate architecture, observed activations and causal evidence. Activation mapping then explains token occurrence identity and compatible projections.

Check what a result establishes

Arithmetic across notations requires baseline competence before circuit analysis. Sparse feature recovery separates reconstruction from recovery of known generating features. Glow Worm measures byte prediction rather than chatbot quality.

Related evaluation lessons from security

The signal and protocol studies use different tasks, but make useful comparisons about distribution shift, false alarms and resource budgets. The security catalogue is a separate map of defensive proposals.

Articles in this reading path

All reading paths