A research library for mechanistic interpretability and Theory of Mind in large language models
-
Updated
Aug 31, 2026 - Python
A research library for mechanistic interpretability and Theory of Mind in large language models
Does clinical co-occurrence strength predict when a medical VLM's internal representation of a finding is decodable but not causally used? Interpretability study on CheXagent-2-3b using activation probing and patching.
Code and processed results for a cross-model causal analysis of subject–verb agreement control in Phi-2, Llama-3.2-3B, and Qwen2.5-3B.
To associate your repository with the activation-patch topic, visit your repo's landing page and select "manage topics."