A complete end-to-end pipeline for LLM interpretability with sparse autoencoders (SAEs) using Llama 3.2, written in pure PyTorch and fully reproducible.
-
Updated
Mar 23, 2025 - Python
A complete end-to-end pipeline for LLM interpretability with sparse autoencoders (SAEs) using Llama 3.2, written in pure PyTorch and fully reproducible.
Real-time 3D visualisation of SAE feature activations inside GPT-2, token by token
A toolbox for exploring the informational nature of LLMs and language
Fast XAI with interactions at large scale. SPEX can help you understand the output of your LLM, even if you have a long context!
[Under Review] Not All Tokens Are Equally Useful for Steering: Robust Directions and Prefix Steering
Automates attribution-graph analysis via probe prompting: circuit-trace a prompt, auto-generate concept probes, profile feature activations, cluster supernodes.
Spectral Attention Divergence — dynamical systems probe for LLM inference via dual-path attention comparison and delay-coordinate attractor reconstruction
[ICLR 2026] AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
Turn a knob inside a small open model instead of writing a prompt. A reproducible RepE and CAA steering harness with an honest benchmark, including the failures.
A minimal mech-interp project for steering LLMs
Public research on LLM internals, Jacobian lenses, SAE steering, nonlinear dynamics, evaluation, and inspectable AI systems.
Mechanistic Interpretability Research Knowledge Base + Notes
Replication of 'From Reasoning to Answer' (EMNLP 2025) — Reasoning-Focus Heads + Activation Patching on DeepSeek-R1-Distill-Qwen-7B
Personal website of Xudong Zhu, featuring research, publications, open-source projects, and technical writing on LLM and representation learning.
Personal profile of Xudong Zhu. PhD researcher in representation learning.
This project explores methods to detect and mitigate jailbreak behaviors in Large Language Models (LLMs). By analyzing activation patterns—particularly in deeper layers—we identify distinct differences between compliant and non-compliant responses to uncover a jailbreak "direction." Using this insight, we develop intervention strategies that modify
Probing whether multilingual LLMs encode kinship distinctions as language-specific directions or universal shared representations
Normative, machine-first definition of Interpretive SEO (SEO interprétatif), aligned with the Interpretive Governance standard.
Mechanistic interpretability × inference research on GPT-2-small — a sparse autoencoder benchmarked against quantization, converted into a live low-overhead safety monitor.
Leakage-resistant experiments testing contextual transport and latent-coordinate structure across Qwen3-4B and Phi-4-mini, with frozen controls and reproducible reports.
To associate your repository with the llm-interpretability topic, visit your repo's landing page and select "manage topics."