A curated list of resources for activation engineering
-
Updated
Oct 2, 2025
A curated list of resources for activation engineering
Official code for Steering Large Language Models using Conceptors, presented at the NeurIPS 2024 MINT Workshop.
🔓 Ablate — directional ablation (abliteration) toolkit for open-source LLMs. Automatic censorship/refusal removal via residual-stream direction ablation, with KL-guided search, an LLM-judge harness, and one-call push to the Hub. pip install ablate-llm
Runtime control of LLM agent behaviors through activation steering vectors. More calibrated than prompting.
Located the number "five" inside Qwen2.5-32B and DeepSeek-V4-Flash (304B) and replaced it with "four" at 128 neurons/layer. The model still reads 5 but computes with it as 4, and in every test we ran it never noticed. Full records, reproducible on low-end hardware.
Iterative Sparse Matrix Steering: Closed-Form Subspace Alignment for Multi-Layer LLM Control (No SGD required).
Python library for activation steering, representation engineering, and inference-time control of transformer language models.
A closed-loop control system for Large Language Models that steers internal activation states in real-time to prevent mode collapse and toxicity
How meaning moves through a transformer - found, traced, and tested across four model scales.
To associate your repository with the activation-engineering topic, visit your repo's landing page and select "manage topics."