Located the number "five" inside Qwen2.5-32B and DeepSeek-V4-Flash (304B) and replaced it with "four" at 128 neurons/layer. The model still reads 5 but computes with it as 4, and in every test we ran it never noticed. Full records, reproducible on low-end hardware.
ai-safety interpretability llm model-editing knowledge-editing mechanistic-interpretability llm-security causal-intervention open-source-llm qwen deepseek activation-patching representation-engineering activation-steering activation-engineering interpretability-ai concept-editing
-
Updated
Oct 5, 2026 - Python