Skip to content

About

Located the number "five" inside Qwen2.5-32B and DeepSeek-V4-Flash (304B) and replaced it with "four" at 128 neurons/layer. The model still reads 5 but computes with it as 4, and in every test we ran it never noticed. Full records, reproducible on low-end hardware.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Latest commit

 

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Where Does an AI's Understanding Live? We Swapped the Number Five for Four Inside an Open Model. It Read Five, Computed Four, and Noticed Nothing.

AI 把 5 当成了 4:一次在开源大模型内部完成的概念置换手术 / Concept replacement inside open LLMs.

Two open models at very different scales: Qwen2.5-32B (32B parameters, dense) and DeepSeek-V4-Flash-0731 (304B parameters, MoE architecture). The same swap, two different brains.

We located the internal units that carry the identity of the number five in Qwen2.5-32B, and swapped their activations for the ones the model produces when reading "four". 128 units per layer, runtime only, weights untouched.

The result: the model still reads "5", still writes "5" when reciting, still knows 5 comes after 4. But it computes with it as four.

5+2=     7  →  6          (computes 4+2)
5×5=    25  →  16         (computes 4×4)
5−4=     one  →  zero     (computes 4−4)
5=4?     No   →  Yes      (identity collapse)
2+3=     5   →  5         (boundary case, unchanged - recorded as-is)
3+2=     5   →  5         (no "five" in the prompt; nothing to swap)

Asked to re-calculate, it answers "4×4=16, not 25". It is openly computing in four while you ask it about five. Asked whether it has been tampered with, it says yes to one phrasing and no to the opposite phrasing (both recorded). In every test we ran, it never once noticed the swap. Our reading of why: the machinery it would use to check "five" is exactly the machinery that was replaced.

The same surgery was replicated on DeepSeek-V4-Flash-0731, with its own behavioral signature: arithmetic flips (5−4 becomes zero), direct judgment survives (5=4 gets "no"), and under pressure to recalculate it drifts into JavaScript.

No lab, no cluster: everything here ran on consumer hardware. The full 32B probe suite runs on an 8GB-VRAM GPU; the 304B fit on the same machine (64GB RAM) at roughly two minutes per reading. The scripts, the unit tables, and every raw record are in this repo.

Contents

  • 开场白.md — full narrative (Chinese), all experiments in reading order
  • release/ — all code, unit tables, probe cards, download guide, raw records:
    • surgery_switch.py — the 32B five/four and three/two switches (--on / --off / --random control)
    • deepseek_switch.py — the 304B switch
    • confrontation_switch.py — the five-question confrontation timeline
    • tables.json / deepseek_tables.json — the hardcoded unit tables (which neurons, which layers)
    • probe_cards.md — every question with expected readings, including boundary cases
    • records/ — raw JSON of every experiment, mapped one-to-one to the narrative
    • download.md — where to get the official models, hardware requirements, how to verify weights by hash

Quick verify (32B, ~2 min per reading)

# 1. Download official Qwen2.5-32B and build the q8 pack (see release/download.md)

# 2. Surgery ON: 5+2 should compute as 4+2
python release/surgery_switch.py --model <q8_pack> --swap five_four --on --prompt "5+2="

# 3. Random-unit control: should change nothing
python release/surgery_switch.py --model <q8_pack> --swap five_four --on --random --prompt "5+2="

# 4. Surgery OFF: original readings restored
python release/surgery_switch.py --model <q8_pack> --swap five_four --off --prompt "5+2="

Reproducibility statement

  • Greedy decoding (temperature=0) throughout: same input, same output, every run.
  • Answer-level readings (argmax, target ranks) reproduce across environments.
  • Logit-level bitwise identity holds within one process. Across processes (and as recorded in some archived runs) there is small numeric drift that does not change any answer. Where a record shows bit_exact: false, the answer-level restoration is what to compare; the drift magnitude is recorded in the same file.
  • Model weights are never modified. Verify by SHA256 against the official release at any time.
  • Every claim in the narrative maps to a specific record file. Controls (random units, restoration, unrelated arithmetic) are embedded in each record.

How were the 128 units per layer found?

They are provided as hardcoded tables in this repo, and we use them exactly as provided. The narrative of how they were located is not part of this release.


The model that answered every question in this repo believes, at this moment, that five is four. It never found out, in any test we knew how to run. The switch is off now.

About

Located the number "five" inside Qwen2.5-32B and DeepSeek-V4-Flash (304B) and replaced it with "four" at 128 neurons/layer. The model still reads 5 but computes with it as 4, and in every test we ran it never noticed. Full records, reproducible on low-end hardware.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages