Skip to content

fix(#499): CPU starvation, not memory, was the cause — and training never shows it - #511

Open
Polichinel wants to merge 1 commit into
developmentfrom
fix/499-cpu-not-memory-was-the-cause
Open

Polichinel wants to merge 1 commit into
developmentfrom
fix/499-cpu-not-memory-was-the-cause

Conversation

@Polichinel

Copy link
Copy Markdown
Collaborator

Correcting a finding I published as settled in the post-mortem yesterday, under a "Root cause" heading, and then acted on by telling the operator to keep a machine.

What happened

A pod with 6.8 effective CPUs but 57 GB of RAM trained at 1.37× baseline — 205 min against ~150. I certified it healthy on that basis, and recorded it as evidence that the morning's failure had been memory, not CPU.

Then evaluation started:

Origin 1/13   1484/1484 [2:16:45<00:00, 32.87s/step]

32.87 s/step, against 0.36 on a machine that fits. One origin took 2 h 16 min; thirteen would have been ~29 hours. Killed, and bright_starship restarted on a 13.6-CPU pod.

What the three machines actually say

RAM effective CPUs training evaluation
31 GB 6.8 — 0.11 steps/s
57 GB 6.8 1.4× baseline 0.03 steps/s
87 GB 13.6 baseline 2.8 steps/s

Doubling the RAM did not rescue 6.8 cores — it was slightly worse, if anything. Memory is not the discriminator. Cores are.

My second diagnosis (CPU, from cpu.pressure 55.6 against memory.pressure 0.00) was correct, and I talked myself out of it on evidence that only covered training — the phase where the bottleneck does not bite.

The generalisable finding

Training and evaluation are bound by different resources. Training is GPU-bound and survives on few cores. Posterior sampling is CPU-bound and collapses.

So a machine that looks healthy while training tells you nothing about what it will do at origin 1, and certifying one on training progress is certifying it on the wrong thing.

This is the third time in two days I extrapolated whole-run behaviour from an unrepresentative sample — after the 84 s/lesson figure taken from a cold first lesson, and the 11,368× compression ratio taken from a 2-lesson model. Same shape, three different quantities. Recorded in the post-mortem as one pattern rather than three slips.

Changes

Post-mortem — §2.4's root cause replaced; new §2.4a gives the full sequence of three diagnoses and what settled it. §11 corrected: "we do not know that 6.8 vCPU is sufficient" → "6.8 cores are NOT sufficient, on three data points".

Guide — ground rule 4 now says judge a machine on evaluation, not training; the health table names which progress bar to watch and warns that training progress looks fine on a machine that cannot evaluate; the selection rule carries the three-machine table; the failure-mode row says 20–30 hours rather than "8× as long".

Why this is its own PR

It contradicts something already merged (#508). Burying that in a later documentation commit would make the record harder to follow than the mistake was.

Verification

  • 8036 passed, 0 failed; ruff check . clean; bash docs/validate_docs.sh passes
  • The 32.87 s/step figure is from the run log on the pod, not inferred
  • bright_starship is now running on a 13.6-CPU pod and will complete normally

🤖 Generated with Claude Code

https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

…ever shows it

Correcting a finding I published as settled in the post-mortem yesterday, with a
"Root cause" heading, and then acted on by telling the operator to keep a machine.

WHAT HAPPENED. A pod with 6.8 effective CPUs but 57 GB of RAM trained at 1.37x
baseline — 205 min against ~150 — so I certified it healthy and recorded that as
evidence that the morning's failure had been memory, not CPU. Then evaluation
started:

    Origin 1/13  1484/1484 [2:16:45<00:00, 32.87s/step]

32.87 s/step against 0.36 on a machine that fits. Thirteen origins would have
been ~29 hours. Killed; bright_starship restarted on a 13.6-CPU pod.

WHAT THE THREE MACHINES ACTUALLY SAY:

    31 GB / 6.8 cores   ->  0.11 steps/s
    57 GB / 6.8 cores   ->  0.03 steps/s    training 1.4x baseline
    87 GB / 13.6 cores  ->  2.8  steps/s

Doubling RAM did not rescue 6.8 cores; if anything it was slightly worse. Memory
is not the discriminator. Cores are. My SECOND diagnosis (CPU, from cpu.pressure
55.6 against memory.pressure 0.00) was correct and I talked myself out of it on
evidence that only covered training — the phase where the bottleneck does not
bite.

THE GENERALISABLE FINDING, which is why this is worth a commit of its own:
training and evaluation are bound by DIFFERENT RESOURCES. Training is GPU-bound
and survives on few cores. Posterior sampling is CPU-bound and collapses. So a
pod that looks healthy while training tells you nothing about what it will do at
origin 1, and certifying a machine on training progress is certifying it on the
wrong thing.

This is the third time in two days I extrapolated whole-run behaviour from an
unrepresentative sample — after the 84 s/lesson figure taken from a cold first
lesson, and the 11,368x compression ratio taken from a 2-lesson model. Same
shape, three different quantities. It is recorded as such in the post-mortem
rather than as three separate slips.

CHANGES. Post-mortem §2.4 replaced: the root cause is now CPU, with a new §2.4a
giving the full sequence of three diagnoses and what settled it. §11 corrected —
"we do not know that 6.8 vCPU is sufficient" becomes "6.8 cores are NOT
sufficient, on three data points".

Guide: ground rule 4 now says judge a machine on EVALUATION, not training; the
health table says which progress bar to watch and warns that training progress
will look fine on a machine that cannot evaluate; the selection rule carries the
three-machine table showing RAM does not compensate for cores; and the
failure-mode row says 20-30 hours rather than "8x as long".

Verified: 8036 passed, ruff clean, validate_docs.sh passes.

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant