From 36b60f1ce9baa49b69a5f9aa70c484f601368c8a Mon Sep 17 00:00:00 2001 From: Polichinl Date: Tue, 29 Sep 2026 01:00:31 +0200 Subject: [PATCH] =?UTF-8?q?fix(#499):=20CPU=20starvation,=20not=20memory,?= =?UTF-8?q?=20was=20the=20cause=20=E2=80=94=20and=20training=20never=20sho?= =?UTF-8?q?ws=20it?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Correcting a finding I published as settled in the post-mortem yesterday, with a "Root cause" heading, and then acted on by telling the operator to keep a machine. WHAT HAPPENED. A pod with 6.8 effective CPUs but 57 GB of RAM trained at 1.37x baseline — 205 min against ~150 — so I certified it healthy and recorded that as evidence that the morning's failure had been memory, not CPU. Then evaluation started: Origin 1/13 1484/1484 [2:16:45<00:00, 32.87s/step] 32.87 s/step against 0.36 on a machine that fits. Thirteen origins would have been ~29 hours. Killed; bright_starship restarted on a 13.6-CPU pod. WHAT THE THREE MACHINES ACTUALLY SAY: 31 GB / 6.8 cores -> 0.11 steps/s 57 GB / 6.8 cores -> 0.03 steps/s training 1.4x baseline 87 GB / 13.6 cores -> 2.8 steps/s Doubling RAM did not rescue 6.8 cores; if anything it was slightly worse. Memory is not the discriminator. Cores are. My SECOND diagnosis (CPU, from cpu.pressure 55.6 against memory.pressure 0.00) was correct and I talked myself out of it on evidence that only covered training — the phase where the bottleneck does not bite. THE GENERALISABLE FINDING, which is why this is worth a commit of its own: training and evaluation are bound by DIFFERENT RESOURCES. Training is GPU-bound and survives on few cores. Posterior sampling is CPU-bound and collapses. So a pod that looks healthy while training tells you nothing about what it will do at origin 1, and certifying a machine on training progress is certifying it on the wrong thing. This is the third time in two days I extrapolated whole-run behaviour from an unrepresentative sample — after the 84 s/lesson figure taken from a cold first lesson, and the 11,368x compression ratio taken from a 2-lesson model. Same shape, three different quantities. It is recorded as such in the post-mortem rather than as three separate slips. CHANGES. Post-mortem §2.4 replaced: the root cause is now CPU, with a new §2.4a giving the full sequence of three diagnoses and what settled it. §11 corrected — "we do not know that 6.8 vCPU is sufficient" becomes "6.8 cores are NOT sufficient, on three data points". Guide: ground rule 4 now says judge a machine on EVALUATION, not training; the health table says which progress bar to watch and warns that training progress will look fine on a machine that cannot evaluate; the selection rule carries the three-machine table showing RAM does not compensate for cores; and the failure-mode row says 20-30 hours rather than "8x as long". Verified: 8036 passed, ruff clean, validate_docs.sh passes. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K --- docs/runpod_run_guide.md | 30 +++++++--- ...tmortem_runpod_first_deployment_2026-09.md | 55 +++++++++++++++---- 2 files changed, 66 insertions(+), 19 deletions(-) diff --git a/docs/runpod_run_guide.md b/docs/runpod_run_guide.md index f5fba401..20739cbe 100644 --- a/docs/runpod_run_guide.md +++ b/docs/runpod_run_guide.md @@ -28,9 +28,12 @@ *Why: keys are injected at container start. Adding one to a running pod does nothing, and restarting does not re-inject it — the environment was fixed at creation.* -4. **Smoke-test new hardware at throwaway length before committing a budget.** - *Why: this is the single practice that saved the first campaign. Forty minutes and under a - dollar caught a machine that would have consumed the entire budget.* +4. **Smoke-test new hardware at throwaway length, and judge it on EVALUATION, not training.** + *Why: training and evaluation are bound by different resources. Training survives on few cores; + posterior sampling does not. A machine that trains at 1.4x baseline can then take 2 h 16 min + for a single origin — 29 hours for thirteen. That happened on 2026-09-28, on a pod certified as + healthy from its training progress alone. Forty minutes and under a dollar is what this costs; + the alternative cost most of a night.* 5. **Publish credentials never go on rented hardware.** *Why: a read credential and a write credential are different decisions. The datafactory key @@ -92,8 +95,18 @@ Machines that satisfied it in practice: `RTX PRO 4500 SE`, `RTX PRO 4500`, `RTX A6000`, `RTX 4090`, `RTX 6000 Ada`, `RTX 5090`, `L40`, `L40S`. -**Do not take `PRO 6000 MIG 24GB`** (31 GB RAM, 6.8 effective CPUs). It is the cheapest listing -and it was 25× slower — 0.11 posterior-sampling steps/s against 2.8 on a machine that fits. +**The vCPU floor is the one that bites, and RAM does not compensate for it.** Measured +2026-09-28 on three machines: + +| RAM | effective CPUs | training | evaluation | +|---|---|---|---| +| 31 GB | 6.8 | — | 0.11 steps/s | +| **57 GB** | **6.8** | 1.4x baseline | **0.03 steps/s** — 2 h 16 min for ONE origin | +| 87 GB | 13.6 | baseline | 2.8 steps/s | + +Doubling the RAM did not rescue 6.8 cores. **Do not take `PRO 6000 MIG 24GB`** (31 GB, 6.8 cores) +or anything else below the floor, however cheap — a machine at 6.8 cores needs ~29 hours for the +13 origins, so it is the most expensive listing on the page, not the cheapest. Availability churns on a scale of seconds; listings vanish mid-form. Hold to the *rule* rather than a favourite model, or the hunt becomes the bottleneck. @@ -236,7 +249,10 @@ export WANDB_MODE=offline WANDB_SILENT=true Watch the posterior-sampling rate during evaluation. -| observed rate | verdict | +Watch the rate reported by **Drawing Posterior Samples**, during evaluation. Training progress +(`month/s`) is **not** the test and will look fine on a machine that cannot evaluate. + +| observed rate, evaluation | verdict | |---|---| | **≥ 2 steps/s** | healthy — proceed | | **< 0.5 steps/s sustained** | terminate the pod and take another | @@ -357,7 +373,7 @@ it. A stopped pod still bills for its volume, at double the running rate. | `.netrc` missing or wrong login | Refuses in preflight, seconds in | Step 2.3; check the login is *yours* | | `total_lessons` still a throwaway value | Refuses in preflight | Restore to 300, or re-clone | | `REGION` is not `"land"` | Refuses in preflight | Wrong branch or an edited config | -| Sampling collapses to ~0.1 steps/s | Runs, produces correct output, takes ~8× as long | Terminate; the machine is undersized (Ground rule 1) | +| Sampling collapses below ~0.5 steps/s | Runs and would produce correct output, but 13 origins take **20-30 hours** instead of one | Terminate and restart the model elsewhere. Training progress looking normal does not contradict this — the phases have different bottlenecks | | Pod dies mid-run | Everything on `/workspace` survives; the run does not | Restart the model; the volume persists | | Fewer than 13 parquets | `STATUS` is `FAILED:collapse` | `run.log` names the origin and the reason | | SSH refused after a restart | Port changed | Re-read the Connect tab | diff --git a/reports/postmortem_runpod_first_deployment_2026-09.md b/reports/postmortem_runpod_first_deployment_2026-09.md index 38b0a819..771b970a 100644 --- a/reports/postmortem_runpod_first_deployment_2026-09.md +++ b/reports/postmortem_runpod_first_deployment_2026-09.md @@ -106,19 +106,49 @@ Had I trusted the CPU diagnosis, I would have rejected every 8-vCPU instance on rest of the day — and instance availability churned so violently (§2.7) that this would have cost hours of the operator's evening for no reason. -**Root cause:** a 31 GB memory limit, approached but never breached, against a workload whose -posterior cube and input volume need more. The kernel reclaimed rather than killed, so nothing -failed — it only slowed, by a factor of 25. +**Root cause: CPU starvation during evaluation.** Posterior sampling is CPU-bound; training is +not. A machine with 6.8 effective cores trains acceptably and then collapses when it starts +drawing samples. -**Symptom mistaken for cause:** CPU saturation. Real (580% of 6.8 cores, 55% stall pressure) and -entirely downstream of the reclaim. +### 2.4a How I got this wrong three times, and what finally settled it -**How we know, and it was luck:** a later pod with the same 6.8 effective CPUs but 57 GB of RAM -ran at ~70% of full speed. Had every pod that day been either good or bad on *both* axes, the -wrong diagnosis would have survived the campaign and become a rule. +This is the most instructive thing in the document, so the sequence is worth keeping. -*Rule: `cpu.pressure` rising while `memory.current` approaches `memory.max` is a memory finding, -not a CPU finding. Distinguish them by varying one at a time, which we did only by accident.* +**First diagnosis — memory.** `memory.current` sat at 20.4 GB of a 31 GB limit and climbing. Fits +a cgroup-thrashing story. Wrong. + +**Second diagnosis — CPU.** `memory.pressure` read `avg10=0.00` while `cpu.pressure` read `55.6` +and the process sat at 580% of 6.8 cores. Right, as it turns out, but I abandoned it. + +**Third diagnosis — back to memory, and this is the one that reached the first version of this +document as a confident finding with a "Root cause" heading.** A later pod with the same 6.8 CPUs +but **57 GB** of RAM appeared to run at ~70% of full speed, which seemed to exonerate CPU. I wrote +that up, and used it to tell the operator to keep the machine. + +**It was wrong, and it was wrong for a reason I had already written down elsewhere in this +document: I extrapolated whole-run behaviour from the training phase.** Training on that pod took +205 minutes against ~150 baseline — 1.37x, tolerable, and all I had looked at. Then evaluation +began: + +| pod | RAM | effective CPUs | training | evaluation | +|---|---|---|---|---| +| morning, abandoned | 31 GB | 6.8 | — | 0.11 steps/s | +| bright_starship | **57 GB** | **6.8** | 1.37x baseline | **0.03 steps/s** | +| the other six | 87 GB | 13.6 | baseline | 2.8 steps/s | + +Origin 1 of 13 took **2 h 16 min** at 32.87 s/step. Thirteen origins would have been **29 hours** +— on a machine I had certified as fine. It was killed and the model restarted on a 13.6-CPU pod. + +**Memory is not the discriminator. Cores are.** 57 GB did not rescue 6.8 cores; it made the +collapse slightly worse, if anything. The second diagnosis was right and I talked myself out of it +on evidence that only covered the phase where the bottleneck does not bite. + +**The generalisable finding:** *training and evaluation have different bottlenecks, so a pod that +looks healthy while training tells you nothing about what it will do at origin 1.* Certify a +machine on posterior-sampling throughput or do not certify it. + +*Rule: judge a rented machine on evaluation throughput, never on training progress. The two +phases are bound by different resources and only one of them is where the money is lost.* ### 2.5 The guard that cannot see the container @@ -447,8 +477,9 @@ shape from a plausible-looking artefact instead of asking the session that owns here depends on the remaining five, but the per-model average may move. `runpod_cost_and_time_note_2026-09.md` uses the same three and must be reissued with this document if it changes. -- **We do not know that 6.8 vCPU is generally sufficient** — only that one pod with 57 GB of RAM - ran at ~70% speed. The RAM/CPU interaction is inferred from two data points. +- **6.8 effective cores are NOT sufficient**, on three data points now: two machines at that + count collapsed in evaluation regardless of having 31 GB or 57 GB of RAM. We do not know where + between 6.8 and 13.6 the threshold sits. - **The 2 steps/s health threshold is n=1 good machine and n=1 bad one.** It separates those two cleanly, which is what an operator needs, but it is not a hardware expectation: a laptop 4070 does the bare forward at ~15 steps/s, so even a healthy pod spends most of its time off the