diff --git a/docs/runpod_run_guide.md b/docs/runpod_run_guide.md index f5fba401..20739cbe 100644 --- a/docs/runpod_run_guide.md +++ b/docs/runpod_run_guide.md @@ -28,9 +28,12 @@ *Why: keys are injected at container start. Adding one to a running pod does nothing, and restarting does not re-inject it — the environment was fixed at creation.* -4. **Smoke-test new hardware at throwaway length before committing a budget.** - *Why: this is the single practice that saved the first campaign. Forty minutes and under a - dollar caught a machine that would have consumed the entire budget.* +4. **Smoke-test new hardware at throwaway length, and judge it on EVALUATION, not training.** + *Why: training and evaluation are bound by different resources. Training survives on few cores; + posterior sampling does not. A machine that trains at 1.4x baseline can then take 2 h 16 min + for a single origin — 29 hours for thirteen. That happened on 2026-09-28, on a pod certified as + healthy from its training progress alone. Forty minutes and under a dollar is what this costs; + the alternative cost most of a night.* 5. **Publish credentials never go on rented hardware.** *Why: a read credential and a write credential are different decisions. The datafactory key @@ -92,8 +95,18 @@ Machines that satisfied it in practice: `RTX PRO 4500 SE`, `RTX PRO 4500`, `RTX A6000`, `RTX 4090`, `RTX 6000 Ada`, `RTX 5090`, `L40`, `L40S`. -**Do not take `PRO 6000 MIG 24GB`** (31 GB RAM, 6.8 effective CPUs). It is the cheapest listing -and it was 25× slower — 0.11 posterior-sampling steps/s against 2.8 on a machine that fits. +**The vCPU floor is the one that bites, and RAM does not compensate for it.** Measured +2026-09-28 on three machines: + +| RAM | effective CPUs | training | evaluation | +|---|---|---|---| +| 31 GB | 6.8 | — | 0.11 steps/s | +| **57 GB** | **6.8** | 1.4x baseline | **0.03 steps/s** — 2 h 16 min for ONE origin | +| 87 GB | 13.6 | baseline | 2.8 steps/s | + +Doubling the RAM did not rescue 6.8 cores. **Do not take `PRO 6000 MIG 24GB`** (31 GB, 6.8 cores) +or anything else below the floor, however cheap — a machine at 6.8 cores needs ~29 hours for the +13 origins, so it is the most expensive listing on the page, not the cheapest. Availability churns on a scale of seconds; listings vanish mid-form. Hold to the *rule* rather than a favourite model, or the hunt becomes the bottleneck. @@ -236,7 +249,10 @@ export WANDB_MODE=offline WANDB_SILENT=true Watch the posterior-sampling rate during evaluation. -| observed rate | verdict | +Watch the rate reported by **Drawing Posterior Samples**, during evaluation. Training progress +(`month/s`) is **not** the test and will look fine on a machine that cannot evaluate. + +| observed rate, evaluation | verdict | |---|---| | **≥ 2 steps/s** | healthy — proceed | | **< 0.5 steps/s sustained** | terminate the pod and take another | @@ -357,7 +373,7 @@ it. A stopped pod still bills for its volume, at double the running rate. | `.netrc` missing or wrong login | Refuses in preflight, seconds in | Step 2.3; check the login is *yours* | | `total_lessons` still a throwaway value | Refuses in preflight | Restore to 300, or re-clone | | `REGION` is not `"land"` | Refuses in preflight | Wrong branch or an edited config | -| Sampling collapses to ~0.1 steps/s | Runs, produces correct output, takes ~8× as long | Terminate; the machine is undersized (Ground rule 1) | +| Sampling collapses below ~0.5 steps/s | Runs and would produce correct output, but 13 origins take **20-30 hours** instead of one | Terminate and restart the model elsewhere. Training progress looking normal does not contradict this — the phases have different bottlenecks | | Pod dies mid-run | Everything on `/workspace` survives; the run does not | Restart the model; the volume persists | | Fewer than 13 parquets | `STATUS` is `FAILED:collapse` | `run.log` names the origin and the reason | | SSH refused after a restart | Port changed | Re-read the Connect tab | diff --git a/reports/postmortem_runpod_first_deployment_2026-09.md b/reports/postmortem_runpod_first_deployment_2026-09.md index 38b0a819..771b970a 100644 --- a/reports/postmortem_runpod_first_deployment_2026-09.md +++ b/reports/postmortem_runpod_first_deployment_2026-09.md @@ -106,19 +106,49 @@ Had I trusted the CPU diagnosis, I would have rejected every 8-vCPU instance on rest of the day — and instance availability churned so violently (§2.7) that this would have cost hours of the operator's evening for no reason. -**Root cause:** a 31 GB memory limit, approached but never breached, against a workload whose -posterior cube and input volume need more. The kernel reclaimed rather than killed, so nothing -failed — it only slowed, by a factor of 25. +**Root cause: CPU starvation during evaluation.** Posterior sampling is CPU-bound; training is +not. A machine with 6.8 effective cores trains acceptably and then collapses when it starts +drawing samples. -**Symptom mistaken for cause:** CPU saturation. Real (580% of 6.8 cores, 55% stall pressure) and -entirely downstream of the reclaim. +### 2.4a How I got this wrong three times, and what finally settled it -**How we know, and it was luck:** a later pod with the same 6.8 effective CPUs but 57 GB of RAM -ran at ~70% of full speed. Had every pod that day been either good or bad on *both* axes, the -wrong diagnosis would have survived the campaign and become a rule. +This is the most instructive thing in the document, so the sequence is worth keeping. -*Rule: `cpu.pressure` rising while `memory.current` approaches `memory.max` is a memory finding, -not a CPU finding. Distinguish them by varying one at a time, which we did only by accident.* +**First diagnosis — memory.** `memory.current` sat at 20.4 GB of a 31 GB limit and climbing. Fits +a cgroup-thrashing story. Wrong. + +**Second diagnosis — CPU.** `memory.pressure` read `avg10=0.00` while `cpu.pressure` read `55.6` +and the process sat at 580% of 6.8 cores. Right, as it turns out, but I abandoned it. + +**Third diagnosis — back to memory, and this is the one that reached the first version of this +document as a confident finding with a "Root cause" heading.** A later pod with the same 6.8 CPUs +but **57 GB** of RAM appeared to run at ~70% of full speed, which seemed to exonerate CPU. I wrote +that up, and used it to tell the operator to keep the machine. + +**It was wrong, and it was wrong for a reason I had already written down elsewhere in this +document: I extrapolated whole-run behaviour from the training phase.** Training on that pod took +205 minutes against ~150 baseline — 1.37x, tolerable, and all I had looked at. Then evaluation +began: + +| pod | RAM | effective CPUs | training | evaluation | +|---|---|---|---|---| +| morning, abandoned | 31 GB | 6.8 | — | 0.11 steps/s | +| bright_starship | **57 GB** | **6.8** | 1.37x baseline | **0.03 steps/s** | +| the other six | 87 GB | 13.6 | baseline | 2.8 steps/s | + +Origin 1 of 13 took **2 h 16 min** at 32.87 s/step. Thirteen origins would have been **29 hours** +— on a machine I had certified as fine. It was killed and the model restarted on a 13.6-CPU pod. + +**Memory is not the discriminator. Cores are.** 57 GB did not rescue 6.8 cores; it made the +collapse slightly worse, if anything. The second diagnosis was right and I talked myself out of it +on evidence that only covered the phase where the bottleneck does not bite. + +**The generalisable finding:** *training and evaluation have different bottlenecks, so a pod that +looks healthy while training tells you nothing about what it will do at origin 1.* Certify a +machine on posterior-sampling throughput or do not certify it. + +*Rule: judge a rented machine on evaluation throughput, never on training progress. The two +phases are bound by different resources and only one of them is where the money is lost.* ### 2.5 The guard that cannot see the container @@ -447,8 +477,9 @@ shape from a plausible-looking artefact instead of asking the session that owns here depends on the remaining five, but the per-model average may move. `runpod_cost_and_time_note_2026-09.md` uses the same three and must be reissued with this document if it changes. -- **We do not know that 6.8 vCPU is generally sufficient** — only that one pod with 57 GB of RAM - ran at ~70% speed. The RAM/CPU interaction is inferred from two data points. +- **6.8 effective cores are NOT sufficient**, on three data points now: two machines at that + count collapsed in evaluation regardless of having 31 GB or 57 GB of RAM. We do not know where + between 6.8 and 13.6 the threshold sits. - **The 2 steps/s health threshold is n=1 good machine and n=1 bad one.** It separates those two cleanly, which is what an operator needs, but it is not a hardware expectation: a laptop 4070 does the bare forward at ~15 steps/s, so even a healthy pod spends most of its time off the