Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
30 changes: 23 additions & 7 deletions docs/runpod_run_guide.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,9 +28,12 @@
*Why: keys are injected at container start. Adding one to a running pod does nothing, and
restarting does not re-inject it — the environment was fixed at creation.*

4. **Smoke-test new hardware at throwaway length before committing a budget.**
*Why: this is the single practice that saved the first campaign. Forty minutes and under a
dollar caught a machine that would have consumed the entire budget.*
4. **Smoke-test new hardware at throwaway length, and judge it on EVALUATION, not training.**
*Why: training and evaluation are bound by different resources. Training survives on few cores;
posterior sampling does not. A machine that trains at 1.4x baseline can then take 2 h 16 min
for a single origin — 29 hours for thirteen. That happened on 2026-09-28, on a pod certified as
healthy from its training progress alone. Forty minutes and under a dollar is what this costs;
the alternative cost most of a night.*

5. **Publish credentials never go on rented hardware.**
*Why: a read credential and a write credential are different decisions. The datafactory key
Expand Down Expand Up @@ -92,8 +95,18 @@
Machines that satisfied it in practice: `RTX PRO 4500 SE`, `RTX PRO 4500`, `RTX A6000`,
`RTX 4090`, `RTX 6000 Ada`, `RTX 5090`, `L40`, `L40S`.

**Do not take `PRO 6000 MIG 24GB`** (31 GB RAM, 6.8 effective CPUs). It is the cheapest listing
and it was 25× slower — 0.11 posterior-sampling steps/s against 2.8 on a machine that fits.
**The vCPU floor is the one that bites, and RAM does not compensate for it.** Measured
2026-09-28 on three machines:

| RAM | effective CPUs | training | evaluation |
|---|---|---|---|
| 31 GB | 6.8 | — | 0.11 steps/s |
| **57 GB** | **6.8** | 1.4x baseline | **0.03 steps/s** — 2 h 16 min for ONE origin |
| 87 GB | 13.6 | baseline | 2.8 steps/s |

Doubling the RAM did not rescue 6.8 cores. **Do not take `PRO 6000 MIG 24GB`** (31 GB, 6.8 cores)
or anything else below the floor, however cheap — a machine at 6.8 cores needs ~29 hours for the
13 origins, so it is the most expensive listing on the page, not the cheapest.

Availability churns on a scale of seconds; listings vanish mid-form. Hold to the *rule* rather
than a favourite model, or the hunt becomes the bottleneck.
Expand Down Expand Up @@ -236,7 +249,10 @@ export WANDB_MODE=offline WANDB_SILENT=true

Watch the posterior-sampling rate during evaluation.

| observed rate | verdict |
Watch the rate reported by **Drawing Posterior Samples**, during evaluation. Training progress
(`month/s`) is **not** the test and will look fine on a machine that cannot evaluate.

| observed rate, evaluation | verdict |
|---|---|
| **≥ 2 steps/s** | healthy — proceed |
| **< 0.5 steps/s sustained** | terminate the pod and take another |
Expand Down Expand Up @@ -357,7 +373,7 @@ it. A stopped pod still bills for its volume, at double the running rate.
| `.netrc` missing or wrong login | Refuses in preflight, seconds in | Step 2.3; check the login is *yours* |
| `total_lessons` still a throwaway value | Refuses in preflight | Restore to 300, or re-clone |
| `REGION` is not `"land"` | Refuses in preflight | Wrong branch or an edited config |
| Sampling collapses to ~0.1 steps/s | Runs, produces correct output, takes ~8× as long | Terminate; the machine is undersized (Ground rule 1) |
| Sampling collapses below ~0.5 steps/s | Runs and would produce correct output, but 13 origins take **20-30 hours** instead of one | Terminate and restart the model elsewhere. Training progress looking normal does not contradict this — the phases have different bottlenecks |
| Pod dies mid-run | Everything on `/workspace` survives; the run does not | Restart the model; the volume persists |
| Fewer than 13 parquets | `STATUS` is `FAILED:collapse` | `run.log` names the origin and the reason |
| SSH refused after a restart | Port changed | Re-read the Connect tab |
Expand Down
55 changes: 43 additions & 12 deletions reports/postmortem_runpod_first_deployment_2026-09.md
Original file line number Diff line number Diff line change
Expand Up @@ -106,19 +106,49 @@ Had I trusted the CPU diagnosis, I would have rejected every 8-vCPU instance on
rest of the day — and instance availability churned so violently (§2.7) that this would have cost
hours of the operator's evening for no reason.

**Root cause:** a 31 GB memory limit, approached but never breached, against a workload whose
posterior cube and input volume need more. The kernel reclaimed rather than killed, so nothing
failed — it only slowed, by a factor of 25.
**Root cause: CPU starvation during evaluation.** Posterior sampling is CPU-bound; training is
not. A machine with 6.8 effective cores trains acceptably and then collapses when it starts
drawing samples.

**Symptom mistaken for cause:** CPU saturation. Real (580% of 6.8 cores, 55% stall pressure) and
entirely downstream of the reclaim.
### 2.4a How I got this wrong three times, and what finally settled it

**How we know, and it was luck:** a later pod with the same 6.8 effective CPUs but 57 GB of RAM
ran at ~70% of full speed. Had every pod that day been either good or bad on *both* axes, the
wrong diagnosis would have survived the campaign and become a rule.
This is the most instructive thing in the document, so the sequence is worth keeping.

*Rule: `cpu.pressure` rising while `memory.current` approaches `memory.max` is a memory finding,
not a CPU finding. Distinguish them by varying one at a time, which we did only by accident.*
**First diagnosis — memory.** `memory.current` sat at 20.4 GB of a 31 GB limit and climbing. Fits
a cgroup-thrashing story. Wrong.

**Second diagnosis — CPU.** `memory.pressure` read `avg10=0.00` while `cpu.pressure` read `55.6`
and the process sat at 580% of 6.8 cores. Right, as it turns out, but I abandoned it.

**Third diagnosis — back to memory, and this is the one that reached the first version of this
document as a confident finding with a "Root cause" heading.** A later pod with the same 6.8 CPUs
but **57 GB** of RAM appeared to run at ~70% of full speed, which seemed to exonerate CPU. I wrote
that up, and used it to tell the operator to keep the machine.

**It was wrong, and it was wrong for a reason I had already written down elsewhere in this
document: I extrapolated whole-run behaviour from the training phase.** Training on that pod took
205 minutes against ~150 baseline — 1.37x, tolerable, and all I had looked at. Then evaluation
began:

| pod | RAM | effective CPUs | training | evaluation |
|---|---|---|---|---|
| morning, abandoned | 31 GB | 6.8 | — | 0.11 steps/s |
| bright_starship | **57 GB** | **6.8** | 1.37x baseline | **0.03 steps/s** |
| the other six | 87 GB | 13.6 | baseline | 2.8 steps/s |

Origin 1 of 13 took **2 h 16 min** at 32.87 s/step. Thirteen origins would have been **29 hours**
— on a machine I had certified as fine. It was killed and the model restarted on a 13.6-CPU pod.

**Memory is not the discriminator. Cores are.** 57 GB did not rescue 6.8 cores; it made the
collapse slightly worse, if anything. The second diagnosis was right and I talked myself out of it
on evidence that only covered the phase where the bottleneck does not bite.

**The generalisable finding:** *training and evaluation have different bottlenecks, so a pod that
looks healthy while training tells you nothing about what it will do at origin 1.* Certify a
machine on posterior-sampling throughput or do not certify it.

*Rule: judge a rented machine on evaluation throughput, never on training progress. The two
phases are bound by different resources and only one of them is where the money is lost.*

### 2.5 The guard that cannot see the container

Expand Down Expand Up @@ -447,8 +477,9 @@ shape from a plausible-looking artefact instead of asking the session that owns
here depends on the remaining five, but the per-model average may move.
`runpod_cost_and_time_note_2026-09.md` uses the same three and must be reissued with this
document if it changes.
- **We do not know that 6.8 vCPU is generally sufficient** — only that one pod with 57 GB of RAM
ran at ~70% speed. The RAM/CPU interaction is inferred from two data points.
- **6.8 effective cores are NOT sufficient**, on three data points now: two machines at that
count collapsed in evaluation regardless of having 31 GB or 57 GB of RAM. We do not know where
between 6.8 and 13.6 the threshold sits.
- **The 2 steps/s health threshold is n=1 good machine and n=1 bad one.** It separates those two
cleanly, which is what an operator needs, but it is not a hardware expectation: a laptop 4070
does the bare forward at ~15 steps/s, so even a healthy pod spends most of its time off the
Expand Down
Loading