Skip to content

docs(#499): the first RunPod deployment — post-mortem, cost note, operator guide - #508

Merged
Polichinel merged 2 commits into
developmentfrom
docs/runpod-deployment-postmortem-guide-and-cost
Sep 28, 2026
Merged

Polichinel merged 2 commits into
developmentfrom
docs/runpod-deployment-postmortem-guide-and-cost

Conversation

@Polichinel

Copy link
Copy Markdown
Collaborator

Three documents from one effort, split by audience, plus tools/podrun v0.1.0 and a correction to a number #507 merged.

What is here

reports/postmortem_runpod_first_deployment_2026-09.md what happened and what must change
reports/runpod_cost_and_time_note_2026-09.md what it cost — written for a manager
docs/runpod_run_guide.md how to do it
tools/podrun/ the runner, v0.1.0 PROVISIONAL

They were one file. It was trying to be a narrative, a cost table and a procedure at once and doing none of them well. Each fact now has one owner and the others cross-reference — the guide is the how, the post-mortem is the why, following fao_delivery_runbook.md's precedent.

tools/podrun is marked provisional in four places (package docstring, a banner in the script, the guide's metadata, and again where the guide first invokes it), and states what promoting it out of 0.1.0 would require: a second model family through it, the forecasting partition through it, and either CI exercising it or a maintainer other than its author using it unaided.

The correction — please read this part

#507 merged a wrong number under my name. The eight config comments, run_integration_tests.sh and docs/CICs/IntegrationTestRunner.md all say "one lesson is 84 s, so 300 lessons is ~7 h", stated as measured fact.

It is wrong. The 84 s came from the first lesson of a cold two-lesson smoke run, which carries warm-up and is not representative of the other 299. Three completed models give 202, 253 and 272 minutes end to end — 40–54 s per lesson, ~4 h per model.

Corrected in all three places, with the error named in the text rather than quietly overwritten.

No operational harm. The recommended --timeout 30000 was over-provisioned against the wrong number and remains safe against the right one.

This repeats, within hours, the lesson the post-mortem's own §2.6 records — never characterise against a throwaway model's output. It is now §7 item 8.

Independent review — three peer sessions, each found something real

views-datafactory found the sharpest one. The credential preflight check I wrote passes only because urllib follows a 308 and carries the Authorization header across it — the exact pattern their #388 removed. Verified on a live pod:

(bare)       -> 308      Authorization in req.headers:           True
/            -> 200      Authorization in req.unredirected_hdrs: False
/.zmetadata  -> 200

Now points at .zmetadata, which answers 200 directly. They also raised the client floor to >=1.13.0, where the credential-handling fixes landed — before that the client could carry a netrc credential across a redirect to another host.

views-hydranet caught that the guide stated two different runtimes, which is how the 84 s error surfaced. Also: the disk_guard problem is worse than I described — even taught to read the cgroup it counts only the posterior cube, missing the ~2.3 GB input volume and the torch context, so it would still under-report peak by 2–3×. The 2 steps/s health threshold is pod-relative, not a hardware expectation (a 4070 laptop does ~15). And both q95 figures were measured at S=16, so the claim is stability across datasets, not across sample counts — the document now says so.

views-faoapi caught an overreach headed for a manager: "the upload can run from a machine Simon controls" is a procedure claim I was not entitled to make. The architecture supports it (UPLOAD_ENABLED defaults False, no store client constructed when disarmed) but there is one entrypoint, no --no-upload, and disarming means editing a committed delivery declaration. Restated as architecture plus an open question. They also caught that the ensemble and delivery steps are per-run and memory-bound, so the per-model extrapolation is silent about them — and that the two documents disagreed on how many models were complete.

Where I pushed back

views-hydranet noted my ~4 h does not match fimbulthul's ~5.5 h. True, but not a correction: the guide documents rented pods, and 4 h is measured on that hardware, n=3. Kept my figure, added theirs as context.

Verification

  • 8036 passed, 0 failed; the skip delta against the previous run is pre-existing environmental skips (views_hydranet absent from the test env, synthetic targets, C-74)
  • ruff check . clean
  • bash docs/validate_docs.sh passes
  • tests/test_tools_layout.py passes with the new group
  • Downloaded predictions (56 MB) deliberately not committed

Still open

The post-mortem is Status: DRAFT — five of eight models were running when this was written. If their timings move the average, both it and the cost note get reissued. Its §10 lists seven open items, including three issues I owe (the disk_guard bug, the >=1.9.0 floor still in the model requirements, and a register entry for plain HTTP from rented hardware).

🤖 Generated with Claude Code

https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

…rator guide

Three documents from one effort, split by audience because they were fighting
each other in a single file. Plus `tools/podrun` v0.1.0, and a correction to a
number that #507 merged.

- reports/postmortem_runpod_first_deployment_2026-09.md — what happened and what
  must change. Owns the narrative, the misdiagnoses, and the failures.
- reports/runpod_cost_and_time_note_2026-09.md — what it cost. Written for a
  manager, cost and time only, one clearly-labelled extrapolation.
- docs/runpod_run_guide.md — how to do it. Owns the procedure and the gates.
- tools/podrun/ — the runner, marked v0.1.0 PROVISIONAL in four places, with the
  conditions for promoting it out of 0.1.0 stated.

Each fact has one owner; the others cross-reference. The guide is the how, the
post-mortem is the why, per fao_delivery_runbook.md's precedent.

THE CORRECTION. #507 merged "one lesson is 84 s, so 300 lessons is ~7 h" into the
eight config comments, run_integration_tests.sh and the runner's CIC, stated as
measured fact. It is wrong. The 84 s came from the FIRST lesson of a cold
two-lesson smoke run, which carries warm-up and is not representative of the
other 299. Three completed models give 202, 253 and 272 minutes end to end —
40-54 s per lesson, ~4 h per model. Corrected in all three places, with the
error named rather than quietly overwritten. No operational harm: the
recommended --timeout 30000 was over-provisioned against the wrong number and
remains safe against the right one.

This repeats, within hours, the lesson the post-mortem's own §2.6 records —
never characterise against a throwaway model's output. It is now §7 item 8.

REVIEW. Three peer sessions reviewed independently and each found something real.

views-hydranet: the guide stated two different runtimes (the above); the
disk_guard problem is worse than described — even taught to read the cgroup it
counts only the posterior cube, missing the ~2.3 GB input volume and the torch
context, so it would still under-report peak by 2-3x; the 2 steps/s threshold is
pod-relative, not a hardware expectation (a 4070 laptop does ~15); both q95
figures were measured at S=16, so the claim is stability across datasets, NOT
across sample counts; and C-151 must be cited as views-models C-151 because
views-hydranet's C-151 is an unrelated entry.

views-faoapi: "the upload can run from a machine Simon controls" was a procedure
claim I was not entitled to — the architecture supports it (UPLOAD_ENABLED
defaults False, no store client constructed when disarmed) but there is one
entrypoint, no --no-upload, and disarming edits a committed delivery
declaration. Restated as architecture + open question. Also: the ensemble and
delivery steps are per-run and memory-bound, so the per-model extrapolation is
silent about them; and the two documents disagreed on how many models were
complete.

views-datafactory: the preflight check I wrote passes only because urllib follows
a 308 and carries the Authorization header across it — the exact pattern their
#388 removed. Verified on a live pod: bare URL 308, /.zmetadata 200, and
Authorization sits in req.headers not unredirected_hdrs. Now points at
.zmetadata. Also raised the client floor to >=1.13.0, where the
credential-handling fixes landed.

Verified: 8036 passed 0 failed, ruff clean, validate_docs.sh passes,
test_tools_layout passes with the new group. Downloaded predictions deliberately
not committed.

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
Polichinel added a commit that referenced this pull request Sep 28, 2026
…nsit, and we put it on rented hardware (#510)

The datafactory speaks plain HTTP. HTTP Basic sends the credential base64-encoded
on every chunk request, and base64 is encoding, not encryption. views-datafactory
accepted that as their C-318 when the audience was a trusted circle on trusted
networks, which on fimbulthul was reasonable.

On 2026-09-28 we changed the audience without changing the mechanism: the
credential went onto five rented machines in datacentres we do not control. The
operator chose knowingly to use his personal login rather than provision a
throwaway, having been told what it meant. That is recorded, not second-guessed —
the work was owed and the server was gone.

What makes the residual risk outlive the run is two properties of the credential
itself: no expiry and no per-host registration. It stays valid until a person
rotates it by hand, and it authenticates from anywhere. So a pod image, a
snapshot or a detached volume that survives a campaign carries a live permanent
credential, and only housekeeping closes that.

Tier 3: a real but unquantified interception risk rather than a demonstrated
compromise, with cheap known mitigations. The trigger is deliberately the general
case — any run on hardware we do not own — not RunPod specifically, because the
next one may be a collaborator's machine or a CI runner.

Mitigations in preference order: a throwaway login retired at campaign end (about
three commands for whoever administers the data server); TLS on the data server,
which removes the class entirely; or deleting rented volumes and images at the
end and rotating afterwards.

Raised by the views-datafactory session reviewing the RunPod post-mortem (#508).
Cross-refs C-151, views-datafactory C-318, and #509 — the client floor still
permits a version that carried this credential across redirects to other hosts.

Header: 160 -> 161 entries, 151 -> 152 concerns, Accepted 4 -> 5, T3 64 -> 65.
Verified: 8036 passed, register header tests green.


Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

Co-authored-by: Claude Opus 5 <[email protected]>
Bug review of tools/podrun/pod_run_model.sh before merging #508. Eight
successful runs had exercised the happy path; this is about the paths they did
not.

CRITICAL, and verified empirically rather than reasoned:

    find <no matches> -print0 | tar --null -T -   ->  exit 0, valid 22-byte archive

So compress_draws could write an empty posterior and the script would proceed to
STATUS=OK. The only other signal was a small number in MANIFEST that a human had
to notice. On the one model family this has run the find pattern matches; on the
next one it might not, and `__init__.py` already warns that no other family has
been through it. Now counted before and verified after:

  - refuse if find matches zero lr_* draw files
  - after tar, list the archive and require the entry count to equal the count found

HIGH: two config guards checked file TEXT, not the parsed value. A regex took the
first `'total_lessons': <digits>` anywhere in the file and a grep matched
`REGION = "land"` anywhere including comments. That is the defect this repo
shipped once already — #501's "the guard that was not one", where a substring
assertion was satisfied by a comment recording the region's history. Both now
import the config module and read the value the interpreter sees. Verified: reads
300 and "land" from violet_visitor, and the file does contain "300 lessons" in a
comment, which is the surface a text check exposes.

HIGH: `SRC=$(ls -d ...predictions_calibration_*)` had no existence guard. An empty
result meant `cd "$SRC/.."` -> `cd /..`, i.e. the filesystem root. It happened to
fail loud because find errors on an empty path argument, which is an accident to
depend on. Now refuses explicitly.

MEDIUM: STATUS was never cleared at start, so a stale FAILED from an earlier
attempt shadowed a fresh run for its whole multi-hour duration -- and the script's
own header advertises `cat STATUS` as the way to watch from outside. Cleared now.

MEDIUM: output directories were not cleared between attempts. Parquets are named
from the SOURCE run's timestamp, so an old set and a new set could coexist; if
they summed to 13 the count check would pass while the manifest covered two
different training runs. Both output dirs are now cleared per attempt.

MEDIUM: the model and converter existence checks had no stage of their own, so
die() reported whatever stage ran last -- FAILED:verify_env for "no such model".

MEDIUM: die()'s explanation went only through the tee subshell, which can lose
its last buffered lines on a hard kill, at exactly the moment the reason is
wanted. The reason is now also written directly to $OUT/FAILURE, as STATUS and
STAGE already were.

LOW-MEDIUM: no lock, so two invocations for the same model on one pod raced on
the clone, the log and STATUS. I nearly caused this myself today when a chain and
a manual assignment could both have taken bright_starship. Now takes $OUT/.lock
and releases it on exit.

Known and NOT fixed, recorded in __init__.py and the post-mortem's open items:
the script has no automated tests, and MANIFEST's git sha and lesson count are
unchecked, so they would ship blank rather than refuse. Both are accepted at
v0.1.0.

Verified: all three guards fire on constructed inputs (empty archive refused;
counts match on real files; empty SRC refused), the config check reads real
values through the interpreter, bash -n clean, ruff clean, test_tools_layout
passes.

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
@Polichinel

Copy link
Copy Markdown
Collaborator Author

Bug review — tools/podrun/pod_run_model.sh

Scoped review of the one genuinely unreviewed thing in this PR: 157 lines of bash that run unattended on machines that bill by the second. The prose had already had three expert reviews; the script had had none, and eight successful runs is evidence about the happy path, not about the ninth input.

Seven issues, all fixed in 11294261.

1. CRITICAL — the draws archive could be empty and still report STATUS: OK. Verified empirically, not reasoned:

find <no matches> -print0 | tar --null -T -   ->  exit 0, valid 22-byte archive, 0 entries

Nothing downstream checked it. The only other signal was a small number in MANIFEST that a human had to notice. The find pattern matches on the one model family this has run; __init__.py already warns no other family has been through it, and this is exactly a next-family failure. Now counted before and the archive listed after, with the counts required to agree.

2. HIGH — two guards checked config text, not the parsed value. A regex took the first 'total_lessons': <digits> anywhere in the file; a grep matched REGION = "land" anywhere, comments included. That is #501's "the guard that was not one" reproduced — a substring assertion satisfied by a comment recording the value's history. Both now import the config and read what the interpreter sees.

3. HIGH — SRC had no existence guard, so an empty glob meant cd "$SRC/.." → cd /... It failed loud only because find errors on an empty path argument, which is an accident to rely on.

4–7. STATUS never cleared between attempts (a stale FAILED shadowed a live run for hours, and the header advertises cat STATUS as the way to watch); output dirs not cleared, so parquets from two training runs could sum to 13 and pass the count check under one manifest; two checks had no stage of their own so die() named the wrong one; die()'s reason went only through the tee subshell, which can lose it on a hard kill — now written directly alongside STATUS and STAGE; and no lock, which I nearly tripped myself today when a chain and a manual assignment could both have claimed bright_starship.

Known and not fixed, recorded in __init__.py and the post-mortem's open items: no automated tests, and MANIFEST's git sha and lesson count are unchecked so they would ship blank rather than refuse. Accepted at v0.1.0.

Guards verified against constructed inputs — empty archive refused, counts match on real files, empty SRC refused, config check reads 300 and land through the interpreter.

# ── 3. the run ────────────────────────────────────────────────────────────────────
stage train_and_evaluate
cd "$REPO/models/$MODEL" || die "cannot enter model dir"
export WANDB_MODE=offline WANDB_SILENT=true
START=$(date +%s)
"$VENV/bin/python" main.py -r calibration -t -e || die "main.py exited non-zero"
echo "run took $(( ($(date +%s) - START) / 60 )) minutes"
# ── 4. collapse to the deliverable ────────────────────────────────────────────────
stage collapse
# Clear a previous attempt first. Parquets are named from the SOURCE run's timestamp, so an
# old set and a new set can coexist; if they happened to sum to 13 the count check below
# would pass while the manifest covered two different training runs.
rm -rf "$OUT/parquet"
mkdir -p "$OUT/parquet"
cd "$REPO" || die "cannot enter repo"
"$VENV/bin/python" -m tools.collapse.collapse_predictions \
"models/$MODEL" --run-type calibration --out-dir "$OUT/parquet" || die "collapse failed"
N=$(ls -1 "$OUT/parquet"/*.parquet 2>/dev/null | wc -l)
[ "$N" -eq 13 ] || die "expected 13 parquets, got $N"
stage verify_parquet
"$VENV/bin/python" - "$OUT/parquet" <<'PY' || die "parquet verification failed"
import sys, glob, pandas as pd, numpy as np
fs = sorted(glob.glob(sys.argv[1] + "/*.parquet"))

🤖 Generated with Claude Code

@Polichinel
Polichinel merged commit a2bfe3a into development Sep 28, 2026
6 checks passed
@Polichinel
Polichinel deleted the docs/runpod-deployment-postmortem-guide-and-cost branch September 28, 2026 22:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant