docs(#499): the first RunPod deployment — post-mortem, cost note, operator guide - #508
Conversation
…rator guide Three documents from one effort, split by audience because they were fighting each other in a single file. Plus `tools/podrun` v0.1.0, and a correction to a number that #507 merged. - reports/postmortem_runpod_first_deployment_2026-09.md — what happened and what must change. Owns the narrative, the misdiagnoses, and the failures. - reports/runpod_cost_and_time_note_2026-09.md — what it cost. Written for a manager, cost and time only, one clearly-labelled extrapolation. - docs/runpod_run_guide.md — how to do it. Owns the procedure and the gates. - tools/podrun/ — the runner, marked v0.1.0 PROVISIONAL in four places, with the conditions for promoting it out of 0.1.0 stated. Each fact has one owner; the others cross-reference. The guide is the how, the post-mortem is the why, per fao_delivery_runbook.md's precedent. THE CORRECTION. #507 merged "one lesson is 84 s, so 300 lessons is ~7 h" into the eight config comments, run_integration_tests.sh and the runner's CIC, stated as measured fact. It is wrong. The 84 s came from the FIRST lesson of a cold two-lesson smoke run, which carries warm-up and is not representative of the other 299. Three completed models give 202, 253 and 272 minutes end to end — 40-54 s per lesson, ~4 h per model. Corrected in all three places, with the error named rather than quietly overwritten. No operational harm: the recommended --timeout 30000 was over-provisioned against the wrong number and remains safe against the right one. This repeats, within hours, the lesson the post-mortem's own §2.6 records — never characterise against a throwaway model's output. It is now §7 item 8. REVIEW. Three peer sessions reviewed independently and each found something real. views-hydranet: the guide stated two different runtimes (the above); the disk_guard problem is worse than described — even taught to read the cgroup it counts only the posterior cube, missing the ~2.3 GB input volume and the torch context, so it would still under-report peak by 2-3x; the 2 steps/s threshold is pod-relative, not a hardware expectation (a 4070 laptop does ~15); both q95 figures were measured at S=16, so the claim is stability across datasets, NOT across sample counts; and C-151 must be cited as views-models C-151 because views-hydranet's C-151 is an unrelated entry. views-faoapi: "the upload can run from a machine Simon controls" was a procedure claim I was not entitled to — the architecture supports it (UPLOAD_ENABLED defaults False, no store client constructed when disarmed) but there is one entrypoint, no --no-upload, and disarming edits a committed delivery declaration. Restated as architecture + open question. Also: the ensemble and delivery steps are per-run and memory-bound, so the per-model extrapolation is silent about them; and the two documents disagreed on how many models were complete. views-datafactory: the preflight check I wrote passes only because urllib follows a 308 and carries the Authorization header across it — the exact pattern their #388 removed. Verified on a live pod: bare URL 308, /.zmetadata 200, and Authorization sits in req.headers not unredirected_hdrs. Now points at .zmetadata. Also raised the client floor to >=1.13.0, where the credential-handling fixes landed. Verified: 8036 passed 0 failed, ruff clean, validate_docs.sh passes, test_tools_layout passes with the new group. Downloaded predictions deliberately not committed. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
…nsit, and we put it on rented hardware (#510) The datafactory speaks plain HTTP. HTTP Basic sends the credential base64-encoded on every chunk request, and base64 is encoding, not encryption. views-datafactory accepted that as their C-318 when the audience was a trusted circle on trusted networks, which on fimbulthul was reasonable. On 2026-09-28 we changed the audience without changing the mechanism: the credential went onto five rented machines in datacentres we do not control. The operator chose knowingly to use his personal login rather than provision a throwaway, having been told what it meant. That is recorded, not second-guessed — the work was owed and the server was gone. What makes the residual risk outlive the run is two properties of the credential itself: no expiry and no per-host registration. It stays valid until a person rotates it by hand, and it authenticates from anywhere. So a pod image, a snapshot or a detached volume that survives a campaign carries a live permanent credential, and only housekeeping closes that. Tier 3: a real but unquantified interception risk rather than a demonstrated compromise, with cheap known mitigations. The trigger is deliberately the general case — any run on hardware we do not own — not RunPod specifically, because the next one may be a collaborator's machine or a CI runner. Mitigations in preference order: a throwaway login retired at campaign end (about three commands for whoever administers the data server); TLS on the data server, which removes the class entirely; or deleting rented volumes and images at the end and rotating afterwards. Raised by the views-datafactory session reviewing the RunPod post-mortem (#508). Cross-refs C-151, views-datafactory C-318, and #509 — the client floor still permits a version that carried this credential across redirects to other hosts. Header: 160 -> 161 entries, 151 -> 152 concerns, Accepted 4 -> 5, T3 64 -> 65. Verified: 8036 passed, register header tests green. Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K Co-authored-by: Claude Opus 5 <[email protected]>
Bug review of tools/podrun/pod_run_model.sh before merging #508. Eight successful runs had exercised the happy path; this is about the paths they did not. CRITICAL, and verified empirically rather than reasoned: find <no matches> -print0 | tar --null -T - -> exit 0, valid 22-byte archive So compress_draws could write an empty posterior and the script would proceed to STATUS=OK. The only other signal was a small number in MANIFEST that a human had to notice. On the one model family this has run the find pattern matches; on the next one it might not, and `__init__.py` already warns that no other family has been through it. Now counted before and verified after: - refuse if find matches zero lr_* draw files - after tar, list the archive and require the entry count to equal the count found HIGH: two config guards checked file TEXT, not the parsed value. A regex took the first `'total_lessons': <digits>` anywhere in the file and a grep matched `REGION = "land"` anywhere including comments. That is the defect this repo shipped once already — #501's "the guard that was not one", where a substring assertion was satisfied by a comment recording the region's history. Both now import the config module and read the value the interpreter sees. Verified: reads 300 and "land" from violet_visitor, and the file does contain "300 lessons" in a comment, which is the surface a text check exposes. HIGH: `SRC=$(ls -d ...predictions_calibration_*)` had no existence guard. An empty result meant `cd "$SRC/.."` -> `cd /..`, i.e. the filesystem root. It happened to fail loud because find errors on an empty path argument, which is an accident to depend on. Now refuses explicitly. MEDIUM: STATUS was never cleared at start, so a stale FAILED from an earlier attempt shadowed a fresh run for its whole multi-hour duration -- and the script's own header advertises `cat STATUS` as the way to watch from outside. Cleared now. MEDIUM: output directories were not cleared between attempts. Parquets are named from the SOURCE run's timestamp, so an old set and a new set could coexist; if they summed to 13 the count check would pass while the manifest covered two different training runs. Both output dirs are now cleared per attempt. MEDIUM: the model and converter existence checks had no stage of their own, so die() reported whatever stage ran last -- FAILED:verify_env for "no such model". MEDIUM: die()'s explanation went only through the tee subshell, which can lose its last buffered lines on a hard kill, at exactly the moment the reason is wanted. The reason is now also written directly to $OUT/FAILURE, as STATUS and STAGE already were. LOW-MEDIUM: no lock, so two invocations for the same model on one pod raced on the clone, the log and STATUS. I nearly caused this myself today when a chain and a manual assignment could both have taken bright_starship. Now takes $OUT/.lock and releases it on exit. Known and NOT fixed, recorded in __init__.py and the post-mortem's open items: the script has no automated tests, and MANIFEST's git sha and lesson count are unchecked, so they would ship blank rather than refuse. Both are accepted at v0.1.0. Verified: all three guards fire on constructed inputs (empty archive refused; counts match on real files; empty SRC refused), the config check reads real values through the interpreter, bash -n clean, ruff clean, test_tools_layout passes. Co-Authored-By: Claude Opus 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
Bug review —
|
Three documents from one effort, split by audience, plus
tools/podrunv0.1.0 and a correction to a number #507 merged.What is here
reports/postmortem_runpod_first_deployment_2026-09.mdreports/runpod_cost_and_time_note_2026-09.mddocs/runpod_run_guide.mdtools/podrun/They were one file. It was trying to be a narrative, a cost table and a procedure at once and doing none of them well. Each fact now has one owner and the others cross-reference — the guide is the how, the post-mortem is the why, following
fao_delivery_runbook.md's precedent.tools/podrunis marked provisional in four places (package docstring, a banner in the script, the guide's metadata, and again where the guide first invokes it), and states what promoting it out of 0.1.0 would require: a second model family through it, the forecasting partition through it, and either CI exercising it or a maintainer other than its author using it unaided.The correction — please read this part
#507 merged a wrong number under my name. The eight config comments,
run_integration_tests.shanddocs/CICs/IntegrationTestRunner.mdall say "one lesson is 84 s, so 300 lessons is ~7 h", stated as measured fact.It is wrong. The 84 s came from the first lesson of a cold two-lesson smoke run, which carries warm-up and is not representative of the other 299. Three completed models give 202, 253 and 272 minutes end to end — 40–54 s per lesson, ~4 h per model.
Corrected in all three places, with the error named in the text rather than quietly overwritten.
No operational harm. The recommended
--timeout 30000was over-provisioned against the wrong number and remains safe against the right one.This repeats, within hours, the lesson the post-mortem's own §2.6 records — never characterise against a throwaway model's output. It is now §7 item 8.
Independent review — three peer sessions, each found something real
views-datafactory found the sharpest one. The credential preflight check I wrote passes only because urllib follows a 308 and carries the
Authorizationheader across it — the exact pattern their #388 removed. Verified on a live pod:Now points at
.zmetadata, which answers 200 directly. They also raised the client floor to>=1.13.0, where the credential-handling fixes landed — before that the client could carry a netrc credential across a redirect to another host.views-hydranet caught that the guide stated two different runtimes, which is how the 84 s error surfaced. Also: the
disk_guardproblem is worse than I described — even taught to read the cgroup it counts only the posterior cube, missing the ~2.3 GB input volume and the torch context, so it would still under-report peak by 2–3×. The 2 steps/s health threshold is pod-relative, not a hardware expectation (a 4070 laptop does ~15). And both q95 figures were measured at S=16, so the claim is stability across datasets, not across sample counts — the document now says so.views-faoapi caught an overreach headed for a manager: "the upload can run from a machine Simon controls" is a procedure claim I was not entitled to make. The architecture supports it (
UPLOAD_ENABLEDdefaultsFalse, no store client constructed when disarmed) but there is one entrypoint, no--no-upload, and disarming means editing a committed delivery declaration. Restated as architecture plus an open question. They also caught that the ensemble and delivery steps are per-run and memory-bound, so the per-model extrapolation is silent about them — and that the two documents disagreed on how many models were complete.Where I pushed back
views-hydranet noted my ~4 h does not match fimbulthul's ~5.5 h. True, but not a correction: the guide documents rented pods, and 4 h is measured on that hardware, n=3. Kept my figure, added theirs as context.
Verification
views_hydranetabsent from the test env, synthetic targets, C-74)ruff check .cleanbash docs/validate_docs.shpassestests/test_tools_layout.pypasses with the new groupStill open
The post-mortem is
Status: DRAFT— five of eight models were running when this was written. If their timings move the average, both it and the cost note get reissued. Its §10 lists seven open items, including three issues I owe (thedisk_guardbug, the>=1.9.0floor still in the model requirements, and a register entry for plain HTTP from rented hardware).🤖 Generated with Claude Code
https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K