ci: validate every package on PRs, stop cancelling main runs, and add a global complexity ceiling - #1184
Merged
Conversation
smol-toml 1.9.0 (released 2026-09-22) builds every parsed table with
Object.create(null). The config loader treats only objects whose
prototype is Object.prototype as tables, and it sorts
Object.entries(table) with the default comparator, which stringifies each
value. A packed install has no lockfile and resolves the ^1.7.1 range to
1.9.0, so `ultrafuzz init` failed while loading the shipped defaults
("project must be a table; run must be a table; ..."). This is the
packed-install failure that turned the package-gates lane red in main run
36495460966.
Rebuild the parsed document with structuredClone, which returns the same
data as ordinary objects, so the loader sees what it saw with 1.7.
`__proto__` keys stay own data properties.
Co-Authored-By: Claude Opus 5.5 <[email protected]>
…el main runs Pull requests ran only the five runtime lanes. package-gates, cli and benchmark-history-typecheck (which holds the workspace typecheck) first ran after merge, so their regressions landed on main: the latest main run, 36495460966, fails package-gates. Every lane now runs on every event. The lane table keeps its test that it runs each gate scripts/validate-release.mjs defines exactly once; the pull-request filter, PULL_REQUEST_REQUIRED_GATES and the --event argument are gone. The lanes no longer wait for the build and super-linter jobs, and the PR-only runtime and CLI smoke runners are deleted because the lanes run those suites in full. release-gates still requires every job. Measured on existing runs: a pull request took about 60 minutes (8.0 min for build and smoke, then 51.5 min for the slowest lane in run 36279711018); on the latest main push the new PR lanes took 12.9 (package-gates), 41.5 (cli) and 9.8 minutes against 26.8-39.5 minutes for the runtime lanes. All pushes to main shared one concurrency group; 24 of the last 40 main push runs were cancelled (as of 2026-09-29). cancel-in-progress: false alone would not stop that, because GitHub also cancels a pending run in the group when a newer run queues, and main pushes arrive in bursts (three within 9 seconds on 2026-09-28). Runs other than pull requests now get their own group, keyed by github.run_id. Delete the ci.yml line-by-line assertions in packages/modal/test/ci-config.test.ts and the timeout_minutes === 120 change detector. The remaining guard parses ci.yml and checks that release validation is not gated off pull requests. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Every runtime and CLI test that launches a run copies a sealed execution snapshot and fsyncs each file, twice (at write and again when the snapshot is sealed). The deep analysis measured 55,998 fsyncs taking 222 s of a 163-261 s test locally; under eatmydata the same test took 23-33 s, and a CLI test went from 169 s to 35 s. A CI runner is discarded after the job, so durability across a crash buys nothing there. Install eatmydata in the release-validation job and run validate:release under it; its LD_PRELOAD reaches every child the gate commands spawn. Co-Authored-By: Claude Opus 5.5 <[email protected]>
`pnpm -w knip` leaves out knip's exports, types and duplicates categories, so unused exports accumulate unseen; on this tree knip reports 52 unused exports and 38 unused exported types. Report them in a continue-on-error step that reuses the pinned knip script. Several open pull requests delete dead code; once they land, these categories can move into the blocking knip step. Co-Authored-By: Claude Opus 5.5 <[email protected]>
config, topology, prompts, evals, evmbench and modal run a bare `vitest run` with no config file, so every test gets vitest's 5 s default. Main went red on it in runs 34967618488 (modal ci-config and modal-documents: "Test timed out in 5000ms") and 34634052856 (modal worker-result), and #1126 raised individual cases instead. Pass --testTimeout=30000 in the six package test scripts. Cases that set their own timeout keep it. Co-Authored-By: Claude Opus 5.5 <[email protected]>
cli.test.ts, lifecycle-commands.test.ts and audit-profile-commands.test.ts
created their project with mkdtemp and never removed it. On main, one run-
launching test ("status surfaces a terminal product and live workflow
lifecycle divergence") left 446 MB and 34,091 entries in TMPDIR; the sealed
snapshot directories are mode dr-x, so a plain `rm -rf` cannot delete them.
A CI lane accumulated every test's run directories until the job ended.
temporaryRoot(prefix, t) creates the canonical directory and registers a
t.after hook that restores owner write permission on the way down and
removes it, together with the `<root>-fake-bin` directory the fake engine
fixtures create beside it. The tests pass their context to tempProject.
Co-Authored-By: Claude Opus 5.5 <[email protected]>
…a recorded baseline The complexity and size rules ran only in the diff-limited strict lint, which reports a message only when its line changed. ESLint reports these rules on a function's first line (max-lines on the first line past the limit), so adding branches inside an over-budget function passed, while editing only its signature failed, forcing an unrelated refactor on a fix. Move complexity, max-depth, max-lines, max-lines-per-function, max-nested-callbacks, max-params and max-statements into the always-on config, with the same relaxed values for tests and templates, and record the 1,003 existing violations (197 files; 701 of them in 108 packages/*/src files) with ESLint's bulk suppressions in eslint-suppressions.json. `pnpm -w lint` now fails when a file gains a violation of a rule and when a recorded count is higher than the file's violations; `pnpm -w lint:prune` lowers the counts. The diff-limited strict lint keeps the type-aware strict rules, no-console and the TODO/FIXME check. It passes --pass-on-unpruned-suppressions because it sees only changed lines, so its counts are below the recorded ones by design. ESLint writes the suppressions file without a final newline, so prettier ignores it. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…elog Co-Authored-By: Claude Opus 5.5 <[email protected]>
Super-Linter lints each changed Markdown file in full with its default markdownlint rules, including MD013 at 400 characters. Four of the 109 tracked Markdown files have longer lines: CHANGELOG.md keeps each entry on one line (44 such lines), and docs/reference/agent-adapter-boundaries.md, docs/explanation/provider-harness-research.md and docs/CODE_OF_CONDUCT.md have long table rows or paragraphs. Any pull request that touched one of them failed `External static analysis` on lines it did not change, and no merged pull request has edited CHANGELOG.md since Super-Linter was added. Use Super-Linter's default markdownlint rules from the pinned v8.7.0 template, with MD013 off, as .github/linters/.yamllint.yml already does for YAML line length. markdownlint-cli 0.49.0 (the version Super-Linter runs) reports 46 MD013 errors on this branch's changed Markdown with the default template and none with this file. Co-Authored-By: Claude Opus 5.5 <[email protected]>
The smoke-isolation test asserted `scripts.test === "vitest run"`, so adding
--testTimeout=30000 failed package-gates in this pull request's first run. The
property the test protects, that the unit-test script never runs the real
Modal smoke, is still asserted by `not.toContain("smoke")` two lines later.
Co-Authored-By: Claude Opus 5.5 <[email protected]>
Co-Authored-By: Claude Opus 5.5 <[email protected]>
…e-bin sibling The per-test cleanup was only measured, so a regression would bring the temporary-directory leak back silently. The new test builds the shape a launched run leaves (read-only files in dr-x directories) plus the `<root>-fake-bin` sibling inside a subtest and asserts both are gone once the subtest ends. It fails if the helper skips restoring owner write permission (the hook throws EACCES) or skips the sibling. Co-Authored-By: Claude Opus 5.5 <[email protected]>
This was referenced Sep 29, 2026
Merged
Merged
New violations still fail because a file's count for a rule rises above the recorded count. Fixed violations leave a stale count, which lint:prune removes; failing on them would make every parallel pull request that deletes code rewrite eslint-suppressions.json and conflict with the others. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…fail it `pnpm -w lint` passes --pass-on-unpruned-suppressions, so a fixed violation whose count was not pruned is accepted. `lint:fix` repeated the file globs without that flag, so in exactly that state it applied its fixes and then exited 2 with "There are suppressions left that do not occur anymore". Defining it as `pnpm -w lint --fix` keeps the two scripts from drifting. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Since lint passes --pass-on-unpruned-suppressions, a recorded count only goes down when someone runs `pnpm -w lint:prune`. Until then the room a fixed violation leaves lets a new violation of the same rule in the same file pass. contributing.md, the eslint.config.js comment and the changelog said a new violation always fails, and contributing.md promised periodic pruning on main that nothing performs. They now state the actual rule: lint fails when a file has more violations of a rule than its recorded count. The changelog also claimed every CLI test deletes its temporary project. Only the tests in cli.test.ts, lifecycle-commands.test.ts and audit-profile-commands.test.ts do; those include every CLI test that launches a run, so the entry now says that. Co-Authored-By: Claude Opus 5.5 <[email protected]>
The informational knip step exited 1 whenever it found anything, which is every run until the dead-code removals land. continue-on-error kept the job green, but each Build gates check still carried a "Process completed with exit code 1." failure annotation. --no-exit-code (knip 6.33.0) prints the same report (52 unused exports, 38 types, 4 duplicates on this tree) and exits 0. continue-on-error stays so a knip crash in this step still cannot block. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…t covers The lane comment and docs/reference/development.md said eatmydata makes fsync a no-op for the whole lane. It does so in the test processes, which is where startRun and runCli copy and fsync the sealed snapshot, but smithersCommandEnv rebuilds each Smithers engine child's environment from an allowlist that has no LD_PRELOAD, so engine processes still sync. `sudo apt-get install -y eatmydata` without `apt-get update` only worked because the current image already ships the package. GitHub's notice on this PR's Build gates job says ubuntu-latest migrates to Ubuntu 26 from 2026-10-19; on an image without eatmydata and without package lists, that install would fail every lane. The step now skips apt entirely when eatmydata is present, as it is today, and otherwise updates the index before installing. Co-Authored-By: Claude Opus 5.5 <[email protected]>
The guard parsed ci.yml but checked only the release-validation job's own
`if`. Adding `if: github.event_name != 'pull_request'` to the "Validate release
lane" step, the step that runs the suites, still passed it; the text guard it
replaced caught that. It also pinned the exact `if` string of the release-gates
requirement step, so an equivalent rewrite failed it.
The guard now rejects an event or pull_request condition on the job and on
every one of its steps, and checks that the requirement step tests
`needs.release-validation.result` without an event condition. Mutations run
against the old and new test: a step-level PR `if` passes the old guard and
fails the new one; a job-level PR `if` and an event-gated requirement step
fail both; wrapping the requirement condition in `${{ }}` fails only the old.
Co-Authored-By: Claude Opus 5.5 <[email protected]>
…y ceiling #1005 deliberately added no baseline or suppression registry, and #960 asks for a complexity ceiling in the style of modem-dev/hunk#861: one global maximum set at today's worst function, with nothing grandfathered, lowered as hotspots are simplified. Drop eslint-suppressions.json and the always-on size budgets, keep the diff-limited strict budgets for changed lines, and fail any function above complexity 90 (synchronizeLinkedWorkflowRun is at 90 once the pending simplification PRs land). The changelog entry moves to the consolidated release notes. Co-Authored-By: Claude Opus 5.5 <[email protected]>
| // with no suppressions, so no function can grow past the worst one today; lower | ||
| // it as the most complex functions are simplified. Changed lines are held to the | ||
| // stricter budgets below by `pnpm -w lint:strict:ci`. | ||
| const complexityCeilingConfigs = [{ files: sourceFiles, rules: { complexity: ["error", 90] } }]; |
There was a problem hiding this comment.
Existing function exceeds new ceiling If
verifyCurrentCampaignTimeoutEvidence still has the 122-point complexity stated in the PR description, the new 90-point limit makes the required pnpm -w lint gate fail. That function is unchanged from the supplied base, so this PR cannot pass that gate against the base until the separate simplification lands or the ceiling is adjusted.
Prompt To Fix With AI
This is a comment left during a code review.
Path: eslint.config.js
Line: 16
Comment:
**Existing function exceeds new ceiling** If `verifyCurrentCampaignTimeoutEvidence` still has the 122-point complexity stated in the PR description, the new 90-point limit makes the required `pnpm -w lint` gate fail. That function is unchanged from the supplied base, so this PR cannot pass that gate against the base until the separate simplification lands or the ceiling is adjusted.
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
aviggiano
added a commit
that referenced
this pull request
Sep 30, 2026
…t helpers dynamic-runtime.ts is already over the strict lint's 500-line budget on main (668 counted lines), and ESLint reports max-lines only on the 501st counted line. lint:strict:ci keeps that report only when its line is one the branch changed. Before the rebase it fell on resolvedConfigForRuntimeRoot, 11 lines past this PR's hunk. #1198 adds lines above it (optional-input lowering), so on current main it falls inside renderRuntimePrompt, which this PR extracts, and the gate fails over a file size this PR did not create. Move renderRuntimePrompt, byte for byte, below resolvedConfigForRuntimeRoot and promptGraphContext, the helpers it calls or whose return type it takes. The report now falls inside promptGraphContext, which this PR does not change. This is the fix #1062 used for lifecycle-inspection.ts; no eslint-disable is added, since #1184 chose not to keep a suppression baseline. Co-Authored-By: Claude Opus 5.5 <[email protected]>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
package-gates(artifacts, config, topology, prompts, evals, modal, dashboard, security, references, packed install),cli, andbenchmark-history-typecheck(which holds the workspace typecheck) ran only on pushes tomain, so their regressions first failed after merge.maincancelled each other. 24 of the last 40mainpush runs were cancelled (as of 2026-09-29), so mostmaincommits never got a complete validation.mainis red inpackage-gates, and no pull request caused it. smol-toml 1.9.0, published 2026-09-22T17:15Z, breaks the config loader, and the packed-install consumer resolves dependencies without the workspace lockfile.package-gateslast passed onmainin run 35789754192 (2026-09-22) and has failed since run 36490897026 (2026-09-28), for example in 36495460966. Runningpackage-gateson pull requests would not have prevented this; it would have turned every open pull request red at once, which is why this PR fixes the loader first.mainwent red on it in runs 34967618488 (modalci-configandmodal-documents) and 34634052856 (modalworker-result).main, the single teststatus surfaces a terminal product and live workflow lifecycle divergenceleft 446 MB and 34,091 entries inTMPDIR.External static analysisfor this reason alone. This PR's first run hit the same 46 errors, on CHANGELOG.md and on existing 621-character table rows inagent-adapter-boundaries.md.Root cause
scripts/ci/release-validation-lanes.mjsfiltered lanes by event. Only the runtime lanes hadpull_request: true.Every push shared one concurrency group (
github.ref) withcancel-in-progress: true. Changing onlycancel-in-progresswould not fix it:mainpushes arrive in bursts, for example three within 9 seconds at 22:11 on 2026-09-28.ESLint reports
complexity,max-lines-per-functionandmax-statementson a function's first line (andmax-lineson the first line past the limit), and the diff plugin keeps a message only when its line changed. Two consequences:Packed-install failure. smol-toml 1.9.0 builds every parsed table with
Object.create(null). The config loader has two assumptions that break on such tables:isPlainObjectaccepts onlyObject.prototype.Object.entries(table).sort()stringifies values, which a null-prototype object cannot do.The packed-install consumer has no lockfile, so it resolves
^1.7.1to 1.9.0, andultrafuzz initfails withproject must be a table; run must be a table; .... The workspace lockfile still pins 1.7.1, which is why only the packed install caught it.Leaked run directories. The CLI tests'
tempProject()calledmkdtempSyncand nothing ever removed the directory. Sealed snapshot directories are modedr-x, so a plainrm -rfcannot delete them anyway.Change
One commit per concern:
fix(config): normalize TOML tables once, at the parse call.parseProjectConfigTomlwraps the parse instructuredClone(parse(text)), which returns the same data as ordinary objects. The rest of the loader is unchanged.__proto__keys stay own data properties (checked with 1.9.0 directly).ci: run every lane on every event.PULL_REQUEST_REQUIRED_GATESand no--eventargument. It keeps its test that it runs each gatescripts/validate-release.mjsdefines exactly once.release-validationneeds onlyrelease-validation-lanes, so the lanes start without waiting for the build and super-linter jobs.release-gatesstill requires every job.Build gates, and its budget drops from 30 to 15 minutes. It took 2.8-3.6 minutes on every run of this PR.github.run_id), so only a pull request's superseded runs are cancelled.ci.ymlline-by-line assertions inpackages/modal/test/ci-config.test.ts(185 lines) and thetimeout_minutes === 120change detector are deleted. The remaining guard parsesci.ymland checks that neither therelease-validationjob nor any of its steps has an event orpull_requestcondition, and thatrelease-gatesrequires the lanes' result without one.ci: runvalidate:releaseundereatmydata. The job installs it only when the image lacks it (command -v eatmydata || { sudo apt-get update && sudo apt-get install -y eatmydata; }); the current image ships eatmydata 131-1ubuntu1, so the step does not touch apt today.LD_PRELOADreaches the test processes, which is wherestartRunandrunClicopy and fsync the sealed snapshot (both run in-process in the runtime and CLI tests).smithersCommandEnv(packages/runtime/src/smithers.ts) builds their environment from an allowlist withoutLD_PRELOAD, so those still sync. I checked this by reading the code, not by tracing.ci: report unused exports without blocking.pnpm -w knip --include exports,types,duplicates --no-exit-codereports 52 unused exports, 38 unused exported types and 4 duplicate exports on this tree and exits 0. The first version exited 1 undercontinue-on-error, which put aProcess completed with exit code 1.failure annotation on every Build gates run; the run on this head has none.continue-on-errorstays so a knip crash in this step still cannot block. The step can become blocking after the dead-code PRs land.test: add--testTimeout=30000to thevitest runscripts of config, topology, prompts, evals, evmbench and modal. A follow-up commit deletes the modal assertionscripts.test === "vitest run", which failedpackage-gatesin this PR's first run. The property it guarded, that the unit-test script never runs the real Modal smoke, is still asserted bynot.toContain("smoke").test(cli): clean up per test.temporaryRoot(prefix, t)(inpackages/cli/test/temporary-root.ts) registers at.afterhook. The hook restores owner write permission and removes the root and the<root>-fake-bindirectory that the fake-engine fixtures create beside it. The tests incli.test.ts,lifecycle-commands.test.tsandaudit-profile-commands.test.tspass their context in; those files hold every CLI test that launches a run.temporary-root.test.tsbuilds the shape a launched run leaves (read-only files indr-xdirectories) plus the sibling, and asserts both are gone when the test ends.build(lint): a global complexity ceiling (Add cyclomatic complexity lint #960), no baseline.pnpm -w lintnow fails any function underpackages/orscripts/above cyclomatic complexity 90. That is today's maximum once the open simplification batch lands (synchronizeLinkedWorkflowRun; the 122-pointverifyCurrentCampaignTimeoutEvidenceis deleted by refactor(runtime): delete host re-implementations of registry artifact gates #1178). Nothing is grandfathered or suppressed, following chore: adopt external static-analysis gates #1005's no-baseline stance and the approach Add cyclomatic complexity lint #960 links to (chore(lint): enforce cyclomatic complexity ceiling modem-dev/hunk#861). Lower the ceiling ineslint.config.jsas the worst functions are simplified. Changed lines stay under the stricter diff-limited budgets (lint:strict:ci: complexity 20, size rules, type-aware strict rules). An earlier revision of this PR baselined 1,003 existing violations ineslint-suppressions.json; that was dropped because chore: adopt external static-analysis gates #1005 deliberately added no baseline or suppression registry.ci: turn off MD013 only..github/linters/.markdown-lint.ymlis super-linter's pinned v8.7.0 default markdownlint template with only MD013 turned off, the same trade-off.github/linters/.yamllint.ymlalready makes for YAML line length. Only 4 of the 109 tracked Markdown files have lines over 400 characters (CHANGELOG.md,agent-adapter-boundaries.md,provider-harness-research.md,CODE_OF_CONDUCT.md).docs: CHANGELOG entries.docs/reference/development.mdanddocs/reference/agent-adapter-boundaries.mdare updated for the lane change, anddevelopment.mdsays which processeseatmydatacovers.Deliberately not built (and why)
A tighter always-on budget. A ceiling of 90 enforces little by itself; the previous revision of this PR measured 1,003 violations of the 20/80/500 budgets. Tightening needs either a baseline (which chore: adopt external static-analysis gates #1005 ruled out) or refactoring the hotspots first, so the ceiling starts at the current maximum.
No per-event lane policy. With every lane on every event, a PR-required gate list has nothing to configure, so it is deleted rather than set to "all".
The lane matrix stays in
release-validation-lanes.mjs. With the event filter gone the script only prints a constant, and inlining it intoci.ymlwould delete the selection job and the script's entry point. test: model-free end-to-end campaign with controller kill and resume on the real pinned engine #1187 edits that script now, so that is a follow-up.Lane timeouts stay at 120 minutes. Runner variance is 2-4x between runs, so lowering them needs more evidence than this PR's runs.
The test: give Modal validation cases individual timeout budgets #1126 per-case vitest timeouts (15-30 s) are left alone. Some of them are in files other open PRs edit. They are now as tight as or tighter than the new default, and locally 5 modal cases timed out at exactly those budgets under load, so removing them is a reasonable follow-up.
packages/runtime/test/temporary-root.tskeeps removing its roots at process exit. Per-test cleanup there needs a check that no runtime test shares a root across tests. Because runtime shards still accumulate run directories, the step that deletes the Android and .NET SDKs to reclaim disk stays.The workspace lockfile stays on smol-toml 1.7.1. A dependency bump belongs in its own PR, and the fix works with both versions.
The snapshot
changed while readingfalse positive is not fixed here.workflow-integrity.tscomparesnlinkandctimeof the source files it copies, and those change whenever another process hard-links the same pnpm-store file. It can fail a real launch on a developer machine; it is a runtime fix in files other open items own.No further CI restructuring.
package-gatesstill repeats thedependency-advisoriesandci-scriptsgates that the build job already runs, and the network advisory check is still inside the build job. Both deserve their own change.Verification
Each fix was checked against unmodified
origin/main(b6dd1da9) and against this branch.origin/mainpackages/config/test/toml-null-prototype.test.ts, withsmol-toml'sparsemocked to return 1.9-shaped tablesfailed default config validationerror as CInode scripts/validate-packed-install.mjs(fresh registry resolution)packed init failed ... project must be a tablepacked install validation passed for 12 packages at 0.1.0scripts/ci/release-validation-lanes.test.tsrun against each lane scriptruntime-supporting,runtime-1..4)vitest run:Test timed out in 5000msTMPDIRleft behindTMPDIRleft by the 8-run testreport bundle --require-verified fails closed on every present invalid event journal, run aloneTMPDIRitself)packages/cli/test/temporary-root.test.tsis new, so it cannot run on main. It passes on this branch and fails under both of two mutations of the helper: skipping the permission restore makes the hook throwEACCES, and skipping the sibling fails its assertion.PR-gating guard, mutation-tested. I ran the previous and the current version of the guard in
release-validation-lanes.test.tsagainst mutated copies ofci.yml:ci.ymlif: github.event_name != 'pull_request'on the "Validate release lane" stepifon therelease-validationjobgithub.event_name != 'pull_request' &&added to therelease-gatesrequirement step${{ }}(equivalent)Other checks run on the final head:
pnpm -w lint,npx prettier --check .,CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:ci,pnpm -w knipandnode scripts/docs-check.mjspass.pnpm -w knip --include exports,types,duplicates --no-exit-codeexits 0 with the report; without--no-exit-codeit exits 1.ci.ymlis clean (without shellcheck installed locally; super-linter runs it in CI). markdownlint-cli 0.49.0 with.github/linters/.markdown-lint.ymlis clean on the changed Markdown files.bun test scripts/ci/release-validation-lanes.test.tspasses 3/3. Earlier in this PR: the modalci-config.test.tspasses 12/12,pnpm --filter @ultrafuzz/{config,modal} typecheckand the CLItsc -p tsconfig.test.jsonpass.pnpm -w test:ci-scriptspasses 95 of 96 locally. The failure isscripts/ci/safe-archive-bun.test.ts > repeatedly extracts a synthetic archive without stalling Bun callbacks, which hits its own 15 s budget at load average 14 on this shared machine. It fails the same way onorigin/mainhere, covers code this PR does not touch, and passes 96/96 in the Build gates job of every CI run of this PR.agent-adapter-boundaries.md. markdownlint-cli 0.49.0 reproduces those 46 errors locally with the pinned default template, and reports none with.github/linters/.markdown-lint.yml.node-provider.test.tsandpublic-worker.test.ts. The suite took 716 s here at load average 33-40, against 130 s on the CI runner, where those files passed. The one file I changed,smoke.test.ts, passes 14/14.CI effect, measured on this PR's six runs.
lint:fix, the knip and eatmydata steps, and the guard test. It passed every job; its Build gates job ran 96/96 CI script tests and carries no failure annotation, and in the shard 1 log theInstall eatmydatastep printed/usr/bin/eatmydataand skipped apt.Durations are for the
Validate release lanestep, in minutes. The baseline is the same lane in seven recent runs: main 36495460966 and 35789754192, and PR runs 36279711018, 36491801277, 36491760245, 36498046011 and 36497862407. The CLI and package lanes have only two baseline runs, because they never ran on PRs.eatmydataeffect on those lanes from runner variance.Local run of the three changed CLI test files (
TMPDIRisolated,eatmydata, 98 min at load average 30-46):TMPDIRheld 0 entries afterwards.WORKFLOW_SUBMISSION_FAILED: workflow execution file ... changed while reading, raised whileultrafuzz runcopied the sealed snapshot.packages/runtime/src/workflow-integrity.ts:2334comparesnlinkandctimeNsbefore and after the read.pnpm installs (other agents', and one of mine) add links to them.Risk / compatibility
More runner time. Pull requests use three more lane runners, and the lanes also run when lint fails. Wall time did not grow materially: 27.8-39.2 min across this PR's first five runs (run 6 queued for runners for 19 min), against 36.6-38.8 min for recent PR runs that ran five lanes. The CLI lane is now the critical path, so the next PR-latency win is the CLI suite itself (serial
--test-concurrency=1, and a full run launch per test).More concurrent runners. Pushes to
mainno longer queue behind each other, so bursts use more runners at once.Pull requests now depend on fresh registry resolution. The packed-install gate in
package-gatesinstalls without the lockfile. A newly published semver-compatible release that breaks the install, as smol-toml 1.9.0 did, now fails every open pull request at once instead of onlymain.Merge order with the open work-item PRs. From
git merge-treeand a lint of each merged tree, heads as of 2026-09-29 04:27 UTC; the heads keep moving:packages/cli/test/cli.test.ts, on test headers this PR changed toasync (t)/tempProject(t).packages/runtime/scripts/run-pr-smoke-tests.mjs, which this PR deletes.scripts/ci/release-validation-lanes.{mjs,test.ts},packages/cli/package.jsonanddocs/reference/development.md.tempProject()call withoutt, which will not compile once this merges; it needstwhen it rebases.mainthe maximum complexity is still 122, so its own lint only passes once refactor(runtime): delete host re-implementations of registry artifact gates #1178 has landed.eatmydata install. Today the step does not run apt. GitHub's notice on these jobs says
ubuntu-latestmoves to Ubuntu 26 from 2026-10-19; if that image lacks eatmydata, the step runsapt-get updateand installs it, which depends on the Ubuntu mirrors.The build job's name on PRs changed from
PR build and Node.js 24 runtime smoketoBuild gates.mainhas no branch protection, and the org ruleset has no required status checks (checked through the API), so nothing keys on the old name.No product runtime behaviour changes, apart from the config loader accepting smol-toml 1.9 tables.
Closes #960
Refs #923, Refs #997, Refs #1000
Refs #923rather thanCloses. The 30-minute budget in #923 no longer exists: #994 splitcliinto its own lane at 75 minutes, and #1137 raised it to 120. The measurements above showeatmydatadoes not materially shorten the CLI suite, and it remains the longest lane.🤖 Generated with Claude Code
The PR appears safe to merge; no outstanding finding or new actionable issue was identified.
Fix with agent prompt
Summary
The PR makes every release-validation lane run on pull requests, prevents main pushes from cancelling one another, and adds a global complexity ceiling. It also normalizes parsed TOML tables, adjusts test timeouts and temporary-directory cleanup, and updates CI linting and documentation.
Diagram
%%{init: {'theme': 'neutral'}}%% flowchart LR Select[Select release lanes] --> Lanes[Run every validation lane] Build[Build gates] --> Final[Release gates] Static[External static analysis] --> Final Select --> Final Lanes --> FinalReviews (8) · Last reviewed commit: "Merge origin/main into claude/w21-ci-qua..."