Skip to content

ci: validate every package on PRs, stop cancelling main runs, and add a global complexity ceiling - #1184

Merged
aviggiano merged 20 commits into
mainfrom
claude/w21-ci-quality
Sep 29, 2026
Merged

aviggiano merged 20 commits into
mainfrom
claude/w21-ci-quality

Conversation

@aviggiano

@aviggiano aviggiano commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

  • Pull requests skipped three of the eight release lanes. package-gates (artifacts, config, topology, prompts, evals, modal, dashboard, security, references, packed install), cli, and benchmark-history-typecheck (which holds the workspace typecheck) ran only on pushes to main, so their regressions first failed after merge.
  • Pushes to main cancelled each other. 24 of the last 40 main push runs were cancelled (as of 2026-09-29), so most main commits never got a complete validation.
  • main is red in package-gates, and no pull request caused it. smol-toml 1.9.0, published 2026-09-22T17:15Z, breaks the config loader, and the packed-install consumer resolves dependencies without the workspace lockfile. package-gates last passed on main in run 35789754192 (2026-09-22) and has failed since run 36490897026 (2026-09-28), for example in 36495460966. Running package-gates on pull requests would not have prevented this; it would have turned every open pull request red at once, which is why this PR fixes the loader first.
  • The complexity lint was both leaky and blocking (Add cyclomatic complexity lint #960). The complexity and size rules ran only in the changed-lines strict lint.
  • Vitest packages used the 5 s default timeout. main went red on it in runs 34967618488 (modal ci-config and modal-documents) and 34634052856 (modal worker-result).
  • CLI tests leaked their run directories. On main, the single test status surfaces a terminal product and live workflow lifecycle divergence left 446 MB and 34,091 entries in TMPDIR.
  • The lanes are dominated by fsync. The analysis profiled one runtime test: 222 s of about 260 s wall time was spent in 55,998 fsyncs.
  • Super-linter failed every pull request that edited CHANGELOG.md. It lints each changed Markdown file in full with MD013 at 400 characters, and CHANGELOG.md has many entries longer than that on a single line. No merged PR has edited CHANGELOG.md since super-linter arrived in chore: adopt external static-analysis gates #1005, and most of the open fix PRs fail External static analysis for this reason alone. This PR's first run hit the same 46 errors, on CHANGELOG.md and on existing 621-character table rows in agent-adapter-boundaries.md.

Root cause

  • scripts/ci/release-validation-lanes.mjs filtered lanes by event. Only the runtime lanes had pull_request: true.

  • Every push shared one concurrency group (github.ref) with cancel-in-progress: true. Changing only cancel-in-progress would not fix it:

    • GitHub also cancels a pending run in a group when a newer run queues.
    • main pushes arrive in bursts, for example three within 9 seconds at 22:11 on 2026-09-28.
  • ESLint reports complexity, max-lines-per-function and max-statements on a function's first line (and max-lines on the first line past the limit), and the diff plugin keeps a message only when its line changed. Two consequences:

    • Adding a branch inside a function that is already over budget passes.
    • Editing only that function's signature fails, which forces an unrelated refactor onto a fix.
  • Packed-install failure. smol-toml 1.9.0 builds every parsed table with Object.create(null). The config loader has two assumptions that break on such tables:

    • isPlainObject accepts only Object.prototype.
    • Object.entries(table).sort() stringifies values, which a null-prototype object cannot do.

    The packed-install consumer has no lockfile, so it resolves ^1.7.1 to 1.9.0, and ultrafuzz init fails with project must be a table; run must be a table; .... The workspace lockfile still pins 1.7.1, which is why only the packed install caught it.

  • Leaked run directories. The CLI tests' tempProject() called mkdtempSync and nothing ever removed the directory. Sealed snapshot directories are mode dr-x, so a plain rm -rf cannot delete them anyway.

Change

One commit per concern:

  1. fix(config): normalize TOML tables once, at the parse call. parseProjectConfigToml wraps the parse in structuredClone(parse(text)), which returns the same data as ordinary objects. The rest of the loader is unchanged. __proto__ keys stay own data properties (checked with 1.9.0 directly).
  2. ci: run every lane on every event.
    • The lane table has no event filter, no PULL_REQUEST_REQUIRED_GATES and no --event argument. It keeps its test that it runs each gate scripts/validate-release.mjs defines exactly once.
    • release-validation needs only release-validation-lanes, so the lanes start without waiting for the build and super-linter jobs. release-gates still requires every job.
    • The PR-only runtime and CLI smoke runners and their package scripts are deleted, because the lanes run those suites in full.
    • The build job is now always named Build gates, and its budget drops from 30 to 15 minutes. It took 2.8-3.6 minutes on every run of this PR.
    • Runs other than pull requests get their own concurrency group (github.run_id), so only a pull request's superseded runs are cancelled.
    • The ci.yml line-by-line assertions in packages/modal/test/ci-config.test.ts (185 lines) and the timeout_minutes === 120 change detector are deleted. The remaining guard parses ci.yml and checks that neither the release-validation job nor any of its steps has an event or pull_request condition, and that release-gates requires the lanes' result without one.
  3. ci: run validate:release under eatmydata. The job installs it only when the image lacks it (command -v eatmydata || { sudo apt-get update && sudo apt-get install -y eatmydata; }); the current image ships eatmydata 131-1ubuntu1, so the step does not touch apt today.
    • LD_PRELOAD reaches the test processes, which is where startRun and runCli copy and fsync the sealed snapshot (both run in-process in the runtime and CLI tests).
    • It does not reach the Smithers engine processes: smithersCommandEnv (packages/runtime/src/smithers.ts) builds their environment from an allowlist without LD_PRELOAD, so those still sync. I checked this by reading the code, not by tracing.
  4. ci: report unused exports without blocking. pnpm -w knip --include exports,types,duplicates --no-exit-code reports 52 unused exports, 38 unused exported types and 4 duplicate exports on this tree and exits 0. The first version exited 1 under continue-on-error, which put a Process completed with exit code 1. failure annotation on every Build gates run; the run on this head has none. continue-on-error stays so a knip crash in this step still cannot block. The step can become blocking after the dead-code PRs land.
  5. test: add --testTimeout=30000 to the vitest run scripts of config, topology, prompts, evals, evmbench and modal. A follow-up commit deletes the modal assertion scripts.test === "vitest run", which failed package-gates in this PR's first run. The property it guarded, that the unit-test script never runs the real Modal smoke, is still asserted by not.toContain("smoke").
  6. test(cli): clean up per test. temporaryRoot(prefix, t) (in packages/cli/test/temporary-root.ts) registers a t.after hook. The hook restores owner write permission and removes the root and the <root>-fake-bin directory that the fake-engine fixtures create beside it. The tests in cli.test.ts, lifecycle-commands.test.ts and audit-profile-commands.test.ts pass their context in; those files hold every CLI test that launches a run. temporary-root.test.ts builds the shape a launched run leaves (read-only files in dr-x directories) plus the sibling, and asserts both are gone when the test ends.
  7. build(lint): a global complexity ceiling (Add cyclomatic complexity lint #960), no baseline. pnpm -w lint now fails any function under packages/ or scripts/ above cyclomatic complexity 90. That is today's maximum once the open simplification batch lands (synchronizeLinkedWorkflowRun; the 122-point verifyCurrentCampaignTimeoutEvidence is deleted by refactor(runtime): delete host re-implementations of registry artifact gates #1178). Nothing is grandfathered or suppressed, following chore: adopt external static-analysis gates #1005's no-baseline stance and the approach Add cyclomatic complexity lint #960 links to (chore(lint): enforce cyclomatic complexity ceiling modem-dev/hunk#861). Lower the ceiling in eslint.config.js as the worst functions are simplified. Changed lines stay under the stricter diff-limited budgets (lint:strict:ci: complexity 20, size rules, type-aware strict rules). An earlier revision of this PR baselined 1,003 existing violations in eslint-suppressions.json; that was dropped because chore: adopt external static-analysis gates #1005 deliberately added no baseline or suppression registry.
  8. ci: turn off MD013 only. .github/linters/.markdown-lint.yml is super-linter's pinned v8.7.0 default markdownlint template with only MD013 turned off, the same trade-off .github/linters/.yamllint.yml already makes for YAML line length. Only 4 of the 109 tracked Markdown files have lines over 400 characters (CHANGELOG.md, agent-adapter-boundaries.md, provider-harness-research.md, CODE_OF_CONDUCT.md).
  9. docs: CHANGELOG entries. docs/reference/development.md and docs/reference/agent-adapter-boundaries.md are updated for the lane change, and development.md says which processes eatmydata covers.

Deliberately not built (and why)

  • A tighter always-on budget. A ceiling of 90 enforces little by itself; the previous revision of this PR measured 1,003 violations of the 20/80/500 budgets. Tightening needs either a baseline (which chore: adopt external static-analysis gates #1005 ruled out) or refactoring the hotspots first, so the ceiling starts at the current maximum.

  • No per-event lane policy. With every lane on every event, a PR-required gate list has nothing to configure, so it is deleted rather than set to "all".

  • The lane matrix stays in release-validation-lanes.mjs. With the event filter gone the script only prints a constant, and inlining it into ci.yml would delete the selection job and the script's entry point. test: model-free end-to-end campaign with controller kill and resume on the real pinned engine #1187 edits that script now, so that is a follow-up.

  • Lane timeouts stay at 120 minutes. Runner variance is 2-4x between runs, so lowering them needs more evidence than this PR's runs.

  • The test: give Modal validation cases individual timeout budgets #1126 per-case vitest timeouts (15-30 s) are left alone. Some of them are in files other open PRs edit. They are now as tight as or tighter than the new default, and locally 5 modal cases timed out at exactly those budgets under load, so removing them is a reasonable follow-up.

  • packages/runtime/test/temporary-root.ts keeps removing its roots at process exit. Per-test cleanup there needs a check that no runtime test shares a root across tests. Because runtime shards still accumulate run directories, the step that deletes the Android and .NET SDKs to reclaim disk stays.

  • The workspace lockfile stays on smol-toml 1.7.1. A dependency bump belongs in its own PR, and the fix works with both versions.

  • The snapshot changed while reading false positive is not fixed here. workflow-integrity.ts compares nlink and ctime of the source files it copies, and those change whenever another process hard-links the same pnpm-store file. It can fail a real launch on a developer machine; it is a runtime fix in files other open items own.

  • No further CI restructuring. package-gates still repeats the dependency-advisories and ci-scripts gates that the build job already runs, and the network advisory check is still inside the build job. Both deserve their own change.

Verification

Each fix was checked against unmodified origin/main (b6dd1da9) and against this branch.

Change On origin/main On this branch
Config fix: packages/config/test/toml-null-prototype.test.ts, with smol-toml's parse mocked to return 1.9-shaped tables Fails at import time with the same failed default config validation error as CI Passes; the full config suite passes 95/95
Config fix: node scripts/validate-packed-install.mjs (fresh registry resolution) Fails with packed init failed ... project must be a table packed install validation passed for 12 packages at 0.1.0
Lanes: scripts/ci/release-validation-lanes.test.ts run against each lane script 2 pass, 1 fail Passes
Lanes: matrix emitted for a pull request 5 lanes (runtime-supporting, runtime-1..4) 8 lanes, the same set for every event
Vitest timeout: throwaway 6 s test (not committed) Bare vitest run: Test timed out in 5000ms Passes under the new package script
CLI cleanup: TMPDIR left behind One test: 446 MB, 34,091 entries All tests of the three changed files: 0 entries
CLI cleanup: TMPDIR left by the 8-run test report bundle --require-verified fails closed on every present invalid event journal, run alone 3.5 GB 4 KB (the empty TMPDIR itself)

packages/cli/test/temporary-root.test.ts is new, so it cannot run on main. It passes on this branch and fails under both of two mutations of the helper: skipping the permission restore makes the hook throw EACCES, and skipping the sibling fails its assertion.

PR-gating guard, mutation-tested. I ran the previous and the current version of the guard in release-validation-lanes.test.ts against mutated copies of ci.yml:

Mutation of ci.yml Previous guard Current guard
None Passes Passes
if: github.event_name != 'pull_request' on the "Validate release lane" step Passes Fails
The same if on the release-validation job Fails Fails
github.event_name != 'pull_request' && added to the release-gates requirement step Fails Fails
The requirement step's condition wrapped in ${{ }} (equivalent) Fails Passes

Other checks run on the final head:

  • pnpm -w lint, npx prettier --check ., CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:ci, pnpm -w knip and node scripts/docs-check.mjs pass.
  • pnpm -w knip --include exports,types,duplicates --no-exit-code exits 0 with the report; without --no-exit-code it exits 1.
  • actionlint 1.7.12 on ci.yml is clean (without shellcheck installed locally; super-linter runs it in CI). markdownlint-cli 0.49.0 with .github/linters/.markdown-lint.yml is clean on the changed Markdown files.
  • bun test scripts/ci/release-validation-lanes.test.ts passes 3/3. Earlier in this PR: the modal ci-config.test.ts passes 12/12, pnpm --filter @ultrafuzz/{config,modal} typecheck and the CLI tsc -p tsconfig.test.json pass.
  • pnpm -w test:ci-scripts passes 95 of 96 locally. The failure is scripts/ci/safe-archive-bun.test.ts > repeatedly extracts a synthetic archive without stalling Bun callbacks, which hits its own 15 s budget at load average 14 on this shared machine. It fails the same way on origin/main here, covers code this PR does not touch, and passes 96/96 in the Build gates job of every CI run of this PR.
  • Super-linter. The first CI run failed only on MD013: 46 errors, all on existing lines of CHANGELOG.md and agent-adapter-boundaries.md. markdownlint-cli 0.49.0 reproduces those 46 errors locally with the pinned default template, and reports none with .github/linters/.markdown-lint.yml.
  • Full modal suite, run locally. 5 of 677 tests timed out at their own 15 s and 30 s per-case budgets, in node-provider.test.ts and public-worker.test.ts. The suite took 716 s here at load average 33-40, against 130 s on the CI runner, where those files passed. The one file I changed, smoke.test.ts, passes 14/14.
  • Locally I ran the full config and modal suites and the three changed CLI files, but not the full runtime or CLI suites, nor the other packages' suites. This PR's CI runs are the full runs.

CI effect, measured on this PR's six runs.

  • Run 1 is 36503175731, which failed only the modal script-string assertion and super-linter, both fixed since.
  • Runs 2-5 are 36506450057, 36510165058, 36512609711 and 36516391013. Each passed every job.
  • Run 6 is 36520602838, on the final head. Its commits change only text, lint:fix, the knip and eatmydata steps, and the guard test. It passed every job; its Build gates job ran 96/96 CI script tests and carries no failure annotation, and in the shard 1 log the Install eatmydata step printed /usr/bin/eatmydata and skipped apt.

Durations are for the Validate release lane step, in minutes. The baseline is the same lane in seven recent runs: main 36495460966 and 35789754192, and PR runs 36279711018, 36491801277, 36491760245, 36498046011 and 36497862407. The CLI and package lanes have only two baseline runs, because they never ran on PRs.

Lane Run 1 Run 2 Run 3 Run 4 Run 5 Run 6 Baseline range (median)
runtime shard 1 15.9 19.0 18.5 21.9 22.1 18.6 24.4-34.5 (28.1)
runtime shard 2 14.9 20.9 20.9 22.0 18.7 17.9 26.3-29.2 (27.1)
runtime shard 3 13.0 12.6 19.9 19.9 14.9 14.1 22.2-27.4 (26.6)
runtime shard 4 22.3 15.1 21.6 21.3 12.9 21.7 20.5-35.4 (27.7)
runtime supporting 19.4 19.2 18.9 14.8 19.0 18.6 21.5-44.7 (23.9)
CLI 36.6 37.2 25.3 36.1 36.2 36.4 34.7, 40.4
package gates 9.1 7.0 6.7 11.8 11.7 8.0 11.3, 11.6
benchmark history and typecheck 1.9 2.2 2.9 3.0 2.9 2.9 2.0, 2.6
  • The runtime lanes got faster consistently.
    • Shards 1-3 came in below every baseline run in all six runs, 19-53% under their medians.
    • Shard 4 was inside its range four times and below it twice.
    • The supporting lane was 19-38% under its median, below every baseline run.
  • The CLI and package lanes are mixed.
    • The CLI lane was inside its two-run baseline five times and well below it once. Its per-test sum was 9-10% below main's in runs 1 and 2 (39.6 and 40.3 min, against 44.1 on main 36495460966), and 38% below in run 3.
    • With two baseline runs, I cannot separate an eatmydata effect on those lanes from runner variance.
  • PR wall time was 39.2, 39.1, 27.8, 37.7 and 37.8 min for runs 1-5, with all eight lanes. Run 6 took 58.0 min end to end, but its lane-selection job waited 19 min for a runner (created 04:12:38, started 04:31:40); from the lanes' start to the end it took 38.4 min. The four most recent successful PR runs before this PR took 36.6-38.8 min, and ran only the five runtime lanes after the 8-minute build-and-smoke prerequisite. The CLI lane is the critical path.

Local run of the three changed CLI test files (TMPDIR isolated, eatmydata, 98 min at load average 30-46):

  • 98 of 101 tests and subtests passed, and TMPDIR held 0 entries afterwards.
  • The 3 failures were all WORKFLOW_SUBMISSION_FAILED: workflow execution file ... changed while reading, raised while ultrafuzz run copied the sealed snapshot.
    • The check at packages/runtime/src/workflow-integrity.ts:2334 compares nlink and ctimeNs before and after the read.
    • The source files are pnpm-store hard links shared by every worktree on this machine, and concurrent pnpm installs (other agents', and one of mine) add links to them.
    • The same test passed on main and on this branch in CI, and failed a different subtest each time it was re-run here, so I don't attribute it to this change. It is a real launch-time false positive; see "Deliberately not built".

Risk / compatibility

Closes #960
Refs #923, Refs #997, Refs #1000

Refs #923 rather than Closes. The 30-minute budget in #923 no longer exists: #994 split cli into its own lane at 75 minutes, and #1137 raised it to 120. The measurements above show eatmydata does not materially shorten the CLI suite, and it remains the longest lane.

🤖 Generated with Claude Code

RetriggerConfidence Score: 5/5

The PR appears safe to merge; no outstanding finding or new actionable issue was identified.

Fix All in Claude CodeFindings

  1. P1 Existing function exceeds new ceiling ▶
Fix with agent prompt
### Issue 1
eslint.config.js:undefined-16
If `verifyCurrentCampaignTimeoutEvidence` still has the 122-point complexity stated in the PR description, the new 90-point limit makes the required `pnpm -w lint` gate fail. That function is unchanged from the supplied base, so this PR cannot pass that gate against the base until the separate simplification lands or the ceiling is adjusted.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Summary

The PR makes every release-validation lane run on pull requests, prevents main pushes from cancelling one another, and adds a global complexity ceiling. It also normalizes parsed TOML tables, adjusts test timeouts and temporary-directory cleanup, and updates CI linting and documentation.

  • Release lanes now start independently of build and static-analysis jobs, while the final gate still requires their results.
  • CLI test cleanup covers sealed run directories and sibling fake-bin directories.
  • The prior complexity-ceiling concern no longer applies: the function identified in that thread is absent from the current tree.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  Select[Select release lanes] --> Lanes[Run every validation lane]
  Build[Build gates] --> Final[Release gates]
  Static[External static analysis] --> Final
  Select --> Final
  Lanes --> Final
Loading

Reviews (8) · Last reviewed commit: "Merge origin/main into claude/w21-ci-qua..."

aviggiano and others added 8 commits September 29, 2026 00:02
smol-toml 1.9.0 (released 2026-09-22) builds every parsed table with
Object.create(null). The config loader treats only objects whose
prototype is Object.prototype as tables, and it sorts
Object.entries(table) with the default comparator, which stringifies each
value. A packed install has no lockfile and resolves the ^1.7.1 range to
1.9.0, so `ultrafuzz init` failed while loading the shipped defaults
("project must be a table; run must be a table; ..."). This is the
packed-install failure that turned the package-gates lane red in main run
36495460966.

Rebuild the parsed document with structuredClone, which returns the same
data as ordinary objects, so the loader sees what it saw with 1.7.
`__proto__` keys stay own data properties.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…el main runs

Pull requests ran only the five runtime lanes. package-gates, cli and
benchmark-history-typecheck (which holds the workspace typecheck) first ran
after merge, so their regressions landed on main: the latest main run,
36495460966, fails package-gates. Every lane now runs on every event. The lane
table keeps its test that it runs each gate scripts/validate-release.mjs
defines exactly once; the pull-request filter, PULL_REQUEST_REQUIRED_GATES and
the --event argument are gone.

The lanes no longer wait for the build and super-linter jobs, and the PR-only
runtime and CLI smoke runners are deleted because the lanes run those suites in
full. release-gates still requires every job. Measured on existing runs: a pull
request took about 60 minutes (8.0 min for build and smoke, then 51.5 min for
the slowest lane in run 36279711018); on the latest main push the new PR lanes
took 12.9 (package-gates), 41.5 (cli) and 9.8 minutes against 26.8-39.5 minutes
for the runtime lanes.

All pushes to main shared one concurrency group; 24 of the last 40 main push
runs were cancelled (as of 2026-09-29). cancel-in-progress: false alone would
not stop that, because GitHub also cancels a pending run in the group when a
newer run queues, and main pushes arrive in bursts (three within 9 seconds on
2026-09-28). Runs other than pull requests now get their own group, keyed by
github.run_id.

Delete the ci.yml line-by-line assertions in packages/modal/test/ci-config.test.ts
and the timeout_minutes === 120 change detector. The remaining guard parses
ci.yml and checks that release validation is not gated off pull requests.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Every runtime and CLI test that launches a run copies a sealed execution
snapshot and fsyncs each file, twice (at write and again when the snapshot is
sealed). The deep analysis measured 55,998 fsyncs taking 222 s of a 163-261 s
test locally; under eatmydata the same test took 23-33 s, and a CLI test went
from 169 s to 35 s. A CI runner is discarded after the job, so durability
across a crash buys nothing there.

Install eatmydata in the release-validation job and run validate:release
under it; its LD_PRELOAD reaches every child the gate commands spawn.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
`pnpm -w knip` leaves out knip's exports, types and duplicates categories, so
unused exports accumulate unseen; on this tree knip reports 52 unused exports
and 38 unused exported types. Report them in a continue-on-error step that
reuses the pinned knip script. Several open pull requests delete dead code;
once they land, these categories can move into the blocking knip step.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
config, topology, prompts, evals, evmbench and modal run a bare `vitest run`
with no config file, so every test gets vitest's 5 s default. Main went red on
it in runs 34967618488 (modal ci-config and modal-documents: "Test timed out in
5000ms") and 34634052856 (modal worker-result), and #1126 raised individual
cases instead. Pass --testTimeout=30000 in the six package test scripts. Cases
that set their own timeout keep it.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
cli.test.ts, lifecycle-commands.test.ts and audit-profile-commands.test.ts
created their project with mkdtemp and never removed it. On main, one run-
launching test ("status surfaces a terminal product and live workflow
lifecycle divergence") left 446 MB and 34,091 entries in TMPDIR; the sealed
snapshot directories are mode dr-x, so a plain `rm -rf` cannot delete them.
A CI lane accumulated every test's run directories until the job ended.

temporaryRoot(prefix, t) creates the canonical directory and registers a
t.after hook that restores owner write permission on the way down and
removes it, together with the `<root>-fake-bin` directory the fake engine
fixtures create beside it. The tests pass their context to tempProject.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…a recorded baseline

The complexity and size rules ran only in the diff-limited strict lint, which
reports a message only when its line changed. ESLint reports these rules on a
function's first line (max-lines on the first line past the limit), so adding
branches inside an over-budget function passed, while editing only its
signature failed, forcing an unrelated refactor on a fix.

Move complexity, max-depth, max-lines, max-lines-per-function,
max-nested-callbacks, max-params and max-statements into the always-on config,
with the same relaxed values for tests and templates, and record the 1,003
existing violations (197 files; 701 of them in 108 packages/*/src files) with
ESLint's bulk suppressions in eslint-suppressions.json. `pnpm -w lint` now
fails when a file gains a violation of a rule and when a recorded count is
higher than the file's violations; `pnpm -w lint:prune` lowers the counts.

The diff-limited strict lint keeps the type-aware strict rules, no-console and
the TODO/FIXME check. It passes --pass-on-unpruned-suppressions because it
sees only changed lines, so its counts are below the recorded ones by design.
ESLint writes the suppressions file without a final newline, so prettier
ignores it.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano requested a review from a team as a code owner September 29, 2026 00:27
Comment thread packages/cli/test/temporary-root.ts
aviggiano and others added 4 commits September 29, 2026 00:32
Super-Linter lints each changed Markdown file in full with its default
markdownlint rules, including MD013 at 400 characters. Four of the 109
tracked Markdown files have longer lines: CHANGELOG.md keeps each entry on
one line (44 such lines), and docs/reference/agent-adapter-boundaries.md,
docs/explanation/provider-harness-research.md and docs/CODE_OF_CONDUCT.md
have long table rows or paragraphs. Any pull request that touched one of
them failed `External static analysis` on lines it did not change, and no
merged pull request has edited CHANGELOG.md since Super-Linter was added.

Use Super-Linter's default markdownlint rules from the pinned v8.7.0
template, with MD013 off, as .github/linters/.yamllint.yml already does for
YAML line length. markdownlint-cli 0.49.0 (the version Super-Linter runs)
reports 46 MD013 errors on this branch's changed Markdown with the default
template and none with this file.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
The smoke-isolation test asserted `scripts.test === "vitest run"`, so adding
--testTimeout=30000 failed package-gates in this pull request's first run. The
property the test protects, that the unit-test script never runs the real
Modal smoke, is still asserted by `not.toContain("smoke")` two lines later.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…e-bin sibling

The per-test cleanup was only measured, so a regression would bring the
temporary-directory leak back silently. The new test builds the shape a
launched run leaves (read-only files in dr-x directories) plus the
`<root>-fake-bin` sibling inside a subtest and asserts both are gone once the
subtest ends. It fails if the helper skips restoring owner write permission
(the hook throws EACCES) or skips the sibling.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
New violations still fail because a file's count for a rule rises above the
recorded count. Fixed violations leave a stale count, which lint:prune removes;
failing on them would make every parallel pull request that deletes code
rewrite eslint-suppressions.json and conflict with the others.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Comment thread package.json Outdated
aviggiano and others added 5 commits September 29, 2026 04:10
…fail it

`pnpm -w lint` passes --pass-on-unpruned-suppressions, so a fixed violation
whose count was not pruned is accepted. `lint:fix` repeated the file globs
without that flag, so in exactly that state it applied its fixes and then
exited 2 with "There are suppressions left that do not occur anymore".
Defining it as `pnpm -w lint --fix` keeps the two scripts from drifting.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Since lint passes --pass-on-unpruned-suppressions, a recorded count only goes
down when someone runs `pnpm -w lint:prune`. Until then the room a fixed
violation leaves lets a new violation of the same rule in the same file pass.
contributing.md, the eslint.config.js comment and the changelog said a new
violation always fails, and contributing.md promised periodic pruning on
main that nothing performs. They now state the actual rule: lint fails when
a file has more violations of a rule than its recorded count.

The changelog also claimed every CLI test deletes its temporary project. Only
the tests in cli.test.ts, lifecycle-commands.test.ts and
audit-profile-commands.test.ts do; those include every CLI test that launches
a run, so the entry now says that.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
The informational knip step exited 1 whenever it found anything, which is
every run until the dead-code removals land. continue-on-error kept the job
green, but each Build gates check still carried a "Process completed with
exit code 1." failure annotation. --no-exit-code (knip 6.33.0) prints the same
report (52 unused exports, 38 types, 4 duplicates on this tree) and exits 0.
continue-on-error stays so a knip crash in this step still cannot block.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…t covers

The lane comment and docs/reference/development.md said eatmydata makes fsync
a no-op for the whole lane. It does so in the test processes, which is where
startRun and runCli copy and fsync the sealed snapshot, but
smithersCommandEnv rebuilds each Smithers engine child's environment from an
allowlist that has no LD_PRELOAD, so engine processes still sync.

`sudo apt-get install -y eatmydata` without `apt-get update` only worked
because the current image already ships the package. GitHub's notice on this
PR's Build gates job says ubuntu-latest migrates to Ubuntu 26 from
2026-10-19; on an image without eatmydata and without package lists, that
install would fail every lane. The step now skips apt entirely when eatmydata
is present, as it is today, and otherwise updates the index before installing.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
The guard parsed ci.yml but checked only the release-validation job's own
`if`. Adding `if: github.event_name != 'pull_request'` to the "Validate release
lane" step, the step that runs the suites, still passed it; the text guard it
replaced caught that. It also pinned the exact `if` string of the release-gates
requirement step, so an equivalent rewrite failed it.

The guard now rejects an event or pull_request condition on the job and on
every one of its steps, and checks that the requirement step tests
`needs.release-validation.result` without an event condition. Mutations run
against the old and new test: a step-level PR `if` passes the old guard and
fails the new one; a job-level PR `if` and an event-gated requirement step
fail both; wrapping the requirement condition in `${{ }}` fails only the old.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…y ceiling

#1005 deliberately added no baseline or suppression registry, and #960 asks for
a complexity ceiling in the style of modem-dev/hunk#861: one global maximum set
at today's worst function, with nothing grandfathered, lowered as hotspots are
simplified. Drop eslint-suppressions.json and the always-on size budgets, keep
the diff-limited strict budgets for changed lines, and fail any function above
complexity 90 (synchronizeLinkedWorkflowRun is at 90 once the pending
simplification PRs land). The changelog entry moves to the consolidated
release notes.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Comment thread eslint.config.js
// with no suppressions, so no function can grow past the worst one today; lower
// it as the most complex functions are simplified. Changed lines are held to the
// stricter budgets below by `pnpm -w lint:strict:ci`.
const complexityCeilingConfigs = [{ files: sourceFiles, rules: { complexity: ["error", 90] } }];

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Existing function exceeds new ceiling If verifyCurrentCampaignTimeoutEvidence still has the 122-point complexity stated in the PR description, the new 90-point limit makes the required pnpm -w lint gate fail. That function is unchanged from the supplied base, so this PR cannot pass that gate against the base until the separate simplification lands or the ceiling is adjusted.

Prompt To Fix With AI
This is a comment left during a code review.
Path: eslint.config.js
Line: 16

Comment:
**Existing function exceeds new ceiling** If `verifyCurrentCampaignTimeoutEvidence` still has the 122-point complexity stated in the PR description, the new 90-point limit makes the required `pnpm -w lint` gate fail. That function is unchanged from the supplied base, so this PR cannot pass that gate against the base until the separate simplification lands or the ceiling is adjusted.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Fix in Claude Code

@aviggiano
aviggiano merged commit 625451c into main Sep 29, 2026
1 of 5 checks passed
@aviggiano
aviggiano deleted the claude/w21-ci-quality branch September 29, 2026 06:39
@aviggiano aviggiano changed the title ci: validate every package on PRs, stop cancelling main runs, and ratchet complexity with a lint baseline ci: validate every package on PRs, stop cancelling main runs, and add a global complexity ceiling Sep 29, 2026
aviggiano added a commit that referenced this pull request Sep 30, 2026
…t helpers

dynamic-runtime.ts is already over the strict lint's 500-line budget on
main (668 counted lines), and ESLint reports max-lines only on the 501st
counted line. lint:strict:ci keeps that report only when its line is one
the branch changed. Before the rebase it fell on
resolvedConfigForRuntimeRoot, 11 lines past this PR's hunk. #1198 adds
lines above it (optional-input lowering), so on current main it falls
inside renderRuntimePrompt, which this PR extracts, and the gate fails
over a file size this PR did not create.

Move renderRuntimePrompt, byte for byte, below resolvedConfigForRuntimeRoot
and promptGraphContext, the helpers it calls or whose return type it
takes. The report now falls inside promptGraphContext, which this PR does
not change. This is the fix #1062 used for lifecycle-inspection.ts; no
eslint-disable is added, since #1184 chose not to keep a suppression
baseline.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add cyclomatic complexity lint

1 participant