Skip to content

Proposal: run SmartOS and AIX in node-daily-master instead of node-test-pull-request #4457

Description

@codebytere

I pulled a month of node-test-pull-request history out of the Jenkins API (450 runs across 167 PRs, 2026-08-06 to 2026-09-05) to see where PR CI time goes. Most of it is waiting for an executor on two platforms rather than running tests, so i'd like to propose running SmartOS and AIX from node-daily-master only and dropping them from the PR job.

job runs median p90 of which waiting for an executor (median / p90) slowest configuration's own run (median)
node-test-commit-smartos 362 246 min 453 min ~215 / ~409 min 46 min (smartos23-x64)
node-test-commit-aix 407 152 min 334 min ~87 / ~261 min 69 min (aix73-power9)
node-test-commit-linux (10 configs) 421 94 min 305 min ~22 / ~194 min 62 min (rhel8-x64)
node-test-commit-osx 433 91 min 188 min ~5 / ~105 min 65 min (macos15-x64)
node-test-commit-linux-containered 435 66 min 154 min ~17 / ~91 min 43 min
node-test-commit-windows-fanned 428 58 min 192 min
node-test-commit-arm 431 52 min 112 min ~0 / ~12 min 45 min (rhel8-arm64)
node-test-commit-plinux 442 35 min 69 min ~0 / ~22 min 33 min
node-test-commit-linuxone 447 30 min 51 min ~0 / ~6 min 30 min
node-test-linter 445 15 min 20 min
node-test-pull-request, end to end 437 245 min 577 min

"Waiting for an executor" is the matrix job's duration minus its slowest configuration's execution time (the API doesn't expose queue time for matrix children directly); aborted runs are excluded. Over the same window 13% of PR runs finished SUCCESS and another 18% UNSTABLE, the average PR needed 2.7 runs, and when exactly one child job was non-green in a run it was linuxone 34 times (nearly all yellow from marked-flaky tests), aix 19, osx 17, linux 17, windows 9, smartos 4. Raw per-configuration numbers (exec median/p90, failure rate) are available if useful.

What the change would mean in practice: SmartOS and AIX stay Tier 2 exactly as BUILDING.md defines it (full test coverage maintained, failures block releases), keep running in node-daily-master and release CI, and failures there get filed the way other daily-only jobs are handled. PRs stop holding a jenkins-workspace executor for 2-3 h while a single smartos23 executor works through the queue, which by these numbers moves the median PR run from ~4 h to ~1.5 h (osx/linux become the long pole) with no new hardware, and takes some pressure off the executor deadlocks discussed in #4435. ppc64le and s390x are also Tier 2 but have no queueing to speak of, so i'd leave them in the PR job.

Mechanically i think this is either a parameter on node-test-commit that node-test-pull-request sets to skip those two phases, or a second multijob without them for the PR trigger; whichever the WG prefers. Two questions before this goes on an agenda: is there history on why these two gate PRs specifically rather than the daily, and would anyone want a lighter variant first (e.g. keep AIX and move only SmartOS, which is the larger share of the wait)?

Refs: #4435
Refs: nodejs/node#65826 (separate, cuts console log volume from the same jobs)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions