Repository navigation
test(e2e): the zh-Hant descriptive-copy wrapping test slows down on CI and ended a browser session #305
Description
Activity
Two further data points, both after the failure in the table above.
- The re-run on the pull request, same code. Commit
3d75404, run 36889099844, attempt 2. No code and no test was edited; only the failed job,e2e (shard 1/4), was re-run, once. Thezh-Hantcase took 4.9 s. - The first run on
mainafter the merge, no re-run. Commit0fbad79, run 36901166882, first attempt. Thezh-Hantcase took 5.0 s. - In both runs shard 1 finished with 403 passed.
Set against the earlier 3.0 s, 7.9 s, 18.5 s, 31.2 s sequence, this looks like large variability between CI runs rather than a steady slowdown: on the same commit the same test took 31.2 s in one attempt and 4.9 s in the next.
A pass at about 5 s is not evidence that the instability is gone. The cause is still not known. The timeout is not raised before the investigation, and this issue stays open.
- The re-run on the pull request, same code. Commit
One more observation, no action taken: in the CI run of #314 at
c145fee(run 37167130970, shard 1),i18n-fr.spec.ts"the other shipped locales are unaffected › zh-TW still reaches zh-Hant" failed its first attempt after 10.7 s withlocator('.toolbar')not visible within 8 s onopenApp, then passed on the retry in 1.3 s. Same shape as the zh-Hant cases above: a zh-Hant page that does not reach the toolbar in time on the runner, then an ordinary pass. The change under test touched no product code between that commit and the previously green one, and the main run after the merge (run 37168149050) passed the test on its first attempt.Another observation, no action taken: on
mainat938aae3(run 37190315128),playback-visual.spec.ts"travel · L0 elision" in the phone project failed its first attempt after 9.5 s withlocator('.toolbar')not visible within 8 s onopenApp, then passed on the retry in 4.3 s. The same tree had passed it on the first attempt in the pull request run (37188245053). The failure had the same visible symptom as the two earlier observations—a toolbar did not become visible within 8 seconds—but it occurred in a non-locale playback test, so this observation does not establish a shared cause.Another observation, no action taken: on
mainat70a9c37(run 37714159715, attempt 1),frame-label-locale-switch.spec.ts:142"every hyphenated locale is a door that opens both ways" in the chromium project failed its first attempt after 13.7 s: aftersetLocale('zh-Hant'), the document language stayedenfor the whole 8 s wait, then the test passed on the retry in 7.5 s. The same tree had passed it with no retry in the pull request run (37712395950, PR #336).This failure is in the locale-switch wait, not in the
waitForAppReadypath whose diagnostic report this issue is waiting for, so no report was produced. It may be a related symptom; this observation does not establish a shared cause.Two more observations, no action taken: on
mainatc4640bc(the v0.22.0 merge of PR #341; run 37800583718, attempt 1), two chromium tests failed their first attempt and passed on the retry. The run's other jobs and tests passed. The pull request run on the same tree (37797891345, head3d310d1) ran both tests with no retry, and PR #341 changed neither spec nore2e/support/.toolbar-locale-width.spec.ts:72"zh-Hans: the Tier-1 toolbar row is no worse than English, and clips nothing" (shard 2): the first attempt started at 15:34:28 UTC and failed after 10.8 s in the locale-switch wait: after switching to zh-Hans, the document language stayedenfor the whole 8 s. The retry, on another worker, passed in 1.3 s.descriptive-copy-wrapping.spec.ts:98"CJK descriptive copy wraps instead of overflowing › zh-Hant: no menu blurb or palette tooltip overflows" (shard 5): the first attempt started at 15:27:22 UTC and failed after 14.3 s inwaitForAppReady("app not ready after 8003 ms, waiting for .toolbar"). The retry, on another worker, passed in 7.0 s.
The second case is the first natural
waitForAppReadyfailure since the diagnostic report was added. Itsapp-ready-diagnosticsattachment reads, in full:app not ready after 8003 ms, waiting for .toolbar stage: renderer did not answer page: open; renderer did not answer within 1000 ms url: http://localhost:5173/ document: DOMContentLoaded yes, load yes, readyState unknown dom: #root children unknown, storage gate unknown, html lang "?" dir "?" toolbar: unknown; canvas: unknown page errors: none console errors: none failed requests: none 4xx/5xx responses: noneSo at the 8 s mark the page had fired DOMContentLoaded and load, but its renderer did not answer the report's own query within 1,000 ms, and no page error, console error or failed request was recorded. This records what was observed; it does not establish a cause, or a cause shared with the locale-switch case above.
One more observation, no action taken: on
mainat10098d6(the v0.24.0 merge of PR #343; run 37885887869, attempt 1), one mobile test failed its first attempt and passed on the retry. The run's other jobs and tests passed. The pull request run on the same tree (37884186416, head4e17549) ran it with no retry, and PR #343 changed neither that spec nore2e/support/.playback-a11y-background.spec.ts:280"playback — Slice 3c: run leaves the document untouched › playing then pausing then resetting a run leaves the GraphDoc digest and undo state untouched" (mobile, shard 2): the first attempt started at 05:08:07 UTC and failed after 9.3 s inwaitForAppReady("app not ready after 8060 ms, waiting for .toolbar"). The retry, on another worker, passed in 6.3 s.
Its
app-ready-diagnosticsreport reads, in full:app not ready after 8060 ms, waiting for .toolbar stage: toolbar and canvas visible page: open; renderer answered in 676 ms url: http://localhost:5173/ document: DOMContentLoaded yes, load yes, readyState complete dom: #root children 2, storage gate no, html lang "en" dir "ltr" toolbar: present yes, visible yes; canvas: present yes, visible yes page errors: none console errors: none failed requests: none 4xx/5xx responses: noneUnlike the
c4640bccase, the renderer answered (in 676 ms) and, when the report was taken, the toolbar and the canvas were both present and visible; the wait for.toolbarhad still timed out at 8,060 ms. No page error, console error or failed request was recorded. This records what was observed; it does not establish a cause, or a cause shared with the earlier cases.Observation from PR #345 CI, run 38032960035 (head
30bf3d4, attempt 1), jobe2e (shard 3/5), recorded as evidence only.- Test:
e2e/i18n-zh-hans.spec.ts:91— "a Traditional-Chinese browser › is never quietly given Simplified Chinese" (chromium). - First attempt: failed in
openApp→waitForAppReadyafter 8004 ms, waiting for.toolbar(stage "toolbar and canvas visible"). - Retry ci: GitHub Actions — verification only (no deploy) #1: passed (1.7 s). The shard reported 362 passed and 1 flaky.
- Diagnostics at the timeout: renderer answered in 403 ms; DOMContentLoaded and load done,
readyStatecomplete;html lang="zh-Hant",dir="ltr"; no storage gate; toolbar present and visible, canvas present and visible; page errors none, console errors none, failed requests none, 4xx/5xx responses none.
This is consistent with the existing zh-Hant symptom recorded here; no cause is established by this run.
- Test:
A second observation from PR #345 CI, following #305 (comment), recorded as evidence only: run 38034892101 (head
1b2824d, attempt 1), jobe2e (shard 2/6).- Test:
e2e/whats-new.spec.ts:1002— "every shipped language carries the notice and every release note › zh-Hant: the sentence names the version, and the panel has text for every item" (chromium). - First attempt: failed in
boot→waitForAppReadyafter 8003 ms, waiting for.toolbar, at the stage "renderer did not answer" (no answer within 1000 ms); page errors none, console errors none, failed requests none, 4xx/5xx responses none. - Retry ci: GitHub Actions — verification only (no deploy) #1: passed (1.6 s). The shard reported 1 flaky.
- The same test passed on its first attempt in the other 17 languages (zh-Hans immediately before zh-Hant, and the 14 after it).
No cause is established by this run.
- Test:
Diagnosis, part 1: the boot path and a local cold-start reproduction (no change made)
Work branch
fix/app-ready-305frommain24d0d91. Nothing in the app, the tests or CI is changed yet.What "ready" waits for today
waitForAppReady(e2e/support/appReady.ts) waits for.toolbarto be visible, then.canvas, each with the 8 s expect budget, afterpage.goto('/')resolved onload. The boot behind it: the boot module opens the storage session,startAppawaitsinitI18n()(the locale's UI catalog and its template-label dictionary, both dynamic imports), sets<html lang>, and only then mounts React. The diagnostic report is taken AFTER the deadline, so "toolbar visible" in a report means the toolbar was there a moment after the 8 s ran out.The CI evidence side by side
Run Test Ran right after Report at the deadline Retry 38032960035 ( 30bf3d4)i18n-zh-hans.spec.ts:91(zh-TW browser):82"a bare zh browser" (zh-Hans), 1.2 srenderer answered in 403 ms; toolbar and canvas visible; lang="zh-Hant"1.7 s 38034892101 ( 1b2824d)whats-new.spec.ts:1002zh-Hantthe zh-Hans case of the same test, 1.2 s renderer did not answer within 1000 ms 1.6 s 37800583718 ( c4640bc, v0.22.0)descriptive-copy-wrapping.spec.ts:98zh-Hant— renderer did not answer within 1000 ms 7.0 s 37885887869 ( 10098d6, v0.24.0)playback-a11y-background.spec.ts:280(mobile, English)— renderer answered in 676 ms; toolbar and canvas visible; lang="en"6.3 s In both PR #345 cases the tests just before and after took 1.2 to 1.5 s each, so the runner as a whole was not slow at that moment; only the failing boot stalled (8.9 s and 11.2 s), and the retry, which runs on a fresh worker and so in a fresh browser process, booted at once. In 38032960035 the failing test was the first Traditional-Chinese page of its shard; in 38034892101 a zh-Hant page (
whats-new.spec.ts:544) had already booted normally earlier in the same shard.Local reproduction
A cold-start tool outside the suite (
scratch/, its own dev server, 1 worker, retries 0): every start is a fresh browser context opened with a browser language (zh-TW, zh-CN or en-US), seeded like the fixture; an init script timestamps<html lang>, the first child of#root,.toolbar/.canvasvisible, long tasks, frame gaps and the app's twolocal()CJK punctuation faces; Node pings the renderer every 100 ms. Times are from navigation start.Condition (cold starts) First start on a fresh dev server Every other start: toolbar visible Longest long task Renderer pings over 1 s fresh dependency cache ( --force), zh-Hant first (12)3,252 ms 423–555 ms 150–253 ms 0 fresh dependency cache, English first (6) 2,045 ms 318–364 ms 109–148 ms 0 warm dependency cache, English first (6) 2,010 ms 321–365 ms 108–149 ms 0 a new browser process per start, zh-Hans → zh-Hant → en (9) 2,062 ms 303–369 ms 106–155 ms 0 renderer CPU throttled 4x (9) 3,017 ms 1,142–1,422 ms 627–848 ms one each in 4 of 9 starts, all Chinese renderer CPU throttled 8x (9) 4,528 ms 2,720–3,205 ms 1,484–1,927 ms one each in 6 of 9 starts, all Chinese - Only the first page a fresh dev server serves is slow (its module transforms), whatever the language; a language's own first load adds nothing measurable (zh-Hant first seen third: 351–354 ms).
- After the catalog arrives (
langset), the first render is one long task; it is a little longer in Chinese than in English (at 4x about 825–848 ms against 627–653 ms) and is the window in which a 1 s probe can go unanswered: at 4x and 8x every unanswered probe was on a Chinese start, none on an English one. - The
local()CJK punctuation face reachesloadedin the same frame as the visible toolbar: it is resolved during that first render. - None of 51 local cold starts came near 8 s. The longest, the first start at 8x CPU, was 4.5 s.
What this does and does not establish
It does not reproduce the stall and does not name its cause. It shows that an ordinary boot, even on a renderer 8 times slower, is a few seconds, so the CI stalls of 8 s and more are not the boot simply running slowly; and that the report taken after the deadline cannot say where the time went. Not established: whether the time on CI is spent waiting for the dev server, in the first-render task (for example font work on a runner whose disk cache is cold), or elsewhere. The next step is to measure the boot stages on the CI runner itself.
Diagnosis, part 2: one targeted measurement on the CI runner
Following #305 (comment). One throwaway workflow run on a branch that was then deleted (the workflow never reached
main): run 38041796233 (1f57d1b),windows-latest, attempt 1, no retry, 30 tests passed. The ordinary e2e, dist and PWA suites were not run.What it did
30 cold document starts in one browser process, a fresh context each, opened with a browser language in the order of both PR #345 failures: zh-Hans → zh-Hant → en, ten times. Each start booted exactly as the suite does,
page.goto('/')thenwaitForAppReadywith its unchanged 8 s condition, and carried the new boot trace (Node's and the renderer's times kept apart, and a late trace if the renderer misses the probe).Result
- 30 of 30 starts ready, 0 stalls: no report, no unanswered probe, no late trace.
- Ready (from
gototo the visible canvas): median 997 ms; per language, without the first start, zh-Hans 995 ms, zh-Hant 999 ms, en 997 ms. - In every start the toolbar and the canvas became visible in the same frame, right after one first-render long task that began a few ms after
<html lang>was set. - The two slowest starts were the first two: ci: GitHub Actions — verification only (no deploy) #1 zh-Hans, 2,048 ms (the fresh dev server's first page; long task 718 ms) and feat(state): Slice 1 — trigger + passive activation + delay (loop-state/1) #2, the first zh-Hant start of the browser process, 2,317 ms: navigation to DOMContentLoaded 819 ms, a 911 ms first-render long task (median of all 30: 235 ms) and the run's slowest request, 843 ms, the IBM Plex Mono
latin-400font file. The other nine zh-Hant starts had long tasks of 216 to 292 ms. - Longest frame gap 926 ms (feat(state): Slice 1 — trigger + passive activation + delay (loop-state/1) #2); median 238 ms. The slowest request of every start was the same font file (median 153 ms).
What this does and does not establish
The 8 s stall did not reproduce in these 30 starts on the CI runner, and its cause is not established. What the run shows is a first-use cost of several hundred ms on the runner, in the first Traditional-Chinese start of a browser process, far below 8 s; whether it is related to the stalls is not known. The boot-trace diagnostics are going into the ordinary e2e so that the next real failure reports where its time went; no readiness condition, timeout or retry policy changes. This issue stays open.
- added a commit that references this issue
on Oct 10, 2026
Problem
On CI, the Traditional Chinese case of the descriptive-copy wrapping test has taken longer on each of the last four runs, with no related code change. On the fourth run it failed: the first attempt ran out of the 30 s test timeout, and the automatic retry ended with the browser session closed.
The test is
e2e/descriptive-copy-wrapping.spec.ts, "CJK descriptive copy wraps instead of overflowing", thezh-Hantcase. It opens a fresh browser context with the localezh-TW, opens every toolbar menu, and hovers every palette chip.Measured
The same test in the same job,
e2e (shard 1/4), on the Windows runner.zh-Hanscasezh-Hantcasejacaseword-breaktest in the same filemain3f07c11maind97ae725d8b35c3d75404On a local machine at
3d75404the whole file passes, 25 of 25, and the three cases take 1.7 s, 1.7 s and 1.6 s.What the failed run showed
Internal server error, session closed.Not known
The cause. Nothing measured so far says what the runner spends the time on, or why the browser session ended. A font is one guess among several and is not established.
Before any change
The timeout is not raised first. A longer timeout would hide a cost that has grown tenfold in four runs and would not explain a closed browser session. Measure on the CI runner first:
zh-TWand for the other localesAcceptance
Context
Seen on pull request 304, whose change cannot reach this test. That one failed job was re-run once, by decision, without editing code or tests.