From da1e83515055f7647c5034e35d527d810816349c Mon Sep 17 00:00:00 2001
From: Claude
Date: Mon, 24 Aug 2026 02:13:58 +0000
Subject: [PATCH 1/3] =?UTF-8?q?prompt:=20scope=20the=20test-performance=20?=
=?UTF-8?q?board=20=E2=80=94=20run=20times,=20hangs,=20no=5Frun,=20one-tap?=
=?UTF-8?q?=20fix=20prompts?=
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
Deep-research scoping for a standing dashboard of the testing/integration
infrastructure's run times (PR smoke gates, unit tests, import time), with
history to flag drops, kill-timer/hang events, and the NO_RUN census β every
row carrying its own one-tap Claude prompt.
- docs/pyautoheart/test_performance_board_assessment.md β the assessment:
Heart owns measurement and publishes a `performance` block in board.json
(schema v2, additive, rows carry their own prompts); the Brain board
consumes the headline verbatim via the existing heart_blockers contract.
Three data planes: Actions-API wall-clock scrape (proven by
morning_health.yml), per-script timings via the delegated runner (one
PyAutoHands change since #260-#263), and Heart's tracked baselines (fix
script_timing first β its bug prompt is filed and unissued).
- draft/feature/pyautoheart/test_performance_board.md β the actionable
phased prompt.
- dashboard regenerated (intake --apply dashboard).
Co-Authored-By: Claude Fable 5
Claude-Session: https://claude.ai/code/session_01EoDPz2LevKeBaDwqFKtZrU
---
dashboard.html | 25 +-
dashboard.md | 42 +--
.../test_performance_board_assessment.md | 281 ++++++++++++++++++
.../pyautoheart/test_performance_board.md | 91 ++++++
4 files changed, 410 insertions(+), 29 deletions(-)
create mode 100644 docs/pyautoheart/test_performance_board_assessment.md
create mode 100644 draft/feature/pyautoheart/test_performance_board.md
diff --git a/dashboard.html b/dashboard.html
index 1bf29b17..9070c653 100644
--- a/dashboard.html
+++ b/dashboard.html
@@ -123,10 +123,10 @@
π
PyAutoMindDashboard
Intent. Priority. Flow.
Every task the Mind is holding. Tap a task's π and its /start_dev command is on your clipboard β paste it into a Claude Code chat to route Claude straight to that task. Recent is the same work by date β what has been happening rather than what to do next.
155 filed prompts, not started β sorted most-pickable first (priority, then size). 25 of them belong to an epic and are listed only under Epics below.
+
156 filed prompts, not started β sorted most-pickable first (priority, then size). 25 of them belong to an epic and are listed only under Epics below.
diff --git a/dashboard.md b/dashboard.md
index 26e04989..33518992 100644
--- a/dashboard.md
+++ b/dashboard.md
@@ -11,11 +11,11 @@ Every task the Mind is holding, on one page: what is in flight, what is parked,
| [In flight](#in-flight) (`active/`) | 1 |
| [Parked](#parked) (`parked.md`) | 3 |
| [Planned](#planned) (`planned.md`) | 6 |
-| [Backlog](#backlog) (`draft/`) | 155 |
+| [Backlog](#backlog) (`draft/`) | 156 |
## Start here
-**Highest priority** (filed as `high`) β showing 12 of 17
+**Highest priority** (filed as `high`) β showing 12 of 18
π TRIAGE: needs manual review before routing β medium Β· safe Β· high
@@ -73,6 +73,14 @@ Every task the Mind is holding, on one page: what is in flight, what is parked,
+π Test-performance section on the Heart board β run times, hangs, NO_RUNβ¦ β pyautoheart Β· large Β· supervised Β· high
+
+```
+/start_dev draft/feature/pyautoheart/test_performance_board.md
+```
+
+
+
π Deep research: Can we speed up Delaunay in PyAutoArray? β autoarray Β· too-large Β· supervised Β· high
```
@@ -105,14 +113,6 @@ Every task the Mind is holding, on one page: what is in flight, what is parked,
-π Re-baseline the MGE imaging JIT profiling regression value β autolens_workspace_developer Β· too-large Β· supervised Β· high
-
-```
-/start_dev draft/test/autolens_workspace_developer/mge_jit_regression_rebaseline.md
-```
-
-
-
**Quick wins** (small enough, and safe enough to run unattended)
π Defer the eager scipy.sparse import in derivative_util (~0.10 s of import) β libraries Β· small Β· safe Β· normal
@@ -235,10 +235,10 @@ Scoped but not started; some are not yet prompt files. Full detail in [`planned.
## Backlog
-**155** filed prompts, not started. Each section is sorted most-pickable first (priority, then size). **25** of them belong to an epic and are listed only under [Epics](#epics) below.
+**156** filed prompts, not started. Each section is sorted most-pickable first (priority, then size). **25** of them belong to an epic and are listed only under [Epics](#epics) below.
-feature β 29
+feature β 30π Numba CPU likelihood phase 1: batched MGE convolution + operated-matrix caching β autoarray Β· medium Β· supervised Β· high
@@ -280,6 +280,14 @@ Scoped but not started; some are not yet prompt files. Full detail in [`planned.
+π Test-performance section on the Heart board β run times, hangs, NO_RUNβ¦ β pyautoheart Β· large Β· supervised Β· high
+
+```
+/start_dev draft/feature/pyautoheart/test_performance_board.md
+```
+
+
+
π Which other searches need prior-support handling β coverage audit after Prodigy β autofit Β· medium Β· supervised Β· medium
```
@@ -1328,6 +1336,7 @@ The 50 newest things to happen to the work in hand, newest first β issued, par
| Date | Event | Task |
|------|-------|------|
+| 2026-08-24 | filed | Test-performance section on the Heart board β run times, hangsβ¦ |
| 2026-08-23 | filed | pynufft removal: unswept downstream residue (1 hard break + stale⦠|
| 2026-08-23 | filed | Properly time and profile the smoke/release script surface |
| 2026-08-23 | filed | Phase 3: stop installing pynufft in Hands/Heart CI and PyAutoCTI⦠|
@@ -1337,12 +1346,12 @@ The 50 newest things to happen to the work in hand, newest first β issued, par
| 2026-08-22 | filed | Untrack the generated FITS test artifacts in autoarray |
| 2026-08-22 | filed | The reconstruction noise map describes a different estimator than the⦠|
| 2026-08-22 | filed | Remove pynufft + legacy TransformerNUFFTPyNUFFT |
-| 2026-08-22 | filed | Point-source JSON datasets record no resolution regime |
β¦ 10 more (40 left)
| Date | Event | Task |
|------|-------|------|
+| 2026-08-22 | filed | Point-source JSON datasets record no resolution regime |
| 2026-08-22 | filed | Is Intel macOS a supported platform, and what is the numpy-only⦠|
| 2026-08-22 | filed | Defer the eager scipy.sparse import in derivative_util (~0.10 s of⦠|
| 2026-08-22 | filed | Bug: fix the tracer.fits existence guard in autolens_workspace⦠|
@@ -1352,12 +1361,12 @@ The 50 newest things to happen to the work in hand, newest first β issued, par
| 2026-08-20 | filed | Numba CPU likelihood phase 1: batched MGE convolution +β¦ |
| 2026-08-19 | filed | status.sh --repos sources a file that no longer exists |
| 2026-08-19 | filed | jax 0.11 breaks beta/gamma message log_partition under jit⦠|
-| 2026-08-19 | filed | autolens_workspace_test jax_likelihood pins: 4 scripts fail smoke on⦠|
β¦ 10 more (30 left)
| Date | Event | Task |
|------|-------|------|
+| 2026-08-19 | filed | autolens_workspace_test jax_likelihood pins: 4 scripts fail smoke on⦠|
| 2026-08-19 | filed | autofit_profiling: bootstrap the repo + general PyAutoFit profiling⦠|
| 2026-08-19 | filed | autoreduce 0.9 on PyPI never got the Python 3.12 floor |
| 2026-08-19 | issued | @PyAutoFitTransformedMessage.factor_gradient crashes on first⦠|
@@ -1367,12 +1376,12 @@ The 50 newest things to happen to the work in hand, newest first β issued, par
| 2026-08-19 | filed | PyAutoConf rename leftovers in Brain functional surfaces |
| 2026-08-19 | filed | Explore: dashboardify the Brain's operational surfaces with pasteable⦠|
| 2026-08-19 | filed | Deduplicate repos_sync.py's check/write pairs |
-| 2026-08-19 | filed | Bug in autocti_workspace: the dataset_1d results/database example⦠|
β¦ 10 more (20 left)
| Date | Event | Task |
|------|-------|------|
+| 2026-08-19 | filed | Bug in autocti_workspace: the dataset_1d results/database example⦠|
| 2026-08-18 | parked | single-source-density-design |
| 2026-08-18 | parked | prior-message-collapse-design |
| 2026-08-18 | filed | @PyAutoFitTransformedMessage.logpdf/pdf omit the transform⦠|
@@ -1382,12 +1391,12 @@ The 50 newest things to happen to the work in hand, newest first β issued, par
| 2026-08-14 | filed | Three jax_likelihood pins are stale by ~1.24e-4 and fail the smoke⦠|
| 2026-08-09 | found | isothermal-ell-sph-oversampling-at-the-cusp |
| 2026-08-08 | parked | pyautoreduce-slacs1430-acs-comparison |
-| 2026-08-08 | filed | Regenerate autolens_workspace markdown/ so the MGE pages show⦠|
β¦ 10 more (10 left)
| Date | Event | Task |
|------|-------|------|
+| 2026-08-08 | filed | Regenerate autolens_workspace markdown/ so the MGE pages show⦠|
| 2026-08-07 | filed | autofit.plot functions accept **kwargs and silently discard them |
| 2026-08-07 | filed | Regenerate setup_notebook-drifted notebooks in⦠|
| 2026-08-06 | filed | Triage: Convolver "No blurring_image provided" warning in canonical⦠|
@@ -1397,7 +1406,6 @@ The 50 newest things to happen to the work in hand, newest first β issued, par
| 2026-08-04 | filed | dataset/imaging/jwst_lw is untracked because the gitignore was never⦠|
| 2026-08-04 | filed | cosmos_web_ring stores boolean masks as float64, wasting ~3.4 MB of⦠|
| 2026-08-04 | filed | autolens_workspace_developer: broad stale-API rot (56 symbols, no CI) |
-| 2026-08-04 | filed | aplt.Output stale-API drift in the remaining workspace repos |
diff --git a/docs/pyautoheart/test_performance_board_assessment.md b/docs/pyautoheart/test_performance_board_assessment.md
new file mode 100644
index 00000000..951dba38
--- /dev/null
+++ b/docs/pyautoheart/test_performance_board_assessment.md
@@ -0,0 +1,281 @@
+# The test-performance board β scoping assessment
+
+Date: 2026-08-24. The deep-research write-up for "the speed of AI development
+is heavily dependent on the speed of tests": a dashboard that always shows the
+run times of the testing/integration infrastructure (PR smoke gates on the
+`*_workspace_test` and normal workspaces, HowTo, unit tests, import time), with
+history to flag performance drops, kill-timer/hang events surfaced, NO_RUN
+scripts listed with their reasons, and one-tap Claude prompts on every row so
+"speed this one up" is a paste, not an archaeology session. Actionable prompt:
+`draft/feature/pyautoheart/test_performance_board.md`.
+
+## The problem, with receipts
+
+- Manual timing archaeology is the current tool. The jax_grad budget work
+ records that "the diagnosis had to be rebuilt from CI job logs by hand"
+ (`complete/2026/08/jax-grad-smoke-timeout-budget.md`), and the 2026-08-23
+ slow-vs-stall audit hand-scraped `[PASS] β s` lines per run.
+- The cost of not watching is measured: four `autogalaxy_workspace_test` runs
+ burned ~24h of runner time at the 6-hour Actions ceiling before the kill
+ timer existed (`draft/bug/ci/jax_vmap_jit_compile_stall.md`); the
+ autolens_workspace_test PR gate is ~11m20s wall-clock of which 553s is
+ scripts, at ~17 runs/week (`draft/test/workspaces/slowest_smoke_gate_scripts.md`,
+ `draft/test/pyautoheart/smoke_relevance_gate.md`).
+- The existing markers are untrustworthy: "**a SLOW marker is not evidence of
+ slowness**" β the first SLOW-marked entry ever measured was wrong by ~50Γ,
+ and every 2026-07-14 marker records no timing at all
+ (`complete/2026/08/jax-compile-stall-slow-vs-stall-audit.md`).
+- The tracked signal is broken: Heart's `script_timing` baselines are orphaned
+ by path-derived slugs (no history accumulated for the moved jax_grad scripts
+ since 2026-07-24) and every stored history is one value repeated seven times
+ (`draft/bug/pyautoheart/script_timing_baselines_orphaned_and_window_filled.md`
+ β filed 2026-08-04, **never issued**).
+
+## What is recommended
+
+**A "β± Performance" section on the Heart board, published in the Heart's
+`board.json` (schema v2, additive) with every row carrying its own one-tap
+prompt β and a headline row on the Brain board consuming it verbatim, exactly
+the way `heart_blockers` already works.** Not a seventh board.
+
+Why Heart, not Brain:
+
+- The doctrine is explicit and repeated: "**measurement lives in Heart;
+ hygiene acts**" and "new standing signals (import cost, CLI noise) become
+ Heart *legs*, not a new repo"
+ (`PyAutoBrain/agents/conductors/hygiene/AGENTS.md:143-149`). Heart already
+ owns every timing signal in the organism: `script_timing`, `test_run`,
+ `import_time`, `unit_test_timing`, `workspace_testmode_timing`.
+- The consumption seam already exists and is tested: the Heart board publishes
+ structured blockers (`{text, severity, repo, repo_url, run_url, prompt}`)
+ and the Brain board renders them verbatim, never re-deriving the prompt
+ (`PyAutoBrain/board/_board.py:228-249`, `tests/test_board.py:43-53`).
+ A `performance` block rides the same contract; schema v2 is additive by
+ design (`complete/2026/08/actionable-health-board.md`).
+- The cadence already lines up: Heart renders at 05:00 UTC, the Brain board
+ reads it at 05:30 (`PyAutoBrain/.github/workflows/brain_board.yml:20-25`).
+- Timing rows are **advisory, never gating** β the precedent is the
+ `import_time` leg ("advisory dashboard section, NOT in the readiness gating
+ set", `complete/2026/07/import-time-heart-leg.md`). The four-tier
+ GREEN/STALE/YELLOW/RED verdict is untouched.
+
+Why not a Brain-board section (considered): the Brain board holds the exact
+code to copy β `gh_json` (`_board.py:136-150`), the overnight scrape
+(`:167-211`), the self-carrying history (`:456-466`) and `sparkline()` β but
+putting the *measurement* there splits ownership against the doctrine, and the
+Brain "owns no state, no health checks". The Brain board's role is the morning
+headline: "N gates slowed / 1 hang flagged", chip β the Heart page. That row
+can ride `draft/feature/pyautobrain/brain_board_follow_ups.md`.
+
+## How the timing information gets calculated and populated
+
+Three data planes, in cost order. "Just scrape recently completed PRs" β the
+suggestion that seeded this scoping β is plane A, it works today, and it is
+already proven in production.
+
+### Plane A β workflow wall-clock, scraped from the Actions API (ships first)
+
+The default `GITHUB_TOKEN` reads other public PyAutoLabs repos' workflow runs:
+`PyAutoMind/.github/workflows/morning_health.yml:68-101` already does exactly
+this daily against PyAutoHeart/PyAutoBrain/PyAutoHands/PyAutoFit with
+`permissions: contents: read` and no PAT. Per run the API serves everything
+the board needs:
+
+- **duration** = `updated_at β run_started_at` (also served directly as
+ `run_duration_ms` by `GET /actions/runs/{id}/timing`; ignore the `billable`
+ fields β they are 0 on public repos). **Trap, already recorded by the Hands
+ board: use `run_started_at`, never `created_at`** β re-attempted runs
+ otherwise report multi-day durations (`complete/2026/08/release-board.md`).
+- **queue delay** = `run_started_at β created_at`.
+- **conclusion** (`success | failure | cancelled | β¦`) β see the kill-timer
+ section for disambiguating `cancelled`.
+- **PR association**: `event: pull_request` + `head_branch` β so "what a
+ contributor waits on a PR" is directly measurable, per gate, per week.
+- **per-job and per-step timings** via `GET /actions/runs/{id}/jobs` β matrix
+ leg identity is in the job `name` (`"pytest (3.12)"`), and the *Run smoke
+ tests* step is separable from checkout/install overhead. Fan out to `/jobs`
+ only for runs worth the detail (the slowest of the day, every `cancelled`
+ one) β that keeps the request budget at ~1 call per tracked workflow per
+ render, trivially inside the 1000 req/hr limit.
+
+The tracked-workflow list is declared config, not code (the tenant-firewall
+rule): a `performance:` block in `PyAutoHeart/config/repos.yaml` naming each
+`repo:workflow` pair β every workspace's `Smoke Tests` caller, the HowTo
+gates, the organ self-test gates (`tests.yml`), Heart's weekly
+`workspace-smoke.yml`, `release-integrate.yml`, and the Brain's
+`nightly-release.yml`. One channel note: since #122 the validation channels
+are split so runs attribute to the **caller** workflow
+(`complete/2026/07/split-validation-channels.md`) β query the callers.
+
+What plane A yields per gate: last-N durations, p50/max, conclusion mix,
+queue delay, and the two contributor-facing numbers that matter β median PR
+gate latency and its trend.
+
+### Plane B β per-script timings, promoted from log lines to a standing dataset
+
+The runner already prints `[PASS] β s` per entry and a
+`=== Smoke test summary ===` terminal line; today that data evaporates into
+job logs (retained ~90 days) and is only recovered by hand-scraping. The
+smoke-runner delegation (`complete/2026/08/smoke-runner-delegation.md`) turned
+all ten workspace runners into thin shims over `autohands/run_python.py` β
+**so per-script timing recording is now one PyAutoHands change, not ten
+repo sweeps**: extend the existing report machinery (`--report-dir` is
+already load-bearing β "without it the gate runs to completion and always
+exits 0") to emit a `smoke_timings.json` (entry, status, seconds, cap in
+force, exit code), uploaded as a run artifact and/or written to
+`$GITHUB_STEP_SUMMARY`. The Heart render fetches the latest artifact per gate.
+
+This is precisely the open question item 4 of
+`draft/research/ci/smoke_timing_and_profiling.md` already poses ("should the
+runner record per-script timings routinely, so this is a standing dataset
+rather than a periodic archaeology exercise?") β this assessment's answer is
+**yes, via the delegated runner**. The `retime.yml` classifier and its
+verdict vocabulary (STALL/SLOW/NEITHER/AMBIGUOUS/ERROR, `retime_results.json`)
+stay as the on-demand deep probe, reached through `smoke-tests.yml`'s
+`runner`/`runner-args`/`script-timeout` inputs.
+
+**STALL β SLOW is a first-class dimension, not a footnote.** "A slow script
+has a tight timing distribution. A stalling one is bimodal" β a healthy
+compile of `rectangular_mge.py` is 3.1s; a stalled one exceeds 300s, same
+commit, same runner image (`jax-compile-stall-slow-vs-stall-audit.md`). A
+single wall-clock number per entry cannot express this; the board should carry
+the retime verdict where one exists, and flag bimodality from the standing
+dataset where it doesn't.
+
+### Plane C β baselines and regression flagging (the "history" that must be durable)
+
+Two history mechanisms exist in the board family, with different guarantees:
+
+1. **Self-carrying published artifact** β the Brain board fetches its own
+ previous `board.json` at render, appends today, caps at 30 entries
+ (`PyAutoBrain/board/_board.py:456-466`). Free, no commits, idempotent per
+ date β but lossy (a publish gap loses everything) and 30 days max. Right
+ for the plane-A per-gate daily aggregates and their sparklines.
+2. **Heart's tracked rolling baselines** (`~/.pyauto-heart/`, the
+ `script_timing` mechanism) β durable, but currently broken as filed. Right
+ for per-script baselines once fixed.
+
+The fix is a prerequisite, and its prompt is already written:
+`draft/bug/pyautoheart/script_timing_baselines_orphaned_and_window_filled.md`
+(rename-aware slugs or a loud no-baseline signal; real 7-run accumulation;
+record the source run id per duration). **Issue it first.**
+
+For flagging drops, do not invent thresholds β reuse the profiling
+conductor's considered doctrine: a regression counts only if it is newer than
+its pin, **at least 2.0Γ the pinned value AND at least 1.0s above it in
+absolute terms**, with sticky pins that never move without an explicit
+`--repin` ("host load alone has produced 7Γ errors in this corpus and an
+alarm that cries wolf gets ignored" β
+`PyAutoBrain/agents/conductors/profiling/AGENTS.md`). The comparability-key
+lesson from the compile-warm dashboard transfers too: never pool different
+hosts under one label β for CI the key is runner image Γ Python leg Γ event
+type (`complete/2026/08/compile-warm-baseline-dashboard.md`).
+
+Unit tests and import time ride the same planes: the organ/library `tests.yml`
+gates are plane-A rows (they are seconds-to-a-minute today β the board's job
+is to notice when that stops being true), and the existing `import_time` leg
+(advisory, off-tick, subprocess-measured) surfaces its red/yellow counts as a
+row with a `/hygiene` chip β the standing leg promotion that
+`hygiene/AGENTS.md:117-119` already names as the deferred follow-up.
+
+## Kill-timer and hang events on the board
+
+The kill timer (PyAutoHands `build_util.timeout_for` + `kill_group`; 300s
+smoke default, 900s `jax_grad/`, 1800s release; exit 124; `TIMEOUT` status
+with the truncated output tail and the cap in force attached) and the
+in-process watchdog (faulthandler dump at 80% of `BUILD_SCRIPT_TIMEOUT`,
+heartbeat lines) already leave ingestible traces. The board renders, per gate:
+
+- **Per-script TIMEOUT rows** (plane B): entry, cap, count over the window,
+ and the retime verdict if one exists. Chip:
+ `/bug kill timer: