Skip to content

Add manual RTX benchmark action with Pages dashboard - #18

Merged
KSEGIT merged 11 commits into
playwright-model-expansionfrom
benchmark-action
Sep 27, 2026
Merged

KSEGIT merged 11 commits into
playwright-model-expansionfrom
benchmark-action

Conversation

@KSEGIT

@KSEGIT KSEGIT commented Sep 26, 2026

Copy link
Copy Markdown
Owner

What this adds

A benchmark you start by hand from the Actions tab. It runs our RTX-worker agent tests and publishes the results to a GitHub Pages dashboard.

  • .github/workflows/benchmark.yml (workflow_dispatch only):
    • A GitHub-hosted runner joins the tailnet (tailscale/github-action@v4) and connects to the worker over SSH.
    • It pauses the live container, runs a loopback-only test llama-server, and runs the chosen suites.
    • It always starts production again (worker.sh down, with retries), and fails if live /health does not return 200.
    • Settings come from the rtx-benchmark Environment. Inputs can override them. Secrets are never inputs, and the worker host and user are masked in the logs.
  • Suites (tests/bench/run_suites.py):
    • fixture: the existing Playwright tasks.
    • long_context: a long-prompt recall probe, sized from a measured 22.2 tokens per line.
    • live_web: DuckDuckGo image search, saving a real image file.
    • concurrency: runs on a 2-slot server, serial and in parallel.
    • VRAM is sampled on the worker.
  • Results (tests/bench/collect.py): one run summary per run. The publish job adds it to the gh-pages history and rebuilds data/index.json.
  • Dashboard (bench-site/): static HTML and JS with inline SVG. It shows the latest run, pass rate and time over runs, peak VRAM against context (per suite), and a table of all runs. It works in light and dark mode and at phone width. It is XSS-safe (it renders text only). The committed sample data is marked "sample".
  • Docs:
    • docs/benchmarks.md: every finding so far in one place.
    • docs/benchmark-action.md: setup (Tailscale, SSH key, Environment, Pages) and how to use the action.

Setup needed before the first run

See docs/benchmark-action.md:

  1. Create the rtx-benchmark Environment and limit its deployment branches to main.
  2. Add the secrets (TS_OAUTH_* or TS_AUTHKEY, WORKER_SSH_KEY, WORKER_KNOWN_HOSTS, and preferably WORKER_HOST/WORKER_SSH_USER) and the variables.
  3. Set the Pages source to "GitHub Actions".
  4. Set up a Tailscale ACL so tag:ci can reach only port 22 on the worker.

A run stops the live Bonsai container for as long as it takes. The workflow has not been run yet, and no step here has contacted the worker.

Verification

  • python3 -m unittest discover -s tests -q: 533 tests, OK (22 skipped), on Python 3.14 and 3.11.
  • shellcheck tests/bench/worker.sh: clean. actionlint: clean.
  • Timing tests: 96 parallel runs, 0 failures.
  • I ran the publish pipeline by hand on the fixtures: merge, summarize, rsync, index.
  • I checked the dashboard in a browser: desktop light, and dark mode at phone width.

Known follow-ups

  • Dashboard polish:
    • The VRAM axis uses uneven tick values.
    • Chart text is small on phones.
    • A suite with only one run shows dots with no line.
  • A worker-side watchdog to start production again if the runner is lost mid-run.

🤖 Generated with Claude Code

https://claude.ai/code/session_01ERiF61xcNaEQNrhaQZvdyq

KSEGIT and others added 9 commits September 26, 2026 10:56
…ntext

Final review fixes: robust timing tests, host/user masked in logs, step
timeouts and a repetitions cap so restore fits the job, phase-2 failures
keep phase-1 results, per-suite ctx on the VRAM chart, GPU name variable,
docs corrected.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ERiF61xcNaEQNrhaQZvdyq
Copilot AI lite review requested due to automatic review settings September 26, 2026 17:28
@coderabbitai

coderabbitai Bot commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: 659574df-c244-4736-a38e-dbbd0005b518

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@KSEGIT
KSEGIT merged commit 3a4f5de into playwright-model-expansion Sep 27, 2026
3 checks passed
KSEGIT added a commit that referenced this pull request Sep 27, 2026
@KSEGIT
KSEGIT deleted the benchmark-action branch September 27, 2026 00:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants