Add manual RTX benchmark action with Pages dashboard - #18
Merged
Merged
Conversation
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01ERiF61xcNaEQNrhaQZvdyq
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01ERiF61xcNaEQNrhaQZvdyq
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01ERiF61xcNaEQNrhaQZvdyq
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01ERiF61xcNaEQNrhaQZvdyq
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01ERiF61xcNaEQNrhaQZvdyq
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01ERiF61xcNaEQNrhaQZvdyq
… timeout Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01ERiF61xcNaEQNrhaQZvdyq
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01ERiF61xcNaEQNrhaQZvdyq
…ntext Final review fixes: robust timing tests, host/user masked in logs, step timeouts and a repetitions cap so restore fits the job, phase-2 failures keep phase-1 results, per-suite ctx on the VRAM chart, GPU name variable, docs corrected. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01ERiF61xcNaEQNrhaQZvdyq
Contributor
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this adds
A benchmark you start by hand from the Actions tab. It runs our RTX-worker agent tests and publishes the results to a GitHub Pages dashboard.
.github/workflows/benchmark.yml(workflow_dispatchonly):tailscale/github-action@v4) and connects to the worker over SSH.llama-server, and runs the chosen suites.worker.sh down, with retries), and fails if live/healthdoes not return 200.rtx-benchmarkEnvironment. Inputs can override them. Secrets are never inputs, and the worker host and user are masked in the logs.tests/bench/run_suites.py):fixture: the existing Playwright tasks.long_context: a long-prompt recall probe, sized from a measured 22.2 tokens per line.live_web: DuckDuckGo image search, saving a real image file.concurrency: runs on a 2-slot server, serial and in parallel.tests/bench/collect.py): one run summary per run. The publish job adds it to thegh-pageshistory and rebuildsdata/index.json.bench-site/): static HTML and JS with inline SVG. It shows the latest run, pass rate and time over runs, peak VRAM against context (per suite), and a table of all runs. It works in light and dark mode and at phone width. It is XSS-safe (it renders text only). The committed sample data is marked "sample".docs/benchmarks.md: every finding so far in one place.docs/benchmark-action.md: setup (Tailscale, SSH key, Environment, Pages) and how to use the action.Setup needed before the first run
See
docs/benchmark-action.md:rtx-benchmarkEnvironment and limit its deployment branches tomain.TS_OAUTH_*orTS_AUTHKEY,WORKER_SSH_KEY,WORKER_KNOWN_HOSTS, and preferablyWORKER_HOST/WORKER_SSH_USER) and the variables.tag:cican reach only port 22 on the worker.A run stops the live Bonsai container for as long as it takes. The workflow has not been run yet, and no step here has contacted the worker.
Verification
python3 -m unittest discover -s tests -q: 533 tests, OK (22 skipped), on Python 3.14 and 3.11.shellcheck tests/bench/worker.sh: clean.actionlint: clean.Known follow-ups
🤖 Generated with Claude Code
https://claude.ai/code/session_01ERiF61xcNaEQNrhaQZvdyq