Skip to content

Bring merged #17 and #18 into main - #19

Merged
KSEGIT merged 24 commits into
mainfrom
playwright-model-expansion
Sep 27, 2026
Merged

KSEGIT merged 24 commits into
mainfrom
playwright-model-expansion

Conversation

@KSEGIT

@KSEGIT KSEGIT commented Sep 27, 2026

Copy link
Copy Markdown
Owner

Why this PR exists

PRs #16, #17 and #18 were stacked. They were merged within seconds of each other, so GitHub had not yet moved #17 and #18 onto main:

This PR brings the reviewed content of #17 and #18 into main. It adds no new code.

playwright-model-expansion contains benchmark-action, and git merge-tree against main shows no conflicts. The only commit on main that this branch lacks is the merge commit of #16.

After merge, benchmark appears under Actions. See docs/benchmark-action.md for the one-time setup before the first run.

🤖 Generated with Claude Code

https://claude.ai/code/session_01ERiF61xcNaEQNrhaQZvdyq

KSEGIT and others added 24 commits September 23, 2026 20:15
Serial runs on the same two-slot server passed 3/4. The one failure had
the same Name-field error as the concurrent runs, so concurrency does
not cause it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ERiF61xcNaEQNrhaQZvdyq
64K loads on the RTX 3070 Ti with ~870 MiB VRAM spare and handles a
60K-token prompt. The live image task passed 3/4 once old page views
were pruned; headless Chrome needed a normal user agent.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ERiF61xcNaEQNrhaQZvdyq
…ntext

Final review fixes: robust timing tests, host/user masked in logs, step
timeouts and a repetitions cap so restore fits the job, phase-2 failures
keep phase-1 results, per-suite ctx on the VRAM chart, GPU name variable,
docs corrected.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ERiF61xcNaEQNrhaQZvdyq
With PW_MCP_OUTPUT_MAX_SIZE=0 the launcher expanded an empty array under
set -u. bash before 4.4 (macOS /bin/bash 3.2, used by the macOS CI runner)
treats that as unbound and exits 1. Use the ${a[@]+"${a[@]}"} form and
add a test that runs the launcher under the old system bash.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ERiF61xcNaEQNrhaQZvdyq
Add manual RTX benchmark action with Pages dashboard
Copilot AI lite review requested due to automatic review settings September 27, 2026 00:21
@coderabbitai

coderabbitai Bot commented Sep 27, 2026

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Note

Currently processing new changes in this PR. This may take a few minutes, please wait...

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: 750d231a-791b-4e16-9de5-13826fb286bd

📥 Commits

Reviewing files that changed from the base of the PR and between 461e9ff and 3a4f5de.

⛔ Files ignored due to path filters (1)
  • models.lock.tsv is excluded by !**/*.tsv
📒 Files selected for processing (78)
  • .env.example
  • .github/workflows/benchmark.yml
  • README.md
  • bench-site/app.js
  • bench-site/data/index.json
  • bench-site/index.html
  • bench-site/style.css
  • claude-bonsai.sh
  • docker/Dockerfile
  • docs/benchmark-action.md
  • docs/benchmarks.md
  • docs/browser-agent-models.md
  • docs/browser-agent-validation.md
  • docs/playwright-agent-benchmark.md
  • fetch-models.sh
  • models.ini.in
  • qwen3.5-chat-template.jinja
  • setup-opencode.sh
  • start-server.sh
  • tests/__init__.py
  • tests/bench/__init__.py
  • tests/bench/collect.py
  • tests/bench/run_suites.py
  • tests/bench/worker.sh
  • tests/fixtures/agent-benchmark.html
  • tests/fixtures/bench/run-2x24k/concurrency-parallel-1.json
  • tests/fixtures/bench/run-2x24k/concurrency-parallel-2.json
  • tests/fixtures/bench/run-2x24k/concurrency-parallel-3.json
  • tests/fixtures/bench/run-2x24k/concurrency-parallel-4.json
  • tests/fixtures/bench/run-2x24k/concurrency-serial-1.json
  • tests/fixtures/bench/run-2x24k/concurrency-serial-2.json
  • tests/fixtures/bench/run-2x24k/concurrency-serial-3.json
  • tests/fixtures/bench/run-2x24k/concurrency-serial-4.json
  • tests/fixtures/bench/run-2x24k/meta.json
  • tests/fixtures/bench/run-2x24k/vram-concurrency.json
  • tests/fixtures/bench/run-32k/fixture.json
  • tests/fixtures/bench/run-32k/live_web-1.json
  • tests/fixtures/bench/run-32k/live_web-2.json
  • tests/fixtures/bench/run-32k/meta.json
  • tests/fixtures/bench/run-32k/vram-fixture.json
  • tests/fixtures/bench/run-32k/vram-live_web.json
  • tests/fixtures/bench/run-64k/live_web-1.json
  • tests/fixtures/bench/run-64k/live_web-2.json
  • tests/fixtures/bench/run-64k/live_web-3.json
  • tests/fixtures/bench/run-64k/long_context.json
  • tests/fixtures/bench/run-64k/meta.json
  • tests/fixtures/bench/run-64k/vram-live_web.json
  • tests/fixtures/bench/run-64k/vram-long_context.json
  • tests/fixtures/bench/run-errors/errors.json
  • tests/fixtures/bench/run-errors/long_context.json
  • tests/fixtures/bench/run-errors/meta.json
  • tests/fixtures/bench/run-errors/vram-long_context.json
  • tests/fixtures/bench/run-missing-meta/fixture.json
  • tests/fixtures/bench/runs-index/broken.json
  • tests/fixtures/bench/runs-index/run-2026092501.json
  • tests/fixtures/bench/runs-index/run-2026092502.json
  • tests/fixtures/bench/runs-index/run-2026092503.json
  • tests/live_image_agent.py
  • tests/long_context_probe.py
  • tests/playwright_agent_bench.py
  • tests/runtime_clients.py
  • tests/smoke_agent_fixture.py
  • tests/test_agent_presets.py
  • tests/test_bench_collect.py
  • tests/test_benchmark_workflow.py
  • tests/test_chat_template.py
  • tests/test_claude_bonsai.py
  • tests/test_live_image_agent.py
  • tests/test_long_context_probe.py
  • tests/test_model_manifest.py
  • tests/test_opencode_config.py
  • tests/test_playwright_agent_bench.py
  • tests/test_run_suites.py
  • tests/test_runtime_clients.py
  • tests/test_script_permissions.py
  • tests/test_server_startup.py
  • tests/test_worker_script.py
  • version.sh
 __________________________________________
< Thread behaviour? I don't even know her. >
 ------------------------------------------
  \
   \   (\__/)
       (•ㅅ•)
       /   づ
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@KSEGIT
KSEGIT merged commit 8deddc6 into main Sep 27, 2026
2 of 4 checks passed

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@KSEGIT
KSEGIT deleted the playwright-model-expansion branch September 27, 2026 00:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants