Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
24 commits
Select commit Hold shift + click to select a range
8c29b60
Add Qwen3.6 and Gemma browser-agent candidates and benchmark
KSEGIT Sep 23, 2026
07f239e
Record RTX Bonsai before and after browser baseline
KSEGIT Sep 24, 2026
fb0a64e
Document RTX comparison provenance and verified downloads
KSEGIT Sep 24, 2026
9bac94c
Harden browser agent prompt and Qwen3.5 Responses template
KSEGIT Sep 24, 2026
cc9eecf
Ignore GGUF template final newline in parity check
KSEGIT Sep 24, 2026
038e78c
Record RTX browser and client validation
KSEGIT Sep 25, 2026
5357505
Record Bonsai RTX compatibility results
KSEGIT Sep 25, 2026
c99febb
Record RTX 32K and two-slot agent tests
KSEGIT Sep 25, 2026
638f878
Record serial control for two-slot Qwen 9B job runs
KSEGIT Sep 25, 2026
c1f912b
Record 64K context and live DuckDuckGo image test for Qwen 9B
KSEGIT Sep 26, 2026
447581d
Add manual RTX benchmark workflow and worker control script
KSEGIT Sep 26, 2026
18ad524
Document benchmark findings and the benchmark action setup
KSEGIT Sep 26, 2026
d2100ef
Make benchmark restore of the live container retry on SSH failure
KSEGIT Sep 26, 2026
d7321ff
Add live-web agent, long-context probe and benchmark suite runner
KSEGIT Sep 26, 2026
7cbc843
Correct benchmark action docs against the real workflow
KSEGIT Sep 26, 2026
c14ad84
Add benchmark results collector and Pages dashboard
KSEGIT Sep 26, 2026
23f309d
Fit long-context probe to measured tokens and kill whole job trees on…
KSEGIT Sep 26, 2026
c7b44a9
Read concurrency VRAM from the single file the suite runner writes
KSEGIT Sep 26, 2026
0edcd5c
Mask worker details, cap run time before restore, record per-suite co…
KSEGIT Sep 26, 2026
9a28395
Fix Playwright launcher abort on macOS bash 3.2 when eviction is off
KSEGIT Sep 26, 2026
2693d69
Merge branch 'playwright-model-expansion' into benchmark-action
KSEGIT Sep 26, 2026
eb7dfb3
Merge branch 'runtime-modernisation' into playwright-model-expansion
KSEGIT Sep 26, 2026
3a453fd
Merge branch 'playwright-model-expansion' into benchmark-action
KSEGIT Sep 26, 2026
3a4f5de
Merge pull request #18 from KSEGIT/benchmark-action
KSEGIT Sep 27, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
27 changes: 27 additions & 0 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -105,27 +105,54 @@ BONSAI_API_KEY=bonsai-change-me
#QCM_QWEN9_CTX=8192
#QCM_QWEN4_CTX=8192
#QCM_GRANITE_CTX=8192
#QCM_QWEN36_CTX=8192
#QCM_GEMMA4_CTX=8192
#QCM_AGENT_UBATCH=256
#QCM_AGENT_CACHE_RAM=1024
# Existing Qwen3.5/Granite agent presets keep their prior 256-token reuse
# granularity. New Qwen3.6/Gemma start with Prism's default 0; compare both
# with the dedicated prompt-prefix probe before changing production settings.
#QCM_AGENT_CACHE_REUSE=256
# GPU layer limits can be reduced for partial offload when a model does not fit.
#BONSAI_NGL=99
#QCM_QWEN9_NGL=99
#QCM_QWEN4_NGL=99
#QCM_GRANITE_NGL=99
# For Qwen3.6, leaving NGL unset lets Prism --fit choose a hybrid GPU/CPU
# tensor split. Setting it overrides that fit choice: sweep measured values on
# the RTX 3070 Ti rather than copying the dense-model 99-layer setting.
# With explicit NGL, FIT_TARGET cannot guarantee the reserved VRAM margin.
#QCM_QWEN36_NGL=12 # example sweep value only; not a tested default
#QCM_QWEN36_FIT_TARGET=1536
#QCM_GEMMA4_NGL=99
# Flash attention: on/off/auto. Bonsai retains its prior on setting.
#BONSAI_FA=on
#QCM_QWEN_FA=auto
#QCM_GRANITE_FA=auto
#QCM_QWEN36_FA=auto
#QCM_GEMMA4_FA=auto
# Begin with f16; q8_0 is the first conservative quantised-cache comparison.
# Quantised V requires flash attention. Verify tools and long turns after changes.
#QCM_QWEN_CTK=f16
#QCM_QWEN_CTV=f16
#QCM_GRANITE_CTK=f16
#QCM_GRANITE_CTV=f16
#QCM_QWEN36_CTK=f16
#QCM_QWEN36_CTV=f16
#QCM_GEMMA4_CTK=f16
#QCM_GEMMA4_CTV=f16
# Qwen agent default: no thinking. Compare on and a bounded budget separately.
# Use the runtime reasoning option; enable_thinking CLI kwargs are deprecated.
#QCM_QWEN_REASONING=off
#QCM_QWEN_REASONING_BUDGET=-1
# Qwen3.6 and Gemma are separate text/tool-only presets; no mmproj is loaded.
# Select one pinned quant per alias and download that exact file with
# ./fetch-models.sh --model <id> [--quant <quant>]. Do not switch quant or
# model within each browser action; measure a whole workflow.
#QCM_QWEN36_QUANT=Q4_K_XL
#QCM_GEMMA4_QUANT=QAT_Q4_0
#QCM_QWEN36_REASONING=off
#QCM_QWEN36_REASONING_BUDGET=-1

# --- ./update.sh (Linux/Docker only) -----------------------------------------
# These matter because update.sh rolls the stack back when it decides a model
Expand Down
Loading
Loading