You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
v1: op_consolidated.py → op_consolidated_v1.py, ConsolidatedV1Scheduler, strategy/flag consolidated_v1 / --consolidated-v1. --consolidated and ConsolidatedScheduler remain as aliases (ImpactGuard).
v2 (--consolidated-v2): v1 posterior, scored by mean + tau * max(0, draw - mean), tau 0.65. It is optimistic (a draw below the mean scores the mean) and tempered (explores less). Leads the no-Elo precedence. It is wired into the Elo ballot, the fan-out, _register_arms, --elo-all/--hail-mary and the banners.
Why this ingredient
I ran a tournament of all 41 op schedulers on the 4 bandit_env environments (table in the op_consolidated_v2.py docstring). v1 was the only scheduler in the leading group on all four. Its competitors beat it only where they explored less: MOSS, FPL, and greedy on rotting.
On 20 paired seeds, v2 vs v1:
env
v1
v2
gain
v2 wins
stationary
1729
1759
+1.8%
19/20
decaying
4612
4685
+1.6%
20/20
rotting
3502
3558
+1.6%
17/20
fatigue150
1524
1544
+1.3%
16/20
Each gain is at least 3 SE of the paired difference.
A v1-vs-v1 control stays within 2 SE.
Cost is about 5µs per select at K=200.
Changes that did not help: global discount, per-arm staleness decay, other caps, other prior strengths.
Tests
tests/test_consolidated_v2_scheduler.py uses scripted draws, with a v1 control that picks differently on the same draws. It also covers the all-pessimistic adversarial case, tau limits, inheritance from v1, and end-to-end wiring.
v2 is added to the convergence harness: share floor 0.93 and recovery floor 0.95, both measured.
Precedence and ballot pins are updated. The stale las_vegas pin is fixed.
Known
Two tests fail on master too: test_every_exported_scheduler_is_covered and test_every_ballot_name_is_mapped. The cause is that --las-vegas is never offered by operator_strategy_pool(). This is logged in TODO and left out of scope here.
These results come from synthetic environments only. A real-target A/B on fuzzgoat is logged in TODO.
Version the consolidated scheduler as v1 while adding and integrating a higher-performing consolidated v2 scheduler with backward-compatible aliases.
New Features:
Add the --consolidated-v2 operator scheduler using optimistic, tempered Thompson scoring with v1-compatible learning and priors.
Bug Fixes:
Preserve legacy consolidated scheduler imports, constructor usage, CLI aliases, and statistics keys while transitioning to versioned schedulers.
Enhancements:
Rename the consolidated scheduler to v1 and expose versioned scheduler names and statistics.
Wire consolidated v1 and v2 through scheduler registration, shared learning fan-out, Elo ballots, no-Elo precedence, sweeps, and feature reporting, with v2 ahead of v1.
Update scheduler documentation and record synthetic benchmark results for v2 and the v1 comparison.
Documentation:
Document the v1 rename, v2 behavior, benchmark results, and follow-up real-target evaluation.
v1: op_consolidated.py -> op_consolidated_v1.py, ConsolidatedV1Scheduler,
strategy/flag consolidated_v1 (--consolidated kept as alias).
v2: v1 posterior scored by mean + tau * max(0, draw - mean), tau 0.65.
Tournament of all 41 op schedulers on 4 bandit_env environments picked
the ingredients; v2 beats v1 on all four (+1.3-1.8%, 20 paired seeds,
>=3 SE; v1-vs-v1 control within 2 SE). Leads no-Elo precedence.
Also pins las_vegas in the fallback-precedence test; its missing ballot
wiring is logged in TODO.
Co-Authored-By: Claude Opus 5.5 <[email protected]>
Claude-Session: https://claude.ai/code/session_019ape9xt5N7dHzpFqQcsNwk
The PR version-controls the consolidated operator scheduler, preserves the old v1 API/CLI aliases, introduces v2 as an inherited v1 posterior with optimistic tempered Thompson scoring, and wires both versions through CLI handling, Elo, fan-out, registration, precedence, diagnostics, convergence tests, and documentation.
Sequence diagram for consolidated v2 operator selection
sequenceDiagram
participant Fuzzer
participant Operators
participant V2 as ConsolidatedV2Scheduler
participant RNG as RandPool
Fuzzer->>Operators: select_op(ops)
Operators->>V2: select_op(ops)
V2->>V2: _indices(ops)
V2->>V2: _prior(idx)
V2->>RNG: betavariate_array(a, b)
RNG-->>V2: Thompson draws
V2->>V2: score = mean + tau * max(0, draw - mean)
V2-->>Operators: selected operator
Operators-->>Fuzzer: selected operator
Loading
Flow diagram for no-Elo consolidated precedence
flowchart TD
Start[Operator selection with Elo disabled]
V2{consolidated_v2 enabled}
V1{consolidated_v1 enabled}
ChooseV2[Use ConsolidatedV2Scheduler]
ChooseV1[Use ConsolidatedV1Scheduler]
Other[Continue to lower fallback schedulers]
Start --> V2
V2 -->|yes| ChooseV2
V2 -->|no| V1
V1 -->|yes| ChooseV1
V1 -->|no| Other
Loading
File-Level Changes
Change
Details
Files
Split the existing consolidated scheduler into an explicitly versioned v1 implementation while preserving the legacy import and CLI alias.
Moved the implementation to ConsolidatedV1Scheduler with versioned stats keys and updated references.
Kept ConsolidatedScheduler and --consolidated as compatibility aliases.
Renamed strategy, sweep, export, documentation, and test coverage identifiers to consolidated_v1.
Integrated both scheduler versions throughout CLI configuration, runtime construction, Elo arbitration, selection precedence, and operator registration.
Added --consolidated-v1 and --consolidated-v2; --elo-all/--hail-mary enable both versions.
Registered both schedulers for format priors, shared record fan-out, strategy banners, and no-Elo selection.
Placed v2 ahead of v1 in the no-Elo fallback precedence and Elo strategy ballot.
Updated regression fixtures and precedence/ballot pins, including the existing las_vegas mapping gap.
Trigger a new review: Comment @sourcery-ai review on the pull request.
Continue discussions: Reply directly to Sourcery's review comments.
Generate a GitHub issue from a review comment: Ask Sourcery to create an
issue from a review comment by replying to it. You can also reply to a
review comment with @sourcery-ai issue to create an issue from it.
Generate a pull request title: Write @sourcery-ai anywhere in the pull
request title to generate a title at any time. You can also comment @sourcery-ai title on the pull request to (re-)generate the title at any time.
Generate a pull request summary: Write @sourcery-ai summary anywhere in
the pull request body to generate a PR summary at any time exactly where you
want it. You can also comment @sourcery-ai summary on the pull request to
(re-)generate the summary at any time.
Generate reviewer's guide: Comment @sourcery-ai guide on the pull
request to (re-)generate the reviewer's guide at any time.
Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
pull request to resolve all Sourcery comments. Useful if you've already
addressed all the comments and don't want to see them anymore.
Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
request to dismiss all existing Sourcery reviews. Especially useful if you
want to start fresh with a new review - don't forget to comment @sourcery-ai review to trigger a new review!
Fuzzer keeps `consolidated` in its positional slot (pre-v2 name of v1);
`consolidated_v1`/`consolidated_v2` are appended. ConsolidatedScheduler
is exported and reports the pre-v2 stats keys.
Co-Authored-By: Claude Opus 5.5 <[email protected]>
Claude-Session: https://claude.ai/code/session_019ape9xt5N7dHzpFqQcsNwk
The real-target calibration is still pending even though v2 is wired into --elo all, --hail-mary, and the highest no-Elo precedence. Synthetic bandit environments do not validate target coverage or throughput; run the planned paired fuzzgoat A/B with a clang-built target and record the result before shipping this integration.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
op_consolidated.py→op_consolidated_v1.py,ConsolidatedV1Scheduler, strategy/flagconsolidated_v1/--consolidated-v1.--consolidatedandConsolidatedSchedulerremain as aliases (ImpactGuard).--consolidated-v2): v1 posterior, scored bymean + tau * max(0, draw - mean), tau 0.65. It is optimistic (a draw below the mean scores the mean) and tempered (explores less). Leads the no-Elo precedence. It is wired into the Elo ballot, the fan-out,_register_arms,--elo-all/--hail-maryand the banners.Why this ingredient
I ran a tournament of all 41 op schedulers on the 4
bandit_envenvironments (table in theop_consolidated_v2.pydocstring). v1 was the only scheduler in the leading group on all four. Its competitors beat it only where they explored less: MOSS, FPL, and greedy on rotting.On 20 paired seeds, v2 vs v1:
Changes that did not help: global discount, per-arm staleness decay, other caps, other prior strengths.
Tests
tests/test_consolidated_v2_scheduler.pyuses scripted draws, with a v1 control that picks differently on the same draws. It also covers the all-pessimistic adversarial case, tau limits, inheritance from v1, and end-to-end wiring.las_vegaspin is fixed.Known
mastertoo:test_every_exported_scheduler_is_coveredandtest_every_ballot_name_is_mapped. The cause is that--las-vegasis never offered byoperator_strategy_pool(). This is logged in TODO and left out of scope here.🤖 Generated with Claude Code
https://claude.ai/code/session_019ape9xt5N7dHzpFqQcsNwk
Generated by Claude Code
Summary by Sourcery
Version the consolidated scheduler as v1 while adding and integrating a higher-performing consolidated v2 scheduler with backward-compatible aliases.
New Features:
--consolidated-v2operator scheduler using optimistic, tempered Thompson scoring with v1-compatible learning and priors.Bug Fixes:
Enhancements:
Documentation:
Tests:
Chores: