Skip to content

Rename consolidated scheduler to v1; add consolidated_v2 - #53

Merged
daedalus merged 3 commits into
masterfrom
claude/peaceful-edison-grmqia
Oct 1, 2026
Merged

daedalus merged 3 commits into
masterfrom
claude/peaceful-edison-grmqia

Conversation

@daedalus

@daedalus daedalus commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

What

  • v1: op_consolidated.py → op_consolidated_v1.py, ConsolidatedV1Scheduler, strategy/flag consolidated_v1 / --consolidated-v1. --consolidated and ConsolidatedScheduler remain as aliases (ImpactGuard).
  • v2 (--consolidated-v2): v1 posterior, scored by mean + tau * max(0, draw - mean), tau 0.65. It is optimistic (a draw below the mean scores the mean) and tempered (explores less). Leads the no-Elo precedence. It is wired into the Elo ballot, the fan-out, _register_arms, --elo-all/--hail-mary and the banners.

Why this ingredient

I ran a tournament of all 41 op schedulers on the 4 bandit_env environments (table in the op_consolidated_v2.py docstring). v1 was the only scheduler in the leading group on all four. Its competitors beat it only where they explored less: MOSS, FPL, and greedy on rotting.

On 20 paired seeds, v2 vs v1:

env v1 v2 gain v2 wins
stationary 1729 1759 +1.8% 19/20
decaying 4612 4685 +1.6% 20/20
rotting 3502 3558 +1.6% 17/20
fatigue150 1524 1544 +1.3% 16/20
  • Each gain is at least 3 SE of the paired difference.
  • A v1-vs-v1 control stays within 2 SE.
  • Cost is about 5µs per select at K=200.

Changes that did not help: global discount, per-arm staleness decay, other caps, other prior strengths.

Tests

  • tests/test_consolidated_v2_scheduler.py uses scripted draws, with a v1 control that picks differently on the same draws. It also covers the all-pessimistic adversarial case, tau limits, inheritance from v1, and end-to-end wiring.
  • v2 is added to the convergence harness: share floor 0.93 and recovery floor 0.95, both measured.
  • Precedence and ballot pins are updated. The stale las_vegas pin is fixed.

Known

  • Two tests fail on master too: test_every_exported_scheduler_is_covered and test_every_ballot_name_is_mapped. The cause is that --las-vegas is never offered by operator_strategy_pool(). This is logged in TODO and left out of scope here.
  • These results come from synthetic environments only. A real-target A/B on fuzzgoat is logged in TODO.

🤖 Generated with Claude Code

https://claude.ai/code/session_019ape9xt5N7dHzpFqQcsNwk


Generated by Claude Code

Summary by Sourcery

Version the consolidated scheduler as v1 while adding and integrating a higher-performing consolidated v2 scheduler with backward-compatible aliases.

New Features:

  • Add the --consolidated-v2 operator scheduler using optimistic, tempered Thompson scoring with v1-compatible learning and priors.

Bug Fixes:

  • Preserve legacy consolidated scheduler imports, constructor usage, CLI aliases, and statistics keys while transitioning to versioned schedulers.

Enhancements:

  • Rename the consolidated scheduler to v1 and expose versioned scheduler names and statistics.
  • Wire consolidated v1 and v2 through scheduler registration, shared learning fan-out, Elo ballots, no-Elo precedence, sweeps, and feature reporting, with v2 ahead of v1.
  • Update scheduler documentation and record synthetic benchmark results for v2 and the v1 comparison.

Documentation:

  • Document the v1 rename, v2 behavior, benchmark results, and follow-up real-target evaluation.

Tests:

  • Add targeted v2 scoring, parameter validation, inheritance, reproducibility, convergence, precedence, ballot, and end-to-end wiring coverage.
  • Update scheduler and compatibility tests for versioned names and legacy aliases.

Chores:

  • Mark the synthetic scheduler tournament follow-up complete and track real-target v2 evaluation separately.

claude added 2 commits October 1, 2026 16:22
v1: op_consolidated.py -> op_consolidated_v1.py, ConsolidatedV1Scheduler,
strategy/flag consolidated_v1 (--consolidated kept as alias).

v2: v1 posterior scored by mean + tau * max(0, draw - mean), tau 0.65.
Tournament of all 41 op schedulers on 4 bandit_env environments picked
the ingredients; v2 beats v1 on all four (+1.3-1.8%, 20 paired seeds,
>=3 SE; v1-vs-v1 control within 2 SE). Leads no-Elo precedence.

Also pins las_vegas in the fallback-precedence test; its missing ballot
wiring is logged in TODO.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Claude-Session: https://claude.ai/code/session_019ape9xt5N7dHzpFqQcsNwk
ImpactGuard flagged the rename as removing ConsolidatedScheduler. The
old module path and package name now alias ConsolidatedV1Scheduler.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Claude-Session: https://claude.ai/code/session_019ape9xt5N7dHzpFqQcsNwk
@sourcery-ai

sourcery-ai Bot commented Oct 1, 2026

Copy link
Copy Markdown

Reviewer's Guide

The PR version-controls the consolidated operator scheduler, preserves the old v1 API/CLI aliases, introduces v2 as an inherited v1 posterior with optimistic tempered Thompson scoring, and wires both versions through CLI handling, Elo, fan-out, registration, precedence, diagnostics, convergence tests, and documentation.

Sequence diagram for consolidated v2 operator selection

sequenceDiagram
    participant Fuzzer
    participant Operators
    participant V2 as ConsolidatedV2Scheduler
    participant RNG as RandPool

    Fuzzer->>Operators: select_op(ops)
    Operators->>V2: select_op(ops)
    V2->>V2: _indices(ops)
    V2->>V2: _prior(idx)
    V2->>RNG: betavariate_array(a, b)
    RNG-->>V2: Thompson draws
    V2->>V2: score = mean + tau * max(0, draw - mean)
    V2-->>Operators: selected operator
    Operators-->>Fuzzer: selected operator
Loading

Flow diagram for no-Elo consolidated precedence

flowchart TD
    Start[Operator selection with Elo disabled]
    V2{consolidated_v2 enabled}
    V1{consolidated_v1 enabled}
    ChooseV2[Use ConsolidatedV2Scheduler]
    ChooseV1[Use ConsolidatedV1Scheduler]
    Other[Continue to lower fallback schedulers]

    Start --> V2
    V2 -->|yes| ChooseV2
    V2 -->|no| V1
    V1 -->|yes| ChooseV1
    V1 -->|no| Other
Loading

File-Level Changes

Change Details Files
Split the existing consolidated scheduler into an explicitly versioned v1 implementation while preserving the legacy import and CLI alias.
  • Moved the implementation to ConsolidatedV1Scheduler with versioned stats keys and updated references.
  • Kept ConsolidatedScheduler and --consolidated as compatibility aliases.
  • Renamed strategy, sweep, export, documentation, and test coverage identifiers to consolidated_v1.
src/fuzzer_tool/core/schedulers/op_consolidated.py
src/fuzzer_tool/core/schedulers/op_consolidated_v1.py
src/fuzzer_tool/core/schedulers/__init__.py
src/fuzzer_tool/services/fuzzer.py
src/fuzzer_tool/services/operators.py
src/fuzzer_tool/cli/commands.py
tests/test_consolidated_scheduler.py
tests/test_consolidated_v1_scheduler.py
tests/support/operator_env.py
Added v2 as a subclass of v1 that modifies only candidate scoring to reduce pessimistic Thompson exploration.
  • Scores each Beta draw as mean + 0.65 * max(0, draw - mean) with validation for tau in (0, 1].
  • Retains v1 posterior updates, category priors, capped evidence, format priors, and shared reward fan-out through inheritance.
  • Adds scripted-draw, limit, inheritance, reproducibility, diagnostics, and end-to-end wiring tests.
  • Adds synthetic tournament and paired-seed benchmark results, while documenting that real-target validation remains outstanding.
src/fuzzer_tool/core/schedulers/op_consolidated_v2.py
tests/test_consolidated_v2_scheduler.py
docs/DEEP_DIVE.md
CHANGELOG.md
docs/TODO.md
Integrated both scheduler versions throughout CLI configuration, runtime construction, Elo arbitration, selection precedence, and operator registration.
  • Added --consolidated-v1 and --consolidated-v2; --elo-all/--hail-mary enable both versions.
  • Registered both schedulers for format priors, shared record fan-out, strategy banners, and no-Elo selection.
  • Placed v2 ahead of v1 in the no-Elo fallback precedence and Elo strategy ballot.
  • Updated regression fixtures and precedence/ballot pins, including the existing las_vegas mapping gap.
src/fuzzer_tool/cli/commands.py
src/fuzzer_tool/services/fuzzer.py
src/fuzzer_tool/services/operators.py
tests/test_regression_elo_all.py
tests/test_regression_scheduler_fallback_precedence.py
tests/test_regression_scheduler_operator_reach.py
tests/test_regression_track_op_effect_coverage.py
Updated scheduler convergence and supporting documentation to reflect versioned names and v2 performance floors.
  • Added v2 reliability and recovery thresholds to the convergence harness.
  • Renamed environment and explanatory references from the unversioned implementation to v1.
  • Recorded the known out-of-scope las_vegas test failures and pending real-target A/B.
tests/test_scheduler_convergence.py
tests/support/bandit_env.py
tests/test_bandit_env_fatigue.py
tests/test_canary_scheduler.py
tests/test_corral_scheduler.py
docs/DEEP_DIVE.md
docs/TODO.md

Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@daedalus
daedalus marked this pull request as ready for review October 1, 2026 17:18
Copilot AI balanced review requested due to automatic review settings October 1, 2026 17:18

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry @daedalus, you've used your own review budget of 250,000 diff characters for the last 7 days.

You can request another review in 1 day and 22 hours by commenting @sourcery-ai review. Upgrade to get a review now.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Constructor and scheduler compatibility aliases currently break existing API consumers.

Review effort: Balanced
Findings: 1 High severity · 2 Medium severity

Open (3)
What changed in this PR

Renames the consolidated scheduler to v1, adds optimistic-tempered v2, and wires both through CLI, Elo, fallback selection, diagnostics, and tests.

Changes:

  • Adds ConsolidatedV2Scheduler.
  • Renames and preserves v1 compatibility aliases.
  • Expands scheduler wiring, documentation, and regression coverage.
File Description
CHANGELOG.md Records v2 and v1 rename.
docs/​DEEP_DIVE.md Documents both schedulers.
docs/​TODO.md Records benchmarks and follow-ups.
src/​fuzzer_tool/​cli/​commands.py Adds flags and presets.
src/​fuzzer_tool/​core/​schedulers/​__init__.py Exports scheduler classes.
src/​fuzzer_tool/​core/​schedulers/​op_consolidated.py Provides legacy alias module.
src/​fuzzer_tool/​core/​schedulers/​op_consolidated_v1.py Defines renamed v1 scheduler.
src/​fuzzer_tool/​core/​schedulers/​op_consolidated_v2.py Implements v2 scoring.
src/​fuzzer_tool/​core/​schedulers/​op_corral.py Updates references.
src/​fuzzer_tool/​services/​fuzzer.py Constructs and feeds both schedulers.
src/​fuzzer_tool/​services/​operators.py Adds ballot and fallback dispatch.
tests/​support/​bandit_env.py Updates environment references.
tests/​support/​operator_env.py Adds both ballot names.
tests/​test_bandit_env_fatigue.py Updates v1 reference.
tests/​test_canary_scheduler.py Uses renamed v1 class.
tests/​test_consolidated_v1_scheduler.py Updates v1 tests and wiring.
tests/​test_consolidated_v2_scheduler.py Adds v2 behavior and wiring tests.
tests/​test_corral_scheduler.py Updates source reference.
tests/​test_regression_elo_all.py Covers both Elo entries.
tests/​test_regression_scheduler_fallback_precedence.py Pins v2-before-v1 precedence.
tests/​test_regression_scheduler_operator_reach.py Covers both scheduler classes.
tests/​test_regression_track_op_effect_coverage.py Covers both tracking flags.
tests/​test_scheduler_convergence.py Adds v2 convergence thresholds.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread src/fuzzer_tool/services/fuzzer.py Outdated
Comment thread src/fuzzer_tool/core/schedulers/__init__.py
Comment thread src/fuzzer_tool/core/schedulers/op_consolidated_v1.py
Fuzzer keeps `consolidated` in its positional slot (pre-v2 name of v1);
`consolidated_v1`/`consolidated_v2` are appended. ConsolidatedScheduler
is exported and reports the pre-v2 stats keys.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Claude-Session: https://claude.ai/code/session_019ape9xt5N7dHzpFqQcsNwk

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Real-target calibration and architecture documentation remain incomplete.

Review effort: Balanced
Findings: 1 Low severity

Open (1)
Resolved since last review (3)
Previously missed (1)

In code that hasn't changed since last review

Medium severity Complete real-target calibration before shipping v2 integration

src/​fuzzer_tool/​core/​schedulers/​op_consolidated_v2.py:71

The real-target calibration is still pending even though v2 is wired into --elo all, --hail-mary, and the highest no-Elo precedence. Synthetic bandit environments do not validate target coverage or throughput; run the planned paired fuzzgoat A/B with a clang-built target and record the result before shipping this integration.

Comment thread src/fuzzer_tool/core/schedulers/op_consolidated_v2.py

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The implementation preserves compatibility and has focused algorithm, wiring, and regression coverage.

Review effort: Balanced
Findings: None

Resolved since last review (1)

@daedalus
daedalus merged commit de0c5fc into master Oct 1, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants