Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -37,6 +37,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

### Added

- **`--consolidated-v2`** (`core/schedulers/op_consolidated_v2.py`): consolidated v1 scored by an optimistic, tempered Thompson draw. +1.3-1.8% over v1 on all four `bandit_env` environments (20 paired seeds). Leads no-Elo precedence.
- **Power Doppler schedule** (`--schedule doppler`, `core/power_doppler.py`): per-seed ensembles of mutant hit counts; mean + SVD wall filter, CFAR χ² flow detection; flow power scales seed energy to `[1, max_mult]`.
- **OS / network scheduler ports**: seed arms `mlfq`, `stride`, `eevdf`, `bfq`, `sfq`, `codel`, `aimd`, `p2c` (`--seed-<name>-scheduler`) and op arms `op_stride`, `op_p2c` (`--op-stride`, `--op-p2c`); `core/fair_queue.py` gains `Stride` and `EEVDF`. Elo arms, in `--hail-mary`; seed arms also run without `--elo`. Falsification and adversarial tests. No paired benchmark yet.
- **Ten op mutators** for in-tree targets and text decoders: `json_mutate` (fuzzgoat), `sql_mutate` (sqlite SQL path), `ecdsa_field_mutate` (secp256k1), `recompress_lz4`, `recompress_png_idat` (format band, sniffer-gated); `encoding_wrap`, `escape_mutate`, `ascii_float` (structural); `utf16_transcode`, `nest_bomb` (radamsa). Tests: `tests/test_{json_mutate,sql_mutate,ecdsa_field_mutate,recompress_roundtrip,text_codec,nest_bomb}.py`. Discovery effect on fuzzgoat unmeasured.
Expand All @@ -56,6 +57,10 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- **`core/group_testing.py`** + `tools/bench_group_testing.py`: non-adaptive pooled which-items-matter inference (COMP/DD, exact for any d) and a binary-splitting baseline. Not wired into `tmin`/colorizer (see handover section 6.5 result).
- **`--second-order-blend W`** (default 0 = off): second-order operator chain `P(next | prev2, prev)` in `MonteCarloScheduler` (`core/op_chain2.py`, sparse, capped at 4096 contexts), backing off to `--pairwise-blend` on unseen contexts. Synthetic A/B only; real-target A/B not run.

### Changed

- **`--consolidated` → `--consolidated-v1`** (`op_consolidated.py` → `op_consolidated_v1.py`, `ConsolidatedScheduler` → `ConsolidatedV1Scheduler`, strategy `consolidated` → `consolidated_v1`, stats keys `consolidated_v1_*`). `--consolidated` kept as an alias. Elo ratings saved under `consolidated` do not carry over.

### Documentation

- **Combinatorics gap analysis** (`docs/handover/handover_combinatorics_permutations_2026-09-02.md` section 6): ranked remaining candidates (FIC-style isolation over covering-array rows, covering-array extensions, orthogonal designs, rank/unrank, group testing) and what to skip. Docs only.
Expand Down
2 changes: 2 additions & 0 deletions docs/DEEP_DIVE.md
Original file line number Diff line number Diff line change
Expand Up @@ -220,6 +220,8 @@ For production and sensitive binaries using AFL family fuzzers is the best cours
- **KL-UCB variants** (`--kl-ducb`, `--kl-ducb-gamma`, `--kl-swucb`, `--kl-swucb-window`): the same two policies with the Gaussian confidence width replaced by the empirical-Bernoulli KL upper bound (Cappé, Garivier, Maillard, Munos, Stoltz, 2013). The cost-adjusted surprisal rewards handed to `record()` are bounded in [0, 1] with a mass at zero — an operator that found no new coverage gets exactly 0 — so the true tail is heavier than the Gaussian one and the Gaussian width under-covers. The KL bound is the smallest q in [p, 1] with KL(p||q) >= xi*log(n)/N, solved by bisection in `core/schedulers/_kl_ucb.py`; it is a strict tightening of the Gaussian form, so it explores less and concentrates on the best arm. It is a separate scheduler rather than a flag on DUCB/SWUCB because the two forms are mutually exclusive and `--elo` arbitrates them as distinct strategies. `tools/measure_klucb_signal.py` runs all four side by side on a Bernoulli environment. Internally, D-UCB and KL-D-UCB share `DiscountedUCBBase` for discounted state and width computation; SW-UCB and KL-SW-UCB share `WindowedUCBBase` for windowed state — both in `core/schedulers/ucb_common.py`.
- **MOSS** (`--moss`, `--moss-gamma`): Audibert & Bubeck's minimax-optimal UCB, anytime form (Degenne & Perchet 2016). The width is `c*sqrt(log+(t/(K*n_i))/n_i)`: zero once an operator holds its fair share t/K of pulls, so the policy stops re-opening every one of ~200 operators as t grows. On a 150-arm environment at fuzzing rates (0.002-0.08) it scored 3992 successes in 60k pulls against Consolidated 3928, D-UCB 756 and SW-UCB 659 (uniform ~560); the discounted/windowed UCBs keep too few pulls per arm at that K to separate good operators from bad. Undiscounted by default, so it adapts slowly to decaying yields (0.34 tail recovery on `DecayingBest`); `--moss-gamma 0.99995` raises that to 0.72 at a 22% cost on 150 arms. Numpy state, 21.5µs per select+record at K=200. Measurements in `core/schedulers/op_moss.py`.
- **Bayes-UCB** (`--bayes-ucb`): Kaufmann et al. 2012. Scores an operator by the `1 - 1/(t (log t)^c)` quantile of its Beta posterior. Deterministic (no rng), `supports_priors = True` so format-operator priors bias the first quantiles. On the Elo ballot and in the fallback chain after MOSS. **Not A/B validated.** Details in `core/schedulers/op_bayes_ucb.py`.
- **Consolidated v1** (`--consolidated-v1`, alias `--consolidated`): flat Thompson over Beta evidence with a category-shrunk prior and a capped pseudocount (forgets by rescaling past 200). Only scheduler in the leading group on all four `bandit_env` environments. `core/schedulers/op_consolidated_v1.py`.
- **Consolidated v2** (`--consolidated-v2`): v1's posterior scored by `mean + tau * max(0, draw - mean)`, tau 0.65 — optimistic (a below-mean draw scores the mean) and tempered (less exploration, the edge MOSS/FPL/greedy had over v1). 20 paired seeds vs v1: +1.8% stationary, +1.6% decaying, +1.6% rotting, +1.3% Fatigue150, each ≥3 SE; a v1-vs-v1 control stays within 2 SE. +~5µs per select at K=200 (the Beta draw dominates). Leads the no-Elo precedence, ahead of v1. Synthetic only; real-target A/B pending. Tournament of all 41 operator schedulers in `core/schedulers/op_consolidated_v2.py`.
- **BO-GP-UCB** (`--bo-gp-ucb`, `--bo-gp-length-scale`, `--bo-gp-noise`): Bayesian Optimization with Expected Improvement acquisition and noisy Gaussian Process posterior. EI(op) = (μ(op) - f_max)Φ(z) + σ(op)φ(z) where z = (μ(op) - f_max)/σ(op). Cholesky decomposition for stable GP posterior inference. The `noise` parameter adds σ² to the kernel diagonal, distinguishing observation noise from epistemic uncertainty — higher noise increases posterior variance and encourages exploration. Unlike UCB schedulers that balance exploration/exploitation via a β parameter, EI naturally balances both without a separate knob. `supports_priors = False` — GP posterior is computed from observations, not Beta priors. `tools/measure_bo_gp_ucb_signal.py` runs it side by side with other schedulers.
- **Badness-indexed exploration floor** (`core/badness_floor.py`; `--op-katz`, `--op-kuramoto`): both schedulers floor each arm at `lambda(badness)/n`, `lambda` interpolating `explore_floor` (SUPERCRITICAL corpus) to `max_explore_floor` (SUBCRITICAL, stalled). Badness comes from `Fuzzer._current_scheduling_badness`; a raising source falls back to the static floor.
- **PLL monitor** (`--pll`, `core/analyzers/analyzer_pll.py`, 2026-09-24): observation layer over `core/pll.py`. Exec times are pushed per execution (one array append), discovery-rate deltas at each stats tick; each series buffers 256 samples, bootstraps a `PhaseLockedLoop` from `detect_periodicity` (retries on a miss), and logs lock/unlock transitions with the stall-recovery flag. `stall_lift = P(stall | transition) / P(stall)`. Consumers: stats field `pll: t:<period><L|u> d:…` and a "PLL" line per series under Spectral Diagnostics (tracked period next to the batch FFT). Read-only. Loops, counters and pending samples persist (`pll` state section; `PhaseLockedLoop.to_dict/from_dict`). fuzzgoat: exec-time series bootstraps at period ≈ 5, stays unlocked.
Expand Down
4 changes: 3 additions & 1 deletion docs/TODO.md
Original file line number Diff line number Diff line change
Expand Up @@ -49,7 +49,9 @@
- [ ] **PLL: correlate lock transitions with stalls on real campaigns** (2026-09-24) — `--pll` logs transitions and `stall lift`; collect over ≥ 3 targets before any behavioural consumer or gain tuning. Exec-time periods near 5 samples sit outside the tuned 15-30 range.
- [ ] **A/B `wfc_reorder_learned` now that it learns** (2026-09-24) — admission never called `notify_new_coverage`, so its tables were always empty in real runs and every prior `--wfc` measurement of it is void. Re-run paired `--wfc` cells on an isobmff/riff target (needs `ffmpeg_read`).
- [ ] **A/B the badness floor on `--op-kuramoto`** (2026-09-24) — now wired like `--op-katz`, unmeasured. Re-run the `tests/support/bandit_env.py` lock-in harness with regime pinned SUBCRITICAL vs SUPERCRITICAL; then paired cells on fuzzgoat.
- [ ] **Re-run the scheduler tournament on `RottingArms` / `Fatigue150`** (2026-09-24) — both now in `tests/support/bandit_env.py`. `Fatigue150` is a reconstruction, not the original: `op_consolidated.py`'s 150-arm numbers are not comparable until re-measured. Then non_ucb step 3 (reward vs own pulls from the ablation CSV `operator` column).
- [x] **Re-run the scheduler tournament on `RottingArms` / `Fatigue150`** (2026-09-24, run 2026-10-01) — all 41 operator schedulers, 4 environments, 3 seeds: table in `core/schedulers/op_consolidated_v2.py`. Consolidated (now v1) leads overall; produced `--consolidated-v2`. Still open: non_ucb step 3 (reward vs own pulls from the ablation CSV `operator` column).
- [ ] **A/B `--consolidated-v2` vs `--consolidated-v1` on a real target** (2026-10-01) — +1.3-1.8% on all four synthetic environments; paired `bench_paired` cells on fuzzgoat (clang) next. Fatigue150 headroom is unlock detection (v1/v2 notice an unlock ~1500 rounds late; greedy oracle 1692 vs v2 1544); staleness decay and global discount both lost.
- [ ] **`--las-vegas` is never offered** (2026-10-01) — in `_FALLBACK_PRECEDENCE` and the dispatch chain but missing from `operator_strategy_pool()`, the `_track_op_effect` gate, `test_regression_scheduler_operator_reach` and `test_regression_track_op_effect_coverage._KWARGS`; the last two fail on main. Wire it like `--moss`.
- [ ] **A/B `--op-tpe`** (2026-09-24) — BO-3 shipped unmeasured. Paired cells vs `--bo-gp-ucb` under `--elo`; sweep `gamma` (0.1/0.25/0.5) and window.
- [ ] **A/B `--op-afl-det`** (2026-09-24) — T1-1 arm shipped unmeasured. Compare cold-start edge curves (first 5 min) on fuzzgoat/png with the arm on vs off; the case for it is cold start and plateaus, not steady state.
- [ ] **Entropy LOO vs Shapley disagreement, then A/B `--entropy-loo`** (2026-09-24) — measure how often LOO and permutation-sampled Shapley rank seeds differently on real corpora (near-duplicate clusters are the known misprice) before building Shapley; then paired A/B. Also open: use the score to protect seeds in minimization.
Expand Down
23 changes: 18 additions & 5 deletions src/fuzzer_tool/cli/commands.py
Original file line number Diff line number Diff line change
Expand Up @@ -457,7 +457,8 @@ def cmd_fuzz(args):
args.successive_elim = True
args.las_vegas = True
args.canary_scheduler = True
args.consolidated = True
args.consolidated_v1 = True
args.consolidated_v2 = True
args.moss = True
args.bayes_ucb = True
args.contextual = True
Expand Down Expand Up @@ -724,7 +725,8 @@ def cmd_fuzz(args):
shaped_reward_floor=getattr(args, "shaped_reward_floor", 0.0),
continuum_reward=getattr(args, "continuum_reward", False),
continuum_reward_floor=getattr(args, "continuum_reward_floor", 0.0),
consolidated=getattr(args, "consolidated", False),
consolidated_v1=getattr(args, "consolidated_v1", False),
consolidated_v2=getattr(args, "consolidated_v2", False),
moss=getattr(args, "moss", False),
moss_gamma=getattr(args, "moss_gamma", 1.0),
bayes_ucb=getattr(args, "bayes_ucb", False),
Expand Down Expand Up @@ -2058,7 +2060,8 @@ def cmd_sweep(args):
"seed_p2c_scheduler",
"softmax",
"topk",
"consolidated",
"consolidated_v1",
"consolidated_v2",
"moss",
"bayes_ucb",
"contextual",
Expand Down Expand Up @@ -3452,12 +3455,22 @@ def main() -> int:
),
)
fuzz_parser.add_argument(
"--consolidated-v1",
"--consolidated",
action="store_true",
help=(
"Enable the consolidated operator scheduler: Thompson sampling with "
"Enable the consolidated_v1 operator scheduler: Thompson sampling with "
"a category-shrunk prior and capped evidence. Takes precedence over "
"every other operator scheduler when Elo is off"
"every other operator scheduler but consolidated_v2 when Elo is off"
),
)
fuzz_parser.add_argument(
"--consolidated-v2",
action="store_true",
help=(
"Enable the consolidated_v2 operator scheduler: consolidated_v1 scored "
"by an optimistic, tempered Thompson draw. Takes precedence over every "
"other operator scheduler when Elo is off"
),
)
fuzz_parser.add_argument(
Expand Down
4 changes: 4 additions & 0 deletions src/fuzzer_tool/core/schedulers/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,8 @@
from fuzzer_tool.core.schedulers.op_canary import CanaryScheduler
from fuzzer_tool.core.schedulers.op_cmaes import CMAESScheduler
from fuzzer_tool.core.schedulers.op_consolidated import ConsolidatedScheduler
from fuzzer_tool.core.schedulers.op_consolidated_v1 import ConsolidatedV1Scheduler
from fuzzer_tool.core.schedulers.op_consolidated_v2 import ConsolidatedV2Scheduler
from fuzzer_tool.core.schedulers.op_contextual import ContextualLinUCBScheduler
from fuzzer_tool.core.schedulers.op_corral import CorralScheduler
from fuzzer_tool.core.schedulers.op_cucb import CUCBScheduler
Expand Down Expand Up @@ -44,6 +46,8 @@
"FPLScheduler",
"CMAESScheduler",
"ConsolidatedScheduler",
"ConsolidatedV1Scheduler",
"ConsolidatedV2Scheduler",
Comment thread
daedalus marked this conversation as resolved.
"C2UCBScheduler",
"MonteCarloScheduler",
"MOptScheduler",
Expand Down
Loading
Loading