Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,10 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

### Fixed

- **EEVDF pick scanned ineligible flows** (`core/fair_queue.py`): one deadline heap popped every flow with an earlier deadline but `ve > V` (5000 pops at 5000 flows). Now a `ve` heap feeds a deadline heap; amortized O(log n). Test: `test_eevdf_pick_does_not_scan_ineligible_flows`.
- **Seed-arm ledgers grew without bound**: non-corpus (Markov) parents were recorded, and departed seeds never left `ArmCounts`. Only corpus parents are recorded now; ledgers trim to 2x the live corpus.
- **Round robin was O(n^2) per pick** (`seed_round_robin`, `op_round_robin`): `x in list` per registered arm. Set membership now: 103 ms -> 0.7 ms per pick at 5000 seeds.
- **`cuckoo_seed_filter` broke 12 corpus-minimization / lineage tests** (`services/corpus_manager.py` read it unguarded) and was missing, with `swap_walk`, from `_HAIL_MARY_FLAGS`.
- **Stack depth was always 0** (`adapters/afl_shim.c`): `__afl_max_stack_depth` was reset and copied to SHM offset 0 but never assigned, so `read_stack_depth()` returned 0 and the stack-depth boost in `core/schedules.py` never fired. The shim now samples the frame address in `__afl_map_loc` (base = first sample after reset, live write). Tests: `tests/test_shim_stack_depth.py`. Per-edge cost unmeasured; boost effect on discovery untested.

### Added (sancov)
Expand Down
2 changes: 1 addition & 1 deletion docs/DEEP_DIVE.md
Original file line number Diff line number Diff line change
Expand Up @@ -208,7 +208,7 @@ For production and sensitive binaries using AFL family fuzzers is the best cours
| `p2c` / `op_p2c` | two uniform draws, higher Beta(1,1) mean wins | uniform with no signal |
| `op_stride` | stride with Beta(1,1) posterior mean as tickets | RR with no signal |

`sfq` groups siblings only under `--lineage` (else every seed is its own flow). Pick cost at 5000 seeds: 4-40 us vs 320 us for DRR. Not persisted. Unmeasured: no paired benchmark yet.
`sfq` groups siblings only under `--lineage` (else every seed is its own flow). Only corpus parents are recorded (Markov-generated inputs are skipped), and every arm trims its ledger to 2x the live corpus. EEVDF keeps two heaps (`ve` order, then deadline order) so ineligible flows with early deadlines are never scanned. Pick cost at 5000 seeds: 4-40 us vs 320 us for DRR. Not persisted. Unmeasured: no paired benchmark yet.
- **Seed-level energy multiplier** (`SeedScorer`): AFL++ power schedules (FAST/COE/RARE/MMOPT/LIN/QUAD) scale `mutations_per_input` per seed — fast seeds with high coverage get more mutation attempts, heavily-fuzzed seeds get fewer. Honggfuzz power factors (novelty decay, density, fertility, freshness, CMP progress, entropy penalty, timeout penalty) applied multiplicatively on top of schedule scoring
- **Elo arbitration** (`--elo`): combined operator + seed scheduling via Bayesian Elo rating system. All available strategies (bandit, mopt, replicator, cem, exp3, eps_greedy, hierarchical, gp_ucb, bo_gp_ucb, softmax, topk, cmaes, contextual, ducb, swucb, cucb for operators; weighted, pareto, format, ga, qea, bayesian, markov for seeds) run in shadow; Elo Thompson-samples which to trust each iteration. Ratings decayed periodically to model non-stationarity. Competition is **group-scoped**: operator schedulers only compete against other operator schedulers, seed strategies only against other seed strategies (namespaced `seed_<name>`), and the random fallback (stall-recovery `random_stall`, random seed pick) is the only non-competitor in either group. `--elo all` enables every scheduler and mutation-stack feature (metropolis, shapley, mi-guided, secretary, wfc, lineage, `--schedule fast`) rather than merely listing them as available; the per-run convergence report prints only schedulers actually selected. **Schedulers are independent**: each scheduler in `core/schedulers/` never imports or calls another scheduler (they compete; only Elo arbitrates between them), sharing only neutral dependencies like the operator→category taxonomy in `core/operator_categories.py` (derived from the operator registry; imported by the hierarchical and gp-ucb schedulers)
- **EXP3 adversarial bandit** (`--exp3`): adversarial bandit algorithm for operator selection in non-stationary reward environments. Uses importance-weighted rewards with exponential weight updates — automatically tracks the best operator even when reward distributions shift over time. Exploration rate `gamma` controls the trade-off (default 0.1, range [0,1]). Window decay discounts old observations to adapt to changing conditions.
Expand Down
4 changes: 2 additions & 2 deletions docs/TODO.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,9 +26,9 @@

## Scheduling
- [ ] **A/B `--schedule doppler`** (2026-10-01) — power Doppler seed energy (`core/power_doppler.py`) is unit-tested and wired; never measured. Paired `bench_paired.py` vs `fast` on clang-built fuzzgoat (Hard Rule 52). Open: (a) ~100 µs/exec at 2k edges, mostly `get_edge_counts()` dict → numpy; a numpy accessor on `ShmCoverage` (exposes `_active_columns`, Hard Rule 34 — needs approval) removes it; (b) ensembles span picks, so a seed needs 32 picks to score — tune `ensemble` vs pick rate; (c) unused: `flow_edges()` (input-sensitive edge set) could feed position/operator targeting; (d) scores not persisted across `--resume`.
- [ ] **A/B the OS / network scheduler ports** (2026-09-30) — `mlfq`, `stride`, `eevdf`, `bfq`, `sfq`, `codel`, `aimd`, `p2c` seed arms and `op_stride`, `op_p2c` are unit- and wiring-tested only. Run `tools/lib/bench_paired.py` per arm vs `seed_round_robin` / `drr` on fuzzgoat (clang, ASAN). Untuned: MLFQ allotment/boost, BFQ budget range, CoDel target/interval, AIMD alpha/beta/loss run. Open: (a) `sfq` groups only direct siblings (`parent_key`), and only under `--lineage`; a lineage-root flow would group whole families; (b) no arm persists state; (c) `seed_round_robin.select_seed` is O(n^2) per pick (`s in seed_ids` on a list): ~100 ms at 5000 seeds.
- [ ] **A/B the OS / network scheduler ports** (2026-09-30) — `mlfq`, `stride`, `eevdf`, `bfq`, `sfq`, `codel`, `aimd`, `p2c` seed arms and `op_stride`, `op_p2c` are unit- and wiring-tested only. Run `tools/lib/bench_paired.py` per arm vs `seed_round_robin` / `drr` on fuzzgoat (clang, ASAN). Untuned: MLFQ allotment/boost, BFQ budget range, CoDel target/interval, AIMD alpha/beta/loss run. Open: (a) `sfq` groups only direct siblings (`parent_key`), and only under `--lineage`; a lineage-root flow would group whole families; (b) no arm persists state.
- [ ] **A/B the effector/token/chunk/changed/rare_mask position arms** (2026-09-30) — wired and unit-tested; a 3k-exec fuzzgoat run confirms live signal (`changed`: 673 moved / 96 unmoved / 2231 unmeasured rounds; `rare_mask`: target on 328/3000 rounds) but fuzzgoat cannot rank position arms (see `pos_fibonacci` entry). Run `tools/lib/bench_paired.py` `pos-arena-{token,chunk,changed,rare-mask}` vs `pos-arena-uniform` on png_read (`chunk` needs a container corpus). Open: (a) `effector` had no drained map in 3k execs (SkipDet skipped 111/112 seeds) -- measure its reach on long runs before A/B, it is not subset-testable (needs the det stage); (b) `changed` credits only ~25% of rounds: parents without a recorded path hash (spliced/generated inputs) -- record one at admission or accept; (c) neither `changed` nor `rare_mask` persists state.
- [ ] **Pre-existing, found 2026-09-30:** `test_regression_hail_mary_gates.py` fails on base (`cuckoo_seed_filter`, `swap_walk` missing from `_HAIL_MARY_FLAGS` or an exclusion); `tools/build_targets.sh` ASAN variants fail in the cloud container (non-ASAN builds fine).
- [ ] **Pre-existing, found 2026-09-30:** `tools/build_targets.sh` ASAN variants fail in the cloud container (non-ASAN builds fine).
- [ ] **Target arena unmeasured** (2026-09-30) — `--target-arena` (Elo over target schedules + Gale-Shapley) has unit tests and a 2-target smoke run only (uninstrumented, no signal). Needs a paired A/B vs `--target-schedule weighted` on ≥2 instrumented ASAN fuzzgoat builds (this container lacks the ASAN runtime). Gale-Shapley batch (32) and target taste (fewest tries) are untuned. `auction` arm (max-weight on Thompson draws, 2026-09-30) same status: A/B it against `gale_shapley` in that run (Elo share + edges); if it wins, consider sharing one yield-row table between the two arms (each keeps its own LRU today, 2x memory and record cost).
- [ ] **E3 on ffmpeg_read** (2026-09-27) — png is null (`docs/learnings/2026-09-27-alphabeta-vs-mcts-png.md`: 12W/8L, p=0.50, Δ inside the rep noise). Vendor + build `ffmpeg_read` (`tools/vendor_ffmpeg.sh`, `tools/build_targets.sh`). No target set in `tools/lib/eval_set.py` contains ffmpeg: add a new named set (e.g. `ffmpeg_signal`, register it in `TARGET_SETS`; do not edit an existing set) holding the built `ffmpeg_read.so` (or `ffmpeg_read_nosan.so`) path with `-m 65536`, then run `--arms elo-mcts,elo-alphabeta --set ffmpeg_signal --reps 2 --lock-single-thread`. If null again, retire `--alphabeta` in favour of `--mcts`.
- [ ] **`--vendor-tracecmp` cannot configure vendored libpng** (2026-09-27) — `tools/build_targets.sh` passes `-fsanitize-coverage=trace-pc-guard` in `CFLAGS` to libpng's `./configure`; configure's link test has no provider for `__sanitizer_cov_trace_pc_guard*`, so it dies with "C compiler cannot create executables" (and the `-L../zlib` zlib probe fails the same way against the instrumented `libz.a`). Worked around by hand: configure plain, then `make libpng16.la CFLAGS=<instrumented>`. Fix in the script the same way. Also: `tools/vendor_libpng.sh` fetches zlib as a GitHub tarball, which the sandbox proxy 403s; `git clone --branch v1.3.1` works.
Expand Down
2 changes: 2 additions & 0 deletions src/fuzzer_tool/cli/commands.py
Original file line number Diff line number Diff line change
Expand Up @@ -2040,6 +2040,8 @@ def cmd_sweep(args):
"seed_canary_scheduler",
"seed_round_robin_scheduler",
"seed_drr_scheduler",
"cuckoo_seed_filter",
"swap_walk",
"seed_mlfq_scheduler",
"seed_stride_scheduler",
"seed_eevdf_scheduler",
Expand Down
70 changes: 43 additions & 27 deletions src/fuzzer_tool/core/fair_queue.py
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,10 @@
DeficitRR cost O(n)/pick carries unspent credit, seeds
WeightedFairQueue cost O(n)/pick tightest fairness, targets
Stride counts O(log n) deterministic tickets, seeds/ops
EEVDF cost O(log n) lag-bounded, new flows join at V
EEVDF cost O(log n)* lag-bounded, new flows join at V

(*) amortized: each flow crosses from the ve heap to the deadline heap once
per service.
Comment on lines +16 to +19

Picks are O(n) because callers hand over the live flow set every call; the
flow sets here (targets, corpus) are rebuilt per pick by the caller anyway.
Expand Down Expand Up @@ -332,13 +335,26 @@ class EEVDF:
Unlike WFQ a flow ahead of the clock is never served early; unlike DRR
a new flow joins at ``V`` (lag 0), so it neither catches up nor waits.
Flat cost and weight are plain round robin.

Two heaps keep a pick amortized O(log n): deadline order is not
eligibility order, so one deadline heap would pop every ineligible flow
with an earlier deadline. ``pending`` holds flows by ``ve``; those that
fall at or below ``V`` move to ``ready``, ordered by deadline::

pending (ve) --ve <= V--> ready (deadline) --pop--> serve
^ |
+------------------ charge: ve += cost / w --------+

Each flow moves once per service. ``V`` only drops when a weight
changes or a flow leaves; a ready flow left ineligible is sent back.
"""

def __init__(self, slice_: float = NEUTRAL_COST) -> None:
self._slice = neutral(slice_)
self._ve: dict[Hashable, float] = {}
self._w: dict[Hashable, float] = {}
self._heap: list[tuple[float, int, Hashable]] = []
self._pending: list[tuple[float, int, Hashable]] = [] # (ve, seq, flow)
self._ready: list[tuple[float, int, Hashable]] = [] # (deadline, seq, flow)
self._seq = 0
self._sum_wve = 0.0
self._sum_w = 0.0
Expand Down Expand Up @@ -370,28 +386,28 @@ def _tick(self) -> int:
return self._seq

def _pop_eligible(self) -> Hashable:
"""Pop the earliest-deadline flow with ve <= V; ineligible ones go back.
"""Promote flows with ve <= V, then pop the earliest eligible deadline.

The lowest ``ve`` is always <= the weighted mean, so the scan ends;
if float drift ever says otherwise, the lowest ``ve`` is served.
The lowest ``ve`` is always <= the weighted mean, so ``ready`` is
non-empty; if float drift ever says otherwise, the lowest ``ve`` is
served.
"""
vt = self.virtual_time
slack = 1e-9 * max(1.0, abs(vt))
stash = []
entry = None
while self._heap:
head = heapq.heappop(self._heap)
if self._ve[head[2]] <= vt + slack:
entry = head
break
stash.append(head)

if entry is None:
entry = min(stash, key=lambda e: self._ve[e[2]])
stash.remove(entry)
for held in stash:
heapq.heappush(self._heap, held)
return entry[2]
limit = vt + 1e-9 * max(1.0, abs(vt))

# pending -> ready: flows whose eligible time the clock has reached.
while self._pending and self._pending[0][0] <= limit:
ve, seq, flow = heapq.heappop(self._pending)
heapq.heappush(self._ready, (ve + self._slice / self._w[flow], seq, flow))

while self._ready:
_, seq, flow = heapq.heappop(self._ready)
if self._ve[flow] <= limit:
return flow
# V dropped (weight change) since promotion: back to pending.
heapq.heappush(self._pending, (self._ve[flow], seq, flow))

return heapq.heappop(self._pending)[2]

def _charge(self, flow: Hashable, cost: float, w: float) -> None:
"""Advance ``ve`` by cost / w; a changed weight re-enters the sums."""
Expand All @@ -403,7 +419,7 @@ def _charge(self, flow: Hashable, cost: float, w: float) -> None:
ve += cost / w
self._ve[flow] = ve
self._w[flow] = w
heapq.heappush(self._heap, (ve + self._slice / w, self._tick(), flow))
heapq.heappush(self._pending, (ve, self._tick(), flow))

def _rebuild(self, flows: Sequence[Hashable], weight: Callable[[Hashable], float]) -> None:
"""Drop departed flows, then join new ones at the surviving clock."""
Expand All @@ -423,9 +439,9 @@ def _rebuild(self, flows: Sequence[Hashable], weight: Callable[[Hashable], float
self._sum_w += w
self._sum_wve += w * vt

seqs = {k: s for _, s, k in self._heap}
self._heap = [
(self._ve[k] + self._slice / self._w[k], seqs.get(k) or self._tick(), k) for k in flows
]
heapq.heapify(self._heap)
# Everything restarts in pending; the next pick promotes the eligible.
seqs = {k: s for _, s, k in self._pending + self._ready}
self._pending = [(self._ve[k], seqs.get(k) or self._tick(), k) for k in flows]
heapq.heapify(self._pending)
self._ready = []
self._seen = list(flows)
16 changes: 16 additions & 0 deletions src/fuzzer_tool/core/schedulers/_arm_counts.py
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,12 @@

from __future__ import annotations

from collections.abc import Sequence

#: The ledger is trimmed to the live corpus once it tracks this many times it.
PRUNE_FACTOR = 2
PRUNE_SLACK = 8


def reward(success: bool, weight: float) -> float:
"""Clamp a weighted success into [0, 1]; NaN and failures are 0."""
Expand Down Expand Up @@ -40,5 +46,15 @@ def mean(self, name: str) -> float:
s, f = self._counts.get(name, (0.0, 0.0))
return (s + 1.0) / (s + f + 2.0)

def _trim(self, live_ids: Sequence[str]) -> None:
"""Bound the ledger by the live candidate set: departed keys are dropped.

O(1) per call until the ledger outgrows the corpus, then one O(n) pass.
"""
if len(self._counts) <= PRUNE_FACTOR * len(live_ids) + PRUNE_SLACK:
return
live = set(live_ids)
self._counts = {k: v for k, v in self._counts.items() if k in live}

def bandit_stats(self) -> dict[str, tuple[float, float]]:
return {k: (s, f) for k, (s, f) in sorted(self._counts.items())}
4 changes: 3 additions & 1 deletion src/fuzzer_tool/core/schedulers/op_round_robin.py
Original file line number Diff line number Diff line change
Expand Up @@ -29,7 +29,9 @@ def select_op(self, ops: list[str]) -> str:
return ops[0]

# Filter to only registered operators in preferred order
available = [op for op in self._operator_order if op in ops]
# Set membership: `in ops` on the list was O(n^2) per pick.
live = set(ops)
available = [op for op in self._operator_order if op in live]
if not available:
# Fallback to registration order if none match available
available = [op for op in ops if op in self._operator_counts]
Expand Down
1 change: 1 addition & 0 deletions src/fuzzer_tool/core/schedulers/seed_aimd.py
Original file line number Diff line number Diff line change
Expand Up @@ -57,6 +57,7 @@ def __init__(
def select_seed(self, seed_ids: list[str]) -> str:
if len(self._window) > PRUNE_FACTOR * len(seed_ids) + PRUNE_SLACK:
self._prune(set(seed_ids))
self._trim(seed_ids)
return self._stride.pick(seed_ids, self._tickets)

def record(self, seed_id: str, success: bool, weight: float = 1.0) -> None:
Expand Down
1 change: 1 addition & 0 deletions src/fuzzer_tool/core/schedulers/seed_bfq.py
Original file line number Diff line number Diff line change
Expand Up @@ -65,6 +65,7 @@ def max_budget(self) -> int:
def select_seed(self, seed_ids: list[str], weight_fn: Callable[[str], float] = _unit) -> str:
if not seed_ids:
return ""
self._trim(seed_ids)
if len(seed_ids) == 1:
return seed_ids[0]
if self._seen is None or seed_ids != self._seen:
Expand Down
1 change: 1 addition & 0 deletions src/fuzzer_tool/core/schedulers/seed_codel.py
Original file line number Diff line number Diff line change
Expand Up @@ -57,6 +57,7 @@ def max_gap(self) -> int:
def select_seed(self, seed_ids: list[str]) -> str:
if not seed_ids:
return ""
self._trim(seed_ids)
if len(seed_ids) == 1:
return seed_ids[0]
if seed_ids != self._order:
Expand Down
1 change: 1 addition & 0 deletions src/fuzzer_tool/core/schedulers/seed_eevdf.py
Original file line number Diff line number Diff line change
Expand Up @@ -40,4 +40,5 @@ def select_seed(
cost_fn: Callable[[str], float] = _unit,
weight_fn: Callable[[str], float] = _unit,
) -> str:
self._trim(seed_ids)
return self._eevdf.pick(seed_ids, cost_fn, weight_fn)
1 change: 1 addition & 0 deletions src/fuzzer_tool/core/schedulers/seed_mlfq.py
Original file line number Diff line number Diff line change
Expand Up @@ -51,6 +51,7 @@ def __init__(self, levels: int = MLFQ_LEVELS, boost_period: int = BOOST_PERIOD)
def select_seed(self, seed_ids: list[str]) -> str:
if not seed_ids:
return ""
self._trim(seed_ids)
if len(seed_ids) == 1:
return seed_ids[0]
if self._seen is None or seed_ids != self._seen:
Expand Down
1 change: 1 addition & 0 deletions src/fuzzer_tool/core/schedulers/seed_p2c.py
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,7 @@ def __init__(self, rng) -> None:
def select_seed(self, seed_ids: list[str]) -> str:
if not seed_ids:
return ""
self._trim(seed_ids)
if len(seed_ids) == 1:
return seed_ids[0]

Expand Down
4 changes: 3 additions & 1 deletion src/fuzzer_tool/core/schedulers/seed_round_robin.py
Original file line number Diff line number Diff line change
Expand Up @@ -72,7 +72,9 @@ def select_seed(self, seed_ids: list[str]) -> str:

# Filter to only registered seeds in preferred (registration)
# order, same fallback chain as RoundRobinScheduler.select_op.
available = [s for s in self._seed_order if s in seed_ids]
# Set membership: `in seed_ids` on the list was O(n^2) per pick.
live = set(seed_ids)
available = [s for s in self._seed_order if s in live]
if not available:
available = [s for s in seed_ids if s in self._seed_counts]
if not available:
Expand Down
1 change: 1 addition & 0 deletions src/fuzzer_tool/core/schedulers/seed_sfq.py
Original file line number Diff line number Diff line change
Expand Up @@ -64,6 +64,7 @@ def select_seed(
) -> str:
if not seed_ids:
return ""
self._trim(seed_ids)
if len(seed_ids) == 1:
return seed_ids[0]

Expand Down
1 change: 1 addition & 0 deletions src/fuzzer_tool/core/schedulers/seed_stride.py
Original file line number Diff line number Diff line change
Expand Up @@ -34,4 +34,5 @@ def __init__(self) -> None:
self._stride = Stride()

def select_seed(self, seed_ids: list[str], weight_fn: Callable[[str], float] = _unit) -> str:
self._trim(seed_ids)
return self._stride.pick(seed_ids, weight_fn)
Loading
Loading