Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,11 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

### Fixed

- **`--mod-solving trace` clobbered `targets`** (`services/fuzzer.py`): the trace block reassigned the directed-targets parameter, disabling the Katz channel and directing at the fuzz target. Renamed the local.
- **Doppler never scored corpora > 64 seeds picked in turn** (`core/power_doppler.py`): LRU evicted every partial frame. New seeds now wait for a slot; abandoned frames are scored early and freed. Flow-edge ids are int64 arrays under a global cap (frozensets could reach hundreds of MiB).
- **Doppler mixed targets' edges** (`services/fuzzer.py`): multi-target frames are keyed per target. Without SHM, `--schedule doppler` now falls back to `base` with a warning instead of reporting enabled.
- **Seed-arm ledgers recorded standalone-QEA parents**: `seed_meta` is not corpus membership. Gated on the seed picker's cached corpus key map, memoized per parent.

- **EEVDF pick scanned ineligible flows** (`core/fair_queue.py`): one deadline heap popped every flow with an earlier deadline but `ve > V` (5000 pops at 5000 flows). Now a `ve` heap feeds a deadline heap; amortized O(log n). Test: `test_eevdf_pick_does_not_scan_ineligible_flows`.
- **Seed-arm ledgers grew without bound**: non-corpus (Markov) parents were recorded, and departed seeds never left `ArmCounts`. Only corpus parents are recorded now; ledgers trim to 2x the live corpus.
- **Round robin was O(n^2) per pick** (`seed_round_robin`, `op_round_robin`): `x in list` per registered arm. Set membership now: 103 ms -> 0.7 ms per pick at 5000 seeds.
Expand Down
2 changes: 1 addition & 1 deletion docs/DEEP_DIVE.md
Original file line number Diff line number Diff line change
Expand Up @@ -85,7 +85,7 @@ For production and sensitive binaries using AFL family fuzzers is the best cours
- **AFLGo SHM-tail distance channel**: compiled into EVERY shim-linked target since `__AFL_DISTANCE_MODE` defaulted to 1 (2026-08-24; `-D__AFL_DISTANCE_MODE=0` opts out). Inert unless directed mode uploads a distance table (`DistanceTableShm`/`__AFL_DIST_SHM_ID`): without one, sum/count stay 0 and every reader takes the Python-side path. `build_targets.sh --distance` additionally builds the trace-pc-instrumented `*_dist.so`/`*_dist_asan.so` variants, where the shim accumulates per-block distances in `__sanitizer_cov_trace_pc()` — the PC (relative to the dladdr-derived object base) probes an open-addressing table of `{key, dist}` entries (packed 12-byte layout; the 4-byte header holds the slot *capacity*, a power of two ≥ 2×entries so empty slots exist, and the builder hash-inserts at `key % capacity` with linear probing to mirror the shim's probe — uploaded by the fuzzer at startup via `DistanceTableShm`/`__AFL_DIST_SHM_ID`), accumulating sum/count into the 16-byte SHM **tail** (after the edge table: `u64 dist_sum`, `u64 dist_count`), written at reset, at process exit (subprocess runs never call reset), and per-iteration in in-process modes via `__afl_dist_flush` (direct_lite has no process boundary, so the runner flushes the tail after each `run_one`). Per-execution `avg_distance = sum/count/100` is read straight from the tail and preferred over Python-side computation; blocks without a table entry don't count (AFLGo semantics). The table's PC keys are recovered by scanning text for `call __sanitizer_cov_trace_pc` sites (modern clang emits no `__sancov_pcs` for trace-pc) mapped to valued blocks via the CFGs — `TargetDistance.pc_distance_table()`. The shim's sanitizer-coverage callbacks are hidden-visibility so a libasan LD_PRELOAD cannot interpose over them in PIE builds. `tools/gen_distance_table.py` emits the table as C or text for inspection. Without the table (count==0) everything degrades to the Python-side path. Works in subprocess, direct_lite, and persistent modes (ASAN direct_lite requires libasan preloaded at fuzzer-process start — the `use_direct_lite` gate). The periodic stats line shows live distance when directed mode is active: `dist: avg:<tail avg> min:<observed min> max:<observed max>` (or `no-data`). `build_targets.sh --distance` builds both `*_dist.so` (no-ASAN) and `*_dist_asan.so` (ASAN) variants with the cmplog shim linked in, so `--cmplog` keeps them in direct_lite mode. Startup reports `[*] Distance instrumentation: detected` when the target carries the channel (the shim's `__afl_dist_flush` or a defined `__sanitizer_cov_trace_pc`), mirroring the AFL-instrumentation check. With `--elo` in directed mode, `aflgo` joins the Elo-arbitrated seed-strategy pool — a distance-pure arm picking `P(seed) ∝ exp(-2·norm_dist)` (distinct from the generic `weighted` arm, which blends distance with speed/size/entropy).
- **AFLGo distance-annealed schedule** (`--schedule go`, requires `--target-functions`): wires the precomputed `avg_distance` (per-seed distance to directed targets) and `_anneal_progress` (exploration/exploitation annealing variable) into `SeedScorer.score()` for mutation budget scaling. During exploration phase (`anneal_progress` ≈ 0): uniform energy. During exploitation phase (`anneal_progress` → 1): `energy *= exp(β · (1 - norm_dist))` where `β = anneal_progress * 5`, capping at 100x. Seeds near the target get exponentially more mutations as the campaign matures. Previously these metrics only influenced seed selection but not mutation intensity.
- **Power schedules** (`--schedule base|fast|coe|rare|mopt|lin|quad|go|aflgo|entropic|doppler`): AFL++ power schedules ported to control mutation budget per seed via `SeedScorer`. Each schedule modifies a base score (100) by frequency-based factors. Honggfuzz-style novelty decay, density, fertility, freshness, and entropy factors are applied multiplicatively on top. `entropic` (libFuzzer `-entropic`) scales energy by `1 + log2(1 + rare)`, where `rare` is the larger of `rare_edge_count`/`tc_ref` already collected for RARE/honggfuzz scoring — an approximation of libFuzzer's feature-frequency Shannon entropy using signal the fuzzer already tracks.
- **Power Doppler schedule** (`--schedule doppler`, `core/power_doppler.py`): ultrasound power Doppler on coverage. Slow time = successive mutants of one seed (32 per ensemble); pixel = edge; sample = `log2(1 + hits)` (raw SHM counts, not buckets). Wall filter: mean removal (static path), then SVD components whose participation ratio spans ≥ half the seed's edges (and ≥ 4) are dropped as clutter — an early reject moving the whole path at once ("flash"). CFAR: residual power per edge vs `σ² · χ²_{dof}(1 − 10⁻³)`, `σ²` = median residual variance (floor 10⁻³). Seed power = summed flow power / dof; energy = `log1p(p)/log1p(max p)` ∈ [0, 1], scaled to `[1, max_mult]` like `katz`. Unscored or static seeds stay 1×. Gram eigendecomposition (n×n) replaces the full SVD (~6× faster). Bounded: 64 open ensembles × 2048 edges (float32, ≤16 MiB), 4096 scores (LRU). SHM coverage only; cost ~100 µs/exec at 2k live edges (dict→array conversion dominates), zero when off. Unmeasured — see `docs/TODO.md`.
- **Power Doppler schedule** (`--schedule doppler`, `core/power_doppler.py`): ultrasound power Doppler on coverage. Slow time = successive mutants of one seed (32 per ensemble); pixel = edge; sample = `log2(1 + hits)` (raw SHM counts, not buckets). Wall filter: mean removal (static path), then SVD components whose participation ratio spans ≥ half the seed's edges (and ≥ 4) are dropped as clutter — an early reject moving the whole path at once ("flash"). CFAR: residual power per edge vs `σ² · χ²_{dof}(1 − 10⁻³)`, `σ²` = median residual variance (floor 10⁻³). Seed power = summed flow power / dof; energy = `log1p(p)/log1p(max p)` ∈ [0, 1], scaled to `[1, max_mult]` like `katz`. Unscored or static seeds stay 1×. Gram eigendecomposition (n×n) replaces the full SVD (~6× faster). Bounded: 64 open ensembles × 2048 edges (float32, ≤16 MiB); when full, new seeds wait instead of evicting partial frames (LRU eviction never closed a frame once >64 seeds were picked in turn), and a frame untouched for 4 frames' worth of samples is scored early (≥ 3 samples) and freed. Scores: 4096 (LRU), flow-edge ids in int64 arrays capped at 2²⁰ total (8 MiB). Multi-target frames are keyed per target (edge ids are per-target). SHM coverage only — without it the schedule falls back to `base` with a warning; cost ~100 µs/exec at 2k live edges (dict→array conversion dominates), zero when off. Unmeasured — see `docs/TODO.md`.
- **Favored set / cull_queue** (`core/schedules.py`, `services/fuzzer.py`): AFL-style `top_rated` minimal-set-cover selection. For each edge, the cheapest seed covering it is selected, then a greedy cover builds the favored set. FAST and COE schedules apply energy bonuses to favored seeds (`_fast_factor`, `_coe_factor`, `coe_skip`). `_cull_queue()` runs periodically during fuzzing and updates `self._favored`; the score call site passes `favored=(seed_key in self._favored)` so the scheduler actually uses it.
- **Bayesian seed quality** (`--bayesian`): `BayesianSeedQuality` (`core/seed_quality.py`) maintains a Beta-Bernoulli posterior per seed over `P(outcome = new_coverage)`. Thompson sampling naturally balances explore/exploit without a manual temperature knob — unexplored seeds have high posterior variance and get sampled. The `record_outcome()` feedback loop is now wired in `fuzz_one()` (was previously a dead code path with all posteriors stuck at Beta(1,1)). State is persisted to `seed_quality.json` and restored on resume.

Expand Down
4 changes: 2 additions & 2 deletions docs/TODO.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@
- [ ] **ptrace breakpoints only on dominator-tree leaves** (2026-09-26) — a hit block implies its dominators ran, so `ptrace_coverage.py` could place int3 on leaves only and infer the rest. Edges `(prev, curr)` are not implied the same way; measure breakpoint count and edge-set loss on fuzzgoat first. See `docs/learnings/2026-09-26-dominators-chk-worst-case.md`.

## Scheduling
- [ ] **A/B `--schedule doppler`** (2026-10-01) — power Doppler seed energy (`core/power_doppler.py`) is unit-tested and wired; never measured. Paired `bench_paired.py` vs `fast` on clang-built fuzzgoat (Hard Rule 52). Open: (a) ~100 µs/exec at 2k edges, mostly `get_edge_counts()` dict → numpy; a numpy accessor on `ShmCoverage` (exposes `_active_columns`, Hard Rule 34 — needs approval) removes it; (b) ensembles span picks, so a seed needs 32 picks to score — tune `ensemble` vs pick rate; (c) unused: `flow_edges()` (input-sensitive edge set) could feed position/operator targeting; (d) scores not persisted across `--resume`.
- [ ] **A/B `--schedule doppler`** (2026-10-01) — power Doppler seed energy (`core/power_doppler.py`) is unit-tested and wired; never measured. Paired `bench_paired.py` vs `fast` on clang-built fuzzgoat (Hard Rule 52). Open: (a) ~100 µs/exec at 2k edges, mostly `get_edge_counts()` dict → numpy; a numpy accessor on `ShmCoverage` (exposes `_active_columns`, Hard Rule 34 — needs approval) removes it; (b) ensembles span picks, so a seed needs 32 picks to score — tune `ensemble` vs pick rate; (c) unused: `flow_edges()` (input-sensitive edge set) could feed position/operator targeting; (d) scores not persisted across `--resume`; (e) `STALE_FRAMES = 4` (abandoned-frame window) and early-close scoring of partial frames are untuned — log `stats()['refused']` on fuzzgoat.
- [ ] **A/B the OS / network scheduler ports** (2026-09-30) — `mlfq`, `stride`, `eevdf`, `bfq`, `sfq`, `codel`, `aimd`, `p2c` seed arms and `op_stride`, `op_p2c` are unit- and wiring-tested only. Run `tools/lib/bench_paired.py` per arm vs `seed_round_robin` / `drr` on fuzzgoat (clang, ASAN). Untuned: MLFQ allotment/boost, BFQ budget range, CoDel target/interval, AIMD alpha/beta/loss run. Open: (a) `sfq` groups only direct siblings (`parent_key`), and only under `--lineage`; a lineage-root flow would group whole families; (b) no arm persists state.
- [ ] **A/B the effector/token/chunk/changed/rare_mask position arms** (2026-09-30) — wired and unit-tested; a 3k-exec fuzzgoat run confirms live signal (`changed`: 673 moved / 96 unmoved / 2231 unmeasured rounds; `rare_mask`: target on 328/3000 rounds) but fuzzgoat cannot rank position arms (see `pos_fibonacci` entry). Run `tools/lib/bench_paired.py` `pos-arena-{token,chunk,changed,rare-mask}` vs `pos-arena-uniform` on png_read (`chunk` needs a container corpus). Open: (a) `effector` had no drained map in 3k execs (SkipDet skipped 111/112 seeds) -- measure its reach on long runs before A/B, it is not subset-testable (needs the det stage); (b) `changed` credits only ~25% of rounds: parents without a recorded path hash (spliced/generated inputs) -- record one at admission or accept; (c) neither `changed` nor `rare_mask` persists state.
- [ ] **Pre-existing, found 2026-09-30:** `tools/build_targets.sh` ASAN variants fail in the cloud container (non-ASAN builds fine).
Expand Down Expand Up @@ -91,7 +91,7 @@
- [ ] **Point in-script cross-references at `tools/benchmark.py`** (2026-09-26) — harness docstrings and `bench.sh`/`bench_sweep.sh` headers still name each other by path (`tools/lib/bench_paired.py`, `tools/lib/bench_diff.py`). Correct, but a reader never learns the entry point from them.
- [ ] **Three operators are still unexamined by the enumeration harness** (2026-09-12) — `avif_chunk_mutate`, `golomb` and `pgs_chunk_mutate` report `too_deep` once `_walk_operator` samples them in spread order, where the lexicographic walk called them `over_budget` and so hid the fact that they exceed `max_depth=16`. Raising the harness depth admits them; the depth cap exists so a truncated path is reported rather than silently walked, so raise it deliberately and re-measure the census rather than removing it.
- [ ] **ffmpeg campaign throughput is admission-bound** (2026-09-26) — Docker `ffmpeg` image, full build, 4 cores: raw target 841 eps, campaign 3 eps with `-c --elo all --lineage-backtrack`, 6 eps with `-c` alone; 55 s to first exec. Not Docker (`--shm-size=2g` unchanged). Profile with `--profile-hotpath` on the same image.
- [ ] **`--mod-solving trace` clobbers `targets`** (2026-09-26) — `Fuzzer.__init__` reassigns the directed-mode `targets` parameter to `multi_targets or [target]` inside the SMT trace block, which then disables the Katz channel and feeds `_distance_targets` the fuzz target itself. Rename the local.
- [x] **`--mod-solving trace` clobbers `targets`** (2026-09-26, fixed 2026-10-01: local renamed `div_targets`; `test_regression_trace_mode_keeps_katz`) — `Fuzzer.__init__` reassigns the directed-mode `targets` parameter to `multi_targets or [target]` inside the SMT trace block, which then disables the Katz channel and feeds `_distance_targets` the fuzz target itself. Rename the local.
- [ ] **Pre-existing red tests** (2026-09-25) — `test_katz_channel::TestScores::test_ensure_scores_caches_until_dirty_and_due`, `test_regression_build_lib_pairing::test_targets_are_actually_built` (red on base 2026-09-26); `test_regression_track_op_effect_coverage::test_every_ballot_name_is_mapped` (no kwargs for `fewa`, `softmax`, `topk`); `test_kruskal_count` / `test_seed_round_robin` `TestFuzzerWiring::test_constructor_flag_is_*last*` (constructor order); `test_regression_no_op_mutations::test_every_selectable_operator_is_reachable`. Red on master before the LRU/bayes-ucb change.

## Scheduling
Expand Down
8 changes: 5 additions & 3 deletions src/fuzzer_tool/core/fair_queue.py
Original file line number Diff line number Diff line change
Expand Up @@ -15,8 +15,9 @@
Stride counts O(log n) deterministic tickets, seeds/ops
EEVDF cost O(log n)* lag-bounded, new flows join at V

(*) amortized: each flow crosses from the ve heap to the deadline heap once
per service.
(*) heap work only, amortized: each flow crosses from the ve heap to the
deadline heap once per service. The whole pick is O(n + log n): it also
compares the caller's flow list against the last one.

Picks are O(n) because callers hand over the live flow set every call; the
flow sets here (targets, corpus) are rebuilt per pick by the caller anyway.
Expand Down Expand Up @@ -336,7 +337,8 @@ class EEVDF:
a new flow joins at ``V`` (lag 0), so it neither catches up nor waits.
Flat cost and weight are plain round robin.

Two heaps keep a pick amortized O(log n): deadline order is not
Two heaps keep a pick's heap work amortized O(log n) (the pick itself
stays O(n + log n), see module note): deadline order is not
eligibility order, so one deadline heap would pop every ineligible flow
with an earlier deadline. ``pending`` holds flows by ``ve``; those that
fall at or below ``V`` move to ``ready``, ordered by deadline::
Expand Down
Loading