From af8677e792e8e37e256bc75e32476d8da6690a23 Mon Sep 17 00:00:00 2001 From: Claude Date: Thu, 1 Oct 2026 01:04:23 +0000 Subject: [PATCH] Fix PR #47 review findings - Doppler: abandon horizon doubles when a dropped seed returns; a fixed horizon thrashed once the corpus cycle outlasted it. - Seed-arm corpus memo: hits re-validated by slot identity, misses not memoized; length-keyed memo went stale on in-place trims. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01Rvj2gdMAa4HA8MGHULfDJE --- CHANGELOG.md | 3 +++ docs/DEEP_DIVE.md | 2 +- docs/TODO.md | 2 +- src/fuzzer_tool/core/power_doppler.py | 21 ++++++++++++++++++- src/fuzzer_tool/services/fuzzer.py | 29 ++++++++++++++++++--------- tests/test_os_net_scheduler_wiring.py | 28 ++++++++++++++++++++++++++ tests/test_power_doppler.py | 24 ++++++++++++++++++++++ 7 files changed, 96 insertions(+), 13 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index c3968339..d1c75c27 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -9,6 +9,9 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ### Fixed +- **Doppler horizon thrashed on slow corpus cycles** (`core/power_doppler.py`): a cycle longer than the fixed abandon horizon dropped every frame one tick before its seed returned. The horizon now doubles when a dropped seed comes back. +- **Stale corpus-membership memo** (`services/fuzzer.py`): keyed on corpus length, so an in-place trim of the parent kept a "member" verdict. Hits are now re-validated by slot identity; misses are not memoized. + - **`--mod-solving trace` clobbered `targets`** (`services/fuzzer.py`): the trace block reassigned the directed-targets parameter, disabling the Katz channel and directing at the fuzz target. Renamed the local. - **Doppler never scored corpora > 64 seeds picked in turn** (`core/power_doppler.py`): LRU evicted every partial frame. New seeds now wait for a slot; abandoned frames are scored early and freed. Flow-edge ids are int64 arrays under a global cap (frozensets could reach hundreds of MiB). - **Doppler mixed targets' edges** (`services/fuzzer.py`): multi-target frames are keyed per target. Without SHM, `--schedule doppler` now falls back to `base` with a warning instead of reporting enabled. diff --git a/docs/DEEP_DIVE.md b/docs/DEEP_DIVE.md index a275f894..9ec4337a 100644 --- a/docs/DEEP_DIVE.md +++ b/docs/DEEP_DIVE.md @@ -85,7 +85,7 @@ For production and sensitive binaries using AFL family fuzzers is the best cours - **AFLGo SHM-tail distance channel**: compiled into EVERY shim-linked target since `__AFL_DISTANCE_MODE` defaulted to 1 (2026-08-24; `-D__AFL_DISTANCE_MODE=0` opts out). Inert unless directed mode uploads a distance table (`DistanceTableShm`/`__AFL_DIST_SHM_ID`): without one, sum/count stay 0 and every reader takes the Python-side path. `build_targets.sh --distance` additionally builds the trace-pc-instrumented `*_dist.so`/`*_dist_asan.so` variants, where the shim accumulates per-block distances in `__sanitizer_cov_trace_pc()` — the PC (relative to the dladdr-derived object base) probes an open-addressing table of `{key, dist}` entries (packed 12-byte layout; the 4-byte header holds the slot *capacity*, a power of two ≥ 2×entries so empty slots exist, and the builder hash-inserts at `key % capacity` with linear probing to mirror the shim's probe — uploaded by the fuzzer at startup via `DistanceTableShm`/`__AFL_DIST_SHM_ID`), accumulating sum/count into the 16-byte SHM **tail** (after the edge table: `u64 dist_sum`, `u64 dist_count`), written at reset, at process exit (subprocess runs never call reset), and per-iteration in in-process modes via `__afl_dist_flush` (direct_lite has no process boundary, so the runner flushes the tail after each `run_one`). Per-execution `avg_distance = sum/count/100` is read straight from the tail and preferred over Python-side computation; blocks without a table entry don't count (AFLGo semantics). The table's PC keys are recovered by scanning text for `call __sanitizer_cov_trace_pc` sites (modern clang emits no `__sancov_pcs` for trace-pc) mapped to valued blocks via the CFGs — `TargetDistance.pc_distance_table()`. The shim's sanitizer-coverage callbacks are hidden-visibility so a libasan LD_PRELOAD cannot interpose over them in PIE builds. `tools/gen_distance_table.py` emits the table as C or text for inspection. Without the table (count==0) everything degrades to the Python-side path. Works in subprocess, direct_lite, and persistent modes (ASAN direct_lite requires libasan preloaded at fuzzer-process start — the `use_direct_lite` gate). The periodic stats line shows live distance when directed mode is active: `dist: avg: min: max:` (or `no-data`). `build_targets.sh --distance` builds both `*_dist.so` (no-ASAN) and `*_dist_asan.so` (ASAN) variants with the cmplog shim linked in, so `--cmplog` keeps them in direct_lite mode. Startup reports `[*] Distance instrumentation: detected` when the target carries the channel (the shim's `__afl_dist_flush` or a defined `__sanitizer_cov_trace_pc`), mirroring the AFL-instrumentation check. With `--elo` in directed mode, `aflgo` joins the Elo-arbitrated seed-strategy pool — a distance-pure arm picking `P(seed) ∝ exp(-2·norm_dist)` (distinct from the generic `weighted` arm, which blends distance with speed/size/entropy). - **AFLGo distance-annealed schedule** (`--schedule go`, requires `--target-functions`): wires the precomputed `avg_distance` (per-seed distance to directed targets) and `_anneal_progress` (exploration/exploitation annealing variable) into `SeedScorer.score()` for mutation budget scaling. During exploration phase (`anneal_progress` ≈ 0): uniform energy. During exploitation phase (`anneal_progress` → 1): `energy *= exp(β · (1 - norm_dist))` where `β = anneal_progress * 5`, capping at 100x. Seeds near the target get exponentially more mutations as the campaign matures. Previously these metrics only influenced seed selection but not mutation intensity. - **Power schedules** (`--schedule base|fast|coe|rare|mopt|lin|quad|go|aflgo|entropic|doppler`): AFL++ power schedules ported to control mutation budget per seed via `SeedScorer`. Each schedule modifies a base score (100) by frequency-based factors. Honggfuzz-style novelty decay, density, fertility, freshness, and entropy factors are applied multiplicatively on top. `entropic` (libFuzzer `-entropic`) scales energy by `1 + log2(1 + rare)`, where `rare` is the larger of `rare_edge_count`/`tc_ref` already collected for RARE/honggfuzz scoring — an approximation of libFuzzer's feature-frequency Shannon entropy using signal the fuzzer already tracks. -- **Power Doppler schedule** (`--schedule doppler`, `core/power_doppler.py`): ultrasound power Doppler on coverage. Slow time = successive mutants of one seed (32 per ensemble); pixel = edge; sample = `log2(1 + hits)` (raw SHM counts, not buckets). Wall filter: mean removal (static path), then SVD components whose participation ratio spans ≥ half the seed's edges (and ≥ 4) are dropped as clutter — an early reject moving the whole path at once ("flash"). CFAR: residual power per edge vs `σ² · χ²_{dof}(1 − 10⁻³)`, `σ²` = median residual variance (floor 10⁻³). Seed power = summed flow power / dof; energy = `log1p(p)/log1p(max p)` ∈ [0, 1], scaled to `[1, max_mult]` like `katz`. Unscored or static seeds stay 1×. Gram eigendecomposition (n×n) replaces the full SVD (~6× faster). Bounded: 64 open ensembles × 2048 edges (float32, ≤16 MiB); when full, new seeds wait instead of evicting partial frames (LRU eviction never closed a frame once >64 seeds were picked in turn), and a frame untouched for 4 frames' worth of samples is scored early (≥ 3 samples) and freed. Scores: 4096 (LRU), flow-edge ids in int64 arrays capped at 2²⁰ total (8 MiB). Multi-target frames are keyed per target (edge ids are per-target). SHM coverage only — without it the schedule falls back to `base` with a warning; cost ~100 µs/exec at 2k live edges (dict→array conversion dominates), zero when off. Unmeasured — see `docs/TODO.md`. +- **Power Doppler schedule** (`--schedule doppler`, `core/power_doppler.py`): ultrasound power Doppler on coverage. Slow time = successive mutants of one seed (32 per ensemble); pixel = edge; sample = `log2(1 + hits)` (raw SHM counts, not buckets). Wall filter: mean removal (static path), then SVD components whose participation ratio spans ≥ half the seed's edges (and ≥ 4) are dropped as clutter — an early reject moving the whole path at once ("flash"). CFAR: residual power per edge vs `σ² · χ²_{dof}(1 − 10⁻³)`, `σ²` = median residual variance (floor 10⁻³). Seed power = summed flow power / dof; energy = `log1p(p)/log1p(max p)` ∈ [0, 1], scaled to `[1, max_mult]` like `katz`. Unscored or static seeds stay 1×. Gram eigendecomposition (n×n) replaces the full SVD (~6× faster). Bounded: 64 open ensembles × 2048 edges (float32, ≤16 MiB); when full, new seeds wait instead of evicting partial frames (LRU eviction never closed a frame once >64 seeds were picked in turn), and a frame untouched for 4 frames' worth of samples is scored early (≥ 3 samples) and freed; that horizon doubles whenever a dropped seed returns (slow, not abandoned), so it converges past the corpus revisit time. Scores: 4096 (LRU), flow-edge ids in int64 arrays capped at 2²⁰ total (8 MiB). Multi-target frames are keyed per target (edge ids are per-target). SHM coverage only — without it the schedule falls back to `base` with a warning; cost ~100 µs/exec at 2k live edges (dict→array conversion dominates), zero when off. Unmeasured — see `docs/TODO.md`. - **Favored set / cull_queue** (`core/schedules.py`, `services/fuzzer.py`): AFL-style `top_rated` minimal-set-cover selection. For each edge, the cheapest seed covering it is selected, then a greedy cover builds the favored set. FAST and COE schedules apply energy bonuses to favored seeds (`_fast_factor`, `_coe_factor`, `coe_skip`). `_cull_queue()` runs periodically during fuzzing and updates `self._favored`; the score call site passes `favored=(seed_key in self._favored)` so the scheduler actually uses it. - **Bayesian seed quality** (`--bayesian`): `BayesianSeedQuality` (`core/seed_quality.py`) maintains a Beta-Bernoulli posterior per seed over `P(outcome = new_coverage)`. Thompson sampling naturally balances explore/exploit without a manual temperature knob — unexplored seeds have high posterior variance and get sampled. The `record_outcome()` feedback loop is now wired in `fuzz_one()` (was previously a dead code path with all posteriors stuck at Beta(1,1)). State is persisted to `seed_quality.json` and restored on resume. diff --git a/docs/TODO.md b/docs/TODO.md index 334c2206..7abb5720 100644 --- a/docs/TODO.md +++ b/docs/TODO.md @@ -25,7 +25,7 @@ - [ ] **ptrace breakpoints only on dominator-tree leaves** (2026-09-26) — a hit block implies its dominators ran, so `ptrace_coverage.py` could place int3 on leaves only and infer the rest. Edges `(prev, curr)` are not implied the same way; measure breakpoint count and edge-set loss on fuzzgoat first. See `docs/learnings/2026-09-26-dominators-chk-worst-case.md`. ## Scheduling -- [ ] **A/B `--schedule doppler`** (2026-10-01) — power Doppler seed energy (`core/power_doppler.py`) is unit-tested and wired; never measured. Paired `bench_paired.py` vs `fast` on clang-built fuzzgoat (Hard Rule 52). Open: (a) ~100 µs/exec at 2k edges, mostly `get_edge_counts()` dict → numpy; a numpy accessor on `ShmCoverage` (exposes `_active_columns`, Hard Rule 34 — needs approval) removes it; (b) ensembles span picks, so a seed needs 32 picks to score — tune `ensemble` vs pick rate; (c) unused: `flow_edges()` (input-sensitive edge set) could feed position/operator targeting; (d) scores not persisted across `--resume`; (e) `STALE_FRAMES = 4` (abandoned-frame window) and early-close scoring of partial frames are untuned — log `stats()['refused']` on fuzzgoat. +- [ ] **A/B `--schedule doppler`** (2026-10-01) — power Doppler seed energy (`core/power_doppler.py`) is unit-tested and wired; never measured. Paired `bench_paired.py` vs `fast` on clang-built fuzzgoat (Hard Rule 52). Open: (a) ~100 µs/exec at 2k edges, mostly `get_edge_counts()` dict → numpy; a numpy accessor on `ShmCoverage` (exposes `_active_columns`, Hard Rule 34 — needs approval) removes it; (b) ensembles span picks, so a seed needs 32 picks to score — tune `ensemble` vs pick rate; (c) unused: `flow_edges()` (input-sensitive edge set) could feed position/operator targeting; (d) scores not persisted across `--resume`; (e) `STALE_FRAMES = 4` (initial abandoned-frame window, self-widening) and early-close scoring of partial frames are untuned — log `stats()['refused']` on fuzzgoat. - [ ] **A/B the OS / network scheduler ports** (2026-09-30) — `mlfq`, `stride`, `eevdf`, `bfq`, `sfq`, `codel`, `aimd`, `p2c` seed arms and `op_stride`, `op_p2c` are unit- and wiring-tested only. Run `tools/lib/bench_paired.py` per arm vs `seed_round_robin` / `drr` on fuzzgoat (clang, ASAN). Untuned: MLFQ allotment/boost, BFQ budget range, CoDel target/interval, AIMD alpha/beta/loss run. Open: (a) `sfq` groups only direct siblings (`parent_key`), and only under `--lineage`; a lineage-root flow would group whole families; (b) no arm persists state. - [ ] **A/B the effector/token/chunk/changed/rare_mask position arms** (2026-09-30) — wired and unit-tested; a 3k-exec fuzzgoat run confirms live signal (`changed`: 673 moved / 96 unmoved / 2231 unmeasured rounds; `rare_mask`: target on 328/3000 rounds) but fuzzgoat cannot rank position arms (see `pos_fibonacci` entry). Run `tools/lib/bench_paired.py` `pos-arena-{token,chunk,changed,rare-mask}` vs `pos-arena-uniform` on png_read (`chunk` needs a container corpus). Open: (a) `effector` had no drained map in 3k execs (SkipDet skipped 111/112 seeds) -- measure its reach on long runs before A/B, it is not subset-testable (needs the det stage); (b) `changed` credits only ~25% of rounds: parents without a recorded path hash (spliced/generated inputs) -- record one at admission or accept; (c) neither `changed` nor `rare_mask` persists state. - [ ] **Pre-existing, found 2026-09-30:** `tools/build_targets.sh` ASAN variants fail in the cloud container (non-ASAN builds fine). diff --git a/src/fuzzer_tool/core/power_doppler.py b/src/fuzzer_tool/core/power_doppler.py index 4e5f4900..fd61cb2e 100644 --- a/src/fuzzer_tool/core/power_doppler.py +++ b/src/fuzzer_tool/core/power_doppler.py @@ -40,7 +40,8 @@ DEFAULT_MAX_SCORES = 4096 # Flow-edge ids kept across all scores: 8 B each bounds them at 8 MiB. DEFAULT_MAX_FLOW_IDS = 1 << 20 -# Open frame untouched for this many full frames of samples is abandoned. +# Open frame untouched for this many full frames of samples is abandoned +# (initial horizon; doubles whenever a dropped seed comes back). STALE_FRAMES = 4 # CFAR false-alarm probability per edge. FALSE_ALARM = 1e-3 @@ -235,6 +236,8 @@ def __init__( self._max_scores = max_scores self._max_flow_ids = max_flow_ids self._stale_after = max_seeds * ensemble * STALE_FRAMES + # Keys whose frames were dropped as abandoned; bounded like _open. + self._dropped_keys: LRUCache = LRUCache(max_seeds) # Insertion order is recency order; capacity is enforced by hand # (_admit, _trim) so evicted frames can be scored / ids uncounted. self._open: dict[str, _Ensemble] = {} @@ -257,6 +260,7 @@ def observe(self, seed_key: str, hits: Mapping[int, int]) -> None: self._ticks += 1 ens = self._open.pop(seed_key, None) if ens is None: + self._revisit(seed_key) ens = self._admit() if ens is None: self._refused += 1 @@ -279,6 +283,19 @@ def observe(self, seed_key: str, hits: Mapping[int, int]) -> None: del self._open[seed_key] self._close(seed_key, ens) + def _revisit(self, seed_key: str) -> None: + """A dropped seed came back: it was slow, not gone. Widen the horizon. + + A fixed horizon thrashes once the corpus cycle outlasts it, e.g. 13 + seeds in turn vs a 12-tick horizon: every frame is dropped one tick + before its seed returns. Doubling converges past the revisit time. + """ + if seed_key not in self._dropped_keys: + return + + del self._dropped_keys[seed_key] + self._stale_after *= 2 + def _admit(self) -> _Ensemble | None: """Fresh frame if a slot is free or the oldest frame is abandoned. @@ -294,6 +311,7 @@ def _admit(self) -> _Ensemble | None: # Abandoned: score what it has rather than throw it away. del self._open[key] + self._dropped_keys[key] = True if old.n >= _MIN_ENSEMBLE: self._close(key, old) return _Ensemble(self._ensemble, self._max_edges) @@ -356,5 +374,6 @@ def stats(self) -> dict[str, float]: "dropped_edges": self._dropped, "refused": self._refused, "flow_ids": self._flow_ids, + "stale_after": self._stale_after, "max_power": self._peak(), } diff --git a/src/fuzzer_tool/services/fuzzer.py b/src/fuzzer_tool/services/fuzzer.py index 211de7c4..c962bc4f 100644 --- a/src/fuzzer_tool/services/fuzzer.py +++ b/src/fuzzer_tool/services/fuzzer.py @@ -2745,8 +2745,8 @@ def __init__( ) if arm is not None ) - # (parent, corpus size, in corpus?): one membership test per pick. - self._parent_memo: tuple[bytes, int, bool] | None = None + # (parent, corpus list, slot, seed there): one membership test per pick. + self._parent_memo: tuple[bytes, list[bytes], int, bytes] | None = None if self._seed_os_arms: log.info("OS/network seed arms enabled: %d", len(self._seed_os_arms)) # LST override: no seed waits more than lst_revisit seconds between @@ -4548,18 +4548,27 @@ def _in_corpus(self, parent: bytes) -> bool: """Live-corpus membership of *parent*, memoized across its executions. Checked against the seed picker's cached key map (rebuilt only on - corpus change), not ``seed_meta``. The memo is keyed on corpus size - too, so a parent admitted mid-pick is seen. + corpus change), not ``seed_meta``. A hit is memoized with the slot it + sits in and re-validated by identity, O(1): an in-place replacement + (trim) or a rebuilt list misses. A miss is never memoized, so a + parent admitted mid-pick is seen. """ - n = len(self.corpus) + corpus = self.corpus memo = self._parent_memo - if memo is not None and memo[0] is parent and memo[1] == n: - return memo[2] + if memo is not None and memo[0] is parent and memo[1] is corpus: + slot = memo[2] + if slot < len(corpus) and corpus[slot] is memo[3]: + return True key_to_seed, _ = self._seed_picker._corpus_keys() - hit = self._seed_key(parent) in key_to_seed - self._parent_memo = (parent, n, hit) - return hit + seed = key_to_seed.get(self._seed_key(parent)) + if seed is None: + self._parent_memo = None + return False + + slot = corpus.index(seed) + self._parent_memo = (parent, corpus, slot, corpus[slot]) + return True def _seed_key(self, data: bytes) -> str: """Return content hash for *data*.""" diff --git a/tests/test_os_net_scheduler_wiring.py b/tests/test_os_net_scheduler_wiring.py index d8fca0ca..b9b79d06 100644 --- a/tests/test_os_net_scheduler_wiring.py +++ b/tests/test_os_net_scheduler_wiring.py @@ -282,6 +282,34 @@ def test_adversarial_parent_admitted_later_is_recorded(self, tmp_path): for arm in f._seed_os_arms: assert arm.bandit_stats() == {key: (1.0, 0.0)} + def test_regression_parent_replaced_in_place_not_recorded(self, tmp_path): + """Falsification (PR #47 review): same-length in-place replacement retires the parent.""" + kwargs = {kw: True for kw, _a, _c in SEED_ARMS.values()} + f = _real_fuzzer(tmp_path, **kwargs) + f.corpus.append(SEED_A) + f._record_seed_os_arms(SEED_A, success=True, weight=1.0) + + f.corpus[f.corpus.index(SEED_A)] = SEED_B + f._record_seed_os_arms(SEED_A, success=True, weight=1.0) + + key = f._seed_key(SEED_A) + for arm in f._seed_os_arms: + assert arm.bandit_stats()[key] == (1.0, 0.0) + + def test_adversarial_parent_kept_across_corpus_rebuild(self, tmp_path): + """A rebuilt corpus list that still holds the parent keeps recording it.""" + kwargs = {kw: True for kw, _a, _c in SEED_ARMS.values()} + f = _real_fuzzer(tmp_path, **kwargs) + f.corpus.append(SEED_A) + f._record_seed_os_arms(SEED_A, success=True, weight=1.0) + + f.corpus = [SEED_C, SEED_A] + f._record_seed_os_arms(SEED_A, success=True, weight=1.0) + + key = f._seed_key(SEED_A) + for arm in f._seed_os_arms: + assert arm.bandit_stats()[key] == (2.0, 0.0) + def test_op_arms_get_priors_registration_and_records(self): """The op arms sit in _register_arms and in fuzz_one's shared record loop.""" from fuzzer_tool.services import fuzzer as fz diff --git a/tests/test_power_doppler.py b/tests/test_power_doppler.py index 9218b8ba..6aea19a9 100644 --- a/tests/test_power_doppler.py +++ b/tests/test_power_doppler.py @@ -202,6 +202,30 @@ def test_regression_abandoned_frame_yields_its_slot(self): assert pd.flow_edges("gone") == frozenset({PATH}) assert pd.stats()["ensembles"] == 2 + def test_regression_slow_cycle_beyond_horizon_still_scores(self): + # Falsification (PR #47 review): 13 seeds in turn, 1 slot, 12-tick + # horizon. Each frame was dropped just before its seed came back. + ens, n_seeds = 3, 13 + pd = PowerDoppler(ensemble=ens, max_seeds=1) + ids = list(range(PATH)) + for _ in range(ens * 4): + for k in range(n_seeds): + _feed(pd, f"s{k}", _static(n=1), ids) + + assert pd.stats()["ensembles"] > 0 + + def test_adversarial_abandoned_keys_do_not_grow_horizon(self): + # Keys that never return keep the horizon; the dropped-key memory stays bounded. + cap = 2 + pd = PowerDoppler(ensemble=N, max_seeds=cap) + horizon = pd.stats()["stale_after"] + ids = list(range(PATH)) + for k in range(cap * 8): + _feed(pd, f"gone{k}", _static(n=horizon + 1), ids) + + assert pd.stats()["stale_after"] == horizon + assert len(pd._dropped_keys) <= cap + def test_adversarial_tiny_partial_frame_not_scored(self): # Fewer samples than a valid ensemble: dropped, never scored. pd = PowerDoppler(ensemble=N, max_seeds=1)