From 2807e04c8133623cac8b606438dace6568bf1127 Mon Sep 17 00:00:00 2001 From: Claude Date: Thu, 1 Oct 2026 00:06:46 +0000 Subject: [PATCH] Add power Doppler seed energy schedule (--schedule doppler) Slow time = mutants of one seed; pixel = edge; sample = log2(1+hits). Mean removal + SVD wall filter (participation ratio drops path-wide "flash" components), CFAR chi^2 detection, flow power -> [0,1] energy, scaled to [1, max_mult] like katz. Bounded LRU ensembles; Gram eigendecomposition instead of full SVD (~6x faster close). Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_019AERo9TJciJNQMw2uc18qM --- README.md | 2 +- docs/DEEP_DIVE.md | 5 +- docs/TODO.md | 1 + docs/architecture.dot | 2 +- docs/images/architecture.png | Bin 610422 -> 613831 bytes docs/images/architecture.svg | 134 ++++++------ src/fuzzer_tool/cli/commands.py | 19 +- src/fuzzer_tool/core/power_doppler.py | 292 ++++++++++++++++++++++++++ src/fuzzer_tool/core/schedules.py | 10 + src/fuzzer_tool/services/fuzzer.py | 21 ++ tests/test_power_doppler.py | 263 +++++++++++++++++++++++ 11 files changed, 675 insertions(+), 74 deletions(-) create mode 100644 src/fuzzer_tool/core/power_doppler.py create mode 100644 tests/test_power_doppler.py diff --git a/README.md b/README.md index dab31a61..16df63f8 100644 --- a/README.md +++ b/README.md @@ -143,7 +143,7 @@ per-sub-operator reward instead of uniformly (`--no-adaptive-havoc` restores uni | Metropolis admission | `--metropolis` | Accept non-improving inputs with P = exp(−ΔE/T) | | Secretary stopping | `--secretary` | Secretary-problem stopping ranks (display only) | | honggfuzz power factors | `--honggfuzz` | Novelty decay, freshness, fertility, density, entropy and timeout penalties | -| AFL++ power schedules | `--schedule` | FAST/COE/RARE/MMOPT/LIN/QUAD/GO/AFLGO/ENTROPIC seed-level energy | +| AFL++ power schedules | `--schedule` | FAST/COE/RARE/MMOPT/LIN/QUAD/GO/AFLGO/ENTROPIC/DOPPLER seed-level energy | | AFLGo directed annealing | `--schedule aflgo` | Exact AFLGo power factor — symmetric 32×/1/32× energy by distance-to-target with time-based cooling (`--t-x`, `--aflgo-cooling`) | | Entropic power schedule | `--schedule entropic` | libFuzzer `-entropic`: energy ∝ log(1 + rare-feature count) from already-tracked rare-edge ownership | | **K-Scheduler Katz centrality** | auto | On trace-pc targets: whole-program ICFG → horizon graph (contracted visited deletion, DAG) → out-degree Katz with β from node-hit counts; Elo-rated `katz` seed arm plus a clamped `--schedule katz` energy | diff --git a/docs/DEEP_DIVE.md b/docs/DEEP_DIVE.md index cc095b2b..fff705c8 100644 --- a/docs/DEEP_DIVE.md +++ b/docs/DEEP_DIVE.md @@ -79,7 +79,8 @@ For production and sensitive binaries using AFL family fuzzers is the best cours - **AFLGo power schedule** (`--schedule aflgo`, `--aflgo-cooling exp|log|lin|quad`, `--t-x MINUTES`): the exact `calculate_score()` distance section from AFLGo's afl-fuzz.c. Temperature `T` follows the chosen cooling over t_x minutes to exploitation (`exp`: 1/20^progress; `log`: 1/(1+2·ln(1+progress·13358.7268297)); `lin`: 1/(1+19·progress); `quad`: 1/(1+19·progress²)); `p = (1−nd)(1−T) + 0.5T` with `nd = (d−min)/(max−min)` normalized over the observed queue (min/max tracked per run, excluding the no-data sentinel); `factor = 2^(2·log2(32)·(p−0.5))` — symmetric around 1.0: early campaigns treat every seed equally, late campaigns give near-target seeds up to 32× energy and far seeds as little as 1/32×. - **AFLGo SHM-tail distance channel**: compiled into EVERY shim-linked target since `__AFL_DISTANCE_MODE` defaulted to 1 (2026-08-24; `-D__AFL_DISTANCE_MODE=0` opts out). Inert unless directed mode uploads a distance table (`DistanceTableShm`/`__AFL_DIST_SHM_ID`): without one, sum/count stay 0 and every reader takes the Python-side path. `build_targets.sh --distance` additionally builds the trace-pc-instrumented `*_dist.so`/`*_dist_asan.so` variants, where the shim accumulates per-block distances in `__sanitizer_cov_trace_pc()` — the PC (relative to the dladdr-derived object base) probes an open-addressing table of `{key, dist}` entries (packed 12-byte layout; the 4-byte header holds the slot *capacity*, a power of two ≥ 2×entries so empty slots exist, and the builder hash-inserts at `key % capacity` with linear probing to mirror the shim's probe — uploaded by the fuzzer at startup via `DistanceTableShm`/`__AFL_DIST_SHM_ID`), accumulating sum/count into the 16-byte SHM **tail** (after the edge table: `u64 dist_sum`, `u64 dist_count`), written at reset, at process exit (subprocess runs never call reset), and per-iteration in in-process modes via `__afl_dist_flush` (direct_lite has no process boundary, so the runner flushes the tail after each `run_one`). Per-execution `avg_distance = sum/count/100` is read straight from the tail and preferred over Python-side computation; blocks without a table entry don't count (AFLGo semantics). The table's PC keys are recovered by scanning text for `call __sanitizer_cov_trace_pc` sites (modern clang emits no `__sancov_pcs` for trace-pc) mapped to valued blocks via the CFGs — `TargetDistance.pc_distance_table()`. The shim's sanitizer-coverage callbacks are hidden-visibility so a libasan LD_PRELOAD cannot interpose over them in PIE builds. `tools/gen_distance_table.py` emits the table as C or text for inspection. Without the table (count==0) everything degrades to the Python-side path. Works in subprocess, direct_lite, and persistent modes (ASAN direct_lite requires libasan preloaded at fuzzer-process start — the `use_direct_lite` gate). The periodic stats line shows live distance when directed mode is active: `dist: avg: min: max:` (or `no-data`). `build_targets.sh --distance` builds both `*_dist.so` (no-ASAN) and `*_dist_asan.so` (ASAN) variants with the cmplog shim linked in, so `--cmplog` keeps them in direct_lite mode. Startup reports `[*] Distance instrumentation: detected` when the target carries the channel (the shim's `__afl_dist_flush` or a defined `__sanitizer_cov_trace_pc`), mirroring the AFL-instrumentation check. With `--elo` in directed mode, `aflgo` joins the Elo-arbitrated seed-strategy pool — a distance-pure arm picking `P(seed) ∝ exp(-2·norm_dist)` (distinct from the generic `weighted` arm, which blends distance with speed/size/entropy). - **AFLGo distance-annealed schedule** (`--schedule go`, requires `--target-functions`): wires the precomputed `avg_distance` (per-seed distance to directed targets) and `_anneal_progress` (exploration/exploitation annealing variable) into `SeedScorer.score()` for mutation budget scaling. During exploration phase (`anneal_progress` ≈ 0): uniform energy. During exploitation phase (`anneal_progress` → 1): `energy *= exp(β · (1 - norm_dist))` where `β = anneal_progress * 5`, capping at 100x. Seeds near the target get exponentially more mutations as the campaign matures. Previously these metrics only influenced seed selection but not mutation intensity. -- **Power schedules** (`--schedule base|fast|coe|rare|mopt|lin|quad|go|aflgo|entropic`): AFL++ power schedules ported to control mutation budget per seed via `SeedScorer`. Each schedule modifies a base score (100) by frequency-based factors. Honggfuzz-style novelty decay, density, fertility, freshness, and entropy factors are applied multiplicatively on top. `entropic` (libFuzzer `-entropic`) scales energy by `1 + log2(1 + rare)`, where `rare` is the larger of `rare_edge_count`/`tc_ref` already collected for RARE/honggfuzz scoring — an approximation of libFuzzer's feature-frequency Shannon entropy using signal the fuzzer already tracks. +- **Power schedules** (`--schedule base|fast|coe|rare|mopt|lin|quad|go|aflgo|entropic|doppler`): AFL++ power schedules ported to control mutation budget per seed via `SeedScorer`. Each schedule modifies a base score (100) by frequency-based factors. Honggfuzz-style novelty decay, density, fertility, freshness, and entropy factors are applied multiplicatively on top. `entropic` (libFuzzer `-entropic`) scales energy by `1 + log2(1 + rare)`, where `rare` is the larger of `rare_edge_count`/`tc_ref` already collected for RARE/honggfuzz scoring — an approximation of libFuzzer's feature-frequency Shannon entropy using signal the fuzzer already tracks. +- **Power Doppler schedule** (`--schedule doppler`, `core/power_doppler.py`): ultrasound power Doppler on coverage. Slow time = successive mutants of one seed (32 per ensemble); pixel = edge; sample = `log2(1 + hits)` (raw SHM counts, not buckets). Wall filter: mean removal (static path), then SVD components whose participation ratio spans ≥ half the seed's edges (and ≥ 4) are dropped as clutter — an early reject moving the whole path at once ("flash"). CFAR: residual power per edge vs `σ² · χ²_{dof}(1 − 10⁻³)`, `σ²` = median residual variance (floor 10⁻³). Seed power = summed flow power / dof; energy = `log1p(p)/log1p(max p)` ∈ [0, 1], scaled to `[1, max_mult]` like `katz`. Unscored or static seeds stay 1×. Gram eigendecomposition (n×n) replaces the full SVD (~6× faster). Bounded: 64 open ensembles × 2048 edges (float32, ≤16 MiB), 4096 scores (LRU). SHM coverage only; cost ~100 µs/exec at 2k live edges (dict→array conversion dominates), zero when off. Unmeasured — see `docs/TODO.md`. - **Favored set / cull_queue** (`core/schedules.py`, `services/fuzzer.py`): AFL-style `top_rated` minimal-set-cover selection. For each edge, the cheapest seed covering it is selected, then a greedy cover builds the favored set. FAST and COE schedules apply energy bonuses to favored seeds (`_fast_factor`, `_coe_factor`, `coe_skip`). `_cull_queue()` runs periodically during fuzzing and updates `self._favored`; the score call site passes `favored=(seed_key in self._favored)` so the scheduler actually uses it. - **Bayesian seed quality** (`--bayesian`): `BayesianSeedQuality` (`core/seed_quality.py`) maintains a Beta-Bernoulli posterior per seed over `P(outcome = new_coverage)`. Thompson sampling naturally balances explore/exploit without a manual temperature knob — unexplored seeds have high posterior variance and get sampled. The `record_outcome()` feedback loop is now wired in `fuzz_one()` (was previously a dead code path with all posteriors stuck at Beta(1,1)). State is persisted to `seed_quality.json` and restored on resume. @@ -567,7 +568,7 @@ fuzzer-tool rank ./target -d corpus -n 10 --dump top_seeds | `--wfc` | Wave Function Collapse structural generation (chunk reordering, pixel generation) | | `--enable-smt-z3` | Z3-based SMT solving for arithmetic constraint solving on cmplog pairs | | `--hw-perf` | Hardware performance counters via perf_event_open (instructions, branches, misses) | -| `--schedule base\|fast\|coe\|rare\|mopt\|lin\|quad\|go\|aflgo\|entropic` | AFL++ power schedule (`aflgo` = exact AFLGo distance annealing, `entropic` = libFuzzer `-entropic`, log-scaled rare-feature energy) | +| `--schedule base\|fast\|coe\|rare\|mopt\|lin\|quad\|go\|aflgo\|entropic\|doppler` | AFL++ power schedule (`aflgo` = exact AFLGo distance annealing, `entropic` = libFuzzer `-entropic`, log-scaled rare-feature energy, `doppler` = power Doppler flow energy) | | `--aflgo-cooling exp\|log\|lin\|quad` | Cooling schedule for the `aflgo` power factor (default exp) | | `--t-x MINUTES` | AFLGo time-to-exploitation in minutes; temperature cools to 1/20 over this window | | `--markov-order N` | Markov chain order(s), comma-separated (e.g. '0,1,2' for ensemble) | diff --git a/docs/TODO.md b/docs/TODO.md index 95b0749d..6c42edac 100644 --- a/docs/TODO.md +++ b/docs/TODO.md @@ -25,6 +25,7 @@ - [ ] **ptrace breakpoints only on dominator-tree leaves** (2026-09-26) — a hit block implies its dominators ran, so `ptrace_coverage.py` could place int3 on leaves only and infer the rest. Edges `(prev, curr)` are not implied the same way; measure breakpoint count and edge-set loss on fuzzgoat first. See `docs/learnings/2026-09-26-dominators-chk-worst-case.md`. ## Scheduling +- [ ] **A/B `--schedule doppler`** (2026-10-01) — power Doppler seed energy (`core/power_doppler.py`) is unit-tested and wired; never measured. Paired `bench_paired.py` vs `fast` on clang-built fuzzgoat (Hard Rule 52). Open: (a) ~100 µs/exec at 2k edges, mostly `get_edge_counts()` dict → numpy; a numpy accessor on `ShmCoverage` (exposes `_active_columns`, Hard Rule 34 — needs approval) removes it; (b) ensembles span picks, so a seed needs 32 picks to score — tune `ensemble` vs pick rate; (c) unused: `flow_edges()` (input-sensitive edge set) could feed position/operator targeting; (d) scores not persisted across `--resume`. - [ ] **A/B the effector/token/chunk/changed/rare_mask position arms** (2026-09-30) — wired and unit-tested; a 3k-exec fuzzgoat run confirms live signal (`changed`: 673 moved / 96 unmoved / 2231 unmeasured rounds; `rare_mask`: target on 328/3000 rounds) but fuzzgoat cannot rank position arms (see `pos_fibonacci` entry). Run `tools/lib/bench_paired.py` `pos-arena-{token,chunk,changed,rare-mask}` vs `pos-arena-uniform` on png_read (`chunk` needs a container corpus). Open: (a) `effector` had no drained map in 3k execs (SkipDet skipped 111/112 seeds) -- measure its reach on long runs before A/B, it is not subset-testable (needs the det stage); (b) `changed` credits only ~25% of rounds: parents without a recorded path hash (spliced/generated inputs) -- record one at admission or accept; (c) neither `changed` nor `rare_mask` persists state. - [ ] **Pre-existing, found 2026-09-30:** `test_regression_hail_mary_gates.py` fails on base (`cuckoo_seed_filter`, `swap_walk` missing from `_HAIL_MARY_FLAGS` or an exclusion); `tools/build_targets.sh` ASAN variants fail in the cloud container (non-ASAN builds fine). - [ ] **Target arena unmeasured** (2026-09-30) — `--target-arena` (Elo over target schedules + Gale-Shapley) has unit tests and a 2-target smoke run only (uninstrumented, no signal). Needs a paired A/B vs `--target-schedule weighted` on ≥2 instrumented ASAN fuzzgoat builds (this container lacks the ASAN runtime). Gale-Shapley batch (32) and target taste (fewest tries) are untuned. `auction` arm (max-weight on Thompson draws, 2026-09-30) same status: A/B it against `gale_shapley` in that run (Elo share + edges); if it wins, consider sharing one yield-row table between the two arms (each keeps its own LRU today, 2x memory and record cost). diff --git a/docs/architecture.dot b/docs/architecture.dot index 4a63a267..5d096454 100644 --- a/docs/architecture.dot +++ b/docs/architecture.dot @@ -56,7 +56,7 @@ digraph fuzzer_tool { subgraph cluster_sched { label="1 · Scheduling — what to fuzz next"; color="#bcd0c4"; fillcolor="#f2f8f4"; fontcolor="#33604a"; - picker [label="services/seed_picker.py + core/schedules.py\lweights · Pareto front · crowding · saturation gate\lFAST COE RARE MMOPT LIN QUAD GO AFLGO ENTROPIC KATZ\lKRUSKAL-COUNT (coupling walkers · recombination)\l", + picker [label="services/seed_picker.py + core/schedules.py\lweights · Pareto front · crowding · saturation gate\lFAST COE RARE MMOPT LIN QUAD GO AFLGO ENTROPIC KATZ DOPPLER\lKRUSKAL-COUNT (coupling walkers · recombination)\l", fillcolor="#d8eade", fontcolor="#1f4433"]; bandit [label="core/schedulers/ (16) + elo · shapley\lmonte_carlo mopt exp3 gp_ucb cmaes contextual\lducb swucb kl_ducb kl_swucb cucb hierarchical replicator epsilon_greedy mcts katz\l", fillcolor="#d8eade", fontcolor="#1f4433"]; diff --git a/docs/images/architecture.png b/docs/images/architecture.png index a28f9c26e8a36c1769a0b8a786242a4418b5f4fb..a60e9f77a61c2b15abea0210ca185ace7f00aecd 100644 GIT binary patch literal 613831 zcmeFZXH-+$7d?u46+2=91p$d53IfulE7GfU=_^O0c3>|#|kL6~*3pSi3pqy-C<`pYBg`tm`D9O*QI3+ZVNP7OL&`AF z9z7)V!lCC}^}Z9aho2W+|I`9W)5#SR9}^oJf7o}5LRvZ|f%Xb*N>~^xgbrC@c*DdU zbyc6*B{<$4`Dt&!%8_UaEEK}dvIe;B((y`U-dwlFiFDol2aVN=U z-{YFA^wG^?rw-lt>F0$DAKP@H7IX8`_2V7=O{1D{jccw?Dc1N~J9^2Cz6RNC*QlY~ zcwSOwpa|sXX41V4KMIPw1)pVmGxNgqhJ_)&je2_9>fkxIkH5~2Z!y~!-!Ij;)iK!A zhR|wB3pxGIjjv(SMt#S?E|Ys9n1GNW6|fL@_dUY4PEf#<;5sq)R_5mCZ$VI|zo)K5 z1pUwJs230CdS3<`Qhg=i7KF80Owp@?Y)`Jwn9i6sdsjO@KQS?7)YI6w-_mrNLhW`$ z%hkP6_;TV+q@p(K-5<5J1z=+T%wd8$47lE-^@!v!*9Du5fQO>O(VQy=kJKVswx;$w z=uyk^^2$a0`UVDr2!lK*0#&uadcJ{;rJ=qtiem}4JiXO*yULS~U7WAjt6$MZpyA{( zG1v2_ruylHE}TjSrcYOt@yk4V+f~g!W^RvF&T@_^SaK;B6dbtL7SnHacHeQkottKP z%=?japAAw?fjd~1;R-!M=(==9MoiOBVZNKoD^bZzpOdaSuKu15l^Z}9EHTN(CPnmC z&k^*7hw1&ESy);Us4P#sW@F~#5ZypoQLnG>++SHT>YF@TsI|NKfg<3CwO7A?N4QiN zQm=2Xb3I&O)W<v9{dEO2(t#QL&-@}+(Y5&1LTK!{kS?i zAx@J$*ACaxis@-V{;*w0jVqfdTV)9XQT>i+O1^&Vph;L{sIoj-UR5DDSZQnHRdN5w z)&KEm{y$Es3Y;>j>@R9ZM>i*@g9N+Sp)!vd2uoRf_||O*t6?JL&O*6UhX|Ojd0Qk+ z*hwXwL~)q9XN5}x9ku#5()II=h#2@tsjZ1>h^nXlz3;y#@ZKjQIXEC-e}dIufk0hx zF9qXt{Kygda15!q^xAD+N6>Iz3k#K_ z@-0yJrH(sTSe}mTbDB}4qW@(%(k|@&^0##T$Vi1{XhD*(1D;A!grY1+O(C>z>wHAF z2SL^FkudpxBA~y{)3C$!EK!hG6ukbKJ*rhhAe$nYakaJK%@esOCyBr!b0wIV*SsjS zT3xiPBA^dJuaT*dx{;$QA53e2T-#W?6jpw&X3uk~H1VcUwMZ?)DALW%slsP_q`6SD zhmTiU+ICaU{xcR9N?P?#;fz2JeUnk(R-37C4YJ@awRvHld2f0(9}bsc5fHE+i@c$4 zU>|Iw4P4k5VREq2n%<;#&FM15U6x?ksfAT`h{8Qet^vPx})e!1e)=N*etZPA?e3j=#m7X=3dX}L|V zg~g1ZpQ0T&sasmQ8ALrtMp|33SohYlZiq&?#>7_FoVl3=iOwzGJ|Q~#cB7xTli1eY z9jNP)9dDeB+58VoghUbpZOSbO4!>Y+X=$)Ke$P28BLivQ>(orZ1vYnz6DT%9pK}6sq^zLQ);+W6w0Txj@xnjzc%A)shN6#)8;1P0 zf`UZNXp