Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

### Added

- **Power Doppler schedule** (`--schedule doppler`, `core/power_doppler.py`): per-seed ensembles of mutant hit counts; mean + SVD wall filter, CFAR χ² flow detection; flow power scales seed energy to `[1, max_mult]`.
- **OS / network scheduler ports**: seed arms `mlfq`, `stride`, `eevdf`, `bfq`, `sfq`, `codel`, `aimd`, `p2c` (`--seed-<name>-scheduler`) and op arms `op_stride`, `op_p2c` (`--op-stride`, `--op-p2c`); `core/fair_queue.py` gains `Stride` and `EEVDF`. Elo arms, in `--hail-mary`; seed arms also run without `--elo`. Falsification and adversarial tests. No paired benchmark yet.
- **Ten op mutators** for in-tree targets and text decoders: `json_mutate` (fuzzgoat), `sql_mutate` (sqlite SQL path), `ecdsa_field_mutate` (secp256k1), `recompress_lz4`, `recompress_png_idat` (format band, sniffer-gated); `encoding_wrap`, `escape_mutate`, `ascii_float` (structural); `utf16_transcode`, `nest_bomb` (radamsa). Tests: `tests/test_{json_mutate,sql_mutate,ecdsa_field_mutate,recompress_roundtrip,text_codec,nest_bomb}.py`. Discovery effect on fuzzgoat unmeasured.
- **`covering_array_webp`, `covering_array_isobmff`, `covering_array_zip`** operators (`core/mutations/covering_array_container.py`): pairwise sweep of WebP `VP8X` (riff_size, chunk_size, flags, reserved, width, height; 49 rows), ISO-BMFF `ftyp` (size, major brand, minor; 30 rows) and the first ZIP local file header (version, flags, method, sizes, name/extra length; 50 rows). Each gated on its own magic. Selection share unmeasured; ISO-BMFF `tkhd` not covered (variable depth).
Expand Down
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -143,7 +143,7 @@ per-sub-operator reward instead of uniformly (`--no-adaptive-havoc` restores uni
| Metropolis admission | `--metropolis` | Accept non-improving inputs with P = exp(−ΔE/T) |
| Secretary stopping | `--secretary` | Secretary-problem stopping ranks (display only) |
| honggfuzz power factors | `--honggfuzz` | Novelty decay, freshness, fertility, density, entropy and timeout penalties |
| AFL++ power schedules | `--schedule` | FAST/COE/RARE/MMOPT/LIN/QUAD/GO/AFLGO/ENTROPIC seed-level energy |
| AFL++ power schedules | `--schedule` | FAST/COE/RARE/MMOPT/LIN/QUAD/GO/AFLGO/ENTROPIC/DOPPLER seed-level energy |
| AFLGo directed annealing | `--schedule aflgo` | Exact AFLGo power factor — symmetric 32×/1/32× energy by distance-to-target with time-based cooling (`--t-x`, `--aflgo-cooling`) |
| Entropic power schedule | `--schedule entropic` | libFuzzer `-entropic`: energy ∝ log(1 + rare-feature count) from already-tracked rare-edge ownership |
| **K-Scheduler Katz centrality** | auto | On trace-pc targets: whole-program ICFG → horizon graph (contracted visited deletion, DAG) → out-degree Katz with β from node-hit counts; Elo-rated `katz` seed arm plus a clamped `--schedule katz` energy |
Expand Down
5 changes: 3 additions & 2 deletions docs/DEEP_DIVE.md
Original file line number Diff line number Diff line change
Expand Up @@ -84,7 +84,8 @@ For production and sensitive binaries using AFL family fuzzers is the best cours
- **AFLGo power schedule** (`--schedule aflgo`, `--aflgo-cooling exp|log|lin|quad`, `--t-x MINUTES`): the exact `calculate_score()` distance section from AFLGo's afl-fuzz.c. Temperature `T` follows the chosen cooling over t_x minutes to exploitation (`exp`: 1/20^progress; `log`: 1/(1+2·ln(1+progress·13358.7268297)); `lin`: 1/(1+19·progress); `quad`: 1/(1+19·progress²)); `p = (1−nd)(1−T) + 0.5T` with `nd = (d−min)/(max−min)` normalized over the observed queue (min/max tracked per run, excluding the no-data sentinel); `factor = 2^(2·log2(32)·(p−0.5))` — symmetric around 1.0: early campaigns treat every seed equally, late campaigns give near-target seeds up to 32× energy and far seeds as little as 1/32×.
- **AFLGo SHM-tail distance channel**: compiled into EVERY shim-linked target since `__AFL_DISTANCE_MODE` defaulted to 1 (2026-08-24; `-D__AFL_DISTANCE_MODE=0` opts out). Inert unless directed mode uploads a distance table (`DistanceTableShm`/`__AFL_DIST_SHM_ID`): without one, sum/count stay 0 and every reader takes the Python-side path. `build_targets.sh --distance` additionally builds the trace-pc-instrumented `*_dist.so`/`*_dist_asan.so` variants, where the shim accumulates per-block distances in `__sanitizer_cov_trace_pc()` — the PC (relative to the dladdr-derived object base) probes an open-addressing table of `{key, dist}` entries (packed 12-byte layout; the 4-byte header holds the slot *capacity*, a power of two ≥ 2×entries so empty slots exist, and the builder hash-inserts at `key % capacity` with linear probing to mirror the shim's probe — uploaded by the fuzzer at startup via `DistanceTableShm`/`__AFL_DIST_SHM_ID`), accumulating sum/count into the 16-byte SHM **tail** (after the edge table: `u64 dist_sum`, `u64 dist_count`), written at reset, at process exit (subprocess runs never call reset), and per-iteration in in-process modes via `__afl_dist_flush` (direct_lite has no process boundary, so the runner flushes the tail after each `run_one`). Per-execution `avg_distance = sum/count/100` is read straight from the tail and preferred over Python-side computation; blocks without a table entry don't count (AFLGo semantics). The table's PC keys are recovered by scanning text for `call __sanitizer_cov_trace_pc` sites (modern clang emits no `__sancov_pcs` for trace-pc) mapped to valued blocks via the CFGs — `TargetDistance.pc_distance_table()`. The shim's sanitizer-coverage callbacks are hidden-visibility so a libasan LD_PRELOAD cannot interpose over them in PIE builds. `tools/gen_distance_table.py` emits the table as C or text for inspection. Without the table (count==0) everything degrades to the Python-side path. Works in subprocess, direct_lite, and persistent modes (ASAN direct_lite requires libasan preloaded at fuzzer-process start — the `use_direct_lite` gate). The periodic stats line shows live distance when directed mode is active: `dist: avg:<tail avg> min:<observed min> max:<observed max>` (or `no-data`). `build_targets.sh --distance` builds both `*_dist.so` (no-ASAN) and `*_dist_asan.so` (ASAN) variants with the cmplog shim linked in, so `--cmplog` keeps them in direct_lite mode. Startup reports `[*] Distance instrumentation: detected` when the target carries the channel (the shim's `__afl_dist_flush` or a defined `__sanitizer_cov_trace_pc`), mirroring the AFL-instrumentation check. With `--elo` in directed mode, `aflgo` joins the Elo-arbitrated seed-strategy pool — a distance-pure arm picking `P(seed) ∝ exp(-2·norm_dist)` (distinct from the generic `weighted` arm, which blends distance with speed/size/entropy).
- **AFLGo distance-annealed schedule** (`--schedule go`, requires `--target-functions`): wires the precomputed `avg_distance` (per-seed distance to directed targets) and `_anneal_progress` (exploration/exploitation annealing variable) into `SeedScorer.score()` for mutation budget scaling. During exploration phase (`anneal_progress` ≈ 0): uniform energy. During exploitation phase (`anneal_progress` → 1): `energy *= exp(β · (1 - norm_dist))` where `β = anneal_progress * 5`, capping at 100x. Seeds near the target get exponentially more mutations as the campaign matures. Previously these metrics only influenced seed selection but not mutation intensity.
- **Power schedules** (`--schedule base|fast|coe|rare|mopt|lin|quad|go|aflgo|entropic`): AFL++ power schedules ported to control mutation budget per seed via `SeedScorer`. Each schedule modifies a base score (100) by frequency-based factors. Honggfuzz-style novelty decay, density, fertility, freshness, and entropy factors are applied multiplicatively on top. `entropic` (libFuzzer `-entropic`) scales energy by `1 + log2(1 + rare)`, where `rare` is the larger of `rare_edge_count`/`tc_ref` already collected for RARE/honggfuzz scoring — an approximation of libFuzzer's feature-frequency Shannon entropy using signal the fuzzer already tracks.
- **Power schedules** (`--schedule base|fast|coe|rare|mopt|lin|quad|go|aflgo|entropic|doppler`): AFL++ power schedules ported to control mutation budget per seed via `SeedScorer`. Each schedule modifies a base score (100) by frequency-based factors. Honggfuzz-style novelty decay, density, fertility, freshness, and entropy factors are applied multiplicatively on top. `entropic` (libFuzzer `-entropic`) scales energy by `1 + log2(1 + rare)`, where `rare` is the larger of `rare_edge_count`/`tc_ref` already collected for RARE/honggfuzz scoring — an approximation of libFuzzer's feature-frequency Shannon entropy using signal the fuzzer already tracks.
- **Power Doppler schedule** (`--schedule doppler`, `core/power_doppler.py`): ultrasound power Doppler on coverage. Slow time = successive mutants of one seed (32 per ensemble); pixel = edge; sample = `log2(1 + hits)` (raw SHM counts, not buckets). Wall filter: mean removal (static path), then SVD components whose participation ratio spans ≥ half the seed's edges (and ≥ 4) are dropped as clutter — an early reject moving the whole path at once ("flash"). CFAR: residual power per edge vs `σ² · χ²_{dof}(1 − 10⁻³)`, `σ²` = median residual variance (floor 10⁻³). Seed power = summed flow power / dof; energy = `log1p(p)/log1p(max p)` ∈ [0, 1], scaled to `[1, max_mult]` like `katz`. Unscored or static seeds stay 1×. Gram eigendecomposition (n×n) replaces the full SVD (~6× faster). Bounded: 64 open ensembles × 2048 edges (float32, ≤16 MiB), 4096 scores (LRU). SHM coverage only; cost ~100 µs/exec at 2k live edges (dict→array conversion dominates), zero when off. Unmeasured — see `docs/TODO.md`.
- **Favored set / cull_queue** (`core/schedules.py`, `services/fuzzer.py`): AFL-style `top_rated` minimal-set-cover selection. For each edge, the cheapest seed covering it is selected, then a greedy cover builds the favored set. FAST and COE schedules apply energy bonuses to favored seeds (`_fast_factor`, `_coe_factor`, `coe_skip`). `_cull_queue()` runs periodically during fuzzing and updates `self._favored`; the score call site passes `favored=(seed_key in self._favored)` so the scheduler actually uses it.
- **Bayesian seed quality** (`--bayesian`): `BayesianSeedQuality` (`core/seed_quality.py`) maintains a Beta-Bernoulli posterior per seed over `P(outcome = new_coverage)`. Thompson sampling naturally balances explore/exploit without a manual temperature knob — unexplored seeds have high posterior variance and get sampled. The `record_outcome()` feedback loop is now wired in `fuzz_one()` (was previously a dead code path with all posteriors stuck at Beta(1,1)). State is persisted to `seed_quality.json` and restored on resume.

Expand Down Expand Up @@ -587,7 +588,7 @@ fuzzer-tool rank ./target -d corpus -n 10 --dump top_seeds
| `--wfc` | Wave Function Collapse structural generation (chunk reordering, pixel generation) |
| `--enable-smt-z3` | Z3-based SMT solving for arithmetic constraint solving on cmplog pairs |
| `--hw-perf` | Hardware performance counters via perf_event_open (instructions, branches, misses) |
| `--schedule base\|fast\|coe\|rare\|mopt\|lin\|quad\|go\|aflgo\|entropic` | AFL++ power schedule (`aflgo` = exact AFLGo distance annealing, `entropic` = libFuzzer `-entropic`, log-scaled rare-feature energy) |
| `--schedule base\|fast\|coe\|rare\|mopt\|lin\|quad\|go\|aflgo\|entropic\|doppler` | AFL++ power schedule (`aflgo` = exact AFLGo distance annealing, `entropic` = libFuzzer `-entropic`, log-scaled rare-feature energy, `doppler` = power Doppler flow energy) |
| `--aflgo-cooling exp\|log\|lin\|quad` | Cooling schedule for the `aflgo` power factor (default exp) |
| `--t-x MINUTES` | AFLGo time-to-exploitation in minutes; temperature cools to 1/20 over this window |
| `--markov-order N` | Markov chain order(s), comma-separated (e.g. '0,1,2' for ensemble) |
Expand Down
1 change: 1 addition & 0 deletions docs/TODO.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,7 @@
- [ ] **ptrace breakpoints only on dominator-tree leaves** (2026-09-26) — a hit block implies its dominators ran, so `ptrace_coverage.py` could place int3 on leaves only and infer the rest. Edges `(prev, curr)` are not implied the same way; measure breakpoint count and edge-set loss on fuzzgoat first. See `docs/learnings/2026-09-26-dominators-chk-worst-case.md`.

## Scheduling
- [ ] **A/B `--schedule doppler`** (2026-10-01) — power Doppler seed energy (`core/power_doppler.py`) is unit-tested and wired; never measured. Paired `bench_paired.py` vs `fast` on clang-built fuzzgoat (Hard Rule 52). Open: (a) ~100 µs/exec at 2k edges, mostly `get_edge_counts()` dict → numpy; a numpy accessor on `ShmCoverage` (exposes `_active_columns`, Hard Rule 34 — needs approval) removes it; (b) ensembles span picks, so a seed needs 32 picks to score — tune `ensemble` vs pick rate; (c) unused: `flow_edges()` (input-sensitive edge set) could feed position/operator targeting; (d) scores not persisted across `--resume`.
- [ ] **A/B the OS / network scheduler ports** (2026-09-30) — `mlfq`, `stride`, `eevdf`, `bfq`, `sfq`, `codel`, `aimd`, `p2c` seed arms and `op_stride`, `op_p2c` are unit- and wiring-tested only. Run `tools/lib/bench_paired.py` per arm vs `seed_round_robin` / `drr` on fuzzgoat (clang, ASAN). Untuned: MLFQ allotment/boost, BFQ budget range, CoDel target/interval, AIMD alpha/beta/loss run. Open: (a) `sfq` groups only direct siblings (`parent_key`), and only under `--lineage`; a lineage-root flow would group whole families; (b) no arm persists state; (c) `seed_round_robin.select_seed` is O(n^2) per pick (`s in seed_ids` on a list): ~100 ms at 5000 seeds.
- [ ] **A/B the effector/token/chunk/changed/rare_mask position arms** (2026-09-30) — wired and unit-tested; a 3k-exec fuzzgoat run confirms live signal (`changed`: 673 moved / 96 unmoved / 2231 unmeasured rounds; `rare_mask`: target on 328/3000 rounds) but fuzzgoat cannot rank position arms (see `pos_fibonacci` entry). Run `tools/lib/bench_paired.py` `pos-arena-{token,chunk,changed,rare-mask}` vs `pos-arena-uniform` on png_read (`chunk` needs a container corpus). Open: (a) `effector` had no drained map in 3k execs (SkipDet skipped 111/112 seeds) -- measure its reach on long runs before A/B, it is not subset-testable (needs the det stage); (b) `changed` credits only ~25% of rounds: parents without a recorded path hash (spliced/generated inputs) -- record one at admission or accept; (c) neither `changed` nor `rare_mask` persists state.
- [ ] **Pre-existing, found 2026-09-30:** `test_regression_hail_mary_gates.py` fails on base (`cuckoo_seed_filter`, `swap_walk` missing from `_HAIL_MARY_FLAGS` or an exclusion); `tools/build_targets.sh` ASAN variants fail in the cloud container (non-ASAN builds fine).
Expand Down
2 changes: 1 addition & 1 deletion docs/architecture.dot
Original file line number Diff line number Diff line change
Expand Up @@ -56,7 +56,7 @@ digraph fuzzer_tool {
subgraph cluster_sched {
label="1 · Scheduling — what to fuzz next";
color="#bcd0c4"; fillcolor="#f2f8f4"; fontcolor="#33604a";
picker [label="services/seed_picker.py + core/schedules.py\lweights · Pareto front · crowding · saturation gate\lFAST COE RARE MMOPT LIN QUAD GO AFLGO ENTROPIC KATZ\lKRUSKAL-COUNT (coupling walkers · recombination)\l",
picker [label="services/seed_picker.py + core/schedules.py\lweights · Pareto front · crowding · saturation gate\lFAST COE RARE MMOPT LIN QUAD GO AFLGO ENTROPIC KATZ DOPPLER\lKRUSKAL-COUNT (coupling walkers · recombination)\l",
fillcolor="#d8eade", fontcolor="#1f4433"];
bandit [label="core/schedulers/ (16) + elo · shapley\lmonte_carlo mopt exp3 gp_ucb cmaes contextual\lducb swucb kl_ducb kl_swucb cucb hierarchical replicator epsilon_greedy mcts katz\l",
fillcolor="#d8eade", fontcolor="#1f4433"];
Expand Down
Binary file modified docs/images/architecture.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Loading