From 721955dd71dc8ea121a1112973f940b95a463fac Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 27 Sep 2026 14:56:41 +0000 Subject: [PATCH] =?UTF-8?q?docs(todo):=20COMPCOV=20ffmpeg=20A/B=20?= =?UTF-8?q?=E2=80=94=20null=20coverage,=20corpus=20bloat?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 12 seeds x 2k iters, L0/L2 + L0 control, corpus replayed without COMPCOV. Real edges 5W/7L p=0.77; L2 corpus +70% (12W/0L) for the same coverage; ~5% eps cost. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01Qc7F1N9YjmuzLAxiE1XJhp --- docs/TODO.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/TODO.md b/docs/TODO.md index b555c11b..1a95a848 100644 --- a/docs/TODO.md +++ b/docs/TODO.md @@ -18,7 +18,7 @@ - [ ] **Leak severity is conditional on edge count** (2026-08-23) — the fixed header clobber only carries stale coverage forward on targets with more than `(map_size - 24) / 8` live edges (8,189 at the default 65,536-entry map). Measured: entry index IS the edge id, assigned sequentially by `__sanitizer_cov_trace_pc_guard_init`, and `png_read.so` has ~4,600 guards, so nothing in `targets/` reaches the threshold. Verified deterministically by planting an entry at index 20,000: reported as live coverage pre-fix, filtered post-fix. A vendored ffmpeg or grep build would cross it naturally — worth re-running the probe (`docs/learnings/2026-08-23-shmat-sentinel-and-header-clobber.md`) once one is available. - [ ] **GEP yield is optimization-level dependent** (2026-08-28) — the layer-3 `trace-gep` callback only sees indices clang still models as a GEP when the sancov pass runs; at `-O1`+ an array index folded into an addressing mode is already gone (measured: 17 call sites emitted, none firing for the indexed load the test exercises; at `-O0` it fires). Vendored `--vendor-tracecmp` targets build at `-O2`, so the divisor half of trace-div/trace-gep is the half that pays there. Worth measuring real GEP yield on a vendored libpng build before deciding whether the flag earns its stream volume. - [ ] **`reset_bitmap()` is redundant on the runner path** (2026-08-23) — `services/runner.py` calls `shm.reset_edge_map()` immediately before `run_one()`, and `run_one` has exactly one call site, so the full-table memset in `InProcessRunner.reset_bitmap()` (reached from `_run_c_direct`) resets a table the generation bump already invalidated. Dropping the call would save a `shm_size * 8` byte memset per execution on the hot path of the fastest execution mode. Not done here because it wants measuring on a built target, and clang was unavailable in the authoring environment. The header-clobber half of this was a real bug and IS fixed — see `docs/learnings/2026-08-23-shmat-sentinel-and-header-clobber.md`. -- [ ] **COMPCOV review leftovers** (2026-09-27) — ASLR-stable site keys, NUL-bounded string walks and prev_loc-neutral marks are fixed. Open: (a) no `docs/DEEP_DIVE.md` section; (b) levels 0/1/2 are bare ints clamped in three places (`cmplog.py`, `fuzzer.py`, shim) — want one enum; `is_const` int param likewise; (c) `test_compcov.py`'s original tests build with gcc and `test_level_1_ignores_non_const_trace_cmp` asserts `0 == 0` — no positive level-1 test; (d) silent no-op when cmplog auto-detect fails or the shim is `__AFL_PRELOAD_ONLY`; (e) png A/B null (2026-09-27, 10 seeds x 10k execs, levels 0/1/2, final corpus replayed on `png_read_tracecmp` so marks cannot count): real edges median 1254 / 1280 / 1332, L2-L0 7W/3L p=0.34 -- inside the L0-vs-L0 control (7W/3L, median +96, IQR [-58, +197]). eps unchanged. Needs ~40+ seeds or reps to resolve; fuzzgoat A/B still owed (Hard Rule 52). Metric trap: the campaign's own `Edges discovered` counts marks (L2 +1072, 10W/0L), so `bench_paired.py --metric edges` would report a false win; (g) no stock target runs Layer 2: `png_read_tracecmp*` links the shim without `-D__AFL_CMPLOG=1`, so its 3147 trace-cmp sites hit the preload-only shim where COMPCOV is a no-op; the A/B used a hand-built variant (same link line + `-D__AFL_CMPLOG=1 -Wl,-Bsymbolic`) -- wire it into `tools/build_targets.sh`. +- [ ] **COMPCOV review leftovers** (2026-09-27) — ASLR-stable site keys, NUL-bounded string walks and prev_loc-neutral marks are fixed. Open: (a) no `docs/DEEP_DIVE.md` section; (b) levels 0/1/2 are bare ints clamped in three places (`cmplog.py`, `fuzzer.py`, shim) — want one enum; `is_const` int param likewise; (c) `test_compcov.py`'s original tests build with gcc and `test_level_1_ignores_non_const_trace_cmp` asserts `0 == 0` — no positive level-1 test; (d) silent no-op when cmplog auto-detect fails or the shim is `__AFL_PRELOAD_ONLY`; (e) png A/B null (2026-09-27, 10 seeds x 10k execs, levels 0/1/2, final corpus replayed on `png_read_tracecmp` so marks cannot count): real edges median 1254 / 1280 / 1332, L2-L0 7W/3L p=0.34 -- inside the L0-vs-L0 control (7W/3L, median +96, IQR [-58, +197]). eps unchanged. Needs ~40+ seeds or reps to resolve; fuzzgoat A/B still owed (Hard Rule 52). ffmpeg A/B null too (2026-09-27, 12 seeds x 2k iters, L0/L2 + L0 control, stock `ffmpeg_read_nosan.so` -- Layer 2 live, 100k const-cmp sites; corpus replayed in-process with COMPCOV unset): real edges median 13810 / 13858, L2-L0 5W/7L p=0.77, control 3W/9L. L2 bloats the corpus +70% (1215 vs 714 entries, 12W/0L p=0.0005) for the same coverage -- real edges per entry 12.1 vs 22.5 -- and costs ~5% eps (2W/8L, p=0.11). Metric trap: the campaign's own `Edges discovered` counts marks (L2 +1072, 10W/0L), so `bench_paired.py --metric edges` would report a false win; (g) no stock target runs Layer 2: `png_read_tracecmp*` links the shim without `-D__AFL_CMPLOG=1`, so its 3147 trace-cmp sites hit the preload-only shim where COMPCOV is a no-op; the A/B used a hand-built variant (same link line + `-D__AFL_CMPLOG=1 -Wl,-Bsymbolic`) -- wire it into `tools/build_targets.sh`. - [ ] **`test_integration.py` ASAN fixtures build with gcc** (2026-09-27) — violates Hard Rule 4, but gcc's shared libasan is what exposed the cmplog preload link-order abort (`verify_asan_link_order`). Move to clang and keep a `-shared-libasan` case (`test_regression_cmplog_asan_link_order.py` covers it). - [ ] **Slow tests time out under `pytest -n 4 --timeout=30`** (2026-09-27) — `test_regression_scheduler_operator_reach.py` (BayesUCB/BOGPUCB/CMAES, 13-20 s serial) and `test_integration.py::test_fuzzer_eps_minimum` (17 s, throughput-sensitive) pass serially, fail under load. Shrink their budgets or mark `slow`. - [ ] **ptrace breakpoints only on dominator-tree leaves** (2026-09-26) — a hit block implies its dominators ran, so `ptrace_coverage.py` could place int3 on leaves only and infer the rest. Edges `(prev, curr)` are not implied the same way; measure breakpoint count and edge-set loss on fuzzgoat first. See `docs/learnings/2026-09-26-dominators-chk-worst-case.md`.