Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

84 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

SplitLab

A from-scratch A/B experimentation library — power, fixed-horizon inference, CUPED, sequential testing — validated against theory by simulation before it was allowed to touch data, then run against two real experiment corpora, ending in analyst memos that make ship / don't-ship calls on individual experiments.

Start here → Memo #2. A ship call on the largest experiment in the ASOS archive, and the one number its owner must not carry away from it. (all five memos · methodology and 29 limitations)


The call

From memo #2, on a test that ran 129.6 million users through two versions over 131 days:

Call: ship. The effect is large, both ends of the interval are worth acting on, and all four metrics agree. The finding worth your time is the second one: this experiment answered its question in the first week and then ran for another eighteen.

The decision was available on day 6.5. The number was not.

At day 6.5 the lift on the table read +3.14%. The final figure is +1.94% — the early number was 1.6× too large. […]

So if this experiment had been stopped at its earliest defensible point, you'd have shipped the right thing and told finance a number 60% too high.

Sequential monitoring buys you a faster decision. It does not buy you a faster estimate, and the two are easy to conflate.

Risk flag — do not build a business case on an early-stopped effect size. The safe pattern is to separate the two: stop the experiment on the boundary, then quote the effect from a fixed window you committed to in advance.

Effect trajectory and stopping boundary for experiment 2c8a04

The finding behind that risk flag

The textbook explanation for an inflated early estimate is boundary selection: a rule fires when the statistic is large, so conditioning on "we stopped" ought to select for an overstated effect. Measured across all 214 early stops in this archive, that is not what dominates.

Share of early stops that overstate mSPRT O'Brien–Fleming
the absolute effect, vs. the same series' final read 50.6% 48.8%
the relative lift — the number a memo actually quotes 77.6% 82.2%
…because the control mean grew between stop and end 98.8% 96.9%

The effect at a stop is a coin flip. The lift is what inflates, and the cause is the denominator: these are cumulative rates over a maturing cohort, so a user exposed on day 1 has the whole run to convert and one exposed on day 50 does not. Mid-flight the baseline is simply lower, and a day-2 lift is inflated with no stopping rule anywhere near it.

This measurement refuted the explanation the memos originally gave — all three sequential memos had attributed the inflation to the boundary and were rewritten to name the denominator instead. Bounded twice: the final read contains the stopping look's own data, so it understates true conditional bias rather than refuting the theory, and the rates are descriptive, since ASOS series are not independent replicates. Full treatment in WRITEUP.md §6.5 and L28.

Type-I error under repeated monitoring

Naive monitoring across 50 looks rejects a true null 32.3% of the time against a nominal 5%. The same looks under the mSPRT hold at 1.4%, and under O'Brien–Fleming spending at 5.0%. The textbook 1-(1-α)^k bound would have predicted 92% — cumulative looks are correlated, so quoting it overstates the damage nearly threefold.

Two regimes of causal inference

This is one half of a pair. CourtIQ does causal inference where randomization is impossible — NBA player impact, Bayesian RAPM, adjustment for confounding you cannot design away. SplitLab does it where randomization is available: design, then inference on a treatment effect.

The two problems need different instincts. Adjustment asks what you must control for; design asks what you must decide before the data arrives, and how much traffic that decision costs.

What else it found

Result Number Artifact
Sequential replay, 338 ASOS series across 75 experiments O'Brien–Fleming stops early on 38.2%, median 41.5% traffic saved; the mSPRT on 25.1%, median 66.7%. The rules cross over — neither is "the" answer asos_replay.json
Disagreement with the fixed-horizon read 19.8% (mSPRT), 13.3% (OBF). A comparator carrying its own error, not a sensitivity/specificity claim "
SRM incident recovery Both of the maintainers' published edges recovered to +0 days, alarm rate 0.70 inside the band against 0.005 outside, with an independent hour-of-day signature upworthy_srm_recovery.json
Strict A/A instrument (it could have failed) 8 / 203 = 0.0394 against an exact expectation of 0.0501, z = −0.70, detection bound 1.53 pp upworthy_calibration.json
CUPED coverage under an estimated θ 0.9459–0.9533 across 30 cells against a nominal 0.95 — the predicted small-n undercoverage does not happen. The real cost is attenuation, ≈ δ/(2n) cuped_coverage.json

Studies these corpora cannot support, and why

Leading with the constraint is the point of this section. Every item was either killed at the due-diligence gate or measured and found unreachable — none of it was discovered late and quietly dropped.

  • SRM detection scored against ground-truth labels. Killed at the Phase 0 gate. Upworthy's problem field ships in no released file and is a deterministic function of created_at — set for every test in a date window, so there are zero unflagged in-window tests and per-test sensitivity is impossible. Worse, the window was itself found via a sample-ratio-mismatch report, so agreeing with it is replication of a known diagnostic, not validation of ours. The study was redesigned into incident recovery, and only the hour-of-day mechanism check is independent of the label.
  • Any CUPED claim from real data. Neither corpus has user-level pre-period covariates. CUPED is validated by simulation only, and no number in any memo, table or figure is sourced otherwise.
  • A "clear null" memo from Upworthy. Measured, not assumed: against a 10% relative floor the 95% interval's half-width is a median 4.5× that floor and never below 1.1×. No Upworthy test can rule out a material effect in both directions, so the corpus supports "no detected difference" and never "no material difference." The null memo comes from ASOS.
  • Mapping where the normal approximation breaks. Upworthy has essentially no small-n regime — the 5th-percentile package has ~2,010 impressions. The calibration study characterises the realistic operating grid, which is a weaker and more honest claim.
  • Seasonality, in any ASOS memo. The archive publishes relative day offsets and no calendar time, so a 131-day run cannot be checked against seasonal or promotional context.
  • "False-positive rate on real traffic." The hypergeometric A/A split conditions on exactly the null the test assumes, so it cannot fail. It is empirical-grid calibration. The 203 strict identical-stimulus pairs are the genuinely falsifiable instrument, and their 1.53 pp detection bound is quoted everywhere their rate is.

Reproducing

Neither corpus is vendored — both are CC-BY 4.0 but run to 80 MB, so data/ is gitignored and provenance is pinned by digest instead. Every file below downloads without authentication.

ASOS Digital Experiments — OSF osf.io/64jsb (DOI 10.17605/OSF.IO/64JSB), companion to arXiv:2111.10198. Direct: https://osf.io/62t7f/download

File Bytes SHA-256
data/asos/asos_experiments.parquet 1,003,239 bdf88b27185d3f7e65912cbb421129a32ae347c2fb0d9442fe54d74266551524

Upworthy Research Archive — OSF osf.io/jd64p.

File Bytes SHA-256
data/upworthy/exploratory.csv 14,260,949 8368313b060f4015a0c6fb34e6d788163cee29554144aba9390b14922eb9d8ce
data/upworthy/confirmatory.csv 66,517,697 b2a88288c88f2b67d40413c9bdbebe79a1c0b5d9a9354180549f59f3bfff685f
pip install -e '.[dev]'
pytest                          # full suite
python -m validation.run all    # the five simulation studies (no download needed)

# with the two archives placed as above:
python -m corpora.upworthy.ingest
python -m corpora.asos.ingest

python -m validation.report                # the gate
python -m validation.report --reproduce    # + re-run and compare the studies

Upworthy asks that confirmatory analyses follow a peer-reviewed analysis plan. A portfolio project cannot obtain peer review, so the substitute is in-repo pre-registration: every filter threshold and every seed is a literal in corpora/upworthy/ingest.sql and corpora/asos/ingest.sql, committed before any analysis number was computed. Filters flag rows rather than deleting them, so a record reports what a threshold cost instead of asserting it was applied.

Layout

splitlab/      power, fixed_horizon, cuped, sequential, srm, simulate
validation/    five simulation studies that had to pass before real data
corpora/       ingest (plain SQL over DuckDB) and study code for both archives
memos/         five analyst memos and their analysis packs
results/       committed JSON records and the SVGs generated from them
tests/         reference libraries live here, never in splitlab/

Rigor apparatus — the floor, not the headline

splitlab/ has no statsmodels or pingouin dependency; scipy.stats distributions only. Reference implementations appear in tests/, to check the work rather than to do it. 956 tests, 100% coverage across splitlab/, validation/ and corpora/.

Every result artifact carries the git SHA it ran under plus a dirty-tree flag, so the SHA provably describes the code that produced it. The sequence is always commit the code, run it, then commit the artifact — all 17 records carry git_dirty: false. Records carry no wall-clock timestamp, which makes them byte-reproducible.

The report gate (python -m validation.report) re-derives 60 headline numbers in this README, in WRITEUP.md and in the memos directly from the committed artifacts, and fails if the prose and the artifact disagree. Several are stored nowhere as a field — the stopping-estimate ratios above are medians over 338 series cells — so claims carry a derivation rather than a key path. --reproduce re-runs the five simulation studies and compares what they write: figures byte for byte, records on study/seed/config/cells, since git_sha necessarily differs once HEAD has moved. The 12 corpus artifacts need data/ and are reported as skipped, never as passed.

The gate is only worth having if it can fail. Mutating every claim's value showed 13 of the 50 then registered still matching somewhere in their documents — 8 occurs 296 times in WRITEUP.md — so claims now bind the phrase the prose wraps the number in. All 60 currently fail when the value is perturbed.

About

From-scratch A/B experimentation library: power, fixed-horizon inference, CUPED, sequential testing. Validated by simulation, then run against the Upworthy and ASOS experiment archives, ending in analyst memos with ship/no-ship calls.

Topics

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages