A from-scratch A/B experimentation library — power, fixed-horizon inference, CUPED, sequential testing — validated against theory by simulation before it was allowed to touch data, then run against two real experiment corpora, ending in analyst memos that make ship / don't-ship calls on individual experiments.
Start here → Memo #2. A ship call on the largest experiment in the ASOS archive, and the one number its owner must not carry away from it. (all five memos · methodology and 29 limitations)
From memo #2, on a test that ran 129.6 million users through two versions over 131 days:
Call: ship. The effect is large, both ends of the interval are worth acting on, and all four metrics agree. The finding worth your time is the second one: this experiment answered its question in the first week and then ran for another eighteen.
At day 6.5 the lift on the table read +3.14%. The final figure is +1.94% — the early number was 1.6× too large. […]
So if this experiment had been stopped at its earliest defensible point, you'd have shipped the right thing and told finance a number 60% too high.
Sequential monitoring buys you a faster decision. It does not buy you a faster estimate, and the two are easy to conflate.
Risk flag — do not build a business case on an early-stopped effect size. The safe pattern is to separate the two: stop the experiment on the boundary, then quote the effect from a fixed window you committed to in advance.
The textbook explanation for an inflated early estimate is boundary selection: a rule fires when the statistic is large, so conditioning on "we stopped" ought to select for an overstated effect. Measured across all 214 early stops in this archive, that is not what dominates.
| Share of early stops that overstate | mSPRT | O'Brien–Fleming |
|---|---|---|
| the absolute effect, vs. the same series' final read | 50.6% | 48.8% |
| the relative lift — the number a memo actually quotes | 77.6% | 82.2% |
| …because the control mean grew between stop and end | 98.8% | 96.9% |
The effect at a stop is a coin flip. The lift is what inflates, and the cause is the denominator: these are cumulative rates over a maturing cohort, so a user exposed on day 1 has the whole run to convert and one exposed on day 50 does not. Mid-flight the baseline is simply lower, and a day-2 lift is inflated with no stopping rule anywhere near it.
This measurement refuted the explanation the memos originally gave — all three sequential memos had attributed the inflation to the boundary and were rewritten to name the denominator instead. Bounded twice: the final read contains the stopping look's own data, so it understates true conditional bias rather than refuting the theory, and the rates are descriptive, since ASOS series are not independent replicates. Full treatment in WRITEUP.md §6.5 and L28.
Naive monitoring across 50 looks rejects a true null 32.3% of the time
against a nominal 5%. The same looks under the mSPRT hold at 1.4%, and under
O'Brien–Fleming spending at 5.0%. The textbook 1-(1-α)^k bound would have
predicted 92% — cumulative looks are correlated, so quoting it overstates the
damage nearly threefold.
This is one half of a pair. CourtIQ does causal inference where randomization is impossible — NBA player impact, Bayesian RAPM, adjustment for confounding you cannot design away. SplitLab does it where randomization is available: design, then inference on a treatment effect.
The two problems need different instincts. Adjustment asks what you must control for; design asks what you must decide before the data arrives, and how much traffic that decision costs.
| Result | Number | Artifact |
|---|---|---|
| Sequential replay, 338 ASOS series across 75 experiments | O'Brien–Fleming stops early on 38.2%, median 41.5% traffic saved; the mSPRT on 25.1%, median 66.7%. The rules cross over — neither is "the" answer | asos_replay.json |
| Disagreement with the fixed-horizon read | 19.8% (mSPRT), 13.3% (OBF). A comparator carrying its own error, not a sensitivity/specificity claim | " |
| SRM incident recovery | Both of the maintainers' published edges recovered to +0 days, alarm rate 0.70 inside the band against 0.005 outside, with an independent hour-of-day signature | upworthy_srm_recovery.json |
| Strict A/A instrument (it could have failed) | 8 / 203 = 0.0394 against an exact expectation of 0.0501, z = −0.70, detection bound 1.53 pp | upworthy_calibration.json |
| CUPED coverage under an estimated θ | 0.9459–0.9533 across 30 cells against a nominal 0.95 — the predicted small-n undercoverage does not happen. The real cost is attenuation, ≈ δ/(2n) | cuped_coverage.json |
Leading with the constraint is the point of this section. Every item was either killed at the due-diligence gate or measured and found unreachable — none of it was discovered late and quietly dropped.
- SRM detection scored against ground-truth labels. Killed at the Phase 0
gate. Upworthy's
problemfield ships in no released file and is a deterministic function ofcreated_at— set for every test in a date window, so there are zero unflagged in-window tests and per-test sensitivity is impossible. Worse, the window was itself found via a sample-ratio-mismatch report, so agreeing with it is replication of a known diagnostic, not validation of ours. The study was redesigned into incident recovery, and only the hour-of-day mechanism check is independent of the label. - Any CUPED claim from real data. Neither corpus has user-level pre-period covariates. CUPED is validated by simulation only, and no number in any memo, table or figure is sourced otherwise.
- A "clear null" memo from Upworthy. Measured, not assumed: against a 10% relative floor the 95% interval's half-width is a median 4.5× that floor and never below 1.1×. No Upworthy test can rule out a material effect in both directions, so the corpus supports "no detected difference" and never "no material difference." The null memo comes from ASOS.
- Mapping where the normal approximation breaks. Upworthy has essentially no small-n regime — the 5th-percentile package has ~2,010 impressions. The calibration study characterises the realistic operating grid, which is a weaker and more honest claim.
- Seasonality, in any ASOS memo. The archive publishes relative day offsets and no calendar time, so a 131-day run cannot be checked against seasonal or promotional context.
- "False-positive rate on real traffic." The hypergeometric A/A split conditions on exactly the null the test assumes, so it cannot fail. It is empirical-grid calibration. The 203 strict identical-stimulus pairs are the genuinely falsifiable instrument, and their 1.53 pp detection bound is quoted everywhere their rate is.
Neither corpus is vendored — both are CC-BY 4.0 but run to 80 MB, so data/ is
gitignored and provenance is pinned by digest instead. Every file below downloads
without authentication.
ASOS Digital Experiments — OSF osf.io/64jsb
(DOI 10.17605/OSF.IO/64JSB), companion to
arXiv:2111.10198. Direct:
https://osf.io/62t7f/download
| File | Bytes | SHA-256 |
|---|---|---|
data/asos/asos_experiments.parquet |
1,003,239 | bdf88b27185d3f7e65912cbb421129a32ae347c2fb0d9442fe54d74266551524 |
Upworthy Research Archive — OSF osf.io/jd64p.
| File | Bytes | SHA-256 |
|---|---|---|
data/upworthy/exploratory.csv |
14,260,949 | 8368313b060f4015a0c6fb34e6d788163cee29554144aba9390b14922eb9d8ce |
data/upworthy/confirmatory.csv |
66,517,697 | b2a88288c88f2b67d40413c9bdbebe79a1c0b5d9a9354180549f59f3bfff685f |
pip install -e '.[dev]'
pytest # full suite
python -m validation.run all # the five simulation studies (no download needed)
# with the two archives placed as above:
python -m corpora.upworthy.ingest
python -m corpora.asos.ingest
python -m validation.report # the gate
python -m validation.report --reproduce # + re-run and compare the studiesUpworthy asks that confirmatory analyses follow a peer-reviewed analysis plan. A
portfolio project cannot obtain peer review, so the substitute is in-repo
pre-registration: every filter threshold and every seed is a literal in
corpora/upworthy/ingest.sql and
corpora/asos/ingest.sql, committed before any
analysis number was computed. Filters flag rows rather than deleting them, so a
record reports what a threshold cost instead of asserting it was applied.
splitlab/ power, fixed_horizon, cuped, sequential, srm, simulate
validation/ five simulation studies that had to pass before real data
corpora/ ingest (plain SQL over DuckDB) and study code for both archives
memos/ five analyst memos and their analysis packs
results/ committed JSON records and the SVGs generated from them
tests/ reference libraries live here, never in splitlab/
Rigor apparatus — the floor, not the headline
splitlab/ has no statsmodels or pingouin dependency; scipy.stats
distributions only. Reference implementations appear in tests/, to check the
work rather than to do it. 956 tests, 100% coverage across splitlab/,
validation/ and corpora/.
Every result artifact carries the git SHA it ran under plus a dirty-tree flag, so
the SHA provably describes the code that produced it. The sequence is always
commit the code, run it, then commit the artifact — all 17 records carry
git_dirty: false. Records carry no wall-clock timestamp, which makes them
byte-reproducible.
The report gate (python -m validation.report) re-derives 60 headline
numbers in this README, in WRITEUP.md and in the memos directly from the
committed artifacts, and fails if the prose and the artifact disagree. Several
are stored nowhere as a field — the stopping-estimate ratios above are medians
over 338 series cells — so claims carry a derivation rather than a key path.
--reproduce re-runs the five simulation studies and compares what they write:
figures byte for byte, records on study/seed/config/cells, since
git_sha necessarily differs once HEAD has moved. The 12 corpus artifacts need
data/ and are reported as skipped, never as passed.
The gate is only worth having if it can fail. Mutating every claim's value showed
13 of the 50 then registered still matching somewhere in their documents — 8
occurs 296 times in WRITEUP.md — so claims now bind the phrase the prose wraps
the number in. All 60 currently fail when the value is perturbed.