Scrapers don't crash. They decay.
A site renames one class, extraction keeps returning 200 OK, and the pipeline
stays green while the data underneath quietly rots.
THWIP is the thing that notices, shows you exactly what was lost, repairs the collector in place, and writes down the proof.
Built for Into the Scrape-Verse, WeMakeDevs × Bright Data, 17–23 Aug 2026. The hackathon is over; the cron has been stopped and the record is frozen where it stood.
Every scraper is a Spider with an Integrity score, 0–100: how much of what it promised came back real. When Integrity drops, a web spins over its panel covering exactly the share that was lost — and you drag the threads away, one whole strand at a time, to read the values that actually arrived.
Nothing below needs an API key, touches the network, or spends a credit.
| # | Do this | It proves |
|---|---|---|
| 1 | Open the live console | Three real collectors, five recorded incidents, the data they returned |
| 2 | Drag the web off the broken half of the diptych near the top | The reveal is the same data the field chips read, not a caption |
| 3 | Open ?mock=1, press BREAK BODEGA, then RE-WEAVE |
The break→heal→receipt loop, in your browser, in ten seconds |
| 4 | git clone this repo, then npm test |
1,457 tests, node:test, zero dependencies, offline |
| 5 | node tools/evidence-report.js inc_004 |
The full trail of the incident no human touched, with SHA-256 digests |
| 6 | git log --author="thwip watch" --oneline | wc -l |
123 commits authored by the workflow, not by a person |
Three collectors, each created with bdata scraper create against a site the Bright Data
pre-built library does not cover. No Scrapers Library entry is used anywhere in this
project.
| Codename | Target | Collector ID | Fields under contract | Why this site |
|---|---|---|---|---|
| ATLAS | books.toscrape.com | c_mt5mpgolbgqeg9ziv |
title, price, rating, image_url, availability | A sandbox built for scraping; no robots.txt, server-rendered. Rating lives in a CSS class, not in text — the natural infected-field candidate |
| KESTREL | news.ycombinator.com | c_mt5mpiem10zvo4kkwj |
title, points, comments, author | A real site with real churn. robots.txt allows the front page; Crawl-delay: 30 is respected |
| BODEGA | our own demo page | c_mt5mo9mkz9u25zefn |
title, price, rating, image | Breakable on purpose, so a judge can reproduce a break without waiting for the web to move |
Every ID is pinned in collectors.json with its per-field validators.
The creation envelopes Bright Data returned are committed verbatim in
evidence/.
All three targets are public. No login wall, no paywall, no personal data, no
government site. robots.txt was read for each before a single request; the reasoning
per target is in the table above.
One real row, exactly as ATLAS returned it and as it is committed in
data/history.json:
{
"title": "Libertarianism for Beginners",
"price": { "value": 51.33, "currency": "GBP", "symbol": "£" },
"rating": "Two",
"image_url": "https://books.toscrape.com/media/cache/91/a4/91a46253e165d144ef5938f2d456b88f.jpg",
"availability": "In stock"
}That shape is the contract. A price must parse as a number, an image_url must be an
absolute URL, a rating must fall in range — declared per field in collectors.json and
checked on every scan. A field that arrives populated but wrong is scored INFECTED at
half credit, which is the failure mode a null check cannot see and the reason Integrity
exists at all.
The console shows those rows as themselves under THE HAUL, each stamped with the collector that fetched it, when it was scanned, and the Integrity that Spider was at when the row was captured. Provenance travels with the data.
A Spider that lost everything. The web covers exactly the share of the contract that did not come back — drag the threads away and the nulls underneath are readable.
Per scan, per field, from what actually came back:
| State | Meaning | Credit |
|---|---|---|
🟢 LIVE |
arrived and passed its validator | full |
🟡 INFECTED |
arrived and is wrong — out of range, wrong type, a literal "undefined" |
half |
🔴 DEAD |
null, empty, or missing | none |
Integrity is the share of the promised contract that came back real, so a scraper returning thirty rows of confident zeros scores 0 and not 100. This is the whole argument: a pipeline can be green while its data rots, and only a per-field score against a declared contract can see it.
A worked example lives on the live page: Bright Data's own dashboard reports ATLAS at a 6.67% success rate while our console reports 100% Integrity, 20 rows of 20. Both are right. ATLAS follows each product link, so the platform counts fourteen failed child fetches per catalogue page — but every contracted field is already on the catalogue page, so every row comes back complete. A platform success rate measures how many requests completed; Integrity measures how much of what you promised came back real.
Judges were told to look for bdata scraper heal. Here is every time it ran, on the
record, with the Collector ID identical on both sides:
| Incident | Spider | Strain | Integrity | Site is ours? | Opened by | Verdict |
|---|---|---|---|---|---|---|
inc_001 |
KESTREL | RENAMED |
0 → 100 | no — news.ycombinator.com | hand | healed |
inc_002 |
ATLAS | DRIFTED |
90 → 100 | no — books.toscrape.com | hand | healed |
inc_003 |
BODEGA | THROTTLED |
0 → 100 | yes | the cron, unattended | wrong diagnosis, kept |
inc_004 |
BODEGA | RENAMED |
50 → 100 | yes | the cron, and no phase was human | healed |
inc_005 |
KESTREL | THROTTLED |
0 → 50 | no — news.ycombinator.com | the cron, unattended | PARTIAL — still open |
Three of the five happened on sites we do not control. Most self-healing demos break a fixture they wrote themselves; more than half of our evidence is somebody else's HTML changing under us.
On 22 Aug we committed a redesign of our own demo page — the class names moved, the way a
real redesign moves them — and then touched nothing. The cron saw Integrity fall to 50%,
waited for a second consecutive bad scan rather than reacting to one, opened the
incident, diagnosed the strain as RENAMED, re-wove the collector, and verified against a
fresh scrape. 11m 56s from detection to verification, price: null → £18.00 and
rating: null → 4.4, on an unchanged Collector ID. Nobody ran a command, approved a
repair, or edited a record.
node tools/evidence-report.js inc_004That prints the collector id on both sides (asserted identical), the four stage timestamps with computed durations, the value every field held before and after, the verdict, and SHA-256 digests recomputed from the committed files at call time — a check, not a claim.
Overnight on 23 Aug, KESTREL lost all four fields. The cron re-wove it and got title and
author back. points and comments did not come back. Verification scored it
PARTIAL, 2 of 4, and — because the fresh run never cleared the healthy threshold —
left the incident open. It is still open in data/incidents.json.
That is the system working. A repair that reports success while half the contract is missing is exactly the lie this project exists to catch, and it does not get an exemption for being our own.
inc_003 was opened autonomously by the cron overnight, and the heal it fired fixed
nothing — because nothing on the target had broken. The scraper was returning one
wrapped row holding a products array, and our own payload parser was scoring the envelope
instead of the rows. The bug was ours.
The wrong diagnosis is still in data/incidents.json, written exactly as the system
believed it at the time. A monitoring tool that quietly rewrites its own history to look
smarter is precisely the thing this project exists to catch.
- verification is computed from a fresh scrape run after the heal, never from the
report the heal writes about itself —
scripts/verify.js; scripts/repair.jssetsresolvedonly when that fresh run clears the healthy threshold, so a heal that returns nothing leaves the incident open — seeinc_005;- values that come back populated but wrong are caught by the per-field validators and
score
INFECTED, which keeps Integrity below the threshold and the incident open.
Four cases are pinned as tests you can read as prose:
node --test test/pipeline/heal-that-lies.test.js — all nulls, populated garbage,
partial recovery, real recovery.
The proof a re-weave writes: every broken field re-checked against a fresh scrape after the heal, with the value before and after, on an unchanged collector id.
The cron in .github/workflows/watch.yml ran every eight
hours from 21 Aug to 18 Sep 2026 — scanning all three collectors, scoring every field,
opening and healing incidents on its own, committing the results, and republishing the
site. Nothing was manual.
It has now been switched off. The workflow still builds and deploys the console on push,
but the scan job and its schedule are gone, so no further API calls are made and no
further data is committed. What you see in data/ is the final state.
scans 390 incidents 5
rows 14,483 heals 4
collectors 3 unchangedIds true
bot commits 123 span 21 Aug → 18 Sep 2026
git log --author="thwip watch" --oneline | wc -l # 123 commits no human made
node tools/numbers-audit.js # every number the console showstools/numbers-audit.js is deliberately a second implementation: it recomputes every
figure from the committed JSON and shares no code with the console, so a disagreement
between them would show rather than hide.
Eight tools over stdio JSON-RPC, written straight against the MCP spec: no SDK, no build step, no dependencies. Six read the committed record and are instant, free, and never touch the network. Two spend Bright Data credit and say so in their own descriptions, so a well-behaved agent asks before it bills you.
/manual.html —
the console's back page. The test count on it is read from data/meta.json at page load.
claude mcp add thwip -- node mcp/server.jsThen ask the fleet the questions you would ask a colleague: is anything broken? what broke? fix it. prove it.
| Free — reads the committed record | Spends Bright Data credit |
|---|---|
fleet_status · spider_history · incident_log |
scan_fleet — scrapes the fleet for real |
heal_receipt · evidence_report · numbers_audit |
heal_spider — diagnoses, re-weaves, verifies |
→ mcp/README.md has the protocol notes, every tool schema, and a
verbatim transcript of the receipt for the incident no human touched.
git clone https://github.com/GoatWhistle/scrape-verse-hack.git
cd scrape-verse-hack
npm test # 1,457 tests, zero dependencies, offline
python3 -m http.server 8000 # any static serverThen open http://localhost:8000/web/ — the console reads two committed JSON files and
needs no backend. Add ?mock=1 for the CHAOS LAB, where the fleet is synthetic and says
so, but the break, the spread and the receipt are the same code the live page runs.
?mock=1 — the CHAOS LAB. Press BREAK BODEGA and the web spins over the panel;
RE-WEAVE heals it and prints the receipt shown further up.
To scan for real you would need a Bright Data account and bdata login; npm run health
and npm run repair are the two entry points the cron itself used. Everything above
this line works without a key.
scripts/ the product: scan, score, diagnose, heal, verify
health-check.js one scan of every collector, scored field by field
repair.js opens an incident after two consecutive bad scans, heals, closes
verify.js scores a fresh run after the heal — never the report the heal writes
web/ the console — no build step, no framework, no dependencies
js/{data,fleet,sheets,mock}/ modules named for the part of the product they build
css/{base,fleet,sheets,fx,print,mock}/
manual.html the back page: the pitch, the tools, the judge path
mcp/ MCP server — eight tools, stdio JSON-RPC, no SDK
tools/ evidence-report.js and numbers-audit.js — proof, computed twice
test/ 1,457 tests in {pipeline,web,mcp,tools}/, mirroring the source tree
data/ history.json and incidents.json — the frozen record
evidence/ raw Bright Data payloads; the digests are recomputed from these
collectors.json targets, Collector IDs, per-field validators
demo-target/ the shop page BODEGA watches — three variants of the same page
No dependencies and no devDependencies: npm install installs nothing, and CI asserts
that it stays that way. Every JS and CSS file is ≤250 lines, enforced as an ESLint
error and re-checked in CI, tests included.
I used Claude Code (Anthropic) throughout, which is also how Bright Data intends
Scraper Studio to be driven: the whole bdata workflow runs inside a coding agent, in
the terminal, with no dashboard.
It helped me most with ideas and with parts of the code — sketching approaches, drafting the pipeline scripts, writing large stretches of the test suite, and arguing back when a design was weak. It did not build the project on its own.
The console front-end I built by hand: the comic layout, the web that buries a broken panel, the tear-away reveal, the character rig, the diptych and the whole visual system came from me rather than from a prompt. So did the decisions that make the rest of it worth anything — which sites I chose to target and why, the Integrity model and its three field states, my rule that a repair waits for two consecutive bad scans, my rule that verification runs a fresh scrape instead of trusting the report a heal writes about itself, and my choice to leave a wrong diagnosis on the record rather than tidy it away.
Stated plainly rather than left for a judge to find:
- Two of five incidents were invoked by hand (
inc_001,inc_002).inc_004closes the gap end to end, unattended, but two of five is two of five. - One heal only half worked (
inc_005) and one diagnosis was flatly wrong (inc_003). Both are still on the record, unedited. - MTTR is a mean of four samples. Enough to display honestly, not enough to be a trend.
REWEAVINGis a state the console can render and nothing writes —repair.jsran to completion inside one CI job, so no mid-heal record is ever persisted. The branch is reachable only from mock data.- The scratch has no keyboard affordance, by design: it is a pointer-only reveal over data that is also readable as text chips on the same panel. Nothing is behind it that is not elsewhere.
- Frame rate was measured on a contended machine. The console holds ~82fps in three consecutive runs with all animation live, but that number deserves a quiet browser before anyone quotes it.
- Public data only. No login wall, no paywall, no personal data, no government site.
- Never recreate a collector that has an ID. If a run fails,
healit — a needless recreate costs credit and breaks everything downstream that trusted the ID. - Never commit a key.
.env,credentials.jsonandconfig.jsonare gitignored, and no key appears in this repository, in the console, or in the demo video. - Never edit a record to look better. The wrong diagnosis in
inc_003stays. So does the half-finished heal ininc_005.
MIT licensed · built in seven days for Into the Scrape-Verse




