Autonomous media → knowledge pipeline: watch a Twitch broadcast, transcribe it locally, extract structured lore with an LLM, and publish a searchable wiki — unattended.
Live site: trak3r.github.io/binlore · Source domain: Barely Informed News (Case Blackwell)
Most “AI demos” stop at a chat transcript. BIN Lore is an end-to-end production loop: resilient VOD ingest, local speech-to-text, canon-aware LLM extraction, idempotent wiki writes, Quartz compile validation, and GitHub Pages deploy. The public artifact is a living fan wiki; the engineering is a repeatable agentic content factory.
- VOD ingest —
yt-dlpaudio-only download from Twitch, with automatic YouTube-archive fallback when Twitch retention expires - Local transcription —
faster-whispertimestamped transcripts (no cloud STT required) - Canon-aware extraction — LLM prompts seeded with existing characters / segments / storylines so ASR name errors reconcile to wiki canon
- Idempotent wiki updates — episode rundowns, character appearance tables, storyline beats, segment occurrence logs
- Screencap CDN — ffmpeg frame capture hosted on a permanent GitHub Release asset CDN (repo stays binary-light)
- Unattended batch — bare
./binloreruns refresh → transcribe → extract → wiki; disk hygiene, retries, Quartz gates, per-episode commits - Published site — Quartz wiki on GitHub Pages with graph view, backlinks, and full-text search; public deploys only when
productionadvances
Twitch / YouTube VOD
│
▼
┌─────────────────────┐
│ binlore ingest │ audio-only download + Whisper transcript
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ binlore extract │ LLM analysis -> structured lore JSON
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ binlore update-wiki │ characters · segments · storylines · episodes
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ Quartz preview │ validate build locally on main
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ merge -> production │ express publish -> GitHub Pages
└─────────────────────┘
# deps: ffmpeg, Node 22+, Python 3.11+
npm ci
cd tools && python3 -m venv .venv && source .venv/bin/activate && pip install -e . && cd ..
cp tools/.env.example tools/.env # add OPENROUTER_API_KEY (extract only)
./binlore vods
./binlore ingest --latest
./binlore extract --latest
./binlore update-wiki --latest
npx quartz build --serve # http://localhost:8080Unattended backlog (370+ historical streams). Bare invocation does the whole loop:
./binlore # refresh → transcribe → extract → wiki
./binlore process-all --status # backlog counts
./binlore transcribe-all # Whisper only (no LLM)
./binlore process-all --extract # mine existing transcripts onlyLegal Disclaimer: Unofficial, non-commercial fan wiki and documentation project. Not affiliated with, endorsed by, or sponsored by Case Blackwell, Barely Informed News, or Twitch. All character names, likenesses, trademarks, and media assets belong to their respective copyright holders and are referenced under fair use (17 U.S.C. § 107) for commentary, criticism, and archival purposes. Not operated for profit. See
content/disclaimer.mdfor full legal disclosures.
Detail below is for running and extending the pipeline. Skip to Setup & Prerequisites if you already know the shape of the system.
When you ingest a stream, all downloaded media and processed artifacts are saved in:
tools/runs/<vod-id>/
- Twitch VOD IDs: Live streams air on Twitch (
caseblackwell), where Twitch assigns a numeric video ID to each broadcast (e.g.2863722826fromhttps://www.twitch.tv/videos/2863722826). - Finding VOD IDs:
- Wiki Episodes List: The complete Episodes & Broadcast Archive has a dedicated VOD ID column for every stream.
- CLI: Run
./binlore vodsto print recent Twitch streams with their IDs, broadcast dates, and lengths.
- YouTube Archive IDs: Twitch purges VODs after ~60 days. The complete historical backlog of 370+ streams is preserved on the YouTube Archive (
@CaseBlackwellStreams). For archived streams beyond Twitch's retention window, the YouTube video ID (e.g.ZSjvjEED3KA) shown on the episodes list can be passed directly to./binlore ingest <id>.
For example, for VOD 2863722826 (High T Wednesday News):
tools/runs/2863722826/
├── audio.m4a # Downloaded stream audio (high-quality audio-only)
├── meta.json # VOD metadata (title, Twitch ID, duration, air date)
├── transcript.json # Full timestamped Whisper transcript (segments array)
├── transcript.txt # Human-readable transcript with [MM:SS] timestamps
├── transcript.plain.txt # Plain un-timestamped transcript text
└── extraction.json # Structured LLM output (segments, characters, lore)
Why audio instead of full video?
yt-dlp pulls the Twitch Audio_Only stream directly (~150 MB instead of a 6+ GB video file). This saves disk space and allows local Whisper transcription to process significantly faster.
Git tracking:
Media files (*.m4a, *.mp4, etc.) and batch logs are listed in .gitignore to prevent large binary bloat. Raw Whisper transcripts (transcript.json, transcript.txt) and stream metadata (meta.json) in tools/runs/ are tracked in the repository so episodes can be reprocessed or re-analyzed in the future without requiring another VOD scrape or audio re-transcription. The wiki pages in content/ are the public, reviewed canon.
Linux / Remote Server (Ubuntu / Debian):
# Media tools for Twitch VOD download & audio processing
sudo apt update && sudo apt install -y ffmpeg
# Node.js 22+ (required for Quartz wiki build)
curl -fsSL https://deb.nodesource.com/setup_22.x | sudo -E bash -
sudo apt install -y nodejsmacOS:
# Media tools for Twitch VOD download & audio processing
brew install yt-dlp ffmpeg
# Node.js 22+ (required for Quartz)
brew install node@22
export PATH="/opt/homebrew/opt/node@22/bin:$PATH"From the repository root:
npm cicd tools
python3 -m venv .venv
source .venv/bin/activate
pip install -e .Copy the template configuration file:
cp tools/.env.example tools/.envLore extraction uses OpenRouter (not Google AI Studio). Transcription does not need a key.
Policy: one capable model only. There is no automatic fallback to cheaper/weaker models — if the primary model fails after retries, extraction stops.
- Get an API key at https://openrouter.ai/keys.
- Add your key to
tools/.env:OPENROUTER_API_KEY="your-openrouter-api-key-here" - (Optional) Pin a capable model. Default is the OpenRouter free Nemotron 3 Ultra (no purchased credits required):
OPENROUTER_MODEL="nvidia/nemotron-3-ultra-550b-a55b:free" # OPENROUTER_MODEL="nvidia/nemotron-3-super-120b-a12b:free" # OPENROUTER_MODEL="deepseek/deepseek-v4-flash-0731:free" # Paid: OPENROUTER_MODEL="anthropic/claude-sonnet-4.6"
When faster-whisper downloads speech-to-text models (such as small or medium), it downloads weights from Hugging Face Hub. On remote servers (AWS, Hetzner, DigitalOcean), unauthenticated requests can be aggressively rate-limited or throttled by Hugging Face (triggering Warning: You are sending unauthenticated requests to the HF Hub).
Setting a token provides:
- Faster, prioritized downloads with higher bandwidth.
- Higher rate limits, preventing HTTP
429 Too Many Requestserrors on shared datacenter IP ranges. - Clean logs by suppressing unauthenticated hub warnings.
To set it up:
- Create a free account at huggingface.co if needed.
- Generate a free access token with default Read permission at https://huggingface.co/settings/tokens.
- Add it to
tools/.env:HF_TOKEN="hf_your_token_here"
Twitch VODs download without login. The YouTube archive fallback can hit a temporary bot-check on datacenter IPs (Sign in to confirm you’re not a bot).
There is no published fixed timeout. yt-dlp’s session rate-limit text says “up to an hour”; temporary bot-checks often clear in ~1–12 hours if you stop hammering. The batch sleeps and probes until clear (default: wait 1h, then probe every 30m). Override with YTDLP_BOT_COOLDOWN_INITIAL / YTDLP_BOT_COOLDOWN_PROBE in tools/.env.
You can run commands in two ways:
- Directly from repo root (recommended): Use
./binlore <command>(e.g../binlore extract --latest) - From within the virtualenv: Run
source tools/.venv/bin/activateonce, then usebinlore <command>directly.
See the latest streams available on Case Blackwell's Twitch channel:
./binlore vods
./binlore vods --limit 10Download the stream audio and generate a full timestamped transcript using local Whisper:
# Ingest the newest stream automatically
./binlore ingest --latest
# Or ingest with immediate audio cleanup to save disk space
./binlore ingest --latest --clean-audio
# Or ingest a specific VOD by URL or ID
./binlore ingest "https://www.twitch.tv/videos/<vod-id>"
./binlore ingest <vod-id>
# Optional: choose Whisper model size (default is 'small')
# Options: tiny, base, small, medium, large-v3
./binlore ingest --latest --model smallThis creates the run folder in tools/runs/<vod-id>/ and stubs an episode page in content/episodes/YYYY-MM-DD.md.
Run the LLM extraction pipeline over the transcript:
# Preview prompt & token counts without calling the API (dry-run)
./binlore extract --latest --dry-run
# Run extraction using OpenRouter (requires OPENROUTER_API_KEY)
./binlore extract --latest
# Or specify a capable model explicitly
./binlore extract --latest --model nvidia/nemotron-3-ultra-550b-a55b:free
# Or target a specific VOD ID (see “Where Does <vod-id> Come From?” above)
./binlore extract <vod-id> --timeout 180What this does automatically:
- Loads current wiki canon (
content/characters/,content/segments/,content/storylines/) so the model knows established talent (like Munch, Crum, Jeb Nogget, Kendelle, Brandon/Cryptozeus) and reconciles ASR phonetic errors (e.g. "Crumb"$\to$ "Crum", "Noggin"$\to$ "Nogget"). - Saves
tools/runs/<vod-id>/extraction.json. - Populates
content/episodes/YYYY-MM-DD.mdwith:- Stream overview
- Segment rundown table (
| Start | End | Segment | Notes |) - On-air talent detected (speaking vs mentioned) with Quartz wikilinks (
[[characters/crum|Crum]]) - Storyline developments (
[[storylines/crum-dick-punch|Crum's Robotic Gorilla Groin Punch Bet]]) - Timestamped candidate lore notes
Once extraction is complete, populate the rest of the wiki (Character appearance tables, Notable moments, Storyline beat timelines, and Segment occurrence tables) from the extraction data:
# Preview changes without modifying files (dry-run)
./binlore update-wiki --latest --dry-run
# Update wiki pages from the latest extraction
./binlore update-wiki --latest
# Or update for a specific VOD ID
./binlore update-wiki <vod-id>What this does automatically:
- Character pages (
content/characters/<name>.md): Appends to## Appearancesand adds timestamped quotes/facts to## Notable momentswith links back to the episode. - New contributors & correspondents: Automatically creates profiles for newly detected on-air personalities (e.g.
hype-train.md,tommy-biglaw.md) and indexes them incontent/characters/index.md. - Storyline pages (
content/storylines/<slug>.md): Appends new beat entries to the## Key beatstimeline with exact timestamps and episode anchors. - Segment pages (
content/segments/<slug>.md): Appends occurrences to## Known occurrences. - Idempotent: Safe to run repeatedly without creating duplicate rows or notes.
Extract sharp, lightweight video frames for characters and segments via ffmpeg without downloading the entire video, and host them on the permanent GitHub release media-assets CDN:
# Capture key frames for all detected characters & segments in latest VOD
./binlore screencap --latest
# Capture a specific frame by timestamp from a known VOD ID
./binlore screencap <vod-id> --timestamp 01:14:00 --name crum
# Capture and immediately upload to GitHub release media-assets
./binlore screencap --latest --uploadImage assets are served directly from GitHub's Fastly CDN (https://github.com/trak3r/binlore/releases/download/media-assets/<name>.jpg), keeping the git repository 100% lightweight and clean of binary files.
Preview your wiki in your browser with live-reload:
# From repository root:
npx quartz build --serveOpen http://localhost:8080.
Work lands on main (transcripts, wiki edits). The live site only updates when you merge main into production. Do not commit directly to production. Preview locally first with Step 6 if you want a visual check.
1. Land work on main
git status
git add content/
git commit -m "Add notes for episode YYYY-MM-DD"
git push origin mainPushing to main does not deploy the public wiki.
2. Publish — merge main → production
CLI (fast-forward when production has not diverged):
git checkout production && git pull
git merge --ff-only main
git push origin production
git checkout mainOr open a PR with base production and compare main, review the diff, then merge.
3. Confirm deploy
Pushing to production runs .github/workflows/deploy.yml. When that Actions run succeeds, https://trak3r.github.io/binlore/ reflects the new tip of production.
Stream audio files take ~150 MB per 2-hour VOD. Once transcription is finished, you can safely remove the audio files to free up disk space while preserving all transcripts, metadata, and wiki content:
# Preview which files would be deleted and space saved
./binlore clean --dry-run
# Delete audio files across all completed runs
./binlore clean
# Keep the most recent stream's audio and delete older ones
./binlore clean --keep 1
# Clean a specific VOD
./binlore clean <vod-id>To process the entire 370+ episode backlog unattended on a home server, VPS, or cloud instance:
./binloreTranscribe only (no LLM):
./binlore transcribe-allExtract only (already-transcribed backlog):
./binlore process-all --extract- Backlog Discovery: Cross-references
tools/youtube_catalog.jsonwithtools/runs/transcripts. Default./binlorerefreshes the catalog from Twitch, then transcribes missing VODs, then extracts lore for transcripts that lackextraction.json../binlore transcribe-all/--skip-extractstop after Whisper;--extractmines only. - Audio Ingest & Resilient Fallback: Downloads audio using
yt-dlp. If a Twitch VOD has expired (Twitch retention is ~60 days), it automatically falls back to the permanent YouTube archive stream. Skipped when a transcript already exists. - Local Whisper Transcription: Transcribes audio via
faster-whisper(default model:small). Does not write episode stubs in transcribe-only mode. - Immediate Disk Cleanup: Deletes the audio file immediately once transcription finishes and is saved. Peak disk usage is capped to at most one temporary audio file at any moment (~150 MB).
- Lore Extraction (
--extractonly): Sends the transcript and a compact canon roster to OpenRouter (capable model only, JSON). Unknown names are queued totools/runs/unknown-characters.jsonlinstead of auto-creating pages. Storyline beats append to Timeline. On credit/quota exhaustion the batch stops — never falls back to a weaker model. - Wiki Population (
--extractonly): Updatescontent/episodes/<date>.mdand existing character/segment pages. Fixture cold-opens (Pepito, Case, News) do not get a row for “did the usual thing.” Appearance tables are sorted by date;first_seenis backdated. - Catalog Synchronization: Regenerates
content/episodes/index.mdafter extract. - Wiki Compilation (
--extractonly): Quartz build is skipped during transcribe-all. - Automatic Git Commit: Transcript ingest commits
tools/runs/<id>/. Extract commits also includecontent/. - Fault-Tolerant Loop: If an individual stream fails, it cleans up partial files, logs the failure, and continues. OpenRouter credit/quota exhaustion pauses the extract batch cleanly.
Running on a server with limited disk space requires strict hygiene:
- Zero Media Accumulation: Audio files are deleted immediately after transcription completes. The script never leaves audio files waiting for batch completion.
- Cleanup on Error / Interrupt: If a download fails or you press
Ctrl+C(SIGINT/SIGTERM), a signal handler sweeps and deletes any temporary.part,.ytdl, or incomplete.m4afiles. - Pre-flight Disk Monitoring: Before downloading each episode, free disk space is checked against
--min-disk-gb(default:1.0GB). If host disk space drops below this threshold, the script halts safely rather than crashing the filesystem. - Pre-run Sweep: Automatically cleans any orphaned media files in
tools/runs/left by previous manual runs before starting. - Bounded Logs: Structured, single-line logs are written to
tools/runs/batch.log(gitignored), ensuring log files never grow out of control.
# 1. SSH in, then start a named screen session
screen -S binlore
# 2. Start the batch processor (transcribe-only by default; use --extract to mine)
./binlore process-all --extract
# 3. Detach: Ctrl+a, then d
# The processor keeps running after you disconnect SSH.
# 4. Reattach later (same host):
screen -r binlore
# If it says "Attached elsewhere":
screen -dr binloreStructured logs also land in tools/runs/batch.log (gitignored) if you want to tail progress from another shell.
# Print current backlog status and exit
./binlore process-all --status
# Output:
# --- [BIN Lore Backlog Status] ---
# Total catalog streams: 376
# Ingested & Extracted: 20
# Remaining in Backlog: 356 (5.3% complete)
# Free Disk Space: 108.72 GB
# Next in queue: 2026-06-22 — Don't Kier the Reaper, it's MONDAY NEWS
# ---------------------------------
# Preview the queue of unprocessed streams without modifying files
./binlore process-all --dry-run --limit 10
# Live-tail the log file
tail -f tools/runs/batch.log| Flag | Default | Description |
|---|---|---|
--limit N |
all | Process up to N episodes (useful for testing batches, e.g. --limit 5) |
--oldest-first |
True |
Process backlog from oldest to newest (default) |
--newest-first |
False |
Process newest unprocessed items first |
--model MODEL |
small |
faster-whisper model: tiny, base, small, medium, large-v3 |
--extract-model |
nvidia/nemotron-3-ultra-550b-a55b:free |
OpenRouter model id (--openrouter-model is an alias) |
--delay SECONDS |
5.0 |
Cool-down sleep in seconds between episodes |
--timeout SECONDS |
180.0 |
Extraction timeout per model |
--min-disk-gb GB |
1.0 |
Minimum free disk space in GB required before ingesting |
--status |
— | Display backlog progress and disk space, then exit |
--dry-run |
— | Preview the queue without downloading or modifying files |
--keep-audio |
False |
Retain audio files on disk (warning: consumes ~150 MB per episode) |
--skip-extract / --transcribe-only |
— | Whisper only (no LLM). Also: ./binlore transcribe-all |
--extract / --extract-only |
— | Mine already-transcribed episodes only |
| (no mode flag) | full pipeline | Refresh catalog → transcribe → extract → wiki |
--no-clean-existing |
False |
Do not sweep tools/runs/ for old media files on startup |
--no-skip-drafts |
False |
Do not skip episodes marked with draft: true |
--build-quartz / --no-build |
extract only | Quartz is skipped during transcribe-all |
--git-commit / --no-git-commit |
True |
Automatically create a local git commit for each processed episode |
--log-file PATH |
tools/runs/batch.log |
Destination path for the structured log file |
Content lives in content/:
| Folder | What it holds |
|---|---|
content/characters/ |
On-air anchors, correspondents, contributors, and guests (e.g. Munch, Crum, Case Blackwell) |
content/segments/ |
Recurring broadcast formats and desks (e.g. Munch & Crum, News, Hype Train) |
content/storylines/ |
Multi-broadcast storylines and investigative sagas (e.g. Crum's Robotic Gorilla Groin Punch Bet, Beyblade Tournament) |
content/episodes/ |
Per-broadcast episode logs, rundowns, and candidate lore notes |
- Use
[[wikilinks]]between pages (e.g.[[characters/munch|Munch]]). - Always cite timestamps when adding lore facts.
- Use
draft: truein page frontmatter to prevent unfinished pages from publishing.
Voluntary Buy Me a Coffee tips help cover API and hosting costs for ongoing development — no paywall, no gated content.
- Phase 0: Repo bootstrap, Quartz setup, GitHub Pages CI/CD, seed pages
- Phase 1: VOD listing, audio download, local Whisper transcription, runs archive
- Phase 2: LLM segment and lore extraction via OpenRouter (capable models only)
- Phase 3: Unattended server batch pipeline (
binlore process-all) with auto-cleanup & validation - Automated git branch/PR generation for proposed wiki edits
- Broadcast screencap gallery and on-air graphic asset index
MIT — see LICENSE.txt. Quartz is © jackyzha0; binlore tooling and customizations are © Thomas Davis. Wiki content is unofficial fan documentation for personal and non-commercial use.


