Skip to content

Repository files navigation

BIN Lore

Autonomous media → knowledge pipeline: watch a Twitch broadcast, transcribe it locally, extract structured lore with an LLM, and publish a searchable wiki — unattended.

Live site: trak3r.github.io/binlore · Source domain: Barely Informed News (Case Blackwell)

Wiki homepage

Why this exists

Most “AI demos” stop at a chat transcript. BIN Lore is an end-to-end production loop: resilient VOD ingest, local speech-to-text, canon-aware LLM extraction, idempotent wiki writes, Quartz compile validation, and GitHub Pages deploy. The public artifact is a living fan wiki; the engineering is a repeatable agentic content factory.

Features

  • VOD ingest — yt-dlp audio-only download from Twitch, with automatic YouTube-archive fallback when Twitch retention expires
  • Local transcription — faster-whisper timestamped transcripts (no cloud STT required)
  • Canon-aware extraction — LLM prompts seeded with existing characters / segments / storylines so ASR name errors reconcile to wiki canon
  • Idempotent wiki updates — episode rundowns, character appearance tables, storyline beats, segment occurrence logs
  • Screencap CDN — ffmpeg frame capture hosted on a permanent GitHub Release asset CDN (repo stays binary-light)
  • Unattended batch — bare ./binlore runs refresh → transcribe → extract → wiki; disk hygiene, retries, Quartz gates, per-episode commits
  • Published site — Quartz wiki on GitHub Pages with graph view, backlinks, and full-text search; public deploys only when production advances

Pipeline

Twitch / YouTube VOD
         │
         ▼
┌─────────────────────┐
│   binlore ingest    │  audio-only download + Whisper transcript
└──────────┬──────────┘
           │
           ▼
┌─────────────────────┐
│   binlore extract   │  LLM analysis -> structured lore JSON
└──────────┬──────────┘
           │
           ▼
┌─────────────────────┐
│ binlore update-wiki │  characters · segments · storylines · episodes
└──────────┬──────────┘
           │
           ▼
┌─────────────────────┐
│   Quartz preview    │  validate build locally on main
└──────────┬──────────┘
           │
           ▼
┌─────────────────────┐
│ merge -> production │  express publish -> GitHub Pages
└─────────────────────┘

Character page

Quick start

# deps: ffmpeg, Node 22+, Python 3.11+
npm ci
cd tools && python3 -m venv .venv && source .venv/bin/activate && pip install -e . && cd ..
cp tools/.env.example tools/.env   # add OPENROUTER_API_KEY (extract only)

./binlore vods
./binlore ingest --latest
./binlore extract --latest
./binlore update-wiki --latest
npx quartz build --serve          # http://localhost:8080

Unattended backlog (370+ historical streams). Bare invocation does the whole loop:

./binlore                         # refresh → transcribe → extract → wiki
./binlore process-all --status    # backlog counts
./binlore transcribe-all          # Whisper only (no LLM)
./binlore process-all --extract   # mine existing transcripts only

Legal Disclaimer: Unofficial, non-commercial fan wiki and documentation project. Not affiliated with, endorsed by, or sponsored by Case Blackwell, Barely Informed News, or Twitch. All character names, likenesses, trademarks, and media assets belong to their respective copyright holders and are referenced under fair use (17 U.S.C. § 107) for commentary, criticism, and archival purposes. Not operated for profit. See content/disclaimer.md for full legal disclosures.


Operator guide

Detail below is for running and extending the pipeline. Skip to Setup & Prerequisites if you already know the shape of the system.

Where Downloaded VODs & Transcripts Live Locally

When you ingest a stream, all downloaded media and processed artifacts are saved in:

tools/runs/<vod-id>/

Where Does <vod-id> Come From?

  • Twitch VOD IDs: Live streams air on Twitch (caseblackwell), where Twitch assigns a numeric video ID to each broadcast (e.g. 2863722826 from https://www.twitch.tv/videos/2863722826).
  • Finding VOD IDs:
    1. Wiki Episodes List: The complete Episodes & Broadcast Archive has a dedicated VOD ID column for every stream.
    2. CLI: Run ./binlore vods to print recent Twitch streams with their IDs, broadcast dates, and lengths.
  • YouTube Archive IDs: Twitch purges VODs after ~60 days. The complete historical backlog of 370+ streams is preserved on the YouTube Archive (@CaseBlackwellStreams). For archived streams beyond Twitch's retention window, the YouTube video ID (e.g. ZSjvjEED3KA) shown on the episodes list can be passed directly to ./binlore ingest <id>.

For example, for VOD 2863722826 (High T Wednesday News):

tools/runs/2863722826/
├── audio.m4a              # Downloaded stream audio (high-quality audio-only)
├── meta.json              # VOD metadata (title, Twitch ID, duration, air date)
├── transcript.json        # Full timestamped Whisper transcript (segments array)
├── transcript.txt         # Human-readable transcript with [MM:SS] timestamps
├── transcript.plain.txt   # Plain un-timestamped transcript text
└── extraction.json        # Structured LLM output (segments, characters, lore)

Why audio instead of full video? yt-dlp pulls the Twitch Audio_Only stream directly (~150 MB instead of a 6+ GB video file). This saves disk space and allows local Whisper transcription to process significantly faster.

Git tracking: Media files (*.m4a, *.mp4, etc.) and batch logs are listed in .gitignore to prevent large binary bloat. Raw Whisper transcripts (transcript.json, transcript.txt) and stream metadata (meta.json) in tools/runs/ are tracked in the repository so episodes can be reprocessed or re-analyzed in the future without requiring another VOD scrape or audio re-transcription. The wiki pages in content/ are the public, reviewed canon.


Setup & Prerequisites

1. System Dependencies

Linux / Remote Server (Ubuntu / Debian):

# Media tools for Twitch VOD download & audio processing
sudo apt update && sudo apt install -y ffmpeg

# Node.js 22+ (required for Quartz wiki build)
curl -fsSL https://deb.nodesource.com/setup_22.x | sudo -E bash -
sudo apt install -y nodejs

macOS:

# Media tools for Twitch VOD download & audio processing
brew install yt-dlp ffmpeg

# Node.js 22+ (required for Quartz)
brew install node@22
export PATH="/opt/homebrew/opt/node@22/bin:$PATH"

2. Quartz Wiki Setup

From the repository root:

npm ci

3. Python Tools Setup (Ingest & Extraction CLI)

cd tools
python3 -m venv .venv
source .venv/bin/activate
pip install -e .

4. Environment Configuration (tools/.env)

Copy the template configuration file:

cp tools/.env.example tools/.env

OpenRouter API Key & Model Configuration

Lore extraction uses OpenRouter (not Google AI Studio). Transcription does not need a key.

Policy: one capable model only. There is no automatic fallback to cheaper/weaker models — if the primary model fails after retries, extraction stops.

  1. Get an API key at https://openrouter.ai/keys.
  2. Add your key to tools/.env:
    OPENROUTER_API_KEY="your-openrouter-api-key-here"
  3. (Optional) Pin a capable model. Default is the OpenRouter free Nemotron 3 Ultra (no purchased credits required):
    OPENROUTER_MODEL="nvidia/nemotron-3-ultra-550b-a55b:free"
    # OPENROUTER_MODEL="nvidia/nemotron-3-super-120b-a12b:free"
    # OPENROUTER_MODEL="deepseek/deepseek-v4-flash-0731:free"
    # Paid: OPENROUTER_MODEL="anthropic/claude-sonnet-4.6"

Hugging Face Token (Optional, Recommended for Cloud Servers)

When faster-whisper downloads speech-to-text models (such as small or medium), it downloads weights from Hugging Face Hub. On remote servers (AWS, Hetzner, DigitalOcean), unauthenticated requests can be aggressively rate-limited or throttled by Hugging Face (triggering Warning: You are sending unauthenticated requests to the HF Hub).

Setting a token provides:

  • Faster, prioritized downloads with higher bandwidth.
  • Higher rate limits, preventing HTTP 429 Too Many Requests errors on shared datacenter IP ranges.
  • Clean logs by suppressing unauthenticated hub warnings.

To set it up:

  1. Create a free account at huggingface.co if needed.
  2. Generate a free access token with default Read permission at https://huggingface.co/settings/tokens.
  3. Add it to tools/.env:
    HF_TOKEN="hf_your_token_here"

YouTube bot-checks (archive fallback)

Twitch VODs download without login. The YouTube archive fallback can hit a temporary bot-check on datacenter IPs (Sign in to confirm you’re not a bot).

There is no published fixed timeout. yt-dlp’s session rate-limit text says “up to an hour”; temporary bot-checks often clear in ~1–12 hours if you stop hammering. The batch sleeps and probes until clear (default: wait 1h, then probe every 30m). Override with YTDLP_BOT_COOLDOWN_INITIAL / YTDLP_BOT_COOLDOWN_PROBE in tools/.env.


Execution Guide

You can run commands in two ways:

  • Directly from repo root (recommended): Use ./binlore <command> (e.g. ./binlore extract --latest)
  • From within the virtualenv: Run source tools/.venv/bin/activate once, then use binlore <command> directly.

Step 1: List Recent VODs

See the latest streams available on Case Blackwell's Twitch channel:

./binlore vods
./binlore vods --limit 10

Step 2: Ingest a VOD (Download Audio + Transcribe)

Download the stream audio and generate a full timestamped transcript using local Whisper:

# Ingest the newest stream automatically
./binlore ingest --latest

# Or ingest with immediate audio cleanup to save disk space
./binlore ingest --latest --clean-audio

# Or ingest a specific VOD by URL or ID
./binlore ingest "https://www.twitch.tv/videos/<vod-id>"
./binlore ingest <vod-id>

# Optional: choose Whisper model size (default is 'small')
# Options: tiny, base, small, medium, large-v3
./binlore ingest --latest --model small

This creates the run folder in tools/runs/<vod-id>/ and stubs an episode page in content/episodes/YYYY-MM-DD.md.

Step 3: Extract Lore, Characters & Segments (OpenRouter)

Run the LLM extraction pipeline over the transcript:

# Preview prompt & token counts without calling the API (dry-run)
./binlore extract --latest --dry-run

# Run extraction using OpenRouter (requires OPENROUTER_API_KEY)
./binlore extract --latest

# Or specify a capable model explicitly
./binlore extract --latest --model nvidia/nemotron-3-ultra-550b-a55b:free

# Or target a specific VOD ID (see “Where Does <vod-id> Come From?” above)
./binlore extract <vod-id> --timeout 180

What this does automatically:

  1. Loads current wiki canon (content/characters/, content/segments/, content/storylines/) so the model knows established talent (like Munch, Crum, Jeb Nogget, Kendelle, Brandon/Cryptozeus) and reconciles ASR phonetic errors (e.g. "Crumb" $\to$ "Crum", "Noggin" $\to$ "Nogget").
  2. Saves tools/runs/<vod-id>/extraction.json.
  3. Populates content/episodes/YYYY-MM-DD.md with:
    • Stream overview
    • Segment rundown table (| Start | End | Segment | Notes |)
    • On-air talent detected (speaking vs mentioned) with Quartz wikilinks ([[characters/crum|Crum]])
    • Storyline developments ([[storylines/crum-dick-punch|Crum's Robotic Gorilla Groin Punch Bet]])
    • Timestamped candidate lore notes

Step 4: Propagate Lore into the Wiki (binlore update-wiki)

Once extraction is complete, populate the rest of the wiki (Character appearance tables, Notable moments, Storyline beat timelines, and Segment occurrence tables) from the extraction data:

# Preview changes without modifying files (dry-run)
./binlore update-wiki --latest --dry-run

# Update wiki pages from the latest extraction
./binlore update-wiki --latest

# Or update for a specific VOD ID
./binlore update-wiki <vod-id>

What this does automatically:

  • Character pages (content/characters/<name>.md): Appends to ## Appearances and adds timestamped quotes/facts to ## Notable moments with links back to the episode.
  • New contributors & correspondents: Automatically creates profiles for newly detected on-air personalities (e.g. hype-train.md, tommy-biglaw.md) and indexes them in content/characters/index.md.
  • Storyline pages (content/storylines/<slug>.md): Appends new beat entries to the ## Key beats timeline with exact timestamps and episode anchors.
  • Segment pages (content/segments/<slug>.md): Appends occurrences to ## Known occurrences.
  • Idempotent: Safe to run repeatedly without creating duplicate rows or notes.

Step 5: Capture & Host Screencaps (binlore screencap)

Extract sharp, lightweight video frames for characters and segments via ffmpeg without downloading the entire video, and host them on the permanent GitHub release media-assets CDN:

# Capture key frames for all detected characters & segments in latest VOD
./binlore screencap --latest

# Capture a specific frame by timestamp from a known VOD ID
./binlore screencap <vod-id> --timestamp 01:14:00 --name crum

# Capture and immediately upload to GitHub release media-assets
./binlore screencap --latest --upload

Image assets are served directly from GitHub's Fastly CDN (https://github.com/trak3r/binlore/releases/download/media-assets/<name>.jpg), keeping the git repository 100% lightweight and clean of binary files.

Step 6: Preview the Wiki Locally

Preview your wiki in your browser with live-reload:

# From repository root:
npx quartz build --serve

Open http://localhost:8080.

Step 7: Review & Publish to GitHub Pages

Work lands on main (transcripts, wiki edits). The live site only updates when you merge main into production. Do not commit directly to production. Preview locally first with Step 6 if you want a visual check.

1. Land work on main

git status
git add content/
git commit -m "Add notes for episode YYYY-MM-DD"
git push origin main

Pushing to main does not deploy the public wiki.

2. Publish — merge main → production

CLI (fast-forward when production has not diverged):

git checkout production && git pull
git merge --ff-only main
git push origin production
git checkout main

Or open a PR with base production and compare main, review the diff, then merge.

3. Confirm deploy

Pushing to production runs .github/workflows/deploy.yml. When that Actions run succeeds, https://trak3r.github.io/binlore/ reflects the new tip of production.

Step 8: Reclaim Disk Space (binlore clean)

Stream audio files take ~150 MB per 2-hour VOD. Once transcription is finished, you can safely remove the audio files to free up disk space while preserving all transcripts, metadata, and wiki content:

# Preview which files would be deleted and space saved
./binlore clean --dry-run

# Delete audio files across all completed runs
./binlore clean

# Keep the most recent stream's audio and delete older ones
./binlore clean --keep 1

# Clean a specific VOD
./binlore clean <vod-id>

Unattended Batch Processing on a Server

To process the entire 370+ episode backlog unattended on a home server, VPS, or cloud instance:

./binlore

Transcribe only (no LLM):

./binlore transcribe-all

Extract only (already-transcribed backlog):

./binlore process-all --extract

What It Does per Episode

  1. Backlog Discovery: Cross-references tools/youtube_catalog.json with tools/runs/ transcripts. Default ./binlore refreshes the catalog from Twitch, then transcribes missing VODs, then extracts lore for transcripts that lack extraction.json. ./binlore transcribe-all / --skip-extract stop after Whisper; --extract mines only.
  2. Audio Ingest & Resilient Fallback: Downloads audio using yt-dlp. If a Twitch VOD has expired (Twitch retention is ~60 days), it automatically falls back to the permanent YouTube archive stream. Skipped when a transcript already exists.
  3. Local Whisper Transcription: Transcribes audio via faster-whisper (default model: small). Does not write episode stubs in transcribe-only mode.
  4. Immediate Disk Cleanup: Deletes the audio file immediately once transcription finishes and is saved. Peak disk usage is capped to at most one temporary audio file at any moment (~150 MB).
  5. Lore Extraction (--extract only): Sends the transcript and a compact canon roster to OpenRouter (capable model only, JSON). Unknown names are queued to tools/runs/unknown-characters.jsonl instead of auto-creating pages. Storyline beats append to Timeline. On credit/quota exhaustion the batch stops — never falls back to a weaker model.
  6. Wiki Population (--extract only): Updates content/episodes/<date>.md and existing character/segment pages. Fixture cold-opens (Pepito, Case, News) do not get a row for “did the usual thing.” Appearance tables are sorted by date; first_seen is backdated.
  7. Catalog Synchronization: Regenerates content/episodes/index.md after extract.
  8. Wiki Compilation (--extract only): Quartz build is skipped during transcribe-all.
  9. Automatic Git Commit: Transcript ingest commits tools/runs/<id>/. Extract commits also include content/.
  10. Fault-Tolerant Loop: If an individual stream fails, it cleans up partial files, logs the failure, and continues. OpenRouter credit/quota exhaustion pauses the extract batch cleanly.

Disk Space & Hygiene Guarantees

Running on a server with limited disk space requires strict hygiene:

  • Zero Media Accumulation: Audio files are deleted immediately after transcription completes. The script never leaves audio files waiting for batch completion.
  • Cleanup on Error / Interrupt: If a download fails or you press Ctrl+C (SIGINT/SIGTERM), a signal handler sweeps and deletes any temporary .part, .ytdl, or incomplete .m4a files.
  • Pre-flight Disk Monitoring: Before downloading each episode, free disk space is checked against --min-disk-gb (default: 1.0 GB). If host disk space drops below this threshold, the script halts safely rather than crashing the filesystem.
  • Pre-run Sweep: Automatically cleans any orphaned media files in tools/runs/ left by previous manual runs before starting.
  • Bounded Logs: Structured, single-line logs are written to tools/runs/batch.log (gitignored), ensuring log files never grow out of control.

How to Run Unattended over SSH (screen)

# 1. SSH in, then start a named screen session
screen -S binlore

# 2. Start the batch processor (transcribe-only by default; use --extract to mine)
./binlore process-all --extract

# 3. Detach: Ctrl+a, then d
# The processor keeps running after you disconnect SSH.

# 4. Reattach later (same host):
screen -r binlore

# If it says "Attached elsewhere":
screen -dr binlore

Structured logs also land in tools/runs/batch.log (gitignored) if you want to tail progress from another shell.

Checking Status & Monitoring Progress

# Print current backlog status and exit
./binlore process-all --status

# Output:
# --- [BIN Lore Backlog Status] ---
# Total catalog streams: 376
# Ingested & Extracted:  20
# Remaining in Backlog:  356 (5.3% complete)
# Free Disk Space:       108.72 GB
# Next in queue:         2026-06-22 — Don't Kier the Reaper, it's MONDAY NEWS
# ---------------------------------

# Preview the queue of unprocessed streams without modifying files
./binlore process-all --dry-run --limit 10

# Live-tail the log file
tail -f tools/runs/batch.log

CLI Options Reference

Flag Default Description
--limit N all Process up to N episodes (useful for testing batches, e.g. --limit 5)
--oldest-first True Process backlog from oldest to newest (default)
--newest-first False Process newest unprocessed items first
--model MODEL small faster-whisper model: tiny, base, small, medium, large-v3
--extract-model nvidia/nemotron-3-ultra-550b-a55b:free OpenRouter model id (--openrouter-model is an alias)
--delay SECONDS 5.0 Cool-down sleep in seconds between episodes
--timeout SECONDS 180.0 Extraction timeout per model
--min-disk-gb GB 1.0 Minimum free disk space in GB required before ingesting
--status — Display backlog progress and disk space, then exit
--dry-run — Preview the queue without downloading or modifying files
--keep-audio False Retain audio files on disk (warning: consumes ~150 MB per episode)
--skip-extract / --transcribe-only — Whisper only (no LLM). Also: ./binlore transcribe-all
--extract / --extract-only — Mine already-transcribed episodes only
(no mode flag) full pipeline Refresh catalog → transcribe → extract → wiki
--no-clean-existing False Do not sweep tools/runs/ for old media files on startup
--no-skip-drafts False Do not skip episodes marked with draft: true
--build-quartz / --no-build extract only Quartz is skipped during transcribe-all
--git-commit / --no-git-commit True Automatically create a local git commit for each processed episode
--log-file PATH tools/runs/batch.log Destination path for the structured log file

Wiki Structure & Writing Lore

Content lives in content/:

Folder What it holds
content/characters/ On-air anchors, correspondents, contributors, and guests (e.g. Munch, Crum, Case Blackwell)
content/segments/ Recurring broadcast formats and desks (e.g. Munch & Crum, News, Hype Train)
content/storylines/ Multi-broadcast storylines and investigative sagas (e.g. Crum's Robotic Gorilla Groin Punch Bet, Beyblade Tournament)
content/episodes/ Per-broadcast episode logs, rundowns, and candidate lore notes
  • Use [[wikilinks]] between pages (e.g. [[characters/munch|Munch]]).
  • Always cite timestamps when adding lore facts.
  • Use draft: true in page frontmatter to prevent unfinished pages from publishing.

Roadmap

Voluntary Buy Me a Coffee tips help cover API and hosting costs for ongoing development — no paywall, no gated content.

Buy Me a Coffee

  • Phase 0: Repo bootstrap, Quartz setup, GitHub Pages CI/CD, seed pages
  • Phase 1: VOD listing, audio download, local Whisper transcription, runs archive
  • Phase 2: LLM segment and lore extraction via OpenRouter (capable models only)
  • Phase 3: Unattended server batch pipeline (binlore process-all) with auto-cleanup & validation
  • Automated git branch/PR generation for proposed wiki edits
  • Broadcast screencap gallery and on-air graphic asset index

License

MIT — see LICENSE.txt. Quartz is © jackyzha0; binlore tooling and customizations are © Thomas Davis. Wiki content is unofficial fan documentation for personal and non-commercial use.

About

Autonomous media→knowledge pipeline: Twitch/YouTube ingest, local Whisper, LLM lore extraction, and a live Quartz wiki.

Topics

Resources

Code of conduct

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages