Skip to content

Add WD14 tagger as an alternative captioning method - #2

Merged
Patrick16 merged 1 commit into
masterfrom
add-wd14-tagger
Sep 9, 2026
Merged

Patrick16 merged 1 commit into
masterfrom
add-wd14-tagger

Conversation

@Patrick16

Copy link
Copy Markdown
Owner

Summary

Adds a local ONNX-based WD14 tagger as a second captioning method alongside the existing vision-LLM path, selectable via a new Method dropdown shared by both per-download captioning and the standalone "Caption an existing folder" action.

Unlike the vision-LLM path, WD14 needs no server/API key -- it's a local image-classification model (the same family used by kohya_ss/taggui) that outputs a comma-separated booru-style tag list (e.g. solo, blue hair, outdoors) instead of a natural-language sentence.

What changed

  • app/wd14_tagger.py -- model loading (cached in-process; a failed load is cached too, so a bad model/no internet fails once instead of re-attempting a slow network call per image), preprocessing (pad-to-square white fill, resize to the model's own input size, RGB→BGR channel order), inference, and tag formatting (category-based threshold split for general vs. character tags, underscore→space except a kaomoji exception list). Four presets (wd-vit-tagger-v3 default / convnext / swinv2 / eva02-large), or any compatible Hugging Face repo id typed in directly.
  • app/captioning.py -- new Captioner interface (VisionLLMCaptioner / WD14Captioner) so jobs.py calls either backend the same way without branching on method itself. Trigger-word handling now lives here too: baked in as a prompt instruction for the vision-LLM path (unchanged behavior), or simply prepended as the first tag for WD14 (a flat tag list has no sentence grammar for a "role" like subject/style/action to fit into).
  • UI: a "Method" select (Vision LLM / WD14 Tagger) at the top of the captioning card, with WD14-specific fields (model, general/character thresholds) shown only when that method is active.
  • requirements.txt: adds onnxruntime, huggingface_hub, numpy.

Why WD14 as a separate path rather than just another LLM

Booru-style tag lists train more consistently than natural-language captions for anime/booru-style LoRA datasets -- this was flagged as a candidate next step in docs/landscape-and-roadmap-notes.md from an earlier competitive-landscape pass (kohya_ss/taggui both ship WD14 for exactly this reason). It's implemented as a distinct local classifier rather than another vision-LLM prompt because that's genuinely what WD14 is -- an ONNX image classifier with ~10,000 output tags and confidence scores, not a chat model.

Verified live (not just mocked)

  • The preprocessing/inference pipeline was validated against a real downloaded model and a real image before writing the final module code -- produced coherent, sensible tags on the first real try (confirmed pad/resize/channel-order choices were correct, not guessed).
  • Inference is fast once loaded: ~0.2s per image on CPU for the default vit preset.
  • A real download job (booru search + WD14 captioning) produced real, sensible tag files on disk.
  • The standalone caption-folder action with a trigger word enabled correctly prepended it as the first tag on every image (e.g. sks_creature, animal ears, solo, ...).
  • The same flow was also driven through the actual browser UI (not just the API) via a real form submit, with identical results.

Tests

27 new tests: tests/test_wd14_tagger.py (inference logic via a fake ONNX session, plus the model-load caching behavior including the cached-failure path) and tests/test_captioning.py (both Captioner implementations, trigger-word handling). Existing tests/test_jobs.py updated for the new build_captioner() plumbing -- the captioner's LLMClient now lives in app.captioning, not app.jobs, so several tests needed their monkeypatch target updated accordingly. Full suite: 193 passed.

Reviewer notes

  • First use of a given WD14 model downloads it from Hugging Face (roughly 50–800MB depending on the preset) and caches it via huggingface_hub's own cache dir -- needs internet the first time for a given model, nothing after that.
  • general_threshold/character_threshold are user-configurable in the UI; defaults (0.35 / 0.85) match the WD14 community convention.

🤖 Generated with Claude Code

Adds a local ONNX-based WD14 tagger (app/wd14_tagger.py) as a second
captioning backend alongside the existing vision-LLM path, selectable via a
new "Method" dropdown shared by both per-download captioning and the
standalone "Caption an existing folder" action.

Unlike the vision-LLM path, WD14 needs no server/API key -- it's a local
image-classification model (the same family used by kohya_ss/taggui) that
outputs a comma-separated booru-style tag list instead of a sentence.

Preprocessing/inference pipeline validated live against real images before
writing the final code (pad-to-square white fill, resize to the model's own
input size, RGB->BGR channel order, sigmoid outputs, category-based
threshold split for general vs. character tags) -- produced coherent,
sensible tags on the first real try.

- app/wd14_tagger.py: model loading (cached in-process, including caching a
  failed load so a bad model/no internet fails once instead of re-attempting
  a slow network call per image), preprocessing, inference, tag formatting.
  Four presets (vit/convnext/swinv2/eva02-large), or any compatible HF repo
  id typed in directly.
- app/captioning.py: new Captioner interface (VisionLLMCaptioner /
  WD14Captioner) so jobs.py calls either backend the same way. Trigger word
  handling now lives here too: baked in as a prompt instruction for the
  vision-LLM path (unchanged), or simply prepended as the first tag for
  WD14 (a flat tag list has no sentence grammar for a "role" to fit into).

Live-verified end to end (not just mocked): a real download job with
booru search + WD14 captioning produced real, sensible tag files; the
standalone caption-folder action with a trigger word correctly prepended it
as the first tag on every image; the same flow driven through the actual
browser UI (not just the API) produced identical results.

Adds onnxruntime + huggingface_hub + numpy to requirements.txt. 27 new
tests (wd14_tagger inference/caching logic, captioning.py's two backends);
existing jobs tests updated for the new build_captioner() plumbing (the
captioner's LLMClient now lives in app.captioning, not app.jobs).
@Patrick16
Patrick16 merged commit a39ea0e into master Sep 9, 2026
1 check passed
@Patrick16
Patrick16 deleted the add-wd14-tagger branch September 9, 2026 19:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant