Repository navigation
Add WD14 tagger as an alternative captioning method - #2
Merged
Merged
Conversation
Adds a local ONNX-based WD14 tagger (app/wd14_tagger.py) as a second captioning backend alongside the existing vision-LLM path, selectable via a new "Method" dropdown shared by both per-download captioning and the standalone "Caption an existing folder" action. Unlike the vision-LLM path, WD14 needs no server/API key -- it's a local image-classification model (the same family used by kohya_ss/taggui) that outputs a comma-separated booru-style tag list instead of a sentence. Preprocessing/inference pipeline validated live against real images before writing the final code (pad-to-square white fill, resize to the model's own input size, RGB->BGR channel order, sigmoid outputs, category-based threshold split for general vs. character tags) -- produced coherent, sensible tags on the first real try. - app/wd14_tagger.py: model loading (cached in-process, including caching a failed load so a bad model/no internet fails once instead of re-attempting a slow network call per image), preprocessing, inference, tag formatting. Four presets (vit/convnext/swinv2/eva02-large), or any compatible HF repo id typed in directly. - app/captioning.py: new Captioner interface (VisionLLMCaptioner / WD14Captioner) so jobs.py calls either backend the same way. Trigger word handling now lives here too: baked in as a prompt instruction for the vision-LLM path (unchanged), or simply prepended as the first tag for WD14 (a flat tag list has no sentence grammar for a "role" to fit into). Live-verified end to end (not just mocked): a real download job with booru search + WD14 captioning produced real, sensible tag files; the standalone caption-folder action with a trigger word correctly prepended it as the first tag on every image; the same flow driven through the actual browser UI (not just the API) produced identical results. Adds onnxruntime + huggingface_hub + numpy to requirements.txt. 27 new tests (wd14_tagger inference/caching logic, captioning.py's two backends); existing jobs tests updated for the new build_captioner() plumbing (the captioner's LLMClient now lives in app.captioning, not app.jobs).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds a local ONNX-based WD14 tagger as a second captioning method alongside the existing vision-LLM path, selectable via a new Method dropdown shared by both per-download captioning and the standalone "Caption an existing folder" action.
Unlike the vision-LLM path, WD14 needs no server/API key -- it's a local image-classification model (the same family used by kohya_ss/taggui) that outputs a comma-separated booru-style tag list (e.g.
solo, blue hair, outdoors) instead of a natural-language sentence.What changed
app/wd14_tagger.py-- model loading (cached in-process; a failed load is cached too, so a bad model/no internet fails once instead of re-attempting a slow network call per image), preprocessing (pad-to-square white fill, resize to the model's own input size, RGB→BGR channel order), inference, and tag formatting (category-based threshold split for general vs. character tags, underscore→space except a kaomoji exception list). Four presets (wd-vit-tagger-v3default / convnext / swinv2 / eva02-large), or any compatible Hugging Face repo id typed in directly.app/captioning.py-- newCaptionerinterface (VisionLLMCaptioner/WD14Captioner) sojobs.pycalls either backend the same way without branching on method itself. Trigger-word handling now lives here too: baked in as a prompt instruction for the vision-LLM path (unchanged behavior), or simply prepended as the first tag for WD14 (a flat tag list has no sentence grammar for a "role" like subject/style/action to fit into).requirements.txt: addsonnxruntime,huggingface_hub,numpy.Why WD14 as a separate path rather than just another LLM
Booru-style tag lists train more consistently than natural-language captions for anime/booru-style LoRA datasets -- this was flagged as a candidate next step in
docs/landscape-and-roadmap-notes.mdfrom an earlier competitive-landscape pass (kohya_ss/taggui both ship WD14 for exactly this reason). It's implemented as a distinct local classifier rather than another vision-LLM prompt because that's genuinely what WD14 is -- an ONNX image classifier with ~10,000 output tags and confidence scores, not a chat model.Verified live (not just mocked)
vitpreset.sks_creature, animal ears, solo, ...).Tests
27 new tests:
tests/test_wd14_tagger.py(inference logic via a fake ONNX session, plus the model-load caching behavior including the cached-failure path) andtests/test_captioning.py(bothCaptionerimplementations, trigger-word handling). Existingtests/test_jobs.pyupdated for the newbuild_captioner()plumbing -- the captioner'sLLMClientnow lives inapp.captioning, notapp.jobs, so several tests needed their monkeypatch target updated accordingly. Full suite: 193 passed.Reviewer notes
huggingface_hub's own cache dir -- needs internet the first time for a given model, nothing after that.general_threshold/character_thresholdare user-configurable in the UI; defaults (0.35 / 0.85) match the WD14 community convention.🤖 Generated with Claude Code