Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 3 additions & 4 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -37,10 +37,9 @@ HF_TOKEN=
# LDS_UPDATE_CHECK_ENABLED=true # checks GitHub releases on launch (no data sent); set false to disable

# ComfyUI model filenames, if yours are named differently (see docs/comfyui-setup.md)
# LDS_QWEN_EDIT_MODEL=qwen_image_edit_2511_int8_convrot.safetensors
# LDS_ANGLES_LORA=qwen/Qwen-Image-Edit-2511-Multiple-Angles-LoRA.safetensors
# LDS_QWEN_TEXT_ENCODER=qwen_2.5_vl_7b_fp8_scaled.safetensors
# LDS_QWEN_VAE=qwen_image_vae.safetensors
# LDS_QWEN21_MODEL=qwen_image_2.1_int8_convrot.safetensors
# LDS_QWEN21_TEXT_ENCODER=qwen3vl_8b_int8_convrot.safetensors
# LDS_QWEN21_VAE=qwen_image_2.1_vae_bf16.safetensors
# LDS_UPSCALE_MODEL=4xNomosWebPhoto_RealPLKSR.safetensors
# LDS_DEJPG_MODEL=1xDeJPG_OmniSR.pth
# LDS_SAM3_CHECKPOINT=sam3.1_multiplex_fp16.safetensors
8 changes: 4 additions & 4 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,10 +6,10 @@ Turn a character, style, or concept into a ready-to-train LoRA dataset. One refe

## Current state

_Last verified: 2026-08-23_
_Last verified: 2026-10-03_

- **Status:** in active development, released at v0.16.0 (git tag `v0.16.0`). CI green. **Every version bump must ship a GitHub Release** or the in-app update check never fires. The project was renamed twice (lora-dataset-studio → lora-distillery → dataset-deviser); older references under the previous names are stale.
- **Works:** all five stages end to end (preprocess → generate & curate → caption → export → train config); the three dataset types (character, style, concept) with their own shot plans and caption framing; local and cloud paths for every stage; gallery curation with selection carried forward between stages, including shift-click range select; advisory dedupe, quality flags and caption lint; trainer configs for ai-toolkit (incl. SDXL), kohya sd-scripts and musubi-tuner; opt-in private Hugging Face publish. Since 0.15.0: per-image fault isolation in ① (one bad source never ends the batch), a cooperative ⏹ Stop across ①/②/③ with documented resume paths, translated ComfyUI failures (missing node, bad input, unreachable server) and an in-app 🩺 setup check. Since 0.15.1: a ② **shot style** (default = keep the reference's own medium; eight presets + custom text) threaded into both prompt builders, the CLI, ④ metadata and ⑤'s sample prompt, a final-prompt preview, and a prev/next/save-and-next caption editor with the image on screen. Since 0.16.0: **one name box and one trigger box in the header** that ②/③/④/⑤ all read (they used to be per-tab and went stale), ④ stating the pair it stamped, an opt-in ④ **hand-off to [Idiot LoRa Builder](https://github.com/Fablestarexpanse/Idiot-Lora-Builder)** (writes its `.lora-studio/ratings.json` so its grid opens pre-triaged — nothing is launched), a ⑤ advisory when the dataset can't fill the training resolution, every bundled ComfyUI model filename overridable from `.env`, and a `cli build` preflight for the local engine.
- **Status:** in active development, released at v0.16.0 (git tag `v0.16.0`); 0.17.0 (local ② generation moved to Qwen-Image 2.1) is committed but not yet tagged or released. CI green at 0.16.0. **Every version bump must ship a GitHub Release** or the in-app update check never fires. The project was renamed twice (lora-dataset-studio → lora-distillery → dataset-deviser); older references under the previous names are stale.
- **Works:** all five stages end to end (preprocess → generate & curate → caption → export → train config); the three dataset types (character, style, concept) with their own shot plans and caption framing; local and cloud paths for every stage; gallery curation with selection carried forward between stages, including shift-click range select; advisory dedupe, quality flags and caption lint; trainer configs for ai-toolkit (incl. SDXL), kohya sd-scripts and musubi-tuner; opt-in private Hugging Face publish. Since 0.15.0: per-image fault isolation in ① (one bad source never ends the batch), a cooperative ⏹ Stop across ①/②/③ with documented resume paths, translated ComfyUI failures (missing node, bad input, unreachable server) and an in-app 🩺 setup check. Since 0.15.1: a ② **shot style** (default = keep the reference's own medium; eight presets + custom text) threaded into both prompt builders, the CLI, ④ metadata and ⑤'s sample prompt, a final-prompt preview, and a prev/next/save-and-next caption editor with the image on screen. Since 0.16.0: **one name box and one trigger box in the header** that ②/③/④/⑤ all read (they used to be per-tab and went stale), ④ stating the pair it stamped, an opt-in ④ **hand-off to [Idiot LoRa Builder](https://github.com/Fablestarexpanse/Idiot-Lora-Builder)** (writes its `.lora-studio/ratings.json` so its grid opens pre-triaged — nothing is launched), a ⑤ advisory when the dataset can't fill the training resolution, every bundled ComfyUI model filename overridable from `.env`, and a `cli build` preflight for the local engine. Since 0.17.0: the local ② engine is **Qwen-Image 2.1** (plain-English prompts shared with Gemini, ~26 s/shot, `LDS_QWEN21_*` filenames) instead of Qwen-Image-Edit 2511 + the Multiple-Angles LoRA, whose graph stays in `studio/comfy_workflows/legacy/`; a dead ComfyUI mid-batch now fails one shot, not the run.
- **In progress:** nothing half-built — 0.16.0 closed out the hand-off, cross-tab-identity and config-hardening work. `docs/ARCHITECTURE.md` carries the "Future ideas" and "Deferred" sections that hold the real backlog.
- **Known gaps / next steps:** pick up from `docs/ARCHITECTURE.md` → "Future ideas" / "Deferred", and read its *Maintainer principles* before changing anything; `gradio` is pinned `<6` because of a real stuck-loading-overlay regression — unpinning needs that verified fixed upstream; training is never launched for you and is not planned to be; the default export `output_root` writes into the repo root, so a smoke-test export leaves an untracked dataset folder to delete before committing.
- **Deep docs:** `docs/ARCHITECTURE.md` (module map, stage and backend detail, gotchas, backlog — the deep reference), `docs/comfyui-setup.md`. Worklogs under `docs/` are gitignored and local-only by design.
Expand Down Expand Up @@ -73,7 +73,7 @@ python cli.py keys --setup
- ⏹ Stop is cooperative and checked *between* items; every stage calls `JOB.start()` first, and the Stop button must stay `queue=False`.
- Bundled ComfyUI workflows are core-nodes-only, so an unknown node means an out-of-date ComfyUI — `comfy_api` preflights node classes and translates rejection bodies.
- `tests/conftest.py` autouse fixture repoints the output roots at tmp; without it the suite writes real folders into `runs/`.
- ② prompts never hard-code a medium — it comes from `studio/shot_style.py` (default: keep the reference's). Never write "photorealistic"/"hyperrealistic"; a test bans them from every generated prompt. Angle shots stay pure `<sks>` grammar.
- ② prompts never hard-code a medium — it comes from `studio/shot_style.py` (default: keep the reference's). Never write "photorealistic"/"hyperrealistic"; a test bans them from every generated prompt. The local prompt never carries a negation ("without", "do not", a named prop): Qwen-Image 2.1 draws what it is told not to; a test bans it.
- Re-value a Gradio dropdown with `gr.update(value=…)`, not `gr.Dropdown(value=…)` — the constructor form re-validates against the ORIGINAL `choices` and can silently drop the value.
- The dataset **name and trigger are single header components** (`project_name`/`project_trigger`); every stage is wired straight to them. Never reintroduce a per-tab copy, and never mirror components per keystroke — Gradio's unqueued `.input` replies land out of order and settle on a prefix of what was typed.
- Every model filename in a bundled ComfyUI template must be reachable from `comfy_api._MODEL_INPUTS` (plus a setting in `config.py` and a line in `.env.example`) — a test enforces it. Un-mapped means the user can't fix it in `.env` and `doctor` can't check it.
Expand Down
5 changes: 2 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,7 +45,7 @@ Run them in order (each step auto-fills the next) or jump straight to the one yo
|---|---|---|
| ① Restore / upscale | ComfyUI models, or basic Lanczos | — |
| ① Subject isolation | **Built-in SAM3** (no ComfyUI) or ComfyUI SAM3 | — |
| ② Generate shots *(character + concept)* | ComfyUI: Qwen Image Edit 2511 + Multiple-Angles LoRA | Gemini (Nano Banana) |
| ② Generate shots *(character + concept)* | ComfyUI: Qwen-Image 2.1 *(Qwen Research License)* | Gemini (Nano Banana) |
| ③ Caption | Qwen3-VL-8B, JoyCaption, NSFW finetune, **WD + e621 taggers**, LM Studio / Ollama / any OpenAI endpoint | Gemini Flash, Groq free tier |
| ④ Export | always local (+ optional **.zip** and **Hugging Face** publish) | — |
| ⑤ Train config | ai-toolkit (incl. SDXL) / **kohya sd-scripts** / musubi-tuner | — |
Expand Down Expand Up @@ -213,8 +213,7 @@ comic, digital illustration, traditional painting, 3D render, ink line art, or *
where you describe the medium yourself (*"a 1970s screen-printed poster, halftone dots"*).

Changing it rebuilds the shot plan, so hand-edited prompt cells are replaced — edit prompts
after you've settled on a style. Camera-angle shots deliberately keep their bare
turnaround grammar, which is what the angles LoRA was trained on. Use **👁 Preview final
after you've settled on a style. Use **👁 Preview final
prompt** to see exactly what a row will send, including outfit and prop-exclusion. The style
is recorded in the exported `metadata.json` and picked up by ⑤'s sample prompt.

Expand Down
27 changes: 19 additions & 8 deletions app.py
Original file line number Diff line number Diff line change
Expand Up @@ -54,7 +54,7 @@

ENGINE_CHOICES = [
("Cloud — Gemini image model (best identity fidelity, SFW only)", "gemini"),
("Local — ComfyUI Qwen Image Edit 2511 (free, private, uncensored)", "comfyui"),
("Local — ComfyUI Qwen Image 2.1 (free, private, uncensored)", "comfyui"),
]
CLOUD_MODEL_CHOICES = [(f"{m} (~${p:.3f}/img est.)", m) for m, p in CLOUD_IMAGE_PRICES.items()]
CAPTIONER_CHOICES = [(c.label, c.key) for c in CAPTIONERS]
Expand Down Expand Up @@ -303,7 +303,7 @@ def preview_final_prompt(plan_df: pd.DataFrame, engine: str, exclude_props: bool
if exclude_props:
shot = apply_prop_exclusion(shot)
field = "local_prompt" if engine == "comfyui" else "cloud_prompt"
which = "Local (ComfyUI / Qwen-Edit)" if engine == "comfyui" else "Cloud (Gemini)"
which = "Local (ComfyUI / Qwen-Image 2.1)" if engine == "comfyui" else "Cloud (Gemini)"
return (f"**{which} prompt for `{shot.id}`** — exactly what the engine receives, "
f"after outfit and prop-exclusion are folded in:\n\n```\n"
f"{getattr(shot, field)}\n```")
Expand Down Expand Up @@ -1343,7 +1343,15 @@ def do_load_plan(plan_name: str):
raise gr.Error(f"Couldn't read the plan at {path}: {e}. Check the YAML — "
f"each shot needs at least an `id`, `kind`, `local_prompt` "
f"and `cloud_prompt`.") from e
return _shots_to_df(shots), f"✅ Loaded {len(shots)} shots from {path}"
note = f"✅ Loaded {len(shots)} shots from {path}"
# Plans saved before 0.17.0 carry the old engine's `<sks>` LoRA grammar, which
# Qwen-Image 2.1 reads as literal text and renders badly.
stale = sum("<sks>" in s.local_prompt for s in shots)
if stale:
note += (f" ⚠️ {stale} shot(s) still use the old Qwen-Image-Edit `<sks>` prompts, "
f"which the local Qwen-Image 2.1 engine doesn't understand — rebuild the "
f"plan from the dataset type / shot style, or rewrite those local prompts.")
return _shots_to_df(shots), note


def estimate_cost(engine: str, cloud_model: str, df: pd.DataFrame) -> str:
Expand Down Expand Up @@ -1795,15 +1803,18 @@ def _check_for_update():
gen_exclude_props = gr.Checkbox(
value=True,
label="Exclude props/accessories from the reference",
info="Asks the generator to drop bags, held objects and "
"accessories carried in your reference, so they don't get "
"baked into every dataset image. Isolating the source in ① "
"is the more reliable fix. Character-oriented wording — off "
info="Cloud engine only: asks Gemini to drop bags, held objects "
"and accessories carried in your reference, so they don't "
"get baked into every dataset image. The local engine "
"ignores it (naming a prop, even to forbid it, makes "
"Qwen draw it) — isolate the source in ① instead, the more "
"reliable fix either way. Character-oriented wording — off "
"by default for Concept datasets.")
gen_isolate = gr.Checkbox(value=False,
label="Isolate generated angle shots (white background)",
info="Cut generated angle shots onto white too "
"(helps the angles LoRA on back views).")
"(replaces each angle shot's setting "
"with a plain white background).")
gen_iso_backend = gr.Dropdown(ISOLATION_CHOICES,
value=settings.isolation_backend,
label="Isolation backend",
Expand Down
Loading
Loading