Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
53 changes: 44 additions & 9 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,9 @@
[![Tests](https://github.com/Patrick16/DatasetForge/actions/workflows/tests.yml/badge.svg)](https://github.com/Patrick16/DatasetForge/actions/workflows/tests.yml)

A local web tool: given a list of text queries, it downloads images from the web
(search via DuckDuckGo, no API key needed) into a folder you choose. Optionally:
(search via DuckDuckGo by default, no API key needed -- Yandex, Google, and
booru boards are also available, see **Search providers** below) into a folder
you choose. Optionally:

- **Query LLM** — generates N additional search-query variations for each input
query (Ollama / LM Studio / any OpenAI-compatible cloud API).
Expand Down Expand Up @@ -48,12 +50,14 @@ network calls or a live model server -- no external services required.
default alongside jpg/png/webp (only `bmp` is off by default).
- **SafeSearch control** — a dropdown (Off / Moderate / On) next to the format
filters. This isn't a filter *we* apply -- it's passed straight through to
DuckDuckGo's own search, so results it excludes never reach this app at all.
Previously there was no way to change it and it was silently always on
(`ddgs`'s own default, "moderate"), which made the tool unable to fetch
the search source itself (DuckDuckGo/Yandex/Google) or turned into a
`rating:` tag (booru boards), so results it excludes never reach this app
at all. Previously there was no way to change it and it was silently always
on (`ddgs`'s own default, "moderate"), which made the tool unable to fetch
results DuckDuckGo itself considers not-safe-for-work no matter what you
searched for or how your local filters were set. Confirmed live that "off"
vs "on" actually returns a different result set for a borderline query.
searched for or how your local filters were set. Confirmed live for both
DuckDuckGo and Yandex that "off" vs "on" actually returns a different
result set for a borderline query.
- **"Filtered" vs "Errors" are counted separately.** Images your own
format/min-size filters excluded show up under "Filtered" (expected,
working as configured); network failures, bad data, etc. show up under
Expand Down Expand Up @@ -128,6 +132,36 @@ network calls or a live model server -- no external services required.
unload` itself exits 0 with a "Model Not Found" message in that case, which
used to be passed straight through and read exactly like a failure.

## Search providers

Chosen via the "Search source" card. All four confirmed working live during
development (2026-09) except where noted:

| Provider | What it is | SafeSearch mechanism | Notes |
|---|---|---|---|
| **DuckDuckGo** (default) | No API key. Under the hood this actually proxies **Bing's** image index (confirmed by inspecting the `ddgs` library's own provider tagging) -- there is no independent DuckDuckGo image index. | `p` query param | The more reliable of the two ways to reach that same Bing data -- Bing's own direct scraping route in `ddgs` ignores the safesearch argument entirely. |
| **Yandex** | Unofficial scraping of `yandex.com/images`. Historically laxer filtering than Google/Bing. | `family` cookie (`0`=off, `2`=on, unset=default) | Confirmed live: toggling it changed 5/25 results on a borderline query. Width/height aren't extracted from the scrape -- harmless, since every file is re-measured with Pillow after downloading anyway. |
| **Google** (experimental) | Unofficial scraping. | `safe` query param | ⚠️ Confirmed unreliable in testing: Google returned an HTTP 429 bot-check on the *first* request, with full browser-like headers and no prior history. Expect frequent zero results; a block degrades to "0 found" plus a log warning, not a crash. Try Yandex or DuckDuckGo instead if this keeps failing for you. |
| **Booru board** | Danbooru-API-family boards, for tag-driven / anime-style datasets. | An explicit `rating:` tag, not a hidden toggle -- content is opt-in tagged (general/sensitive/questionable/explicit; naming varies by board) rather than filtered by a black-box heuristic. | See below -- access requirements vary a lot by board. |

### Booru board access (confirmed live, 2026-09)

- **e621** — works with zero configuration.
- **Gelbooru** / **Rule34.xxx** — now require an `api_key` + `user_id` (free
account -> Account -> API Access Credentials); both returned HTTP
401/"Missing authentication" without one. This is a change from their
historical open access.
- **Danbooru** — supports `login` + `api_key` (HTTP Basic Auth, the
documented method), but requests were blocked by a Cloudflare bot-check
("Just a moment...") in testing regardless, from at least some networks.
Implemented, but best-effort -- it may just not work for you. Also:
anonymous Danbooru API access is limited to 2 combined tags, which a
multi-word query plus the rating tag can easily exceed.

To add another search source, implement `ImageSearchProvider.search()` under
`app/search/` and register it in `build_search_provider()` in
`app/search/__init__.py`.

## Optional LLM providers

Both blocks (query expansion and captioning) are configured independently right
Expand Down Expand Up @@ -192,16 +226,17 @@ app/
model_registry.py # tracks which (provider, base_url, model) this process has used
model_control.py # unloads a model from Ollama (native API) or LM Studio (`lms` CLI)
search/
__init__.py # build_search_provider() factory (picks provider by SearchConfig)
base.py # ImageSearchProvider interface
duckduckgo.py # DuckDuckGo search (no API key)
yandex.py # Yandex Images scraping
google.py # Google Images scraping (experimental, unreliable)
booru.py # Danbooru/e621/Gelbooru/Rule34
static/
index.html, style.css, app.js # web UI, no build step
tests/ # pytest suite for app/ (see "Tests" above)
```

To add a new search source, implement `ImageSearchProvider.search()` under
`app/search/` and wire it up in `app/jobs.py`.

## Known limitations (MVP)

- The folder picker and model discovery only work when the browser and the
Expand Down
4 changes: 2 additions & 2 deletions app/jobs.py
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@
from .download import download_image, sanitize_folder_name
from .llm_client import LLMClient, build_trigger_instruction
from .models import CaptionFolderRequest, JobCreateRequest
from .search.duckduckgo import DuckDuckGoProvider
from .search import build_search_provider

logger = logging.getLogger(__name__)

Expand Down Expand Up @@ -85,7 +85,7 @@ async def run_job(self, state: JobState) -> None:
expander = LLMClient(req.llm_expansion.llm) if req.llm_expansion.enabled else None
captioner = LLMClient(req.captioning.llm) if req.captioning.enabled else None
try:
search_provider = DuckDuckGoProvider()
search_provider = build_search_provider(req.search)
allowed_formats = {f.lower().lstrip(".") for f in req.filters.formats}
out_root = Path(req.output_folder)

Expand Down
23 changes: 23 additions & 0 deletions app/models.py
Original file line number Diff line number Diff line change
Expand Up @@ -49,6 +49,28 @@ class FilterConfig(BaseModel):
safesearch: Literal["on", "moderate", "off"] = "moderate"


class SearchConfig(BaseModel):
"""Which image search backend to use, and booru-specific extras.

- duckduckgo: default, no config needed (proxies Bing's image index).
- yandex: unofficial scraping, historically laxer filtering than DDG/Google.
- google: unofficial scraping -- confirmed unreliable in testing (Google
blocks plain HTTP scraping aggressively); expect frequent zero results.
- booru: Danbooru-API-family boards (e621/gelbooru/rule34/danbooru).
Content is explicitly rating-tagged rather than hidden behind a
SafeSearch toggle, so `filters.safesearch` maps onto a rating: tag
instead of a provider-side filter flag.
"""

provider: Literal["duckduckgo", "yandex", "google", "booru"] = "duckduckgo"
booru_site: Literal["e621", "gelbooru", "rule34", "danbooru"] = "e621"
# gelbooru/rule34 use an api_key + user_id pair; danbooru uses login + api_key
# (as HTTP Basic Auth). e621 needs none of these.
booru_api_key: Optional[str] = None
booru_user_id: Optional[str] = None
booru_login: Optional[str] = None


class JobCreateRequest(BaseModel):
queries: list[str]
n_per_query: int = Field(default=10, ge=1)
Expand All @@ -57,6 +79,7 @@ class JobCreateRequest(BaseModel):
llm_expansion: LLMExpansionConfig = Field(default_factory=LLMExpansionConfig)
captioning: CaptioningConfig = Field(default_factory=CaptioningConfig)
filters: FilterConfig = Field(default_factory=FilterConfig)
search: SearchConfig = Field(default_factory=SearchConfig)
concurrency: int = Field(default=8, ge=1)


Expand Down
26 changes: 26 additions & 0 deletions app/search/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,26 @@
from __future__ import annotations

from ..models import SearchConfig
from .base import ImageSearchProvider
from .booru import BooruProvider
from .duckduckgo import DuckDuckGoProvider
from .google import GoogleProvider
from .yandex import YandexProvider


def build_search_provider(config: SearchConfig) -> ImageSearchProvider:
"""Construct the right ImageSearchProvider for a job's SearchConfig."""
if config.provider == "duckduckgo":
return DuckDuckGoProvider()
if config.provider == "yandex":
return YandexProvider()
if config.provider == "google":
return GoogleProvider()
if config.provider == "booru":
return BooruProvider(
site=config.booru_site,
api_key=config.booru_api_key,
user_id=config.booru_user_id,
login=config.booru_login,
)
raise ValueError(f"Unknown search provider: {config.provider!r}")
159 changes: 159 additions & 0 deletions app/search/booru.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,159 @@
from __future__ import annotations

import logging

import httpx

from .base import ImageResult

logger = logging.getLogger(__name__)

_UA = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0 Safari/537.36"

# Booru sites don't hide a SafeSearch toggle behind an opaque server-side
# filter -- content is explicitly tagged with a rating (naming varies by
# site) that you query directly, so "off" here genuinely means "no rating
# restriction" rather than "hope the filter didn't trigger".
_GELBOORU_STYLE_RATING = {"on": "rating:general", "moderate": "-rating:explicit", "off": ""}
_DANBOORU_RATING = {"on": "rating:g", "moderate": "-rating:e", "off": ""}
_E621_RATING = {"on": "rating:s", "moderate": "-rating:e", "off": ""}

_GELBOORU_STYLE_SITES = {
"gelbooru": "https://gelbooru.com/index.php",
"rule34": "https://api.rule34.xxx/index.php",
}


class BooruProvider:
"""Image search against a booru-style board (the Danbooru-API family).

Confirmed live during development (2026-09):
- e621: works with zero configuration, no account needed.
- gelbooru / rule34: now require an api_key + user_id (free account,
Account -> API Access Credentials) -- both returned HTTP 401/"Missing
authentication" without one, a change from their historical open access.
- danbooru: blocked by a Cloudflare bot-check ("Just a moment...") even
with a real browser User-Agent, regardless of credentials, from at
least some networks. Implemented (with login+api_key Basic Auth, the
documented method) but best-effort -- it may simply not work for you.
"""

def __init__(
self,
site: str = "e621",
api_key: str | None = None,
user_id: str | None = None,
login: str | None = None,
):
self.site = site
self.api_key = api_key
self.user_id = user_id
self.login = login

async def search(self, query: str, n: int, safesearch: str = "moderate") -> list[ImageResult]:
try:
if self.site == "e621":
return await self._search_e621(query, n, safesearch)
if self.site in _GELBOORU_STYLE_SITES:
return await self._search_gelbooru_style(query, n, safesearch)
if self.site == "danbooru":
return await self._search_danbooru(query, n, safesearch)
logger.error("Unknown booru site: %r", self.site)
return []
except Exception:
logger.exception("Booru search failed for query %r on %s", query, self.site)
return []

async def _search_e621(self, query: str, n: int, safesearch: str) -> list[ImageResult]:
rating = _E621_RATING.get(safesearch, "")
tags = f"{query} {rating}".strip()
params = {"tags": tags, "limit": min(max(n, 1), 320)}
# e621 asks API clients to identify themselves in the User-Agent.
headers = {"User-Agent": "DatasetForge/1.0 (image dataset tool)"}
async with httpx.AsyncClient(timeout=15.0, headers=headers) as client:
resp = await client.get("https://e621.net/posts.json", params=params)
resp.raise_for_status()
data = resp.json()

results = []
for post in data.get("posts", []):
file_info = post.get("file") or {}
url = file_info.get("url")
if not url:
continue # deleted/restricted posts have a null file url
general_tags = (post.get("tags") or {}).get("general", [])
results.append(
ImageResult(
url=url,
title=", ".join(general_tags[:8]) or None,
width=file_info.get("width"),
height=file_info.get("height"),
source_page=f"https://e621.net/posts/{post.get('id')}",
)
)
return results

async def _search_gelbooru_style(self, query: str, n: int, safesearch: str) -> list[ImageResult]:
base_url = _GELBOORU_STYLE_SITES[self.site]
rating = _GELBOORU_STYLE_RATING.get(safesearch, "")
tags = f"{query} {rating}".strip()
params: dict[str, str | int] = {
"page": "dapi", "s": "post", "q": "index", "json": "1",
"limit": min(max(n, 1), 100), "tags": tags,
}
if self.api_key and self.user_id:
params["api_key"] = self.api_key
params["user_id"] = self.user_id
headers = {"User-Agent": _UA}
async with httpx.AsyncClient(timeout=15.0, headers=headers) as client:
resp = await client.get(base_url, params=params)
resp.raise_for_status()
data = resp.json()

posts = data.get("post", []) if isinstance(data, dict) else (data or [])
results = []
for post in posts:
url = post.get("file_url")
if not url:
continue
results.append(
ImageResult(
url=url,
title=post.get("tags"),
width=post.get("width"),
height=post.get("height"),
source_page=f"{base_url}?page=post&s=view&id={post.get('id')}",
)
)
return results

async def _search_danbooru(self, query: str, n: int, safesearch: str) -> list[ImageResult]:
rating = _DANBOORU_RATING.get(safesearch, "")
# Anonymous Danbooru API access is limited to 2 combined tags -- a
# multi-word query plus the rating tag can exceed that even with
# credentials on a free account; let the API's own error surface
# rather than silently truncating what the user typed.
tags = f"{query} {rating}".strip()
params = {"tags": tags, "limit": min(max(n, 1), 200)}
auth = (self.login, self.api_key) if self.login and self.api_key else None
headers = {"User-Agent": _UA}
async with httpx.AsyncClient(timeout=15.0, headers=headers) as client:
resp = await client.get("https://danbooru.donmai.us/posts.json", params=params, auth=auth)
resp.raise_for_status()
data = resp.json()

results = []
for post in data:
url = post.get("file_url")
if not url:
continue
results.append(
ImageResult(
url=url,
title=post.get("tag_string_general"),
width=post.get("image_width"),
height=post.get("image_height"),
source_page=f"https://danbooru.donmai.us/posts/{post.get('id')}",
)
)
return results
Loading
Loading