diff --git a/README.md b/README.md index 10fb34b..fbd53f0 100644 --- a/README.md +++ b/README.md @@ -3,7 +3,9 @@ [![Tests](https://github.com/Patrick16/DatasetForge/actions/workflows/tests.yml/badge.svg)](https://github.com/Patrick16/DatasetForge/actions/workflows/tests.yml) A local web tool: given a list of text queries, it downloads images from the web -(search via DuckDuckGo, no API key needed) into a folder you choose. Optionally: +(search via DuckDuckGo by default, no API key needed -- Yandex, Google, and +booru boards are also available, see **Search providers** below) into a folder +you choose. Optionally: - **Query LLM** — generates N additional search-query variations for each input query (Ollama / LM Studio / any OpenAI-compatible cloud API). @@ -48,12 +50,14 @@ network calls or a live model server -- no external services required. default alongside jpg/png/webp (only `bmp` is off by default). - **SafeSearch control** — a dropdown (Off / Moderate / On) next to the format filters. This isn't a filter *we* apply -- it's passed straight through to - DuckDuckGo's own search, so results it excludes never reach this app at all. - Previously there was no way to change it and it was silently always on - (`ddgs`'s own default, "moderate"), which made the tool unable to fetch + the search source itself (DuckDuckGo/Yandex/Google) or turned into a + `rating:` tag (booru boards), so results it excludes never reach this app + at all. Previously there was no way to change it and it was silently always + on (`ddgs`'s own default, "moderate"), which made the tool unable to fetch results DuckDuckGo itself considers not-safe-for-work no matter what you - searched for or how your local filters were set. Confirmed live that "off" - vs "on" actually returns a different result set for a borderline query. + searched for or how your local filters were set. Confirmed live for both + DuckDuckGo and Yandex that "off" vs "on" actually returns a different + result set for a borderline query. - **"Filtered" vs "Errors" are counted separately.** Images your own format/min-size filters excluded show up under "Filtered" (expected, working as configured); network failures, bad data, etc. show up under @@ -128,6 +132,36 @@ network calls or a live model server -- no external services required. unload` itself exits 0 with a "Model Not Found" message in that case, which used to be passed straight through and read exactly like a failure. +## Search providers + +Chosen via the "Search source" card. All four confirmed working live during +development (2026-09) except where noted: + +| Provider | What it is | SafeSearch mechanism | Notes | +|---|---|---|---| +| **DuckDuckGo** (default) | No API key. Under the hood this actually proxies **Bing's** image index (confirmed by inspecting the `ddgs` library's own provider tagging) -- there is no independent DuckDuckGo image index. | `p` query param | The more reliable of the two ways to reach that same Bing data -- Bing's own direct scraping route in `ddgs` ignores the safesearch argument entirely. | +| **Yandex** | Unofficial scraping of `yandex.com/images`. Historically laxer filtering than Google/Bing. | `family` cookie (`0`=off, `2`=on, unset=default) | Confirmed live: toggling it changed 5/25 results on a borderline query. Width/height aren't extracted from the scrape -- harmless, since every file is re-measured with Pillow after downloading anyway. | +| **Google** (experimental) | Unofficial scraping. | `safe` query param | ⚠️ Confirmed unreliable in testing: Google returned an HTTP 429 bot-check on the *first* request, with full browser-like headers and no prior history. Expect frequent zero results; a block degrades to "0 found" plus a log warning, not a crash. Try Yandex or DuckDuckGo instead if this keeps failing for you. | +| **Booru board** | Danbooru-API-family boards, for tag-driven / anime-style datasets. | An explicit `rating:` tag, not a hidden toggle -- content is opt-in tagged (general/sensitive/questionable/explicit; naming varies by board) rather than filtered by a black-box heuristic. | See below -- access requirements vary a lot by board. | + +### Booru board access (confirmed live, 2026-09) + +- **e621** — works with zero configuration. +- **Gelbooru** / **Rule34.xxx** — now require an `api_key` + `user_id` (free + account -> Account -> API Access Credentials); both returned HTTP + 401/"Missing authentication" without one. This is a change from their + historical open access. +- **Danbooru** — supports `login` + `api_key` (HTTP Basic Auth, the + documented method), but requests were blocked by a Cloudflare bot-check + ("Just a moment...") in testing regardless, from at least some networks. + Implemented, but best-effort -- it may just not work for you. Also: + anonymous Danbooru API access is limited to 2 combined tags, which a + multi-word query plus the rating tag can easily exceed. + +To add another search source, implement `ImageSearchProvider.search()` under +`app/search/` and register it in `build_search_provider()` in +`app/search/__init__.py`. + ## Optional LLM providers Both blocks (query expansion and captioning) are configured independently right @@ -192,16 +226,17 @@ app/ model_registry.py # tracks which (provider, base_url, model) this process has used model_control.py # unloads a model from Ollama (native API) or LM Studio (`lms` CLI) search/ + __init__.py # build_search_provider() factory (picks provider by SearchConfig) base.py # ImageSearchProvider interface duckduckgo.py # DuckDuckGo search (no API key) + yandex.py # Yandex Images scraping + google.py # Google Images scraping (experimental, unreliable) + booru.py # Danbooru/e621/Gelbooru/Rule34 static/ index.html, style.css, app.js # web UI, no build step tests/ # pytest suite for app/ (see "Tests" above) ``` -To add a new search source, implement `ImageSearchProvider.search()` under -`app/search/` and wire it up in `app/jobs.py`. - ## Known limitations (MVP) - The folder picker and model discovery only work when the browser and the diff --git a/app/jobs.py b/app/jobs.py index 7639b25..5032427 100644 --- a/app/jobs.py +++ b/app/jobs.py @@ -12,7 +12,7 @@ from .download import download_image, sanitize_folder_name from .llm_client import LLMClient, build_trigger_instruction from .models import CaptionFolderRequest, JobCreateRequest -from .search.duckduckgo import DuckDuckGoProvider +from .search import build_search_provider logger = logging.getLogger(__name__) @@ -85,7 +85,7 @@ async def run_job(self, state: JobState) -> None: expander = LLMClient(req.llm_expansion.llm) if req.llm_expansion.enabled else None captioner = LLMClient(req.captioning.llm) if req.captioning.enabled else None try: - search_provider = DuckDuckGoProvider() + search_provider = build_search_provider(req.search) allowed_formats = {f.lower().lstrip(".") for f in req.filters.formats} out_root = Path(req.output_folder) diff --git a/app/models.py b/app/models.py index fdd4ac1..3df5421 100644 --- a/app/models.py +++ b/app/models.py @@ -49,6 +49,28 @@ class FilterConfig(BaseModel): safesearch: Literal["on", "moderate", "off"] = "moderate" +class SearchConfig(BaseModel): + """Which image search backend to use, and booru-specific extras. + + - duckduckgo: default, no config needed (proxies Bing's image index). + - yandex: unofficial scraping, historically laxer filtering than DDG/Google. + - google: unofficial scraping -- confirmed unreliable in testing (Google + blocks plain HTTP scraping aggressively); expect frequent zero results. + - booru: Danbooru-API-family boards (e621/gelbooru/rule34/danbooru). + Content is explicitly rating-tagged rather than hidden behind a + SafeSearch toggle, so `filters.safesearch` maps onto a rating: tag + instead of a provider-side filter flag. + """ + + provider: Literal["duckduckgo", "yandex", "google", "booru"] = "duckduckgo" + booru_site: Literal["e621", "gelbooru", "rule34", "danbooru"] = "e621" + # gelbooru/rule34 use an api_key + user_id pair; danbooru uses login + api_key + # (as HTTP Basic Auth). e621 needs none of these. + booru_api_key: Optional[str] = None + booru_user_id: Optional[str] = None + booru_login: Optional[str] = None + + class JobCreateRequest(BaseModel): queries: list[str] n_per_query: int = Field(default=10, ge=1) @@ -57,6 +79,7 @@ class JobCreateRequest(BaseModel): llm_expansion: LLMExpansionConfig = Field(default_factory=LLMExpansionConfig) captioning: CaptioningConfig = Field(default_factory=CaptioningConfig) filters: FilterConfig = Field(default_factory=FilterConfig) + search: SearchConfig = Field(default_factory=SearchConfig) concurrency: int = Field(default=8, ge=1) diff --git a/app/search/__init__.py b/app/search/__init__.py index e69de29..720a942 100644 --- a/app/search/__init__.py +++ b/app/search/__init__.py @@ -0,0 +1,26 @@ +from __future__ import annotations + +from ..models import SearchConfig +from .base import ImageSearchProvider +from .booru import BooruProvider +from .duckduckgo import DuckDuckGoProvider +from .google import GoogleProvider +from .yandex import YandexProvider + + +def build_search_provider(config: SearchConfig) -> ImageSearchProvider: + """Construct the right ImageSearchProvider for a job's SearchConfig.""" + if config.provider == "duckduckgo": + return DuckDuckGoProvider() + if config.provider == "yandex": + return YandexProvider() + if config.provider == "google": + return GoogleProvider() + if config.provider == "booru": + return BooruProvider( + site=config.booru_site, + api_key=config.booru_api_key, + user_id=config.booru_user_id, + login=config.booru_login, + ) + raise ValueError(f"Unknown search provider: {config.provider!r}") diff --git a/app/search/booru.py b/app/search/booru.py new file mode 100644 index 0000000..b8b250e --- /dev/null +++ b/app/search/booru.py @@ -0,0 +1,159 @@ +from __future__ import annotations + +import logging + +import httpx + +from .base import ImageResult + +logger = logging.getLogger(__name__) + +_UA = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0 Safari/537.36" + +# Booru sites don't hide a SafeSearch toggle behind an opaque server-side +# filter -- content is explicitly tagged with a rating (naming varies by +# site) that you query directly, so "off" here genuinely means "no rating +# restriction" rather than "hope the filter didn't trigger". +_GELBOORU_STYLE_RATING = {"on": "rating:general", "moderate": "-rating:explicit", "off": ""} +_DANBOORU_RATING = {"on": "rating:g", "moderate": "-rating:e", "off": ""} +_E621_RATING = {"on": "rating:s", "moderate": "-rating:e", "off": ""} + +_GELBOORU_STYLE_SITES = { + "gelbooru": "https://gelbooru.com/index.php", + "rule34": "https://api.rule34.xxx/index.php", +} + + +class BooruProvider: + """Image search against a booru-style board (the Danbooru-API family). + + Confirmed live during development (2026-09): + - e621: works with zero configuration, no account needed. + - gelbooru / rule34: now require an api_key + user_id (free account, + Account -> API Access Credentials) -- both returned HTTP 401/"Missing + authentication" without one, a change from their historical open access. + - danbooru: blocked by a Cloudflare bot-check ("Just a moment...") even + with a real browser User-Agent, regardless of credentials, from at + least some networks. Implemented (with login+api_key Basic Auth, the + documented method) but best-effort -- it may simply not work for you. + """ + + def __init__( + self, + site: str = "e621", + api_key: str | None = None, + user_id: str | None = None, + login: str | None = None, + ): + self.site = site + self.api_key = api_key + self.user_id = user_id + self.login = login + + async def search(self, query: str, n: int, safesearch: str = "moderate") -> list[ImageResult]: + try: + if self.site == "e621": + return await self._search_e621(query, n, safesearch) + if self.site in _GELBOORU_STYLE_SITES: + return await self._search_gelbooru_style(query, n, safesearch) + if self.site == "danbooru": + return await self._search_danbooru(query, n, safesearch) + logger.error("Unknown booru site: %r", self.site) + return [] + except Exception: + logger.exception("Booru search failed for query %r on %s", query, self.site) + return [] + + async def _search_e621(self, query: str, n: int, safesearch: str) -> list[ImageResult]: + rating = _E621_RATING.get(safesearch, "") + tags = f"{query} {rating}".strip() + params = {"tags": tags, "limit": min(max(n, 1), 320)} + # e621 asks API clients to identify themselves in the User-Agent. + headers = {"User-Agent": "DatasetForge/1.0 (image dataset tool)"} + async with httpx.AsyncClient(timeout=15.0, headers=headers) as client: + resp = await client.get("https://e621.net/posts.json", params=params) + resp.raise_for_status() + data = resp.json() + + results = [] + for post in data.get("posts", []): + file_info = post.get("file") or {} + url = file_info.get("url") + if not url: + continue # deleted/restricted posts have a null file url + general_tags = (post.get("tags") or {}).get("general", []) + results.append( + ImageResult( + url=url, + title=", ".join(general_tags[:8]) or None, + width=file_info.get("width"), + height=file_info.get("height"), + source_page=f"https://e621.net/posts/{post.get('id')}", + ) + ) + return results + + async def _search_gelbooru_style(self, query: str, n: int, safesearch: str) -> list[ImageResult]: + base_url = _GELBOORU_STYLE_SITES[self.site] + rating = _GELBOORU_STYLE_RATING.get(safesearch, "") + tags = f"{query} {rating}".strip() + params: dict[str, str | int] = { + "page": "dapi", "s": "post", "q": "index", "json": "1", + "limit": min(max(n, 1), 100), "tags": tags, + } + if self.api_key and self.user_id: + params["api_key"] = self.api_key + params["user_id"] = self.user_id + headers = {"User-Agent": _UA} + async with httpx.AsyncClient(timeout=15.0, headers=headers) as client: + resp = await client.get(base_url, params=params) + resp.raise_for_status() + data = resp.json() + + posts = data.get("post", []) if isinstance(data, dict) else (data or []) + results = [] + for post in posts: + url = post.get("file_url") + if not url: + continue + results.append( + ImageResult( + url=url, + title=post.get("tags"), + width=post.get("width"), + height=post.get("height"), + source_page=f"{base_url}?page=post&s=view&id={post.get('id')}", + ) + ) + return results + + async def _search_danbooru(self, query: str, n: int, safesearch: str) -> list[ImageResult]: + rating = _DANBOORU_RATING.get(safesearch, "") + # Anonymous Danbooru API access is limited to 2 combined tags -- a + # multi-word query plus the rating tag can exceed that even with + # credentials on a free account; let the API's own error surface + # rather than silently truncating what the user typed. + tags = f"{query} {rating}".strip() + params = {"tags": tags, "limit": min(max(n, 1), 200)} + auth = (self.login, self.api_key) if self.login and self.api_key else None + headers = {"User-Agent": _UA} + async with httpx.AsyncClient(timeout=15.0, headers=headers) as client: + resp = await client.get("https://danbooru.donmai.us/posts.json", params=params, auth=auth) + resp.raise_for_status() + data = resp.json() + + results = [] + for post in data: + url = post.get("file_url") + if not url: + continue + results.append( + ImageResult( + url=url, + title=post.get("tag_string_general"), + width=post.get("image_width"), + height=post.get("image_height"), + source_page=f"https://danbooru.donmai.us/posts/{post.get('id')}", + ) + ) + return results diff --git a/app/search/google.py b/app/search/google.py new file mode 100644 index 0000000..dd9dc10 --- /dev/null +++ b/app/search/google.py @@ -0,0 +1,75 @@ +from __future__ import annotations + +import logging +import re + +import httpx + +from .base import ImageResult + +logger = logging.getLogger(__name__) + +_UA = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0 Safari/537.36" + +# Google embeds original image URLs as plain quoted strings inside inline +# " + + +class TestGoogleProvider: + async def test_extracts_image_urls(self, monkeypatch): + async def handler(request): + return httpx.Response( + 200, text=page_html(["https://example.com/a.jpg", "https://example.com/b.png"]) + ) + + patch_async_client(monkeypatch, handler) + results = await GoogleProvider().search("cats", 10) + assert {r.url for r in results} == {"https://example.com/a.jpg", "https://example.com/b.png"} + + async def test_google_domain_urls_are_excluded(self, monkeypatch): + """Google's own UI chrome/icon URLs shouldn't be mistaken for results.""" + + async def handler(request): + return httpx.Response( + 200, + text=page_html( + ["https://www.google.com/logo.png", "https://example.com/real.jpg"] + ), + ) + + patch_async_client(monkeypatch, handler) + results = await GoogleProvider().search("cats", 10) + assert {r.url for r in results} == {"https://example.com/real.jpg"} + + async def test_stops_at_n_results(self, monkeypatch): + async def handler(request): + return httpx.Response(200, text=page_html([f"https://example.com/{i}.jpg" for i in range(20)])) + + patch_async_client(monkeypatch, handler) + results = await GoogleProvider().search("cats", 5) + assert len(results) == 5 + + async def test_rate_limit_response_returns_empty_list_not_an_error(self, monkeypatch): + async def handler(request): + return httpx.Response(429, text="unusual traffic detected") + + patch_async_client(monkeypatch, handler) + results = await GoogleProvider().search("cats", 5) + assert results == [] + + async def test_captcha_page_with_200_status_is_also_treated_as_blocked(self, monkeypatch): + async def handler(request): + return httpx.Response(200, text="Our systems have detected unusual traffic from your network.") + + patch_async_client(monkeypatch, handler) + results = await GoogleProvider().search("cats", 5) + assert results == [] + + async def test_other_non_200_status_returns_empty_list(self, monkeypatch): + async def handler(request): + return httpx.Response(503) + + patch_async_client(monkeypatch, handler) + results = await GoogleProvider().search("cats", 5) + assert results == [] + + async def test_request_failure_returns_empty_list_instead_of_raising(self, monkeypatch): + async def handler(request): + raise httpx.ConnectError("blocked") + + patch_async_client(monkeypatch, handler) + results = await GoogleProvider().search("cats", 5) + assert results == [] + + async def test_safesearch_off_maps_to_safe_off_param(self, monkeypatch): + captured = {} + + async def handler(request): + captured["safe"] = request.url.params.get("safe") + return httpx.Response(200, text=page_html([])) + + patch_async_client(monkeypatch, handler) + await GoogleProvider().search("cats", 5, safesearch="off") + assert captured["safe"] == "off" diff --git a/tests/test_search_yandex.py b/tests/test_search_yandex.py new file mode 100644 index 0000000..d18be44 --- /dev/null +++ b/tests/test_search_yandex.py @@ -0,0 +1,122 @@ +from __future__ import annotations + +import httpx +import pytest + +from app.search import yandex as yandex_module +from app.search.yandex import YandexProvider + +_real_async_client = httpx.AsyncClient + + +def patch_async_client(monkeypatch, handler) -> None: + def fake_async_client(**kwargs): + return _real_async_client(transport=httpx.MockTransport(handler), cookies=kwargs.get("cookies")) + + monkeypatch.setattr(yandex_module.httpx, "AsyncClient", fake_async_client) + + +def page_html(urls: list[str]) -> str: + items = "".join(f'img_href":"{u}",' for u in urls) + return f"{items}" + + +class TestYandexProvider: + async def test_extracts_image_urls_from_a_single_page(self, monkeypatch): + async def handler(request): + return httpx.Response(200, text=page_html(["https://example.com/a.jpg", "https://example.com/b.jpg"])) + + patch_async_client(monkeypatch, handler) + results = await YandexProvider().search("cats", 10) + + assert {r.url for r in results} == {"https://example.com/a.jpg", "https://example.com/b.jpg"} + + async def test_html_entities_in_urls_are_unescaped(self, monkeypatch): + async def handler(request): + return httpx.Response(200, text=page_html(["https://example.com/a.jpg?x=1&y=2"])) + + patch_async_client(monkeypatch, handler) + results = await YandexProvider().search("cats", 10) + assert results[0].url == "https://example.com/a.jpg?x=1&y=2" + + async def test_stops_once_n_is_reached_without_fetching_more_pages(self, monkeypatch): + call_count = 0 + + async def handler(request): + nonlocal call_count + call_count += 1 + return httpx.Response(200, text=page_html([f"https://example.com/{call_count}-{i}.jpg" for i in range(25)])) + + patch_async_client(monkeypatch, handler) + results = await YandexProvider().search("cats", 10) + + assert call_count == 1 # one page already exceeds n=10, no need for a second fetch + assert len(results) == 10 # ... but the final result list is still capped at n + + async def test_paginates_when_first_page_has_too_few_results(self, monkeypatch): + pages = [ + page_html(["https://example.com/1.jpg", "https://example.com/2.jpg"]), + page_html(["https://example.com/3.jpg", "https://example.com/4.jpg"]), + ] + call_count = 0 + + async def handler(request): + nonlocal call_count + resp = httpx.Response(200, text=pages[call_count]) + call_count += 1 + return resp + + patch_async_client(monkeypatch, handler) + results = await YandexProvider().search("cats", 4) + + assert call_count == 2 + assert len(results) == 4 + + async def test_stops_when_a_page_returns_no_new_urls(self, monkeypatch): + """Guards against looping forever if pagination stops producing + anything new (e.g. markup changed, or fewer results exist than n).""" + call_count = 0 + + async def handler(request): + nonlocal call_count + call_count += 1 + return httpx.Response(200, text=page_html(["https://example.com/same.jpg"])) + + patch_async_client(monkeypatch, handler) + results = await YandexProvider().search("cats", 50) + + # page 1 yields the one (new) URL; page 2 re-yields the same URL with + # nothing new, which is what actually triggers the stop -- so it does + # get fetched once before the loop gives up. + assert call_count == 2 + assert len(results) == 1 + + async def test_family_cookie_off_is_sent_for_safesearch_off(self, monkeypatch): + captured = {} + + async def handler(request): + captured["cookie"] = request.headers.get("cookie") + return httpx.Response(200, text=page_html([])) + + patch_async_client(monkeypatch, handler) + await YandexProvider().search("cats", 5, safesearch="off") + assert "family=0" in (captured["cookie"] or "") + + async def test_no_family_cookie_sent_for_moderate(self, monkeypatch): + captured = {} + + async def handler(request): + captured["cookie"] = request.headers.get("cookie") + return httpx.Response(200, text=page_html([])) + + patch_async_client(monkeypatch, handler) + await YandexProvider().search("cats", 5, safesearch="moderate") + assert not captured["cookie"] + + async def test_request_failure_returns_empty_list_instead_of_raising(self, monkeypatch): + async def handler(request): + raise httpx.ConnectError("blocked") + + patch_async_client(monkeypatch, handler) + results = await YandexProvider().search("cats", 5) + assert results == []