Skip to content

Add Yandex, Google, and booru board search providers - #1

Merged
Patrick16 merged 1 commit into
masterfrom
add-search-providers
Sep 9, 2026
Merged

Patrick16 merged 1 commit into
masterfrom
add-search-providers

Conversation

@Patrick16

Copy link
Copy Markdown
Owner

Summary

Adds three new pluggable image search providers alongside DuckDuckGo, selectable per-job from a new Search source card in the UI, plus a SafeSearch control that now actually reaches all of them.

  • Booru board (app/search/booru.py) — Danbooru-API-family boards: e621 / gelbooru / rule34 / danbooru. SafeSearch maps onto an explicit rating: tag instead of a hidden toggle.
  • Yandex (app/search/yandex.py) — scrapes yandex.com/images. SafeSearch via a family cookie.
  • Google (app/search/google.py) — scrapes Google Images, documented as unreliable (see below).

SearchConfig (app/models.py) + a build_search_provider() factory (app/search/__init__.py) replace the old hardcoded DuckDuckGoProvider() in jobs.py.

Live-verified findings (not just mocked)

  • e621 works with zero configuration.
  • Gelbooru / Rule34 now require an api_key + user_id — both returned HTTP 401 without one, a change from their historical open access. UI has fields for this.
  • Danbooru is blocked by a Cloudflare bot-check even with a real browser UA, regardless of credentials, from at least some networks. Implemented (login + API key, the documented method) but best-effort. Also: anonymous Danbooru API access is capped at 2 combined tags.
  • Yandex: confirmed the family cookie actually changes results (5/25 differed on a borderline query between family=0 and family=2).
  • Google: confirmed it blocks plain HTTP scraping aggressively — instant HTTP 429 with full browser-like headers and no prior request history. Implemented as requested, but a block degrades to zero results + a log warning, not a crash, and the UI note says to expect this.

A real bug found and fixed along the way

While writing tests for the Yandex URL extraction, found that the original regex ([^&]+?) truncated any image URL that itself contained an HTML-escaped & (common in query strings) — it stopped at the first literal & instead of the real " terminator. Fixed to .+? (non-greedy up to the real terminator). Re-ran against the same live-captured Yandex page: recovered 5 more valid URLs (30 vs 25) that were previously silently dropped.

Tests

39 new tests across the new providers (tests/test_search_booru.py, tests/test_search_yandex.py, tests/test_search_google.py) and the factory (tests/test_search_factory.py); existing tests/test_jobs.py/tests/test_models.py updated for the build_search_provider() plumbing. Full suite: 166 passed.

Reviewer notes

  • Google's provider is genuinely unreliable by nature of what it's scraping — that's expected, not a bug in this PR; it's surfaced clearly in the UI and README rather than hidden.
  • Booru credential fields (booru_api_key, booru_user_id, booru_login) are persisted in the browser's localStorage like the existing LLM API keys, consistent with how this app already handles secrets (never written server-side).

🤖 Generated with Claude Code

Adds three new pluggable ImageSearchProvider implementations alongside
DuckDuckGo, selectable per-job from a new "Search source" card in the UI:

- BooruProvider: Danbooru-API-family boards (e621/gelbooru/rule34/danbooru).
  SafeSearch maps onto an explicit rating: tag instead of a hidden toggle.
  Confirmed live: e621 needs zero config; gelbooru/rule34 now require an
  api_key+user_id (a policy change from their historical open access);
  danbooru is blocked by Cloudflare bot-checks even with valid credentials
  on at least some networks, so it's implemented best-effort.

- YandexProvider: scrapes yandex.com/images. SafeSearch via a `family`
  cookie, confirmed live to actually change results (5/25 differed on a
  borderline query). Along the way, fixed a real regex bug in the image-URL
  extraction ([^&]+? was truncating any URL containing its own HTML-escaped
  "&"): the fix recovered 5 more valid URLs from the same live page.

- GoogleProvider: scrapes Google Images. Confirmed live that Google blocks
  plain HTTP scraping aggressively (instant 429 with full browser headers,
  no prior request history) -- implemented and wired up as requested, but
  documented as unreliable; degrades to zero results plus a log warning
  rather than crashing.

Introduces SearchConfig (app/models.py) and a build_search_provider()
factory (app/search/__init__.py) so jobs.py no longer hardcodes
DuckDuckGoProvider. UI: new "Search source" card with a provider dropdown
and booru-specific sub-fields (board, api_key, user_id, login), each with a
short reliability note.

Adds 39 new tests across the new providers and the factory; updates
existing jobs/models tests for the new build_search_provider() plumbing.
All backend behavior (e621, Yandex, and the Google failure path) was also
verified against the real live services, not just mocks.
@Patrick16
Patrick16 merged commit a125917 into master Sep 9, 2026
1 check passed
@Patrick16
Patrick16 deleted the add-search-providers branch September 9, 2026 19:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant