Repository navigation
Add Yandex, Google, and booru board search providers - #1
Merged
Merged
Conversation
Adds three new pluggable ImageSearchProvider implementations alongside DuckDuckGo, selectable per-job from a new "Search source" card in the UI: - BooruProvider: Danbooru-API-family boards (e621/gelbooru/rule34/danbooru). SafeSearch maps onto an explicit rating: tag instead of a hidden toggle. Confirmed live: e621 needs zero config; gelbooru/rule34 now require an api_key+user_id (a policy change from their historical open access); danbooru is blocked by Cloudflare bot-checks even with valid credentials on at least some networks, so it's implemented best-effort. - YandexProvider: scrapes yandex.com/images. SafeSearch via a `family` cookie, confirmed live to actually change results (5/25 differed on a borderline query). Along the way, fixed a real regex bug in the image-URL extraction ([^&]+? was truncating any URL containing its own HTML-escaped "&"): the fix recovered 5 more valid URLs from the same live page. - GoogleProvider: scrapes Google Images. Confirmed live that Google blocks plain HTTP scraping aggressively (instant 429 with full browser headers, no prior request history) -- implemented and wired up as requested, but documented as unreliable; degrades to zero results plus a log warning rather than crashing. Introduces SearchConfig (app/models.py) and a build_search_provider() factory (app/search/__init__.py) so jobs.py no longer hardcodes DuckDuckGoProvider. UI: new "Search source" card with a provider dropdown and booru-specific sub-fields (board, api_key, user_id, login), each with a short reliability note. Adds 39 new tests across the new providers and the factory; updates existing jobs/models tests for the new build_search_provider() plumbing. All backend behavior (e621, Yandex, and the Google failure path) was also verified against the real live services, not just mocks.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds three new pluggable image search providers alongside DuckDuckGo, selectable per-job from a new Search source card in the UI, plus a SafeSearch control that now actually reaches all of them.
app/search/booru.py) — Danbooru-API-family boards:e621/gelbooru/rule34/danbooru. SafeSearch maps onto an explicitrating:tag instead of a hidden toggle.app/search/yandex.py) — scrapesyandex.com/images. SafeSearch via afamilycookie.app/search/google.py) — scrapes Google Images, documented as unreliable (see below).SearchConfig(app/models.py) + abuild_search_provider()factory (app/search/__init__.py) replace the old hardcodedDuckDuckGoProvider()injobs.py.Live-verified findings (not just mocked)
api_key+user_id— both returned HTTP 401 without one, a change from their historical open access. UI has fields for this.familycookie actually changes results (5/25 differed on a borderline query betweenfamily=0andfamily=2).A real bug found and fixed along the way
While writing tests for the Yandex URL extraction, found that the original regex (
[^&]+?) truncated any image URL that itself contained an HTML-escaped&(common in query strings) — it stopped at the first literal&instead of the real"terminator. Fixed to.+?(non-greedy up to the real terminator). Re-ran against the same live-captured Yandex page: recovered 5 more valid URLs (30 vs 25) that were previously silently dropped.Tests
39 new tests across the new providers (
tests/test_search_booru.py,tests/test_search_yandex.py,tests/test_search_google.py) and the factory (tests/test_search_factory.py); existingtests/test_jobs.py/tests/test_models.pyupdated for thebuild_search_provider()plumbing. Full suite: 166 passed.Reviewer notes
booru_api_key,booru_user_id,booru_login) are persisted in the browser'slocalStoragelike the existing LLM API keys, consistent with how this app already handles secrets (never written server-side).🤖 Generated with Claude Code