Skip to content

Russian (Cyrillic, inflected) queries return 0 results: keyword leg is exact-whole-word only, and the hash embedder cannot carry them over minRelevance #9

Description

@vvvvoronenko75-cmd

Package: dsh-library 0.2.11 (npm latest), default config (embedding.provider: hash, hybridWeight 0.6, minRelevance 0.15)
Host: @deepseek-ai/dsh 0.1.5-rc.2, web profile, Windows 11, Node v24.19.0
Related: #2 (closed) — the same failure for Chinese, fixed in 0.2.3 by unigram+bigram tokens for CJK runs and by gating on the hybrid score. Cyrillic (and every other inflected non-ASCII script) still has the pre-#2 behaviour on the keyword leg.

Symptom

Library with one Russian Markdown document (12 109 chars, 16 chunks; library_diagnose self-retrieval is fine). Natural Russian queries return nothing; ASCII queries on the same library work.

query results note
плагины рекомендации 0 the document contains "плагин…" in 10/16 chunks
плагины 0 plural; document has плагин, плагинов, плагинам
плагин 2 exact word form present in 2 chunks
dsh-cost-meter 15 ASCII
GitHub бриф экосистема DSH 8 carried by the ASCII tokens

Root cause (measured with the plugin's own functions on the stored index)

  1. Keyword leg = exact whole-word match. tokenize/cjkTokenize emit one token per Cyrillic run ([\u00c0-\uffff]+) with no sub-word expansion; only CJK runs get unigrams/bigrams (中文查询检索失效:minRelevance 过滤闸只看词法分 + tokenize 对 CJK 无分词边界 #2). Russian is inflected — плагины ≠ плагин ≠ плагинов — so scoreRelevance is 0 for any query whose word forms differ from the document's.
  2. The "semantic" leg is the hash embedder (word tokens + char trigrams, 256 dims). For плагины рекомендации its best cosine against any chunk is 0.098; with hybridWeight 0.6 the best hybrid score is 0.6 × 0.098 + 0.4 × 0 = 0.059 < 0.15, so the gate drops everything. The 中文查询检索失效:minRelevance 过滤闸只看词法分 + tokenize 对 CJK 无分词边界 #2 fix (gate on hybrid) does not help here because the hash leg alone never reaches the threshold for short Cyrillic queries.

Reproduced offline by re-running embedHash/scoreRelevance/cosine over dsh_library.json (my embedHash output matches the stored chunk vectors at cosine 1.0000): every number above comes from that replay.

Suggested fix

Extend the #2 treatment beyond CJK. Two cheap options, either is enough for the keyword leg:

  • Prefix stems for inflected scripts. For a run matching [\u0400-\u04ff]+ (Cyrillic — and arguably any non-ASCII, non-CJK run), emit the whole word plus a truncated stem token (e.g. first 5 letters + *). Replay on this index: плагины рекомендации goes from 0 to 1 passing chunk (the right one, seq 13, kw 0.25), какие плагины стоит поставить from 0 to 5.
  • Char n-grams inside the run (4-grams), the direct analogue of the CJK bigrams — more recall, more noise.

Independently, please document in the README that with provider: hash non-ASCII, inflected languages need exact word forms; the current README line ("lexical-grade embeddings … paraphrases") undersells it — plurals and cases fail, not just paraphrases. A real embedder (provider: ollama with a multilingual model such as bge-m3; the default nomic-embed-text is English-centric) fixes the semantic leg but does nothing for the keyword leg until the tokenizer changes.

Workarounds we use meanwhile

  • Query with the exact word forms that occur in the document, or with ASCII identifiers.
  • Lowering search.minRelevance to 0.05 is not a workaround: it lets through chunks with keyword score 0 (pure hash noise).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions