You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Russian (Cyrillic, inflected) queries return 0 results: keyword leg is exact-whole-word only, and the hash embedder cannot carry them over minRelevance #9
Package: dsh-library 0.2.11 (npm latest), default config (embedding.provider: hash, hybridWeight 0.6, minRelevance 0.15) Host: @deepseek-ai/dsh 0.1.5-rc.2, web profile, Windows 11, Node v24.19.0 Related:#2 (closed) — the same failure for Chinese, fixed in 0.2.3 by unigram+bigram tokens for CJK runs and by gating on the hybrid score. Cyrillic (and every other inflected non-ASCII script) still has the pre-#2 behaviour on the keyword leg.
Symptom
Library with one Russian Markdown document (12 109 chars, 16 chunks; library_diagnose self-retrieval is fine). Natural Russian queries return nothing; ASCII queries on the same library work.
query
results
note
плагины рекомендации
0
the document contains "плагин…" in 10/16 chunks
плагины
0
plural; document has плагин, плагинов, плагинам
плагин
2
exact word form present in 2 chunks
dsh-cost-meter
15
ASCII
GitHub бриф экосистема DSH
8
carried by the ASCII tokens
Root cause (measured with the plugin's own functions on the stored index)
Keyword leg = exact whole-word match.tokenize/cjkTokenize emit one token per Cyrillic run ([\u00c0-\uffff]+) with no sub-word expansion; only CJK runs get unigrams/bigrams (中文查询检索失效:minRelevance 过滤闸只看词法分 + tokenize 对 CJK 无分词边界 #2). Russian is inflected — плагины ≠ плагин ≠ плагинов — so scoreRelevance is 0 for any query whose word forms differ from the document's.
The "semantic" leg is the hash embedder (word tokens + char trigrams, 256 dims). For плагины рекомендации its best cosine against any chunk is 0.098; with hybridWeight 0.6 the best hybrid score is 0.6 × 0.098 + 0.4 × 0 = 0.059 < 0.15, so the gate drops everything. The 中文查询检索失效:minRelevance 过滤闸只看词法分 + tokenize 对 CJK 无分词边界 #2 fix (gate on hybrid) does not help here because the hash leg alone never reaches the threshold for short Cyrillic queries.
Reproduced offline by re-running embedHash/scoreRelevance/cosine over dsh_library.json (my embedHash output matches the stored chunk vectors at cosine 1.0000): every number above comes from that replay.
Suggested fix
Extend the #2 treatment beyond CJK. Two cheap options, either is enough for the keyword leg:
Prefix stems for inflected scripts. For a run matching [\u0400-\u04ff]+ (Cyrillic — and arguably any non-ASCII, non-CJK run), emit the whole word plus a truncated stem token (e.g. first 5 letters + *). Replay on this index: плагины рекомендации goes from 0 to 1 passing chunk (the right one, seq 13, kw 0.25), какие плагины стоит поставить from 0 to 5.
Char n-grams inside the run (4-grams), the direct analogue of the CJK bigrams — more recall, more noise.
Independently, please document in the README that with provider: hash non-ASCII, inflected languages need exact word forms; the current README line ("lexical-grade embeddings … paraphrases") undersells it — plurals and cases fail, not just paraphrases. A real embedder (provider: ollama with a multilingual model such as bge-m3; the default nomic-embed-text is English-centric) fixes the semantic leg but does nothing for the keyword leg until the tokenizer changes.
Workarounds we use meanwhile
Query with the exact word forms that occur in the document, or with ASCII identifiers.
Lowering search.minRelevance to 0.05 is not a workaround: it lets through chunks with keyword score 0 (pure hash noise).
Package: dsh-library 0.2.11 (npm
latest), default config (embedding.provider: hash,hybridWeight 0.6,minRelevance 0.15)Host: @deepseek-ai/dsh 0.1.5-rc.2,
webprofile, Windows 11, Node v24.19.0Related: #2 (closed) — the same failure for Chinese, fixed in 0.2.3 by unigram+bigram tokens for CJK runs and by gating on the hybrid score. Cyrillic (and every other inflected non-ASCII script) still has the pre-#2 behaviour on the keyword leg.
Symptom
Library with one Russian Markdown document (12 109 chars, 16 chunks;
library_diagnoseself-retrieval is fine). Natural Russian queries return nothing; ASCII queries on the same library work.плагины рекомендацииплагиныплагин,плагинов,плагинамплагинdsh-cost-meterGitHub бриф экосистема DSHRoot cause (measured with the plugin's own functions on the stored index)
tokenize/cjkTokenizeemit one token per Cyrillic run ([\u00c0-\uffff]+) with no sub-word expansion; only CJK runs get unigrams/bigrams (中文查询检索失效:minRelevance 过滤闸只看词法分 + tokenize 对 CJK 无分词边界 #2). Russian is inflected —плагины≠плагин≠плагинов— soscoreRelevanceis 0 for any query whose word forms differ from the document's.плагины рекомендацииits best cosine against any chunk is 0.098; withhybridWeight 0.6the best hybrid score is 0.6 × 0.098 + 0.4 × 0 = 0.059 < 0.15, so the gate drops everything. The 中文查询检索失效:minRelevance 过滤闸只看词法分 + tokenize 对 CJK 无分词边界 #2 fix (gate on hybrid) does not help here because the hash leg alone never reaches the threshold for short Cyrillic queries.Reproduced offline by re-running
embedHash/scoreRelevance/cosineoverdsh_library.json(myembedHashoutput matches the stored chunk vectors at cosine 1.0000): every number above comes from that replay.Suggested fix
Extend the #2 treatment beyond CJK. Two cheap options, either is enough for the keyword leg:
[\u0400-\u04ff]+(Cyrillic — and arguably any non-ASCII, non-CJK run), emit the whole word plus a truncated stem token (e.g. first 5 letters +*). Replay on this index:плагины рекомендацииgoes from 0 to 1 passing chunk (the right one, seq 13, kw 0.25),какие плагины стоит поставитьfrom 0 to 5.Independently, please document in the README that with
provider: hashnon-ASCII, inflected languages need exact word forms; the current README line ("lexical-grade embeddings … paraphrases") undersells it — plurals and cases fail, not just paraphrases. A real embedder (provider: ollamawith a multilingual model such asbge-m3; the defaultnomic-embed-textis English-centric) fixes the semantic leg but does nothing for the keyword leg until the tokenizer changes.Workarounds we use meanwhile
search.minRelevanceto 0.05 is not a workaround: it lets through chunks with keyword score 0 (pure hash noise).