I build a Chinese-first toolchain for trustworthy LLM & document pipelines — six small, focused Python libraries. Everything installs from PyPI, runs on CPU, and documents its own limits honestly.
| Install | What you get |
|---|---|
pip install worddael |
Chinese-aware RAG chunking with exact source offsets — chunk.text == source[start:end] always holds |
pip install helan |
Chinese PII detection, masking, and reversible redaction for LLM / RAG pipelines |
pip install gewita |
Parse Chinese academic references, align quotations to sources, verify bibliography records |
pip install tellan |
Reconcile LLM API token usage & costs with tamper-evident audit ledgers |
pip install sothstan |
Statistical behavioral audits for OpenAI-compatible LLM API model claims |
pip install siftan |
Reproducible n-gram screening of training data for benchmark contamination |
All six: Python 3.10+, MIT, zero-required-dependency cores, CI-tested, offline-capable where possible.
Real-world feedback and edge cases are the most valuable contributions right now — open an issue or Discussion in any repo.