Liangdu turns LaTeX papers into aligned English-Chinese study editions. It keeps the original English beside a Chinese translation, preserves math, and outputs browser-friendly HTML so Chinese dictionary extensions can add hover lookup, pinyin, and definitions.
The goal is reading practice, not publication typesetting. Liangdu favors clean parallel text that works with tools such as:
- A two-column HTML study page: English on the left, Chinese on the right.
- A two-column Typst/PDF study edition from the same structured document.
- A translation JSON artifact that can be reviewed, cached, and re-rendered.
- MathJax-rendered inline and display math in the HTML page.
- Parallel text layout for studying an English paper beside a Chinese translation.
- DeepL, OpenAI, Gemini, and Anthropic translation with bounded parallelism and an LRU SQLite cache.
- Automatic provider selection: DeepL quality-optimized translation with GPT-5.6 Terra fallback, or Terra alone when DeepL is not configured.
- Automatic rejection of translations that change math, citations, or numbers.
- Cache keys include source text, context, provider route, prompt, and model settings.
- Source-defined LaTeX macro expansion for simple no-argument commands.
- Numbered citations with a generated reference list.
- Rendered display math in HTML and PDF.
- Source figures, TikZ/PGFPlots figures, and LaTeX tables as full-width images with bilingual captions.
- Title-based output filenames when the input is generic, such as
main.tex. - Links to Chinese dictionary browser extensions for hover lookup and pinyin.
Install from GitHub:
uv tool install git+https://github.com/osteele/liangduOr run from a checkout:
git clone https://github.com/osteele/liangdu.git
cd liangdu
uv syncSet DEEPL_API_KEY, OPENAI_API_KEY, or both, then run:
uv run liangdu build path/to/article.tex --pdfOpen the generated HTML file in a browser with a Chinese dictionary extension
enabled. If the source filename is specific, Liangdu bases the HTML, Typst, and
PDF names on it. If the source is generic, such as main.tex, Liangdu derives a
filename from the document title.
Run a dependency check before the first build:
uv run liangdu doctorFor debugging or custom pipelines, the individual stages are available:
uv run liangdu extract path/to/article.tex --out out/article-study.source.json
uv run liangdu translate out/article-study.source.json --out out/article-study.zh.json
uv run liangdu render out/article-study.zh.json
uv run liangdu pdf out/article-study.zh.jsonTo inspect or keep the generated Typst source:
uv run liangdu typst out/article-study.zh.json --out out/article-study.typ
uv run liangdu pdf out/article-study.zh.json --out out/article-study.pdf --typst-out out/article-study.typThe HTML output remains the primary interactive study format because browser dictionary extensions can operate on it directly. The PDF is generated directly from Typst rather than converted from HTML, so its layout is controlled by the same semantic block model instead of browser CSS.
- Python 3.10 or newer.
DEEPL_API_KEYorOPENAI_API_KEYfor translation. Setting both enables DeepL-first translation with OpenAI structural fallback.typstfor PDF compilation.pandocfor LaTeX math conversion into Typst math.latexmkandpdftoppmfor rendering complex LaTeX tables, TikZ, and PGFPlots figures as images.- A TeX distribution such as MacTeX if your paper uses custom LaTeX packages.
On macOS, the non-Python tools are typically:
brew install typst pandoc popplerInstall MacTeX separately if latexmk is not already available.
Use --limit to translate only the first few blocks while testing credentials,
prompts, and cost:
uv run liangdu translate out/article-study.source.json --out out/sample.zh.json --limit 5Translations are cached by default in .liangdu/translation-cache.sqlite.
Cache entries are keyed by the source text, model, prompt, output format, and
model settings, provider route, glossary, and document context, so development
reruns do not spend credits for unchanged blocks. Use --no-cache to bypass the
cache.
The default --provider auto mode uses these routes:
- With both keys, DeepL's
quality_optimizedmodel is primary and GPT-5.6 Terra retranslates any block that changes protected math, citations, or numbers. - With only
DEEPL_API_KEY, DeepL is used and structurally invalid output stops the run instead of entering the cache. - With only
OPENAI_API_KEY, GPT-5.6 Terra is used.
Select a provider explicitly with --provider deepl, --provider openai,
--provider gemini, or --provider anthropic. The provider API keys are
DEEPL_API_KEY, OPENAI_API_KEY, GEMINI_API_KEY, and ANTHROPIC_API_KEY.
--model overrides the provider default; those defaults are gpt-5.6-terra,
gemini-3.1-pro-preview, and claude-fable-5 for the three LLM providers.
DeepL requests include surrounding document text as context and use XML
protection for structural tokens. Set DEEPL_GLOSSARY_ID or pass
--deepl-glossary-id to apply an existing DeepL glossary.
Run the offline evaluator on one artifact or on several provider candidates:
uv run liangdu evaluate out/deepl.zh.json out/openai.zh.json --out out/evaluation.jsonThe report checks missing translations, duplicate indices, changed math,
citations and numbers, source text copied as a translation, absence of Chinese
characters, unusual length ratios, and inconsistent translations of repeated
source text. It also verifies that multiple candidates contain the same English
blocks. Use --strict to exit non-zero on deterministic errors.
These checks do not score semantic accuracy or fluency and do not rank candidates. They are useful as a quality gate and for finding blocks that need closer review, but they are not a substitute for bilingual judgment.
Prepare a blinded pairwise benchmark from two translated artifacts:
uv run liangdu benchmark prepare \
out/deepl.zh.json out/openai.zh.json \
--out out/comparisons.jsonl --seed 42The command samples across block categories when a limit is set, randomizes
candidate order per block, and writes provider identities to
out/comparisons.jsonl.manifest.json, separate from the judge-visible JSONL.
Identical translations and missing candidates are excluded. Human preference,
error-label, and notes fields remain null or empty until a reviewer supplies
them; Liangdu does not manufacture synthetic gold labels.
Run the optional diagnostic judge on a small pilot first:
uv run liangdu benchmark judge out/comparisons.jsonl \
--out out/judgments.jsonl --limit 10The default judge is OpenAI gpt-5.6-sol. Select --provider gemini or
--provider anthropic to run an independent judge family; their defaults are
gemini-3.1-pro-preview and claude-fable-5. Each eligible comparison uses two
API calls, one for each candidate order. The output normalizes both calls back
to the original blinded labels and returns abstain when their overall or
per-dimension verdicts disagree. Dimensions cover semantic adequacy,
completeness, unsupported additions, naturalness/register, terminology, and
learner readability. Critical errors require agreement across both orders and
include concrete evidence.
Pairs that fail deterministic checks are skipped without an API call unless
--include-deterministic-failures is set. Judge output is diagnostic and must
not be treated as a production acceptance gate or human-equivalent evaluation
until it is calibrated against bilingual review.
Before bilingual labels are available, validate each judge against known synthetic defects. The control answers remain in a separate private manifest:
uv run liangdu benchmark controls out/deepl.zh.json \
--out out/controls.jsonl --seed 42 --limit 20
uv run liangdu benchmark judge out/controls.jsonl \
--provider openai --out out/controls.openai.jsonl \
--include-deterministic-failures
uv run liangdu benchmark judge out/controls.jsonl \
--provider gemini --out out/controls.gemini.jsonl \
--include-deterministic-failuresControls inject omissions, unsupported duplication, negation flips, or number changes into an existing translation. They measure whether a judge detects obvious known defects; they do not measure whether the original translation is correct.
Aggregate one or more judge files with the private manifest:
uv run liangdu benchmark summarize out/comparisons.jsonl \
out/judgments.openai.jsonl out/judgments.gemini.jsonl \
--manifest out/comparisons.jsonl.manifest.json \
--out out/benchmark-report.jsonThe report includes wins, ties, abstentions, order consistency, consensus critical-error counts, pairwise judge agreement, synthetic-control accuracy when applicable, agreement with any later human labels, and document-clustered bootstrap intervals when at least two source documents are present. Its review queue prioritizes judge disagreement, critical-error flags, order effects, and abstentions. Provider rankings remain provisional until a bilingual reviewer labels a sample. See the current model-validation status for the latest machine-only pilot and the remaining reviewer work.
Translation runs with bounded parallelism. The default is --parallel 4, which
is a practical starting point for translation calls. Override it if you hit rate
limits or want more conservative smoke tests:
uv run liangdu translate out/article-study.source.json --out out/article-study.zh.json --parallel 2For the local development fixture:
just sample
just sample-pdfThe core invariant is semantic block alignment, not matching PDF page breaks. Each row pairs one English source block with its Chinese translation. Browser pagination and print layout happen after the paired blocks have been rendered.
Planned pipeline:
LaTeX/PDF source
-> structured blocks
-> context- and glossary-aware translation
-> bilingual block model
-> two-column HTML and Typst/PDF
-> optional Tauri app or video output
This is an early implementation. It handles LaTeX section structure, paragraph blocks, citations, and common display math environments, but it is not yet a full LaTeX renderer. The next major step is glossary generation from each paper's terminology.
- Citation numbers are assigned by Liangdu and may differ from the original PDF.
- Figure and table references are numbered, but section/equation references are
still incomplete and may render as plain text such as
Sec.. - Typst math output is readable but not always as polished as the original LaTeX. MiTeX evaluation and rendered-equation fallback are on the roadmap.
- Complex LaTeX tables, TikZ, and PGFPlots are rendered as images, so they are visually faithful but not editable text in the PDF.
- Provider context and an optional DeepL glossary improve terminology consistency, but Liangdu does not yet generate a glossary from a paper.
The intended real-world test source is:
/Users/osteele/Documents/Papers/bayesian-slo-admission/main.tex