Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Local OCR

Give any AI agent offline eyes — batch-OCR scanned PDFs & images, 100% local 给任意 AI Agent 装上"离线眼睛":扫描 PDF / 图片文档的纯本地批量 OCR

Powered by PaddleOCR-VL-1.5 (GGUF) on LM Studio · no multimodal host model needed · one stdlib-only Python script

License: MIT Python LM Studio pip install


What is this / 这是什么

EN · Give your AI coding agent offline eyes. This is a single-file Python skill (ocr.py, standard library only) that any AI agent — ZCode, trae, Qoder, OpenCode, Cursor, Copilot, Cline, or a plain shell pipeline — can run to read scanned PDFs and text images, through a locally running PaddleOCR-VL-1.5 vision model served by LM Studio. Your host model doesn't need to be multimodal: a text-only LLM that can execute shell commands gains the ability to batch-OCR entire documents, returning paragraphs as-is, math formulas as LaTeX, and tables as Markdown. Everything stays on your machine.

中文 · 给 AI Agent 装上"离线眼睛"。本技能是一个单文件 Python 脚本(ocr.py,仅用标准库,无需 pip install),任何能执行命令行的 AI Agent——ZCode、trae、Qoder、OpenCode、Cursor、Copilot、Cline 或普通脚本流程——都可以调用它读取扫描版 PDF 和文字图片:识别由 LM Studio 在本地运行的 PaddleOCR-VL-1.5 视觉模型完成。你的主模型不需要多模态能力:一个只会执行命令的纯文本 LLM 也能因此批量"读"文档——正文原样输出、数学公式转 LaTeX、表格转 Markdown,且数据全程不出本机。

Highlights / 特性

  • 🤖 Built for AI agents / 为 AI Agent 而生 — drop the folder into ZCode / trae / Qoder / OpenCode / Cursor / Copilot / Cline and the agent knows when & how to call it (see CROSS_EDITOR.md). The host LLM doesn't need vision — a text-only model that can run shell commands gains full document-reading ability.
  • 🔒 100% local & private / 纯本地·隐私友好 — recognition runs on your own machine via LM Studio; images and text never leave it. No cloud OCR, no API key, no upload.
  • 📦 Zero dependencies / 零依赖 — ocr.py uses only the Python standard library (3.7+), nothing to pip install
  • 🧮 LaTeX formulas & Markdown tables — the model's native output format, kept as-is
  • 🛡 Hardened against the model's real failure modes / 比裸调模型可靠得多 — auto-retry on degenerate output, multi-pass recovery from silent omissions. See the next section.
  • ⚡ Fast — about 1–8 s per page on an RTX 4060 laptop GPU (CPU also works, just slower)
  • 🌐 Chinese / English / mixed layouts — references, URLs and DOIs come out verbatim
  • 📤 Machine-readable — --json output with text, token usage, elapsed time and a degenerate flag, so agents can verify result quality programmatically

Why not just call the model directly? / 为什么不直接裸调模型?

Calling the PaddleOCR-VL GGUF naively (a made-up prompt, low context, one shot per page) fails in ways that are silent and hard to debug. This skill hardens every call:

  • 🛡 Degenerate-output guard / 退化输出守卫 — GGUF vision models occasionally loop (S5S5S5…, repeated paragraphs). Every output is screened (exact repeated blocks, near-duplicate long lines, token-cap hits); degenerate results are automatically retried with different temperature + seed, and the cleanest attempt wins. Still suspicious → [WARN] on stderr and "degenerate": true in the JSON output.
  • ♻️ Multi-pass recovery / 防静默漏识 — formula-dense pages sometimes stop early and silently drop an entire column/section (about 30% of pages in our real-paper tests, with no error at all). --passes 2~3 runs the OCR several times and keeps the longest result — reliably recovers what was missed, at ~N× the time.
  • 📏 Official task prompts & correct context — the model is trained on prompts like OCR: (the colon matters; plain OCR is out-of-distribution), and must be loaded with ≥16k context or long pages get truncated mid-sentence around ~1200 output tokens. INSTALL.md covers both.

In our tests, a naive call produced a 1150-repeat runaway that consumed the entire token budget on a paper page; with the guard, all 24 test runs across a 12-page math preprint came out clean in single passes, at a 90%+ key-term hit rate with zero tuning.

Scope check / 边界说明: for full layout analysis of complex documents, the official PaddleOCR pipeline is still the heavier-duty option. This skill is the lightweight end of the spectrum — one file, zero dependencies, reliable on text-heavy documents, and honest about what it can't read.

Quick start / 快速开始

First-time setup (install LM Studio, download the GGUF model, the critical mmproj-* naming, 16k context) is in INSTALL.md — about 10 minutes and ~1.8 GB of disk.

# 0. health check
python ocr.py --check

# 1. single image → stdout
python ocr.py -i image.png

# 2. save to file / JSON output (tokens, elapsed, degenerate flag)
python ocr.py -i image.png -o result.txt
python ocr.py -i image.png --json

# 3. task-specific prompts (official task format, colon required)
python ocr.py -i formula.png -p "Formula Recognition:"

Scanned PDF workflow / 扫描 PDF 批量流程

# Step 1: PDF → PNG (any Python with PyMuPDF installed)
import fitz
doc = fitz.open("scan.pdf")
for i, page in enumerate(doc):
    page.get_pixmap(dpi=150).save(f"pages/page{i+1}.png")   # 150 DPI is enough
# Step 2: OCR every page
for f in pages/page*.png; do
  python ocr.py -i "$f" -o "text/$(basename "$f" .png).txt"
done
# Step 3: merge the .txt files in page order

Formula-dense pages (papers, lecture notes): add --passes 2 or 3. Plain text pages don't need it.

Use as an AI-agent skill / 作为 AI 编辑器技能使用

This is the main way to use Local OCR: copy the whole folder into your editor/agent's skill directory and restart it. The agent reads SKILL.md, learns when OCR is the right tool (batch scanned PDFs, text images, or a host model with no vision), and runs ocr.py itself — no multimodal host model required.

Editor / Agent Global skill folder
ZCode %USERPROFILE%\.zcode\skills\local-ocr\
trae / trae CN %USERPROFILE%\.trae-cn\skills\local-ocr\
Qoder / OpenCode / Cursor / Copilot / Cline see CROSS_EDITOR.md

Task prompts / 任务 Prompt

The model is trained with these exact task prompts — the colon matters, plain OCR is out-of-distribution:

Task --prompt value
General text (default) OCR:
Math formulas Formula Recognition:
Tables Table Recognition:
Charts Chart Recognition:

Configuration / 环境变量

Variable Default Meaning
LMSTUDIO_OCR_API http://127.0.0.1:1234/v1/chat/completions API endpoint
LMSTUDIO_OCR_MODEL paddleocr-vl-1.5 Model ID
LMSTUDIO_OCR_PROMPT OCR: Default prompt
LMSTUDIO_OCR_TEMPERATURE 0.45 Sampling temperature
LMSTUDIO_OCR_RETRIES 2 Extra attempts when output screens as degenerate
LMSTUDIO_OCR_PASSES 1 Runs per image; the longest result wins
LMSTUDIO_OCR_TIMEOUT 180 Request timeout (seconds)

Known limitations / 已知局限(如实说明)

  • It reads text, not geometry. For images whose point is the figure — geometry diagrams, circuit diagrams, optical paths — OCR recovers the printed labels and symbols but not the spatial structure (triangle shapes, point positions, curve shapes). Don't expect it to "understand" a diagram.
  • Whole-page small tables may lose their grid structure. For important tables, crop the region and run it separately with Table Recognition:.
  • Figure-heavy mixed pages occasionally trigger GGUF degeneration; the built-in guard retries automatically, but if [WARN] persists, crop the image or switch to a region-specific prompt.
  • This is a lightweight "raw VL-model call" tool. For full layout analysis of complex documents, the official PaddleOCR pipeline is the heavier-duty option.

Troubleshooting / 故障排查

Symptom Fix
Connection refused LM Studio not running, or Local Server not started (Developer → Start Server)
Model not found Load PaddleOCR-VL-1.5 in My Models
HTTP 400: Model does not support images mmproj file missing or misnamed — see INSTALL.md
Output truncated mid-sentence (~1200 tokens) Model loaded with context < 16k — reload with 16384
[WARN] repetition warning Crop the image / use a region prompt; lms runtime update helps too
Slow (> 10 s/page) Set GPU Offload to Max; doubled time on retried pages is normal

Full guide including uninstall/rollback: INSTALL.md.

Credits

License

MIT

About

Give any AI agent offline eyes: batch-OCR scanned PDFs & images with PaddleOCR-VL-1.5 (GGUF) on LM Studio — works with text-only LLMs, 100% local & private, outputs text / LaTeX / Markdown tables. 给 AI Agent 装上离线眼睛:纯本地批量 OCR,公式转 LaTeX、表格转 Markdown。

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages