Line-level restoration of degraded Arabic manuscripts: a U-Net that takes a damaged handwritten line and returns a readable one, trained on synthetic damage that is calibrated against real manuscript decay, and evaluated by whether an OCR model can actually read the result.
Two models, the same generator, different objectives:
Level 1 — unimodal |
Level 2 — recaware |
|
|---|---|---|
| Objective | L1 + ResNet50 perceptual (+ optional PatchGAN) | the same, plus a frozen HATFormer OCR critic |
| Optimizes for | visual fidelity | legibility |
| Text | not used | used as a loss target, never as an input |
| Trainer | train_unimodal.py |
train_recaware.py |
The Level-2 critic is the point of the project: the restored image is read by a
frozen HATFormer and scored against the true transcription, so gradients flow
back into the pixels that matter for reading. Level 3 (conditional diffusion)
is scaffolded in models/diffusion.py but not trained.
On real damaged lines from a held-out book — the honest test — the image-only model does not help at all, while the recognition-aware one does:
| Input to the reader | CER % | WER % |
|---|---|---|
| Degraded (untouched) | 29.81 | 59.98 |
| Level 1 restored | 30.15 | 59.98 |
| Level 2 restored | 28.15 | 58.44 |
And the ranking inverts between the two ways of measuring — Level 1 wins every pixel metric it is optimized for, and loses on reading:
| PSNR ↑ | SSIM ↑ | CER ↓ | |
|---|---|---|---|
| Degraded | 22.76 | 0.9498 | 35.11 |
| Level 1 | 34.66 | 0.9886 | 28.12 |
| Level 2 | 30.66 | 0.9781 | 27.13 |
(Synthetic test set. CER is read by the frozen HATFormer reader; see the paper for the full protocol, per-book breakdowns, and limitations.)
Provenance: this repo is the restoration half of a larger Arabic-manuscript benchmark project — the corpus lives in AraMS-28k and the annotation pipeline in RefLAM. Everything needed to train, evaluate, and demo restoration is self-contained here except the dataset and the model weights, which are on HuggingFace — see Get the data Get the weights Real damage subset
conda create -n arams-restore python=3.11 && conda activate arams-restore
pip install -r requirements.txtThere is nothing to pip install -e . — the code is plain modules under src/,
and every entry point puts src/ on sys.path itself, so run the commands below
from the repo root. If you import this code from your own script, either do
the same or export PYTHONPATH=src.
transformers==4.37.2 is a behavior-critical pin (the HATFormer critic's
VisionEncoderDecoder + interpolate_pos_encoding path). Runs on CUDA, MPS, or
CPU — every entry point calls pick_device().
Neither is in git — both come from the Hub. Run these from the repo root; the target folder names are what the configs resolve against.
pip install -U "huggingface_hub[cli]"
huggingface-cli login # only needed while these repos are private
# the dataset (~3.4 GB): line images + metadata.csv (the split source of truth)
huggingface-cli download Archatext/AraMS-28k-HTR --repo-type dataset --local-dir AraMS-28k-HTR
# the real damaged book_09 lines (~20 MB): the input to the headline result.
# Small and independent of the corpus above — grab this even if you skip the rest.
huggingface-cli download Archatext/AraMS-Restore --repo-type dataset --local-dir real_damage
# the frozen HATFormer OCR critic / reader (~1.3 GB)
huggingface-cli download Archatext/hatformer-arams28k --local-dir checkpoints/ocr/muharaf_ours
# the restoration generators (~124 MB each), shipped as one zip
huggingface-cli download Archatext/Restoration restore.zip --local-dir checkpoints
unzip -q checkpoints/restore.zip -x "__MACOSX/*" "*.DS_Store" -d checkpoints
rm checkpoints/restore.zipLayouts and the leak rule: data/README.md, checkpoints/README.md.
The damage engine needs no checkpoints and no corpus — any line PNG will do (the
real_damage/ pull above is a 22 MB way to get one):
# see the synthetic damage engine on any line image -> outputs/deg_test/montage.png
python scripts/test_degradation.py --image real_damage/book_09_page_009/line_011.png --n 6
python scripts/test_degradation.py --image <line.png> --seed 0 # reproducible variants
python scripts/test_degradation.py --image <line.png> --gray # 384x384 canvas instead of the stripWith checkpoints/restore/*/best.pt in place:
# restore one image
python scripts/test_restoration.py --image <degraded.png> --out restored.png
# the web demo: drag-and-drop / paste an image, before-after slider
python app/server.py # -> http://127.0.0.1:7860
python app/server.py --checkpoint checkpoints/restore/unimodal/best.ptDrop in a line (or hit Try a random real-damage sample) and compare degraded
vs. restored — stacked as above, or with a before/after slider. The sample shown
is real_damage/book_09_page_010/line_015.png, from a held-out book.
Level 1 first, evaluate it, then Level 2 — the ladder is sequential, and the
L1-vs-L2 comparison is only fair because both configs share image size, epochs,
samples_per_image, and seed.
# Level 1 — image-only U-Net
python src/restoration/train_unimodal.py --smoke # sanity check first
python src/restoration/train_unimodal.py # configs/restore_unimodal.yaml
# Level 2 — recognition-aware (frozen HATFormer critic)
python src/restoration/train_recaware.py --smoke
python src/restoration/train_recaware.py # configs/restore_recaware.yaml
# resume Level 2 from a saved epoch
python src/restoration/train_recaware_resume.py \
--checkpoint checkpoints/restore/recaware/best.pt --start-epoch 8Damage is generated on the fly from clean train-split lines — there is no
pre-degraded dataset to build. samples_per_image: 4 means each clean line
yields 4 distinct degraded variants per epoch, non-deterministic in training and
seeded per line in eval so every run sees identical test damage.
configs/restore_recaware.yaml puts the frozen critic on cuda:1 and falls back
to the main device automatically if there is no second GPU.
# Level 2 on held-out splits: PSNR/SSIM + CER read by the frozen critic
python scripts/eval_recaware.py
python scripts/eval_recaware.py --splits test --max-batches 5 # smoke
# L1 vs L2 on the SAME synthetic test set (one shared judge)
python scripts/compare_restoration.py
python scripts/compare_restoration.py --splits test --limit 300
# the headline: REAL damaged lines, CER before vs after restoration
python scripts/restore_real.py --score
# qualitative sheets: N lines per book, stacked damaged/L1/L2
python scripts/visualize_compare.py --per-book 5
# OCR-only reference numbers (per-book CER of the reader itself)
python scripts/eval_hatformer.py --splits testResults land under outputs/ (git-ignored). Report figures for the two
architectures: python scripts/make_architecture_figures.py.
AraMS-Restore/
├── configs/
│ ├── degradation.yaml # the synthetic damage parameters — calibrate here
│ ├── restore_unimodal.yaml # Level 1
│ ├── restore_recaware.yaml # Level 2 (+ critic checkpoint & leak rule)
│ └── ocr_hatformer.yaml # OCR eval settings (the reader / critic)
│
├── src/ # no wrapper package — import as `restoration.…`
│ ├── restoration/ # the work itself
│ │ ├── degradation/
│ │ │ └── synthetic_degradation.py # the damage engine (erase, tear, holes,
│ │ │ # bleed-through, ink blot, feathering)
│ │ ├── datasets.py # serves (damaged, clean, mask, text) pairs;
│ │ │ # build_strip / build_canvas live here
│ │ ├── losses.py # masked L1 + ResNet50 perceptual + GAN
│ │ ├── models/
│ │ │ ├── unet_pix2pix.py # the generator (+ PatchGAN discriminator)
│ │ │ ├── recognition_aware.py# HATFormerCritic — frozen, differentiable
│ │ │ └── diffusion.py # Level 3 scaffold, untrained
│ │ ├── train_unimodal.py # Level 1 trainer
│ │ ├── train_recaware.py # Level 2 trainer
│ │ └── train_recaware_resume.py
│ ├── ocr/ # ONLY what restoration needs from the OCR side
│ │ ├── model.py # rebuild HATFormer + load its 50,265-token
│ │ │ # Arabic BBPE tokenizer for eval
│ │ └── datasets.py # HATFormer line dataset / canvas helpers
│ ├── eval/{metrics.py,ocr_accuracy.py} # PSNR/SSIM, CER/WER, before-after CER
│ ├── data/
│ │ ├── metadata.py # THE metadata.csv loader (schema + line_id)
│ │ └── normalize.py # THE canonical Arabic normalization (clean_gt)
│ └── utils/{io.py,seed.py} # run logging, seeding
│
├── scripts/ # thin CLI wrappers (see Evaluate above)
├── app/ # Flask before/after demo (server.py + static/)
├── docs/ # README assets (the demo screenshot)
├── data/ # split provenance + the schema docs
├── AraMS-28k-HTR/ # the dataset from the Hub — git-ignored
│ # (images/ + metadata.csv; see data/README.md)
├── real_damage/ # real damaged book_09 lines — the honest test
│ # set; from the Hub, git-ignored
├── test_data/ # a handful of degraded/restored samples
└── checkpoints/ # git-ignored; see its README
A line is resized to height 64 keeping aspect and right-padded to width 1152
(build_strip), padding stays pure black and is excluded by a mask, so the U-Net
sees an actual line rather than a mostly-black square. 1152 = 3×384 tiles cleanly
back into HATFormer's 384×384 RTL-flipped stacked canvas, which is what makes the
Level-2 recognition loss exact: the critic reads the restored strip in precisely
the format it was trained on, differentiably, with no re-alignment. Images are
RGB, not grayscale, so red rubrication survives restoration.
- Training damage comes from
train-split lines only. Evaluation damage is generated from held-out lines with fixed per-line seeds — so test damage is identical across runs, and no evaluation line ever contributes a training pair. - Manuscript-level disjointness: a book is in exactly one split (book_03/05/09 are held out entirely).
- The OCR critic must not have trained on the restoration test books, otherwise it has seen the answers and both the training signal and the evaluation are contaminated.
- One normalization (
clean_gt), applied once, already baked intogt_text— never re-normalize on top of it. - Level 1 before Level 2, with matched hyper-parameters — that comparison is the contribution.
Details and per-split book lists: data/README.md.
If you use this code, the checkpoints, or the real-damage lines:
@misc{aramsrestore2026,
title = {AraMS-Restore: Recognition-Aware Restoration of Damaged
Historical Arabic Manuscripts},
author = {Zellagui Mohamed Diaa, Mohamed Guechaoui, Chaib Souleyman},
year = {2026},
note = {TODO: arXiv ID / venue}
}The line images and transcriptions derive from AraMS-28k — please cite the corpus as well, and the HATFormer and Muharaf work the OCR reader builds on.
Code: MIT (LICENSE). Dataset lines, transcriptions, and model weights:
CC BY-NC-SA 4.0, inherited from AraMS-28k.
