Status: research/educational prototype. Does not diagnose disease or replace a medical professional.
An end-to-end pipeline that takes a large histopathology image, detects tissue regions, extracts patches, classifies each patch for tumor presence with a CNN, reconstructs a slide-level probability heatmap, ranks suspicious regions, and reports quantitative tissue statistics — visualized through an interactive zoomable viewer.
Most "medical AI" demo projects stop at a patch-level classifier trained on a clean, pre-cut dataset. Real digital pathology systems have to deal with images far too large to feed into a network directly (tens of thousands of pixels per side), separate actual tissue from blank slide background, turn per-patch predictions back into something spatially meaningful, and do all of this fast enough to be useful. This project builds that full pipeline, not just the classifier at its center.
Real whole-slide image files (.svs, via OpenSlide) are often 1–4GB each and
introduce format-handling complexity that has nothing to do with the ML/CV
skills this project is meant to demonstrate. Given a 10-day timeline, the
MVP pipeline runs against a synthetic large image: a mosaic stitched
from real PatchCamelyon patches at true whole-slide-like dimensions. This
lets every stage of the pipeline — tissue detection, tiling, inference,
heatmap reconstruction, quantification — be built and validated honestly.
Real OpenSlide/.svs support was originally scoped as a stretch goal and
has since been completed — see the Screenshots section and
ARCHITECTURE.md's "Real WSI support" for what was built and verified.
Binary classification: does a 96x96 histopathology patch contain tumor tissue (metastatic) or healthy tissue?
PatchCamelyon (PCam) — 327,680 color patches (96x96px) extracted from histopathology scans of lymph node sections, labeled for presence of metastatic tissue. A subset is used here to keep training tractable.
Large image (synthetic WSI-scale mosaic)
│
▼
Tissue detection (Otsu threshold on HSV saturation — drop blank background)
│
▼
Patch extraction (only over detected tissue, overlapping tiles)
│
▼
CNN inference (ResNet18, transfer learning) — batched, with a benchmarked
│ single-vs-batched / CPU-vs-GPU comparison
▼
Patch classification + Grad-CAM (per-patch interpretability)
│
▼
Heatmap reconstruction (stitch per-tile probabilities back to full image)
│
▼
Suspicious region ranking + quantification (tissue %, tumor-area %,
│ top-N regions by mean probability)
▼
Interactive viewer (FastAPI + OpenSeadragon: pan/zoom, heatmap overlay,
click a region to inspect the underlying patch + score)
- Baseline classifier — ResNet18 (ImageNet-pretrained), fine-tuned on PCam patches.
- Evaluation — precision, recall, F1, AUC, confusion matrix (not just accuracy — false negatives matter more in a diagnostic setting).
- Explainability — Grad-CAM to visualize which regions of a patch drove the prediction.
- Tissue detection — Otsu thresholding on the HSV saturation channel to separate tissue from blank slide background before wasting inference on empty tiles (a real preprocessing step in digital pathology).
- WSI-scale tiling inference — tile the large image into overlapping patches over tissue regions only, classify each tile, stitch results into a full-image heatmap.
- Quantification — tissue area %, estimated tumor area %, ranked list of most suspicious regions by mean patch probability.
- Performance benchmarking — single-image vs batched inference throughput, and CPU vs GPU if available. This is the piece that connects your systems background to the ML work — most beginner projects skip it entirely.
- Interactive viewer — FastAPI backend serving the image, heatmap overlay, and per-region data; a lightweight HTML/JS frontend using OpenSeadragon for pan/zoom (no React — it adds build complexity without adding capability here). Click a region → see the patch, predicted class, and confidence.
histo-classifier/
├── src/
│ ├── dataset.py # PCam loading & preprocessing
│ ├── model.py # ResNet18 classifier definition
│ ├── train.py # training loop
│ ├── evaluate.py # metrics: precision/recall/F1/AUC/confusion matrix
│ ├── gradcam.py # Grad-CAM implementation
│ ├── tissue_detection.py # Otsu-based tissue vs. background segmentation
│ ├── mosaic.py # builds the synthetic large "WSI-scale" image
│ ├── tiling_inference.py # tiling + batched inference + heatmap stitching
│ ├── quantification.py # tissue %, tumor %, suspicious region ranking
│ └── benchmark.py # single vs batched, CPU vs GPU timing comparison
├── app/
│ ├── server.py # FastAPI backend: serves image/heatmap/regions
│ ├── static/
│ │ └── index.html # OpenSeadragon viewer frontend
│ └── demo_app.py # (fallback) Streamlit demo — see Fallback Plan
├── notebooks/ # exploration notebooks
├── data/ # (not committed — see .gitignore)
├── outputs/ # trained weights, heatmaps, benchmark results
├── ARCHITECTURE.md # pipeline + design decisions in more depth
└── PROGRESS.md # daily build log
python -m venv venv
source venv/bin/activate # or venv\Scripts\activate on Windows
pip install -r requirements.txtOnce outputs/model_v3.pth and outputs/synthetic_wsi.png exist (see
Results below for how they were produced), start the server with:
python scripts/run_demo.pyThis formalizes the startup process used throughout development into one
robust command: it kills any stale server already on the port (a real bug
hit during development — a leftover process silently served outdated
results while a new one failed to start), waits for the server by polling
the actual API rather than guessing a fixed delay, and prints the exact
follow-up command for exposing it publicly via a Cloudflare quick tunnel
if needed. See scripts/run_demo.py's docstring for the full list of
problems this specifically fixes.
Each phase has a checkpoint: if you're behind schedule at that point, cut scope in the stated order rather than rushing every remaining piece.
| Days | Goal | If behind, cut in this order |
|---|---|---|
| 1–3 | Python-for-ML, NumPy, OpenCV, PyTorch, CNN/transfer-learning fundamentals | Reduce depth of theory study, not the coding practice — you learn faster by building |
| 4–6 | Train ResNet18 on PCam subset; evaluate (precision/recall/F1/AUC/confusion matrix) | Shrink subset size / epochs before skipping evaluation metrics |
| 7 | Tissue detection + synthetic mosaic + tiling inference + heatmap | Skip tissue detection (run tiling over the whole image) before skipping the heatmap itself |
| 8 | Quantification + region ranking + performance benchmark | Skip the CPU/GPU comparison before skipping tissue-% / tumor-% quantification |
| 9 | Interactive viewer (FastAPI + OpenSeadragon) | Fall back to app/demo_app.py (Streamlit) — a working simple viewer beats a broken fancy one |
| 10 | README, ARCHITECTURE.md, screenshots/demo, resume bullet, interview prep | Never cut this — an undocumented project is much weaker in an interview than a documented smaller one |
Classification (ResNet18, transfer learning, PatchCamelyon full "train" split as training data — 100,000 of 262,144 patches used — with layer1/ layer2 frozen and stronger augmentation; "test" split held out for evaluation):
| Metric | Value |
|---|---|
| Precision | 91.3% |
| Recall @ default threshold (0.5) | 76.1% |
| AUC | 0.936 |
At the default 0.5 threshold, the model still favors precision over recall — in a screening context, missing a real tumor (false negative) is costlier than a false alarm a pathologist reviews and dismisses. A threshold sweep on the held-out test set found a better operating point:
| Threshold | Precision | Recall |
|---|---|---|
| 0.5 | 91.3% | 76.1% |
| 0.3 | 87.5% | 83.6% |
| 0.15 | 81.9% | 90.8% |
| 0.1 | 78.8% | 93.5% |
0.15 is used as the default threshold throughout the pipeline
(quantification.py) based on this tradeoff.
Iteration history — three training runs, each diagnosing and fixing a real problem, not just re-running the same thing hoping for a better number:
- v1 (20,000 patches from PCam's smaller "val" split, no regularization): 92.0% precision, 62.2% recall @ 0.5, AUC 0.894. Diagnosed clear overfitting — train accuracy reached 98.9% while validation loss climbed from 0.76 to 1.61 across 10 epochs, and recall plateaued at ~72% even at the most aggressive thresholds tested, indicating the model was confidently wrong on real tumor cases, not just uncalibrated.
- v2 (same 20,000 patches, added L2 regularization/weight decay): AUC improved to 0.903, recall ceiling rose from ~72% to ~84% at aggressive thresholds — real evidence regularization improved calibration, not just accuracy. Val_loss still climbed over training, though less severely.
- v3 (100,000 patches from the full "train" split — 5x more data — plus partial layer freezing and stronger augmentation): the real fix. Val_loss stayed in a tight 0.28–0.40 band across all 10 epochs instead of climbing, recall @ 0.5 jumped to 76.1% (+13.9 points over v2), and critically, recall at aggressive thresholds is still rising rather than plateaued (93.5% at threshold 0.1, vs. v2's hard ceiling of 84.1%) — the clearest sign the model is now well-calibrated rather than confidently wrong. This confirmed that dataset size was the dominant lever for this overfitting problem, more so than regularization or partial freezing alone (freezing through layer2 only reduces trainable parameters to ~94%, since ResNet18's parameter mass is concentrated in layer4 — so it's a real contributor here, but not the main one).
Performance benchmark (batched vs. single-tile inference, CPU vs. GPU, 200 tiles, measured on the v2 checkpoint):
| Single-tile | Batched | Speedup from batching | |
|---|---|---|---|
| CPU | 78.4 tiles/sec | 197.8 tiles/sec | 2.52× |
| GPU | 227.0 tiles/sec | 1330.0 tiles/sec | 5.86× |
Batching matters more on GPU than CPU because a GPU has far more idle parallel compute to exploit when fed multiple tiles at once — a CPU has much less parallelism available in the first place, so batching helps less dramatically there.
Pipeline: tissue detection, synthetic-mosaic tiling, heatmap
reconstruction, and quantification all run successfully end-to-end against
the trained model (v2, at the time — worth re-running against v3 for full
consistency). See ARCHITECTURE.md for a documented limitation around
tissue-mask/heatmap disagreement on the synthetic mosaic.
the trained model — see ARCHITECTURE.md for a documented limitation
around tissue-mask/heatmap disagreement on the synthetic mosaic.
The interactive viewer running against the trained model, deployed live on a Kaggle GPU session and accessed via a Cloudflare tunnel:
Mosaic view, with tissue statistics and click-to-inspect:
Heatmap overlay, showing per-tile tumor probability:
The sidebar shows real, live output: 54.24% tissue coverage, 71.51% tumor
area (at the tuned 0.3 threshold — see Results below for why), and a
clicked region's tumor probability (88.3% in this example) fetched from
the /api/region endpoint in real time.
Real WSI file support: the pipeline also runs, unmodified, against
genuine whole-slide image files (.svs) via OpenSlide — not just the
synthetic mosaic. This is the standard OpenSlide test slide
(CMU-1-Small-Region.svs), with the tissue-detection/tiling pipeline
correctly finding and processing only the actual tissue region (the
background is untouched, confirming the tissue mask precisely tracked the
real slide's tissue boundary):
See ARCHITECTURE.md's "Real WSI support" section for how this works and
its current scope limits.
Core classifier, evaluation, tissue detection, tiling/heatmap pipeline,
quantification, and performance benchmarking are complete and verified
against real data and a real trained model. The FastAPI backend for the
interactive viewer is complete and verified — all four endpoints
(/api/stats, /api/image, /api/heatmap, /api/region) tested with
real HTTP requests against the live server. The OpenSeadragon browser
frontend (app/static/index.html) is built but has not yet been visually
tested in an actual browser — see PROGRESS.md for the
daily build log.


