Photobook runs as six processes: a FastAPI backend, a background worker, a
React frontend, and four model services. setup.sh provisions all of them.
| Version | macOS | Debian/Ubuntu | |
|---|---|---|---|
| Python | 3.11+ | brew install [email protected] |
apt install python3 python3-venv |
| Node | 18+ | brew install node |
apt install nodejs npm |
| PostgreSQL | 14+ | brew install postgresql@16 |
apt install postgresql |
| pgvector | 0.5+ | brew install pgvector |
apt install postgresql-16-pgvector |
| Redis | 6+ | brew install redis |
apt install redis-server |
pgvector is required: face similarity search is an indexed nearest-neighbour
query and the migrations will not apply without the extension.
Start the services before setup:
brew services start postgresql@16 && brew services start redis # macOS
sudo systemctl start postgresql redis # Linuxgit clone <repository-url> photobook
cd photobook
./setup.shThis creates virtual environments, installs dependencies, creates the
photobook database, applies migrations, generates a SECRET_KEY, and
installs frontend packages. It is safe to re-run.
Useful flags:
./setup.sh --check # report prerequisites and exit
./setup.sh --no-gpu # skip the SAM and face services./run.sh
backend/.venv/bin/python -m scripts.seed_demo_userThen open http://localhost:5173 and sign in as [email protected] /
demo12345. Upload photographs; the pipeline populates the knowledge graph
from them.
| Service | Port |
|---|---|
| Frontend | 5173 |
Backend API (/api/v1/docs) |
8020 |
| geo-service | 8030 |
| vlm-service | 8031 |
| sam-service | 8032 |
| face-service | 8033 |
No weights are distributed with this repository. What each service needs:
Face recognition — nothing to do. InsightFace downloads buffalo_l
(~300 MB) into services/face-service/models/ on first use.
Segmentation (SAM3) — download the checkpoint and place it at
services/sam-service/sam3.pt (~3.4 GB). Without it, segmentation is skipped
and the rest of the pipeline runs normally. Set DEVICE in
services/sam-service/.env to cuda (NVIDIA), mps (Apple silicon) or cpu.
CPU works but costs roughly 5 s per image encode.
Vision-language model — not bundled and not downloaded. The vlm-service is a thin client for an OpenAI-compatible endpoint. Point it at your own:
# services/vlm-service/.env
BACKEND_TYPE=vllm
LLAMA_SERVER_URLS=http://your-host:8090
MODEL_NAME=autoDevelopment used Qwen3-VL-32B-Instruct served by vLLM on 2×RTX 4090. Any
OpenAI-compatible vision endpoint should work, though prompts in
captions/service.py and search/llm_parser.py are tuned for Qwen3-VL and
may need adjusting for other models.
Without a VLM the system still ingests images, reads EXIF and detects faces, but produces no objects, tags, captions or relationships.
Geographic landmark naming uses the Google Places API. Without a key, GPS
coordinates are still read from EXIF and stored; only the place-name lookup is
skipped. To enable it, put a key in services/geo-service/.env:
GOOGLE_PLACES_API_KEY=your-key
cd backend && ./.venv/bin/python -m pytest -qTests needing photographs skip unless you supply your own — see
backend/tests/data/README.md. Integration tests under tests/api/ also need
the model services running.
SECRET_KEY must be at least 32 characters — backend/.env is missing a
key. Generate one with
python3 -c "import secrets; print(secrets.token_urlsafe(48))".
pgvector is not installed — install the extension package, then
psql -d photobook -c "CREATE EXTENSION vector".
Uploads stay "pending" — the worker is not running or cannot reach Redis.
Check logs/backend-worker.log.
Captions are empty — no VLM is configured. Check
curl localhost:8031/health and LLAMA_SERVER_URLS.