Skip to content

Latest commit

 

History

History
129 lines (95 loc) · 4 KB

File metadata and controls

129 lines (95 loc) · 4 KB

Installation

Photobook runs as six processes: a FastAPI backend, a background worker, a React frontend, and four model services. setup.sh provisions all of them.

Prerequisites

Version macOS Debian/Ubuntu
Python 3.11+ brew install [email protected] apt install python3 python3-venv
Node 18+ brew install node apt install nodejs npm
PostgreSQL 14+ brew install postgresql@16 apt install postgresql
pgvector 0.5+ brew install pgvector apt install postgresql-16-pgvector
Redis 6+ brew install redis apt install redis-server

pgvector is required: face similarity search is an indexed nearest-neighbour query and the migrations will not apply without the extension.

Start the services before setup:

brew services start postgresql@16 && brew services start redis   # macOS
sudo systemctl start postgresql redis                            # Linux

Setup

git clone <repository-url> photobook
cd photobook
./setup.sh

This creates virtual environments, installs dependencies, creates the photobook database, applies migrations, generates a SECRET_KEY, and installs frontend packages. It is safe to re-run.

Useful flags:

./setup.sh --check    # report prerequisites and exit
./setup.sh --no-gpu   # skip the SAM and face services

First run

./run.sh
backend/.venv/bin/python -m scripts.seed_demo_user

Then open http://localhost:5173 and sign in as [email protected] / demo12345. Upload photographs; the pipeline populates the knowledge graph from them.

Service Port
Frontend 5173
Backend API (/api/v1/docs) 8020
geo-service 8030
vlm-service 8031
sam-service 8032
face-service 8033

Model weights

No weights are distributed with this repository. What each service needs:

Face recognition — nothing to do. InsightFace downloads buffalo_l (~300 MB) into services/face-service/models/ on first use.

Segmentation (SAM3) — download the checkpoint and place it at services/sam-service/sam3.pt (~3.4 GB). Without it, segmentation is skipped and the rest of the pipeline runs normally. Set DEVICE in services/sam-service/.env to cuda (NVIDIA), mps (Apple silicon) or cpu. CPU works but costs roughly 5 s per image encode.

Vision-language model — not bundled and not downloaded. The vlm-service is a thin client for an OpenAI-compatible endpoint. Point it at your own:

# services/vlm-service/.env
BACKEND_TYPE=vllm
LLAMA_SERVER_URLS=http://your-host:8090
MODEL_NAME=auto

Development used Qwen3-VL-32B-Instruct served by vLLM on 2×RTX 4090. Any OpenAI-compatible vision endpoint should work, though prompts in captions/service.py and search/llm_parser.py are tuned for Qwen3-VL and may need adjusting for other models.

Without a VLM the system still ingests images, reads EXIF and detects faces, but produces no objects, tags, captions or relationships.

Optional: landmark resolution

Geographic landmark naming uses the Google Places API. Without a key, GPS coordinates are still read from EXIF and stored; only the place-name lookup is skipped. To enable it, put a key in services/geo-service/.env:

GOOGLE_PLACES_API_KEY=your-key

Verifying

cd backend && ./.venv/bin/python -m pytest -q

Tests needing photographs skip unless you supply your own — see backend/tests/data/README.md. Integration tests under tests/api/ also need the model services running.

Troubleshooting

SECRET_KEY must be at least 32 characters — backend/.env is missing a key. Generate one with python3 -c "import secrets; print(secrets.token_urlsafe(48))".

pgvector is not installed — install the extension package, then psql -d photobook -c "CREATE EXTENSION vector".

Uploads stay "pending" — the worker is not running or cannot reach Redis. Check logs/backend-worker.log.

Captions are empty — no VLM is configured. Check curl localhost:8031/health and LLAMA_SERVER_URLS.