Skip to content

Latest commit

 

History

History
292 lines (259 loc) · 17.3 KB

File metadata and controls

292 lines (259 loc) · 17.3 KB

Photobook Architecture Overview

System Architecture

The Photobook application is a full-stack AI-powered photo analysis and organization tool that builds a temporal-spatial Knowledge Graph (KG) from personal photos.

┌─────────────────────────────────────────────────────────────────────────────┐
│                              Frontend (React + Vite)                         │
│  ┌─────────────┐  ┌──────────────┐  ┌────────────┐  ┌─────────────────────┐ │
│  │ DashboardPage│  │PersonsPage  │  │ ObjectsPage│  │ KGVisualizationPage │ │
│  └──────┬──────┘  └──────┬───────┘  └─────┬──────┘  └──────────┬──────────┘ │
│         │                │                 │                    │            │
│         └───────────┴────────┴─────────────┴────────────────────┘            │
│                              │                                               │
│                    React Query / API Client                                  │
└──────────────────────────────┼───────────────────────────────────────────────┘
                               │ HTTP REST
                               ▼
┌─────────────────────────────────────────────────────────────────────────────┐
│                            Backend (FastAPI)                                 │
│  ┌─────────────────────────────────────────────────────────────────────────┐│
│  │                         API Routes (/api/v1)                             ││
│  │  /auth  /images  /photobooks  /persons  /objects  /detections  /kg /jobs││
│  └─────────────────────────────────────────────────────────────────────────┘│
│                               │                                              │
│  ┌────────────────────────────┼────────────────────────────────────────────┐│
│  │                      Service Layer                                       ││
│  │   ImageService   KGService   PersonsService   CaptionsService  etc.     ││
│  └────────────────────────────┼────────────────────────────────────────────┘│
│                               │                                              │
│  ┌────────────────────────────┴────────────────────────────────────────────┐│
│  │                    Background Jobs (arq + Redis)                         ││
│  │   VLM Tasks   Face Tasks   Location Tasks   Scene Graph Extraction      ││
│  └──────────────────────────────────────────────────────────────────────────┘│
└────────────────────────────────┬─────────────────────────────────────────────┘
                                 │
           ┌─────────────────────┼───────────────────────┐
           │                     │                       │
           ▼                     ▼                       ▼
┌─────────────────┐   ┌─────────────────┐   ┌─────────────────┐
│  VLM Service    │   │   Geo Service   │   │  Face Service   │
│  (port 8031)    │   │   (port 8030)   │   │  (port 8033)    │
│                 │   │                 │   │                 │
│  Qwen3-VL-32B   │   │  Google Places  │   │   InsightFace   │
│  llama.cpp      │   │  API Client     │   │   buffalo_l     │
└─────────────────┘   └─────────────────┘   └─────────────────┘

           │                     │                       │
           └─────────────────────┼───────────────────────┘
                                 ▼
┌─────────────────────────────────────────────────────────────────────────────┐
│                            Data Layer                                        │
│  ┌─────────────────────────┐   ┌─────────────────────────────────────────┐  │
│  │   PostgreSQL Database   │   │              Redis                       │  │
│  │   (Knowledge Graph)     │   │   (Job Queue + Circuit Breakers)         │  │
│  └─────────────────────────┘   └─────────────────────────────────────────┘  │
│                                                                              │
│  ┌─────────────────────────────────────────────────────────────────────────┐│
│  │                        File Storage                                      ││
│  │   uploads/images/   uploads/thumbnails/   uploads/faces/                ││
│  └─────────────────────────────────────────────────────────────────────────┘│
└──────────────────────────────────────────────────────────────────────────────┘

Technology Stack

Frontend

  • React 18 - UI framework with hooks
  • TypeScript - Type safety
  • Vite - Build tool and dev server
  • TanStack Query (React Query) - Server state management
  • Framer Motion - Animations
  • Tailwind CSS - Styling
  • Sigma.js + graphology - Knowledge Graph visualization

Backend

  • FastAPI - High-performance async Python framework
  • SQLAlchemy 2.0 - Async ORM with PostgreSQL
  • Pydantic v2 - Data validation
  • arq - Async background job queue (Redis-based)
  • httpx - Async HTTP client for service calls
  • Pillow - Image processing

ML Services

  • VLM Service: Qwen3-VL-32B via llama.cpp for vision-language understanding
  • Geo Service: Google Places API for location resolution
  • Face Service: InsightFace buffalo_l for face detection and embeddings

Infrastructure

  • PostgreSQL - Primary database (Knowledge Graph + user data)
  • Redis - Job queue, circuit breaker state, caching
  • File Storage - Local filesystem for images and thumbnails

Data Flow

1. Image Upload Flow

User uploads image
       │
       ▼
┌──────────────────┐
│  POST /images    │
│  (ImageService)  │
│  1. Save file    │
│  2. Generate     │
│     thumbnail    │
│  3. Extract EXIF │
│  4. Create Node  │
│     (type=IMAGE) │
└────────┬─────────┘
         │
         ▼
┌──────────────────┐
│ Enqueue jobs:    │
│ - vlm_extract_all│
│ - location_match │
│ - face_embed     │
└──────────────────┘

2. VLM Analysis Flow

vlm_extract_all_task
       │
       ▼
┌──────────────────────┐
│  VLM Service /extract│
│  - Detect objects    │
│  - Detect people     │
│  - Extract desc/tags │
│  - Scene context     │
│  - Relationships     │
└──────────┬───────────┘
           │
           ▼
┌──────────────────────┐
│  KG Ingestion        │
│  - Create DETECTION  │
│    nodes for objects │
│  - Create PERSON     │
│    nodes for people  │
│  - Create TAG nodes  │
│  - Create Claims     │
│    (relationships)   │
└──────────────────────┘

3. Knowledge Graph Model

The KG uses a claim-centric architecture:

┌─────────────────────────────────────────────────────────────────┐
│                         Node (kg_nodes)                          │
│                                                                  │
│  ┌────────────────────────────────────────────────────────────┐ │
│  │ id: UUID                                                    │ │
│  │ node_type: IMAGE | PERSON | DETECTION | PLACE | TAG | ...  │ │
│  │ canonical_name: str                                         │ │
│  │ properties: JSONB (flexible metadata)                       │ │
│  │ embedding: JSONB (for similarity search)                    │ │
│  │ user_id: FK → users                                         │ │
│  └────────────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────┘
                              │
                    ┌─────────┴─────────┐
                    ▼                   ▼
            subject_id              object_id
                    │                   │
┌───────────────────┴───────────────────┴─────────────────────────┐
│                         Claim (kg_claims)                        │
│                                                                  │
│  ┌────────────────────────────────────────────────────────────┐ │
│  │ id: UUID                                                    │ │
│  │ subject_id: FK → kg_nodes                                   │ │
│  │ predicate: str (e.g., "depicts_person", "has_tag")          │ │
│  │ object_id: FK → kg_nodes (optional)                         │ │
│  │ literal_value: str (for text values)                        │ │
│  │ confidence: float (0.0 - 1.0)                               │ │
│  │ status: proposed | weakly_supported | accepted | rejected   │ │
│  │ valid_from / valid_to: DateTime (temporal scope)            │ │
│  │ source_type: AI_VISION | USER | SYSTEM | EXTERNAL           │ │
│  │ source_model: str (e.g., "qwen3-vl-32b")                    │ │
│  └────────────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────┘
                              │
                              ▼
┌─────────────────────────────────────────────────────────────────┐
│                      Evidence (kg_evidence)                      │
│                                                                  │
│  ┌────────────────────────────────────────────────────────────┐ │
│  │ claim_id: FK → kg_claims                                    │ │
│  │ image_node_id: FK → kg_nodes (IMAGE type)                   │ │
│  │ bbox: JSONB {x, y, w, h}                                    │ │
│  │ mask_path: str (SAM segmentation mask)                      │ │
│  │ model_output: JSONB (raw VLM output)                        │ │
│  │ confidence: float                                           │ │
│  └────────────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────┘

Key Design Patterns

1. Circuit Breaker Pattern

External services (VLM, Geo, Face) are protected by circuit breakers:

  • Closed: Normal operation
  • Open: Service failures exceeded threshold, fast-fail for X seconds
  • Half-Open: Allow limited requests to test recovery
# core/circuit_breaker.py
vlm_circuit = get_vlm_circuit()
if vlm_circuit.is_open:
    raise ServiceUnavailable("VLM service is temporarily unavailable")

2. Repository Pattern

Services encapsulate database operations:

# kg/service.py
kg_service = KGService(session)
node = await kg_service.create_node(user_id, NodeType.PERSON, ...)
claim = await kg_service.create_claim(user_id, subject_id, predicate, ...)

3. Background Jobs with Retry

Long-running tasks use arq with exponential backoff:

# jobs/tasks/vlm.py
retrier = RetryWithBackoff(max_retries=2, base_delay=1.0)
for attempt in retrier:
    try:
        result = await call_vlm_service()
    except TransientError:
        await retrier.wait()

4. Human-in-the-Loop

All AI-generated claims start with status=proposed and can be:

  • Accepted: User confirms the claim
  • Rejected: User rejects the claim
  • Corrected: User modifies the claim

User actions are tracked in UserAction for provenance and audit.

Ports & Endpoints

Service Port Health Check
Backend API 8000 /health
VLM Service 8031 /health
Geo Service 8030 /health
Face Service 8033 /health
Frontend (dev) 5173 -

Configuration

Environment variables (.env file):

# Database
DATABASE_URL=postgresql+asyncpg://photobook:photobook@localhost:5432/photobook

# Redis
REDIS_URL=redis://localhost:6379

# External Services
VLM_SERVICE_URL=http://localhost:8031
GEO_SERVICE_URL=http://localhost:8030
FACE_SERVICE_URL=http://localhost:8033

# Google Places API
GOOGLE_PLACES_API_KEY=your-api-key

# Auth
SECRET_KEY=your-secret-key

Deployment

Development

./run-fullstack.sh  # Starts backend, workers, and services
cd frontend && npm run dev  # Start frontend dev server

Production Considerations

  • Use Gunicorn/Uvicorn workers behind nginx
  • Use managed PostgreSQL and Redis
  • Deploy ML services on GPU-enabled nodes
  • Implement proper secrets management
  • Set up health monitoring and alerting