AI-Powered 3D Volumetric Food Estimation & Personal Nutrition Intelligence Platform
- 1. Executive Summary & Platform Overview
- 2. System Architecture & End-to-End Data Flow
- 3. Complete Technology Stack Matrix
- 4. Three Core Platform Pillars
- 5. Computer Vision & Mathematical Core
- 6. Repository Directory & Architecture Map
- 7. PlateSense Model Training, Datasets & Active Learning
- 8. Real-Time Webcam Detection & Standalone Gradio UI
- 9. Complete Installation & Quick Start Guide
- 10. REST API Endpoint Reference
- 11. Production Deployment & Dockerization
- 12. Open Source, Cloning & Contribution Guide
PlateSense / Food Caliper is an end-to-end artificial intelligence platform built to solve portion size estimation and nutritional auditing without manual food weighing. Traditional dietary tracking apps rely on subjective human guesses, resulting in macro measurement errors of up to 40%.
This project bridges deep learning object detection (Ultralytics YOLOv8), 3D spatial volume estimation, bulk density mapping, and a production-grade web application built with React 18, Vite, FastAPI, and MySQL.
- Multi-Food Identification: Accurately detects and classifies multiple distinct food items in a single image.
-
Volumetric Estimation: Converts 2D pixel contours into 3D volume (
$cm^3$ or$mL$ ) using physical plate calibration. -
Weight & Macro Synthesis: Estimates exact physical mass (
$grams$ ), calories, protein, carbohydrates, fats, and micronutrients. - Interactive Web Portal: Full-featured React web client with live webcam scanning, calibration sliders, daily calorie rings, macro trends, and downloadable CSV/PDF reports.
The diagram below illustrates how a user request flows through the full platform stack:
sequenceDiagram
autonumber
actor User as User / Browser
participant React as React Web App (:5173)
participant API as FastAPI Backend (:8000)
participant ML as Volumetric Analysis Engine
participant YOLO as YOLOv8 Neural Model (best.pt)
participant DB as MySQL Database (food_caliper_db)
User->>React: Upload Meal Photo + Set Plate Diameter (e.g., 25cm)
React->>API: POST /api/v1/analysis/upload (FormData + user_id)
API->>ML: Forward image file & scale parameters
ML->>YOLO: Run inference on image tensor
YOLO-->>ML: Return Bounding Boxes, Polygon Contours & Class Scores
ML->>ML: 1. Calculate Pixel-to-CM Scale Ratio from Plate Diameter
ML->>ML: 2. Compute 2D Area (cm²) & Estimate Height (cm)
ML->>ML: 3. Estimate 3D Volume (cm³) using Geometric Mesh
ML->>ML: 4. Lookup Density (g/cm³) & Calculate Mass (g)
ML->>ML: 5. Query Nutritional Database for Cal/Protein/Carb/Fat
ML-->>API: Return Structured Analysis JSON Payload
API->>DB: Persist scan record in `analyses` & `analysis_items`
API-->>React: HTTP 200 OK + Analysis Results
React-->>User: Render Bounding Box Overlays, Volume (cm³), Weight (g) & Charts
| Layer | Technology | Purpose & Description |
|---|---|---|
| Frontend Framework | React 18.3, TypeScript 5.8, Vite 5.4 | Single Page Application framework with strict typing and fast HMR bundler. |
| UI & Styling | Tailwind CSS 3.4, Shadcn UI (Radix) | Accessible design primitives with custom dark mode glassmorphism theme. |
| Animations | GSAP 3, Framer Motion 12, Lenis | Inertia smooth scrolling, scroll triggers, and dynamic target cursor. |
| Data Viz & State | TanStack React Query v5, Recharts 2.15 | Asynchronous data fetching, caching, and responsive macro trend visualization. |
| Backend Framework | Python 3.10+, FastAPI, Uvicorn | Asynchronous RESTful API server with high concurrency performance. |
| Database & ORM | MySQL 8.0, SQLAlchemy, Pydantic | Relational database storage for users, meal logs, and nutritional density records. |
| AI / Machine Learning | Ultralytics YOLOv8, PyTorch, OpenCV | Single-stage deep convolutional network for object detection and contour segmentation. |
| Auth & Security | PyJWT, Passlib (Bcrypt) | Secure JSON Web Token authentication with bcrypt password hashing. |
| UI Standalone / Test | Gradio UI (app.py, app2.py) |
Interactive rapid prototyping web UI for local testing and research. |
-
Landing Portal (
/): Interactive hero showcase, bento grid features, technology brand loop, and smooth Lenis scrolling. -
Volumetric Scanner (
/analysis): Drag-and-drop image uploader, live webcam capture, real-time plate diameter calibration slider ($15\text{ cm} - 35\text{ cm}$ ), visual bounding overlays, and instant CSV audit exports. -
Daily Dashboard (
/dashboard): Progress rings tracking daily caloric targets, macro distribution charts, hydration water intake counter, and weekly intake history. -
Audit Reports (
/reports): Searchable historical scan log with date filtering and detailed nutrient breakdowns. -
User Profile (
/profile): Physical metrics configuration (Height, Weight, Age) for automatic BMR/TDEE target calculation.
- FastAPI Web Application (
main.py): Asynchronous endpoint handlers with CORS configuration. - Relational Models (
app/models): SQLAlchemy database models for Users, Meal Analyses, and Analysis Items. - Pydantic Schemas (
app/schemas): Request/response validation schemas for auth tokens, profile updates, and scan payloads. - JWT Middleware: Token generation, validation, and user session management.
-
Inference Model (
best.pt): YOLOv8 neural network trained on annotated food datasets. -
Spatial Calibration Matrix: Translates image pixel dimensions to physical metric units (
$cm$ ). -
3D Mesh Reconstruction: Calculates physical volume (
$cm^3$ ) based on geometric modeling (Ellipsoid, Cylinder, Prism). -
Nutritional Database: Maps food items against
Indian_Food_Nutrition_Processed.csvand MySQL tables for exact nutrient density calculations.
Transforming a 2D photograph into 3D volumetric estimates and mass (
[ 2D Image (px) ] ──(YOLOv8)──> [ Contours & Boxes ] ──(Scale Matrix)──> [ Real Dimensions (cm) ]
│
(Geometric Mesh)
▼
[ Calorie/Macro Data ] ◄──(Nutrition DB)─── [ Mass (g) ] ◄──(Density ρ)─── [ Volume (cm³) ]
Input image
Where
Camera distance varies across photos. To establish physical scale, the system uses a known reference object—by default, the physical dinner plate diameter
-
Pixel Diameter Calculation:
$$D_{\text{plate (px)}} = \max(W_{\text{plate-bbox}}, H_{\text{plate-bbox}})$$ -
Pixel-to-Centimeter Scale Ratio (
$S$ ):$$S = \frac{D_{\text{plate (cm)}}}{D_{\text{plate (px)}}} \quad \left[\frac{\text{cm}}{\text{px}}\right]$$ -
Physical Area Conversion: Given contour area in pixels
$A_{\text{px}}$ derived via Green's theorem on polygon contour vertices:$$A_{\text{cm}^2} = A_{\text{px}} \times S^2$$
Single-view RGB cameras lack direct depth channels. Height
-
Ellipsoidal / Dome-shaped Foods (e.g., Rice bowl, Curry, Salad):
$$V_{\text{cm}^3} = \frac{2}{3} \cdot A_{\text{cm}^2} \cdot H_{\text{est}}$$ -
Planar / Flat Foods (e.g., Pizza, Pancake, Dosa, Roti):
$$V_{\text{cm}^3} = A_{\text{cm}^2} \cdot H_{\text{flat-thickness}}$$ -
Cylindrical / Prismatic Foods (e.g., Cake slice, Sandwiches):
$$V_{\text{cm}^3} = A_{\text{cm}^2} \cdot H_{\text{height}} \cdot K_{\text{taper}}$$
Where
Once volume
Sample Bulk Densities ($\rho$):
- White Rice (cooked):
$0.85 \text{ g/cm}^3$ - Chicken Curry:
$1.05 \text{ g/cm}^3$ - Green Salad:
$0.35 \text{ g/cm}^3$ - Whole Wheat Roti:
$0.65 \text{ g/cm}^3$
Finally, macronutrient values are calculated against the nutritional database:
Food/
├── backend/ # FastAPI Python Backend Service
│ ├── app/
│ │ ├── models/ # SQLAlchemy Database Models (User, Analysis, etc.)
│ │ ├── schemas/ # Pydantic Input/Output Schemas
│ │ ├── routes/ # API Endpoints (Auth, Analysis, User Stats)
│ │ ├── utils/ # JWT Authentication & Password Hashing
│ │ └── database.py # MySQL Connection Engine & Base Class
│ ├── uploads/ # Uploaded Image Files Storage
│ ├── main.py # FastAPI Application Entry Point
│ ├── migrate_db.py # Database Migration Utility
│ ├── requirements.txt # Backend Python Dependencies
│ └── .env # Backend Environment Configuration
│
├── Website/Food_caliper/ # React 18 + Vite Production Web Application
│ ├── src/
│ │ ├── assets/ # Image & Vector Branding Assets
│ │ ├── components/ # Reusable UI, Dock, TargetCursor & Animations
│ │ ├── contexts/ # ThemeContext (Dark/Light mode)
│ │ ├── hooks/ # Custom React Hooks (use-mobile, use-toast)
│ │ ├── pages/ # App Routes (Index, Login, Analysis, Dashboard, Reports, Profile)
│ │ ├── services/ # Central Axios API Client (`apiClient.ts`)
│ │ ├── App.tsx # Main App Routing & Providers Component
│ │ └── main.tsx # React Entry Point
│ ├── package.json # Frontend Dependencies & Scripts
│ ├── tailwind.config.ts # Design System & Token Configuration
│ ├── vite.config.ts # Vite Bundler Settings & Aliases
│ ├── IMPLEMENTATION_GUIDE.md # Full Step-by-Step Setup Guide
│ └── README.md # Frontend Dedicated Documentation
│
├── code/ # Google Colab Training Notebooks
│ └── PlateSense(1).ipynb # YOLOv8 Training & Fine-Tuning Notebook
│
├── images1/ # Training & Evaluation Screenshot Artifacts
├── volumetric_food_analysis.py # Core Computer Vision & Volumetric Math Engine
├── food_calibration_data.py # Food Density & Caloric Reference Database
├── food_detection.py # Standalone YOLO Detection & Webcam Script
├── app.py / app1.py / app2.py # Standalone Gradio Interface Scripts
├── best.pt # Trained YOLOv8 PyTorch Model Weights
├── Indian_Food_Nutrition_Processed.csv # Processed Food Nutrition Dataset
├── database_schema.sql # Raw SQL Schema Script
├── Dockerfile # Production Multi-Stage Container Dockerfile
├── requirements.txt # Root Python Requirements
└── README.md # Master Repository Documentation (This File)
The core machine learning objective is detecting, classifying, and segmenting food items present on a plate in real time. YOLOv8 (Ultralytics) was selected for the following advantages:
- Single-Stage Real-Time Speed: Evaluates bounding coordinates and class probabilities in a single pass.
- High Accuracy: Strong performance even with smaller custom datasets.
- Multi-Task Capability: Native support for object detection, instance segmentation, and classification.
- Ease of Deployment: Simple Python API via the
ultralyticspackage.
Public datasets used for training and benchmarking include:
Each image in the dataset requires:
- Precise bounding boxes surrounding each food item.
- Class labels (e.g., Rice, Curry, Salad, Roti, Waffles).
- Image annotations were created using LabelImg and Roboflow.
Training YOLOv8 requires GPU acceleration. Follow these steps in Google Colab:
- Open Google Colab and upload
code/PlateSense(1).ipynb. - Enable GPU: Navigate to
Runtime → Change runtime type$\rightarrow$ setHardware AcceleratortoGPU$\rightarrow$ click Save. - Verify GPU availability:
!nvidia-smi - Mount Google Drive to preserve dataset files, checkpoints, and model weights across sessions:
from google.colab import drive drive.mount('/content/drive')
- Specify your dataset directory path:
dataset_path = '/content/drive/MyDrive/Data_1_'
pip install ultralytics opencv-python matplotlib numpy pandasOrganize dataset directories in YOLO format:
dataset/
├── images/
│ ├── train/
│ └── val/
└── labels/
├── train/
└── val/
Configure data.yaml to define dataset locations, class counts, and label names:
path: /path/to/dataset
train: images/train
val: images/val
names:
0: rice
1: curry
2: salad
3: roti
4: wafflesfrom ultralytics import YOLO
# Load base model (yolov8n.pt or yolov8s.pt)
model = YOLO("yolov8n.pt")
# Train on custom dataset
model.train(data="data.yaml", epochs=50, imgsz=640)Tip: If accuracy needs improvement, increase training duration to 100-150 epochs and enable image augmentations.
results = model.predict(source="test_image.jpg", conf=0.5)
results.show()Model performance is evaluated using standard computer vision metrics:
- Mean Average Precision (mAP): Evaluates bounding box overlap and class correctness across thresholds.
- Precision & Recall: Measures false positive vs. missed detection trade-offs.
- F1 Score: Harmonic mean of Precision and Recall.
- Inference Speed: Milliseconds per frame.
After training completes, YOLOv8 automatically saves model weight checkpoints in runs/detect/train/weights/:
best.pt: Highest performing model checkpoint.last.pt: Final epoch training state.
Save and export the weights:
model.save("models/food_detection_best.pt")from ultralytics import YOLO
# Load previously trained best weights
model = YOLO("runs/detect/train/weights/best.pt")
# Continue training on updated dataset
model.train(data="dataset/data.yaml", epochs=20, imgsz=640)To continuously improve detection precision over time, implement an Active Learning Loop:
[ New Unseen Food Images ] ──> [ Model Prediction ] ──> [ Human Review ]
│
[ Model Fine-Tuning (best.pt) ] ◄── [ Add to Dataset ] ◄── [ Re-Label Errors ]
- Run inference on newly collected food photographs.
- Review predictions to identify missing items or misclassifications.
- Re-label corrected bounding boxes in Roboflow or LabelImg.
- Merge new annotated samples back into the training split.
- Fine-tune using
best.ptas the starting checkpoint.
food_detection.py supports real-time food detection directly via your computer's webcam feed.
| Key | Action |
|---|---|
q |
Quit webcam live feed |
s |
Save screenshot of current detected frame |
+ / - |
Increase or decrease detection confidence threshold dynamically |
app2.py (or app.py) launches an interactive Gradio web application for quick local experimentation:
python app2.pyAccess the interface at http://localhost:7860.
Features:
- Upload food images and specify the custom
best.ptmodel path. - Adjust Plate Diameter (cm) slider for scale calibration.
- Download structured results in JSON and CSV formats.
The system produces structured outputs containing visual bounding overlays, summary totals, and detailed item attributes.
{
"summary": {
"total_items_detected": 1,
"total_volume_ml": 638.99,
"total_volume_liters": 0.639,
"total_weight_grams": 543.15,
"total_weight_kg": 0.543,
"items_with_components": 0
},
"food_items": [
{
"item_id": 1,
"name": "waffles",
"confidence": 0.8909,
"bounding_box": {
"x_min": 120,
"y_min": 85,
"x_max": 450,
"y_max": 390
},
"volume": {
"volume_ml": 638.99,
"weight_grams": 543.15,
"weight_kg": 0.543,
"area_cm2": 316.2,
"estimated_height_cm": 2.89,
"density_g_per_ml": 0.85
}
}
]
}Ensure the following tools are installed:
- Python: v3.10 or higher
- Node.js: v18.0.0 or higher (npm / bun)
- MySQL Server: v8.0 or higher (or Workbench)
- Log in to MySQL:
mysql -u root -p
- Create database:
CREATE DATABASE IF NOT EXISTS food_caliper_db CHARACTER SET utf8mb4 COLLATE utf8mb4_unicode_ci;
- Database tables are automatically initialized by SQLAlchemy when starting the backend.
- Navigate to the backend folder:
cd backend - Create and activate Python virtual environment:
- Windows:
python -m venv venv venv\Scripts\activate
- macOS / Linux:
python3 -m venv venv source venv/bin/activate
- Windows:
- Install dependencies:
pip install -r requirements.txt
- Create
.envfile insidebackend/.env:DATABASE_URL=mysql+pymysql://root:your_password@localhost:3306/food_caliper_db SECRET_KEY=super-secret-jwt-key-change-this-in-production-32chars ALGORITHM=HS256 ACCESS_TOKEN_EXPIRE_MINUTES=1440 YOLO_MODEL_PATH=../best.pt API_HOST=0.0.0.0 API_PORT=8000 FRONTEND_URL=http://localhost:5173
- Run the server:
API will run at
python main.py
http://localhost:8000. Interactive docs available athttp://localhost:8000/docs.
- Open a new terminal and navigate to the frontend folder:
cd Website/Food_caliper - Install dependencies:
npm install
- Create
.env.localinsideWebsite/Food_caliper/.env.local:VITE_API_URL=http://localhost:8000
- Start the Vite dev server:
npm run dev
- Open browser at:
http://localhost:5173
To run standalone volumetric analysis without the web server:
python app2.pyAccess Gradio interface at http://localhost:7860.
| Method | Endpoint | Purpose | Payload |
|---|---|---|---|
POST |
/api/v1/auth/register |
Register a new user | { username, email, password, full_name } |
POST |
/api/v1/auth/login |
Login user & return JWT token | { email, password } |
GET |
/api/v1/auth/profile/{id} |
Get user profile details | None |
PUT |
/api/v1/auth/profile/{id} |
Update body stats (Height, Weight, Age) | { height_cm, weight_kg, age } |
| Method | Endpoint | Purpose | Query / Form Data |
|---|---|---|---|
POST |
/api/v1/analysis/upload |
Upload image & run YOLO 3D analysis | FormData: file, params: user_id |
GET |
/api/v1/analysis/{id} |
Get details for specific analysis ID | params: user_id |
GET |
/api/v1/analysis/history/all |
Get paginated analysis history | params: user_id, limit, offset |
DELETE |
/api/v1/analysis/{id} |
Delete analysis record | params: user_id |
| Method | Endpoint | Purpose | Query Parameters |
|---|---|---|---|
GET |
/api/v1/user/dashboard |
Get daily summary & macro rings | user_id |
GET |
/api/v1/user/stats/weekly |
Get weekly calorie trends & food categories | user_id |
GET |
/api/v1/user/stats/monthly |
Get monthly historic macro data | user_id |
The root directory includes a production multi-stage Dockerfile:
# Build Docker image
docker build -t platesense-app:latest .
# Run container with environment file
docker run -d \
-p 8000:8000 \
-p 5173:5173 \
--env-file backend/.env \
--name platesense-container \
platesense-app:latestTo set up PlateSense / Food Caliper locally or contribute to development, clone the repository:
# Clone the repository
git clone https://github.com/your-username/PlateSense.git
# Navigate into the project root directory
cd PlateSenseWe welcome open-source contributions from developers, researchers, and computer vision enthusiasts! Follow these steps to contribute:
- Fork the Repository: Click Fork at the top right of the GitHub repository.
- Create a Feature Branch:
git checkout -b feature/AmazingFeature
- Commit your Changes:
git commit -m 'Add some AmazingFeature' - Push to the Branch:
git push origin feature/AmazingFeature
- Open a Pull Request: Submit a PR describing your feature, bug fix, or performance improvement.
PlateSense / Food Caliper • Open Source AI Volumetric Food & Nutrition Platform





