Skip to content

Repository files navigation

KARINA: Knowledge-Augmented Representation for Ingredient-level Nutrition Analysis

Paper DOI License Python PyTorch

Official implementation of "KARINA: Knowledge-Augmented Representation for Ingredient-level Nutrition Analysis from Food Images", published at IEEE CAI 2026 (IEEE Conference on Artificial Intelligence). Developed as part of a master's thesis at National Cheng Kung University.

KARINA estimates five nutritional values (calories, mass, fat, carbohydrates, and protein) from food images. Instead of directly predicting nutrition with a large multimodal model (LMM), KARINA extracts ingredient-level descriptions and qualitative nutrition profiles (e.g., "high fat", "low protein") from a frozen LMM such as GPT-4o, and injects this knowledge into an RGB-D visual model through cross-modal attention.

Demo

demo_v1.mp4

Highlights

  • LMM-derived nutritional knowledge. GPT-4o generates ingredient-level descriptions and qualitative nutrition profiles, which serve as weak semantic supervision for the visual model.
  • Multi-modal feature fusion. A visual feature fusion (VFF) module integrates RGB and depth features with channel/spatial attention, and a language-guided feature enhancement (LGFE) module fuses visual and textual features via bidirectional cross-attention.
  • Mask-based Ingredient-level Augmentation (MIA). Segmentation masks from Grounded SAM enable fine-grained alignment between ingredient regions and their textual descriptions.
  • State-of-the-art results. 13.3% mean PMAE on Nutrition5k (RGB-D), and mean MAE of 40.7 / 12.6 on NutritionVerse-Real / NutritionVerse-Synth.

Model Architecture

KARINA architecture

The framework consists of three components:

  1. Visual Feature Fusion (VFF) — two Swin-B encoders extract multi-scale features from the RGB image and the colorized depth map; features are fused per scale with channel and spatial attention.
  2. Language-Guided Feature Enhancement (LGFE) — LMM-generated ingredient descriptions are encoded by BERT (with special tokens delimiting each nutrition phrase) and fused with visual tokens through deformable self-attention and bidirectional cross-attention.
  3. Multi-task regression heads — five MLP heads predict calories, mass, fat, carbohydrates, and protein, trained with a PMAE loss.

Mask-based Ingredient-level Augmentation

Mask-based Ingredient-level Augmentation

With probability p, one ingredient triplet (annotation, description, mask) is sampled; the surrounding region of the RGB image is suppressed by alpha blending, and the dish-level targets are replaced by the ingredient-level ones, enabling fine-grained alignment between image regions and textual descriptions.

Installation

Tested with Python 3.10, PyTorch ≥ 2.0, and CUDA 11.8+.

git clone https://github.com/jyp-studio/KARINA.git
cd KARINA

pip install -r requirements.txt

# Build the deformable attention CUDA ops
cd models/GroundingDINO/ops
python setup.py build install
python test.py  # all checks should pass
cd ../../..

Download the pre-trained weights and place them under weights/:

  • GroundingDINO Swin-B (groundingdino_swinb_cogcoor.pth)
  • BERT-base-uncased (downloaded automatically by 🤗 Transformers, or specify a local path via --options text_encoder_type=/path/to/bert-base-uncased)

Datasets

  • Nutrition5k — RGB-D dish images with dish- and ingredient-level nutrition annotations.
  • NutritionVerse — Real and Synth subsets.

Update the dataset paths in config/datasets_nutri.json (and the NutritionVerse variants) to point to your local copies.

Preprocessing (optional, outputs can be cached)

  1. Generate ingredient descriptions with GPT-4o

    export OPENAI_API_KEY=<your key>
    python check_visible_ing.py --input <annotations.json> --output <out.json>

    The LMM identifies visible ingredients by selecting from a predefined ingredient vocabulary (open-vocabulary identification) and returns a short description plus qualitative nutrition profiles.

  2. Generate ingredient masks with Grounded SAM (for MIA)

    python generate_SAM_mask.py

Preprocessing v1 vs. v2

Two versions of each preprocessing script are provided. v1 reproduces the thesis and is frozen; v2 is an improved, experimental pipeline with the same input/output format, so the two are drop-in interchangeable downstream.

v1 (thesis) v2 (improved)
Scripts check_visible_ing.py, generate_SAM_mask.py check_visible_ing_v2.py, generate_SAM_mask_v2.py
Nutrition profiles GPT-4o guesses low/medium/high per dish Data-grounded lookup table: per-gram densities from the dataset's own ingredient annotations, bucketed by tertiles over the vocabulary — one consistent profile per ingredient; the LMM guess remains only as a fallback
LMM output parsing Regex extraction from free-form JSON OpenAI Structured Outputs (json_schema, strict, enum-constrained categories)
Reproducibility Default sampling temperature=0 + pinned model snapshot (gpt-4o-2024-08-06)
Identification approach the LMM selects visible ingredients from a predefined vocabulary compiled from Nutrition5k GT labels (check_visible_ing.py). Same open-vocabulary identification by default, plus an optional --use_gt_ingredients mode for the training split: the dish's GT ingredient names are given to the LMM, which only describes quantity/appearance and flags visibility
SAM prompting Box prompt only Box + point prompts: own box center as positive, other ingredients' centers as negatives (prevents bleed-over at the source)
Overlap resolution Binary mask subtraction with degenerate-case fallbacks Pixel-level soft assignment: contested pixels go to the detection with the highest score-weighted SAM probability
Mask verification Heuristics only Optional --clip_verify (zero-shot CLIP check of each masked crop) and --agent_verify (GPT-4o inspects a mask overlay, returns a correct/partial/wrong verdict plus corrective points that are fed back into SAM, up to --agent_rounds iterations)
Default models grounding-dino-tiny + sam-vit-base grounding-dino-base + sam-vit-huge
# v2 captioning: train split conditioned on GT ingredient names
python check_visible_ing_v2.py --input train.json --output train_desc.json --use_gt_ingredients
# v2 captioning: test split, open-vocabulary identification (LMM selects from the vocabulary)
python check_visible_ing_v2.py --input test.json --output test_desc.json

# v2 masks with cheap CLIP screening + LMM agent verification loop
python generate_SAM_mask_v2.py --input meta.json --output meta_masks.json \
    --clip_verify --agent_verify

The pure-logic components of v2 (soft overlap resolution, profile bucketing, resume) are covered by offline unit tests:

python tests/test_v2_offline.py

v2 has not been benchmarked end-to-end; before a full dataset run, spot-check 10–20 dishes by comparing v1 and v2 mask overlays side by side.

Training

bash train_dist.sh

Key options in config/cfg_nutri.py:

Parameter Description
backbone Visual encoder backbone (Swin-B by default)
text_encoder_type Language encoder (BERT-base-uncased)
epochs, lr, batch_size Standard training hyperparameters
mia_prob Probability p of applying Mask-based Ingredient-level Augmentation
mia_alpha Transparency coefficient α used in MIA blending
data_aug_scales Random-resize scales for data augmentation
use_text_enhancer Apply self-attention over text features
use_fusion_layer Enable cross-modal fusion between vision and language
exp_lr Use exponential learning-rate decay

Evaluation

bash test_dist.sh

With --save_results, per-dish predictions are saved as results-*.pkl in the output directory. Grad-CAM visualizations are available via tools/vis_cam.py and tools/rgbdnutrition_cam.py.

Results

Percentage of Mean Absolute Error (PMAE, lower is better) on Nutrition5k:

Input Cal ↓ Mass ↓ Fat ↓ Carbs ↓ Protein ↓ Mean ↓
RGB-D 11.4 8.6 17.1 15.1 14.1 13.3
RGB 12.1 10.2 18.6 16.2 14.7 14.4

Mean Absolute Error (MAE) on NutritionVerse:

Dataset Input Cal ↓ Mass ↓ Fat ↓ Carbs ↓ Protein ↓ Mean ↓
Real RGB 107.4 65.6 7.0 11.1 12.6 40.7
Synth RGB-D 36.2 18.4 1.7 3.0 3.6 12.6

Qualitative Results

Grad-CAM visualization of nutrient prediction attention maps

Grad-CAM heatmaps show that KARINA attends to the ingredients that dominate each nutrient (e.g., meats for protein and fat, fruits and starches for carbohydrates).

Citation

If you find this work useful, please cite:

@inproceedings{pan2026karina,
  author    = {Pan, Chieh-Yu and Chu, Wei-Ta},
  title     = {{KARINA: Knowledge-Augmented Representation for Ingredient-level Nutrition Analysis from Food Images}},
  booktitle = {2026 IEEE Conference on Artificial Intelligence (CAI)},
  year      = {2026},
  month     = {May},
  pages     = {1-6},
  publisher = {IEEE Computer Society},
  address   = {Los Alamitos, CA, USA},
  doi       = {10.1109/CAI68641.2026.11536297}
}

Acknowledgements

This codebase is built upon Open-GroundingDino and adapts code from:

We thank the authors for releasing their code.

License

Released under the Apache License 2.0. Portions of the code are derived from the projects listed above; their copyright and license notices are retained in the source files.