Official implementation of "KARINA: Knowledge-Augmented Representation for Ingredient-level Nutrition Analysis from Food Images", published at IEEE CAI 2026 (IEEE Conference on Artificial Intelligence). Developed as part of a master's thesis at National Cheng Kung University.
KARINA estimates five nutritional values (calories, mass, fat, carbohydrates, and protein) from food images. Instead of directly predicting nutrition with a large multimodal model (LMM), KARINA extracts ingredient-level descriptions and qualitative nutrition profiles (e.g., "high fat", "low protein") from a frozen LMM such as GPT-4o, and injects this knowledge into an RGB-D visual model through cross-modal attention.
demo_v1.mp4
- LMM-derived nutritional knowledge. GPT-4o generates ingredient-level descriptions and qualitative nutrition profiles, which serve as weak semantic supervision for the visual model.
- Multi-modal feature fusion. A visual feature fusion (VFF) module integrates RGB and depth features with channel/spatial attention, and a language-guided feature enhancement (LGFE) module fuses visual and textual features via bidirectional cross-attention.
- Mask-based Ingredient-level Augmentation (MIA). Segmentation masks from Grounded SAM enable fine-grained alignment between ingredient regions and their textual descriptions.
- State-of-the-art results. 13.3% mean PMAE on Nutrition5k (RGB-D), and mean MAE of 40.7 / 12.6 on NutritionVerse-Real / NutritionVerse-Synth.
The framework consists of three components:
- Visual Feature Fusion (VFF) — two Swin-B encoders extract multi-scale features from the RGB image and the colorized depth map; features are fused per scale with channel and spatial attention.
- Language-Guided Feature Enhancement (LGFE) — LMM-generated ingredient descriptions are encoded by BERT (with special tokens delimiting each nutrition phrase) and fused with visual tokens through deformable self-attention and bidirectional cross-attention.
- Multi-task regression heads — five MLP heads predict calories, mass, fat, carbohydrates, and protein, trained with a PMAE loss.
With probability p, one ingredient triplet (annotation, description, mask) is sampled; the surrounding region of the RGB image is suppressed by alpha blending, and the dish-level targets are replaced by the ingredient-level ones, enabling fine-grained alignment between image regions and textual descriptions.
Tested with Python 3.10, PyTorch ≥ 2.0, and CUDA 11.8+.
git clone https://github.com/jyp-studio/KARINA.git
cd KARINA
pip install -r requirements.txt
# Build the deformable attention CUDA ops
cd models/GroundingDINO/ops
python setup.py build install
python test.py # all checks should pass
cd ../../..Download the pre-trained weights and place them under weights/:
- GroundingDINO Swin-B (
groundingdino_swinb_cogcoor.pth) - BERT-base-uncased (downloaded automatically by 🤗 Transformers, or specify a local path via
--options text_encoder_type=/path/to/bert-base-uncased)
- Nutrition5k — RGB-D dish images with dish- and ingredient-level nutrition annotations.
- NutritionVerse — Real and Synth subsets.
Update the dataset paths in config/datasets_nutri.json (and the NutritionVerse variants) to point to your local copies.
-
Generate ingredient descriptions with GPT-4o
export OPENAI_API_KEY=<your key> python check_visible_ing.py --input <annotations.json> --output <out.json>
The LMM identifies visible ingredients by selecting from a predefined ingredient vocabulary (open-vocabulary identification) and returns a short description plus qualitative nutrition profiles.
-
Generate ingredient masks with Grounded SAM (for MIA)
python generate_SAM_mask.py
Two versions of each preprocessing script are provided. v1 reproduces the thesis and is frozen; v2 is an improved, experimental pipeline with the same input/output format, so the two are drop-in interchangeable downstream.
| v1 (thesis) | v2 (improved) | |
|---|---|---|
| Scripts | check_visible_ing.py, generate_SAM_mask.py |
check_visible_ing_v2.py, generate_SAM_mask_v2.py |
| Nutrition profiles | GPT-4o guesses low/medium/high per dish | Data-grounded lookup table: per-gram densities from the dataset's own ingredient annotations, bucketed by tertiles over the vocabulary — one consistent profile per ingredient; the LMM guess remains only as a fallback |
| LMM output parsing | Regex extraction from free-form JSON | OpenAI Structured Outputs (json_schema, strict, enum-constrained categories) |
| Reproducibility | Default sampling | temperature=0 + pinned model snapshot (gpt-4o-2024-08-06) |
| Identification approach | the LMM selects visible ingredients from a predefined vocabulary compiled from Nutrition5k GT labels (check_visible_ing.py). |
Same open-vocabulary identification by default, plus an optional --use_gt_ingredients mode for the training split: the dish's GT ingredient names are given to the LMM, which only describes quantity/appearance and flags visibility |
| SAM prompting | Box prompt only | Box + point prompts: own box center as positive, other ingredients' centers as negatives (prevents bleed-over at the source) |
| Overlap resolution | Binary mask subtraction with degenerate-case fallbacks | Pixel-level soft assignment: contested pixels go to the detection with the highest score-weighted SAM probability |
| Mask verification | Heuristics only | Optional --clip_verify (zero-shot CLIP check of each masked crop) and --agent_verify (GPT-4o inspects a mask overlay, returns a correct/partial/wrong verdict plus corrective points that are fed back into SAM, up to --agent_rounds iterations) |
| Default models | grounding-dino-tiny + sam-vit-base | grounding-dino-base + sam-vit-huge |
# v2 captioning: train split conditioned on GT ingredient names
python check_visible_ing_v2.py --input train.json --output train_desc.json --use_gt_ingredients
# v2 captioning: test split, open-vocabulary identification (LMM selects from the vocabulary)
python check_visible_ing_v2.py --input test.json --output test_desc.json
# v2 masks with cheap CLIP screening + LMM agent verification loop
python generate_SAM_mask_v2.py --input meta.json --output meta_masks.json \
--clip_verify --agent_verifyThe pure-logic components of v2 (soft overlap resolution, profile bucketing, resume) are covered by offline unit tests:
python tests/test_v2_offline.pyv2 has not been benchmarked end-to-end; before a full dataset run, spot-check 10–20 dishes by comparing v1 and v2 mask overlays side by side.
bash train_dist.shKey options in config/cfg_nutri.py:
| Parameter | Description |
|---|---|
backbone |
Visual encoder backbone (Swin-B by default) |
text_encoder_type |
Language encoder (BERT-base-uncased) |
epochs, lr, batch_size |
Standard training hyperparameters |
mia_prob |
Probability p of applying Mask-based Ingredient-level Augmentation |
mia_alpha |
Transparency coefficient α used in MIA blending |
data_aug_scales |
Random-resize scales for data augmentation |
use_text_enhancer |
Apply self-attention over text features |
use_fusion_layer |
Enable cross-modal fusion between vision and language |
exp_lr |
Use exponential learning-rate decay |
bash test_dist.shWith --save_results, per-dish predictions are saved as results-*.pkl in the output directory. Grad-CAM visualizations are available via tools/vis_cam.py and tools/rgbdnutrition_cam.py.
Percentage of Mean Absolute Error (PMAE, lower is better) on Nutrition5k:
| Input | Cal ↓ | Mass ↓ | Fat ↓ | Carbs ↓ | Protein ↓ | Mean ↓ |
|---|---|---|---|---|---|---|
| RGB-D | 11.4 | 8.6 | 17.1 | 15.1 | 14.1 | 13.3 |
| RGB | 12.1 | 10.2 | 18.6 | 16.2 | 14.7 | 14.4 |
Mean Absolute Error (MAE) on NutritionVerse:
| Dataset | Input | Cal ↓ | Mass ↓ | Fat ↓ | Carbs ↓ | Protein ↓ | Mean ↓ |
|---|---|---|---|---|---|---|---|
| Real | RGB | 107.4 | 65.6 | 7.0 | 11.1 | 12.6 | 40.7 |
| Synth | RGB-D | 36.2 | 18.4 | 1.7 | 3.0 | 3.6 | 12.6 |
Grad-CAM heatmaps show that KARINA attends to the ingredients that dominate each nutrient (e.g., meats for protein and fat, fruits and starches for carbohydrates).
If you find this work useful, please cite:
@inproceedings{pan2026karina,
author = {Pan, Chieh-Yu and Chu, Wei-Ta},
title = {{KARINA: Knowledge-Augmented Representation for Ingredient-level Nutrition Analysis from Food Images}},
booktitle = {2026 IEEE Conference on Artificial Intelligence (CAI)},
year = {2026},
month = {May},
pages = {1-6},
publisher = {IEEE Computer Society},
address = {Los Alamitos, CA, USA},
doi = {10.1109/CAI68641.2026.11536297}
}This codebase is built upon Open-GroundingDino and adapts code from:
- IDEA-Research/GroundingDINO
- IDEA-Research/DINO
- microsoft/GLIP
- IDEA-Research/Grounded-Segment-Anything
We thank the authors for releasing their code.
Released under the Apache License 2.0. Portions of the code are derived from the projects listed above; their copyright and license notices are retained in the source files.


