beforeat-food-nutrition-vision-lora is a LoRA adapter for food image understanding. This checkpoint is an R&D prototype toward broader food and nutrition recognition. The current version is trained and evaluated for...
Fonte do modelo
Descrição da fonte
beforeat-food-nutrition-vision-lora is a LoRA adapter for food image understanding. This checkpoint is an R&D prototype toward broader food and nutrition recognition. The current version is trained and evaluated for Food-101 dish-name recognition, not full nutrition estimation.
Fontes
1 fonteVerificado 11 de set.
Artefatos de modelo
3 artefatosTrechos de fonte
2 trechosbeforeat-food-nutrition-vision-loraQwen/Qwen3-VL-4B-InstructThis adapter is intended for research and prototype food recognition workflows where an image of a prepared dish is provided and the model returns a concise dish-name prediction.
Example prompt:
Identify the dish in this image. Reply with only the exact Food-101 class name.
For a production app, the recommended v1 deployment path is backend inference: the mobile app uploads a food image to a server, the server runs the base Qwen3-VL model with this LoRA adapter, and the app receives the predicted dish name.
This checkpoint should not be treated as:
The adapter currently recognizes dish names better than it estimates ingredients, calories, macros, or micronutrients.
This adapter was trained on Food-101 examples formatted as image-and-text instruction data. Food-101 is useful for dish-name research, but this checkpoint should be treated as an R&D prototype and downstream users should review the Food-101 dataset terms before any commercial or public deployment.
No Food-101 images are included in this adapter repository.
The adapter was trained locally with PyTorch and PEFT LoRA on Apple Silicon using MPS acceleration.
LoRA configuration:
q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_projTraining progression:
Evaluation was run on a fixed 500-image Food-101 sample.
| Adapter | Strict Accuracy | Alias/Near Accuracy |
|---|---|---|
| 1000-step Food-101 LoRA | 64.60% | 78.40% |
| 1500-step broad continuation | 79.00% | 84.80% |
| 2000-step broad continuation | 85.20% | 87.80% |
| 2500-step hard-only continuation | 81.80% | 82.80% |
| 2500-step mixed continuation, this adapter | 86.80% | 87.60% |
The 2500-step mixed adapter is the current best strict-label checkpoint. The 2000-step broad adapter remains a close fallback because it has slightly higher alias/near scoring.
Common remaining errors include visually similar food categories such as:
tuna tartare vs beef tartarechocolate cake vs chocolate moussedonuts vs beignetssteak, filet mignon, and prime ribbread pudding, panna cotta, and dessert-adjacent classesimport torch
from peft import PeftModel
from transformers import AutoProcessor, Qwen3VLForConditionalGeneration
base_model_id = "Qwen/Qwen3-VL-4B-Instruct"
adapter_id = "YOUR_USERNAME/beforeat-food-nutrition-vision-lora"
processor = AutoProcessor.from_pretrained(
base_model_id,
trust_remote_code=True,
)
model = Qwen3VLForConditionalGeneration.from_pretrained(
base_model_id,
torch_dtype=torch.float16,
device_map="auto",
trust_remote_code=True,
)
model = PeftModel.from_pretrained(model, adapter_id)
model.eval()
This repository also includes a standalone llama.cpp-compatible GGUF LoRA adapter for offline runtime loading:
beforeat-food-nutrition-vision-lora-f16.gguf33065184 bytes6fa0950b1bc4df4c1018b60ff4e0d824789bbc52a3764d1d38078861a0c2d8621.000fa7cb284cbf133fc426733bd64238a3588a33eRequired runtime files:
Qwen3-VL-4B-Instruct-UD-Q4_K_XL.ggufmmproj-F16.ggufunsloth/Qwen3-VL-4B-Instruct-GGUFConversion command:
PYTHONPATH=tools/llama.cpp \
python tools/llama.cpp/convert_lora_to_gguf.py \
adapters/qwen3vl4b_food101_peft_2500_mixed_hard20 \
--base models/Qwen3-VL-4B-Instruct-bf16-remapped \
--outfile adapters/qwen3vl4b_food101_peft_2500_mixed_hard20/beforeat-food-nutrition-vision-lora-f16.gguf \
--outtype f16 \
--verbose
Validation/runtime command:
tools/llama.cpp/build/bin/llama-mtmd-cli \
-m gguf_base/unsloth-Qwen3-VL-4B-Instruct-GGUF/Qwen3-VL-4B-Instruct-UD-Q4_K_XL.gguf \
--mmproj gguf_base/unsloth-Qwen3-VL-4B-Instruct-GGUF/mmproj-F16.gguf \
--image food-101/images/prime_rib/2421701.jpg \
--lora-scaled adapters/qwen3vl4b_food101_peft_2500_mixed_hard20/beforeat-food-nutrition-vision-lora-f16.gguf:1.0 \
-p "Identify the dish in this image. Reply with only the exact Food-101 class name." \
--jinja \
--temp 0 \
-n 16 \
-c 4096 \
--image-min-tokens 1024 \
--no-warmup \
--no-perf \
--log-verbosity 1
Runtime validation confirmed that llama.cpp accepts the adapter with the Unsloth UD-Q4_K_XL base and mmproj-F16.gguf, loading all 504 LoRA tensors with no missing, incompatible, or unexpected tensor errors.
Small hard-class validation slice:
| Expected | PyTorch PEFT | Base GGUF | GGUF + LoRA |
|---|---|---|---|
| bread pudding | panna cotta | dessert | bread pudding |
| filet mignon | beef fillet | Duck Confit | beef tenderloin |
| prime rib | prime rib | roast beef | prime rib |
| samosa | pork and cabbage roll | pie | samosa |
| fried calamari | fried calamari | fried_octopus | fried calamari |
| huevos rancheros | huevos rancheros | Egg | huevos rancheros |
| escargots | escargots | Fish | escargots |
On this intentionally hard 10-image slice:
Qwen3-VL-4B-Instruct-UD-Q4_K_XL.gguf and mmproj-F16.gguf; other base quantizations should be tested before app release.beforeat-food-nutrition-vision-lora-v2-f16.gguf
gguf · 31,5 MB · SHA-256 ac31fb420d91…1f62 · Hugging Face
--- base_model: Qwen/Qwen3-VL-4B-Instruct library_name: peft pipeline_tag: image-text-to-text tags: - qwen3-vl - vision-language - food-recognition - nutrition - peft - lora - beforeat --- # beforeat-food-nutrition-vision-lora `beforeat-food-nutrition-vision-lora` is a LoRA adapter for food image understanding. This checkpoint is an R&D prototype toward broader food and nutrition recognition. The current version is trained and evaluated for Food-101 dish-name recognition, not full nutrition estimation. ## Model Details - **Adapter name:** `beforeat-food-nutrition-vision-lora` - **Base model:** `Qwen/Qwen3-VL-4B-Instruct` - **Adapter type:** PEFT LoRA - **Task:** food image to dish-name text response - **Current training focus:** Food-101 dish classification via instruction-style visual question answering - **Developed for:** Beforeat food recognition research and prototyping ## Intended Use This adapter is intended for research and prototype food recognition workflows where an image of a prepared dish is provided and the model returns a concise dish-name prediction. Example prompt: ```text Identify the dish in this image. Reply with only the exact Food-101 class name. ``` For a production app, the recommended v1 deployment path is backend inference: the mobile app uploads a food image to a server, the server runs the base Qwen3-VL model with this LoRA adapter, and the app receives the predicted dish name. ## Out-of-Scope Use This checkpoint should not be treated as: - A complete nutrition estimator - A medical, dietary, or allergy safety tool - A reliable portion-size estimator - A production-grade model for all cuisines, restaurant conditions, or user-generated food photos The adapter currently recognizes dish names better than it estimates ingredients, calories, macros, or micronutrients. ## Training Data This adapter was trained on Food-101 examples formatted as image-and-text instruction data. Food-101 is useful for dish-name research, but this checkpoint should be treated as an R&D prototype and downstream users should review the Food-101 dataset terms before any commercial or public deployment. No Food-101 images are included in this adapter repository. ## Training Procedure The adapter was trained locally with PyTorch and PEFT LoRA on Apple Silicon using MPS acceleration. LoRA configuration: - **Rank:** 8 - **Alpha:** 16 - **Dro...
Source context: 57 downloads · 1 likes · Pipeline image-text-to-text · Library peft · Repo fatsam13/beforeat-food-nutrition-vision-lora
| foie gras | beef chop | duck | beef chop |
| croque madame | croque madame | sandwich | croque madame |
| lobster bisque | lobster bisque | soup | cheese soup |