對 Qwen3-VL-8B-Instruct 做 QLoRA 微調的 LoRA adapter,任務是把收據照片直接轉成固定 schema 的結構化 JSON(品項、單價、數量、小計、折扣、服務費、稅金、總額)。
Source du modèle
Extrait de la source
對 Qwen3-VL-8B-Instruct 做 QLoRA 微調的 LoRA adapter,任務是把收據照片直接轉成固定 schema 的結構化 JSON(品項、單價、數量、小計、折扣、服務費、稅金、總額)。
Sources
1 sourceVérifié 11 sept.
Artefacts du modèle
1 artefactExtraits de sources
2 extraits--- license: apache-2.0 base_model: unsloth/Qwen3-VL-8B-Instruct-unsloth-bnb-4bit datasets: - naver-clova-ix/cord-v2 language: - id - en library_name: peft pipeline_tag: image-text-to-text tags: - lora - qlora - unsloth - vision-language-model - document-ai - receipt - qwen3-vl --- # vlm-receipt-extractor:Qwen3-VL-8B QLoRA(收據影像 → 結構化 JSON) 對 [Qwen3-VL-8B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-8B-Instruct) 做 QLoRA 微調的 LoRA adapter,任務是把收據照片直接轉成固定 schema 的結構化 JSON(品項、單價、數量、小計、折扣、服務費、稅金、總額)。 ## 任務說明 輸入一張收據照片,輸出: ```json { "items": [ {"name": "Nasi Goreng", "count": 2, "unit_price": 25000, "price": 50000} ], "subtotal": 50000, "discount": null, "service": null, "tax": 5000, "total": 55000 } ``` Schema 完整對齊 [CORD-v2](https://huggingface.co/datasets/naver-clova-ix/cord-v2) 的 `gt_parse` 標註(不含店名/日期,因為原始標註沒有這兩欄);收據上沒有印的欄位一律輸出 `null`,金額為純數值(已去除千分位符號)。 ## Foundation Model 與訓練設定 - Foundation Model:`unsloth/Qwen3-VL-8B-Instruct-unsloth-bnb-4bit`(Unsloth 預量化的 4-bit 版本,vision tower 維持 bf16) - 方法:Unsloth QLoRA,`finetune_vision_layers=True`、`finetune_language_layers=True`、`finetune_attention_modules=True`、`finetune_mlp_modules=True` - LoRA:r=16、alpha=16、dropout=0、bias=none - 訓練:2 epochs、batch size 2 × gradient accumulation 4(有效 batch size 8)、lr=2e-4(linear schedule)、seed=3407 - 資料:CORD-v2 官方 train split(800 筆),影像最長邊縮至 1280px - 硬體:單張 RTX 4090(24GB),實際訓練耗時 19.2 分鐘、200 steps,峰值 VRAM 10.27GB - 驗證:epoch 1 eval_loss=0.0086 → epoch 2 eval_loss=0.0078(持續下降,無過擬合跡象) ## 微調前後對照(CORD-v2 test split,100 筆,greedy decoding、相同 prompt) | 指標 | 微調前(zero-shot) | 微調後 | |---|---|---| | 合法 JSON 率 | 100.0% | 100.0% | | total exact match | 75.0% | **90.0%** | | 欄位級 micro F1 | 0.744 | **0.930** | 各欄位 F1 進步最多的:`items.unit_price`(0.32→0.91)、`discount`(0.27→0.86)、`service`(0.57→0.92)。 微調前的 zero-shot 模型有兩個系統性問題:(1) 把收據上印尼盾金額的千分位句點(如 `25.000` = 兩萬五千)誤讀成小數點,輸出成 `25.0`;(2) 會自行「腦補」收據上沒印的 unit_price,或把品項名稱合併/截斷。微調後這兩類錯誤大幅減少。完整對照表與案例展示見 repo 內的 `results/comparison.md`。 ## 使用方式 ```python import unsloth # 必須放在最前面:修正 transformers/bitsandbytes 對 Qwen3-VL vision tower 的一個載入 bug,即使下面完全不用 FastVisionModel 也需要 from transformers import Qwen3VLForConditionalGeneration, AutoProcessor from peft import PeftModel from PIL import Image import torch BASE = "unsloth/Qwen3-VL-8B-Instruct-unsloth-bnb-4bit" model = Qwen3VLForConditionalGeneration.from_pretrained(BASE, device_map="cuda") # 不要額外傳 quant...
Source context: 13 downloads · 0 likes · Pipeline image-text-to-text · Library peft · Repo steven0226/vlm-receipt-extractor