Aloe-Vision is a medical Large Vision–Language Model built on Qwen2-VL-Instruct, released in 7B and 72B sizes. The model is trained on a \3.5 M samples balanced mixture across medical vs. general and multimodal vs....
Modellquelle
Quellenauszug
Aloe-Vision is a medical Large Vision–Language Model built on Qwen2-VL-Instruct, released in 7B and 72B sizes. The model is trained on a \3.5 M samples balanced mixture across medical vs. general and multimodal vs....
Quellen
1 QuelleVerifiziert 1. Aug.
Modellartefakte
4 Artefaktemodel-00001-of-00004.safetensors
safetensors · 4,63 GB · SHA-256 670c27a94731…b957 · Hugging Face
Herunterladenmodel-00002-of-00004.safetensors
safetensors · 4,65 GB · SHA-256 7fc73d630536…ba6c · Hugging Face
HerunterladenQuellenauszüge
2 Auszügemodel-00003-of-00004.safetensors
safetensors · 4,59 GB · SHA-256 e9492db14301…ee53 · Hugging Face
model-00004-of-00004.safetensors
safetensors · 1,58 GB · SHA-256 fbff96f3c985…362a · Hugging Face
Herunterladen--- license: cc-by-nc-sa-4.0 library\_name: transformers language: - en pipeline\_tag: image-text-to-text tags: - visual-question-answering - medical - healthcare - biology - multimodal - lvlm - grounding datasets: HPAI-BSC/Aloe-Beta-General-Collection model\_type: qwen2-vl base\_model: Qwen/Qwen2-VL-7B-Instruct --- <p align="center"> <img alt="Aloe-Vision" src="https://cdn-uploads.huggingface.co/production/uploads/63a417e70cf4daf6166777a2/xkm30vCSIz1GK__K3QIQZ.png" width="25%"> </p> </h1> <hr style="margin: 15px"> <div align="center" style="line-height:1.15;"> <a href="https://huggingface.co/datasets/HPAI-BSC/Aloe-Vision-Data" target="_blank" style="margin:2px;"> <img alt="Training Dataset" src="https://img.shields.io/badge/🤗%20Training%20Dataset-Aloe%20Vision%20Data-ffc107" style="vertical-align:middle;"/> </a> <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/deed.en" target="_blank" style="margin:2px;"> <img alt="License" src="https://img.shields.io/badge/license-CC_BY--NC--SA_4.0-green" style="vertical-align:middle;"/> </a> <br/> <a href="https://hpai.bsc.es/" target="_blank" style="margin:2px;"> <img alt="Website" src="https://img.shields.io/badge/Website-HPAI-8A2BE2" style="vertical-align:middle;"/> </a> <a href="https://huggingface.co/HPAI-BSC" target="_blank" style="margin:2px;"> <img alt="Hugging Face Org" src="https://img.shields.io/badge/🤗%20HF-HPAI--BSC-ffc107" style="vertical-align:middle;"/> </a> <a href="https://github.com/HPAI-BSC" target="_blank" style="margin:2px;"> <img alt="GitHub" src="https://img.shields.io/badge/GitHub-HPAI--BSC-%23121011.svg" style="vertical-align:middle;"/> </a> </div> </div> **Aloe-Vision** is a **medical Large Vision–Language Model** built on **Qwen2-VL-Instruct**, released in **7B and 72B** sizes. The model is trained on a **\~3.5 M samples** balanced mixture across **medical vs. general** and **multimodal vs. text-only** sources, rebalanced by **loss-contributing assistant tokens** to avoid long-answer bias. We implement **leakage control of evaluation images in the training data** via **exact 64-bit image-hash matching**, removing any duplicates from the training. **Quality filtering** of the training data combines (1) **LVLM-based sample scoring (1–5 scale)** for image–question–answer coherence and relevance and (2) **answer perplexity checks** to flag trivial or noisy...
Source context: 90 downloads · 1 likes · Pipeline visual-question-answering · Repo HPAI-BSC/Aloe-Vision-7B-AR