μ2Qwen3-4B-Instruct is a multi-scale multi-modal model designed for Radiology Report Generation (RRG). It leverages the μ2Tokenizer, a differentiable intermediate layer that efficiently fuses visual features from 3D...
Modellquelle
Quellenauszug
μ2Qwen3-4B-Instruct is a multi-scale multi-modal model designed for Radiology Report Generation (RRG). It leverages the μ2Tokenizer, a differentiable intermediate layer that efficiently fuses visual features from 3D...
Quellen
1 QuelleVerifiziert 10. Sept.
Modellartefakte
1 ArtefaktQuellenauszüge
2 Auszüge--- license: apache-2.0 language: - en - zh tags: - medical - radiology - multimodal - image-to-text - medical-imaging pipeline_tag: image-to-text --- [](https://arxiv.org/abs/2507.00316) [](https://alpachino.org/proj/u2tokenizer/) [](https://github.com/Siyou-Li/u2Tokenizer) ## Model Description  ## Model Summary **μ²Qwen3-4B-Instruct** is a multi-scale multi-modal model designed for **Radiology Report Generation (RRG)**. It leverages the **μ²Tokenizer**, a differentiable intermediate layer that efficiently fuses visual features from 3D CT scans with textual information. The model is fine-tuned using Direct Preference Optimization (DPO) guided by the GREEN score to ensure clinical accuracy and alignment with expert standards. This model is part of the work described in the paper: [μ²Tokenizer: Differentiable Multi-Scale Multi-Modal Tokenizer for Radiology Report Generation](https://arxiv.org/abs/2507.00316). ## Model Details - **Model Architecture:** - **Image Encoder:** 3D Vision Transformer (ViT3D) initialized from M3D-CLIP. - **Tokenizer:** μ²Tokenizer (Multi-scale, Multi-modal). - **LLM Backbone:** Qwen3-4B-Instruct. - **Input:** 3D CT Scans (NIfTI format) and text prompts. - **Output:** Radiology reports or answers to clinical questions. - **Training Data:** Trained on large-scale CT image-report Synthetic datasets base on CT-RATE. ## How to Get Started with the Model ### Requirements ```bash pip install torch transformers monai==1.3.2 nibabel==5.3.3 ``` ### Inference Code Below is a sample code snippet to generate a report from a CT scan using this model. ```python import torch import nibabel as nib import numpy as np from transformers import AutoModelForCausalLM, AutoTokenizer from monai.transforms import ( Compose, CropForeground, ToTensor, ScaleIntensityRangePercentiles ) from monai.transforms.spatial.functional import resize import torch.nn.functional as F import types import inspect class u2Transform: def __init__(self, mode='bilinear', device="cpu"): self.adaptive_transforms = Compose([ ScaleIn...
Source context: 13 downloads · 1 likes · Pipeline image-to-text · Repo AlpachinoNLP/u2Qwen3-4B-Instruct