dinov3
Source du modèle
Extrait de la source
dinov3
Sources
1 sourceVérifié 18 août
Artefacts du modèle
1 artefactExtraits de sources
2 extraits--- library_name: mlx-image license: other license_name: dinov3 license_link: https://ai.meta.com/resources/models-and-libraries/dinov3-license/ tags: - mlx - mlx-image - vision - dinov3 - image-feature-extraction pipeline_tag: image-feature-extraction base_model: - facebook/dinov3-vit7b16-pretrain-lvd1689m --- # vit_small_plus_patch16_224.dinov3 A [Vision Transformer](https://arxiv.org/abs/2010.11929v2) feature extraction model trained on the LVD-1689M web dataset with [DINOv3](https://arxiv.org/abs/2508.10104). The model was trained in self-supervised fashion. No classification head was trained, only the backbone. This is the **ViT-S+/16** variant (29M parameters), distilled from the DINOv3 ViT-7B teacher model. The S+ variant upgrades the standard ViT-S with a **SwiGLU feed-forward network** instead of a standard MLP, giving it stronger representational capacity over ViT-S with only a modest parameter increase. Disclaimer: This is a porting of the Meta AI DINOv3 model weights to Apple MLX Framework. <!-- <div align="center"> <img width="100%" alt="DINO illustration" src="dino.gif"> </div> --> ## How to use ```bash pip install mlx-image ``` Here is how to use this model for feature extraction: ```python import mlx.core as mx from mlxim.model import create_model from mlxim.io import read_rgb from mlxim.transform import ImageNetTransform transform = ImageNetTransform(train=False, img_size=224) x = transform(read_rgb("image.png")) x = mx.expand_dims(x, 0) model = create_model("vit_small_plus_patch16_224.dinov3") model.eval() embeds = model(x, is_training=False) ``` You can also retrieve embeddings from the layer before the head: ```python import mlx.core as mx from mlxim.model import create_model from mlxim.io import read_rgb from mlxim.transform import ImageNetTransform transform = ImageNetTransform(train=False, img_size=224) x = transform(read_rgb("image.png")) x = mx.expand_dims(x, 0) model = create_model("vit_small_plus_patch16_224.dinov3", num_classes=0) model.eval() embeds = model(x, is_training=False) ``` ## Architecture This model follows the ViT architecture with a patch size of 16. For a 224×224 image this results in **1 class token + 4 register tokens + 196 patch tokens = 201 tokens**. The model can accept larger images provided the image shapes are multiples of the patch size (16). If this condition is not met, the model will...
Source context: 275 downloads · 0 likes · Pipeline image-feature-extraction · Library mlx-image · Repo mlx-vision/vit_small_plus_patch16_224.dinov3-mlxim