Linear segmentation probe on the spatial features of facebook/dinov3-vits16-pretrain-lvd1689m.
Source du modèle
Description de la source
Linear segmentation probe on the spatial features of facebook/dinov3-vits16-pretrain-lvd1689m.
Sources
1 sourceVérifié 8 sept.
Artefacts du modèle
1 artefactExtraits de sources
2 extraitsuv add "canvit-pytorch @ git+https://github.com/m2b3/CanViT-PyTorch.git"
import torch
from canvit_pytorch.probes import SegmentationProbe
probe = SegmentationProbe.from_pretrained("canvit/probe-ade20k-40k-dv3s-144px").eval()
# [B, H, W, D] DINOv3 ViT-S/16 spatial features at 144px input
features = torch.randn(1, 9, 9, 384)
with torch.inference_mode():
logits = probe(features) # [B, num_classes, H, W]
assert logits.shape == (1, 150, 9, 9)
Architecture: Dropout → BatchNorm → Conv1×1.
| Hyperparameter | Value |
|---|---|
| Input size | 144 × 144 px |
| Optimizer | AdamW |
| Peak LR | \( 3 \times 10^{-4} \) |
| Weight decay | \( 10^{-3} \) |
| LR schedule | 1,500-step warmup → cosine decay |
| Batch size | 16 |
| Max steps | 40,000 |
| Dropout | 0.1 |
| Augmentation | RandomResizedCrop scale [0.5, 2] + HFlip |
| Precision | bf16 (AMP) |
--- license: mit library_name: canvit-pytorch pipeline_tag: image-segmentation tags: - dinov3 - ade20k - segmentation-probe datasets: - scene_parse_150 base_model: facebook/dinov3-vits16-pretrain-lvd1689m --- # ADE20K Segmentation Probe — DINOv3 ViT-S/16 @ 144px input Linear segmentation probe on the spatial features of [facebook/dinov3-vits16-pretrain-lvd1689m](https://huggingface.co/facebook/dinov3-vits16-pretrain-lvd1689m). - **Paper**: [arXiv:2603.22570](https://arxiv.org/abs/2603.22570) - **Training code**: [github.com/m2b3/CanViT-specialize](https://github.com/m2b3/CanViT-specialize) ## Usage ```bash uv add "canvit-pytorch @ git+https://github.com/m2b3/CanViT-PyTorch.git" ``` ```python import torch from canvit_pytorch.probes import SegmentationProbe probe = SegmentationProbe.from_pretrained("canvit/probe-ade20k-40k-dv3s-144px").eval() # [B, H, W, D] DINOv3 ViT-S/16 spatial features at 144px input features = torch.randn(1, 9, 9, 384) with torch.inference_mode(): logits = probe(features) # [B, num_classes, H, W] assert logits.shape == (1, 150, 9, 9) ``` ## Training Architecture: `Dropout → BatchNorm → Conv1×1`. | Hyperparameter | Value | |---|---| | Input size | 144 × 144 px | | Optimizer | AdamW | | Peak LR | \\( 3 \times 10^{-4} \\) | | Weight decay | \\( 10^{-3} \\) | | LR schedule | 1,500-step warmup → cosine decay | | Batch size | 16 | | Max steps | 40,000 | | Dropout | 0.1 | | Augmentation | RandomResizedCrop scale [0.5, 2] + HFlip | | Precision | bf16 (AMP) |
Source context: 10 downloads · 0 likes · Pipeline image-segmentation · Library canvit-pytorch · Repo canvit/probe-ade20k-40k-dv3s-144px