nvidia-open-model-license
Fonte do modelo
Trecho da fonte
nvidia-open-model-license
Fontes
1 fonteVerificado 18 de ago.
Artefatos de modelo
1 artefatoTrechos de fonte
2 trechos--- library_name: transformers license: other license_name: nvidia-open-model-license license_link: https://developer.download.nvidia.com/licenses/nvidia-open-model-license-agreement-june-2024.pdf pipeline_tag: image-feature-extraction --- # Model Overview [[**Github**](https://github.com/NVlabs/RADIO)] [[**CVPR 2025**](https://arxiv.org/abs/2412.07679)] [[**CVPR 2024**](https://arxiv.org/abs/2312.06709)] ## Description This model performs visual feature extraction. For instance, RADIO generates image embeddings that can be used by a downstream model to classify images. C-RADIOv2 models are available in multiple sizes: * Base (90M parameters). * Large (320M parameters). * Huge (653M parameters). * Gigantic (1.1B parameters). C-RADIOv2 was trained for 1M steps (400k more steps than v1), using inverse frequency sampling for data balancing, and [PHI Standardization](https://arxiv.org/abs/2410.01680) for teacher distribution balancing. This model is ready for commercial/non-commercial use. ### License/Terms of Use GOVERNING TERMS: Use of this model is governed by the [NVIDIA Open Model License Agreement](https://developer.download.nvidia.com/licenses/nvidia-open-model-license-agreement-june-2024.pdf). ## Deployment Geography Global. ## Use Case The embeddings generated by this model are expected to be used by a downstream application. For example: * Image-level understanding (image classification, curation, etc.). * Dense processing (semantic segmentation, depth estimation, etc.). * Integration into a Vision-Language Model. ## Release Date Huggingface: 03/26/2025 via [RADIO Collection of Models](https://huggingface.co/collections/nvidia/radio-669f77f1dd6b153f007dd1c6). ## References * \[CVPR 2025\] [**RADIOv2.5: Improved Baselines for Agglomerative Vision Foundation Models**](https://arxiv.org/abs/2412.07679) * \[CVPR 2024\] [**AM-RADIO: Agglomerative Vision Foundation Model - Reduce All Domains Into One**](https://arxiv.org/abs/2312.06709) ## Model Architecture **Architecture Type:** Neural Network <br> **Network Architecture:** Vision Transformer <br> ## Input **Input Type(s):** Image <br> **Input Format(s):** Red, Green, Blue (RGB) <br> **Input Parameters:** Two Dimensional (2D) <br> **Other Properties Related to Input:** Image resolutions up to 2048x2028 in increments of 16 pixels <br> ## Output **Output Type(s):** Embeddings <br>...
Source context: 380 downloads · 11 likes · Pipeline image-feature-extraction · Library transformers · Repo nvidia/C-RADIOv2-B