A ViT so400m image encoder from the SigLIP-v2 model by Tschannen et al., converted to the Birder format for image feature extraction. This version preserves the original model weights and architecture for downstream...
Modellquelle
Quellenauszug
A ViT so400m image encoder from the SigLIP-v2 model by Tschannen et al., converted to the Birder format for image feature extraction. This version preserves the original model weights and architecture for downstream...
Quellen
1 QuelleVerifiziert 5. Aug.
Modellartefakte
1 Artefaktvit_so400m_p14_ap_c1_siglip-v2-webli.pt
pt · 1,59 GB · SHA-256 f8ac3bdf028d…7be6 · Hugging Face
HerunterladenQuellenauszüge
2 Auszüge--- tags: - image-feature-extraction - birder - pytorch library_name: birder license: apache-2.0 base_model: - google/siglip2-so400m-patch14-224 --- # Model Card for vit_so400m_p14_ap_c1_siglip-v2-webli A ViT so400m image encoder from the SigLIP-v2 model by Tschannen et al., converted to the Birder format for image feature extraction. This version preserves the original model weights and architecture for downstream tasks. See: <https://huggingface.co/google/siglip2-so400m-patch14-224> for further details. ## Model Details - **Model Type:** Image classification and detection backbone - **Model Stats:** - Params (M): 427.7 - Input image size: 224 x 224 - **Papers:** - An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale: <https://arxiv.org/abs/2010.11929> - Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design: <https://arxiv.org/abs/2305.13035> - SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features: <https://arxiv.org/abs/2502.14786> ## Model Usage ### Image Embeddings ```python import birder from birder.inference.classification import infer_image (net, model_info) = birder.load_pretrained_model("vit_so400m_p14_ap_c1_siglip-v2-webli", inference=True) # Get the image size the model was trained on size = birder.get_size_from_signature(model_info.signature) # Create an inference transform transform = birder.classification_transform(size, model_info.rgb_stats) image = "path/to/image.jpeg" # or a PIL image (out, embedding) = infer_image(net, image, transform, return_embedding=True) # embedding is a NumPy array with shape of (1, 1152) ``` ### Detection Feature Map ```python from PIL import Image import birder (net, model_info) = birder.load_pretrained_model("vit_so400m_p14_ap_c1_siglip-v2-webli", inference=True) # Get the image size the model was trained on size = birder.get_size_from_signature(model_info.signature) # Create an inference transform transform = birder.classification_transform(size, model_info.rgb_stats) image = Image.open("path/to/image.jpeg") features = net.detection_features(transform(image).unsqueeze(0)) # features is a dict (stage name -> torch.Tensor) print([(k, v.size()) for k, v in features.items()]) # Output example: # [('neck', torch.Size([1, 1152, 16, 16]))] ``` ## Citation ```bibtex @misc{dosovitskiy2021imagewo...
Source context: 22 downloads · 0 likes · Pipeline image-feature-extraction · Library birder · Repo birder-project/vit_so400m_p14_ap_c1_siglip-v2-webli