A standalone ONNX vision projector extracted from the FastVLM pipeline. Converts image features into the FastVLM multimodal embedding space and is designed for CPU-based image-to-text and vision+language inference.
Source du modèle
Description de la source
A standalone ONNX vision projector extracted from the FastVLM pipeline. Converts image features into the FastVLM multimodal embedding space and is designed for CPU-based image-to-text and vision+language inference.
vision_projector_v1_standalone.onnxUse this model to convert images into embeddings for a compatible FastVLM language backbone.
Sources
1 sourceVérifié 1 août
Artefacts du modèle
1 artefactvision_projector_v1_standalone.onnx
onnx · 484 MB · SHA-256 0af5d1fb8dc1…7faf · Hugging Face
TéléchargerExtraits de sources
2 extraits--- license: mit language: - en pipeline_tag: image-feature-extraction tags: - onnx - vision - image-encoder - multimodal - fastvlm - cpu base_model: - apple/FastVLM-0.5B --- # Apple FastVLM Vision Projector (512) A standalone ONNX vision projector extracted from the FastVLM pipeline. Converts image features into the FastVLM multimodal embedding space and is designed for CPU-based image-to-text and vision+language inference. ## Model file - `vision_projector_v1_standalone.onnx` ## Usage Use this model to convert images into embeddings for a compatible FastVLM language backbone. ```python import onnxruntime as ort session = ort.InferenceSession("vision_projector_v1_standalone.onnx") # feed in your preprocessed image tensor and run inference ``` ## Base model Derived from [apple/FastVLM-0.5B](https://huggingface.co/apple/FastVLM-0.5B).
import onnxruntime as ort
session = ort.InferenceSession("vision_projector_v1_standalone.onnx")
# feed in your preprocessed image tensor and run inference
Derived from apple/FastVLM-0.5B.
Source context: 0 downloads · 0 likes · Pipeline image-feature-extraction · Repo musk12/apple-fastvlm-vision-projector-512