[\[📃 Tech Report\]]( [\[📂 Github\]](
Source du modèle
Extrait de la source
[[📃 Tech Report]]( [[📂 Github]](
Sources
1 sourceVérifié 2 sept.
Artefacts du modèle
1 artefactExtraits de sources
2 extraits--- license: apache-2.0 library_name: perception-encoder pipeline_tag: image-feature-extraction --- # Model Details [\[📃 Tech Report\]](https://arxiv.org/abs/2504.13181) [\[📂 Github\]](https://github.com/facebookresearch/perception_models/) Perception Encoder (PE) is a state-of-the-art encoder for image and video understanding trained via simple vision-language learning. It was introduced in "[Perception Encoder: The best visual embeddings are not at the output of the network](https://ai.meta.com/research/publications/perception-encoder-the-best-visual-embeddings-are-not-at-the-output-of-the-network/)". **Model Developer**: Meta **Model Overview**: Perception Encoder (PE) is a family of large-scale vision encoder models with state-of-the-art performance on a large variety of vision tasks. By using a robust contrastive pretraining recipe and finetuning on synthetically aligned videos, PE not only outperforms all existing models on classification and retrieval, but it also internally produces strong, general features that scale for downstream tasks. PE unlocks the ability for large-scale contrastive pretraining to transfer to downstream tasks with alignment tuning to capitalize on those general features. <img src="https://huggingface.co/facebook/PE-Core-G14-448/resolve/main/docs/pe_image1.png" style="width: 100%; margin: 0 auto; display: block;" /> ## Perception Encoder: Core PE core is our base model trained with our robust image pretraining schedule and finetuned on the data generated by our synthetic video data engine. #### Model Configurations PE core curently comes in 3 sizes. PE core G is the main checkpoint, with L and B models distilled from it. | Scale | Tower | Params | Width | Depth | MLP | Heads | CLIP Dim | Resolution / Context Len | |:-----:|:------:|:------:|:-----:|:-----:|:----:|:-----:|:--------:|:-------------------------:| | **B/16** | Vision | 0.09B | 768 | 12 | 3072 | 12 | 1024 | 224px | | | Text | 0.31B | 1024 | 24 | 4096 | 16 | 1024 | 32 tokens | | **L/14** | Vision | 0.32B | 1024 | 24 | 4096 | 16 | 1024 | 336px | | | Text | 0.31B | 1024 | 24 | 4096 | 16 | 1024 | 32 tokens | | **G/14** | Vision | 1.88B | 1536 | 50 | 8960 | 16 | 1280 | 448px | | | Text | 0.47B | 1280 | 24 | 5120 | 20 | 1280 | 72 tokens | All PE core models use an attention pooling block with 8 heads on top of the vision tower. The L and B models _additional...
Source context: 359 downloads · 0 likes · Pipeline image-feature-extraction · Library perception-encoder · Repo facebook/PE-Spatial-L14-448