> Two interns holding hands, symbolizing the integration of InternViT and InternLM.
Fonte do modelo
Trecho da fonte
Two interns holding hands, symbolizing the integration of InternViT and InternLM.
Fontes
1 fonteVerificado 28 de ago.
Artefatos de modelo
6 artefatosTrechos de fonte
2 trechosmodel-00003-of-00006.safetensors
safetensors · 4,64 GB · SHA-256 130ceb8a965d…99c1 · Hugging Face
model-00004-of-00006.safetensors
safetensors · 4,63 GB · SHA-256 371e350a0244…475b · Hugging Face
Baixarmodel-00005-of-00006.safetensors
safetensors · 4,63 GB · SHA-256 27112b385014…26ca · Hugging Face
Baixarmodel-00006-of-00006.safetensors
safetensors · 1,22 GB · SHA-256 020446149369…cc6d · Hugging Face
Baixar--- license: mit datasets: - laion/laion2B-en - laion/laion-coco - laion/laion2B-multi - kakaobrain/coyo-700m - conceptual_captions - wanng/wukong100m pipeline_tag: visual-question-answering --- # Model Card for InternVL-Chat-V1.5-Int8 <p align="center"> <img src="https://cdn-uploads.huggingface.co/production/uploads/64119264f0f81eb569e0d569/D60YzQBIzvoCvLRp2gZ0A.jpeg" alt="Image Description" width="300" height="300" /> </p> > _Two interns holding hands, symbolizing the integration of InternViT and InternLM._ \[[InternVL 1.5 Technical Report](https://arxiv.org/abs/2404.16821)\] \[[Paper](https://arxiv.org/abs/2312.14238)\] \[[GitHub](https://github.com/OpenGVLab/InternVL)\] \[[Chat Demo](https://internvl.opengvlab.com/)\] \[[中文解读](https://zhuanlan.zhihu.com/p/675877376)] We introduce InternVL 1.5, an open-source multimodal large language model (MLLM) to bridge the capability gap between open-source and proprietary commercial models in multimodal understanding. We introduce three simple designs: 1. Strong Vision Encoder: we explored a continuous learning strategy for the large-scale vision foundation model---InternViT-6B, boosting its visual understanding capabilities, and making it can be transferred and reused in different LLMs. 2. Dynamic High-Resolution: we divide images into tiles ranging from 1 to 40 of 448 × 448 pixels according to the aspect ratio and resolution of the input images, which supports up to 4K resolution input. 3. High-Quality Bilingual Dataset: we carefully collected a high-quality bilingual dataset that covers common scenes, document images, and annotated them with English and Chinese question-answer pairs, significantly enhancing performance in OCR- and Chinese-related tasks. ## Model Details - **Model Type:** multimodal large language model (MLLM) - **Model Stats:** - Architecture: [InternViT-6B-448px-V1-5](https://huggingface.co/OpenGVLab/InternViT-6B-448px-V1-5) + MLP + [InternLM2-Chat-20B](https://huggingface.co/internlm/internlm2-chat-20b) - Image size: dynamic resolution, max to 40 tiles of 448 x 448 (4K resolution). - Params: 25.5B - **Training Strategy:** - Pretraining Stage - Learnable Component: ViT + MLP - Data: Please see our technical report. - SFT Stage - Learnable Component: ViT + MLP + LLM - Data: Please see our technical report. ## Released Models | Model | Vision Foundation Model | Release D...
Source context: 14 downloads · 1 likes · Pipeline visual-question-answering · Library transformers · Repo wyseow/InternVL-Chat-V1-5-Int8-OL