[\[🆕 Blog\]]( [\[📜 InternVL 1.0 Paper\]]( [\[📜 InternVL 1.5 Report\]]( [\[🗨️ Chat Demo\]](
Modellquelle
Quellenauszug
[[🆕 Blog]]( [[📜 InternVL 1.0 Paper]]( [[📜 InternVL 1.5 Report]]( [[🗨️ Chat Demo]](
Quellen
1 QuelleVerifiziert 28. Aug.
Modellartefakte
2 Artefaktemodel-00001-of-00002.safetensors
safetensors · 4,62 GB · SHA-256 ceea2c6af60b…6ca7 · Hugging Face
Herunterladenmodel-00002-of-00002.safetensors
safetensors · 3,11 GB · SHA-256 0191d413ddd6…a20e · Hugging Face
HerunterladenQuellenauszüge
2 Auszüge--- license: mit datasets: - laion/laion2B-en - laion/laion-coco - laion/laion2B-multi - kakaobrain/coyo-700m - conceptual_captions - wanng/wukong100m pipeline_tag: visual-question-answering --- # Model Card for Mini-InternVL-Chat-4B-V1-5 <center> <p><img src="https://cdn-uploads.huggingface.co/production/uploads/64119264f0f81eb569e0d569/pvfKc16O-ej91632FHaIK.png" style="width:80%;" alt="image/png"></p> </center> [\[🆕 Blog\]](https://internvl.github.io/blog/) [\[📜 InternVL 1.0 Paper\]](https://arxiv.org/abs/2312.14238) [\[📜 InternVL 1.5 Report\]](https://arxiv.org/abs/2404.16821) [\[🗨️ Chat Demo\]](https://internvl.opengvlab.com/) [\[🤗 HF Demo\]](https://huggingface.co/spaces/OpenGVLab/InternVL) [\[🚀 Quick Start\]](#model-usage) [\[🌐 Community-hosted API\]](https://rapidapi.com/adushar1320/api/internvl-chat) [\[📖 中文解读\]](https://zhuanlan.zhihu.com/p/675877376) You can run multimodal large models using a 1080Ti now. We are delighted to introduce the Mini-InternVL-Chat series. In the era of large language models, many researchers have started to focus on smaller language models, such as Gemma-2B, Qwen-1.8B, and InternLM2-1.8B. Inspired by their efforts, we have distilled our vision foundation model [InternViT-6B-448px-V1-5](https://huggingface.co/OpenGVLab/InternViT-6B-448px-V1-5) down to 300M and used [InternLM2-Chat-1.8B](https://huggingface.co/internlm/internlm2-chat-1_8b) or [Phi-3-mini-128k-instruct](https://huggingface.co/microsoft/Phi-3-mini-128k-instruct) as our language model. This resulted in a small multimodal model with excellent performance. As shown in the figure below, we adopted the same model architecture as InternVL 1.5. We simply replaced the original InternViT-6B with InternViT-300M and InternLM2-Chat-20B with InternLM2-Chat-1.8B / Phi-3-mini-128k-instruct. For training, we used the same data as InternVL 1.5 to train this smaller model. Additionally, due to the lower training costs of smaller models, we used a context length of 8K during training.  ## Model Details - **Model Type:** multimodal large language model (MLLM) - **Model Stats:** - Architecture: [InternViT-300M-448px](https://huggingface.co/OpenGVLab/InternViT-300M-448px) + MLP + [Phi-3-mini-128k-instruct](https://huggingface.co/micros...
Source context: 18 downloads · 0 likes · Pipeline visual-question-answering · Library transformers · Repo radna/mini_intern_chat_triton