apache-2.0
Fonte do modelo
Trecho da fonte
apache-2.0
Fontes
1 fonteVerificado 20 de ago.
Artefatos de modelo
1 artefatoTrechos de fonte
2 trechos--- license: apache-2.0 license_name: apache-2.0 license_link: https://www.apache.org/licenses/LICENSE-2.0 language: - en - hi - multilingual - gu - mr - te - fr - ja - zh - ml pipeline_tag: image-text-to-text tags: - gemma4 - quatfit - multimodal - vision - audio - text generation - image-text-to-text - agentic - coding - fasterllm --- <p align="center"> <img src="https://huggingface.co/Quatfit/Quatfit-Mini/resolve/main/banner.png" alt="Quatfit Mini Banner" width="100%"> </p> <h1 align="center">Quatfit Mini</h1> <h3 align="center">Gemma 4–Based 8B Multimodal Model · Up to 4× Faster Inference · 131K Context</h3> <p align="center"> <a href="https://huggingface.co/Quatfit/Quatfit-Mini"><img src="https://img.shields.io/badge/🤗%20Hugging%20Face-Model%20Hub-FFD21E?style=flat-square" alt="Hugging Face"></a> <a href="https://huggingface.co/Quatfit/Quatfit-Mini-GGUF"><img src="https://img.shields.io/badge/⚙️%20GGUF%20Builds-Quatfit--Mini--GGUF-4052F5?style=flat-square" alt="GGUF Builds"></a> <a href="https://huggingface.co/Quatfit/Quatfit-Mini/resolve/main/Quatfit-Mini_Technical_Report.pdf"><img src="https://img.shields.io/badge/📄%20Technical%20Report-PDF-8A2BE2?style=flat-square" alt="Technical Report"></a> <a href="https://huggingface.co/Quatfit/Quatfit-Mini"><img src="https://img.shields.io/badge/⚡%204×%20Faster-FP32%20%7C%20GGUF-00C853?style=flat-square" alt="Performance"></a> <a href="https://www.apache.org/licenses/LICENSE-2.0"><img src="https://img.shields.io/badge/📜%20License-Apache%202.0-E53935?style=flat-square" alt="License"></a> </p> --- ## Quatfit Mini **Quatfit Mini** is an **8-billion-parameter multimodal model built on Google's Gemma 4 architecture** and further optimized by **Quatfit AI Research** for efficient deployment, long-context reasoning, and agentic AI workflows. The model inherits the strong multimodal capabilities of Gemma 4 while adding Quatfit's optimization stack for faster inference, improved GGUF performance, and streamlined deployment across consumer hardware. **Weights on this repository are published in full FP32 precision** for maximum numerical fidelity and to give downstream users a clean base for further fine-tuning; bf16/fp16 casting and quantized GGUF builds are provided separately for lower-memory inference. ### Highlights - Built on **Google Gemma 4** - Native multimodal reasoning (Text + Image + Audio...
Source context: 51 downloads · 7 likes · Pipeline image-text-to-text · Repo Quatfit/Quatfit-Mini