weight: fp8 channel, activation: fp8 dynamic
Fuente del modelo
Extracto de la fuente
weight: fp8 channel, activation: fp8 dynamic
Fuentes
1 fuenteVerificado 31 ago
Artefactos del modelo
1 artefactoExtractos de fuentes
2 extractos--- library_name: transformers license: apache-2.0 license_link: https://ai.google.dev/gemma/docs/gemma_4_license pipeline_tag: any-to-any tags: - fp8 - gemma-4 - lightvl --- ## gemma-4-E4B-it-fp8 **weight: fp8 channel, activation: fp8 dynamic** **15G -> 12G memory decrease** **speedup 15%** **Start the vLLM server** **vllm serve Hyper-AI/gemma-4-E4B-it-fp8 --max-model-len 32768** **lightvl Developed by Myself is a lightweight Vision-Language Model (VLM) quantization toolkit supporting FP8, INT8, FP8-Block. It integrates with vLLM for high-throughput inference and supports Qwen3-VL, Qwen3.5, InternVL-Chat, and Gemma-4 models.** **fast quant your model step by step:** 1、 pip3 install lightvl 2、 lightvl YOUR_HF_MODEL_PATH 3、 the output quant model path is YOUR_HF_MODEL_PATH-fp8 <div align="center"> <img src=https://ai.google.dev/gemma/images/gemma4_banner.png> </div> <p align="center"> <a href="https://huggingface.co/collections/google/gemma-4" target="_blank">Hugging Face</a> | <a href="https://github.com/google-gemma" target="_blank">GitHub</a> | <a href="https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/" target="_blank">Launch Blog</a> | <a href="https://ai.google.dev/gemma/docs/core" target="_blank">Documentation</a> <br> <b>License</b>: <a href="https://ai.google.dev/gemma/docs/gemma_4_license" target="_blank">Apache 2.0</a> | <b>Authors</b>: <a href="https://deepmind.google/models/gemma/" target="_blank">Google DeepMind</a> </p> Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on small models) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages. Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. The models are available in four distinct sizes: **E2B**, **E4B**, **26B A4B**, and **31B**. Their diverse sizes make them deployable in environments ranging from high-end phones to laptops and servers, democratizing access to state-of-the-art AI. Gemma 4 introduces key **capability and architectural advancements**: * **Reasoning** – All models in the...
Source context: 247 downloads · 0 likes · Pipeline any-to-any · Library transformers · Repo Hyper-AI/gemma-4-E4B-it-fp8