[!TIP]
KV-cache quantization without any fork (recommended, 2026): upstream
llama.cpp/Ollama now cover this natively — use -ctk q8_0 -ctv q8_0
(~half KV memory, negligible quality loss: perplexity +0.002–0.05) or
-ctk q4_0 -ctv q4_0 (~quarter memory, ≈7.6% perplexity increase). In
Ollama: OLLAMA_KV_CACHE_TYPE=q8_0 with OLLAMA_FLASH_ATTENTION=1. Keep
K and V types symmetric to stay on the fast fused Flash-Attention path.
Since April 2026, mainline llama.cpp also applies Hadamard rotation to
KV activations (PR #21038),
which greatly improves low-bit KV quality (opt-out:
LLAMA_ATTN_ROT_DISABLE=1).
The RotorQuant/TurboQuant fork flow below is experimental/legacy: the
TurboQuant llama.cpp PR was closed without merging (June 2026) and the fork
is unmaintained relative to mainline. It is NOT required to use this model.
BF16 (kept full-precision in every non-GGUF variant; GGUF uses mmproj-F16 split file)
Audio
Parakeet-TDT-0.6B-v2
BF16 (same rationale)
Video
Parakeet-TDT-0.6B-v2 + frame sampler
BF16 (≤ 2 min, 256 frames @ 2 FPS)
NVIDIA's official FP8 / NVFP4 recipe keeps both encoders + the cross-modal
MLP projectors in BF16 to preserve multimodal accuracy. We follow that
convention in every quantized variant we ship.
Runtime quirks
llama.cpp
Use llama-mtmd-cli for multimodal inference; pass --mmproj mmproj-F16.gguf
(see majentik/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-mmproj-F16).
Do NOT use CUDA 13.2 — produces gibberish. Pin CUDA 12.x or
use the Metal/CPU paths.
Ollama
Text-only; multimodal is blocked because Ollama doesn't yet support
the mmproj split-file pattern.
Reasoning mode
enable_thinking defaults to True. To disable extended reasoning
(e.g., for latency-sensitive cases), pass enable_thinking=False
to the chat template / generate call. No separate "no-think"
variant card exists — this is a runtime flag, not a model variant.
Quant trade-off (GGUF lane)
Quant
Approx size
Use case
Recommendation
Q2_K
~17 GB
Lossy, low-RAM CPU/edge
Resource-constrained inference
Q3_K_M
~19 GB
Smaller-than-Q4, modest quality drop
Edge devices with ~16 GB RAM
IQ4_XS
~16 GB
Importance-quant 4-bit, smaller than Q4_K_M
Best size/quality at 4-bit
Q4_K_M
~23 GB
Balanced default
Recommended for most users
Q5_K_M
~24 GB
Higher fidelity than Q4
Quality-sensitive applications
Q6_K
~28 GB
Approaching FP16 quality
High-fidelity CPU/edge
Q8_0
~32 GB
Near-lossless reference
(Current variant — Q8_0 — is bolded.)
Variants in this family
(Showing 56 sibling variants under majentik/nemotron3-nano-omni-30b-*. The current variant — TurboQuant-GGUF-Q8_0 — is bolded.)
---
license: other
license_name: nvidia-open-model-license
license_link: https://developer.download.nvidia.com/licenses/nvidia-open-model-license-agreement-june-2024.pdf
base_model: nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16
tags: [nemotron, multimodal, mamba2, moe, quantized, turboquant, gguf, llama.cpp,
llama-mtmd, multimodal-via-mmproj]
library_name: gguf
pipeline_tag: image-text-to-text
language: [en]
datasets: [nvidia/Nemotron-Image-Training-v3]
inference: false
---
> [!TIP]
> **KV-cache quantization without any fork (recommended, 2026):** upstream
> llama.cpp/Ollama now cover this natively — use `-ctk q8_0 -ctv q8_0`
> (~half KV memory, negligible quality loss: perplexity +0.002–0.05) or
> `-ctk q4_0 -ctv q4_0` (~quarter memory, ≈7.6% perplexity increase). In
> Ollama: `OLLAMA_KV_CACHE_TYPE=q8_0` with `OLLAMA_FLASH_ATTENTION=1`. Keep
> K and V types symmetric to stay on the fast fused Flash-Attention path.
> Since April 2026, mainline llama.cpp also applies Hadamard rotation to
> KV activations ([PR #21038](https://github.com/ggml-org/llama.cpp/pull/21038)),
> which greatly improves low-bit KV quality (opt-out:
> `LLAMA_ATTN_ROT_DISABLE=1`).
>
> The RotorQuant/TurboQuant fork flow below is **experimental/legacy**: the
> TurboQuant llama.cpp PR was closed without merging (June 2026) and the fork
> is unmaintained relative to mainline. It is NOT required to use this model.
<!-- kv-upstream-note -->
# Nemotron-3-Nano-Omni-30B-A3B-Reasoning - TurboQuant GGUF Q8_0
GGUF Q8_0 quantization of `Nemotron-3-Nano-Omni-30B-A3B-Reasoning` (`nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16`) with TurboQuant weight method.
The `Q8_0.gguf` binary in this repo is loaded by `llama.cpp` / `llama-mtmd-cli`.
For multimodal inference (text + image + audio + video) pair this with the
multimodal projector: [`majentik/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-mmproj-F16`](https://huggingface.co/majentik/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-mmproj-F16).
For the matched-KV stack — TurboQuant weights + TurboQuant KV-cache modifier —
see [`majentik/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-TurboQuant-GGUF-Q8_0-TQ-KV`](https://huggingface.co/majentik/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-TurboQuant-GGUF-Q8_0-TQ-KV).
For the runtime KV-cache modifier itself (weight-agnostic), see
[`majentik/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-TurboQuant`](https://huggingface.co/maj...