nvidia-open-model-agreement
Model source
Source description
INT4 weight-only quantization of
nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16.
Sized to fit a single ≥ 24 GB consumer / workstation GPU.
Sources
1 sourceVerified Sep 26
Model artifacts
14 artifactsSource excerpts
2 excerpts| Property | Value |
|---|
| Base model | nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16 |
| Active parameters / token | ~3B (of 31B total) |
| Modality | text + image + audio + video → text |
| Quantization | INT4 weight-only |
| Approx. on-disk size | ~22 GB |
| Context length | up to 256k tokens |
| Languages | English |
Loaded and verified with vLLM ≥ 0.20.0 (native nemotron_v3 /
NanoNemotronVL path). Round-trip correctness: bit-exact within INT4
quantization step (per-layer dequantize MAE ≈ 1e-5).
needle-1M-bench-mvp 50K| Metric | Score |
|---|---|
| Overall recall | 90.0 % |
| Paper-anchored recall | 80.0 % |
| Synthetic-codes recall | 100.0 % |
| Haystack tokens | 50,566 |
| Max output tokens | 2048 |
| Scorer | strip_think_includes (centralized) |
Single miss is the deepest needle (depth 49,496 / 50,566). All 9 other depths score 100 %.
Leaderboard:
drawais/needle-1M-bench-mvp.
Per-row YAML:
.eval_results/nemotron-omni-30b-a3b-w4a16-50k.yaml.
from vllm import LLM, SamplingParams
llm = LLM(
model="drawais/Nemotron-3-Nano-Omni-30B-A3B-W4A16",
trust_remote_code=True,
max_model_len=65536,
)
params = SamplingParams(temperature=0.6, top_p=0.95, max_tokens=4096)
print(llm.generate(["Hello, world!"], params)[0].outputs[0].text)
vllm serve drawais/Nemotron-3-Nano-Omni-30B-A3B-W4A16 \
--trust-remote-code \
--max-model-len 65536 \
--gpu-memory-utilization 0.94
Then point any OpenAI client at the local endpoint:
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="dummy")
print(client.chat.completions.create(
model="drawais/Nemotron-3-Nano-Omni-30B-A3B-W4A16",
messages=[{"role": "user", "content": "Hello"}],
max_tokens=64,
).choices[0].message.content)
For multimodal usage (image / audio / video), reasoning controls,
recommended --reasoning-parser nemotron_v3, tool-calling flags, and
per-modality serving recommendations, follow the upstream
Nemotron-3-Nano-Omni model card.
If audio inputs are used: pip install vllm[audio].
~22 GB on disk for the weights. Total VRAM should leave headroom for KV cache and multimodal-encoder activations; recommended:
--max-model-len and text-only usagetrust_remote_code=True is required.
Source model © NVIDIA Corporation, released under the
NVIDIA Open Model Agreement.
This artifact is a Derivative Work as defined in that agreement.
See LICENSE and NOTICE for full text and
required attribution.
NVIDIA Open Model Agreement (Release Date: April 2, 2026).
Commercially usable. You are free to create and distribute Derivative Works. NVIDIA does not claim ownership of outputs.
The full agreement text is included in LICENSE. The
attribution notice required by Section 3(c) is in NOTICE:
Licensed by NVIDIA Corporation under the NVIDIA Open Model Agreement.
model-00003-of-00014.safetensors
safetensors · 1.32 GB · SHA-256 e453ffe818ec…cac0 · Hugging Face
model-00004-of-00014.safetensors
safetensors · 1.25 GB · SHA-256 f7b770244d48…d95a · Hugging Face
Downloadmodel-00005-of-00014.safetensors
safetensors · 1.31 GB · SHA-256 faa293d5bc0f…9b0b · Hugging Face
Downloadmodel-00006-of-00014.safetensors
safetensors · 1.32 GB · SHA-256 5f8b4036fc04…3448 · Hugging Face
Downloadmodel-00007-of-00014.safetensors
safetensors · 1.32 GB · SHA-256 3fc14259ab7e…851a · Hugging Face
Downloadmodel-00008-of-00014.safetensors
safetensors · 1.32 GB · SHA-256 a01227c20fae…019a · Hugging Face
Downloadmodel-00009-of-00014.safetensors
safetensors · 1.32 GB · SHA-256 3509b0ae6bb7…93ef · Hugging Face
Downloadmodel-00010-of-00014.safetensors
safetensors · 1.26 GB · SHA-256 43b70b617eb6…f240 · Hugging Face
Downloadmodel-00011-of-00014.safetensors
safetensors · 1.32 GB · SHA-256 7219d0d4509d…c667 · Hugging Face
Downloadmodel-00012-of-00014.safetensors
safetensors · 1.32 GB · SHA-256 00a6c43a4de5…862d · Hugging Face
Downloadmodel-00013-of-00014.safetensors
safetensors · 1.97 GB · SHA-256 49c2d382baeb…41cd · Hugging Face
Downloadmodel-00014-of-00014.safetensors
safetensors · 2.37 GB · SHA-256 a628a1a4894b…41e3 · Hugging Face
Download--- license: other license_name: nvidia-open-model-agreement license_link: https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-agreement/ base_model: nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16 tags: - nvidia - nemotron - multimodal - quantized - 4-bit - int4 language: - en library_name: transformers pipeline_tag: any-to-any --- # Nemotron-3-Nano-Omni-30B-A3B — INT4 INT4 weight-only quantization of [`nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16`](https://huggingface.co/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16). Sized to fit a single ≥ 24 GB consumer / workstation GPU. | Property | Value | |---|---| | Base model | nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16 | | Active parameters / token | ~3B (of 31B total) | | Modality | text + image + audio + video → text | | Quantization | INT4 weight-only | | Approx. on-disk size | ~22 GB | | Context length | up to 256k tokens | | Languages | English | ## Validation Loaded and verified with vLLM ≥ 0.20.0 (native `nemotron_v3` / NanoNemotronVL path). Round-trip correctness: bit-exact within INT4 quantization step (per-layer dequantize MAE ≈ 1e-5). ### Score on `needle-1M-bench-mvp` 50K | Metric | Score | |---|---| | Overall recall | **90.0 %** | | Paper-anchored recall | 80.0 % | | Synthetic-codes recall | 100.0 % | | Haystack tokens | 50,566 | | Max output tokens | 2048 | | Scorer | `strip_think_includes` (centralized) | Single miss is the deepest needle (depth 49,496 / 50,566). All 9 other depths score 100 %. Leaderboard: [`drawais/needle-1M-bench-mvp`](https://huggingface.co/datasets/drawais/needle-1M-bench-mvp). Per-row YAML: [`.eval_results/nemotron-omni-30b-a3b-w4a16-50k.yaml`](https://huggingface.co/datasets/drawais/needle-1M-bench-mvp/blob/main/.eval_results/nemotron-omni-30b-a3b-w4a16-50k.yaml). ## Load (vLLM, text) ```python from vllm import LLM, SamplingParams llm = LLM( model="drawais/Nemotron-3-Nano-Omni-30B-A3B-W4A16", trust_remote_code=True, max_model_len=65536, ) params = SamplingParams(temperature=0.6, top_p=0.95, max_tokens=4096) print(llm.generate(["Hello, world!"], params)[0].outputs[0].text) ``` ## Serve (vLLM, OpenAI-compatible) ```bash vllm serve drawais/Nemotron-3-Nano-Omni-30B-A3B-W4A16 \ --trust-remote-code \ --max-model-len 65536 \ --gpu-memory-utilization 0.94 ``` Then point any OpenAI client at the local endp...
Source context: 537 downloads · 0 likes · Pipeline any-to-any · Library transformers · Repo drawais/Nemotron-3-Nano-Omni-30B-A3B-W4A16