license: mit basemodel: microsoft/Phi-4-mini-instruct-onnx basemodelrelation: quantized pipelinetag: text-generation libraryname: onnxruntime-genai language: ar zh cs da nl en fi fr de he hu it ja ko "no" pl pt ru es...
Model source
Source description
This repository repackages the gpu/gpu-int4-rtn-block-32 configuration from Microsoft's
official ONNX release
microsoft/Phi-4-mini-instruct-onnx.
It is not a newly trained model. This is an unofficial community repackage; Microsoft is
the original author and license holder.
Sources
1 sourceVerified Jul 17
Model artifacts
1 artifactgpu/gpu-int4-rtn-block-32/model.onnx
onnx · 280 KB · SHA-256 f8ff300b8571…5250 · Hugging Face
DownloadSource excerpts
3 excerpts| Field | Value |
|---|
| Upstream model | microsoft/Phi-4-mini-instruct-onnx |
| Upstream source revision | fc04c8f93df696602fd9f300a30d1bf2e3081347 |
| Export tool/script | Microsoft ONNX Runtime GenAI model builder (upstream Phi ONNX bundle) |
| Quantization recipe | ONNX Runtime GenAI RTN INT4 (gpu-int4-rtn-block-32) |
All files are under gpu/gpu-int4-rtn-block-32/:
| File | Size | Description |
|---|---|---|
model.onnx | ~280 KB | INT4 ONNX graph |
model.onnx.data | ~3 GB | External weights |
genai_config.json | ~2 KB | ONNX Runtime GenAI session config |
config.json | ~2 KB | Model config |
configuration_phi3.py | ~11 KB | Phi-3/4 config class |
tokenizer.json + tokenizer_config.json | ~15 MB | Tokenizer |
vocab.json + merges.txt | ~6 MB | Vocabulary / merges |
| + |
A 4-bit, GPU-targeted build of Phi-4-Mini-Instruct for local text generation and
translation. Targets NVIDIA CUDA and DirectML (Windows) execution paths where a capable
GPU is available. A CPU variant is in
tonythethompson/phi-4-mini-instruct-cpu-int4-onnx.
onnxruntime-genai-cuda) or DirectML (onnxruntime-genai-directml).Export tooling, precision, and quantization are recorded in the Source table above. This packaging mirror does not publish independent parity benchmarks; validate on your target execution provider before production use.
MIT — inherited from
microsoft/Phi-4-mini-instruct-onnx. This packaging repo adds no new license terms.
--- license: mit base_model: microsoft/Phi-4-mini-instruct-onnx base_model_relation: quantized pipeline_tag: text-generation library_name: onnxruntime-genai language: - ar - zh - cs - da - nl - en - fi - fr - de - he - hu - it - ja - ko - "no" - pl - pt - ru - es - sv - th - tr - uk tags: - onnx - onnxruntime-genai - phi4 - phi4-mini - int4 - gpu - text-generation - translation --- # Phi-4 Mini Instruct — GPU INT4 (ONNX Runtime GenAI) This repository repackages the `gpu/gpu-int4-rtn-block-32` configuration from Microsoft's official ONNX release [microsoft/Phi-4-mini-instruct-onnx](https://huggingface.co/microsoft/Phi-4-mini-instruct-onnx). It is not a newly trained model. This is an unofficial community repackage; Microsoft is the original author and license holder. ## Source | Field | Value | |---|---| | Upstream model | [microsoft/Phi-4-mini-instruct-onnx](https://huggingface.co/microsoft/Phi-4-mini-instruct-onnx) | | Upstream source revision | `fc04c8f93df696602fd9f300a30d1bf2e3081347` | | Export tool/script | Microsoft ONNX Runtime GenAI model builder (upstream Phi ONNX bundle) | | Quantization recipe | ONNX Runtime GenAI RTN INT4 (`gpu-int4-rtn-block-32`) | ## Files All files are under `gpu/gpu-int4-rtn-block-32/`: | File | Size | Description | |---|---|---| | `model.onnx` | ~280 KB | INT4 ONNX graph | | `model.onnx.data` | ~3 GB | External weights | | `genai_config.json` | ~2 KB | ONNX Runtime GenAI session config | | `config.json` | ~2 KB | Model config | | `configuration_phi3.py` | ~11 KB | Phi-3/4 config class | | `tokenizer.json` + `tokenizer_config.json` | ~15 MB | Tokenizer | | `vocab.json` + `merges.txt` | ~6 MB | Vocabulary / merges | | `added_tokens.json` + `special_tokens_map.json` | <1 KB | Special tokens | ## Intended Use A 4-bit, GPU-targeted build of Phi-4-Mini-Instruct for local text generation and translation. Targets NVIDIA CUDA and DirectML (Windows) execution paths where a capable GPU is available. A CPU variant is in [`tonythethompson/phi-4-mini-instruct-cpu-int4-onnx`](https://huggingface.co/tonythethompson/phi-4-mini-instruct-cpu-int4-onnx). ## Runtime Notes - Designed for ONNX Runtime GenAI compatible runtimes. - GPU execution: CUDA (`onnxruntime-genai-cuda`) or DirectML (`onnxruntime-genai-directml`). - Context length: 128K tokens (inherited from Phi-4-M...
Source context: 0 downloads · 0 likes · Pipeline text-generation · Library onnxruntime-genai · Repo tonythethompson/Phi-4-Mini-Instruct-GPU-INT4-ONNX
Source context: 0 downloads · 0 likes · Pipeline text-generation · Library onnxruntime-genai · Repo tonythethompson/trackdub-phi-4-mini-gpu-int4
added_tokens.jsonspecial_tokens_map.json| <1 KB |
| Special tokens |