Fine-tuned unsloth/orpheus-3b-0.1-pretrained for Turkish text-to-speech using LoRA on a SNAC 24 kHz audio-token vocabulary. Trained on Red Hat OpenShift AI with Kubeflow Trainer v2 (TrainJob) · RHOAIENG-62326.
Fonte do modelo
Trecho da fonte
Fine-tuned unsloth/orpheus-3b-0.1-pretrained for Turkish text-to-speech using LoRA on a SNAC 24 kHz audio-token vocabulary. Trained on Red Hat OpenShift AI with Kubeflow Trainer v2 (TrainJob) · RHOAIENG-62326.
Fontes
1 fonteVerificado 22 de set.
Artefatos de modelo
2 artefatosTrechos de fonte
2 trechos--- language: tr license: apache-2.0 base_model: unsloth/orpheus-3b-0.1-pretrained tags: - text-to-speech - tts - turkish - orpheus - snac - lora - openshift-ai - kubeflow-trainer library_name: transformers inference: false --- # Orpheus-3B Turkish TTS (v2) Fine-tuned [unsloth/orpheus-3b-0.1-pretrained](https://huggingface.co/unsloth/orpheus-3b-0.1-pretrained) for **Turkish text-to-speech** using LoRA on a SNAC 24 kHz audio-token vocabulary. Trained on Red Hat OpenShift AI with **Kubeflow Trainer v2 (TrainJob)** · [RHOAIENG-62326](https://redhat.atlassian.net/browse/RHOAIENG-62326). **Use case:** airline-style Turkish announcements where general-purpose TTS lacks Turkish phonology. | | | |---|---| | **Base model** | `unsloth/orpheus-3b-0.1-pretrained` (Llama-3 backbone + SNAC codec tokens) | | **Fine-tuning** | LoRA r=32, α=64 on attention + MLP projections | | **Training data** | [afkfatih/turkish-tts-combined-raw](https://huggingface.co/datasets/afkfatih/turkish-tts-combined-raw) — 20,000 samples | | **Merged weights** | LoRA from `checkpoint-8800` (best eval_loss) | | **Audio codec** | [hubertsiuzdak/snac_24khz](https://huggingface.co/hubertsiuzdak/snac_24khz) @ 24 kHz | ## Model hierarchy & specs ``` unsloth/orpheus-3b-0.1-pretrained # English-biased Orpheus base (frozen) └── LoRA adapters (r=32, α=64) # ~185M trainable params └── merged → this repo # full weights, Turkish TTS ``` | | | |---|---| | Architecture | Llama-3 3B causal LM + SNAC audio token head | | Parameters | ~3B total; ~185M trainable during LoRA fine-tune | | Vocab | 128,256 text + SNAC codec tokens from offset 128,266 | | Audio | SNAC 24 kHz · 7 tokens/frame · 3 codebooks | | Precision | bfloat16 + Flash Attention 2 | | LoRA | r=32, α=64, dropout 0.05 | | Target modules | `q/k/v/o_proj`, `gate/up/down_proj` | Orpheus frames TTS as **causal LM over SNAC audio tokens**: Turkish text → control tokens → autoregressive codec tokens → 24 kHz waveform. ## Dataset specs | | | |---|---| | Source | [afkfatih/turkish-tts-combined-raw](https://huggingface.co/datasets/afkfatih/turkish-tts-combined-raw) | | Samples used | 20,000 (of ~81k available) | | Format | WAV + Turkish transcript | | Preprocessing | Resample 24 kHz mono → SNAC encode → token sequence | | Seq length | max 4,096 tokens | | Split | 95% train / 5% eval (from preprocessed pool) | | Platform | 2× NVIDIA A100 80...
Source context: 448 downloads · 2 likes · Pipeline text-to-speech · Library transformers · Repo AbDhumal/orpheus-3b-turkish-tts-v2