Full fine-tune of CohereLabs/cohere-transcribe-arabic-07-2026 (2.07B, cohereasr Conformer encoder-decoder) for multi-dialect Arabic speech recognition (undiacritized output).
Modellquelle
Quellenbeschreibung
Full fine-tune of CohereLabs/cohere-transcribe-arabic-07-2026
(2.07B, cohere_asr Conformer encoder-decoder) for multi-dialect Arabic
speech recognition (undiacritized output).
Quellen
1 QuelleVerifiziert 13. Sept.
Modellartefakte
1 ArtefaktQuellenauszüge
2 AuszügePrivate / internal model. Evaluate on your own data before production use.
| WER | CER | |
|---|---|---|
| Base (cohere-transcribe-arabic, zero-shot) | 0.457 | 0.174 |
| This model (full fine-tune) | 0.357 | 0.137 |
~22% relative WER cut over the (already strong) Arabic-specialized base. This is the model the base should be fine-tuned into: an earlier LoRA attempt on the same data made it worse (0.457 → 0.510, overfit). A full fine-tune with a low LR and best-checkpoint selection was the fix.
⚠️ This model overfits easily. Eval loss bottomed at step 1000 (0.342) and rose afterward, while WER stayed flat — the saved weights are that best checkpoint (
load_best_model_at_end). Watch WER, not just loss, if you train further.
Same 932-clip held-out test set, same clean_text scoring (strip tashkil + tags,
keep punctuation + dialect spelling) — so every row is directly comparable.
| Model | Params | Zero-shot WER | Fine-tuned WER | CER (best) |
|---|---|---|---|---|
| whisper-large-v3-turbo 🏆 | 809M | 0.590 | 0.344 | 0.115 |
| cohere-transcribe-arabic | 2.0B | 0.457 | 0.357 | 0.137 |
| whisper-medium | 769M | 0.717 | 0.358 | 0.123 |
| nemotron-3.5-asr (streaming) | 638M | 0.592 | 0.422 | — |
| whisper-small | 244M | ~0.77 | 0.428 | 0.151 |
| qwen3-asr-0.6b | 938M | 0.756 |
whisper-large-v3-turbo (WER 0.344), with cohere-transcribe-arabic
a close second (0.357).cohere-transcribe-arabic (0.457, Arabic-specialized). A full
fine-tune (all ~2B params, low LR) now improves it to 0.357; an earlier 32 GB
LoRA attempt had instead degraded it (0.510, overfit) — full-parameter tuning
with best-checkpoint selection was the fix.nemotron-3.5-asr.Fine-tuned on oddadmix/dialectal-arabic-lahgtna-v2-smaller-augmented (private).
augmentation column. Derived from
oddadmix/dialectal-arabic-lahgtna-v2-smaller.[laughter], [exhale], [inhale], [mumble], [cough], timestamps, ...).⚠️ Leakage note: the set is augmentation-expanded from shared source clips, so train/test may share source audio → held-out numbers can be optimistic. Use a leakage-free split for true numbers.
CohereLabs/cohere-transcribe-arabic-07-2026 · full fine-tune (no LoRA)per_device 2 × grad_accum 16) · LR 5e-6 (kept low —
strong specialist) · 2 epochs (2300 steps) · best @ step 1000 (eval_loss 0.342)The exact fine-tuning code is bundled in this repo (train_cohere_full.py,
normalize.py, eval_hf_asr.py, ds_zero2.json, setup.sh) plus requirements.txt.
See FINETUNE.md for the full walkthrough + lessons learned.
Trained on oddadmix/dialectal-arabic-lahgtna-v2-smaller-augmented (private) —
swap in any HF audio dataset with audio + text columns. normalize.py is the
shared text cleaning (strip tashkil + non-verbal tags, keep dialectal letters گ ڨ چ).
bash setup.sh # deps (transformers from source) + HF login
python train_cohere_full.py --output_dir cohere-ar-full \
--per_device_train_batch_size 2 --gradient_accumulation_steps 16 \
--per_device_eval_batch_size 2 --learning_rate 5e-6
python eval_hf_asr.py --model_type cohere --model cohere-ar-full # WER / CER
For less memory / multi-GPU, accelerate launch train_cohere_full.py --deepspeed ds_zero2.json ....
# needs transformers from source (ships the cohere_asr architecture)
import torch, torchaudio
from transformers import AutoProcessor, CohereAsrForConditionalGeneration
repo = "oddadmix/cohere-transcribe-arabic-07-2026-dialectal"
proc = AutoProcessor.from_pretrained(repo)
model = CohereAsrForConditionalGeneration.from_pretrained(
repo, torch_dtype=torch.bfloat16).to("cuda").eval()
wav, sr = torchaudio.load("clip.wav") # 16 kHz mono
inp = proc(wav.mean(0).numpy(), sampling_rate=16000,
language="ar", return_tensors="pt").to(model.device)
inp["input_features"] = inp["input_features"].to(model.dtype) # keep ids long
ids = model.generate(**inp, max_new_tokens=256)
print(proc.tokenizer.batch_decode(ids, skip_special_tokens=True)[0])
Decode one bare array at a time. Passing a list +
padding=Truetriggers a batched/padded path that badly degrades this model's generation (WER 0.45 → 0.65+).
eval_loss bottomed at step 1000 and rose while WER held —
more steps would have overfit. load_best_model_at_end keeps the right weights.--- language: ar license: apache-2.0 library_name: transformers pipeline_tag: automatic-speech-recognition base_model: CohereLabs/cohere-transcribe-arabic-07-2026 datasets: - oddadmix/dialectal-arabic-lahgtna-v2-smaller-augmented tags: - automatic-speech-recognition - arabic - dialectal-arabic - asr - transformers - cohere_asr metrics: - wer - cer model-index: - name: cohere-transcribe-arabic-07-2026-dialectal results: - task: type: automatic-speech-recognition dataset: name: dialectal-arabic-lahgtna-v2 (test) type: oddadmix/dialectal-arabic-lahgtna-v2-smaller-augmented metrics: - type: wer value: 0.357 - type: cer value: 0.137 --- # cohere-transcribe-arabic-07-2026-dialectal Full fine-tune of [`CohereLabs/cohere-transcribe-arabic-07-2026`](https://huggingface.co/CohereLabs/cohere-transcribe-arabic-07-2026) (2.07B, `cohere_asr` Conformer encoder-decoder) for **multi-dialect Arabic** speech recognition (undiacritized output). > **Private / internal model.** Evaluate on your own data before production use. ## Results (932-clip held-out test set) | | WER | CER | |---|---|---| | Base (cohere-transcribe-arabic, zero-shot) | 0.457 | 0.174 | | **This model (full fine-tune)** | **0.357** | **0.137** | ~22% relative WER cut over the (already strong) Arabic-specialized base. This is the model the base *should* be fine-tuned into: an earlier **LoRA** attempt on the same data made it **worse** (0.457 → 0.510, overfit). A **full** fine-tune with a low LR and best-checkpoint selection was the fix. > ⚠️ **This model overfits easily.** Eval loss bottomed at **step 1000** (0.342) and > rose afterward, while WER stayed flat — the saved weights are that best checkpoint > (`load_best_model_at_end`). Watch **WER**, not just loss, if you train further. ## Model comparison — all Arabic ASR models Same 932-clip held-out test set, same `clean_text` scoring (strip tashkil + tags, keep punctuation + dialect spelling) — so every row is directly comparable. | Model | Params | Zero-shot WER | Fine-tuned WER | CER (best) | |---|---|---|---|---| | **whisper-large-v3-turbo** 🏆 | 809M | 0.590 | **0.344** | 0.115 | | **cohere-transcribe-arabic** | 2.0B | **0.457** | 0.357 | 0.137 | | whisper-medium | 769M | 0.717 | 0.358 | 0.123 | | nemotron-3.5-asr (streaming) | 638M | 0.592 | 0.422 | — | | whisper-small | 244M | ~0.77 | 0.428 | 0.151 | | qwen3-asr-0.6b | 938M | 0.756...
Source context: 56 downloads · 2 likes · Pipeline automatic-speech-recognition · Library transformers · Repo oddadmix/cohere-transcribe-arabic-07-2026-dialectal
| 0.676 |
| 0.408 |
| qwen3-asr-1.7b | 1.7B | — | training | — |