---
language:
- lus
license: apache-2.0
pipeline_tag: automatic-speech-recognition
base_model: openai/whisper-large-v3-turbo
tags:
- generated_from_trainer
datasets:
- andrewbawitlung/MiZonal-v3.0
metrics:
- wer
- cer
model-index:
- name: whisper-large-v3-turbo-mizonal3-E5-lus-v2026.06
results:
- task:
name: Automatic Speech Recognition
type: automatic-speech-recognition
dataset:
name: MiZonal v3.0
type: andrewbawitlung/MiZonal-v3.0
config: default
split: test
metrics:
- name: Wer
type: wer
value: 13.8215
- name: Cer
type: cer
value: 2.6263
- name: Real Time Factor
type: rtf
value: 0.0540
---

> **Disclaimer / Notice:** Details for these are in Peer Review and publications of the paper will be made available soon for more details.
# whisper-large-v3-turbo-mizonal3-E5-lus-v2026.06
This model is a fine-tuned version of [openai/whisper-large-v3-turbo](https://huggingface.co/openai/whisper-large-v3-turbo) on the **MiZonal v3.0** dataset.
It achieves the following results on the evaluation set:
- Wer: 13.8215
- Cer: 2.6263
- Real Time Factor: 0.0540
## Quick Inference
```python
import torch
import librosa
from transformers import WhisperProcessor, WhisperForConditionalGeneration
device = "cuda" if torch.cuda.is_available() else "cpu"
processor = WhisperProcessor.from_pretrained("andrewbawitlung/whisper-large-v3-turbo-mizonal3-E5-lus-v2026.06")
model = WhisperForConditionalGeneration.from_pretrained("andrewbawitlung/whisper-large-v3-turbo-mizonal3-E5-lus-v2026.06").to(device)
audio, sr = librosa.load("your_audio.wav", sr=16000)
input_features = processor(audio, sampling_rate=16000, return_tensors="pt").input_features.to(device)
with torch.no_grad():
predicted_ids = model.generate(input_features, max_new_tokens=256)
transcription = processor.batch_decode(predicted_ids, skip_special_tokens=True)[0]
print(transcription)
```
## Model description
### Experiment Configurations
This repository is part of a series of experiments. The different configurations are:
- **E1 (Baseline):** Standard training configuration.
- **E2 (Noise):** Training with background noise augmentation.
- **E3 (Speed):** Training with speed perturbation augmentation.
- **E4 (SpecAug):** Training with SpecAugment (time and frequency masking).
- **E5 (Combined):** Training with a combination of all...