A Whisper-based model that detects and localizes vocal bursts (laughs, coughs, sneezes, sighs, gasps, cries, screams, etc.) in audio, returning precise start/end timestamps for each event.
Modellquelle
Quellenbeschreibung
A Whisper-based model that detects and localizes vocal bursts (laughs, coughs, sneezes, sighs, gasps, cries, screams, etc.) in audio, returning precise start/end timestamps for each event.
model_v2.pt(972 MB), fine-tuned on real in-the-wild audio. The original (v1) is trained on synthetic soundscapes only and is — it is kept for reproducibility, documented under .
Quellen
1 QuelleVerifiziert 7. Aug.
Modellartefakte
4 Artefaktemodel_v2.pt
Quellenauszüge
2 Auszügemodel_v2.ptmodel.pt⚠️ Two things are easy to get wrong, so they are stated up front:
inference.py still auto-downloads model.pt when you do not pass a checkpoint. Pass model_v2.pt explicitly.inference.py's built-in post-processing defaults are still the v1-era values (threshold=0.65, merge_gap=0.3, min_dur=0.5). Pass the v2 values explicitly — they dominate the measured F1 (see below).threshold = 0.50 # was 0.65 in v1
merge_gap = 0.10 # was 0.30 in v1
min_duration = 0.10 # was 0.50 in v1 <-- the one that matters
Ground-truth bursts have a median duration of ~180 ms. A min_duration of 0.5 s therefore
discards ~96 % of real bursts before matching. On one identical checkpoint, only changing
post-processing moved event F1 from 0.243 to 0.598 — a larger effect than any training change
made for v2. If you read older instructions in this card recommending 0.65 / 0.3 / 0.5, those are
the v1 numbers and are not recommended any more.
from huggingface_hub import hf_hub_download
from inference import load_model, detect_vocal_bursts # inference.py from this repo
# 1. Download the recommended checkpoint
ckpt = hf_hub_download("laion/vocalburst-locator", "model_v2.pt")
# 2. Load it (v1 would be loaded if you omit `checkpoint`)
model, fe, device = load_model("cuda", checkpoint=ckpt) # or "cpu"
# 3. Detect, with the v2 post-processing values
events = detect_vocal_bursts(
"audio.mp3",
model=model, fe=fe, device=device,
threshold=0.50,
merge_gap=0.10,
min_dur=0.10,
)
for ev in events:
print(f"{ev['start']:.2f}s - {ev['end']:.2f}s (confidence: {ev['confidence']:.2f})")
Command line equivalent:
python inference.py audio.mp3 \
--checkpoint "$(python -c 'from huggingface_hub import hf_hub_download; print(hf_hub_download("laion/vocalburst-locator","model_v2.pt"))')" \
--threshold 0.50 --merge-gap 0.10 --min-dur 0.10 --device cuda
Raw state dict (if you build the model yourself — same WhisperSegmenter state dict as v1, 485 tensors, LoRA already merged):
import torch
sd = torch.load("model_v2.pt", map_location="cpu")
model.load_state_dict(sd) # same keys as model.pt
Re-measured on a held-out set of 992 real, in-the-wild expressive-speech clips, each checkpoint given a post-processing sweep to find its best possible operating point:
| event F1 @ IoU 0.5 | precision | recall | best threshold | |
|---|---|---|---|---|
model.pt (v1, synthetic) | 0.152 | 0.469 | 0.207 | 0.80 |
model_v2.pt | 0.607 | 0.678 | 0.669 | 0.50 |
4.0x higher F1 on real audio. Note also how v1 fails: it only reaches usable precision at threshold 0.80, where recall collapses to 0.21 — on real recordings it is very unsure, and buying precision costs it four fifths of the events. v2 operates at 0.50 with recall 0.67.
| Your audio | Checkpoint |
|---|---|
| Expressive speech, in-the-wild (default choice) | model_v2.pt |
| In-the-wild audio that also contains music, SFX or non-speech backgrounds | model_v2_mixed.pt |
| Synthetic soundscapes / reproducing the original results | model.pt (v1, superseded) |
On real audio the two v2 weights are statistically indistinguishable; they differ only on the
synthetic-soundscape domain. Both use the same post-processing values above. Details in the
model_v2_mixed.pt section.
🔗 Ensemble: pair this detector with the captioner
laion/vocalburst-captioning-whisper— locate bursts here, then caption each detected segment with that model. See the threshold study below.
pip install torch transformers soundfile librosa huggingface_hub
Detected 3 vocal burst(s) in audio.mp3:
1. 2.14s - 3.82s (duration: 1.68s, confidence: 0.89)
2. 8.50s - 9.12s (duration: 0.62s, confidence: 0.74)
3. 15.30s - 16.94s (duration: 1.64s, confidence: 0.92)
JSON output (--json):
{
"file": "audio.mp3",
"events": [
{"start": 2.14, "end": 3.82, "confidence": 0.89, "duration": 1.68},
{"start": 8.5, "end": 9.12, "confidence": 0.74, "duration": 0.62},
{"start": 15.3, "end": 16.94, "confidence": 0.92, "duration": 1.64}
]
}
This model performs binary frame-level segmentation on audio: for each 20ms frame in a 30-second audio clip, it predicts whether a vocal burst is occurring. Post-processing then groups these frame-level predictions into discrete events with timestamps and confidence scores.
Audio (16kHz, 30s) → Whisper-small Encoder (LoRA rank-8 merged) → 1500 frame embeddings
→ Linear(768→384) + GELU + Dropout
→ Conv1d(384, kernel=7) + GELU + Dropout (temporal smoothing)
→ Linear(384→1) → sigmoid → 1500 probabilities
→ Post-processing → [(start, end, confidence), ...]
The model uses OpenAI's Whisper-small encoder as the audio feature backbone. During training, the encoder was adapted using LoRA (rank 8, alpha 16) on the q_proj and v_proj attention matrices. The LoRA weights have been merged into the base weights, so no adapter library is needed at inference time. All three checkpoints (model.pt, model_v2.pt, model_v2_mixed.pt) share this architecture and load with identical code.
| File | Size | Description |
|---|---|---|
model_v2.pt | 972 MB | Recommended. Fine-tuned on real in-the-wild expressive speech |
model_v2_mixed.pt | 972 MB | v2 trained on a mix of real speech + synthetic soundscapes (keeps the synthetic domain) |
model.pt | 972 MB | v1, synthetic-only training. Superseded — see Previous version |
head_only.pt | 5.3 MB | v1 segmentation head weights only (use with your own Whisper-small encoder) |
inference.py | - | Standalone inference script with CLI and Python API |
train.py | - | Full training script (supports frozen/LoRA/fine-tuning modes) |
generate_dataset.py |
| Parameter | Recommended (v2) | inference.py built-in default | Description |
|---|---|---|---|
threshold | 0.50 | 0.65 | Detection confidence threshold (0-1). Higher = fewer false positives, lower = fewer missed events. |
merge_gap | 0.10 | 0.3 | Merge predicted segments closer than this (seconds). Prevents a single event from being split into fragments. |
min_dur | 0.10 | 0.5 | Discard predicted events shorter than this (seconds). The v1 default of 0.5 discards ~96 % of real bursts. |
checkpoint | model_v2.pt | model.pt (auto-downloaded) | Which weights to load. |
device | auto | auto | "cpu", "cuda", or etc. Auto-detects GPU if available. |
The "built-in default" column is what the script uses if you pass nothing; it has been left at the v1 values for backwards compatibility. Pass the recommended column explicitly.
Imagine the model is a security guard watching for vocal bursts. It has to make a decision for every moment of audio: "Is this a vocal burst, or not?"
There are four possible outcomes:
REALITY
Vocal Burst Not a VB
┌─────────────┬─────────────┐
MODEL Yes │ True Pos ✓ │ False Pos ✗ │ ← "False alarm"
SAYS: │ (correct!) │ (oops) │
├─────────────┼─────────────┤
No │ False Neg ✗ │ True Neg ✓ │ ← "Missed it"
│ (missed!) │ (correct!) │
└─────────────┴─────────────┘
Precision = Of everything the model flagged, how many were real? TP / (TP + FP)
Recall = Of all real vocal bursts, how many did the model catch? TP / (TP + FN)
F1 Score = The harmonic mean of precision and recall — balances both into one number.
threshold — The confidence cutoffThe model outputs a confidence score (0 to 1) for every 20ms frame. The threshold decides: "How confident must the model be before we call it a vocal burst?"
low threshold → Model flags almost everything
✓ High recall (catches most VBs)
✗ Low precision (many false alarms)
Think: paranoid security guard
high threshold → Model only flags when very sure
✓ High precision (almost no false alarms)
✗ Low recall (misses quieter/ambiguous VBs)
Think: lazy security guard
For model_v2.pt the swept best operating point on real audio is 0.50. (For v1 on synthetic
data it was 0.65; for v1 on real audio it was 0.80, where recall collapses — see
Previous version.)
min_dur — Minimum event durationAfter grouping confident frames into events, discard any event shorter than min_dur.
min_dur = 0.1s → Recommended for v2 on real audio
✓ Keeps short coughs/gasps and the ~180 ms median real burst
✗ Slightly more short false positives
min_dur = 0.5s → The old v1 default
✓ Filters noise spikes in synthetic soundscapes
✗ Discards ~96 % of real bursts
min_dur = 1.0s → Only keeps long events
✗ Misses almost everything on real audio
This is the single most impactful knob. On synthetic soundscapes, mixed-in bursts are long
(0.5–3 s) and a large min_dur cheaply removes false positives — which is why v1 shipped 0.5.
On real recordings the ground-truth median burst is ~180 ms, so the same setting throws away
the majority of true events.
merge_gap — Gap tolerance for mergingIf two detected segments are separated by less than merge_gap, merge them into one event.
merge_gap = 0.0s → No merging. A laugh with a brief pause becomes 2 events.
Result: Over-counting (more events than expected)
merge_gap = 0.1s → Recommended for v2. Bridges frame-level dropouts without
swallowing neighbouring bursts.
merge_gap = 1.0s → Even 1-second gaps get bridged....
pt · 927 MB · SHA-256 3b8b2fef6976…7e3d · Hugging Face
--- license: apache-2.0 tags: - audio - segmentation - vocal-burst - whisper - speech - sound-event-detection language: - en pipeline_tag: audio-classification library_name: transformers base_model: openai/whisper-small --- # Vocal Burst Locator A Whisper-based model that **detects and localizes vocal bursts** (laughs, coughs, sneezes, sighs, gasps, cries, screams, etc.) in audio, returning precise start/end timestamps for each event. ## ⭐ Start here: use `model_v2.pt` **The recommended default checkpoint is [`model_v2.pt`](https://huggingface.co/laion/vocalburst-locator/blob/main/model_v2.pt)** (972 MB), fine-tuned on real in-the-wild audio. The original `model.pt` (v1) is trained on synthetic soundscapes only and is **superseded** — it is kept for reproducibility, documented under [Previous version — v1](#previous-version--v1-modelpt). ⚠️ Two things are easy to get wrong, so they are stated up front: 1. `inference.py` still **auto-downloads `model.pt`** when you do not pass a checkpoint. Pass `model_v2.pt` explicitly. 2. `inference.py`'s built-in post-processing defaults are still the v1-era values (`threshold=0.65, merge_gap=0.3, min_dur=0.5`). Pass the v2 values explicitly — **they dominate the measured F1** (see below). ### Recommended post-processing (v2) ```python threshold = 0.50 # was 0.65 in v1 merge_gap = 0.10 # was 0.30 in v1 min_duration = 0.10 # was 0.50 in v1 <-- the one that matters ``` Ground-truth bursts have a **median duration of ~180 ms**. A `min_duration` of 0.5 s therefore discards **~96 % of real bursts** before matching. On one identical checkpoint, only changing post-processing moved event F1 from **0.243 to 0.598** — a larger effect than any training change made for v2. If you read older instructions in this card recommending `0.65 / 0.3 / 0.5`, those are the v1 numbers and are **not** recommended any more. ### Copy-pasteable usage ```python from huggingface_hub import hf_hub_download from inference import load_model, detect_vocal_bursts # inference.py from this repo # 1. Download the recommended checkpoint ckpt = hf_hub_download("laion/vocalburst-locator", "model_v2.pt") # 2. Load it (v1 would be loaded if you omit `checkpoint`) model, fe, device = load_model("cuda", checkpoint=ckpt) # or "cpu" # 3. Detect, with the v2 post-processing values events = detect_vocal_bursts( "audio.mp3", model=model, fe=fe, d...
Source context: 58 downloads · 1 likes · Pipeline audio-classification · Library transformers · Repo laion/vocalburst-locator
| - |
| Synthetic training data generator |
download_sources.py | - | Downloads source audio from HuggingFace datasets |
config.json | - | Model configuration and training hyperparameters |
vocalburst_threshold_report.html | - | Interactive ensemble threshold study report |
"cuda:0"