Best improved LCNN checkpoint from a comparative study of three neural anti-spoofing architectures. Supersedes caa-speech-detection-asvspoof2019/lcnn.
Model source
Source description
Best improved LCNN checkpoint from a comparative study of three neural anti-spoofing architectures. Supersedes caa-speech-detection-asvspoof2019/lcnn.
Version: lcnn_v7 (CQT + label-smoothing + cosine schedule + grad-clip)
Light CNN over Constant-Q Transform (CQT) spectrograms.
Sources
1 sourceVerified Aug 7
Model artifacts
1 artifactSource excerpts
2 excerpts| Component | Value |
|---|
| Input | CQT, 84 bins, 12 bins/octave, fmin=32.7 Hz, hop=160 |
| Channels | [32, 48, 64, 128] |
| Kernel sizes | [5, 5, 3, 3] |
| FC hidden | 64 |
| Dropout | 0.3 |
| Label smoothing | 0.1 |
| Parameters | ~300 k |
Reference: Lavrentyeva et al., "Audio Replay Attack Detection with Deep Learning Frameworks", Interspeech 2017.
| Hyperparameter | Value |
|---|---|
| Epochs | 50 (stopped early at 30) |
| Batch size | 256 |
| Learning rate | 5e-5, cosine schedule (eta_min=5e-6) |
| Weight decay | 1e-3 |
| Gradient clip norm | 0.3 |
| Early stopping patience | 10 |
Dataset: ASVspoof 2019 LA train split (~25k utterances). No data augmentation.
Baseline to beat: EER 8.09% (LFCC+GMM).
| Split | EER | tandem min t-DCF | In-the-Wild EER |
|---|---|---|---|
| Dev | 0.708% | — | — |
| Eval | 3.26% | 0.4930 | 33.41% |
Dev EER improved from 0.902% (baseline LCNN) to 0.708%. Eval EER improved from 8.43% to 3.26% — a 61% relative reduction.
See learning_curves/lcnn_v4_cqt_vs_v7.png for the training trajectory comparing v4_cqt and v7.
Install dependencies from the source repository, then:
import torch
from src.models.lcnn.model import LCNNModel
config = {
"sample_rate": 16000,
"transform_type": "cqt",
"n_bins": 84,
"bins_per_octave": 12,
"hop_length": 160,
"fmin": 32.7,
"max_audio_seconds": 4.0,
"channels": [32, 48, 64, 128],
"kernel_sizes": [5, 5, 3, 3],
"fc_hidden": 64,
"dropout": 0.3,
"label_smoothing": 0.1,
}
model = LCNNModel(config)
state = torch.load("best.pt", map_location="cpu")
model.load_state_dict(state)
model.eval()
# waveform: (B, T) float32 at 16 kHz, T up to 64000
with torch.no_grad():
logits = model({"waveform": waveform})["logits"]
probs = torch.softmax(logits, dim=-1) # [:, 0] = spoof, [:, 1] = bonafide
@inproceedings{wang2020asvspoof,
title = {{ASVspoof} 2019: A Large-Scale Public Database of Synthesized, Converted and Replayed Speech},
author = {Wang, Xin and others},
booktitle = {Computer Speech \& Language},
volume = {64},
year = {2020}
}
@inproceedings{lavrentyeva2017audio,
title = {Audio Replay Attack Detection with Deep Learning Frameworks},
author = {Lavrentyeva, Galina and Novoselov, Sergey and Malykh, Egor and Kozlov, Alexander and Kudashev, Oleg and Shchemelinin, Vadim},
booktitle = {Interspeech},
year = {2017}
}
--- license: mit tags: - audio-classification - anti-spoofing - asvspoof - deepfake-detection datasets: - asvspoof2019 pipeline_tag: audio-classification metrics: - eer - t-dcf --- # lcnn-v7-cqt — Improved LCNN for ASVspoof 2019 LA Best improved LCNN checkpoint from a comparative study of three neural anti-spoofing architectures. Supersedes [caa-speech-detection-asvspoof2019/lcnn](https://huggingface.co/caa-speech-detection-asvspoof2019/lcnn). **Version:** lcnn_v7 (CQT + label-smoothing + cosine schedule + grad-clip) ## Architecture Light CNN over Constant-Q Transform (CQT) spectrograms. | Component | Value | |---|---| | Input | CQT, 84 bins, 12 bins/octave, fmin=32.7 Hz, hop=160 | | Channels | `[32, 48, 64, 128]` | | Kernel sizes | `[5, 5, 3, 3]` | | FC hidden | 64 | | Dropout | 0.3 | | Label smoothing | 0.1 | | Parameters | ~300 k | Reference: Lavrentyeva et al., *"Audio Replay Attack Detection with Deep Learning Frameworks"*, Interspeech 2017. ## Training | Hyperparameter | Value | |---|---| | Epochs | 50 (stopped early at 30) | | Batch size | 256 | | Learning rate | 5e-5, cosine schedule (eta_min=5e-6) | | Weight decay | 1e-3 | | Gradient clip norm | 0.3 | | Early stopping patience | 10 | Dataset: ASVspoof 2019 LA train split (~25k utterances). No data augmentation. ## Results Baseline to beat: **EER 8.09%** (LFCC+GMM). | Split | EER | tandem min t-DCF | In-the-Wild EER | |---|---|---|---| | Dev | 0.708% | — | — | | **Eval** | **3.26%** | **0.4930** | **33.41%** | Dev EER improved from 0.902% (baseline LCNN) to 0.708%. Eval EER improved from 8.43% to 3.26% — a 61% relative reduction. See `learning_curves/lcnn_v4_cqt_vs_v7.png` for the training trajectory comparing v4_cqt and v7. ## Usage Install dependencies from the source repository, then: ```python import torch from src.models.lcnn.model import LCNNModel config = { "sample_rate": 16000, "transform_type": "cqt", "n_bins": 84, "bins_per_octave": 12, "hop_length": 160, "fmin": 32.7, "max_audio_seconds": 4.0, "channels": [32, 48, 64, 128], "kernel_sizes": [5, 5, 3, 3], "fc_hidden": 64, "dropout": 0.3, "label_smoothing": 0.1, } model = LCNNModel(config) state = torch.load("best.pt", map_location="cpu") model.load_state_dict(state) model.eval() # waveform: (B, T) float32 at 16 kHz, T up to 64000 with torch.no_grad(): logits = model({"waveform": waveform})["logits"...
Source context: 6 downloads · 0 likes · Pipeline audio-classification · Repo caa-speech-detection-asvspoof2019/lcnn-v7-cqt