Package profile
README
ComfyUI custom node for speech-to-text using whisper.cpp (C API via ctypes).
Full pipeline: ASR → Vocal Separation → VAD → Alignment → Diarization — zero whisperx.
Node
Sources
5 sourcesVerified Sep 16
whisper_full_params — every parameter exposed as a node widgetshow_advance_cpp (core whisper params) + show_advance_ext (UVR/alignment/diarization detail)cd ComfyUI/custom_nodes/
git clone --recurse-submodules https://github.com/obvirm/ComfyUI-WhisperCPP
cd ComfyUI-WhisperCPP
pip install -r requirements.txt
That's it! DLLs are automatically downloaded from GitHub Releases on first use. No build step required.
If you want to build the DLLs yourself (e.g., for GPU-specific optimizations):
# whisper.cpp — auto-detect GPU backends
python build_whisper_cpp.py
# BSRoformer DLL
python build_bs_roformer.py
# cpp-annote DLL
python build_cpp_annote.py
# whisper.cpp
python build_whisper_cpp.py --gpu vulkan # Force Vulkan
python build_whisper_cpp.py --gpu cuda # Force CUDA
python build_whisper_cpp.py --gpu cpu # CPU only
# BSRoformer
python build_bs_roformer.py --gpu # Auto-detect GPU
python build_bs_roformer.py # CPU only
# cpp-annote
python build_cpp_annote.py --cuda # CUDA ONNX Runtime
python build_cpp_annote.py --gpu # Auto-detect GPU
python build_cpp_annote.py # CPU only
| Module | Purpose | Binding | File |
|---|---|---|---|
| whisper.cpp | ASR (speech-to-text) | whisper_lib.py | whisper.dll + ggml DLLs |
| BSRoformer.cpp | UVR vocal separation | bs_roformer_lib.py | bs_roformer.dll |
| cpp-annote | VAD + Diarization | cpp_annote_lib.py | cpp_annote.dll + onnxruntime.dll |
All modules follow the same pattern: C API header → shared library → ctypes binding.
Audio Input
│
├─ [UVR] BSRoformer.dll → vocal separation (optional)
│ 16kHz → Resample(44100) → DLL → mono → Resample(16000)
│
├─ [Pre-filter] RMS energy → speech detection (optional)
│ Only transcribe sections with audio energy > threshold
│
├─ [VAD] cpp-annote.dll → speech segmentation (optional)
│
├─ whisper.dll → transcribe (per-segment if VAD/pre-filter on)
│
├─ [Alignment] sherpa-onnx CTC → word timestamps (default ON)
│
└─ [Diarization] cpp-annote DLL → speaker labels (optional)
│
└─ 7 output sockets
On first use, each module automatically downloads its DLLs from GitHub Releases:
1. _find_library() → DLL not found
2. gpu_detect.detect_gpu() → NVIDIA/Vulkan/OpenCL/CPU
3. auto_download.download_module("whisper", dir, has_gpu)
→ Download whisper.dll, ggml-base.dll, ggml-cpu.dll, ggml.dll
→ (GPU) also download ggml-vulkan.dll
4. _find_library() → found!
pyproject.toml (no manual updates)Nodes in this pack
1 nodeVerified Sep 16
Verified Aug 2