VRGDG_LoadAudioSplit_HUMO_TranscribeV3
A powerful ComfyUI custom node that automates audio segmentation, transcription, and project management for HuMo / VRGDG-style workflows. This node takes an audio clip, splits it into 16 evenly timed segments, optionally transcribes each segment using OpenAI Whisper, manages metadata and output folders, and provides visual workflow instructions through ComfyUI popups. VRGDGLoadAudioSplitHUMOTranscribeV3 simplifies large audio-to-video pipelines by: - Splitting an input audio file into scene-length chunks (e.g., 4 seconds each). - Managing multi-run workflows for long audio clips. - Automatically detecting project state and versioning output folders. - Optionally transcribing each scene using OpenAI Whisper-Large-V3. - Sending color-coded popup instructions to ComfyUI. | Name | Type | Description | |------|------|--------------| | audio | AUDIO | The input audio clip (waveform + sample rate). | | trigger | any | A signal input used for chaining workflow steps. | | scenedurationseconds | FLOAT | Duration of each scene (default: 4.0s). | | folderpath | STRING | Output folder for project files. | | enableautoqueue | BOOLEAN | Automatically queue next runs if needed. | | language | STRING | Language for Whisper transcription (default: English). | | enablelyrics | BOOLEAN | Enable automatic transcription. | | usecontextonly | BOOLEAN | Use provided context fields instead of transcription. | | overlaplyricseconds | FLOAT | Overlap between segments for smoother lyrics merging. | | fallbackwords | STRING | Backup words for empty transcriptions. | | context1–context16 | STRING | Opt…
VRGDG · 12 Ausgaben