Paketprofil
README
Video → audio → JSON transcript pipeline for ComfyUI. Pairs with ComfyUI-MiMoASR (Xiaomi MiMo-V2.5-ASR) for the ASR step.
This pack provides the pieces of a video-to-text workflow: video loading, audio extraction (), dict bridging, JSON packaging, and disk persistence. The ASR itself is delegated to MiMoASR, which gives you strong Mandarin / English code-switching, dialect support (Wu, Cantonese, Hokkien, Sichuanese, Henan, Northeastern Mandarin), and noisy-audio robustness.
Quellen
3 QuellenQuellenauszüge
1 AuszugSource context: Repo aadebuger/ComfyUI-VideoToText
ffmpegAUDIO| Node | Purpose |
|---|---|
| 🎙️ Load Video (ASR) | Pick a video from ComfyUI/input/ or pass an absolute path. |
| 🎙️ Extract Audio (ffmpeg) | Demux audio track to a 16 kHz mono WAV. On-disk cache keyed by mtime + params. |
| 🎙️ Audio From Path | Bridge: WAV file → ComfyUI AUDIO dict. Optional resample / mono. Also reports duration. |
| 🎙️ Build Transcript JSON | Wrap ASR text + source metadata into a structured v2.0 JSON. |
| 🎙️ Save Transcript JSON | Validate + write JSON to ComfyUI/output/ with timestamped filename. |
This pack does not ship an ASR model. Three things must be true before you can run a workflow:
transformers==4.49.0 downgrade and flash-attn build. Verify those nodes register before you even start with this pack.ffmpeg on your server's PATH — $ which ffmpeg must return a path.torchaudio in ComfyUI's venv — installed in Step 2 below.Quick gut check:
| Required | How to check | |
|---|---|---|
ffmpeg | any recent | $ ffmpeg -version |
torchaudio | any | $ python -c "import torchaudio; print(torchaudio.__version__)" |
| ComfyUI-MiMoASR | installed | $ curl -s http://SERVER:8188/object_info | grep -o MiMoASRLoader |
PreviewAny (from rgthree-comfy) | installed | $ curl -s http://SERVER:8188/object_info | grep -o PreviewAny |
Each step has a verification at the end. Don't skip the verifications.
This pack depends on MiMoASRLoader and MiMoASRTranscribe from a separate repo. Follow its full install guide before continuing here. Stop and verify:
$ curl -s http://SERVER:8188/object_info | python3 -c "
import json, sys
d = json.load(sys.stdin)
for n in ('MiMoASRLoader', 'MiMoASRTranscribe'):
print(f' {n}: {\"OK\" if n in d else \"MISSING\"}')"
Both must print OK. If MISSING, fix MiMoASR before continuing.
$ cd ComfyUI/custom_nodes
$ git clone https://github.com/aadebuger/ComfyUI-VideoToText.git
Or, if installing from a tarball release:
$ cd ComfyUI/custom_nodes
$ tar xzf /path/to/ComfyUI-VideoToText-v0.2.0.tar.gz
Verify:
$ ls ComfyUI/custom_nodes/ComfyUI-VideoToText/__init__.py
# Must exist. If it's nested inside ComfyUI-VideoToText/ComfyUI-VideoToText/, you have a double-wrap — flatten it.
$ cd ComfyUI && source .venv/bin/activate
$ uv pip install torchaudio
Verify:
$ python -c "import torchaudio; print('torchaudio', torchaudio.__version__)"
# Must print a version string. Errors here → re-run install.
ffmpeg on the server (if missing)$ which ffmpeg || sudo apt install ffmpeg
$ ffmpeg -version
The same launch incantation MiMoASR needs — both HF_HOME and MIMO_ASR_REPO env vars must be set:
$ pkill -9 -f main.py; sleep 3
$ cd ComfyUI && source .venv/bin/activate
$ export HF_HOME=/big_disk/hf_cache
$ export MIMO_ASR_REPO=/opt/repos/MiMo-V2.5-ASR
$ nohup python main.py --listen 0.0.0.0 --port 8188 > ~/comfy_start.log 2>&1 &
$ sl
Knoten in diesem Paket
5 Knoten (Plural)Verifiziert 11. Sept.
Verifiziert 11. Sept.