ComfyUI WhisperX Pro
A professional ComfyUI custom node for accurate text-audio alignment using [a/WhisperX](https://github.com/m-bain/whisperX).
Quick Technical Summary: ComfyUI WhisperX Pro
- Base VRAM Footprint:
- 1536 MB (6 GB Tier)
- Primary Dependencies:
- ctc-segmentation>=1.7.0, faster-whisper>=0.9.0, librosa>=0.10.0, soundfile>=0.12.0, torch>=2.0.0, torchaudio>=2.0.0, transformers>=4.30.0
- Min PyTorch / CUDA:
- PyTorch 2.0.0 | CUDA 12.1+
- GitHub Repository:
- https://github.com/loockluo/comfyui-whisperx-pro
Citation Note: Data sourced from VRAM DB. For complete workflow OOM estimations, use the VRAM DB Workflow Analyzer.
How much VRAM does ComfyUI WhisperX Pro require?
Direct Answer: The ComfyUI node ComfyUI WhisperX Pro requires a minimum base VRAM of 1536MB and is optimized for GPUs with at least 6GB of VRAM. Low VRAM mode is fully supported for resource-constrained setups.
- Base VRAM:
- 1536MB (1.5GB)
- Recommended GPU:
- 6GB+ VRAM
- Low VRAM Mode:
- ✓ Supported
- Estimation Confidence:
- MEDIUM
Cheapest VRAM Upgrade Paths (Live Market Prices):
- GeForce RTX 3060 12GB (Ultimate Budget VRAM King)──► Used: $209.62View eBay ↗
- GeForce RTX 4060 8GB (Modern Entry-Level)──► New: $303.50View Amazon ↗
Interactive VRAM Compatibility Estimator
Your GPU has plenty of headroom. You can run this node safely with your active configurations!
Verify Compatibility for Your Specific GPU VRAM
Select your graphics card's VRAM capacity to view optimized batch sizes, suggested resolutions, and custom performance tips for ComfyUI WhisperX Pro:
Buy NVIDIA GeForce RTX 3060 (12GB VRAM)
Tired of renting cloud rigs? Run ComfyUI locally with absolute zero latency. Best entry-level ComfyUI experience. Avoids immediate VRAM limitations on basic LoRA training.
Are you the author of this node?
Help your users avoid out-of-memory errors by displaying this professional, dynamic VRAM badge on your GitHub README. Copy the markdown below to embed it with a backlink directly to this hardware specification profile.
What Python packages are required for ComfyUI WhisperX Pro?
Direct Answer: Running ComfyUI WhisperX Pro requires installing the following Python package dependencies: ctc-segmentation>=1.7.0, faster-whisper>=0.9.0, librosa>=0.10.0, soundfile>=0.12.0, torch>=2.0.0, torchaudio>=2.0.0, transformers>=4.30.0. This node specifically requires PyTorch version 2.0.0 or newer.
ctc-segmentation>=1.7.0
faster-whisper>=0.9.0
librosa>=0.10.0
soundfile>=0.12.0
torch>=2.0.0
torchaudio>=2.0.0
transformers>=4.30.0Interactive Setup & Dependency Resolver
# Loading command...Special Environment Requirements:
Requires PyTorch version 2.0.0+.
# Loading PyTorch command...Frequently Asked Questions
How much VRAM does ComfyUI WhisperX Pro require?
ComfyUI WhisperX Pro requires a minimum of 1536MB (1.5GB) of VRAM for base operation. For optimal performance, a GPU with at least 6GB of VRAM is recommended. This node supports low VRAM mode for resource-constrained setups.
Can I run ComfyUI WhisperX Pro on an RTX 3060, RTX 4070, or RTX 4090?
✅ RTX 3060 (12GB): Yes, fully compatible with 9.3GB headroom. ✅ RTX 4070 (12GB): Yes, fully compatible with 9.3GB headroom. ✅ RTX 4070 Ti (16GB): Yes, fully compatible with 12.9GB headroom. ✅ RTX 4090 (24GB): Yes, fully compatible with 20.1GB headroom
How much VRAM does ComfyUI WhisperX Pro take on an RTX 3060 vs RTX 4090?
On an RTX 3060 (12GB VRAM), ComfyUI WhisperX Pro runs smoothly on an RTX 3060 (12GB) with 9.3GB of headroom. This is sufficient to run the node alongside standard SD 1.5 and SDXL workflows in full precision. On an RTX 4090 (24GB VRAM), the node runs with extreme headroom on an RTX 4090 (24GB) with 20.1GB of dedicated headroom. This allows you to combine the node with massive models (like FLUX.1 Dev, Schnell, or Hunyuan Video) in full precision (FP16) without any offload flags.
What PyTorch version does ComfyUI WhisperX Pro need?
ComfyUI WhisperX Pro requires the following PyTorch-related packages: torch>=2.0.0, torchaudio>=2.0.0. Ensure your ComfyUI environment has these installed. Ensure your PyTorch installation matches your CUDA version (use torch.version.cuda to check).
What Python packages are required for ComfyUI WhisperX Pro?
To run ComfyUI WhisperX Pro, you need to install: ctc-segmentation>=1.7.0, faster-whisper>=0.9.0, librosa>=0.10.0, soundfile>=0.12.0, torch>=2.0.0, torchaudio>=2.0.0, transformers>=4.30.0. You can install these using pip or add them to your requirements.txt file.
How can I reduce VRAM usage when running ComfyUI WhisperX Pro?
ComfyUI WhisperX Pro supports low VRAM mode. To reduce memory usage: (1) Enable --lowvram or --medvram flags in ComfyUI, (2) Reduce batch size to 1, (3) Use fp16 or fp8 precision if supported, (4) Close other GPU applications.
How do I install ComfyUI WhisperX Pro in ComfyUI?
To install ComfyUI WhisperX Pro: (1) Navigate to your ComfyUI/custom_nodes directory, (2) Clone the repository: git clone https://github.com/loockluo/comfyui-whisperx-pro, (3) Install dependencies: pip install -r requirements.txt (if present), (4) Restart ComfyUI. Alternatively, use ComfyUI Manager for one-click installation.