ComfyUI_Step_Audio_EditX_SM
Step_Audio_EditX:the first open-source LLM-based audio model excelling at expressive and iterative audio editing—encompassing emotion, speaking style, and paralinguistics—alongside robust zero-shot text-to-speech (TTS) capabilities,try it in comfyUI
How much VRAM does ComfyUI_Step_Audio_EditX_SM require?
Direct Answer: The ComfyUI node ComfyUI_Step_Audio_EditX_SM requires a minimum base VRAM of 4096MB and is optimized for GPUs with at least 8GB of VRAM. Low VRAM mode is not supported for this node.
- Base VRAM:
- 4096MB (4.0GB)
- Recommended GPU:
- 8GB+ VRAM
- Low VRAM Mode:
- ✗ Not supported
- Estimation Confidence:
- MEDIUM
Interactive VRAM Compatibility Estimator
Your GPU has plenty of headroom. You can run this node safely with your active configurations!
Deploy on High-Performance GPUs
Need more VRAM to run ComfyUI with this node? Rent low-cost, high-performance GPUs on Vast.ai instantly.
Deploy on Cloud GPUs
Need more VRAM to run ComfyUI with this node? Rent low-cost, high-performance cloud GPUs on RunPod instantly.
What Python packages are required for ComfyUI_Step_Audio_EditX_SM?
Direct Answer: Running ComfyUI_Step_Audio_EditX_SM requires installing the following Python package dependencies: accelerate, conformer, diffusers, funasr>=1.1.3, librosa, modelscope, numpy, nvidia-cuda-nvrtc-cu12, omegaconf, onnxruntime, onnxruntime-gpu, openai-whisper==20240930, pillow, protobuf, sentencepiece, six, torch, torchaudio, torchvision, transformers==4.53.3. Ensure your ComfyUI environment has these packages active before launching.
accelerate
conformer
diffusers
funasr>=1.1.3
librosa
modelscope
numpy
nvidia-cuda-nvrtc-cu12
omegaconf
onnxruntime
onnxruntime-gpu
openai-whisper==20240930
pillow
protobuf
sentencepiece
six
torch
torchaudio
torchvision
transformers==4.53.3Interactive Setup & Dependency Resolver
# Loading command...Frequently Asked Questions
How much VRAM does ComfyUI_Step_Audio_EditX_SM require?
ComfyUI_Step_Audio_EditX_SM requires a minimum of 4096MB (4.0GB) of VRAM for base operation. For optimal performance, a GPU with at least 8GB of VRAM is recommended. Low VRAM mode is not supported for this node.
Can I run ComfyUI_Step_Audio_EditX_SM on an RTX 3060, RTX 4070, or RTX 4090?
✅ RTX 3060 (12GB): Yes, fully compatible with 6.8GB headroom. ✅ RTX 4070 (12GB): Yes, fully compatible with 6.8GB headroom. ✅ RTX 4070 Ti (16GB): Yes, fully compatible with 10.4GB headroom. ✅ RTX 4090 (24GB): Yes, fully compatible with 17.6GB headroom
What PyTorch version does ComfyUI_Step_Audio_EditX_SM need?
ComfyUI_Step_Audio_EditX_SM requires the following PyTorch-related packages: torch, torchaudio, torchvision. Ensure your ComfyUI environment has these installed. Ensure your PyTorch installation matches your CUDA version (use torch.version.cuda to check).
What Python packages are required for ComfyUI_Step_Audio_EditX_SM?
To run ComfyUI_Step_Audio_EditX_SM, you need to install: accelerate, conformer, diffusers, funasr>=1.1.3, librosa, modelscope, numpy, nvidia-cuda-nvrtc-cu12, omegaconf, onnxruntime, onnxruntime-gpu, openai-whisper==20240930, pillow, protobuf, sentencepiece, six, torch, torchaudio, torchvision, transformers==4.53.3. You can install these using pip or add them to your requirements.txt file.
How do I install ComfyUI_Step_Audio_EditX_SM in ComfyUI?
To install ComfyUI_Step_Audio_EditX_SM: (1) Navigate to your ComfyUI/custom_nodes directory, (2) Clone the repository: git clone https://github.com/smthemex/ComfyUI_Step_Audio_EditX_SM, (3) Install dependencies: pip install -r requirements.txt (if present), (4) Restart ComfyUI. Alternatively, use ComfyUI Manager for one-click installation.