Step Audio EditX TTS

By saganaki22View on GitHub →

Professional voice cloning and audio editing node for ComfyUI using Step Audio EditX

How much VRAM does Step Audio EditX TTS require?

Direct Answer: The ComfyUI node Step Audio EditX TTS requires a minimum base VRAM of 1536MB and is optimized for GPUs with at least 6GB of VRAM. Low VRAM mode is fully supported for resource-constrained setups.

High (2-4GB)
Base VRAM:
1536MB (1.5GB)
Recommended GPU:
6GB+ VRAM
Low VRAM Mode:
✓ Supported
Estimation Confidence:
MEDIUM

Interactive VRAM Compatibility Estimator

Estimated Total VRAM: 3.00 GBTarget: 8 GB
✅ Comfortable Fit

Your GPU has plenty of headroom. You can run this node safely with your active configurations!

Deploy on High-Performance GPUs

Need more VRAM to run ComfyUI with this node? Rent low-cost, high-performance GPUs on Vast.ai instantly.

🚀 Deploy on Vast.ai

Deploy on Cloud GPUs

Need more VRAM to run ComfyUI with this node? Rent low-cost, high-performance cloud GPUs on RunPod instantly.

🚀 Deploy on RunPod

What Python packages are required for Step Audio EditX TTS?

Direct Answer: Running Step Audio EditX TTS requires installing the following Python package dependencies: accelerate>=0.20.0, gradio==5.49.1, hyperpyyaml>=1.2.2, librosa>=0.10.2, omegaconf>=2.3.0, onnxruntime>=1.17.0, openai-whisper>=20231117, protobuf>=3.20.0, sentencepiece>=0.1.99, soundfile>=0.12.1, torchaudio>=2.0.0, transformers==4.53.3. Ensure your ComfyUI environment has these packages active before launching.

requirements.txt
accelerate>=0.20.0
gradio==5.49.1
hyperpyyaml>=1.2.2
librosa>=0.10.2
omegaconf>=2.3.0
onnxruntime>=1.17.0
openai-whisper>=20231117
protobuf>=3.20.0
sentencepiece>=0.1.99
soundfile>=0.12.1
torchaudio>=2.0.0
transformers==4.53.3

Interactive Setup & Dependency Resolver

Operating System:
Environment Type:
Run this terminal command in your ComfyUI root folder:
# Loading command...

Frequently Asked Questions

How much VRAM does Step Audio EditX TTS require?

Step Audio EditX TTS requires a minimum of 1536MB (1.5GB) of VRAM for base operation. For optimal performance, a GPU with at least 6GB of VRAM is recommended. This node supports low VRAM mode for resource-constrained setups.

Can I run Step Audio EditX TTS on an RTX 3060, RTX 4070, or RTX 4090?

✅ RTX 3060 (12GB): Yes, fully compatible with 9.3GB headroom. ✅ RTX 4070 (12GB): Yes, fully compatible with 9.3GB headroom. ✅ RTX 4070 Ti (16GB): Yes, fully compatible with 12.9GB headroom. ✅ RTX 4090 (24GB): Yes, fully compatible with 20.1GB headroom

What PyTorch version does Step Audio EditX TTS need?

Step Audio EditX TTS requires the following PyTorch-related packages: torchaudio>=2.0.0. Ensure your ComfyUI environment has these installed. Ensure your PyTorch installation matches your CUDA version (use torch.version.cuda to check).

What Python packages are required for Step Audio EditX TTS?

To run Step Audio EditX TTS, you need to install: accelerate>=0.20.0, gradio==5.49.1, hyperpyyaml>=1.2.2, librosa>=0.10.2, omegaconf>=2.3.0, onnxruntime>=1.17.0, openai-whisper>=20231117, protobuf>=3.20.0, sentencepiece>=0.1.99, soundfile>=0.12.1, torchaudio>=2.0.0, transformers==4.53.3. You can install these using pip or add them to your requirements.txt file.

How can I reduce VRAM usage when running Step Audio EditX TTS?

Step Audio EditX TTS supports low VRAM mode. To reduce memory usage: (1) Enable --lowvram or --medvram flags in ComfyUI, (2) Reduce batch size to 1, (3) Use fp16 or fp8 precision if supported, (4) Close other GPU applications.

How do I install Step Audio EditX TTS in ComfyUI?

To install Step Audio EditX TTS: (1) Navigate to your ComfyUI/custom_nodes directory, (2) Clone the repository: git clone https://github.com/Saganaki22/ComfyUI-Step_Audio_EditX_TTS, (3) Install dependencies: pip install -r requirements.txt (if present), (4) Restart ComfyUI. Alternatively, use ComfyUI Manager for one-click installation.