Step Audio EditX TTS
Professional voice cloning and audio editing node for ComfyUI using Step Audio EditX
Quick Technical Summary: Step Audio EditX TTS
- Base VRAM Footprint:
- 1536 MB (6 GB Tier)
- Primary Dependencies:
- accelerate>=0.20.0, gradio==5.49.1, hyperpyyaml>=1.2.2, librosa>=0.10.2, omegaconf>=2.3.0, onnxruntime>=1.17.0, openai-whisper>=20231117, protobuf>=3.20.0, sentencepiece>=0.1.99, soundfile>=0.12.1, torchaudio>=2.0.0, transformers==4.53.3
- Min PyTorch / CUDA:
- PyTorch 2.0+ | CUDA 12.1+
- GitHub Repository:
- https://github.com/Saganaki22/ComfyUI-Step_Audio_EditX_TTS
Citation Note: Data sourced from VRAM DB. For complete workflow OOM estimations, use the VRAM DB Workflow Analyzer.
How much VRAM does Step Audio EditX TTS require?
Direct Answer: The ComfyUI node Step Audio EditX TTS requires a minimum base VRAM of 1536MB and is optimized for GPUs with at least 6GB of VRAM. Low VRAM mode is fully supported for resource-constrained setups.
- Base VRAM:
- 1536MB (1.5GB)
- Recommended GPU:
- 6GB+ VRAM
- Low VRAM Mode:
- ✓ Supported
- Estimation Confidence:
- MEDIUM
Cheapest VRAM Upgrade Paths (Live Market Prices):
- GeForce RTX 3060 12GB (Ultimate Budget VRAM King)──► Used: $209.62View eBay ↗
- GeForce RTX 4060 8GB (Modern Entry-Level)──► New: $303.50View Amazon ↗
Interactive VRAM Compatibility Estimator
Your GPU has plenty of headroom. You can run this node safely with your active configurations!
Verify Compatibility for Your Specific GPU VRAM
Select your graphics card's VRAM capacity to view optimized batch sizes, suggested resolutions, and custom performance tips for Step Audio EditX TTS:
Buy NVIDIA GeForce RTX 3060 (12GB VRAM)
Tired of renting cloud rigs? Run ComfyUI locally with absolute zero latency. Best entry-level ComfyUI experience. Avoids immediate VRAM limitations on basic LoRA training.
Are you the author of this node?
Help your users avoid out-of-memory errors by displaying this professional, dynamic VRAM badge on your GitHub README. Copy the markdown below to embed it with a backlink directly to this hardware specification profile.
What Python packages are required for Step Audio EditX TTS?
Direct Answer: Running Step Audio EditX TTS requires installing the following Python package dependencies: accelerate>=0.20.0, gradio==5.49.1, hyperpyyaml>=1.2.2, librosa>=0.10.2, omegaconf>=2.3.0, onnxruntime>=1.17.0, openai-whisper>=20231117, protobuf>=3.20.0, sentencepiece>=0.1.99, soundfile>=0.12.1, torchaudio>=2.0.0, transformers==4.53.3. Ensure your ComfyUI environment has these packages active before launching.
accelerate>=0.20.0
gradio==5.49.1
hyperpyyaml>=1.2.2
librosa>=0.10.2
omegaconf>=2.3.0
onnxruntime>=1.17.0
openai-whisper>=20231117
protobuf>=3.20.0
sentencepiece>=0.1.99
soundfile>=0.12.1
torchaudio>=2.0.0
transformers==4.53.3Interactive Setup & Dependency Resolver
# Loading command...Frequently Asked Questions
How much VRAM does Step Audio EditX TTS require?
Step Audio EditX TTS requires a minimum of 1536MB (1.5GB) of VRAM for base operation. For optimal performance, a GPU with at least 6GB of VRAM is recommended. This node supports low VRAM mode for resource-constrained setups.
Can I run Step Audio EditX TTS on an RTX 3060, RTX 4070, or RTX 4090?
✅ RTX 3060 (12GB): Yes, fully compatible with 9.3GB headroom. ✅ RTX 4070 (12GB): Yes, fully compatible with 9.3GB headroom. ✅ RTX 4070 Ti (16GB): Yes, fully compatible with 12.9GB headroom. ✅ RTX 4090 (24GB): Yes, fully compatible with 20.1GB headroom
How much VRAM does Step Audio EditX TTS take on an RTX 3060 vs RTX 4090?
On an RTX 3060 (12GB VRAM), Step Audio EditX TTS runs smoothly on an RTX 3060 (12GB) with 9.3GB of headroom. This is sufficient to run the node alongside standard SD 1.5 and SDXL workflows in full precision. On an RTX 4090 (24GB VRAM), the node runs with extreme headroom on an RTX 4090 (24GB) with 20.1GB of dedicated headroom. This allows you to combine the node with massive models (like FLUX.1 Dev, Schnell, or Hunyuan Video) in full precision (FP16) without any offload flags.
What PyTorch version does Step Audio EditX TTS need?
Step Audio EditX TTS requires the following PyTorch-related packages: torchaudio>=2.0.0. Ensure your ComfyUI environment has these installed. Ensure your PyTorch installation matches your CUDA version (use torch.version.cuda to check).
What Python packages are required for Step Audio EditX TTS?
To run Step Audio EditX TTS, you need to install: accelerate>=0.20.0, gradio==5.49.1, hyperpyyaml>=1.2.2, librosa>=0.10.2, omegaconf>=2.3.0, onnxruntime>=1.17.0, openai-whisper>=20231117, protobuf>=3.20.0, sentencepiece>=0.1.99, soundfile>=0.12.1, torchaudio>=2.0.0, transformers==4.53.3. You can install these using pip or add them to your requirements.txt file.
How can I reduce VRAM usage when running Step Audio EditX TTS?
Step Audio EditX TTS supports low VRAM mode. To reduce memory usage: (1) Enable --lowvram or --medvram flags in ComfyUI, (2) Reduce batch size to 1, (3) Use fp16 or fp8 precision if supported, (4) Close other GPU applications.
How do I install Step Audio EditX TTS in ComfyUI?
To install Step Audio EditX TTS: (1) Navigate to your ComfyUI/custom_nodes directory, (2) Clone the repository: git clone https://github.com/Saganaki22/ComfyUI-Step_Audio_EditX_TTS, (3) Install dependencies: pip install -r requirements.txt (if present), (4) Restart ComfyUI. Alternatively, use ComfyUI Manager for one-click installation.