TTS Audio Suite
TTS Audio Suite - Universal multi-engine TTS extension for ComfyUI with unified architecture supporting ChatterBox, F5-TTS, and future engines like RVC. Features modular engine adapters, character voice management, comprehensive SRT subtitle support, and advanced audio processing capabilities.
Quick Technical Summary: TTS Audio Suite
- Base VRAM Footprint:
- 1536 MB (6 GB Tier)
- Primary Dependencies:
- accelerate, av>=12.0.0 # DramaBox official audio reference decoding, bitsandbytes>=0.47.0 # 4-bit quantization support for VibeVoice memory efficiency, cn2an>=0.5.22 # Chinese number to Arabic number conversion, conformer>=0.3.2 # ChatterBox engine, dacite, datasets, echo-tts, einops>=0.7.0 # DramaBox LTX tensor rearrangement, ema-pytorch# F5-TTS exponential moving average, faiss-cpu>=1.7.4, ftfy# MOSS-SoundEffect v2 prompt normalization, funasr>=1.1.3 # FunASR speech processing toolkit, g2p-en>=2.1.0 # English grapheme-to-phoneme conversion, hyperpyyaml# YAML configuration parser, inflect>=7.3.0 # Text normalization for English (used in frontend.py), jieba, json5>=0.12.0 # JSON5 parsing for IndexTTS-2 config files, keras>=2.9.0 # Deep learning framework, langcodes, lingua-language-detector, loguru, modelscope>=1.27.0 # Chinese model hub for IndexTTS-2, monotonic-alignment-search, munch>=4.0.0 # Dictionary access with dot notation, nagisa>=0.2.11 # Japanese tokenizer required by Qwen3-ASR forced aligner, ninja>=1.11.0 # Build tool for CUDA kernel compilation (BigVGAN optimization), numpy>=1.26.4,<2.3.0 # Compatible with both numpy 1.26.4 and 2.x series, omegaconf>=2.3.0, openai-whisper# Mel spectrogram extraction for audio tokenizer, orjson>=3.11.0 # MOSS-TTS remote-code dependency, peft>=0.18.1 # MOSS-TTS LoRA adapter inference/training support (aLoRA/Arrow-era adapter configs), phonemizer# IPA phonemization for multilingual TTS (requires espeak system dependency), praat-parselmouth>=0.4.6 # Praat-based f0 extraction for RVC (pm method), protobuf>=3.20.0 # Protocol buffers (compatible with descript-audiotools), punctuators# ONNX punctuation/truecase post-processing for ASR text, pydub, pypinyin, pyworld>=0.3.5 # World vocoder for RVC harvest/dio methods, pyyaml>=6.0 # DramaBox/LTX configuration support, requests, safetensors>=0.6.2 # MOSS-TTS / HF safetensors checkpoints, sentencepiece>=0.2.1 # Text tokenization, setuptools>=65.0.0 # Provides distutils compatibility for Python 3.12+ (required by FunASR bundled code), sounddevice>=0.4.0, soundfile>=0.12.0, textstat>=0.7.10 # Text statistics and readability, tiktoken>=0.12.0 # MOSS-TTS tokenizer dependency, torch>=2.0.0, torchaudio>=2.0.0, torchdiffeq# F5-TTS differential equations, torchfcpe>=0.0.4 # Fast Context-based Pitch Estimation for RVC (fcpe method), transformers>=5.3.0,<6.0.0 # Main environment for Higgs Audio v3 / Step Audio EditX / MOSS-TTS. Legacy Transformers 4 engines use isolated runtimes., unidecode, vocos# F5-TTS vocoder, wandb# F5-TTS logging, webdataset, x-transformers
- Min PyTorch / CUDA:
- PyTorch 2.0.0 | CUDA 12.1+
- GitHub Repository:
- https://github.com/diodiogod/TTS-Audio-Suite
Citation Note: Data sourced from VRAM DB. For complete workflow OOM estimations, use the VRAM DB Workflow Analyzer.
How much VRAM does TTS Audio Suite require?
Direct Answer: The ComfyUI node TTS Audio Suite requires a minimum base VRAM of 1536MB and is optimized for GPUs with at least 6GB of VRAM. Low VRAM mode is fully supported for resource-constrained setups.
- Base VRAM:
- 1536MB (1.5GB)
- Recommended GPU:
- 6GB+ VRAM
- Low VRAM Mode:
- ✓ Supported
- Estimation Confidence:
- MEDIUM
Cheapest VRAM Upgrade Paths (Live Market Prices):
- GeForce RTX 3060 12GB (Ultimate Budget VRAM King)──► Used: $209.62View eBay ↗
- GeForce RTX 4060 8GB (Modern Entry-Level)──► New: $303.50View Amazon ↗
Interactive VRAM Compatibility Estimator
Your GPU has plenty of headroom. You can run this node safely with your active configurations!
Verify Compatibility for Your Specific GPU VRAM
Select your graphics card's VRAM capacity to view optimized batch sizes, suggested resolutions, and custom performance tips for TTS Audio Suite:
Buy NVIDIA GeForce RTX 3060 (12GB VRAM)
Tired of renting cloud rigs? Run ComfyUI locally with absolute zero latency. Best entry-level ComfyUI experience. Avoids immediate VRAM limitations on basic LoRA training.
Are you the author of this node?
Help your users avoid out-of-memory errors by displaying this professional, dynamic VRAM badge on your GitHub README. Copy the markdown below to embed it with a backlink directly to this hardware specification profile.
What Python packages are required for TTS Audio Suite?
Direct Answer: Running TTS Audio Suite requires installing the following Python package dependencies: accelerate, av>=12.0.0 # DramaBox official audio reference decoding, bitsandbytes>=0.47.0 # 4-bit quantization support for VibeVoice memory efficiency, cn2an>=0.5.22 # Chinese number to Arabic number conversion, conformer>=0.3.2 # ChatterBox engine, dacite, datasets, echo-tts, einops>=0.7.0 # DramaBox LTX tensor rearrangement, ema-pytorch# F5-TTS exponential moving average, faiss-cpu>=1.7.4, ftfy# MOSS-SoundEffect v2 prompt normalization, funasr>=1.1.3 # FunASR speech processing toolkit, g2p-en>=2.1.0 # English grapheme-to-phoneme conversion, hyperpyyaml# YAML configuration parser, inflect>=7.3.0 # Text normalization for English (used in frontend.py), jieba, json5>=0.12.0 # JSON5 parsing for IndexTTS-2 config files, keras>=2.9.0 # Deep learning framework, langcodes, lingua-language-detector, loguru, modelscope>=1.27.0 # Chinese model hub for IndexTTS-2, monotonic-alignment-search, munch>=4.0.0 # Dictionary access with dot notation, nagisa>=0.2.11 # Japanese tokenizer required by Qwen3-ASR forced aligner, ninja>=1.11.0 # Build tool for CUDA kernel compilation (BigVGAN optimization), numpy>=1.26.4,<2.3.0 # Compatible with both numpy 1.26.4 and 2.x series, omegaconf>=2.3.0, openai-whisper# Mel spectrogram extraction for audio tokenizer, orjson>=3.11.0 # MOSS-TTS remote-code dependency, peft>=0.18.1 # MOSS-TTS LoRA adapter inference/training support (aLoRA/Arrow-era adapter configs), phonemizer# IPA phonemization for multilingual TTS (requires espeak system dependency), praat-parselmouth>=0.4.6 # Praat-based f0 extraction for RVC (pm method), protobuf>=3.20.0 # Protocol buffers (compatible with descript-audiotools), punctuators# ONNX punctuation/truecase post-processing for ASR text, pydub, pypinyin, pyworld>=0.3.5 # World vocoder for RVC harvest/dio methods, pyyaml>=6.0 # DramaBox/LTX configuration support, requests, safetensors>=0.6.2 # MOSS-TTS / HF safetensors checkpoints, sentencepiece>=0.2.1 # Text tokenization, setuptools>=65.0.0 # Provides distutils compatibility for Python 3.12+ (required by FunASR bundled code), sounddevice>=0.4.0, soundfile>=0.12.0, textstat>=0.7.10 # Text statistics and readability, tiktoken>=0.12.0 # MOSS-TTS tokenizer dependency, torch>=2.0.0, torchaudio>=2.0.0, torchdiffeq# F5-TTS differential equations, torchfcpe>=0.0.4 # Fast Context-based Pitch Estimation for RVC (fcpe method), transformers>=5.3.0,<6.0.0 # Main environment for Higgs Audio v3 / Step Audio EditX / MOSS-TTS. Legacy Transformers 4 engines use isolated runtimes., unidecode, vocos# F5-TTS vocoder, wandb# F5-TTS logging, webdataset, x-transformers. This node specifically requires PyTorch version 2.0.0 or newer.
accelerate
av>=12.0.0 # DramaBox official audio reference decoding
bitsandbytes>=0.47.0 # 4-bit quantization support for VibeVoice memory efficiency
cn2an>=0.5.22 # Chinese number to Arabic number conversion
conformer>=0.3.2 # ChatterBox engine
dacite
datasets
echo-tts
einops>=0.7.0 # DramaBox LTX tensor rearrangement
ema-pytorch# F5-TTS exponential moving average
faiss-cpu>=1.7.4
ftfy# MOSS-SoundEffect v2 prompt normalization
funasr>=1.1.3 # FunASR speech processing toolkit
g2p-en>=2.1.0 # English grapheme-to-phoneme conversion
hyperpyyaml# YAML configuration parser
inflect>=7.3.0 # Text normalization for English (used in frontend.py)
jieba
json5>=0.12.0 # JSON5 parsing for IndexTTS-2 config files
keras>=2.9.0 # Deep learning framework
langcodes
lingua-language-detector
loguru
modelscope>=1.27.0 # Chinese model hub for IndexTTS-2
monotonic-alignment-search
munch>=4.0.0 # Dictionary access with dot notation
nagisa>=0.2.11 # Japanese tokenizer required by Qwen3-ASR forced aligner
ninja>=1.11.0 # Build tool for CUDA kernel compilation (BigVGAN optimization)
numpy>=1.26.4,<2.3.0 # Compatible with both numpy 1.26.4 and 2.x series
omegaconf>=2.3.0
openai-whisper# Mel spectrogram extraction for audio tokenizer
orjson>=3.11.0 # MOSS-TTS remote-code dependency
peft>=0.18.1 # MOSS-TTS LoRA adapter inference/training support (aLoRA/Arrow-era adapter configs)
phonemizer# IPA phonemization for multilingual TTS (requires espeak system dependency)
praat-parselmouth>=0.4.6 # Praat-based f0 extraction for RVC (pm method)
protobuf>=3.20.0 # Protocol buffers (compatible with descript-audiotools)
punctuators# ONNX punctuation/truecase post-processing for ASR text
pydub
pypinyin
pyworld>=0.3.5 # World vocoder for RVC harvest/dio methods
pyyaml>=6.0 # DramaBox/LTX configuration support
requests
safetensors>=0.6.2 # MOSS-TTS / HF safetensors checkpoints
sentencepiece>=0.2.1 # Text tokenization
setuptools>=65.0.0 # Provides distutils compatibility for Python 3.12+ (required by FunASR bundled code)
sounddevice>=0.4.0
soundfile>=0.12.0
textstat>=0.7.10 # Text statistics and readability
tiktoken>=0.12.0 # MOSS-TTS tokenizer dependency
torch>=2.0.0
torchaudio>=2.0.0
torchdiffeq# F5-TTS differential equations
torchfcpe>=0.0.4 # Fast Context-based Pitch Estimation for RVC (fcpe method)
transformers>=5.3.0,<6.0.0 # Main environment for Higgs Audio v3 / Step Audio EditX / MOSS-TTS. Legacy Transformers 4 engines use isolated runtimes.
unidecode
vocos# F5-TTS vocoder
wandb# F5-TTS logging
webdataset
x-transformersInteractive Setup & Dependency Resolver
# Loading command...Special Environment Requirements:
Requires PyTorch version 2.0.0+.
# Loading PyTorch command...Frequently Asked Questions
How much VRAM does TTS Audio Suite require?
TTS Audio Suite requires a minimum of 1536MB (1.5GB) of VRAM for base operation. For optimal performance, a GPU with at least 6GB of VRAM is recommended. This node supports low VRAM mode for resource-constrained setups.
Can I run TTS Audio Suite on an RTX 3060, RTX 4070, or RTX 4090?
✅ RTX 3060 (12GB): Yes, fully compatible with 9.3GB headroom. ✅ RTX 4070 (12GB): Yes, fully compatible with 9.3GB headroom. ✅ RTX 4070 Ti (16GB): Yes, fully compatible with 12.9GB headroom. ✅ RTX 4090 (24GB): Yes, fully compatible with 20.1GB headroom
How much VRAM does TTS Audio Suite take on an RTX 3060 vs RTX 4090?
On an RTX 3060 (12GB VRAM), TTS Audio Suite runs smoothly on an RTX 3060 (12GB) with 9.3GB of headroom. This is sufficient to run the node alongside standard SD 1.5 and SDXL workflows in full precision. On an RTX 4090 (24GB VRAM), the node runs with extreme headroom on an RTX 4090 (24GB) with 20.1GB of dedicated headroom. This allows you to combine the node with massive models (like FLUX.1 Dev, Schnell, or Hunyuan Video) in full precision (FP16) without any offload flags.
What PyTorch version does TTS Audio Suite need?
TTS Audio Suite requires the following PyTorch-related packages: ema-pytorch# F5-TTS exponential moving average, torch>=2.0.0, torchaudio>=2.0.0, torchdiffeq# F5-TTS differential equations, torchfcpe>=0.0.4 # Fast Context-based Pitch Estimation for RVC (fcpe method). Ensure your ComfyUI environment has these installed. Ensure your PyTorch installation matches your CUDA version (use torch.version.cuda to check).
What Python packages are required for TTS Audio Suite?
To run TTS Audio Suite, you need to install: accelerate, av>=12.0.0 # DramaBox official audio reference decoding, bitsandbytes>=0.47.0 # 4-bit quantization support for VibeVoice memory efficiency, cn2an>=0.5.22 # Chinese number to Arabic number conversion, conformer>=0.3.2 # ChatterBox engine, dacite, datasets, echo-tts, einops>=0.7.0 # DramaBox LTX tensor rearrangement, ema-pytorch# F5-TTS exponential moving average, faiss-cpu>=1.7.4, ftfy# MOSS-SoundEffect v2 prompt normalization, funasr>=1.1.3 # FunASR speech processing toolkit, g2p-en>=2.1.0 # English grapheme-to-phoneme conversion, hyperpyyaml# YAML configuration parser, inflect>=7.3.0 # Text normalization for English (used in frontend.py), jieba, json5>=0.12.0 # JSON5 parsing for IndexTTS-2 config files, keras>=2.9.0 # Deep learning framework, langcodes, lingua-language-detector, loguru, modelscope>=1.27.0 # Chinese model hub for IndexTTS-2, monotonic-alignment-search, munch>=4.0.0 # Dictionary access with dot notation, nagisa>=0.2.11 # Japanese tokenizer required by Qwen3-ASR forced aligner, ninja>=1.11.0 # Build tool for CUDA kernel compilation (BigVGAN optimization), numpy>=1.26.4,<2.3.0 # Compatible with both numpy 1.26.4 and 2.x series, omegaconf>=2.3.0, openai-whisper# Mel spectrogram extraction for audio tokenizer, orjson>=3.11.0 # MOSS-TTS remote-code dependency, peft>=0.18.1 # MOSS-TTS LoRA adapter inference/training support (aLoRA/Arrow-era adapter configs), phonemizer# IPA phonemization for multilingual TTS (requires espeak system dependency), praat-parselmouth>=0.4.6 # Praat-based f0 extraction for RVC (pm method), protobuf>=3.20.0 # Protocol buffers (compatible with descript-audiotools), punctuators# ONNX punctuation/truecase post-processing for ASR text, pydub, pypinyin, pyworld>=0.3.5 # World vocoder for RVC harvest/dio methods, pyyaml>=6.0 # DramaBox/LTX configuration support, requests, safetensors>=0.6.2 # MOSS-TTS / HF safetensors checkpoints, sentencepiece>=0.2.1 # Text tokenization, setuptools>=65.0.0 # Provides distutils compatibility for Python 3.12+ (required by FunASR bundled code), sounddevice>=0.4.0, soundfile>=0.12.0, textstat>=0.7.10 # Text statistics and readability, tiktoken>=0.12.0 # MOSS-TTS tokenizer dependency, torch>=2.0.0, torchaudio>=2.0.0, torchdiffeq# F5-TTS differential equations, torchfcpe>=0.0.4 # Fast Context-based Pitch Estimation for RVC (fcpe method), transformers>=5.3.0,<6.0.0 # Main environment for Higgs Audio v3 / Step Audio EditX / MOSS-TTS. Legacy Transformers 4 engines use isolated runtimes., unidecode, vocos# F5-TTS vocoder, wandb# F5-TTS logging, webdataset, x-transformers. You can install these using pip or add them to your requirements.txt file.
How can I reduce VRAM usage when running TTS Audio Suite?
TTS Audio Suite supports low VRAM mode. To reduce memory usage: (1) Enable --lowvram or --medvram flags in ComfyUI, (2) Reduce batch size to 1, (3) Use fp16 or fp8 precision if supported, (4) Close other GPU applications.
How do I install TTS Audio Suite in ComfyUI?
To install TTS Audio Suite: (1) Navigate to your ComfyUI/custom_nodes directory, (2) Clone the repository: git clone https://github.com/diodiogod/TTS-Audio-Suite, (3) Install dependencies: pip install -r requirements.txt (if present), (4) Restart ComfyUI. Alternatively, use ComfyUI Manager for one-click installation.