LongCat AudioDiT TTS

By saganaki22View on GitHub →

ComfyUI custom nodes for LongCat-AudioDiT TTS - zero-shot and voice cloning diffusion TTS

How much VRAM does LongCat AudioDiT TTS require?

Direct Answer: The ComfyUI node LongCat AudioDiT TTS requires a minimum base VRAM of 1536MB and is optimized for GPUs with at least 6GB of VRAM. Low VRAM mode is fully supported for resource-constrained setups.

High (2-4GB)
Base VRAM:
1536MB (1.5GB)
Recommended GPU:
6GB+ VRAM
Low VRAM Mode:
✓ Supported
Estimation Confidence:
MEDIUM

Interactive VRAM Compatibility Estimator

Estimated Total VRAM: 3.00 GBTarget: 8 GB
✅ Comfortable Fit

Your GPU has plenty of headroom. You can run this node safely with your active configurations!

Deploy on High-Performance GPUs

Need more VRAM to run ComfyUI with this node? Rent low-cost, high-performance GPUs on Vast.ai instantly.

🚀 Deploy on Vast.ai

Deploy on Cloud GPUs

Need more VRAM to run ComfyUI with this node? Rent low-cost, high-performance cloud GPUs on RunPod instantly.

🚀 Deploy on RunPod

What Python packages are required for LongCat AudioDiT TTS?

Direct Answer: Running LongCat AudioDiT TTS requires installing the following Python package dependencies: einops>=0.7.0, huggingface-hub, librosa>=0.10.1, numpy, safetensors>=0.4.0, soundfile, transformers>=4.45.2. Ensure your ComfyUI environment has these packages active before launching.

requirements.txt
einops>=0.7.0
huggingface-hub
librosa>=0.10.1
numpy
safetensors>=0.4.0
soundfile
transformers>=4.45.2

Interactive Setup & Dependency Resolver

Operating System:
Environment Type:
Run this terminal command in your ComfyUI root folder:
# Loading command...

Frequently Asked Questions

How much VRAM does LongCat AudioDiT TTS require?

LongCat AudioDiT TTS requires a minimum of 1536MB (1.5GB) of VRAM for base operation. For optimal performance, a GPU with at least 6GB of VRAM is recommended. This node supports low VRAM mode for resource-constrained setups.

Can I run LongCat AudioDiT TTS on an RTX 3060, RTX 4070, or RTX 4090?

✅ RTX 3060 (12GB): Yes, fully compatible with 9.3GB headroom. ✅ RTX 4070 (12GB): Yes, fully compatible with 9.3GB headroom. ✅ RTX 4070 Ti (16GB): Yes, fully compatible with 12.9GB headroom. ✅ RTX 4090 (24GB): Yes, fully compatible with 20.1GB headroom

What Python packages are required for LongCat AudioDiT TTS?

To run LongCat AudioDiT TTS, you need to install: einops>=0.7.0, huggingface-hub, librosa>=0.10.1, numpy, safetensors>=0.4.0, soundfile, transformers>=4.45.2. You can install these using pip or add them to your requirements.txt file.

How can I reduce VRAM usage when running LongCat AudioDiT TTS?

LongCat AudioDiT TTS supports low VRAM mode. To reduce memory usage: (1) Enable --lowvram or --medvram flags in ComfyUI, (2) Reduce batch size to 1, (3) Use fp16 or fp8 precision if supported, (4) Close other GPU applications.

How do I install LongCat AudioDiT TTS in ComfyUI?

To install LongCat AudioDiT TTS: (1) Navigate to your ComfyUI/custom_nodes directory, (2) Clone the repository: git clone https://github.com/Saganaki22/ComfyUI-LongCat-AudioDIT-TTS, (3) Install dependencies: pip install -r requirements.txt (if present), (4) Restart ComfyUI. Alternatively, use ComfyUI Manager for one-click installation.