V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation

By tiankuan93View on GitHub →

[Original] In the field of portrait video generation, the use of single images to generate portrait videos has become increasingly prevalent. A common approach involves leveraging generative models to enhance adapters for controlled generation. However, control signals can vary in strength, including text, audio, image reference, pose, depth map, etc. Among these, weaker conditions often struggle to be effective due to interference from stronger conditions, posing a challenge in balancing these conditions. In our work on portrait video generation, we identified audio signals as particularly weak, often overshadowed by stronger signals such as pose and original image. However, direct training with weak signals often leads to difficulties in convergence. To address this, we propose V-Express, a simple method that balances different control signals through a series of progressive drop operations. Our method gradually enables effective control by weak conditions, thereby achieving generation capabilities that simultaneously take into account pose, input image, and audio. NOTE: You need to downdload [a/model_ckpts](https://huggingface.co/tk93/V-Express/tree/main) manually.

How much VRAM does V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation require?

Direct Answer: The ComfyUI node V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation requires a minimum base VRAM of 4096MB and is optimized for GPUs with at least 8GB of VRAM. Low VRAM mode is not supported for this node.

Very High (4-8GB)
Base VRAM:
4096MB (4.0GB)
Recommended GPU:
8GB+ VRAM
Low VRAM Mode:
✗ Not supported
Estimation Confidence:
MEDIUM

Interactive VRAM Compatibility Estimator

Estimated Total VRAM: 3.00 GBTarget: 8 GB
✅ Comfortable Fit

Your GPU has plenty of headroom. You can run this node safely with your active configurations!

Deploy on High-Performance GPUs

Need more VRAM to run ComfyUI with this node? Rent low-cost, high-performance GPUs on Vast.ai instantly.

🚀 Deploy on Vast.ai

Deploy on Cloud GPUs

Need more VRAM to run ComfyUI with this node? Rent low-cost, high-performance cloud GPUs on RunPod instantly.

🚀 Deploy on RunPod

What Python packages are required for V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation?

Direct Answer: Running V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation requires installing the following Python package dependencies: av==11.0.0, diffusers==0.24.0, einops==0.4.1, imageio-ffmpeg==0.4.9, insightface==0.7.3, omegaconf==2.2.3, onnxruntime==1.16.3, safetensors==0.4.2, torch==2.0.1, torchaudio==2.0.2, torchvision==0.15.2, tqdm==4.66.1, transformers==4.30.2, xformers==0.0.22. Ensure your ComfyUI environment has these packages active before launching.

requirements.txt
av==11.0.0
diffusers==0.24.0
einops==0.4.1
imageio-ffmpeg==0.4.9
insightface==0.7.3
omegaconf==2.2.3
onnxruntime==1.16.3
safetensors==0.4.2
torch==2.0.1
torchaudio==2.0.2
torchvision==0.15.2
tqdm==4.66.1
transformers==4.30.2
xformers==0.0.22

Interactive Setup & Dependency Resolver

Operating System:
Environment Type:
Run this terminal command in your ComfyUI root folder:
# Loading command...

Frequently Asked Questions

How much VRAM does V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation require?

V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation requires a minimum of 4096MB (4.0GB) of VRAM for base operation. For optimal performance, a GPU with at least 8GB of VRAM is recommended. Low VRAM mode is not supported for this node.

Can I run V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation on an RTX 3060, RTX 4070, or RTX 4090?

✅ RTX 3060 (12GB): Yes, fully compatible with 6.8GB headroom. ✅ RTX 4070 (12GB): Yes, fully compatible with 6.8GB headroom. ✅ RTX 4070 Ti (16GB): Yes, fully compatible with 10.4GB headroom. ✅ RTX 4090 (24GB): Yes, fully compatible with 17.6GB headroom

What PyTorch version does V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation need?

V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation requires the following PyTorch-related packages: torch==2.0.1, torchaudio==2.0.2, torchvision==0.15.2. Ensure your ComfyUI environment has these installed. Ensure your PyTorch installation matches your CUDA version (use torch.version.cuda to check).

What Python packages are required for V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation?

To run V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation, you need to install: av==11.0.0, diffusers==0.24.0, einops==0.4.1, imageio-ffmpeg==0.4.9, insightface==0.7.3, omegaconf==2.2.3, onnxruntime==1.16.3, safetensors==0.4.2, torch==2.0.1, torchaudio==2.0.2, torchvision==0.15.2, tqdm==4.66.1, transformers==4.30.2, xformers==0.0.22. You can install these using pip or add them to your requirements.txt file.

How do I install V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation in ComfyUI?

To install V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation: (1) Navigate to your ComfyUI/custom_nodes directory, (2) Clone the repository: git clone https://github.com/tiankuan93/ComfyUI-V-Express, (3) Install dependencies: pip install -r requirements.txt (if present), (4) Restart ComfyUI. Alternatively, use ComfyUI Manager for one-click installation.