GLM-4V Image Descriptor
Professional AI Image Description Generator Based on Zhipu AI GLM-4V multimodal model, batch generate accurate and detailed descriptions for images in Chinese and English
Quick Technical Summary: GLM-4V Image Descriptor
- Base VRAM Footprint:
- 128 MB (4 GB Tier)
- Primary Dependencies:
- accelerate>=0.21.0 # 模型加速库 / Model acceleration library, bitsandbytes>=0.41.0 # 量化支持库 / Quantization support library, csv# CSV处理库(Python内置) / CSV processing library (built-in Python), glob# 文件名模式匹配(Python内置) / Filename pattern matching (built-in Python), json# JSON处理库(Python内置) / JSON processing library (built-in Python), numpy>=1.21.0 # 数值计算库 / Numerical computing library, os# 操作系统接口(Python内置) / OS interface (built-in Python), pathlib# 路径处理库(Python 3.4+内置) / Path handling library (built-in Python 3.4+), pillow>=9.0.0 # Python图像处理库 / Python Image Library, protobuf>=3.20.0 # Protocol Buffers支持 / Protocol Buffers support, re# 正则表达式库(Python内置) / Regular expression library (built-in Python), sentencepiece>=0.1.97 # 文本分词器 / Text tokenizer, threading# 线程库(Python内置) / Threading library (built-in Python), time# 时间库(Python内置) / Time library (built-in Python), torch>=2.0.0 # PyTorch深度学习框架 / PyTorch deep learning framework, transformers==4.54.0 # Hugging Face Transformers库 / Hugging Face Transformers library
- Min PyTorch / CUDA:
- PyTorch 2.0.0 | CUDA 12.1+
- GitHub Repository:
- https://github.com/linjian-ufo/ComfyUI_GLM4V_voltspark
Citation Note: Data sourced from VRAM DB. For complete workflow OOM estimations, use the VRAM DB Workflow Analyzer.
How much VRAM does GLM-4V Image Descriptor require?
Direct Answer: The ComfyUI node GLM-4V Image Descriptor requires a minimum base VRAM of 128MB and is optimized for GPUs with at least 4GB of VRAM. Low VRAM mode is fully supported for resource-constrained setups.
- Base VRAM:
- 128MB (0.1GB)
- Recommended GPU:
- 4GB+ VRAM
- Low VRAM Mode:
- ✓ Supported
- Estimation Confidence:
- MEDIUM
Cheapest VRAM Upgrade Paths (Live Market Prices):
- GeForce RTX 3060 12GB (Ultimate Budget VRAM King)──► Used: $209.62View eBay ↗
- GeForce RTX 4060 8GB (Modern Entry-Level)──► New: $303.50View Amazon ↗
Interactive VRAM Compatibility Estimator
Your GPU has plenty of headroom. You can run this node safely with your active configurations!
Verify Compatibility for Your Specific GPU VRAM
Select your graphics card's VRAM capacity to view optimized batch sizes, suggested resolutions, and custom performance tips for GLM-4V Image Descriptor:
Buy NVIDIA GeForce RTX 3060 (12GB VRAM)
Tired of renting cloud rigs? Run ComfyUI locally with absolute zero latency. Best entry-level ComfyUI experience. Avoids immediate VRAM limitations on basic LoRA training.
Are you the author of this node?
Help your users avoid out-of-memory errors by displaying this professional, dynamic VRAM badge on your GitHub README. Copy the markdown below to embed it with a backlink directly to this hardware specification profile.
What Python packages are required for GLM-4V Image Descriptor?
Direct Answer: Running GLM-4V Image Descriptor requires installing the following Python package dependencies: accelerate>=0.21.0 # 模型加速库 / Model acceleration library, bitsandbytes>=0.41.0 # 量化支持库 / Quantization support library, csv# CSV处理库(Python内置) / CSV processing library (built-in Python), glob# 文件名模式匹配(Python内置) / Filename pattern matching (built-in Python), json# JSON处理库(Python内置) / JSON processing library (built-in Python), numpy>=1.21.0 # 数值计算库 / Numerical computing library, os# 操作系统接口(Python内置) / OS interface (built-in Python), pathlib# 路径处理库(Python 3.4+内置) / Path handling library (built-in Python 3.4+), pillow>=9.0.0 # Python图像处理库 / Python Image Library, protobuf>=3.20.0 # Protocol Buffers支持 / Protocol Buffers support, re# 正则表达式库(Python内置) / Regular expression library (built-in Python), sentencepiece>=0.1.97 # 文本分词器 / Text tokenizer, threading# 线程库(Python内置) / Threading library (built-in Python), time# 时间库(Python内置) / Time library (built-in Python), torch>=2.0.0 # PyTorch深度学习框架 / PyTorch deep learning framework, transformers==4.54.0 # Hugging Face Transformers库 / Hugging Face Transformers library. This node specifically requires PyTorch version 2.0.0 or newer.
accelerate>=0.21.0 # 模型加速库 / Model acceleration library
bitsandbytes>=0.41.0 # 量化支持库 / Quantization support library
csv# CSV处理库(Python内置) / CSV processing library (built-in Python)
glob# 文件名模式匹配(Python内置) / Filename pattern matching (built-in Python)
json# JSON处理库(Python内置) / JSON processing library (built-in Python)
numpy>=1.21.0 # 数值计算库 / Numerical computing library
os# 操作系统接口(Python内置) / OS interface (built-in Python)
pathlib# 路径处理库(Python 3.4+内置) / Path handling library (built-in Python 3.4+)
pillow>=9.0.0 # Python图像处理库 / Python Image Library
protobuf>=3.20.0 # Protocol Buffers支持 / Protocol Buffers support
re# 正则表达式库(Python内置) / Regular expression library (built-in Python)
sentencepiece>=0.1.97 # 文本分词器 / Text tokenizer
threading# 线程库(Python内置) / Threading library (built-in Python)
time# 时间库(Python内置) / Time library (built-in Python)
torch>=2.0.0 # PyTorch深度学习框架 / PyTorch deep learning framework
transformers==4.54.0 # Hugging Face Transformers库 / Hugging Face Transformers libraryInteractive Setup & Dependency Resolver
# Loading command...Special Environment Requirements:
Requires PyTorch version 2.0.0+.
# Loading PyTorch command...Frequently Asked Questions
How much VRAM does GLM-4V Image Descriptor require?
GLM-4V Image Descriptor requires a minimum of 128MB (0.1GB) of VRAM for base operation. For optimal performance, a GPU with at least 4GB of VRAM is recommended. This node supports low VRAM mode for resource-constrained setups.
Can I run GLM-4V Image Descriptor on an RTX 3060, RTX 4070, or RTX 4090?
✅ RTX 3060 (12GB): Yes, fully compatible with 10.7GB headroom. ✅ RTX 4070 (12GB): Yes, fully compatible with 10.7GB headroom. ✅ RTX 4070 Ti (16GB): Yes, fully compatible with 14.3GB headroom. ✅ RTX 4090 (24GB): Yes, fully compatible with 21.5GB headroom
How much VRAM does GLM-4V Image Descriptor take on an RTX 3060 vs RTX 4090?
On an RTX 3060 (12GB VRAM), GLM-4V Image Descriptor runs smoothly on an RTX 3060 (12GB) with 10.7GB of headroom. This is sufficient to run the node alongside standard SD 1.5 and SDXL workflows in full precision. On an RTX 4090 (24GB VRAM), the node runs with extreme headroom on an RTX 4090 (24GB) with 21.5GB of dedicated headroom. This allows you to combine the node with massive models (like FLUX.1 Dev, Schnell, or Hunyuan Video) in full precision (FP16) without any offload flags.
What PyTorch version does GLM-4V Image Descriptor need?
GLM-4V Image Descriptor requires the following PyTorch-related packages: torch>=2.0.0 # PyTorch深度学习框架 / PyTorch deep learning framework. Ensure your ComfyUI environment has these installed. Ensure your PyTorch installation matches your CUDA version (use torch.version.cuda to check).
What Python packages are required for GLM-4V Image Descriptor?
To run GLM-4V Image Descriptor, you need to install: accelerate>=0.21.0 # 模型加速库 / Model acceleration library, bitsandbytes>=0.41.0 # 量化支持库 / Quantization support library, csv# CSV处理库(Python内置) / CSV processing library (built-in Python), glob# 文件名模式匹配(Python内置) / Filename pattern matching (built-in Python), json# JSON处理库(Python内置) / JSON processing library (built-in Python), numpy>=1.21.0 # 数值计算库 / Numerical computing library, os# 操作系统接口(Python内置) / OS interface (built-in Python), pathlib# 路径处理库(Python 3.4+内置) / Path handling library (built-in Python 3.4+), pillow>=9.0.0 # Python图像处理库 / Python Image Library, protobuf>=3.20.0 # Protocol Buffers支持 / Protocol Buffers support, re# 正则表达式库(Python内置) / Regular expression library (built-in Python), sentencepiece>=0.1.97 # 文本分词器 / Text tokenizer, threading# 线程库(Python内置) / Threading library (built-in Python), time# 时间库(Python内置) / Time library (built-in Python), torch>=2.0.0 # PyTorch深度学习框架 / PyTorch deep learning framework, transformers==4.54.0 # Hugging Face Transformers库 / Hugging Face Transformers library. You can install these using pip or add them to your requirements.txt file.
How can I reduce VRAM usage when running GLM-4V Image Descriptor?
GLM-4V Image Descriptor supports low VRAM mode. To reduce memory usage: (1) Enable --lowvram or --medvram flags in ComfyUI, (2) Reduce batch size to 1, (3) Use fp16 or fp8 precision if supported, (4) Close other GPU applications.
How do I install GLM-4V Image Descriptor in ComfyUI?
To install GLM-4V Image Descriptor: (1) Navigate to your ComfyUI/custom_nodes directory, (2) Clone the repository: git clone https://github.com/linjian-ufo/ComfyUI_GLM4V_voltspark, (3) Install dependencies: pip install -r requirements.txt (if present), (4) Restart ComfyUI. Alternatively, use ComfyUI Manager for one-click installation.