Blog · October 9, 2026
Run MiniMax H3 locally: GPU, VRAM and RAM requirements
MiniMax H3 is open weight, so you can run it on your own hardware. The full model is large: about 124 GB of weights in full precision. Quantized builds run on consumer GPUs, but a 5-second clip can take many minutes, and system RAM becomes the limit.
Model size
How big is MiniMax H3?
| Part | Full precision | Smallest ComfyUI build |
|---|---|---|
| Transformer (33B) | 66.3 GB (BF16) | 16.0 GB (pruned W6A8) |
| Text encoder (Qwen3-VL-32B) | 51.5 GB (BF16) | 15.7 GB (NVFP4-AWQ) |
| Video VAE | 5.2 GB (FP16) | 5.2 GB |
| Audio VAE | 0.6 GB | 0.6 GB |
| Total | 123.6 GB | 42.5 GB |
Sources: Comfy-Org/MiniMax-H3 files and ComfyUI launch post. Read October 9, 2026.
VRAM
MiniMax H3 VRAM and RAM requirements
| GPU memory | Setup | 5 s clip |
|---|---|---|
| 4× 80 GB (H100) | Base H3, BF16, SGLang, 50 steps | 200.1 s at 1344×768 |
| 96 GB (RTX PRO 6000) | FastH3 V2, NVFP4 | 36.5 s at 1344×768 |
| 32 GB (RTX 5090) | FastH3 V2, NVFP4 | 38.6 s at 1344×768 |
| 24 GB (RTX 4090) | FastH3 V2, FP8 | 154.6 s at 1344×768 |
| 12 GB (capped RTX 4090) | FastH3 V2, FP8 | 170.7 s at 768p |
| 12 GB + 32 GB RAM | Base H3, SGLang consumer recipe, 20 steps | 319–356 s at 864×480 |
Sources: Hyperstack, Hao AI Lab, SGLang cookbook. Read October 9, 2026.
MiniMax gives no official minimum VRAM. Low VRAM moves weights to system RAM. The vLLM-Omni recipe for one RTX 5090 asks for at least 200 GiB of host RAM and advises 384 GiB. Community reports go down to 6–8 GB of VRAM with heavy offload, but a 5-second clip then takes 10 minutes or more.
Mac and multi-GPU
Mac, AMD and multi-GPU support
- Mac: ComfyUI did not ship Apple Silicon support at launch. The FastVideo MLX port runs FastH3 with 36 GB of unified memory or more. On an M4 Max, a 5-second 480p clip took about 925 seconds.
- Multi-GPU: SGLang and vLLM-Omni split H3 across GPUs. The GPU count must divide the 56 attention heads: 1, 2, 4, 7 or 8.
- Open weights output: a 768 px short side, 4–15 seconds, 24 fps. 2K output comes from H3-Regenerate-2K, which is API-only (MiniMax README).
Local vs dedicated host
When a dedicated host is the better choice
| Your own GPU | Nuva Lab dedicated host | |
|---|---|---|
| Best for | Tests and personal projects | Production volume for a team |
| Speed | Minutes per clip on one consumer GPU | About 4,000 ten-second 768p clips per day on 8× RTX PRO 6000 |
| Concurrency | One clip at a time | Every GPU serves your jobs |
| Customization | You run LoRA training yourself | We fine-tune and post-train on your data |
| Improvement | Manual | Data flywheel from your reviews |
| US companies | The H3 license excludes the US by default | MiniMax authorization for US companies |
Dedicated MiniMax H3 deployment · Cost per second at full use.
FAQ
Questions
How much VRAM does MiniMax H3 need?
MiniMax states no minimum. Full precision needs four 80 GB GPUs with SGLang. FastH3 and quantized builds run on 12–32 GB cards with offload, and they need a lot of system RAM.
Can I run MiniMax H3 on an RTX 4090?
Yes, with quantized weights. FastH3 V2 in FP8 renders a 5-second 1344×768 clip in about 155 seconds on an RTX 4090, as measured by Hao AI Lab.
Can MiniMax H3 run on a Mac?
Through the FastVideo MLX port, with 36 GB of unified memory or more. A 5-second 480p clip took about 925 seconds on an M4 Max.
Can a US company run MiniMax H3?
The MiniMax H3 Community License lists the US as an excluded territory, so US companies need authorization from MiniMax. Nuva Lab holds a MiniMax authorization for H3 deployments that serve US companies.
Sources and further reading
Read next
Skip the hardware work
Tell us your volume. We'll run MiniMax H3 on a private host for your team.
Get started · 1 min