Blog · October 9, 2026

Run MiniMax H3 locally: GPU, VRAM and RAM requirements

MiniMax H3 is open weight, so you can run it on your own hardware. The full model is large: about 124 GB of weights in full precision. Quantized builds run on consumer GPUs, but a 5-second clip can take many minutes, and system RAM becomes the limit.

Model size

How big is MiniMax H3?

PartFull precisionSmallest ComfyUI build
Transformer (33B)66.3 GB (BF16)16.0 GB (pruned W6A8)
Text encoder (Qwen3-VL-32B)51.5 GB (BF16)15.7 GB (NVFP4-AWQ)
Video VAE5.2 GB (FP16)5.2 GB
Audio VAE0.6 GB0.6 GB
Total123.6 GB42.5 GB

Sources: Comfy-Org/MiniMax-H3 files and ComfyUI launch post. Read October 9, 2026.

VRAM

MiniMax H3 VRAM and RAM requirements

GPU memorySetup5 s clip
4× 80 GB (H100)Base H3, BF16, SGLang, 50 steps200.1 s at 1344×768
96 GB (RTX PRO 6000)FastH3 V2, NVFP436.5 s at 1344×768
32 GB (RTX 5090)FastH3 V2, NVFP438.6 s at 1344×768
24 GB (RTX 4090)FastH3 V2, FP8154.6 s at 1344×768
12 GB (capped RTX 4090)FastH3 V2, FP8170.7 s at 768p
12 GB + 32 GB RAMBase H3, SGLang consumer recipe, 20 steps319–356 s at 864×480

Sources: Hyperstack, Hao AI Lab, SGLang cookbook. Read October 9, 2026.

MiniMax gives no official minimum VRAM. Low VRAM moves weights to system RAM. The vLLM-Omni recipe for one RTX 5090 asks for at least 200 GiB of host RAM and advises 384 GiB. Community reports go down to 6–8 GB of VRAM with heavy offload, but a 5-second clip then takes 10 minutes or more.

Mac and multi-GPU

Mac, AMD and multi-GPU support

  • Mac: ComfyUI did not ship Apple Silicon support at launch. The FastVideo MLX port runs FastH3 with 36 GB of unified memory or more. On an M4 Max, a 5-second 480p clip took about 925 seconds.
  • Multi-GPU: SGLang and vLLM-Omni split H3 across GPUs. The GPU count must divide the 56 attention heads: 1, 2, 4, 7 or 8.
  • Open weights output: a 768 px short side, 4–15 seconds, 24 fps. 2K output comes from H3-Regenerate-2K, which is API-only (MiniMax README).

Local vs dedicated host

When a dedicated host is the better choice

Your own GPUNuva Lab dedicated host
Best forTests and personal projectsProduction volume for a team
SpeedMinutes per clip on one consumer GPUAbout 4,000 ten-second 768p clips per day on 8× RTX PRO 6000
ConcurrencyOne clip at a timeEvery GPU serves your jobs
CustomizationYou run LoRA training yourselfWe fine-tune and post-train on your data
ImprovementManualData flywheel from your reviews
US companiesThe H3 license excludes the US by defaultMiniMax authorization for US companies

Dedicated MiniMax H3 deployment · Cost per second at full use.

FAQ

Questions

How much VRAM does MiniMax H3 need?

MiniMax states no minimum. Full precision needs four 80 GB GPUs with SGLang. FastH3 and quantized builds run on 12–32 GB cards with offload, and they need a lot of system RAM.

Can I run MiniMax H3 on an RTX 4090?

Yes, with quantized weights. FastH3 V2 in FP8 renders a 5-second 1344×768 clip in about 155 seconds on an RTX 4090, as measured by Hao AI Lab.

Can MiniMax H3 run on a Mac?

Through the FastVideo MLX port, with 36 GB of unified memory or more. A 5-second 480p clip took about 925 seconds on an M4 Max.

Can a US company run MiniMax H3?

The MiniMax H3 Community License lists the US as an excluded territory, so US companies need authorization from MiniMax. Nuva Lab holds a MiniMax authorization for H3 deployments that serve US companies.

Sources and further reading

Read next

Skip the hardware work

Tell us your volume. We'll run MiniMax H3 on a private host for your team.

Get started · 1 min