Blog · October 9, 2026

MiniMax H3 inference speed: steps, generation time and how to make it faster

Base MiniMax H3 uses many denoising steps, so one 5-second 768p clip takes minutes on a single GPU. Step distillation, sparse attention and a tuned serving stack bring that down to seconds. This page collects the published numbers and gives the source for each.

Steps

How many steps does MiniMax H3 use?

SetupStepsNotes
MiniMax model card and READMENot statedThe weights are CFG-distilled: no guidance scale and no negative prompt
ComfyUI template20res_multistep sampler, simple scheduler. Simple shots hold up at 12–16 steps
SGLang cookbook5050 sigma points = 49 model calls
ComfyUI Lightning LoRA8 or 48 for text and image to video, 4 for reference to video
FastH3 V28 model callsDistilled with DMD2, 80% sparse video attention

Because the weights are CFG-distilled, each step is one model call. Fewer steps means fewer model calls, and the model calls take most of the time.

Sources: MiniMax H3 model card, ComfyUI docs, SGLang cookbook, FastH3 V2 model card. Read October 9, 2026.

Generation time

MiniMax H3 generation time by GPU

GPUsClipSetupTime
1× B2001344×768, 5 s / 10 s / 15 sBase H3, 50 steps132.5 s / 377.4 s / 678.7 s
4× H100 80GB1344×768, 5.17 sBase H3, 50 steps, SGLang200.1 s
4× H2001344×768Base H3, 50 steps, SGLang74.4 s
8× B30010 sBase H3, 50 steps, vLLM-Omni56.9 s
2× RTX 50901344×768, 124 framesBase H3, 50 steps, layer offload559.7 s

Sources: Hao AI Lab (B200), Hyperstack (H100), SGLang cookbook (H200, RTX 5090), vLLM (B300). Read October 9, 2026.

Times change with the tool, resolution and clip length. Compare two rows only when the setup is the same.

FastH3

FastH3: 8 model calls instead of 49

GPU5 s clip, 832×4805 s clip, 1344×768
RTX PRO 6000 96GB13.5 s36.5 s
RTX 5090 32GB14.8 s38.6 s
RTX 4090 24GB54.6 s154.6 s

FastH3 V2, end to end with audio, measured by Hao AI Lab (October 6, 2026). On one B200, FastH3 V1 renders a 5-second clip in 16.2 s against 132.5 s for base H3, and a 15-second clip in 47.2 s against 678.7 s (source). On eight B300 GPUs with vLLM-Omni, a 10-second clip takes about 8.7 s (source).

A Nuva Lab host with eight RTX PRO 6000 GPUs renders about 4,000 ten-second 768p clips with audio per day at full use. B200, B300 and GB200 hosts are available on request. FastH3 details.

FastH3 covers text to video with audio. Its model card says hard motion, fine detail and some audio can be below base H3. Test it on your own briefs.

Other methods

Other ways to speed up MiniMax H3

01

Turbo and Lightning LoRAs

4-step and 8-step LoRAs cut the step count. ComfyUI uses 8 steps for text and image to video and 4 for reference to video.

02

FP8 weights

vLLM-Omni online FP8 cut peak GPU memory by 38.9% on eight B300 GPUs. The stage time fell by 5.3%.

03

Sparse attention

The MiniMax release has full attention only. FastH3 uses 80–90% sparse video attention. vLLM measured 9.8 s dense against 7.3 s sparse for a 10-second FastH3 clip.

04

More GPUs

SGLang and vLLM-Omni split one clip across GPUs with sequence parallelism. The GPU count must divide the 56 attention heads: 1, 2, 4, 7 or 8.

In production

Production speed is a serving problem

Time per clip is one number. Production also needs batches, queues, retries and steady throughput across a whole day.

On a dedicated MiniMax H3 deployment, every GPU serves only your jobs, so a batch does not wait behind other customers. What the host costs.

FAQ

Questions

How many steps does MiniMax H3 use?

MiniMax does not state a default. The ComfyUI template uses 20 steps and the SGLang cookbook uses 50. FastH3 needs 8 model calls.

How long does MiniMax H3 take to generate a video?

Base H3 takes about 132 seconds for a 5-second 1344×768 clip on one B200 GPU, and about 200 seconds on four H100 GPUs at 50 steps. FastH3 V1 takes 16.2 seconds for the same clip on one B200.

How do I make MiniMax H3 faster?

Use fewer steps with FastH3 or a Turbo LoRA, use FP8 weights, use sparse attention, or split the clip across more GPUs. Test each method on your own prompts, because each one can change the output.

Is FastH3 as good as base MiniMax H3?

It depends on the shot. The FastH3 model card says hard motion, fine detail and some audio can be below base H3. Test it on your own briefs.

Sources and further reading

Read next

Get this speed on your own host

Tell us your volume and clip formats. We'll size a dedicated deployment.

Get started · 1 min