01
Turbo and Lightning LoRAs
4-step and 8-step LoRAs cut the step count. ComfyUI uses 8 steps for text and image to video and 4 for reference to video.
Blog · October 9, 2026
Base MiniMax H3 uses many denoising steps, so one 5-second 768p clip takes minutes on a single GPU. Step distillation, sparse attention and a tuned serving stack bring that down to seconds. This page collects the published numbers and gives the source for each.
Steps
| Setup | Steps | Notes |
|---|---|---|
| MiniMax model card and README | Not stated | The weights are CFG-distilled: no guidance scale and no negative prompt |
| ComfyUI template | 20 | res_multistep sampler, simple scheduler. Simple shots hold up at 12–16 steps |
| SGLang cookbook | 50 | 50 sigma points = 49 model calls |
| ComfyUI Lightning LoRA | 8 or 4 | 8 for text and image to video, 4 for reference to video |
| FastH3 V2 | 8 model calls | Distilled with DMD2, 80% sparse video attention |
Because the weights are CFG-distilled, each step is one model call. Fewer steps means fewer model calls, and the model calls take most of the time.
Sources: MiniMax H3 model card, ComfyUI docs, SGLang cookbook, FastH3 V2 model card. Read October 9, 2026.
Generation time
| GPUs | Clip | Setup | Time |
|---|---|---|---|
| 1× B200 | 1344×768, 5 s / 10 s / 15 s | Base H3, 50 steps | 132.5 s / 377.4 s / 678.7 s |
| 4× H100 80GB | 1344×768, 5.17 s | Base H3, 50 steps, SGLang | 200.1 s |
| 4× H200 | 1344×768 | Base H3, 50 steps, SGLang | 74.4 s |
| 8× B300 | 10 s | Base H3, 50 steps, vLLM-Omni | 56.9 s |
| 2× RTX 5090 | 1344×768, 124 frames | Base H3, 50 steps, layer offload | 559.7 s |
Sources: Hao AI Lab (B200), Hyperstack (H100), SGLang cookbook (H200, RTX 5090), vLLM (B300). Read October 9, 2026.
Times change with the tool, resolution and clip length. Compare two rows only when the setup is the same.
FastH3
| GPU | 5 s clip, 832×480 | 5 s clip, 1344×768 |
|---|---|---|
| RTX PRO 6000 96GB | 13.5 s | 36.5 s |
| RTX 5090 32GB | 14.8 s | 38.6 s |
| RTX 4090 24GB | 54.6 s | 154.6 s |
FastH3 V2, end to end with audio, measured by Hao AI Lab (October 6, 2026). On one B200, FastH3 V1 renders a 5-second clip in 16.2 s against 132.5 s for base H3, and a 15-second clip in 47.2 s against 678.7 s (source). On eight B300 GPUs with vLLM-Omni, a 10-second clip takes about 8.7 s (source).
A Nuva Lab host with eight RTX PRO 6000 GPUs renders about 4,000 ten-second 768p clips with audio per day at full use. B200, B300 and GB200 hosts are available on request. FastH3 details.
FastH3 covers text to video with audio. Its model card says hard motion, fine detail and some audio can be below base H3. Test it on your own briefs.
Other methods
01
4-step and 8-step LoRAs cut the step count. ComfyUI uses 8 steps for text and image to video and 4 for reference to video.
02
vLLM-Omni online FP8 cut peak GPU memory by 38.9% on eight B300 GPUs. The stage time fell by 5.3%.
03
The MiniMax release has full attention only. FastH3 uses 80–90% sparse video attention. vLLM measured 9.8 s dense against 7.3 s sparse for a 10-second FastH3 clip.
04
SGLang and vLLM-Omni split one clip across GPUs with sequence parallelism. The GPU count must divide the 56 attention heads: 1, 2, 4, 7 or 8.
In production
Time per clip is one number. Production also needs batches, queues, retries and steady throughput across a whole day.
On a dedicated MiniMax H3 deployment, every GPU serves only your jobs, so a batch does not wait behind other customers. What the host costs.
FAQ
MiniMax does not state a default. The ComfyUI template uses 20 steps and the SGLang cookbook uses 50. FastH3 needs 8 model calls.
Base H3 takes about 132 seconds for a 5-second 1344×768 clip on one B200 GPU, and about 200 seconds on four H100 GPUs at 50 steps. FastH3 V1 takes 16.2 seconds for the same clip on one B200.
Use fewer steps with FastH3 or a Turbo LoRA, use FP8 weights, use sparse attention, or split the clip across more GPUs. Test each method on your own prompts, because each one can change the output.
It depends on the shot. The FastH3 model card says hard motion, fine detail and some audio can be below base H3. Test it on your own briefs.
Sources and further reading
Tell us your volume and clip formats. We'll size a dedicated deployment.
Get started · 1 min