01
Accelerate
State-of-the-art inference optimization, tuned to your hardware and workload, so every video costs less and arrives sooner.
Flat cost. Dedicated to you.
FastH3 on your dedicated host. Your data, dedicated capacity, one flat hosting cost.
≈$0.014per second of 768p video
or≈4,000768p videos per day
Or feeling ambitious for real-time generation?
Illustrative GPU cost at full utilization: 8× RTX PRO 6000 × $3/hour/GPU; 10-second 768p videos with audio. Excludes service fees. API comparison: H3 768P $0.08/s; Seedance 2.0 720p $0.15/s (Sep 30, 2026).
We take open video models from weight to production
01
State-of-the-art inference optimization, tuned to your hardware and workload, so every video costs less and arrives sooner.
02
Adapt open models to your visual language, formats, and workflows while preserving quality and creative control.
03
Move from weights to a dependable production service with scalable infrastructure, observability, and hands-on support.
04
An agent that assists with, or fully automates, video asset generation, keeps your assets organized, and closes the data loop with your business so what performs shapes what gets made next.
We believe open video models will remain a critical part of the production stack. The Nuva Lab team has contributed to open-source AI for years. We optimize open-weight models for production today and will continue supporting and contributing to the models that come next.
Current Work
MiniMax H3, post-trained and optimized by FastVideo.
90% Sparse≈10× fewer target-video QK/PV block pairs
6.25× Fewer Steps50 → 8 sampling steps
Latency and ThroughputFull-host optimization
Blackwell Native PrecisionBuilt for NVIDIA Blackwell’s native FP4 acceleration
Open model work, production integrations, and ecosystem releases from Nuva Lab and our collaborators.
An optimized serving runtime for FastH3 on NVIDIA Blackwell. A 10-second clip with audio renders in 6.3 s on eight GB200 GPUs (5.4 s with optional NVFP4), with 1.2–1.7× lower median latency and 15–45% higher throughput than other open-source serving stacks. Already serving production traffic at Nuva Lab.
The eight-step FastH3 checkpoint that production serving now runs on, trained with data-free DMD2 distillation and 80% sparse video attention.
The FastVideo ecosystem brings FastH3 to NVIDIA DGX Spark and Apple Silicon for local video generation.
vLLM-Omni details the complete MiniMax H3 serving stack and its integration with FastVideo's four-step FastH3.
Open-weight, four-step sparse-distilled MiniMax H3, developed with FastVideo and NVIDIA collaborators and grounded in production workloads with Nuva Lab.
Flat cost. Dedicated to you. FastH3 on your dedicated host.