01
Accelerate
Reduce latency and serving costs with model-specific optimization tuned to your target hardware and production workload.
FastH3 technical pathProduction grounding and evaluation for FastH3: MiniMax H3, post-trained and optimized by FastVideo.
We take open video models from weight to production
01
Reduce latency and serving costs with model-specific optimization tuned to your target hardware and production workload.
FastH3 technical path02
Adapt open models to your visual language, formats, and workflows while preserving quality and creative control.
FastH3 post-training03
Move from weights to a dependable production service with scalable infrastructure, observability, and hands-on support.
FastH3 resultsWe believe open video models will remain a critical part of the production stack. The Nuva Lab team has contributed to open-source AI for years. We optimize open-weight models for production today and will continue supporting and contributing to the models that come next.
Current Work
MiniMax H3, post-trained and optimized by FastVideo.
90% Sparse≈10× fewer target-video QK/PV block pairs
12.5× Speedup50 → 4 denoising steps
Weights / Attention / Linear
Open model work, production integrations, and ecosystem releases from Nuva Lab and our collaborators.
The FastVideo ecosystem brings FastH3 to NVIDIA DGX Spark and Apple Silicon for local video generation.
vLLM-Omni details the complete MiniMax H3 serving stack and its integration with FastVideo's four-step FastH3.
Open-weight, four-step sparse-distilled MiniMax H3, developed with FastVideo and NVIDIA collaborators and grounded in production workloads with Nuva Lab.