Dedicated deployment

Dedicated MiniMax H3 deployment for enterprise

Nuva Lab runs MiniMax H3 on a private host that serves only your company, tuned for speed. You get the full concurrency of the GPUs, a model customized on your work, agents that run production, and a system that improves with your review data.

What you get

A production system, not an endpoint

01

Private host

Only your workloads run on it. Your references, assets and review history stay with the deployment.

02

Full concurrency

Every GPU serves your jobs. Large batches do not wait in a shared queue.

03

Your model

We fine-tune and post-train H3 on your approved work and run your LoRAs. Fine-tuning and LoRA.

04

Speed

Eight RTX PRO 6000 GPUs render about 4,000 ten-second 768p clips with audio per day at full use. We tune the model and the serving stack for your workload. B200, B300 and GB200 hosts are available on request. Inference speed data.

05

Agents

The Creative Agent runs briefs, batches, retries and checks. The FDE Agent helps deploy and maintain the host. How the agents work.

06

Data flywheel

Your approvals and rejects guide the next tuning round. The loop runs inside your deployment and uses your data only.

Cost

One fixed cost: about $0.014 per second at full use

Illustrative GPU cost: eight RTX PRO 6000 GPUs at $3 per GPU-hour cost $576 per day. At full use they render about 4,000 ten-second 768p clips with audio per day, or about $0.0144 per second of video.

This is GPU cost only. It excludes service fees, storage and network. The MiniMax H3 API list price is $0.08 per second at 768P. Cost guide and calculator.

B200, B300 and GB200 hosts are available on request. Contact us for a quote.

US companies

Authorized for US companies

The MiniMax H3 Community License excludes the US by default. Nuva Lab holds a MiniMax authorization for H3 deployments that serve US companies.

Use of H3 and its derivatives follows the MiniMax H3 license.

Workloads

Built for production video

WorkloadWhat the deployment holds consistent
Performance adsProducts, logos, packaging and brand colors across many variants
Game marketingCharacters, art style and gameplay look across creative tests
Short dramaCharacters, wardrobe and setting across shots and episodes

How it starts

From first call to production

StepWhat happens
1. ScopeWe review your volume, clip formats and review criteria.
2. PilotWe deploy the model and test it on your briefs.
3. CustomizeWe tune on your approved work and measure the acceptance rate.
4. ProductionAgents run your workflows. We keep the deployment upgraded.

FAQ

Questions

What is a dedicated MiniMax H3 deployment?

A private host that runs MiniMax H3 for one company only, tuned for speed. It gives full concurrency, runs your customized weights and keeps your data inside the deployment.

How is it different from the MiniMax H3 API?

The API is a stateless, shared service that charges per generated second. A dedicated host is private, runs your fine-tuned model and improves with your review data.

Can US companies use MiniMax H3?

The MiniMax H3 Community License excludes the US by default. Nuva Lab holds a MiniMax authorization for H3 deployments that serve US companies.

Is my data used to improve models for other customers?

No. The feedback loop runs inside your deployment and uses your data only.

Which GPUs does the host use?

The standard host has eight NVIDIA RTX PRO 6000 GPUs. The speeds and prices on this site are for that host. For B200, B300 or GB200 hosts, contact us.

Sources and further reading

Read next

Get a dedicated MiniMax H3 deployment

Tell us your volume, formats and review process. We'll size a private host for it.

Get started · 1 min