Addis PulseStudio

MiniMax H3 drops open weights with native ComfyUI support

A local-first omni-modal video model targets the RTX 3060 and collapses five production steps into a single pass.

2 min read441 words

What happened

MiniMax released the third-generation H3 video model on August 3, 2026, and made the weights publicly open. ComfyUI added native support for the architecture in the same update window.

Context

The H3 replaces MiniMax’s previous Hailuo 01 and Hailuo 02 lines. Early text-to-video models required separate audio engines and cloud GPUs to reach usable resolution and duration. Moving to a unified omni-modal design with local optimization marks a direct shift from that pipeline.

How it works

H3 uses an omni-modal architecture that processes text, images, video, and audio inputs through a single prompt structure. It outputs 2K resolution clips up to 15 seconds long while generating synchronized stereo audio in the same inference pass. The model accepts first-and-last-frame constraints for shot framing and supports reference-to-video generation to transfer subject or motion from external media. MiniMax designed it to run locally on an NVIDIA RTX 3060.

Our read

The obvious headline here is the open weights and single-pass audio, but the real shift is hardware accessibility. Targeting the RTX 3060 removes the GPU tax that has forced small studios into cloud render farms or multi-card setups. Collapsing five production steps into one omni-modal pass does not eliminate post-production work; it just moves the bottleneck to prompt iteration and local VRAM management. The studio will care about how first-and-last-frame control behaves when chained with LongCat for lip-sync and Whisper for cleanup, but until we test node compatibility and exact memory overhead, this remains a promising spec sheet rather than a workflow replacement.

What this changes

On Monday, we will add the ComfyUI H3 nodes to our local instance and run a test batch on the RTX 3060. We will monitor VRAM consumption during stereo audio generation and measure whether first-and-last-frame control holds consistency across multiple passes. If the model stays under the GPU’s memory ceiling, we will replace our current text-to-video plus external audio sync step with a single ComfyUI graph. The tradeoff is accepting shorter 15-second clips in exchange for losing cloud rendering dependency.

Key takeaways

  • MiniMax H3 dropped on August 3, 2026 with open weights and native ComfyUI support.
  • The model outputs 2K video up to 15 seconds with synchronized stereo audio in one pass.
  • It is optimized for local execution on an NVIDIA RTX 3060 GPU.
  • First-and-last-frame control and reference-to-video replace external pre-visualization tools for shot iteration.
  • Exact VRAM limits, inference speed, context window, commercial licensing terms, and official cloud API pricing are not stated in any source.

Sources

  1. MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video — tier 1
ai newsvideo modelscomfyuiopen source aigpu optimizationstudio hardware

How this post was made

Drafted from clustered primary sources by the models below, then read, edited and approved by a human before it was published. The sources are listed in full at the end of the article.

Drafted
Independent sources
1
cluster pair
gemma4:e2b
cluster label
gemma4:e2b
research brief
qwen3.6:35b
draft article
qwen3.6:35b
seo pack
ornith:9b
Run
editorial-20260803T193710Z