Addis PulseStudio

H3-Class Video Generation on a Single RTX 5090: 15 Seconds in 15, at 448 by 256

Comfy-Org's comfystreamerh3 node fits an 80 GB H3 video pipeline into 32 GB of consumer VRAM. The speed is real. The resolution is the catch.

4 min read792 words

What happened

Comfy-Org published comfystreamerh3, a ComfyUI node that runs H3-class video generation on a single NVIDIA RTX 5090, producing 15 seconds of 448 × 256 video in 15 seconds or less. The workflow stacks six hardware-level optimisations onto FastVideo's FastH3 V2 checkpoint to fit an 80 GB pipeline into 32 GB of VRAM.

Context

MiniMax's H3 Max model topped video-generation quality charts in the month before this write-up, roughly September 2026. FastH3 was developed specifically to optimise H3 for cost, and the author notes that running H3-class models continuously can cost upwards of 6,000 USD per day. FastVideo's Preview V1 generated a 15-second, 1344 × 768 clip in 47.2 seconds on one NVIDIA B200, which the author describes as "much more expensive" than an RTX 5090. The gap between a data-centre accelerator and a 32 GB consumer card is where this workflow sits.

How it works

The pipeline targets one RTX 5090 with 32 GB of VRAM and produces 448 × 256 clips at roughly 1× real-time. Six optimisations do the heavy lifting: four sampling steps per clip; VSA sparse attention that restricts full attention to 20 % of video blocks while still encoding the prompt; ClipProj, a NicoLab28 adapter that swaps H3's 32 B text encoder for a Qwen3-VL-4B, cutting encoder memory from about 15.7 GB to 4.5 GB; a pruned INT8 weight checkpoint from FastVideo; a fused FP4 MLP layer building on ByronLeeeee's work; and Kijai's INT8 video decoder that rebuilds frames in tiles, with NVENC encoding each finished tile while the decoder prepares the next. The workflow retains only the last decoded frame for the next clip rather than caching the full sequence. Combined, these reduce VRAM from roughly 80 GB to below 30 GB.

Our read

The headline is "15 seconds in 15 seconds," but the resolution is 448 × 256. That is not a deliverable. It is a streaming placeholder at best. The 6,000 USD-per-day cost the author cites is real, but it describes continuous high-resolution H3-class inference on data-centre hardware. A rented RTX 5090 at roughly 70 cents per hour changes the unit economics by about two orders of magnitude, at 448 × 256. The SolRefinery upscaling pass (7× pixel density at 3× GPU cost) is the bridge to something presentable, yet the source calls it preliminary testing with no stated availability or licence. The more important second-order effect is architectural: tile-based decode with overlapping NVENC encode, plus single-frame carry-over, turns a batch job into a streaming pipeline. That pattern — decode tile, encode tile, discard — makes "live" generation on consumer hardware plausible, and it is not specific to H3. Any video model with a tileable decoder benefits. The question nobody is asking: at what clip length does retaining only the last decoded frame start to bleed temporal coherence, and is there a known upper bound on segment length?

What this changes

If the studio already has an RTX 5090 or can rent one at 70 cents per hour, comfystreamerh3 is a drop-in ComfyUI node and the 15-second, 448 × 256 pipeline is testable this week without new infrastructure. The immediate use case is internal: prototyping prompts, testing model behaviour, building a live preview in a client dashboard. It is not a client-facing deliverable at that resolution. SolRefinery upscaling is the practical next step, but the source gives it as preliminary and names no licence, so it is not clear whether it is shippable. If the current GPU pool tops out at 24 GB, the workflow will not fit as described and there is no operational change on Monday.

License

The source describes comfystreamerh3 as "released open source" but does not name a specific licence. The licences for FastVideo's FastH3 V2 checkpoint, ClipProj, the FP4 MLP work, and Kijai's INT8 decoder are likewise not stated in the brief. Check each repository's licence file before embedding any of these in a client-deliverable pipeline.

Key takeaways

  • H3-class video generation now runs at roughly 1× real-time on a single 32 GB consumer GPU, but only at 448 × 256.
  • Six stacked optimisations (4 steps, sparse attention, 4B encoder swap, INT8 weights, FP4 MLP, tiled decode) cut VRAM from roughly 80 GB to under 30 GB.
  • The 6,000 USD/day figure applies to continuous high-end inference; a 70-cent-per-hour RTX 5090 rental is a different economic bracket entirely.
  • SolRefinery upscaling (7× pixels at 3× GPUs) is the path to usable resolution but is unverified and unlicensed in the source.
  • No specific licence is named for the ComfyUI node or any of the optimisation components; "open source" is not a licence.

Sources

  1. How I Generated Live Video with MiniMax H3 on a Single GPU — tier 1
video generationcomfyuirtx 5090gpu optimizationai hardware

How this post was made

Drafted from clustered primary sources by the models below, then read, edited and approved by a human before it was published. The sources are listed in full at the end of the article.

Drafted
Independent sources
1
cluster pair
clef:27b
cluster label
gemma4:12b
radar brief
gemma4:12b
research brief
qwen3.8:27b
draft article
qwen3.8:27b
short script
qwen3.8:27b
seo pack
gemma4:12b
Run
editorial-20261010T133900Z