Addis PulseStudio

Wan 3.0 in ComfyUI: 30-Second Generation, Multi-Modal References, Cloud-Only Execution

Alibaba's newest video model generates 30 seconds in a single pass and conditions on up to 20 reference assets. The catch: the source describes only a cloud path, and no licence is named.

4 min read839 words

What happened

Alibaba's Wan 3.0 video generation model is live in ComfyUI as of 25 August 2026, producing up to 30 seconds of footage in a single pass with multi-modal reference conditioning. The ComfyUI blog positions it as a cloud workflow through Comfy Cloud, requiring ComfyUI v0.33.4 or higher.

Context

Video generation until now has meant short clips and cross-fade stitching for anything longer. The 480p draft tier is new in Wan 3.0, which implies prior Alibaba releases did not carry that lower-cost option. The multi-modal reference system (images, video, audio, a document) addressed with @-tokens in the prompt string marks a shift from single-image conditioning to a multi-asset compositional workflow. The ComfyUI integration frames this as a node-graph operation, but the source describes execution through Comfy Cloud. No local-weight download, VRAM requirement, or model parameter count appears in the brief.

How it works

Wan 3.0 accepts a prompt of up to 20,000 characters plus a set of reference assets: 10 images, 5 video clips, 5 audio clips, and one file or URL, all addressed in the prompt with @-tokens (@Image1, @Video2, @Audio1). Resolution is selectable at 480p, 720p, or 1080p. Aspect ratio covers 16:9, 9:16, 4:3, 3:4, 1:1, and an "adaptive" mode where the model selects the framing. Duration runs 2 to 30 seconds, with an "auto" mode that delegates length to the model. Audio generation can be toggled off.

Beyond generation, the model performs instruction-based editing on existing footage: removing or replacing elements, relighting, changing motion, transforming a subject, reshaping the narrative. It can also extend a clip, with input plus output capped at 30 seconds total. Pricing scales with resolution tier; no per-second or per-resolution figure is published.

The workflow runs through ComfyUI v0.33.4 nodes (Node Library or Templates panel) or via Comfy Cloud.

Our read

The 30-second single-pass is the headline. Three quieter details matter more.

The execution path is cloud. The source describes a Comfy Cloud workflow. No downloadable weights, no VRAM figure, no model file size, no GPU requirement. For a studio that has already stood up ComfyUI and a local inference stack on owned hardware, the question of whether this runs on the same box has no answer in the document. Until Alibaba publishes local weights and hardware specs, the 30-second pass is a cloud invoice, not a pipeline change.

The reference system is the actual feature. Ten images, five video clips, five audio clips, one document, all addressable with @-tokens in the prompt. A character-lock or camera-movement conditioning workflow that previously required a separate LoRA or IP-Adapter pass now happens inside the generation call. Fewer checkpoints to version, test, and maintain. The tuning loop shifts from "which checkpoint am I loading" to "which references do I put in the prompt."

The 30-second cap on edit-plus-extend is doing quiet work. A 20-second shot you can relight, swap an element in, and extend by 10. You cannot take a 45-second sequence and extend it. The cap makes this a short-clip refinement tool, not a general editor.

One documentation slip: the source says "up to 20 reference assets" while the itemised list sums to 21. Two assets from the ceiling and you need to know which number is the real limit.

What this changes

For a studio already on ComfyUI with cloud access: the 30-second single pass removes the cross-fade stitching step for short-form deliverables. The @-token reference system means character-lock and camera-movement conditioning no longer needs a separate LoRA or IP-Adapter node in the graph. The 480p tier gives a lower-cost draft for client previews. Instruction-based editing (relight, element swap, motion change) can absorb a slice of the manual rotoscoping or re-shoot budget, within the 30-second input-plus-output cap.

For a studio running a strictly local pipeline with no API calls: nothing changes yet. The source does not confirm a local-weight download path, and no hardware requirements are published. The ComfyUI node exists, but without weights to load, it is a placeholder.

License

No licence is named in the source. The legal terms for commercial use of Wan 3.0 outputs, redistribution, and whether reference uploads are retained or used for retraining are not stated. Check the model card and Comfy Cloud terms of service before building anything commercial on top of this pipeline.

Key takeaways

  • Wan 3.0 generates up to 30 seconds of video in a single pass at 480p, 720p, or 1080p, with audio generation toggleable off.
  • Multi-modal reference conditioning (10 images, 5 video clips, 5 audio clips, 1 document) is addressed via @-tokens in the prompt, replacing separate LoRA or IP-Adapter passes for character and motion conditioning.
  • The source describes a Comfy Cloud execution path; no local weights, VRAM requirements, or model file size are published.
  • Instruction-based video editing and footage extension are capped at 30 seconds total (input plus output).
  • No licence, per-resolution pricing, or terms of service are stated in the source.

Sources

  1. Wan 3.0 in ComfyUI: Native 30-Second Video with Omni-Reference Control — tier 1
video generationcomfyuialibabawan 3.0

How this post was made

Drafted from clustered primary sources by the models below, then read, edited and approved by a human before it was published. The sources are listed in full at the end of the article.

Drafted
Independent sources
1
cluster pair
gemma4:12b
cluster label
gemma4:12b
radar brief
gemma4:12b
research brief
qwen3.8:27b
draft article
qwen3.8:27b
short script
qwen3.8:27b
seo pack
gemma4:12b
Run
editorial-20260825T134800Z