Addis PulseStudio

Gemini Omni 1.1 Flash: Scene Extension, Keyframe Control, and a 360p Draft Tier

Google DeepMind shipped a generative video API update that jumps the scene-continuity context window tenfold, adds a one-third-cost draft tier, and introduces first-and-last-frame interpolation β€” all cloud-only, all on the Gemini API.

4 min read802 words

What happened

Google DeepMind shipped Gemini Omni 1.1 Flash on 2026-08-27, adding scene extension, keyframe interpolation, a 360p draft tier, and 4K output to its generative video API. The model is available via the Gemini API in Google AI Studio, the Enterprise Agent Platform, Google Flow, and the Gemini app for Plus, Pro, and Ultra subscribers.

Context

Prior Gemini Omni models referenced only the final second of prior video for scene continuity. The product line positions itself around bringing "real-world reasoning to generative creation." This release pushes that context window tenfold to ten seconds, introduces a 360p draft tier at one-third the cost of 720p, and adds first-and-last-frame keyframe control. The trajectory is from single-shot generation toward multi-shot workflows: the 40-second extension ceiling and the 3-second video-reference input both target continuity across a sequence, not just within one clip.

How it works

Gemini Omni 1.1 Flash is cloud-hosted and accessed through the Gemini API, Google AI Studio, the Enterprise Agent Platform, Google Flow, or the Gemini app. The model identifier is gemini-omni-1.1-flash.

Scene extension is the headline change. The model now analyzes up to 10 seconds of prior video before generating the next segment, up from 1 second. Extensions come in 10-second increments, capped at 40 seconds cumulative. First-and-last-frame keyframe interpolation lets you specify both endpoints and the model generates continuous video between them.

The 360p tier is a draft mode: up to 60% faster generation and one-third the cost of 720p, based on system throughput. Output scales to 1080p and 4K. A separate input path accepts up to 3 seconds of video as multimodal reference for visual context and character consistency. Neither source describes open weights, a local deployment path, or a self-hosted option.

Our read

The most useful feature is not the 4K output. It is the 360p draft tier. A studio storyboarding or A/B-testing camera angles can generate at one-third the cost of 720p and up to 60% faster. That is a budget valve local ComfyUI previews lack: a recognisable, motion-accurate draft without committing to a full-resolution render. For a small studio running five to ten generations a day in a storyboard phase, the savings compound.

The 10-second context window is a tenfold jump, but the 40-second ceiling is the operationally relevant number. The model can carry a continuous take through a 40-second sequence that local diffusion models, typically capped at 4–8 seconds, cannot match in a single pass. Beyond 40 seconds, stitching resumes. Keyframe interpolation is where the model earns its keep for a small crew: specify the start and end frames, hand it the camera move, and keep character and style work in the local pipeline. The 3-second video-reference input addresses a specific pain: locking a character's appearance across shots without re-prompting from scratch.

What the announcement omits matters as much as what it states. The pricing table is referenced but its contents are not reproduced. API rate limits, concurrent-request caps, and fair-use policies are absent. The 60% speed figure is attributed to system throughput with no peak-load qualifier or SLA. A studio planning a production schedule around these numbers is working with incomplete data.

What this changes

On Monday, the concrete shifts are: use the 360p tier for storyboarding and shot iteration before committing to 720p or 4K renders; assign camera-move shots such as orbits and whip-pans to keyframe interpolation in the cloud while keeping character consistency work in ComfyUI; and use the 3-second video-reference input to lock character appearance across a multi-shot sequence. Confirm whether the Plus subscriber tier is sufficient for production volume or whether the Agent Platform API is required. There is no local or offline deployment described, so every request is a cloud call with latency and egress costs. API rate limits are not stated in the sources, so a batch-rendering pipeline needs to be tested against actual throughput before scheduling around it.

License

Neither source states a licence for the model or for generated video output. The model is cloud-only with no open weights published. Check the model card in Google AI Studio before building anything commercial on Gemini Omni 1.1 Flash.

Key takeaways

  • Scene extension context jumped from 1 second to 10 seconds, with a 40-second cumulative ceiling in 10-second increments.
  • The 360p draft tier generates up to 60% faster and costs one-third of 720p, based on system throughput.
  • First-and-last-frame keyframe interpolation allows the model to generate continuous video between two specified endpoints.
  • The model is cloud-only; no open weights, local deployment path, or self-hosted option is described in either source.
  • Pricing, rate limits, and output licensing are not stated in the available sources.

Sources

  1. Gemini Omni 1.1 Flash lets you build with more control β€” tier 1
  2. Gemini Omni 1.1 Flash β€” tier 3
geminigooglegenerative videovideo production

How this post was made

Drafted from clustered primary sources by the models below, then read, edited and approved by a human before it was published. The sources are listed in full at the end of the article.

Drafted
Independent sources
2
cluster pair
gemma4:12b
cluster label
gemma4:12b
radar brief
gemma4:12b
research brief
qwen3.8:27b
draft article
qwen3.8:27b
short script
qwen3.8:27b
seo pack
gemma4:12b
Run
editorial-20260828T111017Z