Addis PulseStudio

Kyutai Ships OVIE-512: One Image, One New View, 41.6 Milliseconds

A single-frame novel-view-synthesis model, MIT-licensed, with no diffusion steps and no ComfyUI node. Fast, small, and not a video generator.

3 min read756 words

What happened

Kyutai published OVIE-512, a PyTorch model that renders a single novel camera view from one 512×512 input image in 41.6 ms on an H100. It is MIT-licensed, distributed as safetensors, and carries no commercial restrictions.

Context

The OVIE family grew out of a paper (arXiv:2603.23488, "One View Is Enough: In-the-Wild Monocular Pretraining for Novel View Generation") arguing that in-the-wild monocular pretraining with MoGe-2 pseudo multi-view pairs is sufficient for high-quality novel view generation without explicit 3D scene reconstruction. The base OVIE checkpoint runs at 256×256 and completes a view in 8.6 ms. OVIE-512 is that same recipe retrained at 512×512, changing nothing else in the pipeline. It sits in the OVIE collection on Hugging Face alongside the 256 checkpoint, and the associated source lives at github.com/kyutai-labs/ovie.

How it works

OVIE-512 is a fully convolutional encoder-decoder with a ViT bottleneck. The bottleneck regenerates its positional encodings for whatever grid it is fed, which the authors describe as making the architecture resolution-agnostic by construction. In practice, the output resolution is fixed at training time: this checkpoint produces exactly a (1, 3, 512, 512) tensor in the [0, 1] range and cannot emit 768 or 1024 without a separate upscaler.

Inference is a single feed-forward pass. There are no diffusion iterations, no per-scene mesh or SDF construction, no scheduler. You load the model with OVIEModel.from_pretrained("kyutai/ovie-512"), pass an input image and a target camera pose (extrinsics and intrinsics encoded into a 7-element cam_token via extri_intri_to_pose_encoding), and get back one rendered view. Training ran for 250K steps at a global batch size of 256 on an in-the-wild image mix augmented with MoGe-2 pseudo-pairs. The eval script is run with uv run python evaluate.py.

Our read

The paper's title, "One View Is Enough," does more work than it should. It means the training needs only monocular images plus MoGe-2 pseudo-pairs, not that the output is a video. OVIE-512 renders exactly one new frame per forward pass. There is no temporal model, no frame-to-frame coherence mechanism. You can batch 14 target poses through the eval script, but each is an independent render; nothing constrains adjacent poses to look continuous.

The FID delta is the honest number in the card. At like-for-like 256×256 evaluation, OVIE-512 improves PSNR, SSIM, LPIPS, and MEt3R over the base model on both RealEstate10K and DL3DV, but costs roughly one FID point (7.62 vs 6.74 on RealEstate10K; 14.8 vs 13.6 on DL3DV). The authors attribute the loss to sub-512 images in the training mix being upscaled with little true high-frequency detail. That is a data-quality caveat, and it matters when someone quotes the PSNR gain without the FID column.

Sixty-two downloads and zero likes at collection time means no ComfyUI node, no third-party benchmark, no community wrapper. The MIT licence removes the usual "can I actually ship this" friction, but it does not create a workflow. The workflow still has to be built by hand.

What this changes

For a studio running local PyTorch models through ComfyUI, OVIE-512 adds a capability the node ecosystem does not cover: one new camera view from a single still, one forward pass, no diffusion sampling. The integration cost is real. There is no ComfyUI node; you wrap the from_pretrained call in a custom Python node, feed it extrinsics and intrinsics as a tensor, and pipe the 512×512 output into a super-resolution step. Minimum VRAM is not published; the reference hardware is an H100. If the task is "orbit a product from one captured frame," this is the closest off-the-shelf model. If the task is "generate a five-second clip," it is not that.

License

MIT. Commercial use is permitted with no revenue share, no attribution requirement beyond the standard MIT notice, and no model-specific usage clause. Safe to use in client deliverables.

Key takeaways

  • OVIE-512 renders a single novel view from one input image in 41.6 ms on an H100, with no diffusion iterations or per-scene geometry construction.
  • The architecture is resolution-agnostic in design but this checkpoint is fixed at 512×512; reaching 1080p or 4K requires a separate upscaler.
  • At 256×256 eval, OVIE-512 matches or improves on the base model across PSNR, SSIM, LPIPS, and MEt3R, at a cost of roughly one FID point.
  • The MIT licence permits unrestricted commercial use, but no ComfyUI node, community wrapper, or third-party benchmark exists yet.
  • The model is not a video generator; it produces independent frames with no temporal coherence between them.

Sources

  1. kyutai/ovie-512 — image-to-image on Hugging Face — tier 1
novel view synthesiskyutaicomputer vision

How this post was made

Drafted from clustered primary sources by the models below, then read, edited and approved by a human before it was published. The sources are listed in full at the end of the article.

Drafted
Independent sources
1
cluster pair
gemma4:12b
cluster label
gemma4:12b
radar brief
gemma4:12b
research brief
qwen3.8:27b
draft article
qwen3.8:27b
short script
qwen3.8:27b
seo pack
gemma4:12b
Run
editorial-20260930T161307Z