Kyutai Ships OVIE-512: One Image, One New View, 41.6 Milliseconds
A single-frame novel-view-synthesis model, MIT-licensed, with no diffusion steps and no ComfyUI node. Fast, small, and not a video generator.
Category
A single-frame novel-view-synthesis model, MIT-licensed, with no diffusion steps and no ComfyUI node. Fast, small, and not a video generator.
MiMo-V2.6-Flash-MOPD is the MOPD-trained successor to the -RL checkpoint. Its stated purpose is not a benchmark bump β it is fixing tool-call repetition in agentic loops.
inclusionAI's unified multimodal model drops discrete quantization and modality-specific heads, but the two-turn training ceiling and custom inference path mean small studios are watching, not shipping.
A two-checkpoint video generation model with a 3D-aware memory architecture, no stated licence, no ComfyUI node, and no VRAM floor. The lab put it on the hub; the rest is on you.
A four-modality, 1M-context agent model that costs like a 15B model per token β and the first time a hardware company has put that combination under MIT.
Three quantized variants, a 256K context window, and an MIT licence that makes the orchestration story cleaner than the hardware story.
Granite Time Series PatchTST-FM-r2 tops the permissively-licensed zero-shot tier on GIFT-Eval. The conformer swap and the 99-quantile head matter more than the leaderboard position.
H Company pretrained a multimodal encoder from scratch instead of repurposing a generative VLM, and released the weights under Apache-2.0.
VibeVoice-ASR-Streaming-7B folds speaker attribution into the transcription pass, drops the diarisation step, and is MIT-licensed. No benchmarks yet.
K2 Horizon covers 0.9B to 375B under Apache-2.0, but the real release is the training data, intermediate checkpoints, and agentic post-training recipes published alongside.
Hy4-preview lands on Hugging Face with 49B activated parameters, a 1M-token window, and a licence that lets you ship. The benchmark data, however, is internal, and the memory math is the number that should matter more.
Google DeepMind shipped a generative video API update that jumps the scene-continuity context window tenfold, adds a one-third-cost draft tier, and introduces first-and-last-frame interpolation β all cloud-only, all on the Gemini API.
The MCP server that once only talked to a cloud instance now runs beside your local ComfyUI, reads your GPU, and can reach into Blender and DaVinci on the same file system. One agent session covers both paths.
How a new AWS robotics SDK handles cloud storage, frame decoding, and local inference for continuous training loops.
Open weights ship alongside a required ComfyUI version bump. Studio implications for local batching and fine-tuning.
The update achieves expert-level diagnostic performance against physicians by splitting perception and reasoning into parallel workers, validating a latency-preserving architecture for local AV stacks.
The latest runtime release drops legacy torch support, closes H3 memory leaks, and routes Qwen-Image and Grok directly through core.
Meta Superintelligence Labs ships a dense VLM with hybrid attention, DFlash drafting, and day-zero runtime support. Offline video ingestion is feasible on single GPUs, but integration gaps remain.
A new framework isolates visual and semantic domain gaps, enabling diffusion models to generate viable training frames without corrupting foreground structures.
A local-first omni-modal video model targets the RTX 3060 and collapses five production steps into a single pass.