Addis PulseStudio

Xiaomi's MiMo-V2.6-Flash-RL: 309B Parameters, 15B Active, MIT Licensed

A four-modality, 1M-context agent model that costs like a 15B model per token — and the first time a hardware company has put that combination under MIT.

4 min read849 words

What happened

Xiaomi's MiMo team published MiMo-V2.6-Flash-RL to Hugging Face on 21 September 2026. It is a 309-billion-parameter sparse-MoE omnimodal model handling text, image, video, and audio with a one-million-token context window, shipped as safetensors under MIT with 8-bit and fp8 quantisation variants alongside a larger Pro-RL sibling.

Context

The MiMo-V2.6 series is Xiaomi's in-house foundation model line. The model card calls this the "efficiency-balanced checkpoint" of the series, and the Pro-RL sibling published the same day confirms a two-tier release. What is new in V2.6 is the training recipe: rather than running separate RL passes per domain, the team mixed coding, agent, visual, and cybersecurity tasks into a single GRPO run they call "You Only RL Once." For a company that also sells phones and EVs, that single-pass design reads as a compute-budget decision as much as a technical one.

How it works

The core is a Sparse MoE architecture: 309 billion total parameters across many experts, but only 15 billion active per forward pass. Per-token compute is roughly that of a 15B dense model while retaining the parameter capacity of a frontier one. Vision runs through a 681-million-parameter MiMo ViT with 28 layers (24 sliding-window attention, 4 full). Audio goes through a 308-million-parameter AudioTokenizer plus a 127-million-parameter patch encoder. A five-layer speculative decoder handles Multi-Token Prediction to cut generation latency.

Training used fully asynchronous GRPO at 1,568 prompts × 16 rollouts per step, followed by Groupwise Reward Synthesis (offline rubric-building from contrasting rollouts) and Groupwise Advantage Redistribution (online re-ranking of passing trajectories). A final MOPD2 distillation stage mixes autonomous student rollouts with prefix-conditioned teacher and SFT rollouts. Benchmark scores from the model card: 67.9 on DeepSWE v1.1, 73.6 on Toolathlon-Verified, 52.3 on AutomationBench v1.0.6, 61.2 on MiMo Code Bench, 26.0 on ProgramBench.

Our read

The number that matters is not 309B total. It is 15B activated. That ratio puts per-token compute in the range of a mid-size dense model while retaining frontier-class capacity. "Flash" here means cheap per token, not small.

What the model card does not say, but the training recipe implies, is that the single mixed-RL pass is a cost decision. One GRPO sweep across four task families instead of four separate sweeps cuts training compute and wall-clock time substantially. "You Only RL Once" reads like a line item in a capex spreadsheet, not a research methodology name. The team is not optimising for a benchmark table; it is optimising for the GPU budget a phone-and-EV company can justify.

The second-order effect is where this matters for a small studio. A 309B, four-modality, 1M-context agent model under MIT is the first time a hardware company has put that combination on an open licence at a scale previously reserved for closed US hyperscaler APIs. It does not make the model runnable on a workstation: at fp8, 309B is roughly 300 GB of VRAM, 4×H100 territory. But it does make the planning and evaluation layer of a video pipeline callable without a per-token toll. For a studio running ComfyUI, the generation nodes stay the same. What changes is the agent that drives them.

What this changes

If you have 4×H100-80GB or 8×A100-80GB, you can pull the fp8 safetensors and run it locally through the transformers library. The 1M-token context is useful for long multi-session agent scripts that orchestrate a ComfyUI workflow without chunking. If you do not have that hardware, the model card references a Xiaomi MiMo API platform (platform.xiaomimimo.com). Pricing, rate limits, and SLA terms are not stated in either source, so that is a to-investigate item, not a Monday action. There is no ComfyUI node or community adapter for this model in either source; any integration means a custom HTTP bridge or transformers wrapper.

The honest answer for most small studios: nothing changes in your ComfyUI graph. This is an agent and generation model, not a diffusion model. It plans, captions, and evaluates video. It does not render it.

License

MIT. No commercial-use restriction, no model-size ceiling, no attribution requirement beyond the standard MIT notice. The sources do not state whether the licence covers all components (the 681M ViT, the AudioTokenizer) separately or as a single grant; check the model card's licence file before shipping.

Key takeaways

  • 309B total / 15B activated sparse-MoE with 1M-token context and four modalities (text, image, video, audio), published under MIT on 21 September 2026 alongside a Pro-RL sibling.
  • Training used a single mixed RL pass ("You Only RL Once") across coding, agent, visual, and cybersecurity tasks, followed by a GRS + GAR grading loop and MOPD2 distillation.
  • At fp8, inference requires roughly 300 GB of VRAM (4×H100-80GB or 8×A100-80GB minimum); single-workstation use is impractical.
  • Benchmarks from the model card: 67.9 DeepSWE v1.1, 73.6 Toolathlon-Verified, 52.3 AutomationBench v1.0.6, 61.2 MiMo Code Bench, 26.0 ProgramBench.
  • No ComfyUI integration exists in either source; the model serves as a planning and evaluation oracle for a video pipeline, not a generation node.

Sources

  1. XiaomiMiMo/MiMo-V2.6-Flash-RL — text-generation on Hugging Face — tier 1
  2. XiaomiMiMo/MiMo-V2.6-Flash-RL · Hugging Face — tier 3
xiaomimulti-modalopen-sourcemoe

How this post was made

Drafted from clustered primary sources by the models below, then read, edited and approved by a human before it was published. The sources are listed in full at the end of the article.

Drafted
Independent sources
2
cluster pair
gemma4:12b
cluster label
gemma4:12b
radar brief
gemma4:12b
research brief
qwen3.8:27b
draft article
qwen3.8:27b
short script
qwen3.8:27b
seo pack
gemma4:12b
Run
editorial-20260922T120134Z