Addis PulseStudio

XiaomiMiMo Ships a 309B MoE Model Targeted at One Specific Agentic Bug

MiMo-V2.6-Flash-MOPD is the MOPD-trained successor to the -RL checkpoint. Its stated purpose is not a benchmark bump β€” it is fixing tool-call repetition in agentic loops.

4 min read829 words

What happened

XiaomiMiMo published MiMo-V2.6-Flash-MOPD to the Hugging Face hub on 27 September 2026. It is the MOPD-trained successor to the MiMo-V2.6-Flash-RL checkpoint, built to suppress a diagnosed agentic failure mode: tool-call repetition.

Context

The MiMo V2.6 family already carries a Flash variant and a Pro variant. The -RL checkpoint was the reinforcement-learning stage. MOPD is the next training stage: rather than letting the model optimise its own trajectories, it distils from several domain-specialised teachers in a single pass. The tool-call repetition problem β€” an agent in a loop keeps issuing the same or near-identical tool calls without advancing β€” is a diagnosed failure mode in the -RL checkpoint that this release is designed to correct. MOPD2 extends the teacher set into domains where reliable training-time verification is hard: long-horizon game development, scientific research, embodied intelligence.

How it works

The architecture is a Sparse Mixture-of-Experts transformer: 309 billion total parameters, 15 billion activated per forward pass. Context window is 1M tokens. Input accepts four modalities β€” text, image, video, and audio β€” through a 681M-parameter MiMo ViT vision encoder (28 layers: 24 sliding-window attention plus 4 full attention) and a 308M AudioTokenizer paired with a 127M audio patch encoder. A 5-layer Multi-Token Prediction speculative decoder sits on top to cut inference latency.

Training draws from two teacher families: mixRL teachers trained on verifiable tasks and SFT teachers trained on synthetic demonstrations for open-domain work. Each parameter update blends three streams β€” Standard MOPD, Teacher-Prefix OPD, and SFT-Prefix OPD. The tool-call repetition fix is a short specialised-teacher run folded into the normal MOPD pass, not a separate training phase. Weights ship as safetensors in 8-bit and fp8 precision. Loading requires custom code; a stock transformers pipeline will not handle it.

Our read

The interesting move is not the architecture. A 309B-total / 15B-active MoE with four-modality input is a crowded specification by 2026. The interesting move is the failure-mode targeting. Most open-weight releases ship a benchmark table and leave the reader to map it to their use case. XiaomiMiMo names a specific agentic pathology, shows a targeted fix, and ships the checkpoint. That is closer to how a production team debugs a model: identify the worst loop, patch it, move on.

What the model card does not give is the magnitude of the fix. The repetition-rate improvement appears in a chart image (mimo-repetition-rate-flash-en.png) with no numeric axis values in the card text. A reader deciding whether this checkpoint is meaningfully better than the -RL version for an agentic pipeline has to open the image and eyeball it. Small thing, but it tells you the team optimised for the method story, not the reader's decision.

The 1M context and four-modality input are headline specs, but the practical bottleneck is the 309B total parameter count. Even with only 15B activated per pass, the full weight set must reside in memory. No VRAM figure is stated in any source. The 8-bit and fp8 tags confirm quantised checkpoints exist, but a 48 GB or 96 GB single-GPU setup is the realistic floor and likely still insufficient. The API at platform.xiaomimimo.com and the desktop client at mimo.xiaomimimo.com are the intended access paths, and no pricing is listed on the model card.

What this changes

For a ComfyUI-based studio, nothing changes on Monday. The model card documents no ComfyUI node, no API wrapper, and the custom_code tag means a stock text-generation node will not load it without a patched loader. If the studio is wiring an agentic video-editing loop β€” multi-step: scene detection, edit plan, execute, verify β€” the tool-call repetition fix is directly relevant and the XiaomiMiMo API is the entry point. For single-shot captioning, VQA, or scene description, the -RL checkpoint or a smaller model likely suffices and the MOPD delta is not the deciding factor. The MIT licence removes any legal friction if the studio does build a pipeline around it.

License

MiMo-V2.6-Flash-MOPD is released under the MIT licence. Commercial use, modification, and redistribution are permitted without restriction. There is no usage ceiling, no non-commercial clause, and no attribution requirement beyond the licence text itself.

Key takeaways

  • MiMo-V2.6-Flash-MOPD is a 309B-total / 15B-active Sparse MoE multimodal model (text, image, video, audio) with 1M-token context, published under MIT on 27 September 2026 by XiaomiMiMo.
  • Its specific training target is fixing tool-call repetition in agentic loops via on-policy distillation from mixRL and SFT teachers; the magnitude of improvement is shown only in a chart image with no stated numeric values.
  • Local inference is impractical on consumer hardware; the intended access path is the XiaomiMiMo API or desktop client, neither of which lists pricing on the model card.
  • No ComfyUI integration is documented, and the custom_code tag means loading requires a non-standard pipeline or API bridge.
  • The MIT licence permits unrestricted commercial use of the weights and any outputs derived from them.

Sources

  1. XiaomiMiMo/MiMo-V2.6-Flash-MOPD β€” text-generation on Hugging Face β€” tier 1
ai-newsagentic-aimultimodalxiaomimimo

How this post was made

Drafted from clustered primary sources by the models below, then read, edited and approved by a human before it was published. The sources are listed in full at the end of the article.

Drafted
Independent sources
1
cluster pair
gemma4:12b
cluster label
gemma4:12b
radar brief
gemma4:12b
research brief
qwen3.8:27b
draft article
qwen3.8:27b
short script
qwen3.8:27b
seo pack
gemma4:12b
Run
editorial-20260928T225029Z