Xiaomi Ships MiMo-V2.6-Pro-RL: A 1.02T Agent Model Under MIT
A sparse MoE with native video and audio input, 1M-token context, and a single mixed GRPO loop across four domains. Strong on orchestration benchmarks, behind on raw coding. MIT-licensed.
What happened
Xiaomi's MiMo team published MiMo-V2.6-Pro-RL to Hugging Face on 21 September 2026, alongside a lighter MiMo-V2.6-Flash-RL. It is a 1.02-trillion-parameter sparse Mixture-of-Experts model with 42 billion active parameters, 1-million-token context, and native input handling for text, image, video, and audio.
Context
Xiaomi ships the MiMo series as open-weight models alongside its own API platform, desktop app, and studio. The "RL" suffix marks this as the reinforcement-learning-tuned variant. A technical report is linked from the model card. At collection time the page carried 109 likes and zero downloads — a first-day snapshot, not a weeks-in-the-wild read. The model supports English and Chinese.
How it works
The architecture is a sparse MoE: 1.02T total parameters, 42B activated per forward pass. Vision runs through a 681M-parameter MiMo ViT with 28 layers (24 sliding-window, 4 full attention). Audio goes through a 308M-parameter AudioTokenizer and a 127M-parameter patch encoder. A 5-layer speculative decoder handles Multi-Token Prediction to reduce generation latency.
Training used fully asynchronous GRPO at 1,568 prompts × 16 rollouts per step, in a single mixed run covering coding, general agents, visual tasks, and cybersecurity. Two reward components are novel: Groupwise Reward Synthesis builds task-specific rubrics offline from contrasting rollouts, and Groupwise Advantage Redistribution ranks passing trajectories online and shifts advantage toward higher-quality solutions. Post-RL, a multi-prefix distillation step (MOPD2) blends autonomous student rollouts with prefix-conditioned teacher and SFT rollouts.
The model ships as safetensors for the Hugging Face transformers library, with 8-bit and FP8 quantization options.
Our read
The obvious reading is "Xiaomi dropped a frontier model." The more useful one is in the benchmark spread. On AutomationBench v1.0.6, MiMo-V2.6 Pro scores 53.1, ahead of Claude Opus 5 (50.3), GPT-5.6 Sol (45.8), and Claude Fable 5 (46.2). On DeepSWE v1.1 it lands at 71.9, trailing Opus 5 (74.0) and GPT-5.6 Sol (73.0). On Terminal Bench 4.0 it scores 34.9 against Opus 5's 49.0. Strong at multi-step orchestration, competitive in coding, behind on terminal execution.
That ranking is what the model card does not say out loud. Xiaomi is not claiming best coder. It is claiming best agent: the model that plans, calls tools, verifies, and retries inside a 1M-token window. The single mixed GRPO loop across four domains, rather than separate per-domain passes, is the architectural bet. You do not mix cybersecurity and video understanding in one RL loop unless you want general-purpose orchestration.
The practical consequence: this is a director, not a worker. It tells a ComfyUI pipeline what to generate, checks the output against a brief, and iterates. It does not render frames. And the 'custom_code' tag with zero downloads at snapshot time means the serving path is not a drop-in pipeline() call; the exact launch recipe is not stated in any source.
What this changes
If you run ComfyUI locally and want a model that ingests raw footage and returns structured shot lists or continuity checks, this is now a candidate. Call the MiMo API, feed a clip sequence in one call (1M-token context), pipe the text output into your text-to-video graph.
What does not change: 42B active parameters put local serving beyond a single consumer GPU. The model is input-to-text; it does not generate frames. The 'custom_code' tag means a bespoke serving script, not a one-line pipeline() call.
The Monday move: hit the MiMo API with a video-parsing prompt on a clip already in your pipeline, compare the output to what you would write by hand. Close, wire it in. Vague, not yet.
License
The model card states the licence is MIT. Commercial use is permitted, including shipping products built on the model's outputs, with no revenue ceiling and no per-use restriction beyond the standard MIT attribution.
Key takeaways
- MiMo-V2.6-Pro-RL is a 1.02T-parameter sparse MoE (42B active) with native text/image/video/audio input and 1M-token context, published under MIT on 21 September 2026.
- Benchmark positioning is agentic orchestration: ahead of Claude Opus 5 and GPT-5.6 Sol on AutomationBench v1.0.6, behind on DeepSWE v1.1 and Terminal Bench 4.0.
- The single mixed GRPO run across four domains and the GRS/GAR reward system signal a general-purpose agent design, not a domain-specific fine-tune.
- Local serving is out of reach for single-GPU hardware; the practical path is the MiMo API or a rented node for the Flash variant, whose specs are not stated in the brief.
- The 'custom_code' tag and 0-download snapshot mean the inference path is not yet a standard
transformers.pipeline()call; expect a bespoke serving script.
Sources
How this post was made
Drafted from clustered primary sources by the models below, then read, edited and approved by a human before it was published. The sources are listed in full at the end of the article.
- Drafted
- Independent sources
- 1
- cluster pair
- gemma4:12b
- cluster label
- gemma4:12b
- radar brief
- gemma4:12b
- research brief
- qwen3.8:27b
- draft article
- qwen3.8:27b
- short script
- qwen3.8:27b
- seo pack
- gemma4:12b
- Run
- editorial-20260921T223712Z