HomeAI NewsXiaomi’s MiMo-V2.6 is Rewriting the Rules of Open-Source AI

Xiaomi’s MiMo-V2.6 is Rewriting the Rules of Open-Source AI

Matching the industry’s heaviest hitters, Xiaomi’s newly released MiMo-V2.6 Pro and Flash models prove that radically scaled reinforcement learning—built entirely in public—is the future of artificial intelligence.

  • Unprecedented Open-Source Power: The natively omnimodal MiMo-V2.6 series rivals proprietary giants like Claude Opus 5 and GPT-5.6 Sol, with the Pro model achieving an industry-leading open-source score of 46.32 on the Artificial Analysis Intelligence Index.
  • Transparent, Scaled Reinforcement Learning: Built completely in the open, the models achieved massive performance leaps through sample-efficient, highly scalable reinforcement learning that fortified their coding, 3D reasoning, and multi-agent coordination.
  • Entering the “Vibe World”: Moving far beyond basic software engineering, MiMo-V2.6 operates as an autonomous co-creator capable of end-to-end 3D game development, orchestral music composition, dynamic video production, and advanced scientific research.

The artificial intelligence landscape has just experienced a seismic shift. Xiaomi has officially released and open-sourced the MiMo-V2.6 series, a suite of natively omnimodal models that push the Pareto frontier of intelligence and cost outward yet again. The release is anchored by two flagship models: MiMo-V2.6-Pro, the most capable iteration to date, and MiMo-V2.6-Flash, which strikes an optimal balance between efficiency, cost, and raw intelligence. For environments demanding extreme performance, Xiaomi has also introduced MiMo-V2.6-Pro-UltraSpeed, which delivers the same high-quality output up to 20 times faster. By maintaining the API pricing of the V2.5 series while dramatically increasing cognitive capabilities, Xiaomi is redefining what developers can expect from open-source AI.

The performance metrics speak for themselves. MiMo-V2.6-Pro currently stands as the strongest open-source model available, scoring an impressive 46.32 on the Artificial Analysis Intelligence Index and effortlessly surpassing competitors like Kimi K3 and Qwen3.8 Max. In practical applications, the Pro model performs on par with closed-source titans such as Claude Opus 5 and GPT-5.6 Sol across a wide array of agentic benchmarks. However, the true breakthrough lies in Xiaomi’s commitment to the RSI (Reinforcement-Scaled Intelligence) path. By scaling reinforcement learning (RL) compute on verifiable, highly complex tasks, the models continuously expand their capability frontiers through autonomous exploration and environmental feedback.

What makes this leap particularly remarkable is that the entire developmental journey was built in public. Xiaomi live-streamed the production run as they tackled immense research and engineering hurdles. In under six days, both the Flash and Pro models completed 30 RL steps over approximately 750,000 trajectories, costing just $0.85 million and $2.62 million, respectively. The resulting gains were staggering. Average pass rates on training tasks rose by up to 25% in relative terms, while scores on the held-out DeepSWE v1.1 software engineering benchmark skyrocketed—jumping from 48.8 to 65.68 for Flash, and from 58.4 to 72.57 for Pro.

These results were achieved by scaling RL compute across three critical axes. First, Xiaomi implemented larger batches and higher throughput on a fully asynchronous architecture, processing 1,568 samples per update with context lengths up to one million tokens. Second, they introduced a vastly richer multi-task training suite that blended coding, general agent tasks, visual processing, and cyber environments, allowing gains in one capability to naturally reinforce the others. Finally, they deployed significantly more grader compute, closing a self-improvement loop that steered the model toward shorter, more efficient reasoning paths. To ensure stability, the team froze the router to suppress training drift and built a robust defense against reward hacking using adversarial evaluation and anomaly detection. True to their open-source ethos, Xiaomi has released not just the model weights, but the full technical report, training environments, and RL code to the public.

This architectural mastery translates into awe-inspiring practical capabilities. Xiaomi has coined the term “Vibe World” to describe how MiMo-V2.6 extends natural-language programming from writing basic software to constructing fully interactive, multi-dimensional environments. In game development, the model can take a simple text prompt or image, decompose it into sub-tasks, coordinate multiple specialized agents to build 3D scenes, implement interaction logic, and iteratively refine the rendered output until a runnable interactive world is born.

Screenshot

The model’s creative fluency extends well into visual design and video production. MiMo-V2.6 can transform a basic instruction into a fully fleshed-out frontend interface or slide deck, boasting seamless Figma integration and impeccable typographic and aesthetic coherence. For video, it handles everything from shot sequencing and beat-synced editing to generating narration using MiMo-V2.5-TTS. It can even take abstract academic concepts, like Fourier decomposition, and translate them into accessible educational animations. Musically, MiMo-V2.6 has developed profound aesthetic judgment; when tasked with writing a 10-piece orchestral score, the Pro model not only composed the music but independently converted it to MIDI, demonstrating a deep understanding of instrumental arrangement and melodic division.

Beyond the creative arts, MiMo-V2.6-Pro is already proving to be a formidable co-scientist in cutting-edge research. In a recent collaboration with Xiaomi’s materials experts, the model was tasked with designing a novel metal-organic framework (MOF) material capable of adsorbing PFASs—notorious “forever chemicals.” Acting entirely as an autonomous research assistant, MiMo-V2.6-Pro scoured the web, reviewed patents, formulated hypotheses, and conducted “dry experiments.” It autonomously set up simulation environments and calculated binding strengths, effectively identifying highly promising candidates for physical lab testing. It is clear that the MiMo-V2.6 series is not just a tool for generating text, but an omnimodal engine for scientific, digital, and artistic creation.

Helen
Helen
Lead editor at Neuronad covering AI, machine learning, and emerging tech.

Must Read