Process multiple modalities
The Pro and Flash models handle text, images, audio, and video.
Productivity / product dossier
Open omnimodal models with 1M context, plus training code, RL environments and a technical report.
Product brief
MiMo-V2.6 is Xiaomi’s open omnimodal model family for long-horizon agent work. Pro and Flash handle text, images, audio and video with 1M context, while Xiaomi is also releasing the technical report, RL environments and training code behind the public post-training run.
Why we selected it
A technically substantive open model release: omnimodal, 1M-context models accompanied by training code, RL environments, and a technical report. It is especially relevant amid current interest in systems for long-horizo
Product preview saved with our daily selection
Capability scan
Only capabilities supported by the product information we collected are listed here.
The Pro and Flash models handle text, images, audio, and video.
The models are described as having a 1M context window for long-horizon agent work.
Xiaomi is releasing a technical report, RL environments, and training code behind its public post-training run.
Best-fit use cases
FAQ
MiMo-V2.6 is Xiaomi's open omnimodal model family for long-horizon agent work.
They handle text, images, audio, and video.
The description states that Pro and Flash have 1M context.
Xiaomi is releasing a technical report, RL environments, and training code for the public post-training run.