Xiaomi MiMo-V2.6: What the New Open-Weight Models Offer

Reihen von Servern in einem Rechenzentrum
Photo by imgix on Unsplash

Xiaomi has released two new AI models, MiMo-V2.6-Pro and MiMo-V2.6-Flash, with downloadable weights. The flagship currently leads Artificial Analysis’ ranking of open-weight models. For developers, the combination of available weights, a low-cost API, and published training materials matters more than the leaderboard position alone.

Key takeaways

  • MiMo-V2.6-Pro scores 46 on the Artificial Analysis Intelligence Index and currently leads its open-weight category.
  • Xiaomi provides Pro and Flash weights under an MIT license, although self-hosting still requires substantial computing resources.
  • Through Xiaomi’s API, Pro costs $0.435 per million input tokens and $0.87 per million output tokens; Flash costs less.
  • The released training materials make Xiaomi’s methods easier to inspect, but they do not replace independent tests on real workloads.

What Xiaomi actually opened

The distinction between an open model and an open service matters. With an API alone, users send prompts to a provider and depend on its pricing and terms. MiMo-V2.6 weights, by contrast, are available on Hugging Face. The MIT license generally permits commercial use. Companies can run the model on their own or rented infrastructure and adapt it to their work. Releasing weights does not make the original training data fully transparent, however, or make deployment effortless.

Xiaomi describes MiMo-V2.6-Pro as a mixture-of-experts model. It has 1.02 trillion parameters in total, with 42 billion activated for each processing step. That reduces computation per token compared with a dense model of the same total size. The remaining weights still have to be kept available for inference. An ordinary laptop is therefore not a realistic home for the unmodified Pro model. Anyone interested in local AI on smaller hardware should check which smaller versions and quantizations are actually available, rather than assuming that downloadable weights mean a model will run at home.

According to Xiaomi’s model card, the model accepts text, images, audio, and video and offers a context window of up to one million tokens. A large context window is a technical limit, not a promise of reliable answers across a huge file or codebase. What matters is whether the model can retrieve relevant details and use tools over longer sequences of work. Xiaomi focused its training on those workflows: it mixed coding, general agent tasks, visual work, and cybersecurity in one reinforcement learning run, in which attempts receive feedback.

What the independent test does and does not show

Artificial Analysis currently gives MiMo-V2.6-Pro a score of 46 on its Intelligence Index and ranks it first among 114 open-weight models in that comparison. The measurement service lists Xiaomi’s API at $0.435 per million input tokens and $0.87 per million output tokens. An index combines different tasks into one number. It is a reason to run your own tests, not proof that MiMo is better at every coding task, language, or image analysis job. Rankings can also change as new models and measurement methods appear.

Xiaomi’s own results do not show a clean sweep over leading closed models either. On the DeepSWE v1.1 software test, Xiaomi reports 71.9 for Pro and 67.9 for Flash; the same table puts Claude Opus 5 at 74.0. On Terminal Bench 4.0, Xiaomi reports 34.9 for Pro and 49.0 for Claude Opus 5. Those are vendor-reported numbers, and test setup and settings can affect results. The independent index and the vendor benchmarks answer different questions; combining them into a single declaration of a winner would be misleading.

Pro, Flash, or self-hosting?

People who want to try the model can use MiMo Studio or Xiaomi’s API, according to the company. A useful first evaluation is a small set of your own tasks: representative documents, a multistep coding job, and a case where the model should explicitly acknowledge uncertainty. Measure not just correct answers but also latency, cost, data flow, and whether a mistake can be traced. Sensitive original data should go to an external service only after its terms and data location meet the organization’s requirements.

Flash may be the better economic choice for many repeated requests. Xiaomi lists $0.14 per million input tokens and $0.28 per million output tokens. In some agent tests published by Xiaomi, Flash trails Pro by only a few points. Teams planning a high-volume application can therefore test Flash and Pro on identical tasks first. That tells them more than a general comparison of list prices. The recent changes to Claude Opus 5.5 likewise show how quickly prices and capabilities are moving together.

Self-hosting offers greater control over data and model versions, but a model of this size requires specialized hardware, memory, and operational experience. The license does not remove those costs. Xiaomi’s release of a technical report, training code, and task environments is still notable: other teams can examine parts of the method and apply them to smaller models. Whether the reported gains can be reproduced in other settings remains an open research question.

The value is an offer people can test

MiMo-V2.6-Pro presents a new option: an open-weight model can rank near the top in independent testing while being offered at a low API price. That is not a universal instruction to switch. For readers and companies, the sensible next step is a controlled comparison using their own tasks and clear requirements for privacy, speed, and operations. The training release may prove to be the more lasting part of the story: if other groups can verify its benefits, the field gains both more model choices and better knowledge about how capable AI agents are built.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top