Xiaomi MiMo
The English MiMo blog introduces MiMo-V2-Flash as a globally available open-weight foundation language model. Xiaomi describes it as a 309B-total-parameter, 15B-active-parameter Mixture-of-Experts model with hybrid attention, hybrid thinking mode, 256K context, 150 tokens-per-second inference and low API pricing. The release is available through Hugging Face, the MiMo API Platform and MiMo AI Studio.
Editorial verdict
Teams evaluating a low-cost open-weight Chinese model for coding agents, long-context tasks and API/self-hosted comparison.
Avoid using release benchmarks as the only acceptance criterion for production workloads.
MiMo-V2-Flash gives MiMo a concrete English-documented open-weight anchor with price, context, test and deployment details.
$0.1 input / $0.3 output per 1M tokens listed in the English release blog
API Platform billing, AI Studio, Hugging Face self-hosting
Commercial use should follow the current product, API, model license and billing terms.
Review prompt, file, media upload, retention and training-use terms before sensitive workloads.
The release calls out Claude Code, Cursor and Cline-style coding workflows.
Use the published $0.1/$0.3 per-1M-token reference to compare against DeepSeek, Qwen and Kimi pricing.
Hugging Face weights and SGLang inference support make it relevant for local deployment experiments.
Model names, quotas, release status, regional access and commercial terms can change quickly; recheck official sources before procurement or production use.
deepseek-v4-api
qwen
kimi-k2-api
official · en · verified 2026-08-04
Official release with model architecture, pricing, tests, access paths and open-source notes.
other · en · verified 2026-05-17
Model weight entry linked from the release.
docs · en · verified 2026-05-17
Official hosted API access path linked from the release.
other · en · verified 2026-05-17
Official chat and studio access path linked from the release.
Last checked: 2026-08-04