MiniMax
MiniMax M3 was released on June 1, 2026 as MiniMax's flagship for coding, agents and native image/video understanding. Its MiniMax Sparse Attention architecture supports up to a 1M-token API context window with a guaranteed minimum of 512K. MiniMax reports that at 1M context its per-token compute is 1/20 of the previous generation, with more than 9x faster prefill and 15x faster decoding. The API uses the MiniMax-M3 model identifier, supports thinking and non-thinking modes at the same price, applies a higher long-context rate above 512K input tokens, and offers an optional priority service tier for SLA-sensitive workloads. MiniMax calls M3 open-weight, but the checked model page also says full Hugging Face and GitHub publication is coming soon, so current weight availability should be verified separately.
Editorial verdict
Developers evaluating a China-origin frontier model for coding, long-context agentic work and multimodal reasoning.
Avoid it if you need fixed public pricing or a model with fully stable commercial terms already locked in.
M3 should be tracked separately because the official homepage now gives it a distinct model page and positions it as the current frontier coding model.
API & Token Plan access; current pricing should be checked in account
API & Token Plan, Pay-as-you-go API billing, Enterprise sales
Commercial use should follow the current product, API, model license and billing terms.
Review prompt, file, media upload, retention and training-use terms before sensitive workloads.
Evaluate repository repair, terminal execution, kernel optimization and MCP tool use, while reproducing vendor benchmark settings independently.
The 1M context window makes it relevant for large codebases, long documents and multi-step agent tasks.
Use native image and video input for document understanding, visual coding tasks and computer-use agents.
The release reports 12-hour paper reproduction and 24-hour CUDA optimization runs; validate cost, reliability and recovery behavior on your own tasks.
Rechecked the official model page on June 21, 2026. The release article is a dated product snapshot, the live Token Plan can change quotas, and the model page still describes full Hugging Face and GitHub publication as coming soon. Recheck current prices, quotas, weights and priority-tier availability before rollout.
minimax-api
qwen
deepseek
zhipu-glm
kimi
official · en · verified 2026-08-04
Confirms coding and agent positioning, guaranteed 512K and maximum 1M API context, native multimodality, BrowseComp score, long-horizon demonstrations, API identifier and pending full open-source publication.
official · en · verified 2026-06-10
Documents the June 1 release, MSA efficiency, benchmark methodology, real-world agent runs, API modes and access routes.
official · en · verified 2026-06-01
Shows M3 alongside Hailuo 2.3, Speech 2.8 and Music 2.6 on the public homepage.
Last checked: 2026-08-04