Moonshot AI
Kimi K2.7 Code is Moonshot AI's coding-focused agentic model released on June 16, 2026. The official resource positions it as open-source, optimized for long-horizon coding and agentic task execution, with about 30% lower thinking-token usage than K2.6. The Hugging Face model card lists a Mixture-of-Experts architecture with 1T total parameters, 32B activated parameters per token, 384 experts, 8 selected experts per token, 256K context, MLA attention, SwiGLU activation and a 400M-parameter MoonViT vision encoder. It supports text, image and video input, requires thinking mode and preserves reasoning content during multi-step tool calls. The hosted HighSpeed variant targets about 180 output tokens per second and up to 260 tokens per second in short contexts, although Moonshot notes that current capacity is limited.
Editorial verdict
Teams evaluating open-weight coding models for repository-scale agents, Kimi Code workflows or OpenAI-compatible API integration.
Avoid it when you need non-thinking mode, a lightweight self-hosted model, or a general-purpose assistant tuned beyond coding.
K2.7 Code has its own open weights, architecture, benchmarks, API pricing and Kimi Code default-model behavior, so it deserves a dedicated model profile separate from Kimi Code and the broader Kimi API.
Open weights on Hugging Face; Kimi API is $0.19/1M input tokens cache hit, $0.95/1M input cache miss and $4.00/1M output tokens
Open weights, Kimi Code plans, Kimi API pay-as-you-go, Enterprise sales
Commercial use should follow the current product, API, model license and billing terms.
Review prompt, file, media upload, retention and training-use terms before sensitive workloads.
Use it for repository refactors, multi-file feature work and debugging sessions that need persistent reasoning.
Kimi says K2.7 Code is now the default model in Kimi Code with thinking enabled by default.
Call it through Kimi API for coding agents, developer tools, ToolCalls and JSON-mode workflows.
Evaluate local or private deployments through supported inference engines such as vLLM, SGLang and KTransformers.
Checked on June 21, 2026. HighSpeed capacity is currently limited and may fluctuate. Kimi Code plan prices, API rates, parameter constraints, license terms, benchmark methodology and supported inference engines may change; recheck official sources before procurement or deployment.
qwen
deepseek-v4-api
zhipu-glm
minimax-m3
official · en · verified 2026-08-04
Confirms release date, positioning, benchmark table, architecture, access paths, Kimi Code plans and API price table.
official · en · verified 2026-06-17
Confirms open weights, Modified MIT license, architecture, evaluations, deployment engines and usage details.
pricing · en · verified 2026-06-17
Documents API model description, HighSpeed version, context caching, ToolCalls, JSON Mode and Partial Mode.
docs · en · verified 2026-06-21
Confirms hosted model IDs, 256K context, default 32K output, HighSpeed throughput, mandatory thinking, fixed sampling values and multimodal tool-result behavior.
Last checked: 2026-08-04