Moonshot AI
Kimi's international developer API for K3, K2.7 Code, K2.6, K2.5, K2 and Moonshot V1 models with OpenAI-compatible access.
Quick answer
Kimi's developer path is separated through K2.7 Code, K2.6 and the English API platform, so it should be distinct from the consumer assistant and Kimi Code CLI.
The Kimi API now exposes Kimi K3 alongside K2.7 Code, K2.6, K2.5, K2 and Moonshot V1. K3 is Moonshot's 2.8T-parameter, native-vision flagship with a 1M-token context window; the official launch page lists the kimi-k3 model ID and cache-hit, cache-miss and output pricing. K2.7 Code remains the coding-focused agentic model with 256K context, a default 32K maximum output, native text, image and video input, ToolCalls, JSON Mode, Partial Mode, automatic context caching and mandatory thinking mode. The HighSpeed model uses the kimi-k2.7-code-highspeed ID and targets about 180 output tokens per second, reaching up to 260 tokens per second in short-context scenarios. The OpenAI-compatible API supports multimodal tool results, but K2.7 Code fixes its sampling parameters and requires reasoning_content to remain in context across multi-step tool calls. Chat Completions bills both input and output tokens; extracted document text is billed when passed to a model, while file storage and extraction interfaces are temporarily free.
Topic guideA source-backed map of Kimi K3, K2.7 Code, the Kimi API, Kimi Code, Kimi Claw, WebBridge and productivity creation surfaces.
Editorial verdict
Developers evaluating Chinese long-context, coding and agentic models through an international API.
Avoid assuming consumer Kimi behavior or limits match API behavior.
Kimi's developer path is separated through K2.7 Code, K2.6 and the English API platform, so it should be distinct from the consumer assistant and Kimi Code CLI.
Pay-as-you-go API pricing varies by Kimi K2 and Moonshot model
Pay-as-you-go API billing, Platform billing, Enterprise sales
Commercial use should follow the current product, API, model license and billing terms.
Review prompt, file, media upload, retention and training-use terms before sensitive workloads.
Use K2.7 Code for long-horizon coding agents, tool calls and repository-scale workflows.
Use K2.7 Code, K2.6 or Moonshot V1 for 256K long-context prompts, file-like workloads and complex reasoning checks.
Use text, base64 images, uploaded media and image or video tool results in agentic tasks; direct image URLs are not currently supported.
Model names, quotas, release status, regional access and commercial terms can change quickly; recheck official sources before procurement or production use.
qwen
deepseek
zhipu-glm
Official · EN · verified 2026-08-04
Confirms K2.7 Code positioning, benchmarks, architecture, Kimi Code access and API pricing.
Pricing · EN · verified 2026-06-17
Documents K2.7 Code and HighSpeed pricing page, model description, caching, ToolCalls, JSON Mode and Partial Mode.
Documentation · EN · verified 2026-06-21
Documents model IDs, HighSpeed throughput, context and output limits, OpenAI-compatible setup, multimodal tool results, media limits, fixed parameters and tool-call constraints.
Pricing · EN · verified 2026-06-21
Explains token units, input and output billing, document-extraction billing, temporarily free file interfaces and the Token Calculation API.
Pricing · EN · verified 2026-06-21
Confirms campaign dates, qualifying top-up tiers, voucher percentages, one-participation rule, issuance timing, $4,000 cap, 90-day validity and non-refundable top-ups.
Documentation · EN · verified 2026-05-17
Lists K2.6, K2.5, K2 and Moonshot V1 model families.
Last checked: 2026-08-04
No approved reviews yet.