Skip to main content
Chinese AI Tools
ProductsModelsIntegrationsRankingsLatest changes
TopicsAvailability TrackerUse casesSubmit a toolAccount
ZHSearch

Chinese AI Tools

Independent directory for Chinese AI products. Product availability, pricing and terms can change. Verify before commercial use.

Editorial standardsClaim productUpdate infoGet featuredAdvertise

Moonshot AI

Kimi K2.7 Code

Moonshot AI's open-weight, coding-focused agentic model for long-horizon software engineering, Kimi Code and Kimi API workflows.

Globally availableFull English UIPublic APIFreemium

Quick answer

K2.7 Code has its own open weights, architecture, benchmarks, API pricing and Kimi Code default-model behavior, so it deserves a dedicated model profile separate from Kimi Code and the broader Kimi API.

Official sitePricingDocumentation
Pricing
Open weights on Hugging Face; Kimi API is $0.19/1M input tokens cache hit, $0.95/1M input cache miss and $4.00/1M output tokens
Availability
Globally available
API
Public API
Last checked
2026-08-04
Pricing detailsUse-case fitSource evidence

Kimi K2.7 Code is Moonshot AI's coding-focused agentic model released on June 16, 2026. The official resource positions it as open-source, optimized for long-horizon coding and agentic task execution, with about 30% lower thinking-token usage than K2.6. The Hugging Face model card lists a Mixture-of-Experts architecture with 1T total parameters, 32B activated parameters per token, 384 experts, 8 selected experts per token, 256K context, MLA attention, SwiGLU activation and a 400M-parameter MoonViT vision encoder. It supports text, image and video input, requires thinking mode and preserves reasoning content during multi-step tool calls. The hosted HighSpeed variant targets about 180 output tokens per second and up to 260 tokens per second in short contexts, although Moonshot notes that current capacity is limited.

Topic guide

Kimi models, API and agent workspace

A source-backed map of Kimi K3, K2.7 Code, the Kimi API, Kimi Code, Kimi Claw, WebBridge and productivity creation surfaces.

Editorial verdict

Best for

Teams evaluating open-weight coding models for repository-scale agents, Kimi Code workflows or OpenAI-compatible API integration.

Avoid if

Avoid it when you need non-thinking mode, a lightweight self-hosted model, or a general-purpose assistant tuned beyond coding.

Why it matters

K2.7 Code has its own open weights, architecture, benchmarks, API pricing and Kimi Code default-model behavior, so it deserves a dedicated model profile separate from Kimi Code and the broader Kimi API.

Trust: 4/4 sources verified, recently checkedCoverage: 100/100

Pricing

Open weights on Hugging Face; Kimi API is $0.19/1M input tokens cache hit, $0.95/1M input cache miss and $4.00/1M output tokens

Payment

Open weights, Kimi Code plans, Kimi API pay-as-you-go, Enterprise sales

Commercial use

Commercial use should follow the current product, API, model license and billing terms.

Privacy

Review prompt, file, media upload, retention and training-use terms before sensitive workloads.

Use-case fit

Long-horizon coding agents

Strong

Use it for repository refactors, multi-file feature work and debugging sessions that need persistent reasoning.

Kimi Code default model

Strong

Kimi says K2.7 Code is now the default model in Kimi Code with thinking enabled by default.

OpenAI-compatible coding API

Strong

Call it through Kimi API for coding agents, developer tools, ToolCalls and JSON-mode workflows.

Open-weight evaluation

Medium

Evaluate local or private deployments through supported inference engines such as vLLM, SGLang and KTransformers.

Global user checklist

RegistrationConfirmedKimi Code and Kimi API provide hosted access; Hugging Face hosts the open weights.
English UIConfirmedThe official resource page, API pricing page and Hugging Face model card are English-facing.
API and docsConfirmedDocs cover Kimi API access, OpenAI-compatible deployment examples, tool calls, multimodal input and deployment engines.
International paymentConfirmedThe official page lists Kimi Code annual-billing plan prices and Kimi API per-token prices for K2.7 Code.
Commercial usePartialWeights and code are under Modified MIT; review the license and model-use terms before redistribution or regulated use.
Data and privacy termsReviewHosted use can send code, tool traces, images and video to Kimi API or Kimi Code; self-hosting changes the data boundary but adds infrastructure risk.

Checked on June 21, 2026. HighSpeed capacity is currently limited and may fluctuate. Kimi Code plan prices, API rates, parameter constraints, license terms, benchmark methodology and supported inference engines may change; recheck official sources before procurement or deployment.

Pros

  • - Open model weights and code repository are released under a Modified MIT License
  • - 1T total parameters with 32B activated parameters per token and 256K context
  • - Designed for long-horizon coding with stronger coding and agent benchmark scores than K2.6
  • - Supports text, image and video input with a native multimodal architecture
  • - API docs include automatic context caching, ToolCalls, JSON Mode and Partial Mode
  • - Hosted HighSpeed variant targets about 180 tokens per second and up to 260 tokens per second in short contexts

Cons

  • - Thinking mode is mandatory; Kimi Code requests with thinking disabled fall back to K2.6
  • - Self-hosting a 1T-parameter MoE model requires serious inference infrastructure and engine compatibility checks
  • - Benchmarks include Moonshot in-house suites and should be compared with independent workload tests
  • - Modified MIT and model-use terms should be reviewed before commercial redistribution or sensitive deployments
  • - Hosted API requests must use fixed sampling values, auto or none tool choice, and preserved reasoning_content across tool steps

Decision paths

qwen

deepseek-v4-api

zhipu-glm

minimax-m3

Sources

Kimi K2.7 Code official resource

Official · EN · verified 2026-08-04

Confirms release date, positioning, benchmark table, architecture, access paths, Kimi Code plans and API price table.

Kimi K2.7 Code Hugging Face model card

Official · EN · verified 2026-06-17

Confirms open weights, Modified MIT license, architecture, evaluations, deployment engines and usage details.

Kimi K2.7 Code API pricing

Pricing · EN · verified 2026-06-17

Documents API model description, HighSpeed version, context caching, ToolCalls, JSON Mode and Partial Mode.

Kimi K2.7 Code quickstart

Documentation · EN · verified 2026-06-21

Confirms hosted model IDs, 256K context, default 32K output, HighSpeed throughput, mandatory thinking, fixed sampling values and multimodal tool-result behavior.

Last checked: 2026-08-04

Reviews

No approved reviews yet.

Availability snapshot

Availability
Globally available
English UI
Full English UI
API
Public API
Rating
4.7 (0)

Latest updates

Latest changes
Release · 2026-08-28

Kimi Code 0.39.1 fixes input, permissions and startup reliability

Kimi Code 0.39.1 fixes IME first-character loss, persists permission mode per session, improves attachments, prevents large-workspace startup from remaining on Connecting, keeps Skill instructions out of user-message history and increases the update timeout.

Release · 2026-08-14

Kimi Code 0.36.1 improves long-running work controls

Kimi Code 0.36.1 adds plan and goal indicators, a plan viewer, background Bash task filtering and details, subagent status views, fuzzy slash-command search with pinyin support and automatic session titles. The stable release also fixes self-hosted OpenAI-compatible approval history, Windows file watching, MCP OAuth cancellation, large session export and long-session forking.

Release · 2026-07-17

Moonshot launches Kimi K3

Moonshot AI launched Kimi K3 as a hosted 2.8T-parameter, native-vision MoE model with a 1M-token context window. It became available in Kimi, Kimi Work, Kimi Code and the Kimi API as kimi-k3 at $0.30/M cache-hit input tokens, $3/M cache-miss input tokens and $15/M output tokens. At launch, Moonshot scheduled the full-weight release for July 27; that release is now tracked separately. Published limitations include sensitivity to dropped thinking history and possible over-proactiveness on ambiguous tasks.

Release · 2026-06-30

OpenRouter Owl Alpha added

OpenRouter's Owl Alpha is now tracked as a free text model for agentic coding, tool use, automated workflows and complex instruction execution. The live model API lists a 1,048,756-token context window, up to 262,144 output tokens, native tools and structured outputs, zero input and output token prices, and one Stealth INT8 provider route. OpenRouter says it works with Claude Code and OpenClaw, but also warns that the provider may log prompts and completions and use them for model improvement. The underlying model identity and license remain undisclosed, so the profile treats it as a non-sensitive Alpha evaluation route rather than a production-safe default.

Submit a review