DeepSeek
DeepSeek's English API platform for V4-Pro and V4-Flash, with OpenAI-compatible and Anthropic-compatible access.
Quick answer
The English docs now make DeepSeek's current API shape explicit: V4-Pro/V4-Flash, dual OpenAI/Anthropic compatibility, thinking mode, tool calls, caching and agent-tool integrations.
DeepSeek is tracked from the official English API docs. The public IDs deepseek-v4-flash and deepseek-v4-pro currently route to V4-Flash-0731 and V4-Pro-0813. Both support 1M context, 384K maximum output, thinking/non-thinking modes, JSON, tools, caching and OpenAI/Anthropic-compatible endpoints. Peak hours are 01:00-04:00 and 06:00-10:00 UTC. V4-Flash costs $0.007/$0.22/$0.66 per million cache-hit input, cache-miss input and output tokens off-peak, or $0.014/$0.44/$1.32 at peak; V4-Pro costs $0.022/$0.66/$1.98 off-peak or $0.044/$1.32/$3.96 at peak.
Topic guideA source-backed map of DeepSeek V4 models, API pricing, native Harness, third-party agent runtimes and practical integration guides.
Editorial verdict
Developers evaluating Chinese model APIs for coding agents, reasoning, tool use and OpenAI/Anthropic-compatible migration.
Avoid new integrations that still hard-code deepseek-chat or deepseek-reasoner without a migration plan to V4 model names.
The English docs now make DeepSeek's current API shape explicit: V4-Pro/V4-Flash, dual OpenAI/Anthropic compatibility, thinking mode, tool calls, caching and agent-tool integrations.
V4-Flash off-peak: $0.007/M cache-hit input, $0.22/M cache-miss input and $0.66/M output; peak rates are double
Topped-up balance, Granted balance, Platform billing
API commercial use should follow the active platform terms.
Treat prompts and logs as vendor-processed data unless enterprise terms say otherwise.
Use deepseek-v4-pro or deepseek-v4-flash directly rather than relying on retired compatibility aliases.
Official docs list integrations for Claude Code, GitHub Copilot, OpenCode, Kilo Code, OpenClaw and other agent tools.
V4 models document 1M context, 384K max output and context caching with cache-hit usage fields.
Pricing and routed snapshots were rechecked on September 7, 2026. The official table still lists deepseek-v4-flash and deepseek-v4-pro as the stable public IDs, routing to V4-Flash-0731 and V4-Pro-0813 with active peak/off-peak billing.
Qwen is one option when deployment and cloud billing matter.
Kimi API is a cataloged reference for English documentation, token billing and agent-oriented models.
Documentation · EN · verified 2026-05-17
Confirms English docs, OpenAI/Anthropic compatibility, V4 model names and legacy alias deprecation date.
Pricing · EN · verified 2026-09-07
Lists V4-Pro-0813 and V4-Flash-0731 with 1M context, 384K maximum output, features, 2,500/500 concurrency limits and active peak/off-peak prices.
Documentation · EN · verified 2026-05-17
Confirms Claude Code and agent-tool integration paths through the Anthropic-compatible endpoint.
Documentation · EN · verified 2026-05-17
Confirms the 2026-04-24 DeepSeek-V4 update and alias deprecation schedule.
Last checked: 2026-09-07
No approved reviews yet.