Guide
Choose between hosted and open Qwen models, coding-agent access, subscriptions and multimodal APIs without mixing licenses or billing paths.
Published · Updated
Use Qwen Cloud for broad managed multimodal coverage, Qwen Code for terminal-first agent work, Token Plan for predictable supported-tool access, and open Qwen weights only when the license and serving budget are acceptable.
The guide separates model capability, product access, subscription billing and open-weight licensing using official Qwen and Qwen Cloud sources checked August 25, 2026.
Use Qwen3.8-Max for the hosted flagship. Compare Qwen3.8-2.4T-A95B and Qwen3.8-27B as distinct open deployments: the first is Max-class text with a custom license and high infrastructure needs; the second is a dense Apache-2.0 multimodal model.
Qwen3.8-Max through managed access.
Review each weight release's license, modalities and hardware independently.
Use pay-as-you-go Qwen Cloud APIs for production integration and broad model coverage. Use Token Plan only when its supported coding tools, quotas and renewal terms match the workflow. A Token Plan subscription is not automatically interchangeable with API balance.
Managed text, vision, image, video, audio, embedding and rerank APIs.
Subscription route for documented compatible coding agents.
Qwen Code provides interactive terminal, headless automation, IDE integration and an SDK. QwenPaw and other agents are separate runtimes that can use Qwen. Audit permissions, model routing and provider data terms at the runtime boundary.
First-party open-source terminal agent optimized for Qwen models.
Treat Qwen support as an integration, not product ownership.
Image, Wan/HappyHorse video, audio/CosyVoice and embedding/rerank APIs have different units, asynchronous workflows, safety rules and commercial terms. Benchmark and budget each modality independently.
Separate image, video and audio quality, latency and rights checks.
Select embedding and rerank models independently from chat generation.