Chinese AI Tools
ProductsModelsIntegrationsRankingsLatest changes
Availability TrackerUse casesSubmit a toolAccount
ZHSearch

Chinese AI Tools

Independent directory for Chinese AI products. Product availability, pricing and terms can change. Verify before commercial use.

Editorial standardsClaim productUpdate infoGet featuredAdvertise

Alibaba Cloud

Qwen-AgentWorld

Qwen-AgentWorld is positioned by Qwen as a native language world model for general agents. The public release includes Qwen-AgentWorld-35B-A3B open weights and AgentWorldBench, an evaluation benchmark spanning seven agent interaction domains: MCP, Search, Terminal, SWE, Android, Web and OS. The README describes the 35B-A3B model as a MoE language world model with 35B total parameters, 3B active parameters and a 256K context, trained from more than 10M real-world interaction trajectories through CPT, SFT and RL stages. It can be served through SGLang or vLLM with an OpenAI-compatible endpoint, used with Transformers, and evaluated with the included AgentWorldBench scripts.

Globally availableFull English UIPublic APIFree

Editorial verdict

Best for

Researchers and agent builders who need a simulator or benchmark for tool, terminal, SWE, Android, web and OS agent environments.

Avoid if

Avoid treating it as a general chat model or end-user IDE assistant; it is optimized for environment simulation and evaluation workflows.

Why it matters

Qwen-AgentWorld is tracked separately from Qwen Code because its public surface is a language world model plus AgentWorldBench, not a terminal coding product.

Trust: 3/3 sources verified, recently checkedCoverage: 100/100

Pricing

Apache-2.0 open weights and benchmark; self-hosted inference infrastructure required

Payment

Hugging Face download, ModelScope download, Self-hosted inference, Configured judge-model API billing

Commercial use

Commercial use should follow the current product, API, model license and billing terms.

Privacy

Review prompt, file, media upload, retention and training-use terms before sensitive workloads.

Use-case fit

Agent environment simulation

Strong

Use the model to predict environment observations for MCP, terminal, SWE, Android, web, OS and search-style agent trajectories.

World-model benchmark evaluation

Strong

Run AgentWorldBench to score predicted observations across the repository's five evaluation dimensions.

Synthetic training and RL research

Medium

Evaluate simulated RL, controllable perturbations and fictional-world construction before applying those ideas to real agent training.

Global user checklist

RegistrationConfirmedThe GitHub repository, Hugging Face model, ModelScope mirror, blog and arXiv report are public.
English UIConfirmedThe README, quickstart and evaluation instructions are English-facing.
API and docsConfirmedQuickstart examples cover SGLang, vLLM, Transformers and OpenAI-compatible local serving.
International paymentConfirmedThe public release is open-source, but users pay for their own inference hardware and any configured judge-model API.
Commercial usePartialThe README states Apache-2.0 for open weights and AgentWorldBench; verify the model-card license files before production use.
Data and privacy termsPartialSelf-hosting can control inference data, but benchmark data, generated trajectories and external judge calls need internal policy review.

Model names, quotas, release status, regional access and commercial terms can change quickly; recheck official sources before procurement or production use.

Pros

  • - Open weights for Qwen-AgentWorld-35B-A3B and an Apache-2.0 benchmark release
  • - Covers seven agent-environment domains in one model family
  • - SGLang and vLLM examples expose an OpenAI-compatible local API
  • - AgentWorldBench evaluates predicted observations across format, factuality, consistency, realism and quality

Cons

  • - It is a world-model and benchmark release, not a turnkey coding agent or hosted assistant
  • - Running the 35B MoE model with 256K context requires substantial local or cloud inference resources
  • - The README reports 397B research scores, but the listed open-source weights are 35B-A3B

Decision paths

qwen

qwen-code

openclaw

qwenpaw

deer-flow

Sources

Qwen-AgentWorld GitHub repository

official · en · verified 2026-08-04

Confirms the 2026-06-24 release, 35B-A3B open weights, AgentWorldBench, seven domains, SGLang/vLLM deployment and Apache-2.0 license statement.

Qwen-AgentWorld technical report

benchmark · en · verified 2026-06-25

Provides the research report for Qwen-AgentWorld language world models for general agents.

Qwen-AgentWorld Hugging Face collection

other · en · verified 2026-06-25

Hosts the official Qwen-AgentWorld model and AgentWorldBench dataset collection.

Last checked: 2026-08-04

Reviews

Availability snapshot

Availability
available
English UI
full
API
available
Rating
4.1 (0)

Latest updates

Latest changes
Open source · 2026-06-26

DeepReinforce Ornith 1.0 open coding models added

DeepReinforce's Hugging Face organization now exposes the Ornith 1.0 family of MIT-licensed agentic-coding models. The collection includes public 9B, 35B and 397B model cards plus GGUF and FP8 variants. The cards position Ornith as a self-improving model family post-trained from Gemma 4 and Qwen 3.5, using reinforcement learning to optimize both solution rollouts and their scaffolds, with reported results on Terminal-Bench 2.1, SWE-Bench, NL2Repo, SWE Atlas and ClawEval. The profile records vLLM, SGLang, Transformers and OpenAI-compatible local serving paths.

Open source · 2026-06-24

Qwen-AgentWorld model and AgentWorldBench released

Qwen has released Qwen-AgentWorld-35B-A3B open weights and the AgentWorldBench evaluation benchmark. The project positions Qwen-AgentWorld as a native language world model for simulating agentic environments across MCP, Search, Terminal, SWE, Android, Web and OS domains. The README says the 35B MoE model has 3B active parameters and a 256K context, can be served through SGLang or vLLM with an OpenAI-compatible API, and that AgentWorldBench scores predicted environment observations on format, factuality, consistency, realism and quality. The repository also reports 397B research results, but only the 35B-A3B weights are listed as the open-source release.

Submit a review