Chinese AI Tools
ProductsModelsIntegrationsRankingsLatest changes
Availability TrackerUse casesSubmit a toolAccount
ZHSearch

Chinese AI Tools

Independent directory for Chinese AI products. Product availability, pricing and terms can change. Verify before commercial use.

Editorial standardsClaim productUpdate infoGet featuredAdvertise

MiniMax

MiniMax Audio / Speech

MiniMax Audio covers Speech 2.8, Speech 2.6 and Speech-02 models, 40-language speech synthesis, synchronous HTTP/WebSocket TTS, async long-form TTS, voice cloning and official voice-management APIs. Speech 2.8 adds native sound tags for breaths, pauses, hesitations, chuckles and other expressive actions, higher-fidelity cloning from a 10-second sample, and cleaner output with reduced background noise and digital artifacts. MiniMax also documents improved cross-lingual pronunciation and reduced accent bleed for the Mandarin-Japanese pair.

Globally availableFull English UIPublic APIFreemium

Editorial verdict

Best for

Teams evaluating Chinese speech synthesis, voice cloning and multilingual audio generation APIs.

Avoid if

Avoid using cloned voices in production without clear consent, data and commercial-use review.

Why it matters

MiniMax Audio deserves a separate profile because the official API docs cover a mature speech product line beyond general model chat.

Trust: 6/6 sources verified, recently checkedCoverage: 100/100

Pricing

Speech 2.8 Turbo $60/M characters; HD $100/M characters; rapid cloning $1.50/voice; voice design $3/voice

Payment

Audio Subscription, Token Plan, Credits, Pay-as-you-go API billing

Commercial use

Voice cloning, synthetic voice and generated-audio use should be reviewed against current consent and product terms.

Privacy

Review uploaded voice samples, cloned voice retention and generated audio storage before using real voices.

Use-case fit

Multilingual text-to-speech

Strong

Use Speech 2.8 or 2.6 for multilingual TTS, voice chat and online social interaction scenarios.

Long-form audio generation

Strong

Async TTS supports long-form audio tasks such as books or long documents.

Voice cloning and custom voices

Medium

Use voice cloning and voice design only after legal and consent checks.

Conversational character speech

Strong

Use native sound tags to add breaths, hesitations, chuckles and pacing cues to assistants, characters and narration.

Cross-lingual cloned voices

Medium

Evaluate Mandarin-Japanese voice transfer where reducing accent bleed and pronunciation shifts matters.

Global user checklist

RegistrationConfirmedAvailable through the MiniMax international API platform.
English UIConfirmedSpeech docs and product pages are English-facing.
API and docsConfirmedOfficial docs cover TTS, async TTS, voice cloning, voice design and voice management; the Speech 2.8 release documents native sound tags and 10-second cloning.
Commercial useReviewVoice rights, consent and generated-audio usage need explicit review.

Checked June 21, 2026. Recheck model IDs, sound-tag syntax, supported languages, voice-clone requirements, consent controls and subscription quotas before production.

Pros

  • - Speech 2.8 and 2.6 are current documented models
  • - Supports HTTP and WebSocket TTS plus async long-text generation
  • - Voice cloning and voice design APIs are documented
  • - Native sound tags add breaths, pauses, hesitations and expressive vocal actions
  • - Speech 2.8 can clone a voice from a 10-second reference sample
  • - The release targets cleaner studio-style output with less noise and digital distortion

Cons

  • - Voice rights and consent requirements need explicit review
  • - Voice similarity, studio-grade clarity and human-authenticity claims are vendor-reported
  • - The announcement specifically demonstrates Mandarin-Japanese cross-lingual improvement; broader language gains remain to be verified

Decision paths

Use the API Platform for full multimodal access

The API Platform profile covers billing, keys and cross-modal integration.

Compare with SparkDesk for China speech workflows

SparkDesk is a Chinese speech and vertical-scenario reference.

Sources

MiniMax homepage

official · en · verified 2026-08-04

Surfaces Speech 2.8 together with M3, Hailuo 2.3 and Music 2.6.

MiniMax Speech 2.8 news

official · en · verified 2026-06-21

Confirms native sound tags, 10-second high-fidelity cloning, noise and artifact reduction, Mandarin-Japanese cross-lingual improvements and availability through MiniMax Audio and Open Platform.

MiniMax API pricing

pricing · en · verified 2026-06-21

Lists Speech 2.8 Turbo and HD character pricing plus rapid voice-cloning and voice-design prices.

MiniMax models overview

docs · en · verified 2026-06-01

Lists Speech 2.8, Speech 2.6 and Speech-02 model families.

MiniMax speech guide

docs · en · verified 2026-06-01

Documents synchronous TTS and streaming usage.

MiniMax voice clone guide

docs · en · verified 2026-06-01

Documents voice cloning capabilities.

Last checked: 2026-08-04

Reviews

Availability snapshot

Availability
available
English UI
full
API
available
Rating
4.3 (0)

Latest updates

Latest changes
Pricing · 2026-08-04

MiniMax multimodal API pricing refreshed

MiniMax's live API pricing lists permanently discounted M3 rates for context up to 512K at $0.30/M input, $1.20/M output and $0.06/M cache-read tokens; 512K-1M costs $0.60, $2.40 and $0.12, with access above 512K currently limited. M2.7 costs $0.30/M multimodal input and $1.20/M output, while Highspeed costs $0.60 and $2.40. Hailuo 2.3 starts at $0.28 per 768P six-second video, Speech 2.8 Turbo at $60/M characters, rapid voice cloning at $1.50 per voice, voice design at $3 per voice, and Music 2.6 at $0.15 per song with a free-trial label.

Release · 2026-06-21

MiniMax Speech 2.8 profile refreshed

MiniMax Speech 2.8 adds native sound tags for breaths, pauses, hesitations and expressive actions, plus higher-fidelity voice cloning from a 10-second sample and cleaner output with reduced background noise and digital artifacts. MiniMax also highlights improved cross-lingual pronunciation and reduced accent bleed for the Mandarin-Japanese pair, with more languages planned. The model is live through MiniMax Audio and the Open Platform API; cloning quality and human-indistinguishability claims remain vendor-reported and require consent-aware testing.

Open source · 2026-06-17

Confucius4-TTS multilingual zero-shot TTS profile added

NetEase Youdao's Confucius4-TTS is now tracked in AI Audio as an Apache-2.0 open-source multilingual and cross-lingual zero-shot TTS engine. The profile records its speech-encoder-plus-LLM architecture, 14 supported languages, unconstrained voice cloning without reference transcripts, cross-lingual voice transfer, emotion transfer, online demo, Hugging Face and ModelScope model paths, Python inference, fine-tuning workflow and benchmark sources.

Release · 2026-06-01

MiniMax M3 and latest model wave added

MiniMax is now tracked with a dedicated M3 model profile, while the homepage currently highlights M3, Hailuo 2.3, Speech 2.8 and Music 2.6 alongside the API Platform, Token Plan and Agent surfaces.

Submit a review