MiniMax
MiniMax Audio covers Speech 2.8, Speech 2.6 and Speech-02 models, 40-language speech synthesis, synchronous HTTP/WebSocket TTS, async long-form TTS, voice cloning and official voice-management APIs. Speech 2.8 adds native sound tags for breaths, pauses, hesitations, chuckles and other expressive actions, higher-fidelity cloning from a 10-second sample, and cleaner output with reduced background noise and digital artifacts. MiniMax also documents improved cross-lingual pronunciation and reduced accent bleed for the Mandarin-Japanese pair.
Editorial verdict
Teams evaluating Chinese speech synthesis, voice cloning and multilingual audio generation APIs.
Avoid using cloned voices in production without clear consent, data and commercial-use review.
MiniMax Audio deserves a separate profile because the official API docs cover a mature speech product line beyond general model chat.
Speech 2.8 Turbo $60/M characters; HD $100/M characters; rapid cloning $1.50/voice; voice design $3/voice
Audio Subscription, Token Plan, Credits, Pay-as-you-go API billing
Voice cloning, synthetic voice and generated-audio use should be reviewed against current consent and product terms.
Review uploaded voice samples, cloned voice retention and generated audio storage before using real voices.
Use Speech 2.8 or 2.6 for multilingual TTS, voice chat and online social interaction scenarios.
Async TTS supports long-form audio tasks such as books or long documents.
Use voice cloning and voice design only after legal and consent checks.
Use native sound tags to add breaths, hesitations, chuckles and pacing cues to assistants, characters and narration.
Evaluate Mandarin-Japanese voice transfer where reducing accent bleed and pronunciation shifts matters.
Checked June 21, 2026. Recheck model IDs, sound-tag syntax, supported languages, voice-clone requirements, consent controls and subscription quotas before production.
The API Platform profile covers billing, keys and cross-modal integration.
SparkDesk is a Chinese speech and vertical-scenario reference.
official · en · verified 2026-08-04
Surfaces Speech 2.8 together with M3, Hailuo 2.3 and Music 2.6.
official · en · verified 2026-06-21
Confirms native sound tags, 10-second high-fidelity cloning, noise and artifact reduction, Mandarin-Japanese cross-lingual improvements and availability through MiniMax Audio and Open Platform.
pricing · en · verified 2026-06-21
Lists Speech 2.8 Turbo and HD character pricing plus rapid voice-cloning and voice-design prices.
docs · en · verified 2026-06-01
Lists Speech 2.8, Speech 2.6 and Speech-02 model families.
Last checked: 2026-08-04