StepFun
StepAudio covers StepFun's audio models. The English docs list StepAudio 2.5 TTS as a contextual text-to-speech model with natural-language control, emotional arcs and zero-shot voice cloning from about 3 seconds of reference audio. They also list step-tts-2, step-tts-mini, stepaudio-2.5-asr for streaming / near-realtime transcription, and stepaudio-2-asr-pro as a 32B ASR Pro model.
Editorial verdict
Teams evaluating Chinese speech APIs for expressive TTS, voice cloning, dubbing, customer service, NPC dialogue and transcription.
Avoid it when you need a fully validated international audio workflow without testing signup, consent handling and rate limits.
StepAudio is a distinct capability line and should be visible in the AI Audio category, not hidden under the generic StepFun profile.
stepaudio-2.5-tts $0.85 / 10,000 characters; step-tts-2 $0.40 / 10,000 characters; ASR $0.022 / hour; voice cloning $1.50 / voice
Open Platform balance, Step Plan quota for supported audio models
Commercial use should follow StepFun's audio API terms and any voice cloning consent requirements.
Voice cloning and transcription can involve biometric or sensitive audio; consent, retention and data-processing terms need review.
Use contextual TTS for audiobooks, short drama dubbing, ad narration and emotional storytelling.
Use stepaudio-2.5-asr for captions, voice input, meeting transcription and backend batch processing.
Audio docs explicitly mention game NPCs as an audio-driven experience use case.
minimax-audio
qwen-audio
zhipu-glm-audio
seeduplex-audio
docs · en · verified 2026-08-04
Documents StepAudio 2.5 TTS, step-tts models and ASR models.
Last checked: 2026-08-04