Zhipu AI
BigModel's model overview lists GLM-TTS, GLM-TTS-Clone, GLM-ASR, GLM-Realtime and GLM-4-Voice under audio/video models, covering speech synthesis, voice cloning, speech recognition and realtime audio-video interaction.
Editorial verdict
Developers evaluating Chinese speech, voice clone, ASR and realtime multimodal APIs.
Avoid production voice cloning without consent and data-retention review.
Audio is now a documented GLM capability family and should be visible in the audio category.
Usage-based audio API pricing varies by model
Free model where available, Platform billing
Commercial use should follow the current product, API, model license and billing terms.
Review prompt, file, media upload, retention and training-use terms before sensitive workloads.
Use it to test TTS, voice cloning, ASR and realtime audio-video calls.
Model names, quotas, release status, regional access and commercial terms can change quickly; recheck official sources before procurement or production use.
minimax-audio
sparkdesk
zhipu-glm
docs · zh · verified 2026-08-04
Lists GLM-TTS, GLM-TTS-Clone, GLM-ASR, GLM-Realtime and GLM-4-Voice.
docs · zh · verified 2026-05-17
Lists speech, voice clone, ASR and realtime API documentation entries.
Last checked: 2026-08-04