NetEase Youdao
Confucius4-TTS is an LLM-based text-to-speech system from NetEase Youdao for multilingual and cross-lingual speech synthesis. The repository describes a speech encoder plus large language model architecture that preserves speaker identity across languages. It supports Chinese, English, Japanese, Korean, German, French, Spanish, Indonesian, Italian, Thai, Portuguese, Russian, Malay and Vietnamese. The key release claims include unconstrained voice cloning without a reference transcript, cross-lingual voice transfer, zero-shot voice transfer, emotion transfer, online demo access, Hugging Face and ModelScope model downloads, Python inference, and a training path split into Text2Semantic and Semantic2Acoustic modules.
Editorial verdict
Speech researchers and localization teams evaluating open multilingual TTS, cross-lingual dubbing, zero-shot voice cloning and emotion-preserving voice transfer.
Avoid it if you need a managed speech API, turnkey commercial rights, low-resource CPU inference or voice cloning without explicit consent.
Confucius4-TTS belongs in AI Audio because it is an open TTS and voice-cloning engine with model downloads, inference code and benchmark claims.
Apache-2.0 open-source repository with Hugging Face and ModelScope model downloads; local CUDA inference required
GitHub repository, Hugging Face model download, ModelScope model download, Online demo, Local CUDA inference
Commercial use should follow the current product, API, model license and billing terms.
Review prompt, file, media upload, retention and training-use terms before sensitive workloads.
Keep one speaker identity while synthesizing speech across the 14 supported languages.
Use a reference audio prompt to clone a voice without additional model training or reference transcript.
Follow the Text2Semantic and Semantic2Acoustic training paths with TSV data for local experiments.
Model names, quotas, release status, regional access and commercial terms can change quickly; recheck official sources before procurement or production use.
longcat-audiodit
minimax-audio
stepaudio
qwen-audio
mimo-speech
official · en · verified 2026-08-04
Confirms the multilingual zero-shot TTS positioning, 14 languages, online demo, installation, inference, fine-tuning, benchmarks and Apache-2.0 repository license.
docs · zh · verified 2026-06-25
Provides the Chinese documentation mirror for features, setup, inference, training and evaluation.
official · en · verified 2026-06-25
Linked from the official README as the online demo path.
other · en · verified 2026-06-25
Official README links this as the Hugging Face model path.
other · zh · verified 2026-06-25
Official README links this as the ModelScope model path.
Last checked: 2026-08-04