NVIDIA releases Magpie TTS — an open-weights multilingual voice model in 12 languages including Mandarin. Low-latency and self-hostable for HK creators.
NVIDIA has released Magpie TTS Multilingual — a 364M-parameter open-weights text-to-speech model supporting 12 languages including Mandarin, Cantonese-adjacent code-switching, and production-ready deployment via NVIDIA NIM. Here's what Hong Kong creators need to know.
One Model, Twelve Languages
Magpie TTS Multilingual is an open-weights model designed to run on your own infrastructure. At 364M parameters, it's small enough for efficient inference but large enough to produce natural-sounding speech across a wide language set.
The model supports English, Spanish, French, German, Italian, Vietnamese, Mandarin, Hindi, Japanese, and now adds Modern Standard Arabic, Korean, and Brazilian Portuguese. Each language comes with male and female speaker voices through a shared multilingual speaker representation — meaning one checkpoint handles all 12 without swapping models.
For Hong Kong creators working in bilingual or trilingual content (Cantonese, Mandarin, English), this is significant. The model's improved code-switching support for Hindi and Japanese — enabled by IPA grapheme-to-phoneme processing and custom pronunciation dictionaries — signals better handling of mixed-language content overall. Names, technical terms, and code-switched phrases should render more accurately than previous generations.
Low-Latency and Self-Hosted
The key differentiator here is control. Magpie isn't a cloud-only API — it ships with open weights and production deployment tooling through NVIDIA NIM. You can run it on your own hardware, which means:
- Data residency: Audio data never leaves your infrastructure - Latency tuning: Optimize inference for your specific workload - Customization: Fine-tune pronunciation, voices, and domain-specific terminology - Scale on your terms: No per-character API pricing, no rate limits
For HK agencies handling sensitive client content or needing predictable costs for high-volume voiceover work, self-hosted TTS removes a major barrier.
What This Means for Hong Kong Creators
Voiceover production is a recurring cost for Hong Kong's advertising and content creation industry. Traditional voice talent for bilingual spots runs HK$3,000-8,000 per session. AI TTS brings that down dramatically, but cloud APIs add per-minute fees and data privacy concerns.
Magpie TTS changes the calculus. With open weights and self-hosting, agencies running Cooly Studio for AI video production can integrate a zero-cost-per-call TTS layer for multilingual voiceovers. The 364M parameter footprint means it runs on a single consumer GPU — no expensive enterprise hardware required.
The timing matters. As Hong Kong brands push into Southeast Asian and Middle Eastern markets, the ability to generate voiceovers in Vietnamese, Arabic, and Korean from the same model is a practical advantage. One model, one deployment, twelve languages.
How to Get Started
Magpie TTS Multilingual is available now on Hugging Face under an open license. Developers can download the weights, deploy via NVIDIA NIM for production workloads, or run inference directly through popular TTS pipelines. If you're already using Cooly Studio for AI video production, the model integrates naturally into existing voiceover workflows — just generate your script, pass it through Magpie, and layer the output onto your video timeline.
Frequently Asked Questions
Q: Is Magpie TTS free to use? A: Yes, the model weights are open and free to download. Production deployment costs are limited to the hardware you run it on — no per-character or per-minute fees.
Q: Does Magpie TTS support Cantonese? A: Mandarin is supported natively. While Cantonese isn't explicitly listed, the IPA grapheme-to-phoneme processing pipeline and code-switching improvements may handle Cantonese-adjacent content. For pure Cantonese voiceover work, you may need to supplement with other tools.
Q: What hardware do I need to run it? A: At 364M parameters, Magpie runs on a single consumer GPU. A mid-range NVIDIA RTX card is sufficient for real-time inference.
Q: How does it compare to cloud TTS APIs like ElevenLabs or Google Cloud TTS? A: Magpie offers lower latency (no network round-trip) and full data privacy, but requires self-hosting. Cloud APIs offer convenience and broader language support. The right choice depends on your volume, privacy requirements, and infrastructure.
Q: Can I use Magpie TTS with Cooly Studio? A: Yes. Generate your voiceover script, run it through Magpie locally, and import the audio into Cooly Studio for AI video production — seamless integration with no API costs.
