Chinese TTS API Voices

Generate lifelike Chinese speech across a wide range of voices and every major accent, over carrier-grade infrastructure built for voice agents and IVR.

Pay as you go from ~$3 per 1M characters, no commitment

Built on the same infrastructure thousands of teams ship voice on

14,000+
companies build on Telnyx
100+
languages & dialects
1,300+
voices, one API
<500ms
end-to-end latency
VOICES & ACCENTS

Chinese voices & accents

Chinese voices across every major regional accent. Hear them below, or browse the full catalog by country.

DEVELOPERS

Call the Chinese TTS API

Generate Chinese audio in one request. OpenAI-SDK compatible, with streaming for real-time apps.

cURL
curl -X POST https://api.telnyx.com/v2/text-to-speech/speech \
  -H "Authorization: Bearer $TELNYX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "voice": "Telnyx.NaturalHD.zh-CN-1",
    "text": "您的处方药已准备好,请到药房取药。"
  }' --output sample.mp3
Node
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.TELNYX_API_KEY,
  baseURL: "https://api.telnyx.com/v2",
});

const audio = await client.audio.speech.create({
  model: "Telnyx.NaturalHD",
  voice: "zh-CN-1",
  input: "您的处方药已准备好,请到药房取药。",
});
Python
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["TELNYX_API_KEY"],
    base_url="https://api.telnyx.com/v2",
)

audio = client.audio.speech.create(
    model="Telnyx.NaturalHD",
    voice="zh-CN-1",
    input="您的处方药已准备好,请到药房取药。",
)

Voice IDs follow the Telnyx.<Tier>.<Voice> convention, so you can swap the voice without touching the rest of your code.

Read the TTS docs
USE CASES

What teams build with Chinese TTS

Voice agents & IVR

Natural Chinese phone agents that handle calls end-to-end, carrier-grade.

Customer support automation

Bilingual support flows that resolve common requests without a queue.

E-learning

Narrate Chinese courses and lessons in any regional accent.

Audiobooks

Long-form Chinese narration with consistent, lifelike delivery.

Video voiceover & dubbing

Localize content into Chinese at scale with named native voices.

Accessibility

Chinese screen-reading and read-aloud for inclusive products.

ALL VOICES

Browse all Chinese voices by country

Female Chinese TTS Voices

127
telnyx⚡ Hosted

Jing - Clear Coordinator

Ultra Chinese female: Clear Coordinator. Articulate, low-latency ready.

Telnyx.Ultra.6eb8965c-e295-47bd-a9e4-3eeebb3abcff
minimax

News Anchor

MiniMax 02t female, News Anchor. Articulate for voice assistants.

Minimax.speech-02-turbo.Chinese (Mandarin)_News_Anchor
inworld

Jing

An energetic, fast-paced young Chinese female [8bbf].

Inworld.Mini.Jing
azure

Xiaoxiao Dragon HD Flash Latest

Azure DragonHD Mandarin Chinese female, zh-CN-Xiaoxiao:DragonHDFlashLatestNeural. Contr...

Azure.zh-CN-Xiaoxiao:DragonHDFlashLatestNeural
telnyx⚡ Hosted

Hua - Sunny Support

Ultra Chinese female: Sunny Support. Clear, call-optimized.

Telnyx.Ultra.7a5d4663-88ae-47b7-808e-8f9b9ee4127b
minimax

News Anchor

MiniMax 2.6t female, News Anchor. Refined for IVR systems.

Minimax.speech-2.6-turbo.Chinese (Mandarin)_News_Anchor
inworld

Jing

An energetic, fast-paced young Chinese female.

Inworld.Max.Jing
azure

Xiaoxiao2 Dragon HD Flash Latest

Azure DragonHD Mandarin Chinese female, zh-CN-Xiaoxiao2:DragonHDFlashLatestNeural. Cont...

Azure.zh-CN-Xiaoxiao2:DragonHDFlashLatestNeural

Male Chinese TTS Voices

91
telnyx⚡ Hosted

Hao - Friendly Guy

Ultra Chinese male: Friendly Guy. Robust, production-grade.

Telnyx.Ultra.16212f18-4955-4be9-a6cd-2196ce2c11d1
minimax

Reliable Executive

MiniMax 02t male, Reliable Executive. Direct for inbound calls.

Minimax.speech-02-turbo.Chinese (Mandarin)_Reliable_Executive
inworld

Ming

A young adult male with a smooth, clear voice speaking slowly and deliberately i... [2622]

Inworld.Mini.Ming
azure

Yunxiao Dragon HD Flash Latest

Azure DragonHD Mandarin Chinese male, zh-CN-Yunxiao:DragonHDFlashLatestNeural. Lucid, l...

Azure.zh-CN-Yunxiao:DragonHDFlashLatestNeural
telnyx⚡ Hosted

Liu - Plain Talker

Ultra Chinese male: Plain Talker. Fluid, stack-native.

Telnyx.Ultra.653b9445-ae0c-4312-a3ce-375504cff31e
minimax

Reliable Executive

MiniMax 2.6t male, Reliable Executive. Rich for real-time calls.

Minimax.speech-2.6-turbo.Chinese (Mandarin)_Reliable_Executive
inworld

Ming

A young adult male with a smooth, clear voice speaking slowly and deliberately in a qui...

Inworld.Max.Ming
azure

Yunyi Dragon HD Flash Latest

Azure DragonHD Mandarin Chinese male, zh-CN-Yunyi:DragonHDFlashLatestNeural. Natural, c...

Azure.zh-CN-Yunyi:DragonHDFlashLatestNeural

Mainland China Chinese TTS Voices

185
minimax

Reliable Executive

MiniMax 02t male, Reliable Executive. Direct for inbound calls.

Minimax.speech-02-turbo.Chinese (Mandarin)_Reliable_Executive
inworld

Jing

An energetic, fast-paced young Chinese female [8bbf].

Inworld.Mini.Jing
azure

Xiaoxiao Dragon HD Flash Latest

Azure DragonHD Mandarin Chinese female, zh-CN-Xiaoxiao:DragonHDFlashLatestNeural. Contr...

Azure.zh-CN-Xiaoxiao:DragonHDFlashLatestNeural
minimax

Reliable Executive

MiniMax 2.6t male, Reliable Executive. Rich for real-time calls.

Minimax.speech-2.6-turbo.Chinese (Mandarin)_Reliable_Executive
inworld

Jing

An energetic, fast-paced young Chinese female.

Inworld.Max.Jing
azure

Xiaoxiao2 Dragon HD Flash Latest

Azure DragonHD Mandarin Chinese female, zh-CN-Xiaoxiao2:DragonHDFlashLatestNeural. Cont...

Azure.zh-CN-Xiaoxiao2:DragonHDFlashLatestNeural
minimax

Reliable Executive

MiniMax 2.8t male, Reliable Executive. Resonant for inbound calls.

Minimax.speech-2.8-turbo.Chinese (Mandarin)_Reliable_Executive
inworld

Jing

An energetic, fast-paced young Chinese female [b6d0].

Inworld.TTS2.Jing

Hong Kong Chinese TTS Voices

21
minimax

Professional Female Host

MiniMax 02t female, ProfessionalHost(F). Expressive for voice assistants.

Minimax.speech-02-turbo.Cantonese_ProfessionalHost(F)
azure

HiuMaan

Azure Neural Cantonese female, zh-HK-HiuMaanNeural. Polished, carrier-optimized.

Azure.zh-HK-HiuMaanNeural
minimax

Professional Female Host

MiniMax 2.6t female, ProfessionalHost(F). Bright for customer service.

Minimax.speech-2.6-turbo.Cantonese_ProfessionalHost(F)
azure

WanLung

Azure Neural Cantonese male, zh-HK-WanLungNeural. Lucid, stack-native.

Azure.zh-HK-WanLungNeural
minimax

Professional Female Host

MiniMax 2.8t female, ProfessionalHost(F). Grounded for voice interfaces.

Minimax.speech-2.8-turbo.Cantonese_ProfessionalHost(F)
azure

HiuGaai

Azure Neural Cantonese female, zh-HK-HiuGaaiNeural. Warm, built for scale.

Azure.zh-HK-HiuGaaiNeural
minimax

Gentle Lady

MiniMax 02t female, GentleLady. Polished for AI pipelines.

Minimax.speech-02-turbo.Cantonese_GentleLady
minimax

Gentle Lady

MiniMax 2.6t female, GentleLady. Lucid for customer service.

Minimax.speech-2.6-turbo.Cantonese_GentleLady

Taiwan Chinese TTS Voices

3
azure

HsiaoChen

Azure Neural Taiwanese Mandarin female, zh-TW-HsiaoChenNeural. Crisp, infrastructure-na...

Azure.zh-TW-HsiaoChenNeural
azure

YunJhe

Azure Neural Taiwanese Mandarin male, zh-TW-YunJheNeural. Direct, real-time optimized.

Azure.zh-TW-YunJheNeural
azure

HsiaoYu

Azure Neural Taiwanese Mandarin female, zh-TW-HsiaoYuNeural. Clean, production-grade.

Azure.zh-TW-HsiaoYuNeural
POWERED BY TELNYX

Infrastructure for TTS Library, courtesy of Telnyx

The delivery path, model breadth, low-latency pipeline, and economics behind every voice above.

RECOMMENDED STACK

The stack quality Chinese TTS needs

01

Delivery path

24 kHz audio has to survive the phone network. Without a carrier-grade delivery path it gets crushed to 8 kHz, throwing away the quality the model produced. Own the delivery, or the voice degrades before anyone hears it.

02

Model layer

No single engine wins every language, accent, and budget. Front many models (Telnyx, AWS Polly, Rime, Inworld, ElevenLabs, and more) behind one API so you can swap by config instead of re-integrating each time.

03

Real-time pipeline

For voice agents and IVR, TTS alone isn't enough: TTS, STT, LLM, SIP and numbers belong in one path under ~500ms end-to-end, or the back-and-forth feels laggy and robotic.

04

Economics

Predictable per-character, pay-as-you-go pricing, no seats, minimums, or surprise overage tiers, is what keeps quality voice affordable once you scale past a demo.

TRUST & COMPLIANCE

Enterprise-grade trust

SOC 2 Type IIHIPAAPCI DSSGDPRISO 27001 / 27701STIR/SHAKEN A-level

Owned infrastructure in 20+ countries, PSTN reach in 100+ countries, used by 14,000+ companies.

PRICING

Chinese TTS pricing

From ~$3 per 1M charactersPay as you go. No commitment, no seats.
ModelPer characterPer 1M characters
Telnyx (standard)$0.000003 / char~$3 / 1M chars
Telnyx HD$0.000048 / char~$48 / 1M chars
Bring-your-own (ElevenLabs, Azure)Your provider's rateBilled through Telnyx

Qualifying startups can apply to the Telnyx startups program for up to $20K in credits. Volume discounts available on the Growth Plan.

Start building

Start building with Chinese TTS

Chinese voices, carrier-grade delivery, from ~$3 per 1M characters.

Pay as you go. No commitment. Business email required.

FAQ

Chinese text to speech FAQ

Chinese text to speech (TTS) converts written Chinese into natural-sounding spoken audio using AI voices. Telnyx exposes Chinese voices through one API for apps, IVR, and voice agents.

PHONOLOGY & PROSODY

Chinese phonology and prosody

Pitch that carries the dictionary

Mandarin is a tone language[1]: nearly every syllable carries one of four fixed pitch contours: high, rising, dipping, or falling: and changing the contour changes the word. The syllable "ma" means mother (mā), hemp (má), horse (mǎ), or scold (mà). English uses pitch to signal stress and focus[2]; Mandarin pitch is built into the word itself[3]. A TTS system that misshapes a single tone doesn't sound unnatural: it says the wrong word. Producing accurate Mandarin requires inference that resolves tone at the syllable level, with no inter-provider routing degrading the pitch signal.

Evenly chopped, not bouncy

English is stress-timed[1]: stressed syllables land at regular intervals while unstressed syllables compress between them, creating a strong-weak-weak bounce. Mandarin is syllable-timed[2]: syllables stay closer to equal in duration with far less reduction, producing what sounds like a row of similarly sized beats[3] carrying different pitch shapes. A voice engine trained on English stress-timing will squeeze and stretch syllables that should stay even. Getting this right requires models built for Mandarin rhythm, running co-located with the audio pipeline.

Sentence melody on a tonal tightrope

English intonation is relatively free: pitch accents move around a sentence to mark focus or signal questions[1] without changing word identity. Mandarin intonation must ride on top of lexical tones[2], using post-focus pitch compression to convey emphasis while keeping each syllable's tone intact. The same discourse function: focus, question, statement: is realized through different prosodic strategies[3] than English uses. Imposing English-style rising question contours onto Mandarin warps lexical tones into wrong words. This two-layer pitch system demands synthesis infrastructure where tone and intonation are resolved together, not split across providers.