Skip to content

Text-to-Speech

Text → audio bytes. Same three providers, three SDK calls, three big voice catalogs. Every voice ID below was verified against the live docs.

Voice IDs (English US, verified)

Provider Voice ID Kind Notes
Azure Speech en-US-Ava:DragonHDLatestNeural HD Neural Newest voice family.
Azure Speech en-US-JennyNeural Neural Long-standing default.
Azure Speech en-US-AndrewMultilingualNeural Multilingual Neural Cross-language reuse.
GCP TTS en-US-Chirp3-HD-Achernar (F), en-US-Chirp3-HD-Fenrir (M) Chirp 3 HD Newest voice tier.
GCP TTS en-US-Neural2-C (F), en-US-Neural2-A (M) Neural2 Premium standard.
AWS Polly Ruth / Amy / Matthew / Danielle Generative Engine="generative".
AWS Polly Joanna / Salli / Kevin / Kimberly Neural Engine="neural".

The full catalogs are large and drift often — hit the docs URLs above before pinning to production.

Audio formats

Provider Default output Programmatic override
Azure Speech 24 kHz 48 kb/s mono MP3 SpeechConfig.set_speech_synthesis_output_format(...)
GCP TTS Chosen per call AudioConfig(audio_encoding=AudioEncoding.MP3) / .LINEAR16 / .OGG_OPUS
AWS Polly Chosen per call OutputFormat="mp3" / "pcm" / "ogg_vorbis" / "json"

The examples default to MP3 so the raw bytes are directly playable.

Native per-cloud calls

import os
import azure.cognitiveservices.speech as speechsdk

speech_config = speechsdk.SpeechConfig(
    subscription=os.environ["SPEECH_KEY"],
    endpoint=os.environ["SPEECH_ENDPOINT"],
)
speech_config.speech_synthesis_voice_name = os.environ.get(
    "AZURE_TTS_VOICE", "en-US-JennyNeural",
)
# Get bytes back instead of playing to the local speaker.
audio_out = speechsdk.audio.PullAudioOutputStream()
audio_config = speechsdk.audio.AudioOutputConfig(stream=audio_out)
synth = speechsdk.SpeechSynthesizer(
    speech_config=speech_config, audio_config=audio_config,
)
result = synth.speak_text_async("Hello from Azure Speech").get()
audio_bytes = result.audio_data
import os
from google.cloud import texttospeech

client = texttospeech.TextToSpeechClient()
resp = client.synthesize_speech(
    input=texttospeech.SynthesisInput(text="Hello from Cloud TTS"),
    voice=texttospeech.VoiceSelectionParams(
        language_code="en-US",
        name=os.environ.get("GCP_TTS_VOICE", "en-US-Chirp3-HD-Achernar"),
    ),
    audio_config=texttospeech.AudioConfig(
        audio_encoding=texttospeech.AudioEncoding.MP3,
    ),
)
audio_bytes = resp.audio_content
import os, boto3

client = boto3.client("polly", region_name=os.environ.get("AWS_REGION", "us-east-1"))
resp = client.synthesize_speech(
    Text="Hello from Polly",
    VoiceId=os.environ.get("AWS_POLLY_VOICE", "Ruth"),
    OutputFormat="mp3",
    Engine=os.environ.get("AWS_POLLY_ENGINE", "generative"),
)
audio_bytes = resp["AudioStream"].read()

A common signature

def synthesize(text: str) -> bytes: ...

Returns MP3 bytes by default; the voice service writes them straight to the HTTP response.