Text-to-Speech¶
Text → audio bytes. Same three providers, three SDK calls, three big voice catalogs. Every voice ID below was verified against the live docs.
Official docs verified 2026-08-08
- Azure Speech TTS: learn.microsoft.com/…/speech-service/get-started-text-to-speech
- GCP TTS SDK: cloud.google.com/text-to-speech/docs/quickstart-client-libraries
- GCP TTS voice list: cloud.google.com/text-to-speech/docs/list-voices-and-types
- AWS Polly SDK: docs.aws.amazon.com/polly/…/synthesize-speech-python
- AWS Polly voice catalog: docs.aws.amazon.com/polly/…/available-voices
Voice IDs (English US, verified)¶
| Provider | Voice ID | Kind | Notes |
|---|---|---|---|
| Azure Speech | en-US-Ava:DragonHDLatestNeural |
HD Neural | Newest voice family. |
| Azure Speech | en-US-JennyNeural |
Neural | Long-standing default. |
| Azure Speech | en-US-AndrewMultilingualNeural |
Multilingual Neural | Cross-language reuse. |
| GCP TTS | en-US-Chirp3-HD-Achernar (F), en-US-Chirp3-HD-Fenrir (M) |
Chirp 3 HD | Newest voice tier. |
| GCP TTS | en-US-Neural2-C (F), en-US-Neural2-A (M) |
Neural2 | Premium standard. |
| AWS Polly | Ruth / Amy / Matthew / Danielle |
Generative | Engine="generative". |
| AWS Polly | Joanna / Salli / Kevin / Kimberly |
Neural | Engine="neural". |
The full catalogs are large and drift often — hit the docs URLs above before pinning to production.
Audio formats¶
| Provider | Default output | Programmatic override |
|---|---|---|
| Azure Speech | 24 kHz 48 kb/s mono MP3 | SpeechConfig.set_speech_synthesis_output_format(...) |
| GCP TTS | Chosen per call | AudioConfig(audio_encoding=AudioEncoding.MP3) / .LINEAR16 / .OGG_OPUS |
| AWS Polly | Chosen per call | OutputFormat="mp3" / "pcm" / "ogg_vorbis" / "json" |
The examples default to MP3 so the raw bytes are directly playable.
Native per-cloud calls¶
import os
import azure.cognitiveservices.speech as speechsdk
speech_config = speechsdk.SpeechConfig(
subscription=os.environ["SPEECH_KEY"],
endpoint=os.environ["SPEECH_ENDPOINT"],
)
speech_config.speech_synthesis_voice_name = os.environ.get(
"AZURE_TTS_VOICE", "en-US-JennyNeural",
)
# Get bytes back instead of playing to the local speaker.
audio_out = speechsdk.audio.PullAudioOutputStream()
audio_config = speechsdk.audio.AudioOutputConfig(stream=audio_out)
synth = speechsdk.SpeechSynthesizer(
speech_config=speech_config, audio_config=audio_config,
)
result = synth.speak_text_async("Hello from Azure Speech").get()
audio_bytes = result.audio_data
import os
from google.cloud import texttospeech
client = texttospeech.TextToSpeechClient()
resp = client.synthesize_speech(
input=texttospeech.SynthesisInput(text="Hello from Cloud TTS"),
voice=texttospeech.VoiceSelectionParams(
language_code="en-US",
name=os.environ.get("GCP_TTS_VOICE", "en-US-Chirp3-HD-Achernar"),
),
audio_config=texttospeech.AudioConfig(
audio_encoding=texttospeech.AudioEncoding.MP3,
),
)
audio_bytes = resp.audio_content
import os, boto3
client = boto3.client("polly", region_name=os.environ.get("AWS_REGION", "us-east-1"))
resp = client.synthesize_speech(
Text="Hello from Polly",
VoiceId=os.environ.get("AWS_POLLY_VOICE", "Ruth"),
OutputFormat="mp3",
Engine=os.environ.get("AWS_POLLY_ENGINE", "generative"),
)
audio_bytes = resp["AudioStream"].read()
A common signature¶
Returns MP3 bytes by default; the voice service writes them straight to the HTTP response.