Speech-to-Text¶
Take audio bytes, get a text transcript. Every provider has one SDK + one language code + one method call — the shape differs enough that the app dispatches on CHIRON_PROVIDER.
Official docs verified 2026-08-08
- Azure Speech Python SDK: learn.microsoft.com/…/speech-service/get-started-speech-to-text
- GCP Speech-to-Text v2: cloud.google.com/speech-to-text/v2/docs/quickstart-client-libraries
- AWS Transcribe boto3: docs.aws.amazon.com/transcribe/…/getting-started-python-sdk
Model + language IDs¶
| Provider | Package | Recognition model / language | Notes |
|---|---|---|---|
| Azure Speech | azure-cognitiveservices-speech |
Speech service picks the model automatically for the given language locale (e.g. en-US) |
Set locale via speech_recognition_language. |
| GCP Speech-to-Text v2 | google-cloud-speech |
model="long" (default); Chirp 2/3 via chirp_2 |
Uses the _ (auto) recognizer with language_codes=["en-US"]. |
| AWS Transcribe | boto3 transcribe |
Async job; LanguageCode="en-US" |
Job-based — input goes through S3; result is a downloadable JSON. |
Native per-cloud calls¶
Local WAV file → text. In production the audio config accepts a stream instead of a file path.
import os
import azure.cognitiveservices.speech as speechsdk
speech_config = speechsdk.SpeechConfig(
subscription=os.environ["SPEECH_KEY"],
endpoint=os.environ["SPEECH_ENDPOINT"],
)
speech_config.speech_recognition_language = "en-US"
audio_config = speechsdk.audio.AudioConfig(filename="input.wav")
recognizer = speechsdk.SpeechRecognizer(
speech_config=speech_config, audio_config=audio_config,
)
result = recognizer.recognize_once_async().get()
if result.reason == speechsdk.ResultReason.RecognizedSpeech:
print(result.text)
import os
from google.cloud.speech_v2 import SpeechClient
from google.cloud.speech_v2.types import cloud_speech
client = SpeechClient()
config = cloud_speech.RecognitionConfig(
auto_decoding_config=cloud_speech.AutoDetectDecodingConfig(),
language_codes=["en-US"],
model="long",
)
request = cloud_speech.RecognizeRequest(
recognizer=f"projects/{os.environ['GOOGLE_CLOUD_PROJECT']}/locations/global/recognizers/_",
config=config,
content=open("input.wav", "rb").read(),
)
resp = client.recognize(request=request)
text = " ".join(alt.transcript for r in resp.results for alt in r.alternatives[:1])
print(text)
Transcribe is an asynchronous batch service — the audio file must live in S3, and you poll for completion. Real-time streaming is a separate API (start_stream_transcription) documented at docs.aws.amazon.com/transcribe/latest/dg/streaming.html.
import json, time, os, urllib.request, uuid, boto3
client = boto3.client("transcribe", region_name=os.environ.get("AWS_REGION", "us-east-1"))
job = f"chiron-voice-{uuid.uuid4().hex[:8]}"
client.start_transcription_job(
TranscriptionJobName=job,
LanguageCode="en-US",
MediaFormat="wav",
Media={"MediaFileUri": f"s3://{os.environ['CHIRON_AUDIO_BUCKET']}/input.wav"},
)
while True:
j = client.get_transcription_job(TranscriptionJobName=job)["TranscriptionJob"]
if j["TranscriptionJobStatus"] in ("COMPLETED", "FAILED"):
break
time.sleep(2)
uri = j["Transcript"]["TranscriptFileUri"]
with urllib.request.urlopen(uri) as resp:
data = json.loads(resp.read())
text = data["results"]["transcripts"][0]["transcript"]
print(text)
A common signature¶
The voice service exposes:
CHIRON_PROVIDER picks the backend. Downstream (LLM, TTS) is provider-agnostic — the transcribed text hands off cleanly.