Skip to content

Speech-to-Text

Take audio bytes, get a text transcript. Every provider has one SDK + one language code + one method call — the shape differs enough that the app dispatches on CHIRON_PROVIDER.

Model + language IDs

Provider Package Recognition model / language Notes
Azure Speech azure-cognitiveservices-speech Speech service picks the model automatically for the given language locale (e.g. en-US) Set locale via speech_recognition_language.
GCP Speech-to-Text v2 google-cloud-speech model="long" (default); Chirp 2/3 via chirp_2 Uses the _ (auto) recognizer with language_codes=["en-US"].
AWS Transcribe boto3 transcribe Async job; LanguageCode="en-US" Job-based — input goes through S3; result is a downloadable JSON.

Native per-cloud calls

Local WAV file → text. In production the audio config accepts a stream instead of a file path.

import os
import azure.cognitiveservices.speech as speechsdk

speech_config = speechsdk.SpeechConfig(
    subscription=os.environ["SPEECH_KEY"],
    endpoint=os.environ["SPEECH_ENDPOINT"],
)
speech_config.speech_recognition_language = "en-US"

audio_config = speechsdk.audio.AudioConfig(filename="input.wav")
recognizer = speechsdk.SpeechRecognizer(
    speech_config=speech_config, audio_config=audio_config,
)
result = recognizer.recognize_once_async().get()

if result.reason == speechsdk.ResultReason.RecognizedSpeech:
    print(result.text)
import os
from google.cloud.speech_v2 import SpeechClient
from google.cloud.speech_v2.types import cloud_speech

client = SpeechClient()

config = cloud_speech.RecognitionConfig(
    auto_decoding_config=cloud_speech.AutoDetectDecodingConfig(),
    language_codes=["en-US"],
    model="long",
)
request = cloud_speech.RecognizeRequest(
    recognizer=f"projects/{os.environ['GOOGLE_CLOUD_PROJECT']}/locations/global/recognizers/_",
    config=config,
    content=open("input.wav", "rb").read(),
)

resp = client.recognize(request=request)
text = " ".join(alt.transcript for r in resp.results for alt in r.alternatives[:1])
print(text)

Transcribe is an asynchronous batch service — the audio file must live in S3, and you poll for completion. Real-time streaming is a separate API (start_stream_transcription) documented at docs.aws.amazon.com/transcribe/latest/dg/streaming.html.

import json, time, os, urllib.request, uuid, boto3

client = boto3.client("transcribe", region_name=os.environ.get("AWS_REGION", "us-east-1"))

job = f"chiron-voice-{uuid.uuid4().hex[:8]}"
client.start_transcription_job(
    TranscriptionJobName=job,
    LanguageCode="en-US",
    MediaFormat="wav",
    Media={"MediaFileUri": f"s3://{os.environ['CHIRON_AUDIO_BUCKET']}/input.wav"},
)

while True:
    j = client.get_transcription_job(TranscriptionJobName=job)["TranscriptionJob"]
    if j["TranscriptionJobStatus"] in ("COMPLETED", "FAILED"):
        break
    time.sleep(2)

uri = j["Transcript"]["TranscriptFileUri"]
with urllib.request.urlopen(uri) as resp:
    data = json.loads(resp.read())
text = data["results"]["transcripts"][0]["transcript"]
print(text)

A common signature

The voice service exposes:

def transcribe(audio_bytes: bytes, media_format: str = "wav") -> str: ...

CHIRON_PROVIDER picks the backend. Downstream (LLM, TTS) is provider-agnostic — the transcribed text hands off cleanly.