Skip to content

Voice agent

STT → LLM → TTS. The middle step is the E2 call-an-llm service imported verbatim — the voice recipe only adds transcription and speech synthesis on the ends.

The pipeline

     ┌─────────┐    ┌────────────┐    ┌─────────┐
     │  audio  │──► │  STT       │──► │ prompt  │
     │  bytes  │    │ (cloud)    │    │  text   │
     └─────────┘    └────────────┘    └────┬────┘
                                           │
     ┌─────────┐    ┌────────────┐    ┌────▼────────┐
     │  audio  │◄── │  TTS       │◄── │ E2 chat_*   │
     │  bytes  │    │ (cloud)    │    │ (REUSED)    │
     └─────────┘    └────────────┘    └─────────────┘

Per-cloud services

Cloud STT TTS
Azure Azure AI Speech — SpeechRecognizer in azure-cognitiveservices-speech Azure AI Speech — SpeechSynthesizer (same SDK)
GCP Speech-to-Text v2 — google-cloud-speech Text-to-Speech — google-cloud-texttospeech
AWS Amazon Transcribe — boto3 transcribe client Amazon Polly — boto3 polly client

Docs verified 2026-08-08

Every voice-id / model / API call below was verified against these live pages:

Pages

  • STT — audio → text per provider.
  • LLM in the middle — the two-line reuse of E2 (chat_native / chat_claude).
  • TTS — text → audio per provider + a verified voice-id table.
  • Deploy — Terraform composing E1 baseline + E2 call-an-llm; Helm reusing the E2 chart shape.

Validate-only

Same discipline as E2/E3 — no live audio, no live LLM, no live TTS in CI. The "verify" block on each page is the local one-liner to test against your own credentials.