Skip to content

LLM in the middle

This is the reuse chapter. No new SDKs, no new dispatch table — the voice recipe imports E2 call-an-llm's chat_native.py + chat_claude.py verbatim, exactly like E3 RAG does.

The composition

# examples/voice/service/voice.py
import sys
from pathlib import Path

sys.path.insert(0, str(Path(__file__).resolve().parent.parent.parent / "call-an-llm" / "service"))
from chat_native import azure_chat, gcp_chat, aws_chat            # E2 verbatim
from chat_claude import foundry_claude, vertex_claude, bedrock_claude   # E2 verbatim

Then the voice pipeline is:

from stt import transcribe
from tts import synthesize

_NATIVE = {"azure": azure_chat, "gcp": gcp_chat, "aws": aws_chat}


def voice_turn(audio_in: bytes, provider: str = PROVIDER) -> bytes:
    text_in  = transcribe(audio_in)              # STT (this recipe)
    text_out = _NATIVE[provider](text_in)        # LLM (E2 reused)
    return synthesize(text_out)                   # TTS (this recipe)

That's the whole recipe: one function, three calls, and the middle one is imported.

Claude through-line composes identically

from chat_claude import bedrock_claude, foundry_claude, vertex_claude

_CLAUDE = {"azure": foundry_claude, "gcp": vertex_claude, "aws": bedrock_claude}

Swap _NATIVE for _CLAUDE in voice_turn — nothing else changes.

Streaming — end-to-end latency

For a live voice UX, wire the E2 streaming chat between the transcript arrival and the TTS synthesis. Azure and GCP TTS both accept a streaming input; AWS Polly synthesizes to a byte stream (AudioStream) you can flush as it arrives.

For end-to-end (mic-in / speaker-out) you also want STT streaming — that's a separate API on every provider:

  • Azure — SpeechRecognizer.start_continuous_recognition_async with an event listener on Recognized.
  • GCP — client.streaming_recognize(requests=iter([...])) on SpeechClient.
  • AWS — start_stream_transcription on a WebSocket client (a separate amazon-transcribe package).

The docs above cover the batch shape. Streaming is a small step on the same primitives.

The one thing to NOT do

Don't put chat logic in the voice module. Every future Chiron recipe that needs an LLM (evals, agents, multimodal, ...) also imports E2's dispatch — if you fork it here, you have to keep three near-identical copies in sync forever. The two lines above are the entire contract.