LLM in the middle¶
This is the reuse chapter. No new SDKs, no new dispatch table — the voice recipe imports E2 call-an-llm's chat_native.py + chat_claude.py verbatim, exactly like E3 RAG does.
The composition¶
# examples/voice/service/voice.py
import sys
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent.parent / "call-an-llm" / "service"))
from chat_native import azure_chat, gcp_chat, aws_chat # E2 verbatim
from chat_claude import foundry_claude, vertex_claude, bedrock_claude # E2 verbatim
Then the voice pipeline is:
from stt import transcribe
from tts import synthesize
_NATIVE = {"azure": azure_chat, "gcp": gcp_chat, "aws": aws_chat}
def voice_turn(audio_in: bytes, provider: str = PROVIDER) -> bytes:
text_in = transcribe(audio_in) # STT (this recipe)
text_out = _NATIVE[provider](text_in) # LLM (E2 reused)
return synthesize(text_out) # TTS (this recipe)
That's the whole recipe: one function, three calls, and the middle one is imported.
Claude through-line composes identically¶
from chat_claude import bedrock_claude, foundry_claude, vertex_claude
_CLAUDE = {"azure": foundry_claude, "gcp": vertex_claude, "aws": bedrock_claude}
Swap _NATIVE for _CLAUDE in voice_turn — nothing else changes.
Streaming — end-to-end latency¶
For a live voice UX, wire the E2 streaming chat between the transcript arrival and the TTS synthesis. Azure and GCP TTS both accept a streaming input; AWS Polly synthesizes to a byte stream (AudioStream) you can flush as it arrives.
For end-to-end (mic-in / speaker-out) you also want STT streaming — that's a separate API on every provider:
- Azure —
SpeechRecognizer.start_continuous_recognition_asyncwith an event listener onRecognized. - GCP —
client.streaming_recognize(requests=iter([...]))onSpeechClient. - AWS —
start_stream_transcriptionon a WebSocket client (a separateamazon-transcribepackage).
The docs above cover the batch shape. Streaming is a small step on the same primitives.
The one thing to NOT do¶
Don't put chat logic in the voice module. Every future Chiron recipe that needs an LLM (evals, agents, multimodal, ...) also imports E2's dispatch — if you fork it here, you have to keep three near-identical copies in sync forever. The two lines above are the entire contract.