From turn-commit to first audio: LLM ~59%, TTS ~28%, my orchestration ~14% — and one first token took 2.9 s. I was tuning the wrong layer.
Provider-agnostic Go engine for low-latency STT→LLM→TTS voice pipelines — streaming engines behind uniform interfaces, swap vendors with a config change. Owns the hot path.