Skip to content
Voice

Voice agent, cascaded pipeline

Telephony → speech-to-text → agent with tools → text-to-speech, tuned for sub-second turns.

Use it when

Phone or in-app voice for bookings, support, collections and outbound reminders, where you need control over the model, the tools and the voice, and clear logs of what was said.

Structure

The parts, top to bottom

hover a part to see its job

across every level

Hover or tap any part to see what it does. The light shows the order a request moves through.

Flow

What happens, in order

  1. 1Call via SIP / WebRTC
  2. 2Voice activity detection
  3. 3Streaming speech-to-text
  4. 4Agent reasons & calls tools
  5. 5Streaming text-to-speech
  6. 6Interruption handling & handoff
  7. 7Transcript, QA & analytics
Tools

What we typically build it with

Twilio / LiveKitDeepgram / AssemblyAIElevenLabs / CartesiaPipecat / Vapi / RetellOpenAI / Gemini

Trade-offs

Maximum control and observability, and the cheapest per minute (streaming STT now costs $0.002–0.008 per minute). The latency budget must be engineered across every stage: loosely stitched multi-vendor stacks commonly land at a second or more per turn while tightly integrated ones get well under that; the commonly recommended target is a median under 800 ms with p95 under 1.5 s. Semantic end-of-turn detection cuts false interruptions by around 30%.

What we solve

Problems this architecture solves

Generic problem statements with the flow and the outcomes the industry has documented.

Voice agent

Calls missed, callers on hold

Clinics, dealerships and service businesses lose bookings after hours and during peaks. IVR menus frustrate callers and staff repeat the same ten conversations all day.

How the system works

  1. Call arrives
  2. Streaming speech-to-text
  3. Agent reasons & checks calendar / CRM
  4. Streaming text-to-speech
  5. Book, confirm, hand off
Twilio / LiveKitDeepgramElevenLabs / CartesiaOpenAI / GeminiPipecat / Vapi

Outcome: Voice agents replacing IVR resolve a majority of qualified calls automatically (66% at one Parloa customer) and cut wait times by a third or more, at turn latencies under a second. Parloa customer results

Architecture
Related

Other voice patterns

Voice

Voice agent, native speech-to-speech

One multimodal model listens and speaks; the lowest latency and the most natural tone.

Use it when: Conversational experiences where naturalness matters more than fine control, and the provider's voices and languages fit your audience.

Flow

  1. 1Audio stream in
  2. 2Realtime model reasons in speech
  3. 3Tool calls for data & actions
  4. 4Audio stream out
  5. 5Guardrails on transcript

Let's build intelligent systems that drive growth

Tachyon is the engineering partner for teams that need AI in production, not in a deck. Start with a free 60-minute discovery call.

Contact

Tell us the problem, we will map it to the architecture

Every great partnership begins with a conversation. Whether you are exploring possibilities or ready to scale, tell us what you are actually trying to build.

Prefer to talk?

Pick a 60-minute slot. No pitch, just an engineer with honest answers.

Book a call

NDA available on request. We reply within one business day.