A voice agent that answers before the pause gets awkward.

WebRTC and LiveKit streaming, edge voice-activity detection and instant barge-in, so callers can interrupt the agent the way they would interrupt a person.

Live specimen: caller and agent audio ribbons. The barge-in cuts the agent mid-word.

Cascading bots make people wait, then talk over them.

Speech-to-text, then the language model, then text-to-speech, each over HTTP and each waiting for the last. Design target for the old pattern: 2,000 to 4,000 ms of dead air per turn.

The result is an unnatural pause and an agent that cannot be interrupted. This build replaces the relay race with one continuous stream.

Two timelines for the same turn of conversation.

Design targets and benchmark claims, drawn to one shared scale. Segment splits are illustrative.

Fig. 1 Cascading STT, LLM, TTS over HTTP 2,000 to 4,000 ms
STTLLMTTS

Each stage waits for the previous one, so the pauses add up and nothing can cut in.

Fig. 2 WebRTC streaming path via LiveKit < 500 ms
stream

Audio, detection and response overlap in one full-duplex session.

Five stages, all running at once.

  1. 1Full-duplex stream
  2. 2Edge VAD
  3. 3Barge-in event
  4. 4Async tool calls
  5. 5Jitter and loss

Full-duplex WebRTC streaming

Audio flows both ways at once through the LiveKit media server. Nobody has to finish before the other side can start.

Silero VAD at the edge

Voice activity is monitored on 10 to 30 ms chunks, so speech boundaries are found as they happen.

Interruption in under 50 ms

When the caller speaks, an immediate interrupt event clears the client playback buffer. Design target: under 50 ms.

Function calls mid-sentence

Tools run asynchronously while the agent keeps talking, with conversational fillers covering the wait.

Jitter buffer and loss concealment

A jitter buffer and WebRTC packet-loss concealment keep speech smooth over cellular networks.

Four deliverables, ready to run.

LiveKit agent service

Edge VAD and interruption handling in one service, so barge-in works from the first call.

Terraform SIP trunking

Scripts that route Twilio or Telnyx SIP trunks into LiveKit rooms.

Next.js frontend

Audio visualiser and a WebRTC transcription data channel.

Latency benchmarks

Automated network tests under packet loss.

Sub-500 millisecond conversational turnaround, so automated contact centers can finally support natural human barge-in. Callers speak when they want to, and the agent stops listening to itself.

Let's make your voice agent feel human.

Real-time voice infrastructure, from SIP trunk to interruption handling.

Start a conversation