Full-duplex WebRTC streaming
Audio flows both ways at once through the LiveKit media server. Nobody has to finish before the other side can start.
WebRTC and LiveKit streaming, edge voice-activity detection and instant barge-in, so callers can interrupt the agent the way they would interrupt a person.
Speech-to-text, then the language model, then text-to-speech, each over HTTP and each waiting for the last. Design target for the old pattern: 2,000 to 4,000 ms of dead air per turn.
The result is an unnatural pause and an agent that cannot be interrupted. This build replaces the relay race with one continuous stream.
Design targets and benchmark claims, drawn to one shared scale. Segment splits are illustrative.
Each stage waits for the previous one, so the pauses add up and nothing can cut in.
Audio, detection and response overlap in one full-duplex session.
Audio flows both ways at once through the LiveKit media server. Nobody has to finish before the other side can start.
Voice activity is monitored on 10 to 30 ms chunks, so speech boundaries are found as they happen.
When the caller speaks, an immediate interrupt event clears the client playback buffer. Design target: under 50 ms.
Tools run asynchronously while the agent keeps talking, with conversational fillers covering the wait.
A jitter buffer and WebRTC packet-loss concealment keep speech smooth over cellular networks.
Edge VAD and interruption handling in one service, so barge-in works from the first call.
Scripts that route Twilio or Telnyx SIP trunks into LiveKit rooms.
Audio visualiser and a WebRTC transcription data channel.
Automated network tests under packet loss.
Sub-500 millisecond conversational turnaround, so automated contact centers can finally support natural human barge-in. Callers speak when they want to, and the agent stops listening to itself.
Real-time voice infrastructure, from SIP trunk to interruption handling.
Start a conversation