How to Build a Real-Time Voice Agent

QCon San Francisco 2026

Session AI/ML

How to Build a Real-Time Voice Agent

Monday Nov 16 / 01:35PM PST, Ballroom BC at Hyatt Regency, San Francisco

Register

$2,955, Conference (3 days)
Current pricing ends October 13th

Abstract

A voice agent looks like a chatbot with a microphone. Architecturally, it's a real-time distributed system with extremely unforgiving constraints: less than 500ms of end-to-end response latency across a multi-model pipeline, intelligent and natural-sounding conversation, and reliability when scaling to thousands of concurrent calls.

This talk walks through the architecture of a production voice agent end to end: the streaming pipeline that chains speech-to-text, turn detection, an LLM, and text-to-speech into a single system, and the design decisions that make it work under real traffic. We'll dig into where the latency budget actually goes, why colocating and serving your own models can beat just calling inference APIs, and why autoscaling stateful, long-lived audio connections is harder than it looks. 

Topics

AI/ML Voice AI Architecture
Register

$2,955, Conference (3 days). Current pricing ends October 13th. All pass options.

76% senior dev or higher
1:11 speaker ratio
60+ practitioners

QCon San Francisco 2026 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.

Share

From the same track

Monday 16 November

10:35 Ballroom BC Session Distributed Systems Live Resharding Without Regret: Lessons from Building Valkey's Atomic Slot Migration Jacob Murphy Open Source Maintainer @Valkey & Software Engineer @Google Cloud's Memorystore Team 11:45 Ballroom BC Session Architecting Nubank's Global Financial Infrastructure 13:35 Ballroom BC Session AI/ML How to Build a Real-Time Voice Agent Rishabh Bhargava Director of ML @Together AI 14:45 Ballroom BC Session Inside the Disaggregated Architecture Serving LLM Inference Nancy Zheng VP of Engineering of AI Inference @Cerebras 15:55 Ballroom BC Session Architecting Real-Time Timeline Engines for 100 Million Active Users 17:05 Seacliff D Unconference Unconference: Architectures You've Always Wondered About

Current pricing ends October 13th
$2,955, Conference (3 days)

Register