Progressive Failure Modes of Modern AI Serving Systems

QCon San Francisco 2026

Session

Progressive Failure Modes of Modern AI Serving Systems

Tuesday Nov 17 / 10:35AM PST, Ballroom A at Hyatt Regency, San Francisco

Register

$2,835, Conference (3 days)
Current pricing ends September 8th

Abstract

Inference platforms fail in layers. Most organizations focus on model quality while underestimating the systems engineering required to operate production AI workloads safely and reliably at scale.

Before GPU saturation even becomes a problem, teams often expose models directly to ungoverned traffic, lack concurrency controls, fail to measure system behavior, overload memory bandwidth, and eventually destroy latency guarantees and operational stability.

This talk walks through the progressive failure modes of modern AI serving systems and how to architect scalable inference infrastructure that remains observable, resilient, and performant under real production workloads. In this talk, I will walk the attendees through real code paths, production failure scenarios, debugging strategies, and architectural tradeoffs, showing both how these systems fail and how to systematically fix them.

Register

$2,835, Conference (3 days). Current pricing ends September 8th. All pass options.

76% senior dev or higher
1:11 speaker ratio
60+ practitioners

QCon San Francisco 2026 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.

Share

From the same track

Tuesday 17 November

10:35 Ballroom A Session Progressive Failure Modes of Modern AI Serving Systems Abi Aryan AI Infrastructure Engineer and Educator 11:45 Ballroom A Session The Revenge of the Data Scientist: Why Reliable AI Needs Evals, Traces, and Metrics Hamel Husain Machine Learning Engineer, 20+ Years in Applied AI, Machine Learning, and Data Science 13:35 Ballroom A Session Skills, Memory, or Fine-Tuning? The Engineering Loop Behind Self-Improving Agents Abhinav Sinha CEO @Lucidic AI, Previously @Stanford AI Lab, @Citadel and Susquehanna International Group, and @Apple 14:45 Ballroom A Session Lessons from Building a $100M Product in Six Weeks at OpenAI Brian Yang Member of Technical Staff @OpenAI 15:55 Ballroom A Session Performance Engineering in the Age of AI 17:05 Seacliff D Unconference Unconference: Engineering AI Systems

Current pricing ends September 8th
$2,835, Conference (3 days)

Register