Speaker
Abstract
Everyone talks about model quality. Fewer people talk about what happens when that model has to serve real traffic without falling over, or why the system ended up shaped the way it did in the first place.
At Netflix, our ML serving platform runs in two modes. Synchronous serving answers requests in real time, for use cases such as recommendations and search. Asynchronous serving processes predictions in large batches, often on a schedule, for workloads such as Automatic Speech Recognition and Machine Translation. Both modes serve similar models, but they fail in different ways, and some of that traces back to Conway's Law: how ML practitioners and the teams consuming their models are organized and shaped the systems we built.
We'll share some of the incidents we've run into across both modes, what we learned, and what we changed as a result. We'll also talk about our experience rebuilding this platform from the ground up and turning it into a shared offering that other teams at Netflix now build on. Along the way, we'll cover topics such as admission control, concurrency limits, and what's worth measuring in production. This talk is meant to be a practical, honest look at building ML infrastructure, including people and org structure behind it.
Main Takeaways:
- Lessons on architecting for both sync and async inference, and where the two diverge.
- How to build concurrency controls that hold up in production.
- How to measure and observe inference systems correctly.
- How to debug production inference failures at scale.
- How org structure shapes serving architecture, for better or worse.
$2,955, Conference (3 days). Current pricing ends October 13th. All pass options.
QCon San Francisco 2026 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.
Part of the track
Engineering AI Systems Hosted by Melanie Zhao Engineering Lead @BlackRock, Pioneering AI Adoption in Asset ManagementFrom the same track
Tuesday 17 November
10:35 Ballroom A Session AI/ML Skills, Memory, or Fine-Tuning? The Engineering Loop Behind Self-Improving Agents Abhinav Sinha CEO @Lucidic AI, Previously @Stanford AI Lab, @Citadel and Susquehanna International Group, and @Apple As agents become mainstream, everyone wants to improve theirs either by making fewer mistakes on existing tasks or by taking on harder ones. This usually happens once an agent is already deployed in production. 11:45 Ballroom A Session Inside Asimov: BlackRock’s Platform Architecture for Applied AI at Scale Dan Wolf, Sri Annapurna Bhupatiraju BlackRock’s Portfolio Management Group is transforming how investment teams conduct research by combining deep investment expertise with shared AI capabilities. 13:35 Ballroom A Session AI The Revenge of the Data Scientist: Why Reliable AI Needs Evals, Traces, and Metrics Hamel Husain Machine Learning Engineer, 20+ Years in Applied AI, Machine Learning, and Data Science Most teams can now ship an AI prototype by calling a foundation-model API. The hard part is knowing whether that system works when real users, messy data, and business consequences arrive. 14:45 Ballroom A Session AI/ML Lessons from Building a $100M Product in Six Weeks at OpenAI Brian Yang Member of Technical Staff @OpenAI Engineering is shifting from synchronous, line-by-line implementation to an asynchronous model in which agents execute, verify, and retry while humans own intent, architecture, constraints, and release decisions. 15:55 Ballroom A Session Bridging Retrieval and Ranking: From LLM Embeddings and Graph Features to a GPU-Served DLRM Sudarshan Srinivasa Ramanujam Lead Engineer @LinkedIn - Working on Feed Ranking & Retrieval Modeling The LinkedIn Feed retrieval layer selects, within 120 ms, the handful of posts each member sees from hundreds of millions of candidates. This split-second decision shapes what over a billion professionals read and engage with every day. 17:05 Ballroom A Session Failing in Layers: Lessons from Rebuilding Netflix's ML Serving Platform Rajat Shah, Wenbing Bai Everyone talks about model quality. Fewer people talk about what happens when that model has to serve real traffic without falling over, or why the system ended up shaped the way it did in the first place.