Track host
About the track
This track will take you behind the curtain and into the heart of system meltdowns at some of the world's leading software companies in "The stories behind the incidents" track. Learn directly from SREs about real-world, high-impact production failures at scale, including the immediate challenges of triage, diagnosis, and mitigation in complex distributed systems. From these stories, you’ll gain insights into the nature of real incidents and how skilled SREs recover from them.
You’ll learn about the ambiguous, confusing, and uncertain nature of incidents when you’re in the middle of them, and hear the tales of how engineers were able to improvise innovative solutions in order to restore service. You’ll also learn how fundamentally unpredictable incidents are, and, consequently, the importance of preparing to be surprised.
Sessions in this track
Wednesday 19 November. 5 sessions per track, chosen and introduced by the Track Host.
10:35 Seacliff ABC Session Incidents The Human Toll of Incidents & Ways To Mitigate It Kyle Lexmond Production Engineer @Meta, Previously @AWS and @Twitter Have you ever wondered what it's like to respond to a significant incident? Walk through an hour by hour reconstruction of an incident response or two, focusing on what it was like to be "in the room" and the human response to the incidents. 11:45 Seacliff ABC Session Incidents When Incidents Refuse to End Vanessa Huerta Granda Resilience Engineering Manager @Enova, Co-Author of the Howie Guide on Post Incident Analysis, Board Member for the Resilience in Software Foundation As engineers, we’re used to managing failure, but long-running outages hit differently. They stretch teams, systems, and assumptions about how incidents “should” play out. 13:35 Seacliff ABC Session Staff Plus Engineering The Ironies of A^2 I^2 J. Paul Reed Staff Incident Operations Manager @Chime In this talk, we'll explore some of the "ironies" of automation—and now, artificial intelligence—in their interactions with software operators (i.e. you), especially during high consequence, high tempo situations (aka incidents). 14:45 Seacliff ABC Session Incident Response Week-Long Outage: Lifelong Lessons Molly Struve Staff Site Reliability Engineer @Netflix Routine database upgrades should be straightforward, especially with familiar, well-established technology. We were confident heading into our Elasticsearch upgrade, equipped with a solid plan and excited to see performance gains like we had seen from past upgrades. 15:55 Seacliff ABC Session Incident Analysis The Time it Wasn't DNS Sean Klein Principal Technical Program Manager - Modern Incident Analysis @Microsoft Azure In January of 2023, the Microsoft Azure Wide Area Network experienced a global outage. If you were a Microsoft customer at the time, you were impacted by this outage.QCon San Francisco 2025 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.