Speaker
Abstract
This talk presents research collected from the VOID—an open database of public incident reports. Containing over 2,000 reports for almost 700 organizations, the database allows for more structured review and research about software-related incident reporting. Key results from our research challenge standard industry practices for incident response and analysis, like tracking Mean Time To Resolve (MMTR) and using Root Cause Analysis (RCA) methodology. In particular, we demonstrate how unreliable MTTR can be, and how RCA can lead to environments where people are less likely to admit mistakes and speak up about things that could lead to future incidents. We propose alternate metrics (SLOs and cost of coordination data), practices (Near Miss analysis), and mindsets (humans are the solution, not the problem) to help organizations better learn from their incidents, and make their systems safer and more resilient.
Topics
QCon San Francisco 2022 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.
Part of the track
Effective SRE Hosted by Casey Rosenthal CEO, Co-Founder @verica_ioFrom the same track
Wednesday 26 October
10:35 Bayview Session SRE The Endgame of SRE Amy Tobey Senior Principal Engineer and SRE practice Leader @Equinix The containers are deployed and the builds are green. Yaml flows through the system, linted, reviewed, tested, and shipped with ease and regularity. Our intrepid SRE finds themself at a crossroads. The infrastructure is great but teams still struggle to maintain error budgets. 11:50 Bayview Session SRE Did the Chaos Test Pass? Christina Yakomin Senior Site Reliability Engineering Specialist @Vanguard_Group People used to ask me all the time how to figure out if their chaos test has “passed,” and I’d always say “well, that’s a loaded question.” To confirm that a chaos test “passed,” we need to do verification of hypotheses - sometimes you’re trying to prove some system behavior occurred in response… 13:40 Seacliff ABC Session [Panel] SRE: Is it Working? Courtney Nash, Amy Tobey, Christina Yakomin, Sasha Rosenbaum How does SRE mature from a craft with a wide range of skills and levels of expertise to a mature discipline? 14:55 Bayview Session SRE Rethinking Reliability: What You Can (and Can't) Learn From Incidents Courtney Nash Co-founder @The VOID, Previously @Verica, @Holloway, @Fastly, @O’Reilly Media, @Microsoft, & @Amazon This talk presents research collected from the VOID—an open database of public incident reports. Containing over 2,000 reports for almost 700 organizations, the database allows for more structured review and research about software-related incident reporting. 16:10 Bayview Session SRE The Eternal Sunshine of the Toil-Less Prod Sasha Rosenbaum Director of the Cloud Services Black Belt Team @RedHat One of the most important decisions in building an SRE practice is what kind of work should be assigned to the SRE team, and in what percentages.