Track host
About the track
Bugs get triggered. Hardware fails. Networks are unreliable. Accidents and malicious attacks happen. Any of these events can cause a system to stop functioning, or put it into a state so degraded that its performance is no longer acceptable.
Critical systems must continue to function and meet SLAs in the presence of internal or external faults. In this track, we will delve into industry best practices as well as innovative approaches to designing resilient systems.
These approaches can be technical, such as new ways to route around degraded network links. Alternatively, they can be sociotechnical: the more a system depends on its operator, the more important it is that the operator has a clear, unambiguous understanding of the system’s state and how to intervene. Most systems require a mix of both.
In this track, we will delve into each of these areas to provide attendees with the tools they need to build resilient systems and empower operators.
The day in the host's words
Sessions in this track
Tuesday 3 October. 6 sessions per track, chosen and introduced by the Track Host.
10:35 Ballroom BC Session Architecture Disaster Recovery Across a Million Pieces Michelle Brush Engineering Director, SRE @Google, Previously Director of HealtheIntent Architecture @Cerner Corporation & Lead Engineer @Garmin, Author of "2 out of the 97 Things Every SRE Should Know" Data recovery is more than just backing up and restoring a data store. The goal of any disaster recovery effort is getting the system back to working as expected across all of its parts. 11:45 Ballroom BC Session Reliability Designing Fault-Tolerant Software with Control System Transparency Jon Moore Staff Software Engineer @Stripe with over 35 years of software engineering experience across both academia and industry Teams at NASA and JPL that create mission-critical software for spacecraft take a principled approach to fault tolerance. Let's see how those same principles, centered around a concept of transparency, can help us achieve reliability in pragmatic, modern software delivery settings. 13:35 Ballroom BC Session Resiliency How Do We Talk to Each Other? How Surfacing Communication Patterns in Organizations Can Help You Understand and Improve Your Resilience Nora Jones Founder and CEO @jeli_io, Founder of Learning From Incidents (LFI) Online Community and Conference As a system increases in inevitable complexity, it becomes impossible for a single operator to have a clear, unambiguous understanding of what's happening in the system. Understanding the system requires a joint effort between teammates and technology. 14:45 Ballroom BC Session Database How Netflix Ensures Highly-Reliable Online Stateful Systems Joseph Lynch Principal Software Engineer @Netflix Building Highly-Reliable and High-Leverage Infrastructure Across Stateless and Stateful Services Under most stateless services are stateful databases, caches, and systems which form the bedrock applications are built on. 15:55 Ballroom BC Session Resiliency Orchestrating Resilience: Building Modern Asynchronous Systems Sai Pragna Etikyala Technical Lead @Twilio Building asynchronous, event-driven systems can be daunting. Managing states, ensuring resilience, maintaining traceability, and handling a myriad of other challenges often require more effort than building the functionality itself. 17:05 Seacliff D Unconference Unconference: Designing for Resilience What is an unconference? An unconference is a participant-driven meeting. Attendees come together, bringing their challenges and relying on the experience and know-how of their peers for solutions.QCon San Francisco 2023 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.