Engineering Boundaries for Outages You Can't Prevent

Abstract

Details coming soon.


Speaker

Em Ruppe

Em Ruppe

Technical Incident Commander @Chime, Previously @SendGrid and @Twilio, and Product and Training @Jeli.io

Em Ruppe is a Technical Incident Commander on the Incident Ops team at Chime, focusing on response and analysis, and a proud member of the Resilience in Software Foundation. She wrote hundreds of status posts, incident timelines and analyses at SendGrid, and was a founding member of the Incident Command team at Twilio. They also worked in product and training at Jeli.io, helping teams level up how they learn from incidents before it was acquired by Pagerduty. Em is passionate about protecting responders from burnout and making incidents easier for everyone involved.

Read more

From the same track

Session

Adapt or Drift: Resilience Engineering When AI Moves the Operating Point

Resilient systems do not stay resilient by standing still. They survive by adapting. But adaptation comes with risk: under sustained pressure, organizations and systems can slowly drift toward failure while still appearing to operate normally.

Speaker image - Andrew Hatch

Andrew Hatch

Engineering Leader and SRE Manager @Cisco ThousandEyes, With 25+ Years Building Software, Operations, SRE, and Platform Teams Across Australia, India, and the United States

Session

Saturation: How Your Software Will Fail at Scale

Even in the ethereal world of software, everything has a limit. And, once that limit is reached, very bad things can happen. In this talk, we will explore the problem of saturation: when a software system runs into one of these limits.

Speaker image - Lorin Hochstein

Lorin Hochstein

Staff Software Engineer @Airbnb, Writes @surfingcomplexity.blog, Previously @Netflix and Member of the Resilience in Software Foundation

Session

The Freeze Paradox: Why Stopping Deployments Doesn't Stop Failures

When systems fail at Netflix and in an environment this complex, they sometimes will, the instinctive response is to freeze deployments.

Speaker image - Prachi Jain

Prachi Jain

Senior Site Reliability Engineer @Netflix, Expert in Building and Managing Scalable, Reliable Services, Previously @Fastly and @Cisco

Speaker image - Sandhya Narayan

Sandhya Narayan

Technical Program Manager @Netflix, Expert in Information Security, Compliance, and Risk Management, Previously @Adobe, SAP, @eBay, and the Stanford Research Institute

Session

Turning Unseen Micro-Failures into Systemic Safeguards

Details coming soon.