Unconference: Resilience Engineering
From the same track
Adapt or Drift: Resilience Engineering When AI Moves the Operating Point
Monday Nov 16 / 02:45PM PST
Resilient systems do not stay resilient by standing still. They survive by adapting. But adaptation comes with risk: under sustained pressure, organizations and systems can slowly drift toward failure while still appearing to operate normally.
Andrew Hatch
Engineering Leader and SRE Manager @Cisco ThousandEyes, With 25+ Years Building Software, Operations, SRE, and Platform Teams Across Australia, India, and the United States
Saturation: How Your Software Will Fail at Scale
Monday Nov 16 / 11:45AM PST
Even in the ethereal world of software, everything has a limit. And, once that limit is reached, very bad things can happen. In this talk, we will explore the problem of saturation: when a software system runs into one of these limits.
Lorin Hochstein
Staff Software Engineer @Airbnb, Writes @surfingcomplexity.blog, Previously @Netflix and Member of the Resilience in Software Foundation
The Freeze Paradox: Why Stopping Deployments Doesn't Stop Failures
Monday Nov 16 / 10:35AM PST
When systems fail at Netflix and in an environment this complex, they sometimes will, the instinctive response is to freeze deployments.
Prachi Jain
Senior Site Reliability Engineer @Netflix, Expert in Building and Managing Scalable, Reliable Services, Previously @Fastly and @Cisco
Sandhya Narayan
Technical Program Manager @Netflix, Expert in Information Security, Compliance, and Risk Management, Previously @Adobe, SAP, @eBay, and the Stanford Research Institute
Engineering Boundaries for Outages You Can't Prevent
Monday Nov 16 / 03:55PM PST
Details coming soon.
Em Ruppe
Technical Incident Commander @Chime, Previously @SendGrid and @Twilio, and Product and Training @Jeli.io
Turning Unseen Micro-Failures into Systemic Safeguards
Monday Nov 16 / 05:05PM PST
Details coming soon.