Saturation: How Your Software Will Fail at Scale

Abstract

Even in the ethereal world of software, everything has a limit. And, once that limit is reached, very bad things can happen. In this talk, we will explore the problem of saturation: when a software system runs into one of these limits. We'll see a sampling of the many different kinds of system overload, how our systems are always potentially at risk of saturation no matter how well designed they are, and what we can do to prepare for the day when a system in our charge has been pushed beyond what it can normally handle.


Speaker

Lorin Hochstein

Lorin Hochstein

Staff Software Engineer @Airbnb, Writes @surfingcomplexity.blog, Previously @Netflix and Member of the Resilience in Software Foundation

Lorin Hochstein is Staff Software Engineer, Reliability at Airbnb. He was previously Senior Staff Software Engineer at Coupang, Senior Software Engineer at Netflix, Senior Software Engineer at SendGrid Labs, Lead Architect for Cloud Services at Nimbis Services, Computer Scientist at the University of Southern California's Information Sciences Institute, and Assistant Professor in the Department of Computer Science and Engineering at the University of Nebraska–Lincoln.

Lorin has a B.Eng. in Computer Engineering from McGill University, an M.S. in Electrical Engineering from Boston University, and a PhD in Computer Science from the University of Maryland.

Read more
Find Lorin Hochstein at:

From the same track

Session

Adapt or Drift: Resilience Engineering When AI Moves the Operating Point

Monday Nov 16 / 02:45PM PST

Resilient systems do not stay resilient by standing still. They survive by adapting. But adaptation comes with risk: under sustained pressure, organizations and systems can slowly drift toward failure while still appearing to operate normally.

Speaker image - Andrew Hatch

Andrew Hatch

Engineering Leader and SRE Manager @Cisco ThousandEyes, With 25+ Years Building Software, Operations, SRE, and Platform Teams Across Australia, India, and the United States

Session

The Freeze Paradox: Why Stopping Deployments Doesn't Stop Failures

Monday Nov 16 / 10:35AM PST

When systems fail at Netflix and in an environment this complex, they sometimes will, the instinctive response is to freeze deployments.

Speaker image - Prachi Jain

Prachi Jain

Senior Site Reliability Engineer @Netflix, Expert in Building and Managing Scalable, Reliable Services, Previously @Fastly and @Cisco

Speaker image - Sandhya Narayan

Sandhya Narayan

Technical Program Manager @Netflix, Expert in Information Security, Compliance, and Risk Management, Previously @Adobe, SAP, @eBay, and the Stanford Research Institute

Session

Engineering Boundaries for Outages You Can't Prevent

Monday Nov 16 / 03:55PM PST

Details coming soon.

Speaker image - Em Ruppe

Em Ruppe

Technical Incident Commander @Chime, Previously @SendGrid and @Twilio, and Product and Training @Jeli.io

Session

Turning Unseen Micro-Failures into Systemic Safeguards

Monday Nov 16 / 05:05PM PST

Details coming soon.

Session

Unconference: Resilience Engineering

Monday Nov 16 / 01:35PM PST