Reliability

Session Resiliency

The Freeze Paradox: Why Stopping Deployments Doesn't Stop Failures

Monday Nov 16 / 10:35AM PST

When systems fail at Netflix and in an environment this complex, they sometimes will, the instinctive response is to freeze deployments.

Speaker image - Prachi Jain

Prachi Jain

Senior Site Reliability Engineer @Netflix, Expert in Building and Managing Scalable, Reliable Services, Previously @Fastly and @Cisco

Speaker image - Sandhya Narayan

Sandhya Narayan

Technical Program Manager @Netflix, Expert in Information Security, Compliance, and Risk Management, Previously @Adobe, SAP, @eBay, and the Stanford Research Institute

Session Distributed Systems

Live Resharding Without Regret: Lessons from Building Valkey's Atomic Slot Migration

Monday Nov 16 / 10:35AM PST

Sharding is easy. Resharding under heavy load is notoriously difficult. How do you move gigabytes of state across live database nodes without dropping keys, blocking the main event loop, or breaking client abstractions?

Speaker image - Jacob Murphy

Jacob Murphy

Open Source Maintainer @Valkey & Software Engineer @Google Cloud's Memorystore Team

Session Reliability

Saturation: How Your Software Will Fail at Scale

Monday Nov 16 / 02:45PM PST

Even in the ethereal world of software, everything has a limit. And, once that limit is reached, very bad things can happen. In this talk, we will explore the problem of saturation: when a software system runs into one of these limits.

Speaker image - Lorin Hochstein

Lorin Hochstein

Staff Software Engineer @Airbnb, Writes @surfingcomplexity.blog, Previously @Netflix and Member of the Resilience in Software Foundation