Enhancing Reliability Using Service-Level Prioritized Load Shedding at Netflix

QCon San Francisco 2025

Session Resilience

Enhancing Reliability Using Service-Level Prioritized Load Shedding at Netflix

Monday Nov 17 / 05:05PM PST, Ballroom BC at Hyatt Regency, San Francisco

Abstract

How does Netflix maintain a seamless viewing experience for millions of users, especially during traffic spikes or when backend datastores are overloaded? Autoscaling can help during traffic spikes, but it costs money, takes a few minutes to kick in, and capacity may not always be available. Furthermore, if a downstream service or datastore is overloaded, autoscaling may exacerbate the problem. 

In this talk, we will cover service-level prioritized load shedding – our solution for prioritizing requests within a single application instance. This innovative solution ensures requests that are critical to user experience maintain high availability, and allows dynamically re-purposing non-critical capacity to serve critical traffic during times of duress. We will also discuss how we automated per-cluster tuning and validation of load shedding, enabling us to quickly deploy unique configurations to hundreds of clusters. 

Key Takeaways

  1. Understand the evolution of load shedding techniques at Netflix and how service-level prioritization enhances user experience and reliability.
  2. Gain insights from real-world applications and testing scenarios that demonstrate the effectiveness of prioritized load shedding.
  3. Learn how platform engineering and service owners at Netflix effectively collaborated to build a generic prioritized load shedding library, efficiently roll it out, and automate tuning of load shedding thresholds.
  4. Learn how we are continuing to improve the experience for service owners through our continued investments in making the load shedding capability more flexible and easier to operate.
     

Topics

Resilience Platform Engineering Developer Experience
76% senior dev or higher
1:11 speaker ratio
60+ practitioners

QCon San Francisco 2025 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.

Share

From the same track

Monday 17 November

10:35 Seacliff ABC Session Platform Engineering Continuous Delivery for Foundational Platforms Ian Nowland CEO @Junction Labs, Author of O'Reilly's Platform Engineering, Previously SVP Core Engineering at Datadog and Leader of AWS Nitro 11:45 Seacliff ABC Session Beyond Line Charts: Why Some Diversity in Telemetry Visualization Is Long Overdue Yao Yue Founder & Chief Executive Officer @IOP Systems, Platform Engineer, Distributed System Aficionado, Cache Expert 13:35 Ballroom BC Session Microservices Platforms: When Team Topologies Meets Microservices Patterns Chris Richardson Creator of microservices.io, Java Champion, & Core Microservices Thoughtleader 14:45 Seacliff D Unconference Unconference: Modern Platform Engineering and Dev Enablement 15:55 Ballroom BC Session Platform Engineering: Lessons from the Rise and Fall of eBay Velocity Randy Shoup SVP Engineering @Thrive Market, Previously @eBay, @Google, @Stitch Fix 17:05 Ballroom BC Session Resilience Enhancing Reliability Using Service-Level Prioritized Load Shedding at Netflix Anirudh Mendiratta, Benjamin Fedorka