Speaker
Abstract
How does Netflix maintain a seamless viewing experience for millions of users, especially during traffic spikes or when backend datastores are overloaded? Autoscaling can help during traffic spikes, but it costs money, takes a few minutes to kick in, and capacity may not always be available. Furthermore, if a downstream service or datastore is overloaded, autoscaling may exacerbate the problem.
In this talk, we will cover service-level prioritized load shedding – our solution for prioritizing requests within a single application instance. This innovative solution ensures requests that are critical to user experience maintain high availability, and allows dynamically re-purposing non-critical capacity to serve critical traffic during times of duress. We will also discuss how we automated per-cluster tuning and validation of load shedding, enabling us to quickly deploy unique configurations to hundreds of clusters.
Key Takeaways
- Understand the evolution of load shedding techniques at Netflix and how service-level prioritization enhances user experience and reliability.
- Gain insights from real-world applications and testing scenarios that demonstrate the effectiveness of prioritized load shedding.
- Learn how platform engineering and service owners at Netflix effectively collaborated to build a generic prioritized load shedding library, efficiently roll it out, and automate tuning of load shedding thresholds.
- Learn how we are continuing to improve the experience for service owners through our continued investments in making the load shedding capability more flexible and easier to operate.
Topics
QCon San Francisco 2025 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.
From the same track
Monday 17 November
10:35 Seacliff ABC Session Platform Engineering Continuous Delivery for Foundational Platforms Ian Nowland CEO @Junction Labs, Author of O'Reilly's Platform Engineering, Previously SVP Core Engineering at Datadog and Leader of AWS Nitro Platform teams frequently inherit systems that were never architected for their current scale, yet are so foundational that downtime can halt the business. 11:45 Seacliff ABC Session Beyond Line Charts: Why Some Diversity in Telemetry Visualization Is Long Overdue Yao Yue Founder & Chief Executive Officer @IOP Systems, Platform Engineer, Distributed System Aficionado, Cache Expert For decades, visualization of service metrics overwhelmingly converges to line charts. The time-centric nature of real-time telemetry further cemented this phenomenon via storage layouts and domain-specific query languages. 13:35 Ballroom BC Session Microservices Platforms: When Team Topologies Meets Microservices Patterns Chris Richardson Creator of microservices.io, Java Champion, & Core Microservices Thoughtleader When many teams work on a large, complex application, the microservice architecture potentially enables them to work independently and deliver a continuous stream of changes. 14:45 Seacliff D Unconference Unconference: Modern Platform Engineering and Dev Enablement 15:55 Ballroom BC Session Platform Engineering: Lessons from the Rise and Fall of eBay Velocity Randy Shoup SVP Engineering @Thrive Market, Previously @eBay, @Google, @Stitch Fix Once a stock market darling and a pioneering hyperscaler in the 1990s and early 2000s, eBay has been in steady decline since the 2010s. A household name with a flat business, eBay has been unable to make substantive strides in its market reach or its engineering outcomes in the last 15 years. 17:05 Ballroom BC Session Resilience Enhancing Reliability Using Service-Level Prioritized Load Shedding at Netflix Anirudh Mendiratta, Benjamin Fedorka How does Netflix maintain a seamless viewing experience for millions of users, especially during traffic spikes or when backend datastores are overloaded? Autoscaling can help during traffic spikes, but it costs money, takes a few minutes to kick in, and capacity may not always be available.