Speaker
Abstract
When systems fail at Netflix and in an environment this complex, they sometimes will, the instinctive response is to freeze deployments. Stop the bleeding,but blanket freezes carry their own hidden cost: delayed rollouts, missed business windows, and an illusion of safety that disappears the moment the freeze lifts and a backlog of pent-up changes floods production all at once.
The uncomfortable truth is that in a distributed system at Netflix's scale, not deploying is itself a risk. Bugs don't get fixed. Security patches wait. Rollback windows close. Blanket freezes optimize for one failure mode while creating several others.
In this session, Prachi Jain and Sandhya Narayan from Netflix Security Engineering will share how Netflix dismantled the one-size-fits-all freeze model and replaced it with service-aware risk controls, a framework that treats deployment risk as continuous and measurable, not binary. They'll walk through how Netflix:
Classifies services by criticality and blast radius
Feeds deployment confidence scores and test coverage into pipeline gates
Empowers domain teams to make informed deployment decisions in real time, even during major incidents and high-traffic launch events.
Attendees will see the technical machinery: CI/CD risk integration, controlled bypass mechanisms with full auditability, canary rollouts, regional staggering, and the feedback loops that make the system self-improving. The result is a model where resilience comes not from pausing change, but from understanding which changes are safe to make and when.
Key Takeaways:
Why deployment freezes are a resilience anti-pattern at distributed scale, and what to replace them with
How to build a service risk classification framework using deployment confidence, test coverage, and blast radius
Practical implementation: integrating risk profiles into CI/CD pipelines, canary deployments, and real-time monitoring
How to give domain teams deployment autonomy without sacrificing system reliability
$2,835, Conference (3 days). Current pricing ends September 8th. All pass options.
QCon San Francisco 2026 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.
From the same track
Monday 16 November
10:35 Seacliff ABC Session The Freeze Paradox: Why Stopping Deployments Doesn't Stop Failures Prachi Jain, Sandhya Narayan When systems fail at Netflix and in an environment this complex, they sometimes will, the instinctive response is to freeze deployments. 11:45 Seacliff ABC Session Saturation: How Your Software Will Fail at Scale Lorin Hochstein Staff Software Engineer @Airbnb, Writes @surfingcomplexity.blog, Previously @Netflix and Member of the Resilience in Software Foundation Even in the ethereal world of software, everything has a limit. And, once that limit is reached, very bad things can happen. In this talk, we will explore the problem of saturation: when a software system runs into one of these limits. 13:35 Seacliff D Unconference Unconference: Resilience Engineering 14:45 Seacliff ABC Session Adapt or Drift: Resilience Engineering When AI Moves the Operating Point Andrew Hatch Engineering Leader and SRE Manager @Cisco ThousandEyes, With 25+ Years Building Software, Operations, SRE, and Platform Teams Across Australia, India, and the United States Resilient systems do not stay resilient by standing still. They survive by adapting. But adaptation comes with risk: under sustained pressure, organizations and systems can slowly drift toward failure while still appearing to operate normally. 15:55 Seacliff ABC Session Engineering Boundaries for Outages You Can't Prevent Em Ruppe Technical Incident Commander @Chime, Previously @SendGrid and @Twilio, and Product and Training @Jeli.io Details coming soon. 17:05 Seacliff ABC Session Turning Unseen Micro-Failures into Systemic Safeguards Details coming soon.