The Freeze Paradox: Why Stopping Deployments Doesn't Stop Failures

QCon San Francisco 2026

Session Resiliency

The Freeze Paradox: Why Stopping Deployments Doesn't Stop Failures

Monday Nov 16 / 10:35AM PST, Seacliff ABC at Hyatt Regency, San Francisco

Register

$2,955, Conference (3 days)
Current pricing ends October 13th

Abstract

When systems fail at Netflix and in an environment this complex, they sometimes will, the instinctive response is to freeze deployments. Stop the bleeding,but blanket freezes carry their own hidden cost: delayed rollouts, missed business windows, and an illusion of safety that disappears the moment the freeze lifts and a backlog of pent-up changes floods production all at once.

The uncomfortable truth is that in a distributed system at Netflix's scale, not deploying is itself a risk. Bugs don't get fixed. Security patches wait. Rollback windows close. Blanket freezes optimize for one failure mode while creating several others.

In this session, Prachi Jain and Sandhya Narayan from Netflix Security Engineering will share how Netflix dismantled the one-size-fits-all freeze model and replaced it with service-aware risk controls, a framework that treats deployment risk as continuous and measurable, not binary. They'll walk through how Netflix: 

  • Classifies services by criticality and blast radius

  • Feeds deployment confidence scores and test coverage into pipeline gates

  • Empowers domain teams to make informed deployment decisions in real time, even during major incidents and high-traffic launch events.

Attendees will see the technical machinery: CI/CD risk integration, controlled bypass mechanisms with full auditability, canary rollouts, regional staggering, and the feedback loops that make the system self-improving. The result is a model where resilience comes not from pausing change, but from understanding which changes are safe to make and when.

Key Takeaways:

  • Why deployment freezes are a resilience anti-pattern at distributed scale, and what to replace them with

  • How to build a service risk classification framework using deployment confidence, test coverage, and blast radius

  • Practical implementation: integrating risk profiles into CI/CD pipelines, canary deployments, and real-time monitoring

  • How to give domain teams deployment autonomy without sacrificing system reliability

Interview

This session challenges a reflex that most engineering orgs share:  when something breaks, freeze all deployments. We will show why that instinct, while understandable, actually creates new risks at Netflix's scale, from delayed security patches to the dangerous pileup of changes that all land in production the moment a freeze lifts.     

We will walk through the alternative Netflix built: a service-aware risk framework that continuously classifies deployment risk by criticality, blast radius, test coverage, and confidence scores, instead of  treating every deployment as equally dangerous.

For senior developers, this reframes resilience itself. It's no longer about how effectively you can stop change, it's about how precisely you can identify which changes are safe, for which services, at any given moment.

As systems become more distributed and interdependent, the cost of blanket policies keeps rising. A freeze that was reasonable for a monolith with a handful of services becomes a blunt and expensive instrument when you are operating thousands of independently deployed services with different risk profiles.

The rise of AI-assisted development is accelerating this trend: teams are shipping more code, faster, and increasingly deploying AI-driven features whose behavior and failure modes are harder to predict upfront. At the same time, leaders are under more pressure than ever to move quickly without compromising reliability or security. That tension only grows during incidents or major launch events, exactly when teams are tempted to freeze everything. Getting this right means fewer missed business windows, faster security patching, and less firefighting caused by deployment backlogs. It's a resilience strategy that scales with system complexity, and with the pace AI is introducing, instead of fighting against it.

The biggest challenge is that most orgs treat deployment risk as binary: safe or unsafe, frozen or unfrozen. Building a system that treats risk as continuous and measurable requires real investment:  instrumenting services with meaningful confidence scores, understanding blast radius, integrating that data into CI/CD gates, and building auditable bypass mechanisms so teams aren't blocked when a genuinely low-risk change needs to go out during a freeze. Another challenge is organizational, not technical. Giving domain teams real autonomy to make deployment decisions requires trust, clear guardrails, and feedback loops so the system actually improves over time rather than becoming another source of alert fatigue or inconsistent judgment calls.

Start by classifying your services by criticality and blast radius, even if it is informal at first. Simply creating that map will reveal where a blanket freeze policy is doing more harm than good, and it becomes the foundation that everything else in our framework builds on.

QCon consistently brings together practitioners who are solving real, hard problems at scale rather than talking in abstractions. The sessions tend to go deep into the how, not just the what, which means the audience walks away with ideas they can actually apply. And because of that depth, the hallway conversations are often as valuable as the talks themselves.

Topics

Resiliency Reliability Security Risk Awareness Security
Register

$2,955, Conference (3 days). Current pricing ends October 13th. All pass options.

76% senior dev or higher
1:11 speaker ratio
60+ practitioners

QCon San Francisco 2026 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.

Share

From the same track

Monday 16 November

10:35 Seacliff ABC Session Resiliency The Freeze Paradox: Why Stopping Deployments Doesn't Stop Failures Prachi Jain, Sandhya Narayan 11:45 Seacliff ABC Session Reliability Saturation: How Your Software Will Fail at Scale Lorin Hochstein Staff Software Engineer @Airbnb, Writes @surfingcomplexity.blog, Previously @Netflix and Member of the Resilience in Software Foundation 13:35 Seacliff D Unconference Unconference: Resilience Engineering 14:45 Seacliff ABC Session Resilience Adapt or Drift: Resilience Engineering When AI Moves the Operating Point Andrew Hatch Engineering Leader and SRE Manager @Cisco ThousandEyes, With 25+ Years Building Software, Operations, SRE, and Platform Teams Across Australia, India, and the United States 15:55 Seacliff ABC Session Incidents Engineering Boundaries for Outages You Can't Prevent Em Ruppe Technical Incident Commander @Chime, Previously @SendGrid and @Twilio, and Product and Training @Jeli.io 17:05 Seacliff ABC Session It's Never the Thing That Broke: Learning From Incidents at (AI) Speed Vanessa Huerta Granda Resilience Engineering Manager @Enova, Co-Author of the Howie Guide on Post Incident Analysis, Board Member for the Resilience in Software Foundation

Current pricing ends October 13th
$2,955, Conference (3 days)

Register