Speaker
Abstract
At Netflix scale - billions of requests per day and petabytes of key-value data - even small inefficiencies in storage and network paths become expensive. This talk shares how we reduced the footprint of both keys and values through efficient binary tuple encoding and dynamic dictionary-based compression for a high-throughput, storage-agnostic key-value platform, and what we learned when database realities met compression theory.
We will walk through the engineering journey from prototype to production-minded rollout: selecting training data from live workloads, comparing compression strategies, and balancing quality against operational cost. Keys, which must remain ordered and stable, and Values which are highly variable in format and size require different approaches to optimally store. Furthermore, those approaches are data-dependent, so we will show how data shape directly influences encoding effectiveness and where naive approaches fail.
Most importantly, we will focus on system outcomes beyond compression ratio and how engineers gain trust in the approaches with rigorous verification. We will examine the impact on database/storage footprint, compaction and cache behavior, network IO, and p99 latency guardrails. We will also cover reliability patterns required in real systems: synthetic verification, simulation testing, dictionary versioning, compatibility/fallback paths, safe rollout controls, and failure handling when training signals are noisy or incomplete.
Attendees will leave with a practical framework for applying optimal encoding techniques in distributed storage systems: how to choose training pipelines, what signals to monitor, and how to get measurable efficiency gains without sacrificing latency or correctness.
What you will learn:
- Techniques for encoding both Keys and Values efficiently, they require different approaches!
- How to evaluate compression strategies using database/storage metrics (not just compression ratio), including footprint, IO, cache behavior, and tail latency.
- How workload characteristics (value sizes, churn, hot-key skew) should drive training-sample strategy and dictionary lifecycle decisions.
- How to design safe production rollouts with versioning, compatibility/fallback paths, observability, and fast rollback controls.
- How to build a repeatable and high confidence verification approach to compare training pipelines and make evidence-based trade-offs between efficiency and latency.
Interview
Ayushi Singh: This session is co-presented with Joey Lynch - he'll cover the key-encoding side of our work, binary tuple encoding, while I'll focus on the value side: dynamic, workload-aware dictionary compression. Together, key and value encoding are what make this approach work at scale. At Netflix, our key-value storage platform operates at multi-petabyte scale, and my focus is on that value-compression architecture. I'll walk through how I implemented the dictionary training and compression pipeline itself, along with the production safeguards built around it (dictionary versioning, fallback paths, and rollback mechanisms), and how we're rolling it out: a phased, canary-driven rollout rather than a single cutover, so we can validate impact on storage, compaction, and latency incrementally before it's running against full production traffic.
I'll also cover a supplementary design question: where compression should actually live, implementing it at the underlying datastore layer versus at the abstraction layer above it. That choice has real consequences for portability, operability, and how much control you have over compression behavior across different backing stores. Values also vary enormously in size, entropy, and access pattern, so a static, one-size-fits-all compression strategy leaves meaningful storage and infrastructure efficiency on the table.
Ayushi Singh: Storage and network costs represent one of the fastest-growing categories of infrastructure spend industry-wide, so this is no longer just a cost problem - it's a reliability engineering problem. I approach value compression at scale as a first-class reliability decision, not a secondary optimization, because it directly affects compaction throughput and p99 latency. Designing this correctly is what allows systems processing this volume of data to remain both cost-efficient and stable in production.
Ayushi Singh: The most common failure I've seen is evaluating compression success purely by compression ratio, which tells you almost nothing about actual system behavior in production. Another challenge is deciding where compression logic should live: implementing it directly at the datastore layer can be more performant but ties you to that store's internals, while implementing it at the abstraction layer trades some of that performance for portability and consistent behavior across multiple backing stores. There's also a tendency to underestimate the production safeguards this kind of change requires once it's running on live traffic - versioning, fallback paths, and rollback mechanisms are essential, not optional extras.
Ayushi Singh: I want people to leave with a practical playbook for building and testing this architecture in their own production environments. They should walk away understanding the end-to-end encoding pipeline, the non-negotiable safeguards (versioning, fallback, rollback), and how to use a phased, canary-driven approach to validate compression before rolling it out to live traffic.
Ayushi Singh: At its heart, QCon is about the developer journey. The stage belongs to engineers sharing what actually happens when design hits scale - the post-mortems, trade-offs, and critical breakthroughs. It's unmatched peer-to-peer learning designed to help software practitioners make better architectural decisions and accelerate their own technical growth
Topics
$2,955, Conference (3 days). Current pricing ends October 13th. All pass options.
QCon San Francisco 2026 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.
From the same track
Tuesday 17 November
10:35 Ballroom BC Session Data Platforms Orderly Keys, Wild Values: Adaptive Compression for Distributed Key-Value Storage Joseph Lynch, Ayushi Singh At Netflix scale - billions of requests per day and petabytes of key-value data - even small inefficiencies in storage and network paths become expensive. 11:45 Ballroom BC Session Databases How to Build Online Systems with Object Storage Almog Gavra Co-Founder @Responsive.dev - Building Object-Native Databases, Previously @Confluent and @LinkedIn “Diskless” systems that delegate durability to object storage are everywhere, and for three good reasons: 13:35 Ballroom BC Session Autoscaling 800 Valkey Clusters: Distributed Systems Lessons from Scaling Stateful Fleets Kishor Yadav Kommanaboina Staff Software Engineer @Snapchat - Leading Caching and Storage Infrastructure Autoscaling stateless services is a solved problem. In a sharded datastore, a resize is not a capacity change — it's moving ownership of key ranges between nodes while the cluster serves live traffic, and rebalancing takes long enough that reacting to a traffic peak is already too late. 14:45 Seacliff D Unconference Unconference: Distributed Systems in Production 15:55 Ballroom BC Session Distributed Systems When Your Users Are Agents: Lessons from a Distributed Postgres Platform Gwen Shapira Co-Founder and CPO @Nile, Previously Engineering Leader @Confluent, PMC Member @Kafka, & Committer Apache Sqoop Distributed systems are built around assumptions about workload behavior: connections have reasonable lifetimes, retries eventually stop, traffic spikes have recognizable causes, and application code produces somewhat predictable query patterns. 17:05 Ballroom BC Session Adaptive Systems in Production: What Recommendation Systems Can Teach Us About Agents Mallika Rao Senior Engineering Manager @Zocdoc, Previously @Netflix, @Twitter and @Walmart As organizations race to build AI agents, many teams are encountering challenges that feel new: evaluation uncertainty, feedback loops, behavioral drift, exploration versus exploitation, and maintaining user trust in systems that continuously adapt. But these challenges are not new.