Orderly Keys, Wild Values: Adaptive Compression for Distributed Key-Value Storage

QCon San Francisco 2026

Session

Orderly Keys, Wild Values: Adaptive Compression for Distributed Key-Value Storage

Tuesday Nov 17 / 10:35AM PST, Ballroom BC at Hyatt Regency, San Francisco

Register

$2,835, Conference (3 days)
Current pricing ends September 8th

Abstract

At Netflix scale - billions of requests per day and petabytes of key-value data - even small inefficiencies in storage and network paths become expensive. This talk shares how we reduced the footprint of both keys and values through efficient binary tuple encoding and dynamic dictionary-based compression for a high-throughput, storage-agnostic key-value platform, and what we learned when database realities met compression theory.

We will walk through the engineering journey from prototype to production-minded rollout: selecting training data from live workloads, comparing compression strategies, and balancing quality against operational cost. Keys, which must remain ordered and stable, and Values which are highly variable in format and size require different approaches to optimally store. Furthermore, those approaches are data-dependent, so we will show how data shape directly influences encoding effectiveness and where naive approaches fail.

Most importantly, we will focus on system outcomes beyond compression ratio and how engineers gain trust in the approaches with rigorous verification. We will examine the impact on database/storage footprint, compaction and cache behavior, network IO, and p99 latency guardrails. We will also cover reliability patterns required in real systems: synthetic verification, simulation testing, dictionary versioning, compatibility/fallback paths, safe rollout controls, and failure handling when training signals are noisy or incomplete.

Attendees will leave with a practical framework for applying optimal encoding techniques in distributed storage systems: how to choose training pipelines, what signals to monitor, and how to get measurable efficiency gains without sacrificing latency or correctness.

What you will learn:

  1. Techniques for encoding both Keys and Values efficiently, they require different approaches!
  2. How to evaluate compression strategies using database/storage metrics (not just compression ratio), including footprint, IO, cache behavior, and tail latency.
  3. How workload characteristics (value sizes, churn, hot-key skew) should drive training-sample strategy and dictionary lifecycle decisions.
  4. How to design safe production rollouts with versioning, compatibility/fallback paths, observability, and fast rollback controls.
  5. How to build a repeatable and high confidence verification approach to compare training pipelines and make evidence-based trade-offs between efficiency and latency.
     
Register

$2,835, Conference (3 days). Current pricing ends September 8th. All pass options.

76% senior dev or higher
1:11 speaker ratio
60+ practitioners

QCon San Francisco 2026 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.

Share

From the same track

Tuesday 17 November

10:35 Ballroom BC Session Orderly Keys, Wild Values: Adaptive Compression for Distributed Key-Value Storage Joseph Lynch, Ayushi Singh 11:45 Ballroom BC Session When Your Users Are Agents: Lessons from a Distributed Postgres Platform Gwen Shapira Co-Founder and CPO @Nile, Previously Engineering Leader @Confluent, PMC Member @Kafka, & Committer Apache Sqoop 13:35 Ballroom BC Session Adaptive Systems in Production: What Recommendation Systems Can Teach Us About Agents Mallika Rao Senior Engineering Manager @Zocdoc, Previously @Netflix, @Twitter and @Walmart 14:45 Ballroom BC Session How to Build Online Systems with Object Storage Almog Gavra Co-Founder @Responsive.dev - Building Object-Native Databases, Previously @Confluent and @LinkedIn 15:55 Seacliff D Unconference Unconference: Distributed Systems in Production 17:05 Ballroom BC Session Autoscaling 800 Valkey Clusters: Distributed Systems Lessons from Scaling Stateful Fleets Kishor Yadav Kommanaboina Staff Software Engineer @Snapchat - Leading Caching and Storage Infrastructure

Current pricing ends September 8th
$2,835, Conference (3 days)

Register