Speaker
Abstract
Sharding is easy. Resharding under heavy load is notoriously difficult. How do you move gigabytes of state across live database nodes without dropping keys, blocking the main event loop, or breaking client abstractions?
Using Valkey and Redis as case studies, we will survey different resharding architectures and dive deep into Valkey's new Atomic Slot Migration. We'll walk through the practical tradeoffs of these approaches, covering client redirections (MOVED/ASK), fork-based slot snapshotting, and rollback staging. Along the way, we'll shine a light on the rough edge cases that actually matter in production.
$2,835, Conference (3 days). Current pricing ends September 8th. All pass options.
QCon San Francisco 2026 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.
Part of the track
Architectures You've Always Wondered About Hosted by Daniela Miao Co-Founder & CTO @Momento, Systems & Observability Nerd, ex-Lightstep, ex-DynamoDBFrom the same track
Monday 16 November
10:35 Ballroom BC Session Live Resharding Without Regret: Lessons from Building Valkey's Atomic Slot Migration Jacob Murphy Open Source Maintainer @Valkey & Software Engineer @Google Cloud's Memorystore Team Sharding is easy. Resharding under heavy load is notoriously difficult. How do you move gigabytes of state across live database nodes without dropping keys, blocking the main event loop, or breaking client abstractions? 11:45 Ballroom BC Session How to Build a Real-Time Voice Agent Rishabh Bhargava Director of ML @Together AI A voice agent looks like a chatbot with a microphone. 13:35 Ballroom BC Session Architecting Nubank's Global Financial Infrastructure Instant payment frameworks are transforming global finance, but few institutions have faced the infrastructure scale required by Brazil's Pix network. 14:45 Ballroom BC Session Inside the Disaggregated Architecture Serving LLM Inference Every LLM request is really two very different jobs. Reading your prompt (prefill) is compute-intensive, while generating the response token by token (decode) is dominated by memory bandwidth. For years, the industry ran both phases on the same hardware and accepted the compromise. 15:55 Ballroom BC Session Architecting Real-Time Timeline Engines for 100 Million Active Users Details coming soon. 17:05 Seacliff D Unconference Unconference: Architectures You've Always Wondered About