Abstract
Instant payment frameworks are transforming global finance, but few institutions have faced the infrastructure scale required by Brazil's Pix network. At over 130 million customers, managing this transactional volume demands an architecture that cannot afford traditional banking maintenance windows, consistency trade-offs, or runaway cloud costs.
In this session, Cat Swetel breaks down Nubank's runtime environments, database topologies, and zero-downtime designs required to support massive financial scale. We will dive deep into the specific socio-technical and systems engineering trade-offs made to keep pace with explosive growth, handle unpredictable traffic spikes without over-provisioning, and ensure sub-second global availability under rigid regulatory constraints.
$2,835, Conference (3 days). Current pricing ends September 8th. All pass options.
QCon San Francisco 2026 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.
Part of the track
Architectures You've Always Wondered About Hosted by Daniela Miao Co-Founder & CTO @Momento, Systems & Observability Nerd, ex-Lightstep, ex-DynamoDBFrom the same track
Monday 16 November
10:35 Ballroom BC Session Live Resharding Without Regret: Lessons from Building Valkey's Atomic Slot Migration Jacob Murphy Open Source Maintainer @Valkey & Software Engineer @Google Cloud's Memorystore Team Sharding is easy. Resharding under heavy load is notoriously difficult. How do you move gigabytes of state across live database nodes without dropping keys, blocking the main event loop, or breaking client abstractions? 11:45 Ballroom BC Session How to Build a Real-Time Voice Agent Rishabh Bhargava Director of ML @Together AI A voice agent looks like a chatbot with a microphone. 13:35 Ballroom BC Session Architecting Nubank's Global Financial Infrastructure Instant payment frameworks are transforming global finance, but few institutions have faced the infrastructure scale required by Brazil's Pix network. 14:45 Ballroom BC Session Inside the Disaggregated Architecture Serving LLM Inference Every LLM request is really two very different jobs. Reading your prompt (prefill) is compute-intensive, while generating the response token by token (decode) is dominated by memory bandwidth. For years, the industry ran both phases on the same hardware and accepted the compromise. 15:55 Ballroom BC Session Architecting Real-Time Timeline Engines for 100 Million Active Users Details coming soon. 17:05 Seacliff D Unconference Unconference: Architectures You've Always Wondered About