Speaker
Abstract
LinkedIn’s Kubernetes-based compute platform spans more than 500k bare-metal machines, 5M pods, with thousands of developers doing more than 110k deploys a week. We don’t expose raw Kubernetes to developers and instead, provide a curated set of platform offerings.
We extend Kubernetes with APIs for stateless and stateful applications. Deployment and self-service debugging are first-class platform capabilities, so developers can run and troubleshoot their applications without needing to understand the underlying Kubernetes details.
We built the compute platform as a layered system that connects machine lifecycle, Kubernetes cluster management, and workload platforms. These layers coordinate infrastructure maintenance with application needs, allowing us to operate the fleet and its workloads together.
This talk explains how we designed these layers and the contracts between them, and how this architecture supported a centrally run migration of over 25M CPU cores. Existing applications could run on the old and the new platforms simultaneously, allowing gradual migration without application changes. We’ll share the tradeoffs behind what we expose, what we extend, and what we deliberately leave out of Kubernetes, and how these choices help us operate at scale while keeping deployment and debugging simple for developers.
$2,955, Conference (3 days). Current pricing ends October 13th. All pass options.
QCon San Francisco 2026 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.
Part of the track
Real World Platform Engineering Hosted by Daniel Bryant Platform Engineer, Co-Author of "Mastering API Architecture", Java Champion, and InfoQ News ManagerFrom the same track
Wednesday 18 November
10:35 Pacific DEKJ Session Why Most Platform Teams Fail: The Adoption Problem Nobody Wants to Own Shweta Vohra Architecture Leader @Booking.com, Author of "Decoding Platform Engineering Patterns" & "Dear Software and AI Architect", 24+ Years Experience Building Cloud, Platform, and AI Systems We have all seen the moment: the platform goes live, the launch deck looks sharp, the portal is polished, the golden paths are documented, and yet teams quietly continue doing things the old way. Not always because the platform is bad, but because adoption was assumed, not owned. 11:45 Pacific DEKJ Session Beyond the Kubernetes API: Building LinkedIn’s Compute Platform Ronak Nathani Principal Staff Software Engineer, Compute Infra @LinkedIn, Podcast Host @Software Misadventures LinkedIn’s Kubernetes-based compute platform spans more than 500k bare-metal machines, 5M pods, with thousands of developers doing more than 110k deploys a week. We don’t expose raw Kubernetes to developers and instead, provide a curated set of platform offerings. 13:35 Pacific DEKJ Session Decoupling Local Development from Remote Infrastructure Details coming soon. 14:45 Pacific DEKJ Session Platform Engineering Platform Engineering’s Second Act: From Vending Machine to Passport Control Smruti Patel, Alex Mann Three years ago on the QCon SF stage, I made the case for “Acceleration, Autonomy, and Accountability” as the pillars of a successful platform. Those pillars haven't moved. AI has just rewritten what each one requires, and the platform team's job along with it. 15:55 Pacific DEKJ Session Platform Engineering Building a Migration Platform: Moving 100+ Netflix RDBMS Workloads to Aurora PostgreSQL Ammar Khaku, Kshitij Gupta In late 2024, Netflix made a bet: consolidate the vast majority of our relational database use cases onto a single engine: Amazon Aurora PostgreSQL.