Azure Cosmos DB: Low Latency and High Availability at Planet Scale

QCon San Francisco 2022

Session Architecture

Azure Cosmos DB: Low Latency and High Availability at Planet Scale

Wednesday Oct 26 / 01:40PM PDT, Ballroom A

Abstract

Azure Cosmos DB is a fully-managed, multi-tenant, distributed, shared-nothing, horizontally scalable database that provides planet-scale capabilities and multi-model APIs for Apache Cassandra, MongoDB, Gremlin, Tables, and the Core (SQL) APIs. It currently powers many mission-critical services both within Microsoft (such as Microsoft Teams and Active Directory) and across large-scale Fortune 500 organizations (such as Walmart and Adobe). 

This talk covers the internal architecture of Azure Cosmos DB and how it achieves high availability, low latency, and scalability. We will first cover the design of the storage engine, with particular emphasis on ensuring high availability and scalability through partitioning and replication. Next, we will zoom in on the request routing gateway to see how it has evolved to solve the well-known multi-tenant cloud infrastructure challenges of containing noisy neighbors and limiting blast radius. Lastly, we will discuss performance as a feature and as a culture. We will cover what we measure and how we think about SLOs to achieve and maintain low latency. 

Building planet-scale services necessitates solving complex scalability challenges and making numerous tradeoffs across various components in the product. We look forward to sharing our experiences and lessons learned in building Azure Cosmos DB.

Topics

Architecture High Availability Low Latency Scalability Storage Engine Partitioning and Replication Request Routing Gateway Cloud Infrastructure
76% senior dev or higher
1:11 speaker ratio
60+ practitioners

QCon San Francisco 2022 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.

Share

From the same track

Wednesday 26 October

10:35 Ballroom A Session Architecture Amazon DynamoDB: Evolution of a Hyper-Scale Cloud Database Service Akshat Vig Distinguished Engineer @MongoDB, Previously Senior Principal Engineer NoSQL@AWS 11:50 Ballroom A Session Architecture Honeycomb: How We Used Serverless to Speed Up Our Servers Jessica Kerr Principal Developer Evangelist @honeycombio 13:40 Ballroom A Session Architecture Azure Cosmos DB: Low Latency and High Availability at Planet Scale Mei-Chin Tsai, Vinod Sridharan 14:55 Ballroom A Session Architecture From Zero to A Hundred Billion: Building Scalable Real Time Event Processing At DoorDash Allen Wang Software Engineer @DoorDash, previously Lead for real-time data infrastructure team @Netflix 16:10 Ballroom A Session Architecture Magic Pocket: Dropbox’s Exabyte-Scale Blob Storage System Facundo Agriel Software Engineer / Tech Lead @Dropbox, previously @Amazon