Speaker
Abstract
Azure Cosmos DB is a fully-managed, multi-tenant, distributed, shared-nothing, horizontally scalable database that provides planet-scale capabilities and multi-model APIs for Apache Cassandra, MongoDB, Gremlin, Tables, and the Core (SQL) APIs. It currently powers many mission-critical services both within Microsoft (such as Microsoft Teams and Active Directory) and across large-scale Fortune 500 organizations (such as Walmart and Adobe).
This talk covers the internal architecture of Azure Cosmos DB and how it achieves high availability, low latency, and scalability. We will first cover the design of the storage engine, with particular emphasis on ensuring high availability and scalability through partitioning and replication. Next, we will zoom in on the request routing gateway to see how it has evolved to solve the well-known multi-tenant cloud infrastructure challenges of containing noisy neighbors and limiting blast radius. Lastly, we will discuss performance as a feature and as a culture. We will cover what we measure and how we think about SLOs to achieve and maintain low latency.
Building planet-scale services necessitates solving complex scalability challenges and making numerous tradeoffs across various components in the product. We look forward to sharing our experiences and lessons learned in building Azure Cosmos DB.
Topics
QCon San Francisco 2022 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.
Part of the track
Architectures You've Always Wondered About Hosted by Randy Shoup SVP Engineering @Thrive Market, Previously @eBay, @Google, @Stitch FixFrom the same track
Wednesday 26 October
10:35 Ballroom A Session Architecture Amazon DynamoDB: Evolution of a Hyper-Scale Cloud Database Service Akshat Vig Distinguished Engineer @MongoDB, Previously Senior Principal Engineer NoSQL@AWS Amazon DynamoDB is a cloud database service that provides consistent performance at any scale. Hundreds of thousands of customers rely on DynamoDB for its fundamental properties: consistent performance, availability, durability, and a fully managed serverless experience. 11:50 Ballroom A Session Architecture Honeycomb: How We Used Serverless to Speed Up Our Servers Jessica Kerr Principal Developer Evangelist @honeycombio Honeycomb is the state of the art in observability: customers send us lots of data and then compose complex, ad-hoc queries. Most are simple, some are not. Some are REALLY not; this load is both complex, spontaneous, and urgent. 13:40 Ballroom A Session Architecture Azure Cosmos DB: Low Latency and High Availability at Planet Scale Mei-Chin Tsai, Vinod Sridharan Azure Cosmos DB is a fully-managed, multi-tenant, distributed, shared-nothing, horizontally scalable database that provides planet-scale capabilities and multi-model APIs for Apache Cassandra, MongoDB, Gremlin, Tables, and the Core (SQL) APIs. 14:55 Ballroom A Session Architecture From Zero to A Hundred Billion: Building Scalable Real Time Event Processing At DoorDash Allen Wang Software Engineer @DoorDash, previously Lead for real-time data infrastructure team @Netflix At DoorDash, real time events are an important data source to gain insight into our business but building a system capable of handling billions of real time events is challenging. 16:10 Ballroom A Session Architecture Magic Pocket: Dropbox’s Exabyte-Scale Blob Storage System Facundo Agriel Software Engineer / Tech Lead @Dropbox, previously @Amazon Magic Pocket is used to store all of Dropbox’s data.