How GitHub Copilot Serves 400 Million Completion Requests a Day

QCon San Francisco 2024

Session HTTP

How GitHub Copilot Serves 400 Million Completion Requests a Day

Monday Nov 18 / 03:55PM PST, Ballroom A

Abstract

GitHub Copilot is the largest LLM powered Code Completion service in the world, serving hundreds of millions of requests per day with an average response time of under 200ms. This is the story of the architecture which powers this product.

Topics

HTTP Load balancing High Scale
76% senior dev or higher
1:11 speaker ratio
60+ practitioners

QCon San Francisco 2024 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.

Share

From the same track

Monday 18 November

10:35 Ballroom A Session MLOps Supporting Diverse ML Systems at Netflix David Berg, Romain Cledat 11:45 Ballroom A Session Architecture Optimizing Search at Uber Eats Janani Narayanan, Karthik Ramasamy 13:35 Ballroom A Session Architecture Changing the Model: Why and How We Re-Architected Slack Ian Hoffman Staff Software Engineer @Slack, Previously @Chairish 14:45 Seacliff D Unconference Unconference: Architectures You've Always Wondered About 15:55 Ballroom A Session HTTP How GitHub Copilot Serves 400 Million Completion Requests a Day David Cheney Lead, Copilot Proxy @GitHub, Open Source Contributor and Project Member for Go Programming Language, Previously @VMware 17:05 Ballroom A Session Legacy Modernization: Architecting Real-Time Systems Around a Mainframe Jason Roberts, Sonia Mathew