Scaling Large Language Model Serving Infrastructure at Meta

QCon San Francisco 2024

Session

Scaling Large Language Model Serving Infrastructure at Meta

Tuesday Nov 19 / 10:35AM PST, Ballroom BC

Abstract

Running LLMs requires significant computational power, which scales with model size and context length. We will discuss strategies for fitting models to various hardware configurations and share techniques for optimizing inference latency and throughput at Meta.

As we transition from stand-alone LLMs to production grade systems that support LLM at a global scale, we delve into our approach to constructing systems that accommodate dynamic user requests and widespread product adoption. This includes implementing caching strategies and addressing infra latency, efficiency and reliability issues within real data centers of a heterogeneous hardware fleet.

Finally, we will present case studies that demonstrate our methods for achieving a balance between model quality, latency, throughput, reliability, and cost in a complex and demanding environment.

76% senior dev or higher
1:11 speaker ratio
60+ practitioners

QCon San Francisco 2024 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.

Share

From the same track

Tuesday 19 November

10:35 Ballroom BC Session Scaling Large Language Model Serving Infrastructure at Meta Ye (Charlotte) Qi Senior Staff Engineer @Meta 11:45 Ballroom BC Session Generative AI GenAI for Productivity Mandy Gu Senior Software Development Manager @Wealthsimple 13:35 Pacific DEKJ Session LLMOps Navigating LLM Deployment: Tips, Tricks, and Techniques Meryem Arik Co-Founder and CEO @Doubleword (Previously TitanML), Recognized as a Technology Leader in Forbes 30 Under 30, Recovering Physicist 14:45 Seacliff ABC Session AI/ML Search: from Linear to Multiverse Faye Zhang Staff Software Engineer @Pinterest, Tech Lead on GenAI Search Traffic Projects, Speaker, Expert in AI/ML with a Strong Background in Large Distributed System 15:55 Seacliff ABC Session AI/ML 10 Reasons Your Multi-Agent Workflows Fail and What You Can Do About It Victor Dibia Principal Research Software Engineer @Microsoft Research, Core Contributor to AutoGen, Author of "Multi-Agent Systems with AutoGen" book. Previously @Cloudera, @IBMResearch 17:05 Ballroom BC Session Machine Learning A Framework for Building Micro Metrics for LLM System Evaluation Denys Linkov Head of ML @Voiceflow, LinkedIn Learning Instructor, ML Advisor and Instructor, Previously @LinkedIn