Navigating LLM Deployment: Tips, Tricks, and Techniques

QCon San Francisco 2024

Session LLMOps

Navigating LLM Deployment: Tips, Tricks, and Techniques

Tuesday Nov 19 / 01:35PM PST, Pacific DEKJ

Abstract

Self-hosted Language Models are going to power the next generation of applications in critical industries like financial services, healthcare, and defense. Self-hosting LLMs, as opposed to using API-based models, comes with its own host of challenges - as well as needing to solve business problems, engineers need to wrestle with the intricacies of model inference, deployment and infrastructure. In this talk we are going to discuss the best practices in model optimisation, serving and monitoring - with practical tips and real case-studies.

Topics

LLMOps AI/ML Deployment
76% senior dev or higher
1:11 speaker ratio
60+ practitioners

QCon San Francisco 2024 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.

Share

From the same track

Tuesday 19 November

10:35 Ballroom BC Session Scaling Large Language Model Serving Infrastructure at Meta Ye (Charlotte) Qi Senior Staff Engineer @Meta 11:45 Ballroom BC Session Generative AI GenAI for Productivity Mandy Gu Senior Software Development Manager @Wealthsimple 13:35 Pacific DEKJ Session LLMOps Navigating LLM Deployment: Tips, Tricks, and Techniques Meryem Arik Co-Founder and CEO @Doubleword (Previously TitanML), Recognized as a Technology Leader in Forbes 30 Under 30, Recovering Physicist 14:45 Seacliff ABC Session AI/ML Search: from Linear to Multiverse Faye Zhang Staff Software Engineer @Pinterest, Tech Lead on GenAI Search Traffic Projects, Speaker, Expert in AI/ML with a Strong Background in Large Distributed System 15:55 Seacliff ABC Session AI/ML 10 Reasons Your Multi-Agent Workflows Fail and What You Can Do About It Victor Dibia Principal Research Software Engineer @Microsoft Research, Core Contributor to AutoGen, Author of "Multi-Agent Systems with AutoGen" book. Previously @Cloudera, @IBMResearch 17:05 Ballroom BC Session Machine Learning A Framework for Building Micro Metrics for LLM System Evaluation Denys Linkov Head of ML @Voiceflow, LinkedIn Learning Instructor, ML Advisor and Instructor, Previously @LinkedIn