A Framework for Building Micro Metrics for LLM System Evaluation

QCon San Francisco 2024

Session Machine Learning

A Framework for Building Micro Metrics for LLM System Evaluation

Tuesday Nov 19 / 05:05PM PST, Ballroom BC

Abstract

LLM accuracy is a challenging topic to address and is much more multi dimensional than a simple accuracy score. In this talk we’ll dive deeper into how to measure LLM related metrics, going through examples, case studies and techniques beyond just a single accuracy and score. We’ll discuss how to create, track and revise micro LLM metrics to have granular direction for improving LLM models.

Interview

I'm mainly focused on applied research and helping teams build in the LLM and conversational AI space. The goal is to look at industry challenges, create accessible and practical research and guides that help us create better conversational experiences.

Each problem in the AI space, or any use case has unique challenges. There has been a lot of focus on catch-all metrics, but once you've been serving production traffic you'll find edge cases and scenarios you want to measure. This is where micro metrics can help, defining specific outputs and behaviors that you want to track for your use case.

A range between intermediate and senior developer and product lead. The concepts are pretty standard from a product, ML and software perspective - the learning comes from thinking through the provided case studies and how they can be applied to your own use cases.

Figuring out how to prompt and steer multi model models. Speech to speech is very exciting, but businesses need the ability to check for hallucinations and integrate with other services before responding to users.

Topics

Machine Learning LLMs Evaluations Case Study Lessons Learned
76% senior dev or higher
1:11 speaker ratio
60+ practitioners

QCon San Francisco 2024 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.

Share

From the same track

Tuesday 19 November

10:35 Ballroom BC Session Scaling Large Language Model Serving Infrastructure at Meta Ye (Charlotte) Qi Senior Staff Engineer @Meta 11:45 Ballroom BC Session Generative AI GenAI for Productivity Mandy Gu Senior Software Development Manager @Wealthsimple 13:35 Pacific DEKJ Session LLMOps Navigating LLM Deployment: Tips, Tricks, and Techniques Meryem Arik Co-Founder and CEO @Doubleword (Previously TitanML), Recognized as a Technology Leader in Forbes 30 Under 30, Recovering Physicist 14:45 Seacliff ABC Session AI/ML Search: from Linear to Multiverse Faye Zhang Staff Software Engineer @Pinterest, Tech Lead on GenAI Search Traffic Projects, Speaker, Expert in AI/ML with a Strong Background in Large Distributed System 15:55 Seacliff ABC Session AI/ML 10 Reasons Your Multi-Agent Workflows Fail and What You Can Do About It Victor Dibia Principal Research Software Engineer @Microsoft Research, Core Contributor to AutoGen, Author of "Multi-Agent Systems with AutoGen" book. Previously @Cloudera, @IBMResearch 17:05 Ballroom BC Session Machine Learning A Framework for Building Micro Metrics for LLM System Evaluation Denys Linkov Head of ML @Voiceflow, LinkedIn Learning Instructor, ML Advisor and Instructor, Previously @LinkedIn