No More Spray and Pray— Let's Talk About LLM Evaluations

Session AI/ML

No More Spray and Pray— Let's Talk About LLM Evaluations

Abstract

The pace of development in AI in the past year or so has been dizzying, to say the least, with new models and techniques emerging weekly. Yet, amidst the hype, a sobering reality emerges: much of these advancements lack robust empirical evidence. Recent surveys reveal that testing and evaluation practices for generative AI features across organizations remain in their infancy. It’s time we start treating LLM systems like any other software, subjecting them to the same, if not greater level of rigorous testing.

In this talk, we explore the landscape of LLM system evaluation. Specifically, we will focus on challenges with evaluating LLM systems, determining metrics for evaluation, existing tools and techniques, and of course how to evaluate an LLM system. Finally, we will bring all these pieces together by walking through an end-to-end evaluation study of a real LLM system we’ve built.

Key Takeaways:
The main takeaways from this talk are intended to be a better understanding of the importance of evaluation when it comes to LLMs and for attendees to leave with a practical framework for LLM evaluation that they can apply to their projects.
 

Interview

Each day looks different but you'll find me doing one of these: head down researching and writing code for tutorials covering different aspects of building AI applications, creating and delivering hands-on AI workshops for our customers, or advising customers on how to go about building AI applications for different use cases.

Having worked as a data scientist for several years, thoroughly evaluating ML models before shipping them was a big part of the process, especially when building models for customer-facing applications. With LLM-based applications, there has been a tendency to ship fast to get ahead of the race, without spending the time and effort to test them, largely because there are no existing best practices and guidelines to do this and I wanted to change that. 

This session is mainly targeted toward anyone building customer-facing AI applications. 

The main takeaways from this talk are intended to be a better understanding of the importance of evaluation when it comes to LLMs and for attendees to leave with a practical framework for LLM evaluation that they can apply to their projects.

It's here and it's AI (I'm totally not biased!). We have yet to see the full potential of AI, but I am excited to see how it changes how we work, write software, etc.   

Topics

AI/ML LLMs Evaluations Monitoring
76% senior dev or higher
1:11 speaker ratio
60+ practitioners

is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.

Share