Unleashing Llama's Potential: CPU-Based Fine-Tuning

QCon San Francisco 2024

Session AI/ML

Unleashing Llama's Potential: CPU-Based Fine-Tuning

Wednesday Nov 20 / 10:35AM PST, Seacliff ABC

Abstract

Generative AI landscape is rapidly changing as new models are appearing in horizon every few days. However, the hardware and software characteristics of these models have many similar patterns and execution phases.

In this talk, we will use Llama2 as base model to highlight basic characterization. We will present a detailed analysis of Llama2 workload performance on a platform powered by the AMD EPYC Processor. All our analysis was completed using the latest multi-core CPU servers. This includes scalability analysis as well as detailed phase by phase analysis to detect software and hardware bottlenecks at various stages.

Based on this, we will share our recommendations for tuning, optimization, and deployment best practices for the software stack with consideration of the hardware on which it is deployed. We extend our analysis to Llama3 using architecture relevant software optimization and share the best deployment practices relevant to the most AI inference deployment use cases.

Topics

AI/ML LLM Llama Microprocessor Architecture Platform Design Performance Hardware Software EPYC
76% senior dev or higher
1:11 speaker ratio
60+ practitioners

QCon San Francisco 2024 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.

Share

From the same track

Wednesday 20 November

10:35 Seacliff ABC Session AI/ML Unleashing Llama's Potential: CPU-Based Fine-Tuning Anil Rajput, Rema Hariharan 11:45 Seacliff ABC Session AI HW/SW optimization Maximizing Deep Learning Performance on CPUs using Modern Architectures Bibek Bhattarai AI Technical Lead @Intel, Computer Scientist Invested in Hardware-Software Optimization, Building Scalable Data Analytics, Mining, and Learning Systems 13:35 Seacliff ABC Session Hybrid cloud Evaluating and Deploying State-of-the-Art Hardware to Meet the Challenges of Modern Workloads Rebecca Weekly VP of Infrastructure @GEICO 14:45 Seacliff ABC Session RISC-V Optimizing Custom Workloads with RISC-V Ludovic Henry Member of Technical Staff @Rivos, Performance-Minded Engineer, Hardware & Software, Previously @Xamarin, @Microsoft, @Datadog 15:55 Seacliff ABC Session High-Resolution Platform Observability Brian Martin Co-founder and Software Engineer @IOP Systems, Focused on High-Performance Software and Systems, Previously @Twitter