Modern Compute Stack for Scaling Large AI/ML/LLM Workloads

QCon San Francisco 2023

Session Distributed Computing

Modern Compute Stack for Scaling Large AI/ML/LLM Workloads

Tuesday Oct 3 / 01:35PM PDT, Ballroom A

Abstract

Advanced machine learning (ML)  models, particularly large language models (LLMs), require scaling beyond a single machine. As open-source LLMs become more prevalent on platforms and model hubs like HuggingFace (HF), ML practitioners and GenAI developers are increasingly inclined to fine-tune these models with their private data to suit their specific needs.

However, several concerns arise: which compute infrastructure should be used for distributed fine-tuning and training? How can ML workloads be effectively scaled for data ingestion, training/tuning, or inference? How can large models be accommodated within a cluster? And how can CPUs and GPUs be optimally utilized?

Fortunately, an opinionated stack is emerging among ML practitioners, leveraging open-source libraries.

This session focuses on the integration of HuggingFace and Ray AI Runtime (AIR), enabling scaling of model training and data loading. We’ll delve into implementation details, explore the 

Transformer APIs, and demonstrate how Ray AIR facilitates an end-to-end ML workflow, encompassing data ingestion, training/tuning, or inference.

By exploring the integration between HF and Ray AIR, we’ll discuss how Ray’s orchestration capabilities fulfill computation and memory requirements. Also, we’ll showcase how existing HF Transformer APIs, DeepSpeed, and Accelerate code can seamlessly integrate with Ray AIR’s Trainers and demonstrate its capabilities within this emerging component stack. Finally, we’ll demonstrate how to fine-tune an open-source LLM model with HF Transformer APIs and Ray AIR Trainers.

Topics

Distributed Computing Scaling AI/ML/LLM Workloads
76% senior dev or higher
1:11 speaker ratio
60+ practitioners

QCon San Francisco 2023 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.

Share

From the same track

Tuesday 3 October

10:35 Seacliff ABC Session AI/ML Chronon - Airbnb’s End-to-End Feature Platform Nikhil Simha Author of "Chronon Feature Platform", Previously Built Stream Processing Infra @Meta and NLP Systems @Amazon & @Walmartlabs 11:45 Seacliff ABC Session AI/ML Defensible Moats: Unlocking Enterprise Value with Large Language Models Nischal HP Vice President of Data Science @Scoutbee, Decade of Experience Building Enterprise AI 13:35 Ballroom A Session Distributed Computing Modern Compute Stack for Scaling Large AI/ML/LLM Workloads Jules Damji Lead Developer Advocate @Anyscale, MLflow Contributor, and Co-Author of "Learning Spark" 14:45 Pacific DEKJ Session AI/ML Generative Search: Practical Advice for Retrieval Augmented Generation (RAG) Sam Partee Principal Engineer @Redis 15:55 Seacliff D Unconference Unconference: Modern ML 17:05 Ballroom BC Session AI/ML Building Guardrails for Enterprise AI Applications W/ LLMs Shreya Rajpal Founder @Guardrails AI, Experienced ML Practitioner with a Decade of Experience in ML Research, Applications and Infrastructure