Scale Out Batch Inference with Ray

QCon San Francisco 2024

Session

Scale Out Batch Inference with Ray

Monday Nov 18 / 11:45AM PST, Ballroom BC

Abstract

As AI technologies continue to evolve, the demand for processing both structured and unstructured data across diverse industries is rapidly growing. However, scaling AI batch processing across thousands of GPUs presents significant challenges in maintaining scalability, reliability, and observability. These challenges are further amplified when aiming for high-throughput batch data processing with large language models (LLMs), due to their computational demands and complexity.

In this presentation, we will demonstrate how we built a scalable and efficient batch inference stack using Ray at Anyscale. We begin by introducing Ray as a robust, scalable AI compute engine, followed by an in-depth look at RayData, a versatile and high-performance deep learning data processing pipeline. Next, we will introduce vLLM, the leading open-source framework for LLM inference, and illustrate how the combination of RayData and vLLM offers an ideal solution for scalable batch inference.

76% senior dev or higher
1:11 speaker ratio
60+ practitioners

QCon San Francisco 2024 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.

Share

From the same track

Monday 18 November

10:35 Ballroom BC Session Knowledge Graphs Enhance LLMs’ Explainability and Trustworthiness With Knowledge Graphs Leann Chen AI Developer Advocate @Diffbot, Creator of AI and Knowledge Graph Content on YouTube, Passionate About Knowledge Graphs & Generative AI 11:45 Ballroom BC Session Scale Out Batch Inference with Ray Cody Yu Staff Software Engineer and Tech Lead @Anyscale, Ex-Amazonian, vLLM Committer, Apache TVM PMC 13:35 Ballroom BC Session AI/ML Recommender and Search Ranking Systems in Large Scale Real World Applications Moumita Bhattacharya Senior Research Scientist @Netflix, Previously @Etsy, Specialized in Machine Learning, Deep Learning, Big Data, Scala, Tensorflow, and Python 14:45 Ballroom BC Session AI/ML Why Most Machine Learning Projects Fail to Reach Production and How to Beat the Odds Wenjie Zi Senior Machine Learning Engineer and Tech Lead @Grammarly, Specializing in Natural Language Processing, 10+ Years of Industrial Experience in Artificial Intelligence Applications 15:55 Seacliff D Unconference Unconference: AI and ML for Software Engineers 17:05 Ballroom BC Session AI/ML Reinforcement Learning for User Retention in Large-Scale Recommendation Systems Saurabh Gupta, Gaurav Chakravorty