Speaker
Abstract
As deep learning continues to drive advancements across various industries, efficiently navigating the landscape of specialized AI hardware has a huge impact on cost and speed of operation. In addition, unleashing the full potential of these hardware through appropriate software stacks can be daunting.
This talk explores the advancements in modern CPU processors for enhanced AI capabilities and acceleration of underlying computation elements, specifically General Matrix Multiply (GEMM) operations. It will dive deep into the Intel Advanced Matrix Extensions (AMX) built into modern data-center CPUs and how to use them to perform efficient low-precision matrix operations. Additionally, we will explore software tools and frameworks that unlock the full performance of these accelerators, offering actionable insights for kernel developers, framework engineers, and data scientists.
Topics
QCon San Francisco 2024 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.
From the same track
Wednesday 20 November
10:35 Seacliff ABC Session AI/ML Unleashing Llama's Potential: CPU-Based Fine-Tuning Anil Rajput, Rema Hariharan Generative AI landscape is rapidly changing as new models are appearing in horizon every few days. However, the hardware and software characteristics of these models have many similar patterns and execution phases. 11:45 Seacliff ABC Session AI HW/SW optimization Maximizing Deep Learning Performance on CPUs using Modern Architectures Bibek Bhattarai AI Technical Lead @Intel, Computer Scientist Invested in Hardware-Software Optimization, Building Scalable Data Analytics, Mining, and Learning Systems As deep learning continues to drive advancements across various industries, efficiently navigating the landscape of specialized AI hardware has a huge impact on cost and speed of operation. 13:35 Seacliff ABC Session Hybrid cloud Evaluating and Deploying State-of-the-Art Hardware to Meet the Challenges of Modern Workloads Rebecca Weekly VP of Infrastructure @GEICO At GEICO we are on a journey to entirely modernize our Infrastructure. We are building an open-source, cloud-agnostic hybrid stack to run across public and on prem private cloud infrastructure without having to expose vendor specific stacks to our application developers. 14:45 Seacliff ABC Session RISC-V Optimizing Custom Workloads with RISC-V Ludovic Henry Member of Technical Staff @Rivos, Performance-Minded Engineer, Hardware & Software, Previously @Xamarin, @Microsoft, @Datadog This talk will explore how RISC-V architecture can accelerate custom workloads, focusing on AI/ML applications. We’ll start by examining the RISC-V ecosystem and its increasing relevance in the software development landscape. 15:55 Seacliff ABC Session High-Resolution Platform Observability Brian Martin Co-founder and Software Engineer @IOP Systems, Focused on High-Performance Software and Systems, Previously @Twitter Many observability tools fail to provide us with the relevant insights for understanding hardware health and utilization.