Realtime and Batch Processing of GPU Workloads

QCon San Francisco 2025

Session AI Architecture

Realtime and Batch Processing of GPU Workloads

Wednesday Nov 19 / 01:35PM PST, Ballroom A at Hyatt Regency, San Francisco

Abstract

SS&C Technologies runs 47 trillion dollars of assets on our global private cloud. We have the primitives for infrastructure as well as platforms as a service like Kubernetes, Kafka, NiFi, Databases, etc. A year ago we broke ground and went live with AI as a service providing RAG, inference for embeddings, LLM text, image and voice and we needed an efficient and low TCO platform to power the needs of the business. Our centralized AI Gateway has a prioritized job scheduler that we wrote and we will discuss how over 300 production use cases run workloads in a way that provide the SLAs for the demands required while keeping the GPU costs down. We also run on AWS around the globe and will discuss how the platform works in a multi cloud environment also keeping costs down in different ways in AWS while also meeting SLAs.

Topics

AI Architecture Platform Engineering Distributed Systems
76% senior dev or higher
1:11 speaker ratio
60+ practitioners

QCon San Francisco 2025 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.

Share

From the same track

Wednesday 19 November

10:35 Ballroom A Session AI/ML Producing the World's Cheapest Tokens: A How-to Guide Meryem Arik Co-Founder and CEO @Doubleword (Previously TitanML), Recognized as a Technology Leader in Forbes 30 Under 30, Recovering Physicist 11:45 Ballroom A Session Capacity Planning How Netflix Shapes our Fleet for Efficiency and Reliability Joseph Lynch, Argha C 13:35 Ballroom A Session AI Architecture Realtime and Batch Processing of GPU Workloads Joseph Stein Principal Architect of Research & Development @SS&C Technologies, Previous Apache Kafka Committer and PMC Member 14:45 Ballroom A Session Architecture From ms to µs: OSS Valkey Architecture Patterns for Modern AI Dumanshu Goyal Uber Technical Lead @Airbnb Powering $11B Transactions, Formerly @Google and @AWS 15:55 Ballroom A Session Platform Engineering Write-Ahead Intent Log: A Foundation for Efficient CDC at Scale Vinay Chella, Akshat Goel