Producing the World's Cheapest Tokens: A How-to Guide

QCon San Francisco 2025

Session AI/ML

Producing the World's Cheapest Tokens: A How-to Guide

Wednesday Nov 19 / 10:35AM PST, Ballroom A at Hyatt Regency, San Francisco

Abstract

AI inference is expensive, but it doesn’t have to be. In this talk, we’ll break down how to systematically drive down the cost per token across different types of AI workloads. Using real-world examples from data transformation, offline agents, and aggregated insights, we’ll unpack how to measure, optimize, and ultimately produce the world’s cheapest tokens. The session will be hardware-agnostic, featuring analysis of both Nvidia and AMD GPUs, and will include advice which can be implemented by using open-source serving frameworks such as Dynamo, vLLM, and SGLang.

What you'll take away:

  1. Token Economics 101 - Understand what actually drives cost per token
  2. Inference Optimization Tactics that can be used to drive down unit economics depending on the AI workload type
  3. Right GPU, Right Job - Ho two choose hardware and deployment strategy for maximum cost performance

Topics

AI/ML Inference Platform Engineering K8s Scale
76% senior dev or higher
1:11 speaker ratio
60+ practitioners

QCon San Francisco 2025 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.

Share

From the same track

Wednesday 19 November

10:35 Ballroom A Session AI/ML Producing the World's Cheapest Tokens: A How-to Guide Meryem Arik Co-Founder and CEO @Doubleword (Previously TitanML), Recognized as a Technology Leader in Forbes 30 Under 30, Recovering Physicist 11:45 Ballroom A Session Capacity Planning How Netflix Shapes our Fleet for Efficiency and Reliability Joseph Lynch, Argha C 13:35 Ballroom A Session AI Architecture Realtime and Batch Processing of GPU Workloads Joseph Stein Principal Architect of Research & Development @SS&C Technologies, Previous Apache Kafka Committer and PMC Member 14:45 Ballroom A Session Architecture From ms to µs: OSS Valkey Architecture Patterns for Modern AI Dumanshu Goyal Uber Technical Lead @Airbnb Powering $11B Transactions, Formerly @Google and @AWS 15:55 Ballroom A Session Platform Engineering Write-Ahead Intent Log: A Foundation for Efficient CDC at Scale Vinay Chella, Akshat Goel