Speaker
Abstract
Most teams can now ship an AI prototype by calling a foundation-model API. The hard part is knowing whether that system works when real users, messy data, and business consequences arrive. In this talk, I’ll argue that production AI is won or lost in the harness around the model: traces, metrics, labels, test sets, judges, and the discipline to inspect failures directly. Drawing from “The Revenge of the Data Scientist,” I’ll show five common eval pitfalls — generic metrics, unverified judges, weak experimental design, bad labels, and over-automation — and explain how engineering teams can avoid them. The practical takeaway is simple: reliable AI is not a model-only problem. It is an engineering system, and the missing muscle is often data science.
Main Takeaways:
- Production AI quality depends on a harness: tests, traces, metrics, labels, and experiments that tell you when the system is going off track.
- Generic eval dashboards and off-the-shelf metrics rarely diagnose real application failures; teams need error analysis and domain-specific metrics.
- LLM judges should be treated like classifiers: validated against human labels, tuned on development data, and reported with precision/recall rather than blind accuracy.
- The fastest path to better AI systems is still to look at the data: read traces, involve domain experts, and design experiments around real production behavior.
$2,835, Conference (3 days). Current pricing ends September 8th. All pass options.
QCon San Francisco 2026 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.
Part of the track
Engineering AI Systems Hosted by Melanie Zhao Engineering Lead @BlackRock, Pioneering AI Adoption in Asset ManagementFrom the same track
Tuesday 17 November
10:35 Ballroom A Session Progressive Failure Modes of Modern AI Serving Systems Abi Aryan AI Infrastructure Engineer and Educator Inference platforms fail in layers. Most organizations focus on model quality while underestimating the systems engineering required to operate production AI workloads safely and reliably at scale. 11:45 Ballroom A Session The Revenge of the Data Scientist: Why Reliable AI Needs Evals, Traces, and Metrics Hamel Husain Machine Learning Engineer, 20+ Years in Applied AI, Machine Learning, and Data Science Most teams can now ship an AI prototype by calling a foundation-model API. The hard part is knowing whether that system works when real users, messy data, and business consequences arrive. 13:35 Ballroom A Session Skills, Memory, or Fine-Tuning? The Engineering Loop Behind Self-Improving Agents Abhinav Sinha CEO @Lucidic AI, Previously @Stanford AI Lab, @Citadel and Susquehanna International Group, and @Apple As agents become mainstream, everyone wants to improve theirs either by making fewer mistakes on existing tasks or by taking on harder ones. This usually happens once an agent is already deployed in production. 14:45 Ballroom A Session Lessons from Building a $100M Product in Six Weeks at OpenAI Brian Yang Member of Technical Staff @OpenAI Engineering is shifting from synchronous, line-by-line implementation to an asynchronous model in which agents execute, verify, and retry while humans own intent, architecture, constraints, and release decisions. 15:55 Ballroom A Session Performance Engineering in the Age of AI Details coming soon. 17:05 Seacliff D Unconference Unconference: Engineering AI Systems