The Revenge of the Data Scientist: Why Reliable AI Needs Evals, Traces, and Metrics

QCon San Francisco 2026

Session

The Revenge of the Data Scientist: Why Reliable AI Needs Evals, Traces, and Metrics

Tuesday Nov 17 / 11:45AM PST, Ballroom A at Hyatt Regency, San Francisco

Register

$2,835, Conference (3 days)
Current pricing ends September 8th

Abstract

Most teams can now ship an AI prototype by calling a foundation-model API. The hard part is knowing whether that system works when real users, messy data, and business consequences arrive. In this talk, I’ll argue that production AI is won or lost in the harness around the model: traces, metrics, labels, test sets, judges, and the discipline to inspect failures directly. Drawing from “The Revenge of the Data Scientist,” I’ll show five common eval pitfalls — generic metrics, unverified judges, weak experimental design, bad labels, and over-automation — and explain how engineering teams can avoid them. The practical takeaway is simple: reliable AI is not a model-only problem. It is an engineering system, and the missing muscle is often data science.

Main Takeaways:

  1. Production AI quality depends on a harness: tests, traces, metrics, labels, and experiments that tell you when the system is going off track.
  2. Generic eval dashboards and off-the-shelf metrics rarely diagnose real application failures; teams need error analysis and domain-specific metrics.
  3. LLM judges should be treated like classifiers: validated against human labels, tuned on development data, and reported with precision/recall rather than blind accuracy.
  4. The fastest path to better AI systems is still to look at the data: read traces, involve domain experts, and design experiments around real production behavior.
Register

$2,835, Conference (3 days). Current pricing ends September 8th. All pass options.

76% senior dev or higher
1:11 speaker ratio
60+ practitioners

QCon San Francisco 2026 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.

Share

From the same track

Tuesday 17 November

10:35 Ballroom A Session Progressive Failure Modes of Modern AI Serving Systems Abi Aryan AI Infrastructure Engineer and Educator 11:45 Ballroom A Session The Revenge of the Data Scientist: Why Reliable AI Needs Evals, Traces, and Metrics Hamel Husain Machine Learning Engineer, 20+ Years in Applied AI, Machine Learning, and Data Science 13:35 Ballroom A Session Skills, Memory, or Fine-Tuning? The Engineering Loop Behind Self-Improving Agents Abhinav Sinha CEO @Lucidic AI, Previously @Stanford AI Lab, @Citadel and Susquehanna International Group, and @Apple 14:45 Ballroom A Session Lessons from Building a $100M Product in Six Weeks at OpenAI Brian Yang Member of Technical Staff @OpenAI 15:55 Ballroom A Session Performance Engineering in the Age of AI 17:05 Seacliff D Unconference Unconference: Engineering AI Systems

Current pricing ends September 8th
$2,835, Conference (3 days)

Register