Speaker
Abstract
Engineering is shifting from synchronous, line-by-line implementation to an asynchronous model in which agents execute, verify, and retry while humans own intent, architecture, constraints, and release decisions.
Drawing on my experience building OpenAI Ads with Codex from day one and scaling it to more than $100 million in ARR in under six weeks, I’ll share an end-to-end operating model for building a production product with coding agents while avoiding common failure modes. I’ll share what worked, what did not, and the lessons we learned about using coding agents to build quickly without compromising quality or engineering judgment.
To make the workflow concrete, I’ll use a sanitized case study of reducing latency on a complex production request path without compromising correctness or reliability. Performance work creates a particularly useful agentic workflow because progress is measurable. We can establish a baseline, investigate bottlenecks, implement candidate improvements, run benchmarks and regression tests, and iterate based on evidence until the improvement holds.
Although the examples use Codex, the operating model is not Codex-specific. The same principles apply to other agentic coding tools and harnesses that can work with a codebase, use engineering tools, and iterate against tests, benchmarks, and other forms of feedback.
I’ll also separate tokenmaxxing from token economics. The goal is not to minimize tokens or maximize agent activity. It is to spend more inference only when verification shows that it improves accepted production throughput, cycle time, or reliability. Attendees will leave with a practical operating model for moving faster with coding agents without surrendering engineering judgment or accountability.
Main Takeaways:
- Use measurable feedback loops for complex production work. Follow a sanitized latency improvement from baseline and investigation through implementation, benchmarking, regression testing, and rollout.
- Build a repeatable agentic engineering workflow, not a tool-specific trick. See how the practices demonstrated with Codex can be applied with other agentic coding tools and harnesses.
- Practice token economics, not indiscriminate tokenmaxxing. Spend inference where evidence shows it improves cycle time, reliability, and accepted production output.
Topics
$2,955, Conference (3 days). Current pricing ends October 13th. All pass options.
QCon San Francisco 2026 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.
Part of the track
Engineering AI Systems Hosted by Melanie Zhao Engineering Lead @BlackRock, Pioneering AI Adoption in Asset ManagementFrom the same track
Tuesday 17 November
10:35 Ballroom A Session Inference Progressive Failure Modes of Modern AI Serving Systems Abi Aryan AI Infrastructure Engineer and Educator Inference platforms fail in layers. Most organizations focus on model quality while underestimating the systems engineering required to operate production AI workloads safely and reliably at scale. 11:45 Ballroom A Session AI The Revenge of the Data Scientist: Why Reliable AI Needs Evals, Traces, and Metrics Hamel Husain Machine Learning Engineer, 20+ Years in Applied AI, Machine Learning, and Data Science Most teams can now ship an AI prototype by calling a foundation-model API. The hard part is knowing whether that system works when real users, messy data, and business consequences arrive. 13:35 Ballroom A Session AI/ML Skills, Memory, or Fine-Tuning? The Engineering Loop Behind Self-Improving Agents Abhinav Sinha CEO @Lucidic AI, Previously @Stanford AI Lab, @Citadel and Susquehanna International Group, and @Apple As agents become mainstream, everyone wants to improve theirs either by making fewer mistakes on existing tasks or by taking on harder ones. This usually happens once an agent is already deployed in production. 14:45 Ballroom A Session AI/ML Lessons from Building a $100M Product in Six Weeks at OpenAI Brian Yang Member of Technical Staff @OpenAI Engineering is shifting from synchronous, line-by-line implementation to an asynchronous model in which agents execute, verify, and retry while humans own intent, architecture, constraints, and release decisions. 15:55 Ballroom A Session Inside AIROS: BlackRock’s Platform Architecture for Applied AI at Scale Dan Wolf Head of Platform Engineering for PMG Tech @BlackRock BlackRock’s Portfolio Management Group is transforming how investment teams conduct research by combining deep investment expertise with shared AI capabilities. AIROS, the AI Research Operating System, is the platform architecture behind that transformation. 17:05 Seacliff D Unconference Unconference: Engineering AI Systems