Speaker
Abstract
What if everything you know about building distributed systems is backwards? What if instead of putting databases at the bottom of your architecture stack, you put them at the center—not just storing data, but actually orchestrating your application logic, managing your workflows and their state, and providing reliability guarantees? A system that treats PostgreSQL not just as your database, but also as a durability layer for your application runtime. Instead of the typical pattern of building orchestration layers on top of databases, we compile workflow logic directly into database operations.
Drawing from our experience scaling systems at Reddit and Netflix and years of research at Stanford and MIT, this talk will discuss various use cases showing how using the open source Transact library from DBOS and durable computing concepts creates more reliable systems while reducing cost, processing time, and operational complexity.
Interview
Our session is about building reliable software directly on top of the database you already use, without relying on an external workflow coordinator. This is important because introducing extra services often adds new points of failure instead of improving reliability. By leveraging the database as the foundation for durability and recovery, applications become simpler and more reliable, which should always be important for senior developers, as important as security.
Durable computing and reliability are especially critical as we move into 2026 because AI agents are starting to take on more critical tasks on behalf of users. Today, these agents are still fragile and prone to failure. But as they gain the ability to make more real world interactions, the cost of unreliability grows dramatically. Building reliable agents will be one of the most important responsibilities for software leaders in 2026.
Developers and architects face recurring challenges in this space. Modern applications, especially AI agents, fail in unpredictable ways, and most systems don't provide enough observability to diagnose them. Existing workflow engines promise reliability but introduce heavy infrastructure, steep learning curves, and operational overhead. They often create more fragility than they remove. With the rise of AI coding agents, expecting them to design full distributed systems is unrealistic. Instead, using a library-based approach to durability allows AI agents to create durable software with a high chance of success.
We would hope attendees leave with the mindset that reliability should be built in from the start, not bolted on later. The simplest next step is to try out a lightweight library like DBOS Transact in a small part of their system and see how durable execution changes the way they think about reliability.
QCon attracts people who are actually building and operating systems at scale, not just talking about them. You’re sitting next to engineers solving the same problems you’ve dealt with, people who understand the difference between what sounds good in a blog post and what actually works. The conference has this unique ability to spot architectural shifts before they become mainstream, and the speakers aren't just evangelizing – they're sharing real war stories, including the parts that didn't work. Plus, the hallway conversations with practitioners who've been through the same scaling challenges, late-night outages, and architectural regrets are often more valuable than the talks themselves.
(Jeremy here) My favorite QCon story of all time: I was a keynote speaker at QCon Sao Paulo, along with Neal Ford. Neal was doing a keynote on how to give a good technical presentation, and he went first. He then proceeded to give a list of do’s and don’t’s, and I realized that I was doing every don’t and didn’t have any of the do’s. I then had to speak after him! He graciously gave me a copy of his book on the topic after the conference, which I read cover to cover and then completely changed how I give presentations.
The fun epilogue to this story is that five years later we were both keynote speakers back to back at another conference. We both had different topics, but I asked him if he could rate my talk. He gave me perfect marks!
Topics
QCon San Francisco 2025 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.
Part of the track
Architectures You've Always Wondered About Hosted by Khawaja Shams Co-Founder & CEO @Momento, previously @NASA and @AmazonFrom the same track
Monday 17 November
10:35 Ballroom A Session Architecture How to Build an Exchange Frank Yu Director of Engineering @Coinbase, Previously Principal Engineer and Director @FairX These days it is possible to achieve fairly good performance on cloud provisioned systems. We discuss the design of a high performance, strongly consistent system which maintains constant service in the face of regular updates to core logic. 11:45 Ballroom A Session Durability Compiling Workflows into Databases: The Architecture That Shouldn't Work (But Does) Jeremy Edberg, Qian Li What if everything you know about building distributed systems is backwards? 13:35 Pacific DEKJ Session Architecture Parting the Clouds: The Rise of Disaggregated Systems Murat Demirbas Principal Research Scientist @MongoDB Research, Previously Principal Applied Scientist @AWS and a Professor of Computer Science at the University at Buffalo (SUNY) Cloud systems are undergoing an architectural shift. Traditional shared-nothing designs struggle to deliver the elasticity, availability, and operational simplicity that the cloud demands. 14:45 Ballroom A Session Platform Engineering Building Resilient Platforms: Insights from 20+ Years in Mission-Critical Infrastructure Matthew Liste Head of Infrastructure @American Express, Previously @JPMorgan Chase and @Goldman Sachs In this talk, Matthew will describe lessons learned from over 20+ years of building scalable, secure and stable infrastructure platforms for software in financial services (electronic trading, credit card processing etc.), the talk is relevant to anyone building platforms for… 15:55 Ballroom A Session Architecture Architecting a Centralized Platform for Data Deletion at Netflix Vidhya Arvind, Shawn Liu What does it take to safely delete data at Netflix scale? In large-scale systems, data deletion cuts across infrastructure, reliability, and performance complexities. 17:05 Seacliff D Unconference Unconference: Architectures You've Always Wondered About