Speaker: Nancy Zheng
VP of Engineering of AI Inference @Cerebras
Nancy Zheng is the VP of Engineering of AI Inference at Cerebras. She's building the inference serving architecture for the fastest AI in the world. She worked on enabling OpenAI's GPT 5.6 Sol Ultrafast on Cerebras (15x faster!), and now she's building one of the world's first disaggregated inference clouds combining GPUs and Cerebras to get 5x throughput. Previously, Nancy was a founder and led teams at Google and Stripe.
Session
Inside the Disaggregated Architecture Serving LLM Inference
Every LLM request is really two very different jobs. Reading your prompt (prefill) is compute-intensive, while generating the response token by token (decode) is dominated by memory bandwidth. For years, the industry ran both phases on the same hardware and accepted the compromise.