Nancy Zheng

Speaker: Nancy Zheng

VP of Engineering of AI Inference @Cerebras

Nancy Zheng is the VP of Engineering of AI Inference at Cerebras. She's building the inference serving architecture for the fastest AI in the world. She worked on enabling OpenAI's GPT 5.6 Sol Ultrafast on Cerebras (15x faster!), and now she's building one of the world's first disaggregated inference clouds combining GPUs and Cerebras to get 5x throughput. Previously, Nancy was a founder and led teams at Google and Stripe.

Session

Inside the Disaggregated Architecture Serving LLM Inference

Every LLM request is really two very different jobs. Reading your prompt (prefill) is compute-intensive, while generating the response token by token (decode) is dominated by memory bandwidth. For years, the industry ran both phases on the same hardware and accepted the compromise.

Read more

Date

Monday Nov 16 / 02:45PM PST ( 50 minutes )

Location

Ballroom BC

Share