Speaker
Abstract
The LinkedIn Feed retrieval layer selects, within 120 ms, the handful of posts each member sees from hundreds of millions of candidates. This split-second decision shapes what over a billion professionals read and engage with every day. Ideally we would run the full ranking model over the entire corpus, but at the scale of modern recommender systems, that is infeasible, which is why retrieval exists as a distinct, lightweight stage that narrows the corpus before ranking.
At LinkedIn we built a modern GPU Retrieval-as-Ranking engine: instead of a lightweight approximate step, it scores candidates with a full deep learning recommendation model (DLRM) directly at the retrieval stage. Running on GPUs let us scale to O(10⁷) parameters and helped achieve a 2.5% increase in Content Time Spent in online A/B tests.
For Suggested content, where the candidate pool is every post on LinkedIn, a member's affinity to the post's topic was the strongest signal, and LLM embeddings that capture post semantics were a natural fit. However, for in-network content (posts generated by a member's connections), LLMs could not effectively capture the affinity between the viewer and the content creator, an important signal for this use case. Combining the semantic understanding of LLM embeddings with explicit edge features inside a DLRM was key to the success of in-network retrieval.
We will introduce our GPU Retrieval-as-Ranking inference stack and how we scaled it to an index of over one billion documents through a sharded setup. We will dive deep into the design choices behind the serving infrastructure, how edge features are joined to documents at runtime, and the optimizations that let us scale large models in the retrieval stage to production.
Main Takeaways:
- How a single retrieval model, a DLRM that fuses LLM embeddings (post semantics) with graph edge features (viewer-to-author affinity), captures signals that complement each other at the retrieval layer.
- Building efficient GPU-powered retrievers capable of running inference on a large volume of documents with DLRMs at strict latency budgets.
- Keeping the GPU-resident document index fresh and scaling to over 1 billion documents with a maximum staleness target of a couple of minutes.
- A reusable serving pattern that generalizes beyond LinkedIn scale: when a lightweight retrieval step is bottlenecked by a signal it cannot represent, moving a full model onto GPUs at the retrieval stage can outperform adding another ranking layer.
$2,955, Conference (3 days). Current pricing ends October 13th. All pass options.
QCon San Francisco 2026 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.
Part of the track
Engineering AI Systems Hosted by Melanie Zhao Engineering Lead @BlackRock, Pioneering AI Adoption in Asset ManagementFrom the same track
Tuesday 17 November
10:35 Ballroom A Session AI/ML Skills, Memory, or Fine-Tuning? The Engineering Loop Behind Self-Improving Agents Abhinav Sinha CEO @Lucidic AI, Previously @Stanford AI Lab, @Citadel and Susquehanna International Group, and @Apple As agents become mainstream, everyone wants to improve theirs either by making fewer mistakes on existing tasks or by taking on harder ones. This usually happens once an agent is already deployed in production. 11:45 Ballroom A Session Inside Asimov: BlackRock’s Platform Architecture for Applied AI at Scale Dan Wolf, Sri Annapurna Bhupatiraju BlackRock’s Portfolio Management Group is transforming how investment teams conduct research by combining deep investment expertise with shared AI capabilities. 13:35 Ballroom A Session AI The Revenge of the Data Scientist: Why Reliable AI Needs Evals, Traces, and Metrics Hamel Husain Machine Learning Engineer, 20+ Years in Applied AI, Machine Learning, and Data Science Most teams can now ship an AI prototype by calling a foundation-model API. The hard part is knowing whether that system works when real users, messy data, and business consequences arrive. 14:45 Ballroom A Session AI/ML Lessons from Building a $100M Product in Six Weeks at OpenAI Brian Yang Member of Technical Staff @OpenAI Engineering is shifting from synchronous, line-by-line implementation to an asynchronous model in which agents execute, verify, and retry while humans own intent, architecture, constraints, and release decisions. 15:55 Ballroom A Session Bridging Retrieval and Ranking: From LLM Embeddings and Graph Features to a GPU-Served DLRM Sudarshan Srinivasa Ramanujam Lead Engineer @LinkedIn - Working on Feed Ranking & Retrieval Modeling The LinkedIn Feed retrieval layer selects, within 120 ms, the handful of posts each member sees from hundreds of millions of candidates. This split-second decision shapes what over a billion professionals read and engage with every day. 17:05 Ballroom A Session Failing in Layers: Lessons from Rebuilding Netflix's ML Serving Platform Rajat Shah, Wenbing Bai Everyone talks about model quality. Fewer people talk about what happens when that model has to serve real traffic without falling over, or why the system ended up shaped the way it did in the first place.