Breaking Down the Walls to Python & Big Data + AI Performance with Transpilation and LLMs

QCon San Francisco 2026

Session

Breaking Down the Walls to Python & Big Data + AI Performance with Transpilation and LLMs

Tuesday Nov 17 / 10:35AM PST, Seacliff ABC at Hyatt Regency, San Francisco

Register

$2,835, Conference (3 days)
Current pricing ends September 8th

Abstract

Many of us have a love hate relationship with Apache Spark, especially when it comes to PySpark. Especially in Python, the UDF performance can be a dumpster fire. This fire comes from a glorious, toxic mix of data copies, serialization (pickle, json, etc.), and incompatibility with most accelerators (excluding NV Rapids).

This talk looks at the new work being done to allow us to bring the computation to the data rather than the data to the computation. The recently accepted Spark Project Improvement proposal adds basic, and safe, transpilation which can accelerate relatively simple Python UDFs by converting them into catalyst expressions; but safe only gets us so far. We'll look at how we can go beyond "safe" transpilation and use LLMs to dynamically convert large chunks of code for performance, how to pick the appropriate level of safety in doing so, and why this is all a little bit crazy and turned off by default (but the non-LLM transpilation should be safe. probably.)

Come for the performance, stay for the non-deterministic compilers with dynamic property testing.

Register

$2,835, Conference (3 days). Current pricing ends September 8th. All pass options.

76% senior dev or higher
1:11 speaker ratio
60+ practitioners

QCon San Francisco 2026 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.

Share

Current pricing ends September 8th
$2,835, Conference (3 days)

Register