Stream All the Things — Patterns of Effective Data Stream Processing

QCon San Francisco 2024

Session

Stream All the Things — Patterns of Effective Data Stream Processing

Tuesday Nov 19 / 01:35PM PST, Ballroom A

Abstract

Data streaming is a really difficult problem. Despite 10+ years of attempting to simplify it, teams building real-time data pipelines can spend up to 80% of their time optimizing it or fixing downstream output by handling bad data at the lake. All we want is a service that will be reliable, handle all kinds of data, connect with all kinds of systems, be easy to manage, and scale up and down as our systems change.

Oh, it should also have super low latency and result in good data. Is it too much to ask?

In this presentation, we’ll discuss the basic challenges of data streaming and introduce a few design and architecture patterns, such as DLQ, used to tackle these challenges.

We will then explore how to implement these patterns using Apache Flink and discuss the challenges that real-time AI applications bring to our infra. Difficult problems are difficult, and we offer no silver bullets. Still, we will share pragmatic solutions that have helped many organizations build fast, scalable, and manageable data streaming pipelines.

76% senior dev or higher
1:11 speaker ratio
60+ practitioners

QCon San Francisco 2024 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.

Share

From the same track

Tuesday 19 November

10:35 Ballroom A Session Platform Engineering Beyond Durability: Enhancing Database Resilience and Reducing the Entropy Using Write-Ahead Logging at Netflix Prudhviraj Karumanchi, Vidhya Arvind 11:45 Pacific DEKJ Session Architecture OpenSearch Cluster Topologies for Cost-Saving Autoscaling Amitai Stern Engineering Manager @Logz.io, Managing Observability Data Storage of Petabyte Scale, OpenSearch Leadership Committee Member and Contributor 13:35 Ballroom A Session Stream All the Things — Patterns of Effective Data Stream Processing Adi Polak Director, Advocacy and Developer Experience Engineering @Confluent, Author of "Scaling Machine Learning with Spark" and "High Performance Spark 2nd Edition" 14:45 Ballroom A Session Stream and Batch Processing Convergence in Apache Flink Jiangjie (Becket) Qin Principal Staff Software Engineer @LinkedIn, Data Infra Engineer, PMC Member of Apache Kafka & Apache Flink, Previously @Alibaba and @IBM 15:55 Ballroom A Session Data Pipelines Efficient Incremental Processing with Netflix Maestro and Apache Iceberg Jun He Staff Software Engineer @Netflix, Managing and Automating Large-Scale Data/ML Workflows, Previously @Airbnb and @Hulu 17:05 Seacliff D Unconference Unconference: Shift-Left Data Architecture