Apache Beam

A unified model for defining both batch and streaming data-parallel processing pipelines. Provides a general approach to expressing embarrassingly parallel data processing pipelines and supports end users, SDK writers, and runner writers.

First releasedJune 15, 2016
Developed byApache Software Foundation
Open-sourceYes

Interesting facts

The Beam pipeline is executed by one of the distributed processing back-ends, which include Apache Apex, Apache Flink, Apache Spark, and Google Cloud Dataflow.

Sign up for updates
straight to your inbox