apache/beam

Apache Beam is a unified programming model for Batch and Streaming data processing.

View on GitHub ↗Jump to charts ↓Open shareable report →

Data as of . Signed-in members get hourly updates — create a free account.

Summary Information

Updated 21 minutes ago
Added to GitGenius on January 3rd, 2025
Created on February 2nd, 2016
Open Issues & Pull Requests: 3,872 (+0)
GitHub issues: Enabled
Number of forks: 4,669
Total Stargazers: 8,679 (+0)
Total Subscribers: 259 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 3.1 days
Mean response time: 257.5 days
90th percentile: 1202.0 days
Tracked items: 2,645

Maintainer activity

147 people did triage or write work on this repository in the last 12 months.

At least 7% of beam's 147 maintainers work at Google. 42 say where they work, and 10 of those are Google.

Counts unlabeled, assigned, unassigned, milestoned, demilestoned, locked, unlocked over the last 12 months. These are issue and pull request events that require triage or write permission. Commits and code review are not counted. labeled and renamed are excluded because GitHub issue forms record the issue author as the actor. Figures from October 7, 2026. This count is not comparable across projects: each project's automation decides which of these events a person emits.

How this project is maintained

About 17% of issues opened in the past year have never received a reply. 61% of open issues come from outside the core team, a mix of external reports and the maintainers' own roadmap. Work labelled "P2" is answered fastest, typically in about 14 hours, while "P3" waits about 35 months. 52% of tracked open issues have had no activity in three months. 69% of issues opened in the past year have been closed, leaving a working backlog.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 3,438
New in 7 days: 14
Closed in 7 days: 18
Avg open age: 1,267 days
Stale 30+ days: 3,339
Stale 90+ days: 3,247

Recent activity

Opened in 7 days: 14
Closed in 7 days: 15
Comments in 7 days: 13
Events in 7 days: 54

Top labels

  • P3 (3,788)
  • bug (3,331)
  • P2 (2,604)
  • java (2,093)
  • python (1,801)
  • io (1,460)
  • new feature (1,290)
  • awaiting triage (1,242)

Most active issues this week

Sign in to see which issues are moving.
Sign in

Detailed Description

Apache Beam is a unified programming model for defining and executing both batch and streaming data-parallel processing pipelines. The project provides language-specific SDKs for constructing pipelines and multiple runners for executing them on distributed processing backends including Apache Flink, Apache Spark, Google Cloud Dataflow, and Hazelcast Jet. The repository is written primarily in Java and is maintained under the Apache Software Foundation.

The core of Beam's design philosophy centers on three key abstractions: PCollection represents a collection of data that can be bounded or unbounded in size, PTransform represents a computation that transforms input PCollections into output PCollections, and Pipeline manages a directed acyclic graph of PTransforms and PCollections ready for execution. This model evolved from several internal Google data processing projects including MapReduce, FlumeJava, and Millwheel, and was originally known as the Dataflow Model.

Beam supports multiple language-specific SDKs for writing pipelines against its unified model. The repository currently contains SDKs for Java, Python, and Go, with provisions for community contributions of additional SDKs and domain-specific languages. The project also provides multiple execution runners that allow the same pipeline code to execute on different distributed processing backends. Available runners include the DirectRunner for local machine execution, the DataflowRunner for Google Cloud Dataflow, the FlinkRunner for Apache Flink clusters, the SparkRunner for Apache Spark clusters, the JetRunner for Hazelcast Jet clusters, and the Twister2Runner for Twister2 clusters.

The repository demonstrates significant community engagement and maintenance activity.

Beam is classified across a comprehensive range of data processing domains including streaming analytics, ETL workflows, cloud-native platforms, big data processing, batch processing, pipeline frameworks, distributed computing, real-time analytics, parallel processing, data pipelines, event-driven applications, stream processing, real-time computing, and machine learning pipelines. The project supports embarrassingly parallel data processing and is designed to serve three distinct user categories: end users writing pipelines with existing SDKs, SDK writers developing language-specific implementations, and runner writers implementing execution environments for distributed processing.