kubeflow/spark-operator

Kubernetes operator for managing the lifecycle of Apache Spark applications on Kubernetes.

View on GitHub ↗Jump to charts ↓

Data as of . Signed-in members get hourly updates — create a free account.

Summary Information

Updated 36 minutes ago
Added to GitGenius on September 22nd, 2026
Created on January 3rd, 2018
Open Issues & Pull Requests: 159 (+0)
GitHub issues: Enabled
Number of forks: 1,536
Total Stargazers: 3,153 (+0)
Total Subscribers: 71 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 3.0 days
Mean response time: 52.4 days
90th percentile: 110.6 days
Tracked items: 571

Maintainer activity

32 people did triage or write work on this repository in the last 12 months.

Counts unlabeled, assigned, unassigned, milestoned, demilestoned, locked, unlocked over the last 12 months. These are issue and pull request events that require triage or write permission. Commits and code review are not counted. labeled and renamed are excluded because GitHub issue forms record the issue author as the actor. Figures from October 7, 2026. This count is not comparable across projects: each project's automation decides which of these events a person emits.

How this project is maintained

About 17% of issues opened in the past year have never received a reply. 68% of open issues come from outside the core team, a mix of external reports and the maintainers' own roadmap. Almost all tracked open issues have seen activity in the last three months. Only 59% of issues opened in the past year have been closed. Three people close 51% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 80
New in 7 days: 2
Closed in 7 days: 6
Avg open age: 383 days
Stale 30+ days: 44
Stale 90+ days: 25

Recent activity

Opened in 7 days: 1
Closed in 7 days: 5
Comments in 7 days: 4
Events in 7 days: 14

Top labels

  • lifecycle/stale (478)
  • kind/bug (95)
  • kind/feature (74)
  • enhancement (45)
  • lifecycle/frozen (33)
  • question (26)
  • help wanted (12)
  • good first issue (11)

Most active issues this week

Sign in to see which issues are moving.
Sign in

Detailed Description

Spark Operator is a Kubernetes operator that manages the lifecycle of Apache Spark applications on Kubernetes.

The operator solves the problem of running Spark jobs on Kubernetes by providing a custom resource definition that allows users to define and submit Spark applications declaratively. Rather than managing Spark clusters manually or using ad-hoc submission scripts, the operator handles job scheduling, resource allocation, and lifecycle management through Kubernetes-native mechanisms. This approach integrates Spark workloads into the Kubernetes ecosystem, enabling teams to manage data processing jobs using the same tools and patterns they use for other containerized applications.

Teams should adopt this operator if they run Spark workloads on Kubernetes and want to avoid managing separate Spark cluster infrastructure. It suits organizations that have already standardized on Kubernetes for container orchestration and want to extend that platform to batch data processing. The operator is particularly valuable for teams running on Google Cloud Dataproc or other Kubernetes environments where native integration with the control plane is beneficial. It works well for both one-off jobs and recurring batch processing pipelines that need to coexist with other Kubernetes workloads.

The project shows consistent development activity with regular updates to maintain compatibility with evolving Kubernetes versions and Spark releases. The codebase demonstrates active maintenance through ongoing bug fixes and feature enhancements that respond to user needs. The project maintains comprehensive documentation and examples that help new users understand how to deploy and configure the operator. Community contributions are regularly integrated, indicating an engaged user base providing feedback and improvements. The development approach prioritizes stability and backward compatibility while gradually introducing new capabilities for managing Spark applications at scale.