opendatahub-io/opendatahub-operator

Open Data Hub operator to manage ODH component integrations

View on GitHub ↗Jump to charts ↓Open shareable report

Summary Information

Updated 29 minutes ago
Added to GitGenius on January 17th, 2025
Created on February 19th, 2020
Open Issues & Pull Requests: 94 (+0)
Number of forks: 309
Total Stargazers: 107 (+0)
Total Subscribers: 13 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 0.9 hours
Mean response time: 14.0 days
90th percentile: 9.2 days
Tracked items: 23

How this project is maintained

Around half of the issues opened in the past year never receive a reply. Only 4% of issues opened in the past year have been closed. Three people close 60% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 10
New in 7 days: 0
Closed in 7 days: 1
Avg open age: 43 days
Stale 30+ days: 8
Stale 90+ days: 7

Recent activity

Opened in 7 days: 0
Closed in 7 days: 1
Comments in 7 days: 1
Events in 7 days: 8

Top labels

  • priority/high (1)

Detailed Description

The opendatahub-operator is a Kubernetes operator written in Go that serves as the primary control plane for Open Data Hub, a platform designed to manage and deploy data science applications and infrastructure components. The operator uses the DataScienceCluster custom resource definition to declaratively deploy and configure applications such as Jupyter Notebooks, data science pipelines, and other machine learning tools within Kubernetes clusters.

The operator functions as a comprehensive orchestration layer for a multi-tenant, cloud-native data science platform. It manages the lifecycle of numerous integrated components including KServe for model serving, Ray for distributed computing, Training Operator for ML job management, Feast for feature stores, Model Registry for model versioning, TrustyAI for AI explainability, the ODH Dashboard for web-based management, Workbenches for notebook environments, and AI Pipelines for ML workflow orchestration. These components are automatically integrated based on DataScienceCluster configuration rather than requiring separate installation, streamlining deployment complexity.

The platform supports deployment on OpenShift 4.19 or higher and can be installed directly from the community-operators catalog on OperatorHub, with the latest releases available through the Fast channel. The operator also supports multiple deployment methods including Cloud Manager for multi-cloud provisioning and RHAII Mode for specific provider integrations. Installation requires creating a DSCInitialization custom resource followed by a DataScienceCluster resource to enable desired components.

The operator supports extensive configuration through environment variables and flags, with comprehensive documentation covering prerequisites, platform requirements, namespace configuration, and resource allocation. Optional external operators can be installed for advanced functionality including certificate management via OpenShift Cert Manager, job queueing through Red Hat build of Kueue, distributed tracing with OpenTelemetry and Tempo operators, and enhanced observability through Cluster Observability and Perses operators. GPU support is available through NVIDIA GPU Operator and DCGM Exporter for GPU-accelerated workloads.

The repository includes detailed developer documentation covering local development setup, manifest customization, component addition procedures, and comprehensive testing infrastructure including functional tests, end-to-end tests, integration tests via Jenkins pipeline, and Prometheus unit tests for alerts. The documentation provides guidance on upgrade testing, release workflows, troubleshooting, and runtime logging level adjustments. The operator's extensible architecture allows customization of manifest sources for both local development and production operator image builds, supporting diverse deployment scenarios and organizational requirements.