kubeflow/katib

Automated Machine Learning on Kubernetes

View on GitHub ↗Jump to charts ↓Open shareable report

Summary Information

Updated 51 minutes ago
Added to GitGenius on June 18th, 2024
Created on April 3rd, 2018
Open Issues & Pull Requests: 106 (+0)
Number of forks: 536
Total Stargazers: 1,697 (+0)
Total Subscribers: 52 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 25.7 hours
Mean response time: 154.2 days
90th percentile: 567.5 days
Tracked items: 117

How this project is maintained

Around half of the issues opened in the past year never receive a reply. 37% of open issues come from outside the core team, a mix of external reports and the maintainers' own roadmap. Work labelled "area/sdk" is answered fastest, typically in about 5 hours, while "kind/feature" waits about 33 hours. Only 3% of issues opened in the past year have been closed. Three people close 87% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 73
New in 7 days: 0
Closed in 7 days: 0
Avg open age: 1,798 days
Stale 30+ days: 72
Stale 90+ days: 69

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

  • kind/feature (334)
  • kind/bug (319)
  • lifecycle/stale (163)
  • lifecycle/frozen (111)
  • area/katib (105)
  • priority/p1 (86)
  • help wanted (75)
  • area/front-end (60)

Most active issues this week

No issue events were indexed in the last 7 days.

Detailed Description

Kubeflow Katib is a Kubernetes-native automated machine learning platform that enables hyperparameter tuning, early stopping, and neural architecture search for machine learning workloads. Written primarily in Python, the project provides a framework-agnostic approach to AutoML that can optimize applications written in any programming language while natively supporting major ML frameworks including TensorFlow, PyTorch, XGBoost, JAX, and scikit-learn. The platform integrates seamlessly with Kubernetes Custom Resources and provides out-of-the-box support for Kubeflow Training Operator, Argo Workflows, and Tekton Pipelines, making it suitable for distributed training environments.

The project implements a comprehensive suite of search algorithms for different optimization tasks. For hyperparameter tuning, Katib supports Random Search, Grid Search, Bayesian Optimization, Tree of Parzen Estimators (TPE), Multivariate TPE, CMA-ES, Sobol's Quasirandom Sequence, HyperBand, and Population Based Training. For neural architecture search, it provides ENAS and DARTS algorithms. Early stopping capabilities include the Median Stopping Rule. These algorithms are powered by established optimization frameworks including Goptuna, Hyperopt, Optuna, and Scikit Optimize, allowing users to leverage well-tested optimization methodologies.

Katib is designed with accessibility in mind, offering both a control plane for Kubernetes deployments and a Python SDK available through PyPI as kubeflow-katib. This dual approach enables both infrastructure administrators to deploy Katib clusters and data scientists to programmatically create hyperparameter tuning experiments without deep Kubernetes expertise. The project includes comprehensive examples and getting started guides to facilitate adoption.

The repository shows active development and community engagement.

The codebase is classified across multiple optimization and machine learning domains, reflecting its role as a comprehensive experimentation framework for AI model optimization and machine learning workflows. Katib supports reproducible research through its experiment framework capabilities and enables benchmarking of different optimization strategies. The project's integration with the broader Kubeflow ecosystem positions it as a central component for MLOps workflows on Kubernetes, addressing the need for scalable, cloud-native hyperparameter optimization in production environments. The project maintains active community engagement through bi-weekly AutoML and Training Working Group meetings and a dedicated Slack channel, with documented adopters and presentations demonstrating real-world usage.