scikit-learn/scikit-learn

scikit-learn: machine learning in Python

View on GitHub ↗Jump to charts ↓Open shareable report →

Data as of . Signed-in members get hourly updates — create a free account.

Summary Information

Updated 4 minutes ago
Added to GitGenius on April 24th, 2023
Created on August 17th, 2010
Open Issues & Pull Requests: 2,160 (+0)
GitHub issues: Enabled
Number of forks: 27,487
Total Stargazers: 67,499 (+1)
Total Subscribers: 2,130 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 29.6 hours
Mean response time: 488.3 days
90th percentile: 2018.6 days
Tracked items: 2,340

Maintainer activity

23 people did triage or write work on this repository in the last 12 months.

Counts unlabeled, assigned, unassigned, milestoned, demilestoned, locked, unlocked over the last 12 months. These are issue and pull request events that require triage or write permission. Commits and code review are not counted. labeled and renamed are excluded because GitHub issue forms record the issue author as the actor. Figures from October 7, 2026. This count is not comparable across projects: each project's automation decides which of these events a person emits.

How this project is maintained

Practically every issue opened in the past year has drawn a reply. 66% of open issues come from outside the core team, a mix of external reports and the maintainers' own roadmap. Work labelled "Build / CI" is answered fastest, typically in about 10 hours, while "Enhancement" waits about 37 months. 41% of tracked open issues have had no activity in three months. 66% of issues opened in the past year have been closed, leaving a working backlog.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Issue analytics are temporarily unavailable.

Detailed Description

scikit-learn is a Python machine learning library built on top of SciPy and distributed under the 3-Clause BSD license. The project was initiated in 2007 by David Cournapeau as a Google Summer of Code project and has since grown through contributions from numerous volunteers. It is currently maintained by a volunteer team and provides a comprehensive toolkit for machine learning tasks in Python.

The library encompasses a broad range of machine learning functionality across multiple domains. This breadth reflects scikit-learn's position as a general-purpose machine learning framework rather than a specialized tool for any single technique or problem domain.

The project maintains active development with significant community engagement.

The repository's technical infrastructure reflects production-grade standards. The codebase uses GitHub Actions for unit testing, CircleCI for continuous integration, and Codecov for coverage tracking. Code quality is enforced through Ruff for style checking, and nightly wheel builds ensure compatibility across platforms. The project maintains a benchmark suite tracked through asv, allowing performance monitoring across releases.

Dependencies are carefully managed with specified minimum versions. The project requires Python 3.11 or later, NumPy 1.24.1 or higher, and SciPy 1.10.0 or higher. Additional dependencies include Joblib 1.4.0, Narwhals 2.0.1, Threadpoolctl 3.5.0, Matplotlib 3.6.1, scikit-image 0.22.0, Pandas 1.5.0, Seaborn 0.13.0, Pytest 7.1.2, and Plotly 5.22.0, reflecting the library's integration with the broader Python scientific computing ecosystem.

The repository shares contributors with other major Python projects including Microsoft's VSCode, pandas-dev/pandas, and matplotlib/matplotlib, indicating deep integration within the Python data science community. Installation is straightforward through pip or conda, with comprehensive documentation available both for stable releases and development versions. The project actively welcomes new contributors of all experience levels and provides detailed development guides covering code contribution, documentation, and testing procedures. Testing can be controlled through the SKLEARN_SEED environment variable for reproducibility.