krzjoa/awesome-python-data-science

Probably the best curated list of data science software in Python.

View on GitHub ↗Jump to charts ↓

Summary Information

Updated 45 minutes ago
Type:Curated List / Learning ResourceCategory(s):General Awesome Lists & DirectoriesLearning & Resources
Added to GitGenius on September 19th, 2026
Created on December 21st, 2017
Open Issues & Pull Requests: 21 (+0)
GitHub issues: Enabled
Number of forks: 463
Total Stargazers: 3,595 (+0)
Total Subscribers: 61 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 0.0 hours
Mean response time: 119.7 days
90th percentile: 239.4 days
Tracked items: 2

Most active contributors

Sign in to see contributor activity.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 1
New in 7 days: 0
Closed in 7 days: 0
Avg open age: 147 days
Stale 30+ days: 1
Stale 90+ days: 1

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

No label distribution available yet.

Most active issues this week

No issue events were indexed in the last 7 days.

Detailed Description

Awesome Python Data Science is a curated list of data science software in Python organized by domain and capability.

The list addresses the challenge of discovering relevant tools across the fragmented Python data science ecosystem. It organizes software into categories spanning machine learning, deep learning, natural language processing, computer vision, time series analysis, reinforcement learning, graph machine learning, model explanation, optimization, feature engineering, visualization, data manipulation, and supporting infrastructure like deployment, statistics, distributed computing, and data validation. This categorical structure allows practitioners to find established tools within their specific problem domain rather than searching through undifferentiated package indexes.

The list suits anyone building data science projects in Python who wants to understand what established solutions exist before implementing custom code. It works well for teams evaluating tool choices across different stages of a pipeline, from data preparation through model deployment. The breadth of categories means it covers both specialized domains like quantum computing and spatial analysis alongside foundational areas like data frames and pipelines. Developers new to Python data science benefit from seeing what the community considers standard practice in each area.

The project maintains an organized, manually curated collection that reflects deliberate editorial choices about which tools merit inclusion. The structure remains stable across updates, with categories and subcategories preserved to support navigation. Contributions are accepted through a defined process, indicating ongoing community engagement with the curation effort.