awesomedata/awesome-public-datasets

A topic-centric list of HQ open datasets.

View on GitHub ↗Jump to charts ↓

Summary Information

Updated 29 minutes ago
Added to GitGenius on December 30th, 2024
Created on November 20th, 2014
Open Issues & Pull Requests: 159 (+0)
Number of forks: 11,805
Total Stargazers: 78,587 (+1)
Total Subscribers: 2,312 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 11.7 days
Mean response time: 746.3 days
90th percentile: 3519.2 days
Tracked items: 5

How this project is maintained

Around half of the issues opened in the past year never receive a reply. 99% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. Only 0% of issues opened in the past year have been closed.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 84
New in 7 days: 0
Closed in 7 days: 0
Avg open age: 2,004 days
Stale 30+ days: 83
Stale 90+ days: 80

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

No label distribution available yet.

Most active issues this week

No issue events were indexed in the last 7 days.

Detailed Description

The Awesome Public Datasets repository is a curated, topic-centric list of high-quality open data sources maintained at awesomedata/awesome-public-datasets on GitHub. The project was originally incubated at OMNILab at Shanghai Jiao Tong University during Xiaming Chen's Ph.D. studies and is now connected to the BaiYuLan Open AI community. The repository serves as a comprehensive directory of freely available datasets collected and organized from blogs, answers, and user responses, though it notes that while most listed datasets are free, some are not.

A distinctive feature of this repository is its automated generation process. The repository is automatically generated by apd-core, meaning direct modifications to the main file are discouraged. Instead, contributors are directed to a dedicated contribution workflow through the apd-core repository. This architectural choice reflects a commitment to maintaining data quality and consistency across the curated list. The project maintains an active Slack community at awesomedataworld.slack.com where users can receive instant updates about high-quality data resources and engage with other data enthusiasts.

The repository's scope is remarkably broad, spanning numerous domains and disciplines. The README excerpt reveals datasets organized by topic including Agriculture, Architecture, and Biology, with entries covering everything from historical crop yields and soil moisture measurements to genomic data and microscopy images. Each dataset entry includes status indicators showing whether the resource is functioning properly or needs attention, along with metadata links for additional information.

The project is classified across 24 distinct categories including data science, machine learning data, research data, educational resources, and statistical analysis, reflecting its multifaceted utility for researchers, data scientists, and educators. The community-driven nature of the project, combined with its systematic organization and active maintenance, positions it as a valuable resource for anyone seeking high-quality public datasets across diverse fields of study and application.