huggingface/datasets

🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation tools

View on GitHub ↗Jump to charts ↓

Data as of . Signed-in members get hourly updates — create a free account.

Summary Information

Updated 8 minutes ago
Added to GitGenius on November 21st, 2023
Created on March 26th, 2020
Open Issues & Pull Requests: 1,457 (+0)
GitHub issues: Enabled
Number of forks: 3,505
Total Stargazers: 22,043 (+0)
Total Subscribers: 282 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 3.6 days
Mean response time: 185.1 days
90th percentile: 742.6 days
Tracked items: 825

Maintainer activity

7 people did triage or write work on this repository in the last 12 months.

Counts unlabeled, assigned, unassigned, milestoned, demilestoned, locked, unlocked over the last 12 months. These are issue and pull request events that require triage or write permission. Commits and code review are not counted. labeled and renamed are excluded because GitHub issue forms record the issue author as the actor. Figures from October 7, 2026. This count is not comparable across projects: each project's automation decides which of these events a person emits.

How this project is maintained

About 19% of issues opened in the past year have never received a reply. 90% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. Work labelled "documentation" is answered fastest, typically in about 7 hours, while "dataset request" waits about 35 months. 46% of tracked open issues have had no activity in three months. Only 43% of issues opened in the past year have been closed.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 906
New in 7 days: 6
Closed in 7 days: 8
Avg open age: 886 days
Stale 30+ days: 859
Stale 90+ days: 806

Recent activity

Opened in 7 days: 5
Closed in 7 days: 8
Comments in 7 days: 4
Events in 7 days: 5

Top labels

  • bug (614)
  • enhancement (475)
  • dataset request (136)
  • dataset-viewer (88)
  • dataset bug (64)
  • good first issue (52)
  • documentation (28)
  • duplicate (28)

Most active issues this week

Sign in to see which issues are moving.
Sign in

Detailed Description

The Hugging Face Datasets library is a lightweight Python package that serves as the largest hub of ready-to-use datasets for AI models, providing fast and efficient data manipulation tools. The library centers on two core features: one-line dataloaders for public datasets and efficient data preprocessing capabilities. Users can load datasets with simple commands like load_dataset("rajpurkar/squad"), instantly accessing datasets across 467 languages and dialects, including image datasets, audio datasets, text datasets, 3D medical images, video datasets, and agent traces. The preprocessing functionality allows users to prepare data for training and evaluation through simple commands like dataset.map(process_example).

The library supports an extensive range of file formats natively, including CSV, JSON, JSONL, Parquet, Arrow, XML, Text, Webdataset, and more. Multi-modal data support is built in, covering text, audio, image, video, PDF, and NIfTI 3D medical data types. A streaming mode enables users to iterate over data on-the-fly without downloading entire datasets, with performance improvements up to 100x faster when using the Xet backend. The Apache Arrow backend provides zero-copy memory-mapped storage, naturally freeing users from RAM limitations. Additional capabilities include smart caching that automatically reuses processed results, multi-framework interoperability with NumPy, Pandas, Polars, Arrow, PyTorch, TensorFlow, JAX, and Spark, and multi-processing support for fast parallel data processing.

The core architecture provides two main dataset classes: Dataset for in-memory or memory-mapped datasets backed by Apache Arrow with support for indexing and caching, and IterableDataset for lazy, streamable datasets suited for large-scale out-of-core processing. Both classes are wrapped in DatasetDict or IterableDatasetDict for handling multi-split datasets like train, test, and validation splits. Additional features include built-in FAISS and Elasticsearch index support for similarity search, flexible JSON type support for structured data, and direct read-write capabilities to Hugging Face Storage Buckets for mutable large-scale raw data.

The library is designed to enable community contributions, with detailed documentation for adding new datasets to the Hub. Users can upload datasets through web browsers, Python scripts, or Git-based workflows. The project maintains strict code standards using Ruff for linting and requires comprehensive testing and documentation for contributions.