argilla-io/distilabel

Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.

View on GitHub ↗Jump to charts ↓

Summary Information

Updated 49 minutes ago
Added to GitGenius on September 20th, 2026
Created on October 16th, 2023
Open Issues & Pull Requests: 105 (+0)
GitHub issues: Enabled
Number of forks: 255
Total Stargazers: 3,395 (+0)
Total Subscribers: 23 (+0)

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 47
New in 7 days: 0
Closed in 7 days: 0
Avg open age: 652 days
Stale 30+ days: 46
Stale 90+ days: 45

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

  • enhancement (45)
  • bug (28)
  • documentation (17)
  • good first issue (8)
  • help wanted (4)
  • improvement (4)
  • integrations (2)
  • idea (1)

Most active issues this week

No issue events were indexed in the last 7 days.

Detailed Description

Distilabel is a framework for generating synthetic data and collecting AI feedback through scalable, research-backed pipelines.

The tool addresses the challenge of creating high-quality training datasets efficiently by providing a programmatic approach to synthetic data generation and AI-based evaluation. It works by implementing methodologies from verified research papers, allowing engineers to build pipelines that synthesize data and gather feedback from any LLM provider through a unified API. The framework emphasizes data quality as a lever for improving both computational efficiency and model output quality, helping teams maintain high standards throughout their datasets rather than relying solely on compute scaling.

Distilabel suits engineers building traditional NLP systems (classification, extraction) or generative AI applications (instruction following, dialogue generation, model judging) who want to own their fine-tuning data and iterate quickly on research-backed approaches. The tool's unified API for integrating feedback from multiple LLM providers gives teams flexibility in choosing their models while maintaining a consistent pipeline architecture. It is designed for projects where dataset quality directly impacts downstream model performance and where teams need fault tolerance and scalability in their data generation workflows.

The project is maintained by community collaborators who have recently joined to continue development after the original authors moved to other work. Active improvements and fixes are being developed on the develop branch ahead of upcoming releases.