igorbarinov/awesome-data-engineering

A curated list of data engineering tools for software developers

View on GitHub ↗Jump to charts ↓

Summary Information

Updated 36 minutes ago
Added to GitGenius on September 7th, 2026
Created on June 18th, 2015
Open Issues & Pull Requests: 21 (+0)
GitHub issues: Enabled
Number of forks: 1,610
Total Stargazers: 9,029 (+0)
Total Subscribers: 255 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 15.7 days
Mean response time: 16.5 days
90th percentile: 35.0 days
Tracked items: 8

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 2
New in 7 days: 0
Closed in 7 days: 0
Avg open age: 122 days
Stale 30+ days: 1
Stale 90+ days: 1

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

No label distribution available yet.

Most active issues this week

No issue events were indexed in the last 7 days.

Detailed Description

Awesome Data Engineering is a curated list that organizes data engineering tools and resources for software developers.

The list addresses the challenge of discovering relevant tools across the fragmented data engineering landscape. It works by categorizing resources into logical sections spanning the full data pipeline: from databases and data ingestion through stream and batch processing, to monitoring and profiling. This organizational approach helps developers quickly locate tools suited to specific problems rather than searching through undifferentiated tool directories.

Developers building data platforms or pipelines should use this list as a reference when evaluating technology choices. It suits teams at any stage who need to understand what options exist for particular data engineering concerns. The list covers infrastructure components like file systems and serialization formats alongside higher-level concerns like workflow orchestration and data lake management, making it relevant whether you are designing a new system or filling gaps in an existing one.

The project maintains an organized, categorized structure with sections for databases, data ingestion, stream processing, batch processing, dashboards, workflow tools, monitoring, and profiling, alongside community resources including forums, conferences, podcasts, and books. The breadth of categories reflects active curation across the data engineering domain rather than focus on a narrow subset of tools.