apache/datafusion

Apache DataFusion SQL Query Engine

View on GitHub ↗Jump to charts ↓Open shareable report

Summary Information

Updated 25 minutes ago
Added to GitGenius on November 11th, 2024
Created on April 17th, 2021
Open Issues & Pull Requests: 2,074 (+0)
Number of forks: 2,334
Total Stargazers: 9,185 (+0)
Total Subscribers: 115 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 10.4 hours
Mean response time: 81.9 days
90th percentile: 175.8 days
Tracked items: 5,181

How this project is maintained

Around half of the issues opened in the past year never receive a reply. 86% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. Work labelled "regression" is answered fastest, typically in about 3 hours, while "EPIC" waits about 2 days. 52% of tracked open issues have had no activity in three months. Only 5% of issues opened in the past year have been closed.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 1,692
New in 7 days: 42
Closed in 7 days: 40
Avg open age: 419 days
Stale 30+ days: 1,501
Stale 90+ days: 1,280

Recent activity

Opened in 7 days: 36
Closed in 7 days: 35
Comments in 7 days: 50
Events in 7 days: 192

Top labels

  • enhancement (4,038)
  • bug (2,736)
  • good first issue (663)
  • help wanted (178)
  • performance (156)
  • documentation (120)
  • regression (87)
  • development-process (68)

Detailed Description

Apache DataFusion is an extensible query engine written in Rust that leverages Apache Arrow as its in-memory columnar format. The project provides libraries and binaries for developers building fast and feature-rich database and analytic systems, with the ability to customize the engine for particular workloads. The core engine offers SQL and DataFrame APIs, a full query planner, a columnar streaming multi-threaded vectorized execution engine, and support for partitioned data sources. Built-in support includes CSV, Parquet, JSON, and Avro formats, with extensive customization capabilities at nearly all points including data sources, query languages, functions, and custom operators.

The repository maintains active development with significant community engagement.

DataFusion's ecosystem extends beyond the core Rust implementation through multiple language bindings and specialized projects. DataFusion Python provides a Python interface for SQL and DataFrame queries, DataFusion Java offers Java bindings, and DataFusion Comet serves as an accelerator for Apache Spark based on the DataFusion engine. The project is classified across multiple analytical domains including analytics platforms, data integration, data processing, scalable queries, distributed computing, streaming data, ETL workflows, batch processing, query optimization, and real-time analytics.

The crate provides extensive customization through configurable features. Default features include nested expressions for working with complex types, compression support for multiple formats, cryptographic and datetime functions, encoding functions, Parquet and SQL support, regular expression functions, Unicode-aware operations, and logical plan unparsing. Optional features add Apache Avro support, backtrace information in error messages, Parquet Modular Encryption, and serialization capabilities. The project follows Apache Software Foundation licensing under the Apache License 2.0 and maintains a committed Cargo.lock file with regular dependency updates via Dependabot, ensuring reproducible builds and managed dependency evolution.