apache/arrow

Apache Arrow is the universal columnar format and multi-language toolbox for fast data interchange and in-memory analytics

View on GitHub ↗Jump to charts ↓Open shareable report →

Data as of . Signed-in members get hourly updates — create a free account.

Summary Information

Updated 36 minutes ago
Added to GitGenius on January 3rd, 2025
Created on February 17th, 2016
Open Issues & Pull Requests: 2,460 (+1)
GitHub issues: Enabled
Number of forks: 4,354
Total Stargazers: 17,185 (+0)
Total Subscribers: 348 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 19.9 hours
Mean response time: 468.6 days
90th percentile: 1871.8 days
Tracked items: 7,381

How this project is maintained

About 7% of issues opened in the past year have never received a reply. 73% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. Work labelled "Component: Continuous Integration" is answered fastest, typically in under an hour, while "Priority: Major" waits about 90 months. 61% of tracked open issues have had no activity in three months, so the open count overstates what is actively being worked. 70% of issues opened in the past year have been closed, leaving a working backlog.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 2,074
New in 7 days: 34
Closed in 7 days: 52
Avg open age: 816 days
Stale 30+ days: 1,728
Stale 90+ days: 1,560

Recent activity

Opened in 7 days: 25
Closed in 7 days: 43
Comments in 7 days: 23
Events in 7 days: 79

Top labels

  • Type: enhancement (14,688)
  • Component: C++ (10,221)
  • Type: bug (10,002)
  • Component: Python (5,920)
  • Component: R (2,758)
  • Component: Continuous Integration (2,187)
  • Status: stale-warning (2,144)
  • Type: task (2,022)

Most active issues this week

Sign in to see which issues are moving.
Sign in

Detailed Description

Apache Arrow is a columnar data format and cross-language toolkit maintained by the Apache Software Foundation, designed to enable fast data interchange and in-memory analytics across heterogeneous systems. The project provides a standardized in-memory representation of data that allows different programming languages and data systems to efficiently share, process, and analyze information without expensive serialization overhead.

The core of Apache Arrow consists of several interconnected components. The Arrow Columnar Format defines a standard and efficient in-memory representation for various data types, including nested structures. The Arrow IPC Format provides efficient serialization of columnar data and metadata for interprocess communication and cross-environment data exchange. The Arrow Flight RPC protocol, built on top of the IPC format, enables remote services to exchange Arrow data with application-defined semantics, making it suitable for storage servers and database systems. Additionally, ADBC (Arrow Database Connectivity) offers Arrow-powered APIs, drivers, and libraries for accessing databases and query engines.

The repository contains reference implementations across multiple programming languages. The main Apache Arrow repository includes C++ libraries, Python bindings, R libraries, Ruby libraries, and C bindings using GLib. Gandiva, an LLVM-based expression compiler, is integrated into the C++ codebase for optimized query execution. Separate repositories maintain implementations for Go, Java, JavaScript, Julia, Rust, Swift, and .NET, reflecting Arrow's commitment to true cross-language interoperability.

The Arrow libraries provide numerous software components beyond the core format. These include columnar vector and table-like containers supporting flat or nested types, a language-agnostic metadata messaging layer using Google's FlatBuffers, reference-counted off-heap buffer memory management for zero-copy data sharing, IO interfaces for local and remote filesystems, and self-describing binary wire formats for RPC and IPC. The project includes integration tests verifying binary compatibility between implementations and provides readers and writers for widely-used file formats such as Parquet and CSV.

Activity data shows substantial ongoing development and community engagement. The project maintains connections with other major data ecosystem projects including pandas, Apache DataFusion, and Microsoft VSCode through overlapping contributor networks, positioning Arrow as a central infrastructure component for the broader data engineering and analytics ecosystem.