apache/parquet-java

Apache Parquet Java

View on GitHub ↗Jump to charts ↓Open shareable report →

Data as of . Signed-in members get hourly updates — create a free account.

Summary Information

Updated 2 hours ago
Added to GitGenius on September 22nd, 2026
Created on June 10th, 2014
Open Issues & Pull Requests: 725 (+0)
GitHub issues: Enabled
Number of forks: 1,573
Total Stargazers: 3,084 (+0)
Total Subscribers: 94 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 7.9 days
Mean response time: 99.8 days
90th percentile: 163.1 days
Tracked items: 244

How this project is maintained

Roughly one issue in four opened in the past year never receives a reply. 87% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. Work labelled "Type: bug" is answered fastest, typically in about 6 days, while "Component: Parquet" waits about 7 months. 40% of tracked open issues have had no activity in three months. 61% of issues opened in the past year have been closed, leaving a working backlog.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 107
New in 7 days: 1
Closed in 7 days: 1
Avg open age: 519 days
Stale 30+ days: 94
Stale 90+ days: 70

Recent activity

Opened in 7 days: 1
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 1

Top labels

  • Type: enhancement (122)
  • Type: bug (110)
  • Component: Parquet (23)
  • Priority: Major (21)
  • Component: Java (9)
  • Component: Hadoop (2)
  • Priority: Minor (2)
  • Component: Avro (1)

Most active issues this week

Sign in to see which issues are moving.
Sign in

Detailed Description

Apache Parquet Java is a Java implementation of the Apache Parquet column-oriented data file format for efficient storage and retrieval.

Parquet addresses the challenge of storing and querying large volumes of complex, nested data efficiently. It uses a record shredding and assembly algorithm derived from the Dremel paper to represent nested structures, combined with high-performance compression and encoding schemes. This approach allows bulk data to be stored in a columnar format, which enables selective column access and reduces I/O overhead compared to row-oriented formats.

Developers should adopt this tool when building Java applications that need to read or write Parquet files, particularly in data pipeline, analytics, or big data contexts where columnar storage provides performance benefits. The tool is suitable for projects integrating with Hadoop ecosystems or other analytics platforms that support Parquet, since the format is widely recognized across programming languages and tools. The repository contains the canonical Java implementation and is the appropriate choice for JVM-based systems requiring native Parquet support.

The project maintains active development with continuous integration testing against Hadoop 3. Build requirements include Java 17 or higher and Maven, with dependencies on the Thrift compiler for code generation from Parquet's Thrift definitions. The codebase includes an inlined version of the Parquet Thrift IDL that can be synchronized with the canonical format specification maintained in a separate repository.