apache/gravitino

World's most powerful open data catalog for building a high-performance, geo-distributed and federated metadata lake.

View on GitHub ↗Jump to charts ↓Open shareable report →

Data as of . Signed-in members get hourly updates — create a free account.

Summary Information

Updated 41 minutes ago
Added to GitGenius on September 21st, 2026
Created on April 23rd, 2023
Open Issues & Pull Requests: 1,103 (+0)
GitHub issues: Enabled
Number of forks: 943
Total Stargazers: 3,238 (+0)
Total Subscribers: 42 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 10.1 hours
Mean response time: 16.3 days
90th percentile: 30.9 days
Tracked items: 3,552

How this project is maintained

About 9% of issues opened in the past year have never received a reply. 73% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. Work labelled "bug" is answered fastest, typically in about 2 hours, while "subtask" waits about 4 days. 34% of tracked open issues have had no activity in three months. 78% of issues opened in the past year have been closed, leaving a working backlog.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 674
New in 7 days: 94
Closed in 7 days: 109
Avg open age: 325 days
Stale 30+ days: 428
Stale 90+ days: 346

Recent activity

Opened in 7 days: 84
Closed in 7 days: 92
Comments in 7 days: 24
Events in 7 days: 373

Top labels

  • improvement (1,440)
  • bug (825)
  • subtask (795)
  • good first issue (502)
  • 1.3.0 (405)
  • 2.0.0 (403)
  • 1.0.0 (304)
  • feature (304)

Most active issues this week

Sign in to see which issues are moving.
Sign in

Detailed Description

Apache Gravitino is a federated metadata lake that unifies metadata management across diverse data sources and regions through a single API and model.

Organizations managing metadata across multiple systems face fragmentation: metadata lives in different sources like Hive, MySQL, HDFS, and S3, often spread across regions and clouds. Gravitino solves this by providing a unified metadata access layer that connects directly to underlying systems without requiring data movement. Changes in source systems are immediately reflected through Gravitino's connectors, and the tool supports geo-distribution to share metadata across regions and clouds. It integrates with query engines like Trino and Spark without requiring SQL dialect modifications.

Teams should adopt Gravitino when they operate federated data architectures spanning multiple metadata sources, regions, or clouds and need unified governance. It suits organizations building data lakes or lakehouses that require end-to-end data governance including access control, auditing, and discovery across all assets. The tool is particularly valuable for multi-region metadata synchronization in hybrid or multi-cloud setups. It also provides native Iceberg REST catalog service support. AI asset management capabilities are noted as work in progress.

The project maintains active development with regular commits across core metadata management, connector implementations, and API improvements. Work spans multiple areas including governance features, query engine integrations, and documentation expansion. The codebase shows sustained engineering effort on both foundational systems and new capability areas.