apache/pinot

Apache Pinot - A realtime distributed OLAP datastore

View on GitHub ↗Jump to charts ↓Open shareable report

Summary Information

Updated 13 minutes ago
Added to GitGenius on January 3rd, 2025
Created on May 19th, 2014
Open Issues & Pull Requests: 1,421 (+0)
Number of forks: 1,500
Total Stargazers: 6,125 (+0)
Total Subscribers: 223 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 4.0 days
Mean response time: 425.7 days
90th percentile: 2135.1 days
Tracked items: 1,154

How this project is maintained

Around half of the issues opened in the past year never receive a reply. 95% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. Work labelled "query" is answered fastest, typically in about 7 hours, while "feature" waits about 7 days. 42% of tracked open issues have had no activity in three months. Only 3% of issues opened in the past year have been closed.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 976
New in 7 days: 1
Closed in 7 days: 7
Avg open age: 1,323 days
Stale 30+ days: 859
Stale 90+ days: 806

Recent activity

Opened in 7 days: 0
Closed in 7 days: 7
Comments in 7 days: 27
Events in 7 days: 56

Top labels

  • stale (404)
  • bug (338)
  • feature (279)
  • multi-stage (177)
  • good first issue (166)
  • help wanted (134)
  • ingestion (129)
  • troubleshooting (117)

Detailed Description

Apache Pinot is a real-time distributed OLAP datastore engineered to deliver scalable analytics with low latency on massive datasets. Written in Java, it was originally built by engineers at LinkedIn and Uber to power interactive real-time analytic applications. The system is designed to scale horizontally with no upper bound, maintaining constant performance based on cluster size and expected query throughput.

The platform supports dual ingestion modes, accepting batch data from sources like Hadoop HDFS, Amazon S3, Azure ADLS, and Google Cloud Storage, as well as streaming data from Apache Kafka, Apache Pulsar, and AWS Kinesis. This hybrid capability allows organizations to combine batch and streaming sources into unified tables for querying. Pinot also supports upsert operations during real-time ingestion, enabling at-scale data updates with consistency guarantees.

Pinot's query capabilities center on a standard SQL interface accessible through a built-in query editor and REST API. The system can filter and aggregate petabyte-scale datasets with P90 latencies in the tens of milliseconds, making it suitable for interactive UI applications. It supports versatile joins including arbitrary fact-to-dimension and fact-to-fact operations on petabyte datasets. The architecture is column-oriented with various compression schemes such as Run Length and Fixed Bit Length encoding.

The indexing system is pluggable, offering multiple technologies including timestamp, inverted, StarTree, Bloom filter, range, text search, JSON, and geospatial indexes. Built-in multitenancy enables data management and security across isolated logical namespaces, supporting cloud-friendly resource allocation. The platform is cloud-native on Kubernetes with Helm charts providing horizontally scalable and fault-tolerant clustered deployments.

Pinot is particularly well-suited for executing real-time OLAP queries on immutable data with fast aggregations and analytics. It excels at querying time-series data with numerous dimensions and metrics. The system is designed for contexts requiring both real-time stream ingestion and batch processing while maintaining consistent low-latency query performance. Building Pinot uses Maven, with optimized development builds available through the pinot-fastdev profile for faster iteration cycles.