jon-becker/prediction-market-analysis

A framework for collecting and analyzing prediction market data, including the largest publicly available dataset of Polymarket and Kalshi market and trade...

View on GitHub ↗Jump to charts ↓Open shareable report

Summary Information

Updated 38 minutes ago
Added to GitGenius on September 18th, 2026
Created on November 22nd, 2025
Open Issues & Pull Requests: 9 (+0)
GitHub issues: Enabled
Number of forks: 547
Total Stargazers: 3,852 (+0)
Total Subscribers: 31 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 19.6 hours
Mean response time: 11.1 days
90th percentile: 32.2 days
Tracked items: 10

Most active contributors

Sign in to see contributor activity.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 6
New in 7 days: 0
Closed in 7 days: 0
Avg open age: 67 days
Stale 30+ days: 4
Stale 90+ days: 2

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

No label distribution available yet.

Most active issues this week

Sign in to see which issues are moving.

Detailed Description

Prediction Market Analysis is a framework for collecting and analyzing prediction market data that provides the largest publicly available dataset of Polymarket and Kalshi market and trade information.

The framework addresses the challenge of accessing and studying prediction market microstructure by bundling pre-collected datasets alongside tools for gathering new data and running statistical analyses. It collects market metadata, trade history via APIs and blockchain sources, and stores everything in Parquet format with automatic progress tracking. The approach separates data collection, storage, and analysis into distinct phases, allowing researchers to work with historical data or extend the dataset with fresh information.

Researchers studying prediction markets, particularly those analyzing Polymarket and Kalshi, should consider this tool if they need structured access to trade-level data without building collection infrastructure from scratch. The pre-collected dataset eliminates the initial barrier of bootstrapping data collection. The framework is extensible through custom analysis scripts, making it suitable for projects that require both standard statistical outputs and domain-specific investigations. Python 3.9 or later is required.

The project shows active maintenance with regular updates to data collection indexers and analysis capabilities. Documentation is provided for both data schemas and the process of writing new analyses. The maintainer responds to issues and accepts pull requests, indicating ongoing engagement with the user base.