open-compass/opencompass

OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+...

View on GitHub ↗Jump to charts ↓Open shareable report

Summary Information

Updated 49 minutes ago
Added to GitGenius on September 8th, 2026
Created on June 15th, 2023
Open Issues & Pull Requests: 401 (+0)
GitHub issues: Enabled
Number of forks: 862
Total Stargazers: 7,404 (+1)
Total Subscribers: 29 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 0.0 hours
Mean response time: 11.2 days
90th percentile: 11.1 hours
Tracked items: 360

How this project is maintained

Around half of the issues opened in the past year never receive a reply. 95% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. Only 7% of issues opened in the past year have been closed. Three people close 64% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 145
New in 7 days: 2
Closed in 7 days: 0
Avg open age: 558 days
Stale 30+ days: 140
Stale 90+ days: 130

Recent activity

Opened in 7 days: 2
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

  • Bug (2)
  • backlog (2)
  • Enhancement (1)
  • good first issue (1)

Detailed Description

OpenCompass is an LLM evaluation platform that supports a wide range of models including Llama3, Mistral, InternLM2, GPT-4, Qwen, GLM, and Claude across numerous datasets.

The platform addresses the challenge of systematically assessing large language models by providing a unified framework for running evaluations. It enables users to benchmark models against multiple datasets and compare their performance across different capabilities and domains. The tool handles the complexity of configuring diverse models and evaluation datasets through a centralized configuration system.

Organizations evaluating proprietary or open-source language models should consider OpenCompass if they need to run standardized benchmarks across many models simultaneously or compare custom models against established baselines. The platform suits teams building or selecting language models who require reproducible evaluation results and the ability to track performance across different model versions. It works well for researchers and practitioners who want to leverage existing benchmark datasets rather than building evaluation infrastructure from scratch.

The project shows consistent development activity with regular updates and maintenance. The codebase receives ongoing improvements to support newly released models and datasets. Documentation is actively maintained to reflect structural changes and new features. The project maintains engagement with its user community through multiple communication channels and actively incorporates feedback for platform enhancements.