jeinlee1991/chinese-llm-benchmark

非线智能 NoneLinear -...

View on GitHub ↗Jump to charts ↓

Summary Information

Updated 14 minutes ago
Added to GitGenius on September 10th, 2026
Created on June 4th, 2023
Open Issues & Pull Requests: 18 (+0)
GitHub issues: Enabled
Number of forks: 265
Total Stargazers: 6,436 (+0)
Total Subscribers: 70 (+0)

Repository Insights (GitGenius)

Most active contributors

Sign in to see contributor activity.

Related repositories by overlapping contributors

No overlapping-contributor repos identified yet.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 12
New in 7 days: 0
Closed in 7 days: 0
Avg open age: 434 days
Stale 30+ days: 12
Stale 90+ days: 10

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

No label distribution available yet.

Most active issues this week

No issue events were indexed in the last 7 days.

Detailed Description

Chinese LLM Benchmark is a benchmarking and evaluation system for Chinese language large language models.

The project addresses the need to systematically evaluate and compare the capabilities of Chinese-language LLMs across diverse domains and tasks. It provides a structured evaluation framework spanning seven major capability areas including education, healthcare, finance, law, reasoning and mathematics, language and instruction following, and agent and tool use. The evaluation extends to approximately three hundred fine-grained dimensions within these domains. Beyond rankings, the project maintains a defect library exceeding two million entries that documents model failures and limitations, enabling researchers to analyze and improve model performance.

The tool suits organizations and researchers who need comprehensive capability assessment of Chinese LLMs for their specific use cases. It is particularly valuable for those working in regulated industries like finance, healthcare, and law where domain-specific performance matters, as well as for teams developing or fine-tuning private models. The project offers free evaluation services for private models through direct contact with the team. The breadth of model coverage—spanning both commercial systems and open-source alternatives—makes it applicable whether you are selecting among existing models or benchmarking custom implementations.

The project maintains active development with continuous model additions and evaluation updates. The team publishes technical documentation detailing their evaluation methodology and findings. The defect library receives ongoing expansion as new model evaluations are completed. The project provides direct engagement channels for users seeking evaluation services for proprietary models.