esbatmop/mnbvc

MNBVC(Massive Never-ending BT Vast Chinese...

View on GitHub ↗Jump to charts ↓

Summary Information

Updated 20 minutes ago
Added to GitGenius on September 16th, 2026
Created on December 31st, 2022
Open Issues & Pull Requests: 22 (+0)
GitHub issues: Enabled
Number of forks: 296
Total Stargazers: 4,277 (+0)
Total Subscribers: 71 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 3.1 days
Mean response time: 3.1 days
90th percentile: 5.2 days
Tracked items: 9

Most active contributors

Sign in to see contributor activity.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 7
New in 7 days: 0
Closed in 7 days: 0
Avg open age: 430 days
Stale 30+ days: 5
Stale 90+ days: 5

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

No label distribution available yet.

Most active issues this week

No issue events were indexed in the last 7 days.

Detailed Description

MNBVC is a large-scale Chinese language corpus dataset designed for training natural language processing and machine learning models.

The project addresses the need for comprehensive Chinese text data by aggregating diverse sources of pure text content spanning multiple genres and cultural contexts. The dataset encompasses mainstream and niche cultural content, including news articles, essays, novels, books, magazines, academic papers, scripts, forum posts, wiki entries, classical poetry, song lyrics, product descriptions, jokes, anecdotes, and chat logs. This breadth of sources aims to capture the full spectrum of written Chinese across different registers and communities.

Teams building Chinese language models, particularly those training large language models or developing Chinese NLP applications, should consider this dataset if they need diverse, large-scale training material. The project suits organizations working on machine translation, text generation, sentiment analysis, or other Chinese language tasks where exposure to varied writing styles and cultural contexts is valuable. The inclusion of informal content like chat records and niche cultural material distinguishes it from datasets focused solely on formal or mainstream sources.

The project shows active development with regular updates to the corpus. The repository maintains ongoing expansion of the dataset with new sources being continuously integrated. Documentation and data organization reflect a structured approach to corpus management, with clear categorization of different text types and sources.