bradyfu/awesome-multimodal-large-language-models

:sparkles::sparkles:Latest Advances on Multimodal Large Language Models

View on GitHub ↗Jump to charts ↓

Summary Information

Updated 26 minutes ago
Added to GitGenius on September 3rd, 2026
Created on May 19th, 2023
Open Issues & Pull Requests: 110 (+0)
GitHub issues: Enabled
Number of forks: 1,134
Total Stargazers: 18,002 (+0)
Total Subscribers: 291 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 19.2 days
Mean response time: 81.9 days
90th percentile: 230.2 days
Tracked items: 45

How this project is maintained

100% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. Three people close 65% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 44
New in 7 days: 0
Closed in 7 days: 0
Avg open age: 557 days
Stale 30+ days: 43
Stale 90+ days: 43

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

No label distribution available yet.

Most active issues this week

No issue events were indexed in the last 7 days.

Detailed Description

Awesome-Multimodal-Large-Language-Models is a curated collection and research resource documenting advances in multimodal large language models.

The project addresses the need to track and organize the rapidly evolving landscape of multimodal LLMs—models that process and generate content across text, vision, audio, and video modalities. It serves as a centralized reference point for researchers and practitioners seeking to understand the state of the field, featuring surveys, benchmark datasets, evaluation tools, and links to model implementations. The collection emphasizes evaluation frameworks, particularly the MME benchmark series for assessing multimodal capabilities, and documents emerging architectures like the VITA series that target real-time omni-modal interaction.

Developers and researchers working on multimodal systems should adopt this resource to stay informed about recent model architectures, training methodologies, and evaluation standards. It suits teams building vision-language systems, audio-visual models, or video understanding applications who need both a literature overview and practical benchmarking tools. The project is particularly valuable for those implementing or comparing multimodal models, as it centralizes evaluation datasets and citation information that would otherwise require extensive searching across multiple venues.

Almost all open issues are raised by outside users rather than the core team, indicating a substantial base of adopters relying on the resource for real-world work. Issues and pull requests often wait weeks or longer for a first response, suggesting limited capacity for rapid engagement with community contributions.