yunlong10/Awesome-LLMs-for-Video-Understanding

🔥🔥🔥 [IEEE TCSVT] Latest Papers, Codes and Datasets on Vid-LLMs.

View on GitHub ↗Jump to charts ↓

Data as of . Signed-in members get hourly updates — create a free account.

Summary Information

Updated 46 minutes ago
Added to GitGenius on September 21st, 2026
Created on July 16th, 2023
Open Issues & Pull Requests: 7 (+0)
GitHub issues: Enabled
Number of forks: 152
Total Stargazers: 3,289 (+0)
Total Subscribers: 57 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 7.3 days
Mean response time: 9.7 days
90th percentile: 38.9 days
Tracked items: 9

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 4
New in 7 days: 0
Closed in 7 days: 0
Avg open age: 478 days
Stale 30+ days: 3
Stale 90+ days: 3

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

No label distribution available yet.

Most active issues this week

Sign in to see which issues are moving.
Sign in

Detailed Description

Awesome-LLMs-for-Video-Understanding is a curated collection and survey of papers, code, and datasets for video understanding with large language models.

The project addresses the need to organize and synthesize the rapidly growing body of work on video-language models by providing a comprehensive survey that covers video understanding techniques powered by large language models, training strategies, relevant tasks, datasets, benchmarks, and evaluation methods. It establishes a taxonomy for video-language models based on video representation and language model functionality, and classifies video understanding tasks from the perspectives of granularity and language involvement.

Developers and researchers working on video understanding, multimodal learning, or large language model applications should use this resource to understand the landscape of available models, training approaches, and evaluation benchmarks. The project suits anyone building or evaluating video-language systems, as it provides structured access to both foundational concepts and recent advances in the field. The survey includes coverage of applications across various domains and discusses how different architectural choices and training strategies affect model performance.

The project maintains active development with regular updates to reflect new models and benchmarks in the field. The maintainers have expanded the collection substantially to include additional models and benchmarks, redesigned figures and tables for clarity, and introduced new organizational chapters to improve accessibility of the material. A follow-up work on video-language model post-training has been released as a complementary resource.