Awesome-LLMs-for-Video-Understanding is a curated collection and survey of papers, code, and datasets for video understanding with large language models.
The project addresses the need to organize and synthesize the rapidly growing body of work on video-language models by providing a comprehensive survey that covers video understanding techniques powered by large language models, training strategies, relevant tasks, datasets, benchmarks, and evaluation methods. It establishes a taxonomy for video-language models based on video representation and language model functionality, and classifies video understanding tasks from the perspectives of granularity and language involvement.
Developers and researchers working on video understanding, multimodal learning, or large language model applications should use this resource to understand the landscape of available models, training approaches, and evaluation benchmarks. The project suits anyone building or evaluating video-language systems, as it provides structured access to both foundational concepts and recent advances in the field. The survey includes coverage of applications across various domains and discusses how different architectural choices and training strategies affect model performance.
The project maintains active development with regular updates to reflect new models and benchmarks in the field. The maintainers have expanded the collection substantially to include additional models and benchmarks, redesigned figures and tables for clarity, and introduced new organizational chapters to improve accessibility of the material. A follow-up work on video-language model post-training has been released as a complementary resource.