hkuds/videorag

[KDD'2026] "VideoRAG: Chat with Your Videos"

View on GitHub ↗Jump to charts ↓

Summary Information

Updated 24 minutes ago
Added to GitGenius on September 20th, 2026
Created on February 3rd, 2025
Open Issues & Pull Requests: 21 (+0)
GitHub issues: Enabled
Number of forks: 476
Total Stargazers: 3,380 (+0)
Total Subscribers: 53 (+0)

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 20
New in 7 days: 0
Closed in 7 days: 0
Avg open age: 372 days
Stale 30+ days: 19
Stale 90+ days: 17

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

No label distribution available yet.

Most active issues this week

No issue events were indexed in the last 7 days.

Detailed Description

VideoRAG is a retrieval-augmented generation system that enables conversational interaction with video content through large language models.

The system addresses the challenge of understanding and querying long videos by combining video processing with retrieval-augmented generation techniques. VideoRAG processes video content to extract relevant information and uses retrieval mechanisms to fetch pertinent segments or frames when answering user queries, allowing the language model to ground its responses in actual video content rather than relying solely on learned parameters.

The tool suits developers and researchers working with long-form video understanding who need to build applications where users can ask natural language questions about video content. It is particularly relevant for projects involving multi-modal interaction where video comprehension at scale is required. The approach integrates large language models with video-specific retrieval, making it applicable to scenarios where traditional video analysis methods fall short due to video length or complexity.

The project shows active development with ongoing refinement of its core retrieval and video understanding mechanisms. The codebase demonstrates attention to the practical challenges of handling video data within a language model framework. The work has been developed with consideration for reproducibility and integration into larger systems, as evidenced by the structured approach to combining video processing with retrieval components.