QwenLM/Qwen-MM-Plugins

Make any agent harness multimodal-native.

View on GitHub ↗Jump to charts ↓Open shareable report →

Data as of . Signed-in members get hourly updates — create a free account.

Summary Information

Updated 19 minutes ago
Added to GitGenius on September 27th, 2026
Created on July 29th, 2026
Open Issues & Pull Requests: 12 (+0)
GitHub issues: Enabled
Number of forks: 206
Total Stargazers: 3,121 (+0)
Total Subscribers: 14 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 4.9 hours
Mean response time: 3.5 days
90th percentile: 10.7 days
Tracked items: 20

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 7
New in 7 days: 3
Closed in 7 days: 3
Avg open age: 14 days
Stale 30+ days: 0
Stale 90+ days: 0

Recent activity

Opened in 7 days: 3
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

  • enhancement (2)
  • bug (1)
  • windows (1)

Most active issues this week

Sign in to see which issues are moving.
Sign in

Detailed Description

Qwen-MM-Plugins is a plugin framework that enables agents to work natively with multimodal content including images, video, documents, and 3D files.

The tool addresses the challenge of integrating multimodal capabilities into agent workflows. It provides a collection of specialized plugins that handle different types of media processing and reasoning tasks. The core plugin enables reading and understanding images, video, documents, and 3D files. Beyond basic media handling, the framework includes plugins for long-video question-answering, video editing and generation, 3D modeling through Blender and FreeCAD integration, audio-visual memory construction across extended video sequences, music video and movie commentary generation, tutorial video conversion to illustrated documents, physical hardware operation, and 3D spatial reasoning. Plugins can be composed together to build complex multimodal agent behaviors, and the framework separates concerns like model services and web search into standalone plugins.

Developers building agents that need to process and reason over visual, video, or 3D content should consider this tool. It suits projects requiring sophisticated multimodal understanding beyond simple image captioning, such as educational content generation, video analysis and editing, 3D design automation, or hardware control through visual understanding. The project provides a plugin hub with documentation and interactive examples to explore available capabilities.

The project maintains active development with regular plugin additions addressing new use cases. The team engages with the community through dedicated communication channels for discussing workflows and feature requests. Documentation is comprehensive, including cookbooks with video demonstrations and interactive demos for each plugin capability.