facebookresearch/mmf

A modular framework for vision & language multimodal research from Facebook AI Research (FAIR)

View on GitHub ↗Jump to charts ↓

Summary Information

Updated 32 minutes ago
Added to GitGenius on September 12th, 2026
Created on June 27th, 2018
Open Issues & Pull Requests: 150 (+0)
GitHub issues: Enabled
Number of forks: 938
Total Stargazers: 5,631 (+0)
Total Subscribers: 109 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 52.9 days
Mean response time: 391.5 days
90th percentile: 1333.0 days
Tracked items: 4

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 6
New in 7 days: 0
Closed in 7 days: 0
Avg open age: 808 days
Stale 30+ days: 6
Stale 90+ days: 6

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

No label distribution available yet.

Most active issues this week

No issue events were indexed in the last 7 days.

Detailed Description

MMF is a modular framework for vision and language multimodal research built on PyTorch. The framework addresses the need for a unified, reusable codebase for multimodal machine learning projects that combine visual and textual information. It provides reference implementations of state-of-the-art vision and language models, allowing researchers to avoid rebuilding common components for each new project. The framework supports distributed training and is designed to be scalable and fast while remaining unopinionated about specific architectural choices.

Developers adopting MMF should know it serves two primary use cases: bootstrapping new vision and language research projects and providing a starter codebase for challenge competitions. The framework is particularly suited for work involving visual question answering, image captioning, visual dialog, and tasks requiring joint reasoning over images and text. It has been used as the foundation for multiple research projects at Facebook AI Research and has powered official challenge codebases for competitions around datasets like Hateful Memes, TextVQA, and TextCaps.

The project maintains active development with continuous integration checks and comprehensive documentation. The codebase receives regular updates to support new research directions and model implementations. The framework demonstrates sustained investment in both code quality and user accessibility through its maintained documentation site and example implementations.