facebookresearch/tribev2

This repository contains the code to train and evaluate TRIBE v2, a multimodal model for brain response prediction

View on GitHub ↗Jump to charts ↓Open shareable report →

Data as of . Signed-in members get hourly updates — create a free account.

Summary Information

Updated 17 minutes ago
Added to GitGenius on September 21st, 2026
Created on March 24th, 2026
Open Issues & Pull Requests: 50 (+0)
GitHub issues: Enabled
Number of forks: 707
Total Stargazers: 3,277 (+0)
Total Subscribers: 33 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 0.1 hours
Mean response time: 3.2 days
90th percentile: 14.8 days
Tracked items: 31

How this project is maintained

Roughly one issue in three opened in the past year never receives a reply. 100% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. Only 49% of issues opened in the past year have been closed. Three people close 91% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 23
New in 7 days: 0
Closed in 7 days: 0
Avg open age: 125 days
Stale 30+ days: 23
Stale 90+ days: 16

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

No label distribution available yet.

Most active issues this week

Sign in to see which issues are moving.
Sign in

Detailed Description

TRIBE v2 is a multimodal foundation model for predicting brain responses to naturalistic stimuli.

The tool addresses the challenge of understanding how the brain processes complex sensory and linguistic information by training a deep learning model to predict fMRI activity patterns. It works by combining state-of-the-art encoders for vision, audio, and language into a unified Transformer architecture that maps multimodal representations onto the cortical surface. The model predicts brain responses for an average subject across the fsaverage5 cortical mesh, accounting for hemodynamic lag by offsetting predictions by five seconds.

Researchers in computational neuroscience and brain encoding should consider this tool if they need to model how visual, auditory, and linguistic stimuli drive neural activity. The project provides pretrained weights accessible through HuggingFace, making it straightforward to run inference on new video, audio, or text inputs without training from scratch. A Colab demo notebook offers a full walkthrough including brain visualizations. For those needing to train custom models, the repository includes configuration for both local testing and distributed training on Slurm clusters, with dependencies managed through separate installation profiles for inference-only, visualization, or full training setups.

The project maintains active engagement with the research community through documentation of training procedures and contribution guidelines. Development activity shows consistent attention to making the codebase accessible through multiple entry points, from quick-start inference examples to detailed training configurations. The repository includes comprehensive installation instructions tailored to different use cases, reflecting responsiveness to varying user needs. Code organization follows a clear project structure that separates core model logic from training utilities and configuration management.