TRIBE v2 is a multimodal foundation model for predicting brain responses to naturalistic stimuli.
The tool addresses the challenge of understanding how the brain processes complex sensory and linguistic information by training a deep learning model to predict fMRI activity patterns. It works by combining state-of-the-art encoders for vision, audio, and language into a unified Transformer architecture that maps multimodal representations onto the cortical surface. The model predicts brain responses for an average subject across the fsaverage5 cortical mesh, accounting for hemodynamic lag by offsetting predictions by five seconds.
Researchers in computational neuroscience and brain encoding should consider this tool if they need to model how visual, auditory, and linguistic stimuli drive neural activity. The project provides pretrained weights accessible through HuggingFace, making it straightforward to run inference on new video, audio, or text inputs without training from scratch. A Colab demo notebook offers a full walkthrough including brain visualizations. For those needing to train custom models, the repository includes configuration for both local testing and distributed training on Slurm clusters, with dependencies managed through separate installation profiles for inference-only, visualization, or full training setups.
The project maintains active engagement with the research community through documentation of training procedures and contribution guidelines. Development activity shows consistent attention to making the codebase accessible through multiple entry points, from quick-start inference examples to detailed training configurations. The repository includes comprehensive installation instructions tailored to different use cases, reflecting responsiveness to varying user needs. Code organization follows a clear project structure that separates core model logic from training utilities and configuration management.