facebookresearch/vggt

[CVPR 2025 Best Paper Award] VGGT: Visual Geometry Grounded Transformer

View on GitHub ↗Jump to charts ↓Open shareable report

Summary Information

Updated 20 minutes ago
Added to GitGenius on September 4th, 2026
Created on February 18th, 2025
Open Issues & Pull Requests: 275 (+0)
GitHub issues: Enabled
Number of forks: 1,547
Total Stargazers: 14,357 (+0)
Total Subscribers: 482 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 11.7 hours
Mean response time: 6.0 days
90th percentile: 10.3 days
Tracked items: 379

How this project is maintained

Around half of the issues opened in the past year never receive a reply. 100% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. 93% of tracked open issues have had no activity in three months, so the open count overstates what is actively being worked. Only 3% of issues opened in the past year have been closed.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 249
New in 7 days: 0
Closed in 7 days: 0
Avg open age: 386 days
Stale 30+ days: 249
Stale 90+ days: 242

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

No label distribution available yet.

Most active issues this week

No issue events were indexed in the last 7 days.

Detailed Description

VGGT is a visual geometry grounded transformer for 3D scene understanding and reconstruction from video or image sequences.

The tool addresses the challenge of recovering 3D geometric structure from visual input by grounding transformer architectures in visual geometry principles. Rather than treating 3D reconstruction as a purely learned task, VGGT incorporates geometric constraints and reasoning directly into its transformer design, enabling it to predict camera poses, depth maps, and 3D point clouds from image sequences. The approach leverages the visual geometry group's expertise to build geometric priors into the model architecture itself.

Developers working on 3D reconstruction, structure-from-motion, or scene understanding tasks should consider this tool, particularly when working with video sequences or multi-view image collections. The project is well-suited for applications requiring both geometric accuracy and the flexibility of learned representations. A commercial-use-friendly checkpoint is available for production deployments, though access requires an application process. The tool integrates with standard 3D pipelines by supporting COLMAP format output, making it compatible with NeRF and Gaussian splatting libraries.

The project has active development with recent memory optimization improvements that increase the number of input frames processable within a given GPU budget. Training code is available for fine-tuning on custom datasets. A follow-up version with additional capabilities has been released, indicating sustained development momentum beyond the initial research contribution.