depthanything/depth-anything-v2

[NeurIPS 2024] Depth Anything V2. A More Capable Foundation Model for Monocular Depth Estimation

View on GitHub ↗Jump to charts ↓

Summary Information

Updated 43 minutes ago
Added to GitGenius on September 7th, 2026
Created on June 13th, 2024
Open Issues & Pull Requests: 242 (+0)
GitHub issues: Enabled
Number of forks: 908
Total Stargazers: 8,786 (+1)
Total Subscribers: 62 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 5.0 days
Mean response time: 27.3 days
90th percentile: 61.7 days
Tracked items: 149

How this project is maintained

Around half of the issues opened in the past year never receive a reply. 100% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. Only 1% of issues opened in the past year have been closed.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 174
New in 7 days: 0
Closed in 7 days: 0
Avg open age: 601 days
Stale 30+ days: 174
Stale 90+ days: 171

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

No label distribution available yet.

Most active issues this week

No issue events were indexed in the last 7 days.

Detailed Description

Depth Anything V2 is a monocular depth estimation model that predicts depth maps from single images.

The tool addresses the problem of inferring three-dimensional scene structure from monocular input by training a foundation model capable of fine-grained depth prediction across diverse scenes. It operates as a neural network that processes images and outputs corresponding depth maps, with the approach emphasizing robustness and accuracy in both relative and metric depth estimation modes.

Developers should choose this tool if they need depth estimation without stereo input or structured light sensors. It suits applications ranging from 3D reconstruction and autonomous navigation to augmented reality and robotics. The project offers four model variants scaling from 24.8M to 1.3B parameters, allowing trade-offs between inference speed and accuracy depending on deployment constraints. Compared to diffusion-based depth models, the tool provides faster inference, fewer parameters, and higher depth accuracy. The project has expanded beyond single-image estimation to support video depth prediction for extended sequences and metric depth refinement when low-resolution LiDAR prompts are available.

The project maintains active development with recent releases of complementary tools including Video Depth Anything for temporal consistency across long video sequences and Prompt Depth Anything for metric depth at 4K resolution. Integration into major frameworks including Hugging Face Transformers and Apple Core ML Models indicates broad ecosystem adoption. The codebase includes a dedicated benchmark dataset and provides pre-trained checkpoints across all model scales, with smaller metric depth variants available alongside the base relative depth models.