MoGe is a deep learning model for monocular geometry estimation that recovers 3D structure from single images.
The tool addresses the challenge of extracting accurate 3D geometry from photographs without stereo or depth sensors. It produces metric point maps, depth maps, normal maps, and camera field of view estimates from a single forward pass. The model supports flexible input resolutions and aspect ratios, and can optionally incorporate ground-truth field of view information to improve accuracy. Performance is optimized for speed, achieving 60 millisecond latency on high-end GPUs with inference resolution adjustable for faster processing when needed.
Developers working on 3D reconstruction, 3D vision applications, or monocular depth estimation tasks should consider this tool. It suits projects requiring accurate geometry from single images without additional sensor input or multi-view data. The project offers multiple versions with progressive improvements: the original MoGe focuses on metric accuracy across open-domain images, MoGe-2 adds metric scale and sharp detail recovery, and MoGe-3 emphasizes fine-grained point map geometry. Installation requires Python 3.10 or newer and works with standard dependency managers, though macOS is not supported due to underlying library constraints.
The project maintains active development with multiple research iterations published and demonstrated through interactive project pages and live demos. The codebase is structured around a single unified model that handles all geometry estimation tasks in one pass rather than requiring separate specialized models. Documentation includes specific guidance on normal map estimation and flexible configuration for inference speed versus accuracy trade-offs.