CVAT is a computer vision annotation platform that provides open-source, self-hosted tooling for building high-quality visual datasets through image, video, and 3D annotation with AI-assisted labeling and team collaboration features.
The tool addresses the challenge of creating labeled training data for computer vision models by offering a unified interface for multiple annotation types. It integrates AI-powered assistance through custom ML model connections for detection, segmentation, and tracking to accelerate labeling workflows. The platform supports team collaboration with multi-user and multi-organization capabilities, role-based access control, and review workflows, while maintaining full data sovereignty through self-hosted deployment that keeps data within your infrastructure.
Teams should adopt this tool if they need to own their annotation infrastructure and data without relying on external services. It suits research teams, production AI deployments, and organizations with data privacy requirements. The project provides developer-friendly SDKs and APIs for integration into custom pipelines, and the MIT-licensed core allows modification and redistribution. The platform serves as the foundation for commercial offerings, meaning the open-source version benefits from production-grade battle-testing at scale. For teams preferring managed solutions, fully hosted and enterprise variants are available separately.
Development activity shows consistent maintenance and active community engagement. The project has sustained a large user base with millions of Docker pulls and broad adoption across research and production teams. The engineering team actively maintains the codebase and provides comprehensive documentation including installation guides, tutorials, and an academy resource. The transparent development approach on GitHub since its inception demonstrates commitment to the open-source community.