VACE is an all-in-one video creation and editing model that unifies reference-to-video generation, video-to-video editing, and masked video-to-video editing into a single framework.
The tool addresses the fragmentation of video generation and editing workflows by consolidating multiple tasks into one model. Rather than requiring separate tools for different operations, VACE enables users to compose tasks freely, supporting capabilities like moving objects, swapping content, applying reference styles, expanding scenes, and animating static elements. This unified approach streamlines workflows by allowing diverse editing operations within a single inference pipeline.
Developers working on video creation projects should consider VACE if they need flexible, composable video editing and generation in a single model rather than chaining multiple specialized tools. The project provides model implementations in different scales and offers inference code, preprocessing utilities, and Gradio demos for integration. The tool is particularly suited for applications requiring both generative and editing capabilities without the overhead of managing separate models.
The project maintains active development with model releases across multiple scales and platforms. Code for model inference, preprocessing, and interactive demos has been released. The team has published evaluation benchmarks to support reproducibility and community assessment of the approach. Community contributions are being tracked through creative use cases shared on the project page.