DualPipe is a bidirectional pipeline parallelism algorithm that achieves full overlap of forward and backward computation-communication phases during distributed model training.
The tool addresses the inefficiency of traditional pipeline parallelism, which leaves devices idle during certain phases. DualPipe solves this by scheduling micro-batches bidirectionally through the pipeline stages, allowing forward and backward computations to overlap with communication operations. This bidirectional approach significantly reduces pipeline bubbles compared to existing methods like 1F1B and ZB1P. The project also provides DualPipeV, a variant that uses a V-shaped schedule to reduce the number of devices required while maintaining the same bubble reduction benefits.
Teams training large language models with pipeline parallelism should consider DualPipe if they want to improve training efficiency and reduce idle time across their compute cluster. The approach is particularly suited to distributed training scenarios where communication overhead is a bottleneck. Compared to 1F1B, DualPipe reduces pipeline bubbles at the cost of doubling parameter storage per device. DualPipeV offers a middle ground, achieving the same bubble reduction as DualPipe while requiring only half the devices. The implementation requires PyTorch 2.0 or above and demands a custom overlapped_forward_backward method tailored to specific model architectures.
The project maintains focused scope with a small core development team. The codebase is written in Python and includes example implementations alongside the core algorithm. Documentation references the original technical report and provides visual scheduling diagrams to illustrate how the algorithm distributes work across pipeline stages.