deepseek-ai/DualPipe

A bidirectional pipeline parallelism algorithm for computation-communication overlap in DeepSeek V3/R1 training.

View on GitHub ↗Jump to charts ↓

Data as of . Signed-in members get hourly updates — create a free account.

Summary Information

Updated 58 minutes ago
Added to GitGenius on September 23rd, 2026
Created on February 26th, 2025
Open Issues & Pull Requests: 5 (+0)
GitHub issues: Enabled
Number of forks: 332
Total Stargazers: 3,014 (+0)
Total Subscribers: 28 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 4.9 days
Mean response time: 45.5 days
90th percentile: 67.0 days
Tracked items: 11

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 3
New in 7 days: 0
Closed in 7 days: 0
Avg open age: 534 days
Stale 30+ days: 3
Stale 90+ days: 3

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

No label distribution available yet.

Most active issues this week

Sign in to see which issues are moving.
Sign in

Detailed Description

DualPipe is a bidirectional pipeline parallelism algorithm that achieves full overlap of forward and backward computation-communication phases during distributed model training.

The tool addresses the inefficiency of traditional pipeline parallelism, which leaves devices idle during certain phases. DualPipe solves this by scheduling micro-batches bidirectionally through the pipeline stages, allowing forward and backward computations to overlap with communication operations. This bidirectional approach significantly reduces pipeline bubbles compared to existing methods like 1F1B and ZB1P. The project also provides DualPipeV, a variant that uses a V-shaped schedule to reduce the number of devices required while maintaining the same bubble reduction benefits.

Teams training large language models with pipeline parallelism should consider DualPipe if they want to improve training efficiency and reduce idle time across their compute cluster. The approach is particularly suited to distributed training scenarios where communication overhead is a bottleneck. Compared to 1F1B, DualPipe reduces pipeline bubbles at the cost of doubling parameter storage per device. DualPipeV offers a middle ground, achieving the same bubble reduction as DualPipe while requiring only half the devices. The implementation requires PyTorch 2.0 or above and demands a custom overlapped_forward_backward method tailored to specific model architectures.

The project maintains focused scope with a small core development team. The codebase is written in Python and includes example implementations alongside the core algorithm. Documentation references the original technical report and provides visual scheduling diagrams to illustrate how the algorithm distributes work across pipeline stages.