Repository Issue Activity (beta)

deepspeedai/DeepSpeed

Current issue state, recent activity, and per-issue timelines from the indexed issue data.

Open Issues
1,129
New in 7 Days
9
Closed in 7 Days
5
Average Open Age
942 days
Stale 30+ Days
1,109
Stale 90+ Days
1,094
Last 2 Weeks
DateOpenedClosedCommentsEventsOpen Backlog
2026-08-2400000
2026-08-23132123
2026-08-22215143
2026-08-2110284
2026-08-2000283
2026-08-1921123
2026-08-1800111
2026-08-1730003
2026-08-1601011
2026-08-1500000
2026-08-14104142
2026-08-1320026
2026-08-12015173
2026-08-1100242
This Week

Opened: 6

Closed: 5

Comments: 13

Events: 45

Top Labels
bug (1663)
training (963)
enhancement (417)
inference (300)
ci-failure (120)
deepspeed-chat (106)
compression (83)
build (42)
Issue Explorer
IssueAuthorStateLabelsCommentsReactionsUpdated

#8290 HF transformers main injects 'embedding_rowwise' into tp_plan for tied-embedding models; AutoTP rejects the whole plan

Opened 2 days ago
delock
open
No labels
2020 hours ago

#8260 [BUG] reduce_bucket_size=0 passes validation but crashes ZeRO backward with ZeroDivisionError

Opened 7 days ago
fwerkor
closed - completed
No labels
0022 hours ago

#7413 [BUG] FlopsProfiler accumulates metrics when called multiple times

Opened 1 year ago
PalmaLeandro
closed - completed
No labels
201 day ago

#8156 compute_elastic_config raises bare ZeroDivisionError for return_microbatch=True on non-0.2 elasticity versions without explicit world_size

Opened 1 month ago
ErenAta16
closed - completed
No labels
201 day ago

#8297 [BUG] ZeRO-1/2 fail mapping zero-sized parameter inside a partition

Opened 1 day ago
fwerkor
open
No labels
001 day ago

#8197 [REQUEST] OPSD Profile and improve HybridEngine rollout performance

Opened 24 days ago
nathon-lee
open
enhancement
1011 day ago

#8279 [BUG] ZeRO-1/2 fail on trailing zero-sized trainable parameters

Opened 5 days ago
fwerkor
closed - completed
No labels
602 days ago

#8291 [Ulysses SP] process-wide _ulysses_num_kv_heads global breaks a second model with a different head count in the same process

Opened 2 days ago
delock
open
No labels
012 days ago

#8285 [AutoTP] update_mp_params shards vision-tower attributes (patch_embed.embed_dim) without sharding weights, breaking multimodal forward

Opened 3 days ago
delock
open
No labels
003 days ago

#8193 [REQUEST] Dynamic/Adaptive Prefetching Window for ZeRO-3

Opened 26 days ago
sowndappan5
open
enhancement
313 days ago

#8173 [RFC] Support sharding LM heads and adopting Online Softmax

Opened 1 month ago
jinyouzhi
open
enhancement
124 days ago

#102 max_grad_norm is ignored in FP16 training

Opened 7 years ago
LvanderGoten
open
No labels
514 days ago

#7775 [BUG] ZeRO-3 with `zero_quantized_weights=true` incorrectly casts bf16 inputs to fp16, causing BERT training failure

Opened 7 months ago
sdjasj
closed - completed
bug
training
404 days ago

#8276 [BUG] DeepSeek-V3 MLA rank changes trigger 16-byte data_ptr alignment failure in grouped_mm under DeepSpeed

Opened 5 days ago
fwerkor
open
No labels
005 days ago

#8263 Hybrid Engine registers partial inference policies for unsupported Qwen architectures

Opened 7 days ago
LiRunGuo
open
No labels
007 days ago

#8262 ZeRO-3 OPSD rollout deadlocks when data-parallel ranks finish generation at different lengths

Opened 7 days ago
LiRunGuo
open
No labels
007 days ago

#8076 [BUG]ds_z3_config error

Opened 2 months ago
whyseu
closed - completed
bug
training
108 days ago

#8249 IndexError in is_nfs_path() when df wraps a long device name (breaks `import deepspeed`)

Opened 11 days ago
avnerv
open
No labels
109 days ago

#8252 Define affine IR for checkpointing

Opened 10 days ago
delock
open
enhancement
1010 days ago

#8230 [RFC] Universal checkpoint: non-semantic per-parameter geometric shard map

Opened 17 days ago
delock
open
No labels
9110 days ago

#8248 [BUG]DeepSpeed latest version fails during NCCL process-group initialization with CUDA error 600, while 0.17.2 works

Opened 11 days ago
janelu9
open
bug
training
0011 days ago

#7384 [BUG] Ulysses DistributedAttention silently produces incorrect output when #GPUs does not divide global sequence length

Opened 1 year ago
selflein
open
bug
training
3011 days ago

#8231 Cleanup: remove tp_shard process-wide scalar globals (num_kv_heads / tp_grain_size / ...), thread explicitly

Opened 17 days ago
delock
open
No labels
3112 days ago

#7433 [BUG] TypeError: 'staticmethod' object is not callable, in deepcompile (patch_compiled_func.py)

Opened 1 year ago
rookie-ryan
closed - completed
bug
training
1012 days ago

#8224 ZeRO-2 + bf16 silently computes incorrect gradients at multi-rank when a submodule is used more than once per step, a regression from #7665, still present in 0.19.3

Opened 18 days ago
souroosh
open - reopened
No labels
2113 days ago

Rows per page:

1–25 of 3,226