Repository Issue Activity (beta)

kai-scheduler/KAI-Scheduler

Current issue state, recent activity, and per-issue timelines from the indexed issue data.

Open Issues
88
New in 7 Days
9
Closed in 7 Days
8
Average Open Age
37 days
Stale 30+ Days
34
Stale 90+ Days
3
Last 2 Weeks
DateOpenedClosedCommentsEventsOpen Backlog
2026-07-24105204
2026-07-23342142
2026-07-2200274
2026-07-2131395
2026-07-2002244
2026-07-1911001
2026-07-1800003
2026-07-1710000
2026-07-16514165
2026-07-15432157
2026-07-14234229
2026-07-13125219
2026-07-124093128
2026-07-1100000
This Week

Opened: 8

Closed: 8

Comments: 14

Events: 54

Top Labels
enhancement (135)
bug (75)
stale (24)
Q2-2026 (12)
question (11)
good first issue (10)
help wanted (10)
needs-design (9)
Issue Explorer
IssueAuthorStateLabelsCommentsReactionsUpdated

#1966 preempt: JobOrderFn tiebreak has no resource-awareness, falls back to creation-time order

Opened 1 day ago
CoolingCube
open
enhancement
208 hours ago

#1944 Knative scale-to-zero causes PodGroup creation failure when Knative gang scheduling is enabled

Opened 3 days ago
NaelAsbi123
open
No labels
3114 hours ago

#1913 feat(scheduler): allow jobs to override stale gang eviction grace period

Opened 10 days ago
enoodle
closed - completed
enhancement
good first issue
5014 hours ago

#1968 stalegangeviction evicts pods of successfully completing jobs

Opened 19 hours ago
lixiang233
open
bug
1014 hours ago

#1967 [Bug] CUDA_DEVICE_MEMORY_LIMIT is not set by HAMI Core plugin on UMA nodes where nvidia.com/gpu.memory label is missing

Opened 1 day ago
akshnevrekar
closed - not_planned
bug
101 day ago

#1827 perf(scheduler): bound per-task per-node fit-error retention during allocation

Opened 20 days ago
enoodle
closed - completed
No labels
001 day ago

#1964 Segmented PyTorch grouper ignores elastic minReplicas and requires all worker segments

Opened 1 day ago
rotembubrunai
open
No labels
111 day ago

#1927 podgrouper: segmented PyTorch and LWS emit invalid minMember on parent SubGroups

Opened 9 days ago
rotembubrunai
closed - completed
No labels
202 days ago

#1756 [Numa awareness] Support best effort topology manager

Opened 1 month ago
davidLif
open
enhancement
103 days ago

#1945 status-updater floods logs and API server retrying status/patch updates for deleted PodGroups

Opened 3 days ago
david-gang
open
No labels
213 days ago

#1948 Pipeline-only allocation leaks speculative fit errors

Opened 3 days ago
enoodle
open
No labels
113 days ago

#1888 PodGroup schedulingConditions not cleared after the workload schedules

Opened 12 days ago
gshaibi
closed - completed
No labels
113 days ago

#1929 docs/metrics: inconsistent naming between queue_deserved_gpus and queue_quota_* metrics

Opened 9 days ago
david-gang
open
No labels
104 days ago

#1933 How to schedule a pod onto a cordoned node?

Opened 8 days ago
lasse-ii
open
No labels
304 days ago

#1936 Set annotation to ignore quota occupied by workload

Opened 7 days ago
amy
open
enhancement
needs-design
724 days ago

#1939 reclaim: FeasibleNodesForJob reuses the reclaimer's node set to re-home victims, needlessly killing relocatable CPU-only victims

Opened 5 days ago
david-gang
open
No labels
404 days ago

#1873 DRA GPU count overflow can understate queue demand and prevent eligible GPU reclaim

Opened 15 days ago
thc1006
closed - completed
No labels
505 days ago

#1584 Design kai behavior for priority class preemptionPolicy

Opened 2 months ago
davidLif
closed - not_planned
enhancement
105 days ago

#848 Scheduler Assigns Multiple Workloads to Fully-Allocated GPU Node Causing Infinite Retry Loop

Opened 7 months ago
dttung-starling
closed - completed
bug
1935 days ago

#1821 feat: Container level VRAM metrics - integration with HAMi

Opened 21 days ago
dttung2905
open
enhancement
106 days ago

#1934 Can reclaim enforce fair-share (not just deserved quota) between sibling queues?

Opened 8 days ago
lasse-ii
open
No labels
108 days ago

#1930 shared DRA ResourceClaim double-counted per pod makes node unschedulable

Opened 9 days ago
TensorRaya
open
No labels
118 days ago

#1856 NUMA scale tests

Opened 17 days ago
itsomri
closed - completed
enhancement
008 days ago

#1912 Cannot find parameter scheduler.gpuSharing.hamicoreEnabled in values.yaml file

Opened 10 days ago
shakir91
closed - completed
No labels
219 days ago

#1923 queue validator: validateParentChildQuota only sums CPU across siblings

Opened 9 days ago
kshitizlohia1994
open
No labels
009 days ago

Rows per page:

1–25 of 347