Current issue state, recent activity, and per-issue timelines from the indexed issue data.
| Date | Opened | Closed | Comments | Events | Open Backlog |
|---|---|---|---|---|---|
| 2026-09-09 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-08 | 1 | 0 | 0 | 0 | 1 |
| 2026-09-07 | 0 | 0 | 0 | 0 | 41 |
| 2026-09-06 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-05 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-04 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-03 | 1 | 0 | 0 | 0 | 0 |
| 2026-09-02 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-01 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-31 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-30 | 1 | 0 | 0 | 0 | 0 |
| 2026-08-29 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-28 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-27 | 0 | 0 | 0 | 0 | 0 |
Opened: 2
Closed: 0
Comments: 0
Events: 0
| Issue | Author | State | Labels | Comments | Reactions | Updated |
|---|---|---|---|---|---|---|
#2075 AdEMAMix32bit and PagedAdEMAMix32bit allocate a single-size state1 and fail on the first step Opened 1 day ago | caiotheodoro | open | No labels | 0 | 0 | 1 day ago |
#2011 Offer to help improve the MPS backend fallbacks Opened 2 months ago | egeozkoc | open | No labels | 4 | 0 | 6 days ago |
#2067 quantize_4bit fails with "invalid configuration argument" (ops.cu:54) for tensors of exactly >= 2^31 elements Opened 6 days ago | TomMoeras | open | No labels | 1 | 0 | 6 days ago |
#2064 CPU `gemm_4bit_forward` kernel is requested without `backend="cpu"`, so it never loads on a CUDA torch build Opened 10 days ago | jjjsood | open | No labels | 0 | 0 | 10 days ago |
#1261 Wrong doc and function signature for 8-bit optim Opened 2 years ago | gau-nernst | open | Documentation Contributions Welcome Optimizers | 5 | 0 | 12 days ago |
#2047 CPU dequantize_4bit returns shape (1, n) for even-length 1-D inputs; all other backends return (n,) Opened 20 days ago | 2sumtech | open | No labels | 0 | 0 | 20 days ago |
#2045 Can't quantize nemotron 3.5 Opened 21 days ago | supersimple33 | open | No labels | 0 | 0 | 21 days ago |
#2042 Triton 4-bit quantization stores absmax past the end of the tensor unless the block count is a multiple of 8 Opened 23 days ago | truong-v | closed - completed | Intel | 0 | 0 | 22 days ago |
#2041 get_gaudi_sw_version() hangs indefinitely on Windows — module-level subprocess.run with shell pipe to grep Opened 25 days ago | Mcic1980 | open | Duplicate Windows Proposing to Close | 1 | 0 | 23 days ago |
#2010 Adam/AdamW/LAMB/AdEMAMix: weight decay applied in wrong order in default and Triton backends, diverging from CUDA kernel Opened 2 months ago | ErenAta16 | open | Optimizers CUDA | 6 | 0 | 24 days ago |
#2034 `MatMul4Bit` keeps `B` and `quant_state` on `ctx` instead of `save_for_backward`, so gradient checkpointing cannot recompute them Opened 1 month ago | MakazhanAlpamys | open | No labels | 0 | 0 | 1 month ago |
#2025 Triton FP4 dequantization has a truncated constant, disagreeing with CUDA/CPU by 7 ULP Opened 1 month ago | truong-v | closed - completed | Intel | 1 | 0 | 1 month ago |
#2027 `gemm_4bit`: blocksize-alignment warning emitted on every call, and now also during training (new in 0.50.0) Opened 1 month ago | albertvillanova | open | No labels | 1 | 0 | 1 month ago |
#2021 Target-wide `-mavx512*` puts EVEX in the scalar dequant fallbacks; source builds SIGILL on non-AVX-512 x86_64 Opened 1 month ago | pjordanandrsn | open | Linux Build x64 CPU | 4 | 0 | 1 month ago |
#1785 Support quantizing tensors when numel() > INT_MAX Opened 11 months ago | matthewdouglas | open | CUDA | 2 | 2 | 2 months ago |
#1126 Linux binaries should ship with appropriate RPATH Opened 2 years ago | matthewdouglas | open | Medium Priority Linux Build | 1 | 1 | 2 months ago |
#1868 Question: intentional FP16-only path for int8_vectorwise_quant / LLM.int8 activation quant? (BF16 support + removing casts) Opened 7 months ago | sanghyunna | open | No labels | 6 | 0 | 2 months ago |
#2014 ROCm 8.4 binary not found (latest ROCm is 7.14) Opened 2 months ago | aoguntayo | closed - completed | ROCm | 2 | 0 | 2 months ago |
#1849 Failed to quant MoE models with fused expert weights in transformers v5 Opened 8 months ago | ITcarrot | open | Hugging Face Integration | 8 | 8 | 2 months ago |
#2005 Memory leak: nn.parametrize.replace_parameter_4bit + HF gradient_checkpointing_enable() — dequant cache not released on early-stopped recompute Opened 2 months ago | aka-mnaf-zariche | closed - completed | No labels | 0 | 0 | 2 months ago |
#1991 Prebuilt XPU wheel links `libsycl.so.8`; fails to load on oneAPI 2026 (`libsycl.so.9`) Opened 2 months ago | jiqing-feng | closed - completed | Intel Build | 4 | 0 | 2 months ago |
#1995 quantize_blockwise / dequantize_blockwise produce incorrect results for non-contiguous CPU inputs (CPU follow-up to #1690) Opened 2 months ago | egeozkoc | closed - completed | x64 CPU aarch64 | 1 | 0 | 2 months ago |
#1992 Lion applies coupled (L2) weight decay instead of decoupled in 32-bit CUDA, Triton, and the default backend Opened 2 months ago | eaglstun | closed - completed | Optimizers | 1 | 0 | 2 months ago |
#1955 [ROCm] gfx1201 (Navi 48) support - compilation guide and known issues Opened 4 months ago | UFO0506 | closed - completed | No labels | 5 | 0 | 2 months ago |
#1988 pip wheel ships top-level tests/ directory; shadows consumers tests/ package and breaks python -m unittest discovery Opened 2 months ago | jmcloy622-ops | closed - not_planned | No labels | 1 | 0 | 2 months ago |