Repository Issue Activity (beta)

turboderp-org/exllamav2

Current issue state, recent activity, and per-issue timelines from the indexed issue data.

Open Issues
90
New in 7 Days
0
Closed in 7 Days
0
Average Open Age
585 days
Stale 30+ Days
90
Stale 90+ Days
88
Last 2 Weeks
DateOpenedClosedCommentsEventsOpen Backlog
2026-09-1500000
2026-09-14000090
2026-09-1300000
2026-09-1200000
2026-09-1100000
2026-09-1000000
2026-09-0900000
2026-09-0800000
2026-09-0700000
2026-09-0600000
2026-09-0500000
2026-09-0400000
2026-09-0300000
2026-09-0200000
This Week

Opened: 0

Closed: 0

Comments: 0

Events: 0

Top Labels
bug (81)
Issue Explorer
IssueAuthorStateLabelsCommentsReactionsUpdated

#571 Curious about Exllama+TP

Opened 2 years ago
grimulkan
open - reopened
No labels
1502 months ago

#817 Multiple out-of-bounds accesses in QMatrix quant-metadata handling via a crafted GPTQ/EXL2 model

Opened 2 months ago
professor-moody
open
No labels
002 months ago

#814 E8 lattice VQ for KV cache: 2-3 bit with asymmetric K/V — cross-model results

Opened 5 months ago
jagmarques
closed - completed
No labels
204 months ago

#813 [BUG] System freeze during inference with RTX 5060 Ti via Thunderbolt 5 eGPU on Linux (kernel freeze, SSH drop, GPU fan full speed)

Opened 6 months ago
Vassili79
open
bug
106 months ago

#811 [REQUEST] Please add support for Qwen3VL

Opened 11 months ago
sujitvasanth
open
No labels
1011 months ago

#808 [BUG] 2080ti outputs gibberish

Opened 11 months ago
frenzybiscuit
open
bug
1011 months ago

#804 [REQUEST]Please support Torch 2.8

Opened 1 year ago
pivtienduc
open
No labels
2011 months ago

#809 [QUESTION] Inquiry about DynamicGenerator: EOS not returned until max_new_tokens reached, despite stop_conditions

Opened 11 months ago
keds-rnd
open - reopened
bug
3011 months ago

#810 [REQUEST] Make a version that is on Cuda V13

Opened 11 months ago
thebaconstudio
open
No labels
0011 months ago

#806 [QUESTION] I cannot achieve token/s speed as in your tests even with better GPU

Opened 1 year ago
piercarlo62
open
bug
001 year ago

#803 [BUG] [Cause Identified & Workarounds Proposed] Infinite Generation for Llama3 & Other Models with Multiple EOS IDs

Opened 1 year ago
abgulati
open
bug
111 year ago

#800 [BUG] Dry sampling over long contexts

Opened 1 year ago
Ph0rk0z
closed - completed
bug
101 year ago

#802 [BUG] Unable to run examples/chat.py

Opened 1 year ago
homeworkace
open
bug
001 year ago

#795 [BUG] ExllamaV2 version >0.2.8 broken for mistral 7b(v0.2) models on Nvidia 2060

Opened 1 year ago
IceFog72
open
bug
701 year ago

#784 [BUG] Runtime error when trying to load Qwen3 32B

Opened 1 year ago
umar-mq
open
bug
1001 year ago

#798 [REQUEST] Support for Hunyuan-A13B-Instruct

Opened 1 year ago
RodriMora
open
No labels
011 year ago

#797 [BUG] ExLlamaV2Generator import broken across WHL & source repo — Windows 11 + RTX 5090 + CUDA 12.8 build inconsistencies

Opened 1 year ago
mindworksmanagement
open
bug
001 year ago

#777 [BUG]gemma 3 27b exl2 loops nonsense afterwards 2-3 correct paragraphs

Opened 1 year ago
ciprianveg
open
bug
1101 year ago

#749 [REQUEST] Please add support for Gemma3.

Opened 2 years ago
emzaedu
open
No labels
38181 year ago

#793 [BUG] Exllamav2 quickly devolves into endless repetition in versions newer than 2.8.0

Opened 1 year ago
ZhenyaPav
open
bug
101 year ago

#789 [BUG] Remove Sentencepiece

Opened 1 year ago
kingbri1
closed - completed
bug
201 year ago

#724 [REQUEST] Support new SOTA vision model: Qwen 2.5 VL (3B, 7B, 72B)

Opened 2 years ago
ThomasBaruzier
closed - completed
No labels
221 year ago

#792 [QUESTION] Estimate measurements and quantization peak VRAM use in advance

Opened 1 year ago
ThomasBaruzier
open
No labels
001 year ago

#790 [BUG] Silent crash in safetensors call when compiling shards

Opened 1 year ago
r0mar0ma
open
bug
001 year ago

#780 [BUG] Blue Screen MEMORY_MANAGEMENT Error when trying to quantize Gemma3.

Opened 1 year ago
Nrgte
open
bug
201 year ago

Rows per page:

1–25 of 175