Repository Issue Activity (beta)

docling-project/docling

Current issue state, recent activity, and per-issue timelines from the indexed issue data.

Open Issues
895
New in 7 Days
17
Closed in 7 Days
10
Average Open Age
228 days
Stale 30+ Days
808
Stale 90+ Days
762
Last 2 Weeks
DateOpenedClosedCommentsEventsOpen Backlog
2026-08-242617698
2026-08-2300001
2026-08-2210002
2026-08-21106165
2026-08-2000261
2026-08-19228418
2026-08-185131126
2026-08-17615237
2026-08-1600373
2026-08-15108195
2026-08-142512338
2026-08-1340495
2026-08-1210592
2026-08-11157264
This Week

Opened: 11

Closed: 9

Comments: 36

Events: 143

Top Labels
bug (925)
question (489)
enhancement (425)
docx (85)
html (52)
pdf parsing (52)
table structure (34)
ocr (33)
Issue Explorer
IssueAuthorStateLabelsCommentsReactionsUpdated

#4049 Bug: HTML tables are missing from DoclingDocument when parsing arXiv/LaTeXML HTML

Opened 2 days ago
medjmalami
open
bug
html
206 hours ago

#2508 Markdown with embedded coordinates

Opened 10 months ago
brownsloth
closed - completed
question
506 hours ago

#4030 Unlimited-OCR preset uses DeepSeek grounding prompt instead of official '<image>document parsing.'

Opened 6 days ago
zStupan
closed - completed
bug
007 hours ago

#3583 Improve parsing of JATS documents

Opened 2 months ago
ceberam
open - reopened
enhancement
xml
good first issue
1207 hours ago

#3939 Document.save_as_json method does not expose json module-level options making the conversion very strict

Opened 20 days ago
filip-komarzyniec
open
question
docling-document
207 hours ago

#4058 Vector-dense pages (CAD/schematic path storms): word-cell decode runs unconditionally in preprocessing and costs GiBs per page — request config surface to bound/skip it

Opened 10 hours ago
vainkop
open
No labels
908 hours ago

#4025 Complete the parsing of iWork Pages

Opened 6 days ago
ceberam
open - reopened
enhancement
iwork
208 hours ago

#3495 bug: borderless table detected as both TableItem and PictureItem (dead code in _handle_cross_type_overlaps)

Opened 3 months ago
adnane-errazine
closed - completed
bug
layout
308 hours ago

#4028 TableFormer ignores drawn horizontal rules, mis-pairing rows when a cell wraps over several lines

Opened 6 days ago
ciaran-finnegan
open
No labels
1011 hours ago

#4022 TesseractOcrCliModel lang=["auto"] silently ignores detected script on Windows (POSIX-only "script/" prefix check)

Opened 7 days ago
Stonica
closed - completed
bug
1012 hours ago

#1762 `_guess_format` method fails on multi-byte characters in UTF-8 documents

Opened 1 year ago
nizq
closed - completed
bug
1212 hours ago

#3901 Extract avatars from pdf page scan

Opened 28 days ago
NikTerentev
closed - completed
question
1018 hours ago

#4053 Invisible text reaches the exported document: `_visible_text_cells` is only wired to the OCR path, not to `get_segmented_page()`

Opened 21 hours ago
hunter-heidenreich
open
bug
2019 hours ago

#4043 Hyphen-minus stripped when a CLI flag (-flag) wraps across a line break during PDF text extraction

Opened 3 days ago
milosdzajic
open
bug
301 day ago

#2330 Docling misses major chunks in HTML parsing

Opened 11 months ago
NikhilVerma
closed - completed
bug
html
1202 days ago

#2620 [html] Docling produces empty Markdown file when parsing HTML document

Opened 9 months ago
tysonite
closed - completed
bug
html
502 days ago

#2528 Docling-serve Fails to Utilize GPU Despite Available CUDA Provider

Opened 10 months ago
walterkru
open
bug
753 days ago

#3685 [Bee] Complex Excel layouts result in duplicate tables when converted to HTML

Opened 2 months ago
jelslip
open
bug
xlsx
303 days ago

#2584 Make RapidOcrOptions ignore `artifacts_path`

Opened 10 months ago
simonschoe
open
enhancement
ocr
1103 days ago

#4018 [Bee] docling parse generates text as each character separated

Opened 7 days ago
yqliving
open
bug
404 days ago

#4016 PDF: wrapped section_header titles split across two elements — SECTION_HEADER missing from predict_merges (companion to #3881)

Opened 7 days ago
Jason-Zonelex
open
No labels
304 days ago

#4034 DOCX textboxes are silently dropped and the result differs run to run: id() of transient lxml proxies used as node identity

Opened 5 days ago
hellsinger1337
open
bug
docx
305 days ago

#4007 pypdfium2 backend: text-cell coordinates ignore /Rotate while page size and rendering apply it

Opened 8 days ago
joanfabregat
closed - completed
No labels
105 days ago

#4002 VlmPipeline (MLX): ValidationError — VlmPredictionToken.logprob receives bfloat16 array instead of float

Opened 9 days ago
lucas-rafa-94
closed - completed
No labels
105 days ago

#4027 Support the parsing of Apple's Keynotes

Opened 6 days ago
ceberam
open
enhancement
iwork
105 days ago

Rows per page:

1–25 of 1,975