GHSA-cgc7-9qp3-86m3: Medium severity pip/docling-slim vulnerability
Summary
The HTML, JATS, ODS (OpenDocument spreadsheet) and BoxNote backends accept table rowspan / colspan values without an upper bound. A few bytes of input, such as <td rowspan="100000000">, make docling run loops proportional to the declared span and allocate a table grid of the declared size. The result is CPU and memory exhaustion.
Details
- docling/backend/htmlbackend.py (getcellspans) parses span attributes with no upper limit. The cell-filling loop then iterates rowspan × colspan times. - docling/backend/jatsbackend.py and docling/backend/boxnotebackend.py fill their tables the same way. - The OpenDocument spreadsheet path scans the declared span range. - Export (for example exporttomarkdown()) materialises the full grid through TableData.grid in docling-core.
documenttimeout does not bound this. It is checked between pipeline stages, and these backends convert the whole document in a single call. maxfilesize and maxnumpages do not help because the payload is tiny.
Measured on 2.130.0: a 54-byte HTML file with rowspan="1e8" takes about 4.4 s of CPU, and the time grows linearly with the value. A 52-byte file with colspan="3000000" takes about 23 s and reaches 4.5 GB peak memory during Markdown export.
Impact
Denial of service of the converting process from a very small input document. Confidentiality and integrity are not affected.
Proof of concept
html <table><tr><td colspan="3000000">x</td></tr></table>
python from docling.documentconverter import DocumentConverter DocumentConverter().convert("span.html").document.exporttomarkdown()
Patches
Fixed in docling 2.131.0 by #4414. Table spans are clamped to the HTML limits (colspan 1000, rowspan 65534) and to the actual size of the table in the HTML, JATS, BoxNote and OpenDocument spreadsheet backends, so conversion time and memory grow with the real table only.
Workarounds
Upgrade to 2.131.0. For older versions:
Run conversions of untrusted documents in a separate process with memory and CPU-time limits, or restrict allowedformats to formats that are not affected.
Affected Software
Remediation
Recommended actions to resolve this vulnerability, in priority order.
- Upgrade
Upgrade
pip/docling-slimto a version that resolves this vulnerability.Fixed in 2.131.0 - Upgrade
Upgrade
pip/doclingto a version that resolves this vulnerability.Fixed in 2.131.0 - Upgrade
Upgrade
doclingto a version that resolves this vulnerability.Fixed in 2.131.0 - Configuration
Restrict allowed_formats to formats that are not affected.
Docling allowed_formats = formats that are not affected - Compensating control
Run conversions of untrusted documents in a separate process with memory and CPU-time limits.
Event History
Frequently Asked Questions
Which inputs can trigger the resource exhaustion?
HTML, JATS, ODS spreadsheet, and BoxNote inputs containing table cells with extremely large rowspan or colspan values can trigger it. A very small file is sufficient because processing and allocation scale with the declared span rather than the input size.
What does an attacker need to do to exploit this issue?
An attacker needs to provide a crafted document with oversized table span attributes and have it processed by Docling. The reported vector is network-accessible, requires no privileges, and requires user interaction.
Do document_timeout, max_file_size, or max_num_pages prevent exploitation?
No. document_timeout is checked only between pipeline stages, while these backends convert an entire document in one call; file-size and page-count limits do not help because the triggering payload can be only a few bytes.
When is memory exhaustion most likely to occur?
Parsing oversized spans can consume CPU, and exporting the resulting document, such as through export_to_markdown(), can materialize the full table grid through TableData.grid. This can cause substantial memory consumption in addition to the span-processing workload.
Is there evidence of an affected version?
The advisory reports measurements on version 2.130.0, where a 54-byte HTML document with rowspan set to 100000000 used about 4.4 seconds of CPU. Processing time was reported to grow linearly with the declared span value.