CVE-2026-102995: pypdf: Possible large memory usage for large /ToUnicode streams (Follow-up 2)
Impact
An attacker who uses this vulnerability can craft a PDF which leads to large memory consumption. This requires parsing the /ToUnicode entry of a font with unusually large values, for example during text extraction.
Patches
This has been fixed in pypdf==6.18.1.
Workarounds
If you cannot upgrade yet, consider applying the changes from PR #4071.
Other sources
pypdf is a free and open-source pure-python PDF library. Prior to 6.18.1, a crafted PDF can place unusually large source-code or destination-string tokens in a font /ToUnicode mapping, causing pypdf/cmap.py parsebfchar to decode and retain oversized values during operations such as text extraction and consume excessive memory. This is a second follow-up to earlier /ToUnicode resource-consumption fixes and is limited to the remaining token-length path. This issue is fixed in version 6.18.1.
— MITRE
Affected Software
Remediation
Recommended actions to resolve this vulnerability, in priority order.
- Upgrade
Upgrade
pip/pypdfto a version that resolves this vulnerability.Fixed in 6.18.1 - Upgrade
Upgrade
pypdfto a version that resolves this vulnerability.Fixed in 6.18.1 - Compensating control
If upgrading is not yet possible, apply the changes from PR #4071 as a workaround.
Event History
Frequently Asked Questions
Which deployments are exposed to this issue?
Deployments using pypdf versions prior to 6.18.1 are affected when they parse attacker-controlled PDFs and process a font's /ToUnicode mapping. Text-extraction workflows are specifically identified as an operation that can trigger the issue.
What does an attacker need to do to exploit it?
An attacker needs to supply a crafted PDF containing unusually large source-code or destination-string tokens in a font /ToUnicode mapping. No privileges or user interaction are indicated by the supplied CVSS vector.
What is the impact of successful exploitation?
Parsing the malicious mapping can cause pypdf to decode and retain oversized values, leading to excessive memory consumption. This can affect availability of the process handling the PDF.
What can be done if upgrading is not immediately possible?
Apply the changes from PR #4071 as a workaround. Otherwise, reduce exposure by avoiding processing untrusted PDFs in affected text-extraction or other parsing workflows until the fix can be deployed.
How can I determine whether the fix is installed?
Check the installed pypdf version. The issue is fixed in pypdf 6.18.1; versions earlier than 6.18.1 are affected.