CVE-2026-81725: NLTK before 3.10.3 Regular Expression Denial of Service via Pl196xCorpusReader
Summary
Pl196xCorpusReader still parses whole TEI blocks with multiple lazy regexes over attacker-controlled text. A malformed file with many opening tags and no matching closing tags forces repeated rescans and produces quadratic CPU growth in public reader APIs.
Details
- Vulnerability type: Regular-expression denial of service - Affected component: nltk.corpus.reader.pl196x.TEICorpusView.readblock and Pl196xCorpusReader public methods - Affected versions: Published 3.9.4 and current source v3.10.0-rc2 both reproduced. - Patched versions: Not yet patched - Root cause: Lazy .? whole-block regexes rescan untrusted XML-like blocks from each opening-tag position.
The parser uses regexes for paragraphs, sentences, and word tags across the whole <text> block. When the attacker supplies many unmatched opening tags, each attempt scans toward the end of the block and fails, then restarts from the next opening tag. There is near four-times runtime growth each time the number of malformed <p> tags doubled, through normal public calls such as words() and taggedwords().
PoC
Preconditions - The application parses attacker-influenced PL196X or TEI-like corpus files through public reader APIs.
Steps 1. Create a corpus file with a valid header followed by a <text> block that contains many opening tags and no matching closing tags. 2. Instantiate Pl196xCorpusReader on that corpus. 3. Call words() or taggedwords() and measure elapsed time as the malformed tag count doubles. 4. Observe near quadratic growth instead of near-linear behavior.
Minimal reproducible excerpt
text size=1000 0.014s size=2000 0.057s size=4000 0.231s size=8000 0.927s
Impact
A consumer that accepts attacker-influenced corpus files can be forced into heavy CPU use and parser-thread stalling before the application concludes the input contains no valid content.
Remediation
Replace the whole-block lazy-regex parser with a linear parser or bounded tokenizer, and add regression tests that assert near-linear behavior on malformed inputs with many unmatched tags.
Other sources
NLTK before 3.10.3 contains a regular expression denial of service vulnerability in Pl196xCorpusReader that allows attackers to cause quadratic CPU consumption by supplying malformed TEI blocks with many unmatched opening tags. Attackers can exploit lazy regex patterns in the readblock method through public APIs like words() and taggedwords() to force repeated rescans and achieve near-quadratic runtime growth.
— MITRE
Affected Software
Remediation
Recommended actions to resolve this vulnerability, in priority order.
- Upgrade
Upgrade
pip/nltkto a version that resolves this vulnerability.Fixed in 3.10.3 - Upgrade
Upgrade
nltk.corpus.reader.pl196x.PLIST/TEI Pl196xCorpusReaderto a version that resolves this vulnerability.Fixed in 3.10.3 - Configuration
Modify `nltk.corpus.reader.pl196x.TEICorpusView.read_block` and `Pl196xCorpusReader` so TEI blocks are parsed using a linear parser or bounded tokenizer instead of whole-block lazy-regex patterns that rescan attacker-controlled blocks from each opening-tag position.
Pl196xCorpusReader / TEICorpusView.read_block parser strategy for TEI blocks = Replace whole-block lazy `.*?` regex parsing with a linear parser or bounded tokenizer - Compensating control
Apply a CPU/time limit (e.g., per-request or per-parse timeout) around public corpus reader API calls such as `words()` and `tagged_words()` when parsing attacker-influenced PL196X/TEI-like files to mitigate regex DoS causing near-quadratic runtime growth.
- Operational
Add regression tests that assert near-linear behavior (no near-quadratic growth) for malformed inputs with many opening tags and no matching closing tags when calling public methods like `words()` and `tagged_words()`.
Event History
Frequently Asked Questions
Which NLTK versions are affected?
NLTK versions before 3.10.3 are affected. Version 3.10.3 is not identified as affected by the provided information.
What must an attacker be able to do to trigger the issue?
An attacker must be able to supply malformed TEI blocks containing many unmatched opening tags to code that processes them through Pl196xCorpusReader. The vulnerable parsing behavior can be reached through public APIs including words() and tagged_words().
What is the expected impact of successful exploitation?
Successful exploitation causes excessive CPU consumption through repeated regular-expression rescans, with near-quadratic runtime growth. The provided severity vector indicates an availability impact and does not indicate confidentiality or integrity impact.