CVE-2026-81725: NLTK before 3.10.3 Regular Expression Denial of Service via Pl196xCorpusReader

Published Aug 27, 2026
·
Updated

Summary

Pl196xCorpusReader still parses whole TEI blocks with multiple lazy regexes over attacker-controlled text. A malformed file with many opening tags and no matching closing tags forces repeated rescans and produces quadratic CPU growth in public reader APIs.

Details

- Vulnerability type: Regular-expression denial of service - Affected component: nltk.corpus.reader.pl196x.TEICorpusView.readblock and Pl196xCorpusReader public methods - Affected versions: Published 3.9.4 and current source v3.10.0-rc2 both reproduced. - Patched versions: Not yet patched - Root cause: Lazy .? whole-block regexes rescan untrusted XML-like blocks from each opening-tag position.

The parser uses regexes for paragraphs, sentences, and word tags across the whole <text> block. When the attacker supplies many unmatched opening tags, each attempt scans toward the end of the block and fails, then restarts from the next opening tag. There is near four-times runtime growth each time the number of malformed <p> tags doubled, through normal public calls such as words() and taggedwords().

PoC

Preconditions - The application parses attacker-influenced PL196X or TEI-like corpus files through public reader APIs.

Steps 1. Create a corpus file with a valid header followed by a <text> block that contains many opening tags and no matching closing tags. 2. Instantiate Pl196xCorpusReader on that corpus. 3. Call words() or taggedwords() and measure elapsed time as the malformed tag count doubles. 4. Observe near quadratic growth instead of near-linear behavior.

Minimal reproducible excerpt

text size=1000 0.014s size=2000 0.057s size=4000 0.231s size=8000 0.927s

Impact

A consumer that accepts attacker-influenced corpus files can be forced into heavy CPU use and parser-thread stalling before the application concludes the input contains no valid content.

Remediation

Replace the whole-block lazy-regex parser with a linear parser or bounded tokenizer, and add regression tests that assert near-linear behavior on malformed inputs with many unmatched tags.

Other sources

NLTK before 3.10.3 contains a regular expression denial of service vulnerability in Pl196xCorpusReader that allows attackers to cause quadratic CPU consumption by supplying malformed TEI blocks with many unmatched opening tags. Attackers can exploit lazy regex patterns in the readblock method through public APIs like words() and taggedwords() to force repeated rescans and achieve near-quadratic runtime growth.

— MITRE

Affected Software

3 affected componentsFixes available
nltk nltk<3.10.3
nltk nltk<3.10.3
pip/nltk<=3.10.2
3.10.3

Remediation

Recommended actions to resolve this vulnerability, in priority order.

  1. Upgrade

    Upgrade pip/nltk to a version that resolves this vulnerability.

    Fixed in 3.10.3
  2. Upgrade

    Upgrade nltk.corpus.reader.pl196x.PLIST/TEI Pl196xCorpusReader to a version that resolves this vulnerability.

    Fixed in 3.10.3
  3. Configuration

    Modify `nltk.corpus.reader.pl196x.TEICorpusView.read_block` and `Pl196xCorpusReader` so TEI blocks are parsed using a linear parser or bounded tokenizer instead of whole-block lazy-regex patterns that rescan attacker-controlled blocks from each opening-tag position.

    Pl196xCorpusReader / TEICorpusView.read_block parser strategy for TEI blocks = Replace whole-block lazy `.*?` regex parsing with a linear parser or bounded tokenizer
  4. Compensating control

    Apply a CPU/time limit (e.g., per-request or per-parse timeout) around public corpus reader API calls such as `words()` and `tagged_words()` when parsing attacker-influenced PL196X/TEI-like files to mitigate regex DoS causing near-quadratic runtime growth.

  5. Operational

    Add regression tests that assert near-linear behavior (no near-quadratic growth) for malformed inputs with many opening tags and no matching closing tags when calling public methods like `words()` and `tagged_words()`.

Event History

Aug 27, 2026
CVE Published
via MITRE·02:51 PM
Data Sourced
via MITRE·02:51 PM
DescriptionSeverityWeakness
Data Sourced
via NVD·05:21 PM
DescriptionSeverityWeaknessAffected Software
Sep 8, 2026
Advisory Published
via GitHub·08:28 PM
Data Sourced
via GitHub·08:28 PM
DescriptionWeaknessAffected Software

Frequently Asked Questions

1

Which NLTK versions are affected?

NLTK versions before 3.10.3 are affected. Version 3.10.3 is not identified as affected by the provided information.

2

What must an attacker be able to do to trigger the issue?

An attacker must be able to supply malformed TEI blocks containing many unmatched opening tags to code that processes them through Pl196xCorpusReader. The vulnerable parsing behavior can be reached through public APIs including words() and tagged_words().

3

What is the expected impact of successful exploitation?

Successful exploitation causes excessive CPU consumption through repeated regular-expression rescans, with near-quadratic runtime growth. The provided severity vector indicates an availability impact and does not indicate confidentiality or integrity impact.

Contact

SecAlerts Pty Ltd.
132 Wickham Terrace
Fortitude Valley,
QLD 4006, Australia
info@secalerts.co
By using SecAlerts services, you agree to our services end-user license agreement. This website is safeguarded by reCAPTCHA and governed by the Google Privacy Policy and Terms of Service. All names, logos, and brands of products are owned by their respective owners, and any usage of these names, logos, and brands for identification purposes only does not imply endorsement. If you possess any content that requires removal, please get in touch with us.
© 2026 SecAlerts Pty Ltd.
ABN: 70 645 966 203, ACN: 645 966 203