CVE-2026-81726: NLTK through 3.10.3 Path Traversal via Model-Artifact APIs
Summary
Several model-artifact APIs still treat caller-controlled model paths as ordinary filenames even when NLTK path security is enforced. The same outside-root paths are rejected by guarded helpers, but these public read and write flows still use raw file APIs.
Details
- Vulnerability type: File sandbox bypass - Affected component: TransitionParser.train, TransitionParser.parse, AveragedPerceptron.save, AveragedPerceptron.load, PerceptronTagger.savetojson, savemaxentparams - Affected versions: Published 3.9.4 and current source v3.10.0-rc2 both reproduced. - Patched versions: Not yet patched - Root cause: Model import and export helpers use built-in open() on caller-controlled paths instead of pathsec-aware helpers.
TransitionParser.train() writes outside allowed roots, TransitionParser.parse() reads outside allowed roots, AveragedPerceptron bypasses the sandbox in both directions, and adjacent read-side helpers in the same family already show the intended guarded behavior. I confirmed outside-root reads and writes while pathsec.open() or the guarded sibling helpers rejected the same paths.
PoC
Preconditions - The application enables pathsec enforcement and lets untrusted workflows choose model import or export paths.
Steps 1. Enable pathsec.ENFORCE=True and restrict allowed roots to a dedicated sandbox directory. 2. Use public model import or export APIs with paths that point outside that root. 3. Observe the same paths are rejected by negative-control guarded helpers such as pathsec.open(), PerceptronTagger.loadfromjson(), or loadmaxentparams(). 4. Observe the vulnerable APIs still read or write outside-root files successfully.
Minimal reproducible excerpt
text transitiontrainexists True transitionparseloaderreadbytes 13 averagedloadkeys ['bias'] maxentsave wrote ['alwayson.tab', 'labels.txt']
Impact
Consumers that rely on pathsec for local containment can be tricked into reading or overwriting files outside approved roots through normal model persistence and loading APIs.
Remediation
Route all model-path file access through nltk.pathsec.open() or existing pathsec-aware helpers, and add regression tests that pair each vulnerable API with a negative control on the same path.
Other sources
NLTK through 3.10.3 contains a path traversal vulnerability in model-artifact APIs that bypass pathsec enforcement by using raw file operations on caller-controlled paths. Attackers can read or write files outside allowed sandbox roots through TransitionParser, AveragedPerceptron, PerceptronTagger, and maxent parameter APIs when pathsec is enabled.
— MITRE
Affected Software
Remediation
Recommended actions to resolve this vulnerability, in priority order.
- Configuration
Enable pathsec enforcement by setting pathsec.ENFORCE=True so model-path file access is constrained to allowed roots.
pathsec ENFORCE = True - Configuration
Restrict allowed roots to a dedicated sandbox directory for model import/export to prevent outside-root reads/writes.
pathsec allowed roots (sandbox directories) = dedicated sandbox directory - Configuration
Route all model-path file access for the affected APIs (TransitionParser.train/parse, AveragedPerceptron.save/load, PerceptronTagger.save_to_json, save_maxent_params) through nltk.pathsec.open() or existing pathsec-aware helpers instead of using built-in open() on caller-controlled paths.
NLTK pathsec-aware model artifact APIs file access mechanism = use nltk.pathsec.open() / existing pathsec-aware helpers - Operational
Add regression tests that, for each vulnerable API, pair the same outside-root path with a negative control using pathsec.open() (and guarded siblings like PerceptronTagger.load_from_json() and load_maxent_params()) to confirm guarded rejection while the vulnerable public read/write APIs previously bypassed pathsec.
Event History
Frequently Asked Questions
Which applications are realistically exposed to this issue?
Applications using NLTK through 3.10.3 are exposed when they enable pathsec and pass attacker-controlled paths to the affected model-artifact APIs: TransitionParser, AveragedPerceptron, PerceptronTagger, or maxent parameter APIs.
What does an attacker need to exploit the vulnerability?
An attacker needs a way to influence a file path supplied to one of the affected APIs. No privileges or user interaction are required according to the provided vector, but exploitation has high attack complexity.
Does enabling pathsec protect affected applications?
No. The issue specifically bypasses pathsec enforcement because the affected APIs use raw file operations on caller-controlled paths, allowing reads or writes outside configured sandbox roots.
How can I determine whether my application is affected?
Review whether it uses NLTK through 3.10.3, has pathsec enabled, and invokes the listed model-artifact APIs with paths derived from untrusted input. Such paths may permit access outside the intended sandbox root.