CVE-2026-84366: Scrapy: S3DownloadHandler sends signed S3 requests over plaintext HTTP by default
Problem
Scrapy’s S3DownloadHandler sends signed S3 requests over plaintext HTTP by default.
A normal request like s3://bucket/key is converted into http://bucket.s3.amazonaws.com/key unless request.meta["issecure"] is explicitly set. The generated request is then signed with configured AWS credentials, so AWS authorization material can be sent without TLS.
Vulnerable code in scrapy/core/downloader/handlers/s3.py:
python scheme = "https" if request.meta.get("issecure") else "http" url = f"{scheme}://{bucket}.s3.amazonaws.com{path}"
The request is then signed and dispatched:
python self.signer.addauth(awsrequest) request = request.replace(url=url, headers=awsrequest.headers.items())
Impact
Users making Scrapy s3:// requests with AWS credentials are impacted.
A network attacker able to observe traffic between Scrapy and S3, such as a public Wi-Fi attacker, compromised router, ISP/corporate network observer, or local network attacker using ARP spoofing, can read:
text bucket/key path AWS Authorization header X-Amz-Security-Token, if temporary credentials are used S3 object contents S3 response headers
An active MITM attacker can also modify the plaintext S3 response body, status code, and headers before Scrapy processes it. This can cause scraped data poisoning, poisoned exports, HTTP cache poisoning when cache is enabled, and influence over later crawl targets through forged redirects or attacker-controlled links.
Suggested classification: CWE-319: Cleartext Transmission of Sensitive Information.
PoC
A minimal PoC creates an s3:// request, enables fake AWS settings, captures the request produced by S3DownloadHandler, and prints the rewritten URL and auth headers.
Other sources
Scrapy is a high-level web crawling and scraping framework for Python. Prior to 2.17.0, in scrapy/core/downloader/handlers/s3.py, Scrapy's S3DownloadHandler converts an S3-scheme bucket and key request into a plaintext HTTP request to the corresponding S3 endpoint unless request.meta["issecure"] is explicitly enabled, then signs and sends the plaintext request with configured AWS credentials. A network attacker who can observe traffic between Scrapy and S3 can read the bucket and key path, AWS Authorization header, X-Amz-Security-Token when temporary credentials are used, S3 object contents, and S3 response headers. An active man-in-the-middle attacker can also modify the plaintext S3 response body, status code, and headers before Scrapy processes them, causing scraped-data poisoning, poisoned exports, HTTP cache poisoning when caching is enabled, or influence over later crawl targets through forged redirects or attacker-controlled links. Users making S3-scheme requests with AWS credentials are affected. This issue is fixed in version 2.17.0.
— NVD
Affected Software
Remediation
Recommended actions to resolve this vulnerability, in priority order.
- Upgrade
Upgrade
pip/scrapyto a version that resolves this vulnerability.Fixed in 2.17 - Upgrade
Upgrade
scrapy/core/downloader/handlers/s3.py (S3DownloadHandler)to a version that resolves this vulnerability.Fixed in 2.17.0
Event History
Frequently Asked Questions
Who is affected by this issue?
Users of Scrapy versions before 2.17.0 that make S3-scheme requests using configured AWS credentials are affected. Exposure requires a network attacker able to observe or intercept traffic between the Scrapy process and S3.
Is the default S3 request behavior affected?
Yes. Before 2.17.0, S3DownloadHandler uses plaintext HTTP unless request.meta["is_secure"] is explicitly enabled. Requests explicitly configured with is_secure enabled do not use that plaintext behavior.
What can an attacker do if they can intercept the connection?
A passive attacker can read bucket and key paths, AWS Authorization headers, temporary-credential security tokens, object contents, and S3 response headers. An active attacker can alter response bodies, status codes, and headers, potentially poisoning scraped data, exports, or caches and influencing subsequent crawl targets through forged redirects or links.
What should be done if an immediate upgrade is not possible?
Explicitly enable request.meta["is_secure"] for S3-scheme requests to avoid the plaintext HTTP behavior. Upgrade to Scrapy 2.17.0 when possible, where the issue is fixed.