GHSA-6g59-gm2v-qhvq: Infoleak
Summary
The DNS-rebinding / redirect SSRF bypass in PRAI-05 is not limited to the httpx/urllib backend. When crawl4ai (headless Chromium via Playwright) is installed, webcrawl auto-selects provider=crawl4ai, and the headless browser re-resolves DNS and follows redirects on its own — with no per-connection SSRF guard. Runtime-confirmed read-back of the internal canary via both DNS rebinding and redirect (provider: "crawl4ai"). Status: runtime-confirmed addendum to PRAI-05.
Details
Affected component - Package praisonaiagents 4.6.63. Backend selected when crawl4ai+Playwright are installed. - Files: src/praisonai-agents/praisonaiagents/tools/webcrawltools.py (crawlwithcrawl4ai) and src/praisonai-agents/praisonaiagents/tools/crawl4aitools.py (standalone crawl wrappers).
Vulnerable code / root cause
Path: src/praisonai-agents/praisonaiagents/tools/webcrawltools.py
Function: crawlwithcrawl4ai
Snippet: python async with AsyncWebCrawler() as crawler: for url in urls: result = await crawler.arun(url=url) # headless browser: re-resolves DNS, follows redirects
Path: src/praisonai-agents/praisonaiagents/tools/crawl4aitools.py
Function: Crawl4AITools.crawl / crawl4ai
Snippet: python result = await crawler.arun(url=url, config=config) # no issafecrawlurl / no per-connect validation
Issue: the only SSRF check is the single pre-fetch issafecrawlurl() on the initial URL string in webcrawl() (see PRAI-05). The headless browser then resolves and connects independently and follows redirects in-browser — no resolved-IP pinning, no per-hop/per-connect validation. Input (urls) is attacker/agent-controlled; the sink is crawler.arun(url=...); the guard is bypassed by DNS rebinding (TOCTOU) and by redirects (followed in-browser).
Attack flow / Why bypassed / Security boundary Identical to PRAI-05: TOCTOU between the guard's resolution and the browser's connection; redirects followed in-browser; internal response returned as crawl content. See PRAI-05.
Proof of Concept
Environment crawl4ai + Playwright Chromium installed in a dedicated runtime container (127.0.0.1:18081), same controlled internal canary + rebinding DNS. Runnable assets: PraisonAI-Runtime-Repro\runtime-files\ (docker-compose.crawl.yml).
Steps to reproduce 1. SSRF-04-01-Crawl4AI-DNS-Rebind-Trigger → 127.0.0.1:18081: http POST /tool/webcrawl HTTP/1.1 Host: 127.0.0.1:18081 Content-Type: application/json
{"url":"http://rebind.lab:8081/secret"} 2. Redirect variant SSRF-04-03-Crawl4AI-Redirect-Readback: {"url":"http://redirector:8082/redirect-to-internal"}.
Expected result The crawl backend refuses internal destinations regardless of DNS timing/redirects.
Actual result HTTP 200, "provider":"crawl4ai", response content contains PRAISONAIINTERNALSECRETCANARY7f3a91 + FAKEINTERNALTOKENDONOTUSE7f3a91 for both the DNS-rebinding and the redirect payload.
Screenshots
Crawl4AI DNS rebinding read-back
The attacker-controlled rebind.lab URL is accepted by the Crawl4AI-backed webcrawl endpoint. PraisonAI returns the internal canary response body containing PRAISONAIINTERNALSECRETCANARY7f3a91.
<img width="1542" height="763" alt="01-Crawl4AI-Burp-Readback" src="https://github.com/user-attachments/assets/a03b3a1c-e382-4f67-8ea6-b766dceb4a69" />
Crawl4AI DNS rebinding runtime evidence
The runtime log shows the Crawl4AI/Chromium backend resolving rebind.lab to an internal Docker IP during the fetch phase. The internal canary receives GET /secret from a headless Chrome user agent.
<img width="1659" height="946" alt="02-Crawl4AI-DNS-Log-And-Internal-Hit" src="https://github.com/user-attachments/assets/d93db090-3e39-43d0-af3b-5d1d8f614db8" />
Crawl4AI redirect read-back
The attacker-controlled redirector URL is accepted by the Crawl4AI-backed endpoint. PraisonAI follows the redirect and returns the internal canary response body.
<img width="1544" height="771" alt="03-Crawl4AI-Redirect-Readback" src="https://github.com/user-attachments/assets/7d05c8aa-5bda-45bc-b94c-bef0bdcf055c" />
Crawl4AI redirect runtime evidence
The controlled redirector returns 302 -> http://internal-canary:8081/secret, and the internal canary receives GET /secret from the Crawl4AI/Chromium backend.
<img width="1590" height="920" alt="04-Crawl4AI-Redirect-Internal-Hit-Log" src="https://github.com/user-attachments/assets/dd9c0fec-3583-43ab-b83c-4470fa6aa409" />
Impact Same class as PRAI-05: read-back SSRF to internal/metadata services, now on the crawl4ai backend. Broadens the affected surface (the flaw is in the shared validate-without-pinning design, not one backend).
Suggested remediation Resolve-once + IP-pin + per-connect validation must also cover the crawl4ai backend (constrain the headless browser to the validated IP, or allowlist crawl destinations).
Affected Software
Remediation
Recommended actions to resolve this vulnerability, in priority order.
- Upgrade
Upgrade
pip/praisonaiagentsto a version that resolves this vulnerability.Fixed in 1.6.78 - Compensating control
Apply resolve-once, IP-pinning, and per-connect/per-hop SSRF validation to the Crawl4AI/Playwright backend; constrain the headless browser to the validated IP or allowlist crawl destinations so redirects and DNS rebinding cannot reach internal destinations.