GHSA-qg25-6gc4-48mg: Infoleak

Published Oct 8, 2026
·
Updated

Summary

PraisonAI's webcrawl agent tool performs a server-side HTTP fetch of an agent/attacker-influenced URL. SSRF is meant to be prevented by issafecrawlurl(), which resolves the hostname and rejects private/loopback/link-local IPs at validation time. The validated value is the URL string (not a pinned IP); the fetch backend then re-resolves the hostname at connection time. Because validation and connection perform two independent DNS resolutions, a DNS-rebinding domain that returns a public IP during validation and an internal IP during the fetch fully bypasses the guard, and the internal HTTP response body is returned to the caller.

This is SSRF with internal response disclosure (read-back) — not blind SSRF. Runtime-confirmed against PraisonAI 4.6.63; the crawl response returned the controlled internal markers PRAISONAIINTERNALSECRETCANARY7f3a91 / FAKEINTERNALTOKENDONOTUSE7f3a91. Severity High. Reachable by any actor who can influence the URL an agent crawls (e.g. a chat/bot/agent surface).

Details

Affected component - Package: praisonaiagents (PraisonAI), version 4.6.63. - File: src/praisonai-agents/praisonaiagents/tools/webcrawltools.py; tool webcrawl / crawlweb (part of the default bot tool set).

Vulnerable code / root cause

Code point 1 — check-time-only DNS validation, no IP pinning

Path: src/praisonai-agents/praisonaiagents/tools/webcrawltools.py

Function: issafecrawlurl

Snippet: python for info in socket.getaddrinfo(hostname, None): # resolve at CHECK time ip = ipaddress.ipaddress(info[4][0]) if (ip.isloopback or ip.isprivate or ip.islinklocal or ip.ismulticast or ip.isunspecified): return False return True Issue: the guard validates the hostname by resolving it once at check time. It does not pin the resolved IP and does not return/forward that IP to the HTTP client. Any later resolution can differ.

Code point 2 — guard runs, then the URL string is handed to the backend

Function: webcrawl

Snippet: python for u in rawurllist: if issafecrawlurl(u): # validate the URL string urllist.append(u) ... results = crawlwithhttpx(urllist) # or crawlwithcrawl4ai(urllist) Issue: attacker-controlled input (urls) is validated as a string; the backend then fetches that string and re-resolves DNS independently of the guard. There is no shared, pinned IP between check and fetch.

Code point 3 — crawlwithhttpx backend re-resolves (redirect re-validation does not stop rebinding)

Function: crawlwithhttpx

Snippet: python with httpx.Client(followredirects=False, timeout=30.0) as client: for in range(maxredirects + 1): if not issafecrawlurl(current): # re-resolves hostname (CHECK) raise ValueError("Redirect target failed SSRF validation") response = client.get(current) # resolves AGAIN at CONNECT Issue: even with per-hop redirect re-validation, issafecrawlurl(current) and client.get(current) are two separate DNS resolutions of the same hostname. A rebinding domain answers public to the check and internal to the connect → TOCTOU bypass. No IP pinning.

Code point 4 — urllib fallback (same function), no per-hop guard

Snippet: python import urllib.request with urllib.request.urlopen(url, timeout=30) as response: # re-resolves + auto-follows redirects content = response.read().decode('utf-8', errors='ignore') Issue: when httpx is not installed, this fallback inside crawlwithhttpx fetches the URL and auto-follows redirects with no per-hop/per-connect validation. (Results from this function are labelled "provider": "httpx" regardless of which path runs.)

Code point 5 — crawl4ai/Chromium backend (confirmed addendum)

The crawl4ai backend (crawlwithcrawl4ai → crawler.arun(url=url), headless Chromium) is also runtime-confirmed affected (browser re-resolves DNS / follows redirects with no per-connect guard). To keep this report focused on the webcrawl SSRF guard, the backend-specific evidence is in SSRF-04Crawl4AISSRFBackendAddendum.md.

Attack flow 1. Attacker controls a hostname (e.g. rebind.lab) whose authoritative DNS rebinds. 2. Lookup #1 (the guard) → a public IP → issafecrawlurl() returns true. 3. The backend re-resolves → the attacker's DNS now answers an internal/private IP (cloud metadata, loopback, internal service). 4. The backend connects to the internal service and returns its body to the caller → internal data disclosure.

Why existing protection is bypassed - The guard validates the hostname, not a pinned IP; check and connect resolve independently → DNS rebinding (TOCTOU) defeats it on every backend. - Redirect re-validation (httpx path) re-checks the hostname but still re-resolves at connect, so it does not stop rebinding; the urllib fallback and crawl4ai backends have no per-hop guard at all.

Security boundary The server-side fetch reaches internal/loopback/metadata services not exposed to the attacker and returns their content (CVSS Scope: Changed). Reachable wherever an agent can be induced to crawl an attacker-supplied URL (PR:L). An unauthenticated single-request path to webcrawl read-back was not found in 4.6.63 (so PR:N / Critical is not claimed).

Proof of Concept

Environment Real PraisonAI 4.6.63 in a local Docker runtime; a controlled internal canary service (Docker-internal only, not published) returns synthetic markers; a controlled DNS responder implements rebinding for rebind.lab. No public host / real metadata / real secret. Runnable assets: PraisonAI-Runtime-Repro\runtime-files\.

Steps to reproduce 1. Burp Repeater tab PRAI-05-01-DNS-Rebind-Trigger → 127.0.0.1:18080: http POST /tool/webcrawl HTTP/1.1 Host: 127.0.0.1:18080 Content-Type: application/json

{"url":"http://rebind.lab:8081/secret"} 2. Send (PRAI-05-02-DNS-Rebind-Secret-Readback captures the response). If a send returns the "blocked" error, the rebinding DNS auto-resets (~3s) — resend. 3. Redirect variant: PRAI-05-03-Redirect-Trigger / PRAI-05-04-Redirect-Secret-Readback send {"url":"http://redirector:8082/redirect-to-internal"}.

Expected result A safe SSRF guard refuses destinations that resolve to internal/private IPs regardless of DNS timing or redirects, and does not return internal content.

Actual result HTTP 200 with the internal body in the crawl result. Primary evidence is the provider: "httpx" backend returning the internal canary via DNS rebinding: json {"inputurl":"http://rebind.lab:8081/secret", "result":{"content":"{ ... \"secret\": \"PRAISONAIINTERNALSECRETCANARY7f3a91\", \"token\": \"FAKEINTERNALTOKENDONOTUSE7f3a91\" ... }","provider":"httpx"}} The redirect variant returns the same internal markers via a redirect chain (provider: "httpx").

Screenshots

DNS rebinding read-back

The attacker-controlled rebind.lab URL is accepted by webcrawl, and the PraisonAI response contains the internal canary response body.

<img width="1543" height="785" alt="01-DNS-Rebind-Burp-Readback" src="https://github.com/user-attachments/assets/e86e95dd-3d3a-4bef-b1c6-cb897234a212" />

DNS rebinding runtime evidence

The runtime log shows rebind.lab first resolving to an allowed/public IP during validation (guard-pass), then resolving to an internal Docker IP during the actual fetch (fetch-hit). The internal canary receives GET /secret from the PraisonAI container.

<img width="1654" height="828" alt="02-DNS-Rebind-DNS-Log-And-Internal-Hit" src="https://github.com/user-attachments/assets/f1d28046-de34-48f0-9bee-bbe26e598d5f" />

Redirect-based SSRF read-back

The attacker-controlled redirector URL is accepted by webcrawl. PraisonAI follows the redirect and returns the internal canary response body containing PRAISONAIINTERNALSECRETCANARY7f3a91.

<img width="1540" height="772" alt="03-Redirect-Burp-Readback" src="https://github.com/user-attachments/assets/79eec159-4b9f-44be-a9eb-b15aba35ead8" />

Redirect chain runtime evidence

The controlled redirector returns 302 -> http://internal-canary:8081/secret, and the internal canary receives GET /secret, confirming that the server-side client followed the redirect into the internal network.

<img width="1637" height="894" alt="04-Redirect-Internal-Hit-Log" src="https://github.com/user-attachments/assets/1c503e30-9d01-4d32-8658-497f755a969e" />

Reproduction assets

The attached archive contains the local Docker runtime used to reproduce the issue with controlled canary services only. It does not contain real secrets, real cloud metadata access, or third-party API keys.

PraisonAI-Runtime-Repro.zip

Impact SSRF against internal/loopback/cloud-metadata endpoints with disclosure of internal HTTP responses (read-back) to the attacker. Bypasses the project's SSRF protection on every fetch backend.

Suggested remediation 1. Resolve the host once, reject all returned records that are private/loopback/link-local/ULA/CGNAT/metadata, then connect to that exact validated IP (pin it; send the original Host). Do not let the HTTP client / browser re-resolve. 2. Apply the same validation + IP pinning to every backend (httpx, urllib fallback, crawl4ai) and every redirect hop. 3. Disable automatic redirect following (or cap + re-validate each hop with pinning). 4. Treat IPv4-mapped IPv6, decimal/octal/hex IPs, and CGNAT/non-global ranges as unsafe.

Affected Software

1 affected componentFixes available
pip/praisonaiagents<=1.6.77
1.6.78

Remediation

Recommended actions to resolve this vulnerability, in priority order.

  1. Upgrade

    Upgrade pip/praisonaiagents to a version that resolves this vulnerability.

    Fixed in 1.6.78
  2. Compensating control

    For the httpx, urllib fallback, and crawl4ai/Chromium backends, resolve each hostname once, reject IPv4-mapped IPv6, decimal/octal/hex IPs, private, loopback, link-local, ULA, CGNAT, multicast, unspecified, and cloud-metadata ranges, then connect to that exact validated IP without allowing the client or browser to re-resolve DNS; preserve the original Host header. Disable automatic redirect following, or cap redirects and repeat the same validation and IP pinning for every redirect hop.

Event History

Oct 8, 2026
Advisory Published
via GitHub·04:36 PM
Data Sourced
via GitHub·04:36 PM
DescriptionSeverityWeaknessAffected Software

Frequently Asked Questions

1

Who can realistically trigger this issue?

Any actor who can influence a URL that a PraisonAI agent passes to the web_crawl tool can trigger it. This includes users of chat, bot, or agent surfaces that allow attacker-controlled URLs to be crawled.

2

What does an attacker need to bypass the URL safety check?

The attacker needs a DNS-rebinding hostname that resolves to a public IP address during validation and to an internal IP address when the fetch backend connects. The separate DNS lookups allow the request to reach the internal address.

3

Is the SSRF blind, or can the attacker obtain data from internal services?

It is not blind SSRF. The internal HTTP response body is returned to the caller, enabling internal response disclosure.

Contact

SecAlerts Pty Ltd.
132 Wickham Terrace
Fortitude Valley,
QLD 4006, Australia
info@secalerts.co
By using SecAlerts services, you agree to our services end-user license agreement. This website is safeguarded by reCAPTCHA and governed by the Google Privacy Policy and Terms of Service. All names, logos, and brands of products are owned by their respective owners, and any usage of these names, logos, and brands for identification purposes only does not imply endorsement. If you possess any content that requires removal, please get in touch with us.
© 2026 SecAlerts Pty Ltd.
ABN: 70 645 966 203, ACN: 645 966 203