Summary
Several remote-content download paths in Pydantic AI buffered the entire HTTP response body into memory before enforcing any size limit. An application that exposes the local web-fetch tool (webfetchtool, or the WebFetch capability's local fallback) to untrusted prompts can be driven to fetch an attacker-chosen URL that streams a very large body, exhausting process memory and crashing the worker. The same unbounded buffering applied to FileUrl media downloads (ImageUrl, DocumentUrl, VideoUrl, AudioUrl).
This is an availability issue only. SSRF protections (scheme allowlist, private-IP and cloud-metadata blocking) are unaffected; there is no confidentiality or integrity impact.
Details
The download helpers read the full response body before applying content-size controls, so an existing text-length limit only truncated after the whole body was already in memory, and media downloads had no wire-level cap at all. A single large response could grow process memory without bound .
Who Is Affected
You are affected if your application registers the local web-fetch tool (or relies on the WebFetch capability's local fallback) and exposes the agent to untrusted prompts, or if it downloads large remote FileUrls influenced by untrusted input. Applications that only fetch developer-controlled URLs are not exposed to the model-chosen attack path.
Remediation
Upgrade to 2.24.0 or later (v2) or 1.107.2 or later (v1). Patched versions enforce a default 50 MiB cap on web-fetch and FileUrl downloads while streaming; pass None to the limit to restore the previous unbounded behavior.
Credits
Identified during internal review of media-download hardening.
Summary
Pydantic AI's OpenTelemetry instrumentation supports InstrumentationSettings(includecontent=False) to exclude message content — user prompts, model completions, tool call arguments and responses — from exported telemetry, so that agents can be monitored without sending sensitive content to the observability backend.
Retry prompts — the feedback messages Pydantic AI sends back to the model when its output fails validation or an output validator raises ModelRetry — were not covered by this setting when they were not associated with a tool call, as with structured output modes that don't use tool calls (e.g. NativeOutput or PromptedOutput) and output validators applied to text output. Their full content was recorded on model request spans even with includecontent=False. Since validation feedback can quote invalid values from the model's response, message content that the setting was configured to withhold could reach the telemetry backend.
Details
Retry prompts without an associated tool call were serialized into span message attributes (such as genai.input.messages and pydanticai.allmessages) as regular text parts, and this path did not apply the includecontent check that every other message part applies. Retry prompts tied to a tool call, and all other message content, were redacted correctly.
Impact
This does not grant an attacker any new access to the agent or its data: the content is only visible to whoever can read the exported telemetry. The exposure matters when traces are exported to a destination whose audience is broader or less trusted than the agent's own data — a third-party observability vendor or a shared backend — which is the scenario includecontent=False exists for. The disclosed content is limited to retry feedback: validation error details, which can quote invalid values from the model's output, or output-validator ModelRetry messages.
Who Is Affected
You are affected if you enabled instrumentation with includecontent=False and your agent can produce retry prompts outside of tool calls: it uses a structured output mode that doesn't rely on tool calls (such as NativeOutput or PromptedOutput), or output validators on text output. You are not affected if you don't set includecontent=False, or if all of your agent's retries come from tool calls (including the default tool-based structured output mode), which were redacted correctly.
Remediation
Upgrade to a patched version; retry prompt content now honors includecontent=False like all other message content. If you rely on this setting, consider reviewing previously exported traces for retry prompt content that should not have been recorded.
Workaround for Unpatched Versions
If you cannot upgrade, scrub or drop the message attributes (genai.input.messages, genai.output.messages, pydanticai.allmessages) in your telemetry pipeline (for example with an OpenTelemetry Collector processor), or use tool-based structured output modes, whose retry feedback honors includecontent=False.
Credits
Reported privately by the University of Sydney Security Research Team (Liyi Zhou, Ziyue Wang, Strick, Maurice, Chenchen Yu), and later independently discovered and reported publicly, with a fix, by @sean-kim05.
Summary
The Pydantic AI development web chat UI (Agent.toweb(), clai web) does not validate the Host header of incoming requests. A website a developer visits can use DNS rebinding to make requests to a chat UI running on that developer's machine appear same-origin to the browser, causing the served agent to run and to execute its tools with the privileges and credentials of the local process.
Details
Once a name the attacker controls resolves to the loopback address, the browser treats the request as same-origin, so neither an Origin check nor a CSRF token constrains it — a same-origin page can read the served UI and any token in it.
Binding the web UI to localhost — the default — does not prevent this.
Impact
Applications and developers serving an agent through Agent.toweb() or clai web. The consequences depend on the tools the served agent exposes, and can include data disclosure as well as unwanted tool side effects.
Current browser protections reduce but do not remove this exposure: Chromium's Local Network Access gates loopback subresource requests, but does not cover top-level navigations, and Safari does not implement it.
Mitigation
Upgrade to pydantic-ai/pydantic-ai-slim >= 2.30.0, or >= 1.107.5 on the v1 maintenance line.
The fix validates the Host header and rejects anything other than localhost, a loopback/LAN IP address, or an explicitly allowed host, responding 421 Misdirected Request otherwise. If you serve the web chat UI under a real hostname — behind a reverse proxy, tunnel, or similar — name it explicitly:
python app = agent.toweb(allowedhosts=['ui.example.com'])
Summary
Pydantic AI's OpenTelemetry instrumentation supports InstrumentationSettings(includecontent=False) to exclude message content — user prompts, model completions, tool call arguments and responses, and any other message content — from exported telemetry, so that agents can be monitored without sending sensitive content to the observability backend.
The message attributes honoured the setting. Several other parts of the same spans did not, and exported content regardless of it:
- Exceptions. Recorded as OpenTelemetry exception events, with their message and stack trace, on tool, agent run, model request, embedding and image generation spans, and on realtime session spans. The message of the exception carrying a tool's ModelRetry or ToolFailed feedback is the text sent back to the model, which can quote tool arguments and results. When a tool ran out of retries, the error that ended the run was chained from the last retry, so its stack trace repeated that text on the run span. Exceptions raised by tool code were recorded with their full message, as were failures of a model request, embedding or image generation, whose messages carry the provider's error response body. - The error status description. The OpenTelemetry SDK describes a span's ERROR status with the exception's type and message, so the message reached the backend a second time even where the event itself was redacted. - Instructions and the prompted-output template. The modelrequestparameters attribute serialized the request's instruction parts in full, exporting the agent's instructions — both static and those built from dependencies at request time — along with the template used to prompt for structured output.
Details
Content reaches these places because of what they carry rather than through any single missing check, which is why the setting was applied correctly to the message attributes (genai.input.messages, genai.tool.call.arguments, genai.tool.call.result, and so on) while these paths were not: the exception recording and the status description came from the OpenTelemetry SDK's defaults, and the instructions and output template were serialized as part of the request's parameters rather than treated as content.
Impact
This does not grant an attacker any new access to the agent or its data: the content is only visible to whoever can read the exported telemetry. The exposure matters when traces are exported to a destination whose audience is broader or less trusted than the agent's own data — a third-party observability vendor or a shared backend — which is the scenario includecontent=False exists for.
Who Is Affected
You are affected if you enabled instrumentation with includecontent=False. Instructions are exported on every successful run, so no failure is needed to be affected; the exception paths additionally require a tool to raise, or a model, embedding or image generation request to fail. Realtime sessions are affected the same way, through the errors their provider connections report. GHSA-3gh4-cghq-f8v4 stated that retries coming from tool calls were redacted correctly; that was true of the retry prompt in the message attributes but not of the exception event on the tool span. You are not affected if you don't set includecontent=False.
Remediation
Upgrade to a patched version. With includecontent=False, exception events now record only the exception type, error statuses carry no description, and modelrequestparameters is exported without its instruction content or output template. If you rely on this setting, consider reviewing previously exported traces for exception events, for error status descriptions, and for modelrequestparameters.
Workaround for Unpatched Versions
If you cannot upgrade, scrub the exception.message and exception.stacktrace event attributes, the error status description, and the instruction parts and promptedoutputtemplate within modelrequestparameters, on spans emitted by Pydantic AI in your telemetry pipeline (for example with an OpenTelemetry Collector processor). Setting includemodelrequestparameters=False alongside includecontent=False removes both, at the cost of the rest of that attribute.
Credits
Reported privately by @BrianWillows, whose report covered the exception events on tool and agent run spans. The remaining paths were found while fixing it, by a test that drives content through every channel at once and checks the whole exported trace.
Summary
The local web-fetch tool (webfetchtool, also used as the WebFetch capability's local fallback) processed responses with several steps whose running time grows quadratically with the size of certain server-controlled inputs, and ran them on the event loop: decoding the body with whichever charset the server declared, extracting the page title with a backtracking regular expression, and converting the HTML to markdown. An application that exposes this tool to untrusted prompts can be steered to fetch an attacker-controlled page of a megabyte or two that blocks the event loop for minutes, stalling every other coroutine in the process — other agent runs, other requests being served — for the duration.
This is an availability issue only. SSRF protections and the download size limit introduced in GHSA-v2xh-2vp8-57h8 are unaffected; that limit bounds how much is downloaded, not how long the response takes to process.
Details
Title extraction used a backtracking pattern over the raw response body, so a body made of repeated unterminated tag openings cost time proportional to the square of its size. The HTML-to-markdown conversion had the same shape in three of its steps: normalizing whitespace, stripping preformatted blocks, and numbering ordered lists all took time proportional to the square of a run of spaces or a list's length. All of it ran on the event loop, and the regex steps hold the interpreter lock even when moved off it, so the whole process paid for the size of a server-controlled response.
The response body was also decoded on the event loop with the codec named by the charset parameter of the response's Content-Type, looked up in Python's codec registry. That registry includes punycode, whose decoder takes time proportional to the square of its input: a response of about one megabyte labelled charset=punycode blocked the event loop for roughly half a minute, with no HTML required. The registry also includes codecs that aren't text encodings at all, such as rot13 and base64codec; a response labelled with one of those raised an unexpected exception out of the tool, aborting the agent run that fetched it.
Separately, the HTML-to-markdown conversion recursed once per nested element, so a page nested a few hundred elements deep raised a RecursionError out of the tool, aborting the agent run that fetched it. A JSON response nested deeper than the interpreter allows did the same. These only affect that one run.
Who Is Affected
You are affected if your application registers the local web-fetch tool (or relies on the WebFetch capability's local fallback) and exposes the agent to untrusted prompts. alloweddomains narrows the exposure to pages on those domains but does not remove it. Applications that only fetch developer-controlled URLs are not exposed to the model-chosen attack path.
Remediation
Upgrade to a patched version. The title is now found with a single linear scan, the conversion steps above run in linear time, and decoding, title extraction and conversion all run in a worker thread. A charset naming a codec that isn't a text encoding, and a page too deeply nested to convert, are reported back to the model as a failed fetch instead of aborting the run; a JSON body too deeply nested to parse is returned as plain text.
Credits
Reported privately by @BrianWillows, whose report covered the quadratic title extraction. The response decoding, the codecs that are not text encodings, and the quadratic steps in the HTML-to-markdown conversion were found while fixing it.
Summary
When an application using Pydantic AI opts a URL into local network access — either a FileUrl with forcedownload='allow-local', or webfetchtool(allowlocalurls=True) — the cloud-metadata blocklist could be bypassed by appending an IPv6 zone identifier to a metadata address (for example fd00:ec2::254%251). The host ignores the zone identifier on a destination that is not link-local and delivers the request to the metadata endpoint anyway, exposing cloud IAM short-term credentials.
This is an incomplete fix of GHSA-cqp8-fcvh-x7r3 / CVE-2026-46678 and GHSA-cg7w-rg45-pc59 / CVE-2026-48782, themselves follow-ups to CVE-2026-25580. The parent advisory's remediation guaranteed that cloud metadata endpoints are always blocked, even with local access allowed. That guarantee did not hold for zone-scoped spellings of the IPv6 metadata endpoints.
Details
The cloud-metadata guard compared IPv6 addresses against its blocklist by set membership. Python includes the zone identifier in IPv6Address equality and hashing, so a zone-scoped spelling of a blocked address did not match, while the network stack ignores the zone identifier for a destination that is not link-local. The private-range checks, and the IPv4 and transition-form metadata checks, were already unaffected, because they compare by network containment and by packed bytes respectively.
Only the IPv6 cloud metadata endpoints were reachable this way, so the issue requires an IPv6-enabled environment — for example AWS EC2 or EKS with IPv6, GCP IPv6-only instances, or Scaleway.
Who Is Affected
You are affected only if your application opts a URL that is, or could be, influenced by untrusted input into local network access, through either:
- a FileUrl (ImageUrl, AudioUrl, VideoUrl, DocumentUrl) with forcedownload='allow-local'; or - webfetchtool(allowlocalurls=True), where the model chooses the URL.
Both are off by default.
You are not affected through the FileUrl path if you use any of the bundled integrations to ingest user input, because they do not propagate forcedownload from external data:
- Agent.toweb / clai web - VercelAIAdapter - AGUIAdapter / Agent.toagui
webfetchtool is configured by your own application, so a client cannot turn on allowlocalurls.
Applications that only download from developer-controlled URLs are not affected.
Remediation
Upgrade to a patched version. The cloud-metadata and private-IP checks now drop an IPv6 zone identifier before evaluating the address, so every blocklist comparison is made on the address itself. A zone identifier is still carried on the connection, so legitimate link-local fetches under local network access continue to work.
Workaround for Unpatched Versions
Avoid opting into local network access — forcedownload='allow-local' or webfetchtool(allowlocalurls=True) — on any URL that could be influenced by untrusted input. If you must, reject URL hosts containing % before constructing the FileUrl or configuring the tool.
Credits
Reported by @euriconicacio.
Pydantic AI is a Python agent framework for building applications and workflows with Generative AI. From 1.77.0 until 1.107.6 and 2.44.0, the local webfetchtool and the WebFetch local fallback compare blockeddomains entries with a URL hostname before both values are normalized to the form used by getaddrinfo. An attacker-influenced model can use an equivalent IDNA spelling, non-ASCII label separator, case variation, or trailing root label that resolves to a blocked host but does not match the configured string, causing the application to fetch that host with its own privileges. alloweddomains fails closed for unmatched spellings, and private-IP and cloud-metadata protections remain effective. This issue is fixed in versions 1.107.6 and 2.44.0.
Summary
Applications using Pydantic AI's local web-fetch tool can experience excessive CPU and memory use when it converts attacker-controlled HTML. An agent must fetch the affected page; provider-native web fetching is not affected.
Details
Nested block elements cause HTML-to-Markdown conversion to reprocess accumulated text at each level and can greatly expand the intermediate output. The response-body limit bounds downloaded bytes, while the returned-content limit is applied only after conversion. On current releases, conversion runs in a worker thread but can still consume substantial resources and delay other work in the process. Older releases performed conversion on the event loop.
Mitigation
Upgrade to a patched release of pydantic-ai or pydantic-ai-slim. If you cannot upgrade yet, avoid using local web fetching for attacker-controlled HTML.
Pydantic AI is a Python agent framework for building applications and workflows with Generative AI. From 2.10.0 until 2.53.0, streamed requests made through ConcurrencyLimitedModel or limitmodelconcurrency can retain shared concurrency slots because anyio.CapacityLimiter associates an acquired slot with the borrowing task while streaming cleanup can run in a different task. Early stream termination, cancellation, consumer exceptions, or complete streamtext() consumption with debounceby=0.1 can therefore leave capacity occupied, eventually preventing later requests that share the long-lived limiter from proceeding and causing a denial of service. Agent-level maxconcurrency and non-streaming model requests are not affected. This issue is fixed in version 2.53.0.
Summary
When an application using Pydantic AI opts a URL into forcedownload='allow-local' (which disables the default block on private/internal IPs) and runs on a network that routes the affected IPv6 transition forms (NAT64- or ISATAP-configured networks), the cloud-metadata blocklist could be bypassed by encoding the metadata IP in an IPv6 transition form that the previous fix did not decode — IPv4-compatible IPv6 (::a.b.c.d), the NAT64 RFC 8215 local-use prefix (64:ff9b:1::/48), operator-chosen NAT64 prefixes, or ISATAP. The IPv6 wrapper is then delivered to the underlying IPv4 metadata endpoint, exposing cloud IAM short-term credentials.
The bypass is exploitable only in environments whose network actually routes these forms — NAT64-configured networks (IPv6-only or dual-stack-with-NAT64 deployments, including some Kubernetes setups) for the NAT64 variants, or networks with an ISATAP tunnel for ISATAP. A standard dual-stack cloud VM or container does not route them and is not affected in practice. The IPv4-compatible and Teredo variants are deprecated and addressed as defense-in-depth.
This is an incomplete fix of GHSA-cqp8-fcvh-x7r3 / CVE-2026-46678 (itself a follow-up to CVE-2026-25580). The prior remediation decoded only IPv4-mapped IPv6, 6to4, and the NAT64 well-known prefix; the metadata guarantee did not hold for the remaining transition forms.
Severity
MEDIUM — CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:C/C:H/I:N/A:N = 6.8
Same impact metrics and narrow attack surface as the parent advisory (AC:H): exploitation requires the application to have opted into allow-local on a URL influenced by untrusted input, and the NAT64/ISATAP variants additionally require the deployment network to route those forms.
CWE-918: Server-Side Request Forgery (SSRF)
Affected Versions
| Package | Vulnerable | Patched | |---|---|---| | pydantic-ai | >= 1.56.0, < 1.102.0; >= 2.0.0b1, < 2.0.0b3 | 1.102.0; 2.0.0b3 | | pydantic-ai-slim | >= 1.56.0, < 1.102.0; >= 2.0.0b1, < 2.0.0b3 | 1.102.0; 2.0.0b3 |
These transition forms have not been decoded since SSRF protection was introduced in 1.56.0.
Who Is Affected
Users are affected only if their application explicitly opts a FileUrl (ImageUrl, AudioUrl, VideoUrl, DocumentUrl) into forcedownload='allow-local' on a URL that is, or could be, influenced by untrusted input.
Beyond that precondition, the affected encodings only reach a metadata endpoint in environments whose network actually routes them. The broadly-routable IPv4-mapped form was addressed in 1.99.0 (CVE-2026-46678); the additional forms addressed here require a NAT64-configured network (IPv6-only or dual-stack-with-NAT64 deployments, including some Kubernetes setups) for the NAT64 variants, or an ISATAP tunnel for the ISATAP variant. The IPv4-compatible and Teredo forms are deprecated and not routed by modern stacks; they are addressed as defense-in-depth. Most deployments on a standard dual-stack cloud VM or container are therefore not exploitable in practice, but the fix restores the "always blocked" guarantee for the environments that are.
Users are not affected if they use any of the bundled integrations to ingest user input, because they do not propagate forcedownload from external data:
- Agent.toweb / clai web - VercelAIAdapter - AGUIAdapter / Agent.toagui
Applications that only download from developer-controlled URLs are not affected.
Remediation
Upgrade to 1.102.0 or later (or 2.0.0b3 or later on the 2.0 pre-release line). The cloud-metadata and private-IP blocklists now decode the embedded IPv4 of every standardized IPv6 transition form before evaluating it — IPv4-mapped, IPv4-compatible, 6to4, NAT64 across all prefix lengths (including the RFC 8215 local-use prefix and operator-chosen prefixes), ISATAP, and Teredo. The set of always-blocked cloud metadata/credential endpoints has also been expanded across providers.
Workaround for Unpatched Versions
Avoid passing forcedownload='allow-local' on any URL that could be influenced by untrusted input. If developers must, resolve the hostname themselves and validate the result against their own metadata blocklist — including IPv6 transition forms — before constructing the FileUrl.
Credits
Reported by @SnailSploit.