GHSA-6fqq-452j-qhrp: High severity pip/pydantic-ai-slim vulnerability
This issue was posted by Codex Desktop using gpt-6.1-sol on behalf of David.
Summary
Applications that wrap a model with ConcurrencyLimitedModel or limitmodelconcurrency can permanently lose shared concurrency capacity when a streamed request releases its slot from a different task than the one that acquired it. This can happen when a stream ends early, and also when a stream is fully consumed using the default streamtext() debouncing.
In an application that exposes an affected streaming endpoint to network clients and shares a long-lived model limiter across requests, a client can repeatedly start a stream and disconnect. The completed requests retain their slots, eventually preventing subsequent requests that share the limiter from proceeding.
Agent-level maxconcurrency and non-streaming model requests are not affected by this defect.
Details
The built-in limiter uses anyio.CapacityLimiter, which associates each acquired slot with its borrowing task. Pydantic AI's streaming lifecycle can acquire the slot on the task consuming the stream and run cleanup on another internal task. The limiter rejects that release, so the slot remains occupied even after the request has ended. Cleanup can raise a RuntimeError; a later request on the borrowing task can also fail because that task still holds a slot.
Early termination includes stopping iteration, a consumer exception, and cancellation. Fully consuming streamtext() with its default debounceby=0.1 can also reach the cross-task release path. Fully consumed streams must therefore not be assumed safe.
Mitigation
Upgrade to a patched release of pydantic-ai or pydantic-ai-slim. If you cannot upgrade yet, use the agent-level maxconcurrency setting instead of a concurrency-limited model, or avoid streaming runs through a concurrency-limited model.
Affected Software
Remediation
Recommended actions to resolve this vulnerability, in priority order.
- Upgrade
Upgrade
pip/pydantic-ai-slimto a version that resolves this vulnerability.Fixed in 2.53.0 - Upgrade
Upgrade
pip/pydantic-aito a version that resolves this vulnerability.Fixed in 2.53.0 - Configuration
Use the agent-level max_concurrency setting instead of wrapping the model with ConcurrencyLimitedModel or limit_model_concurrency.
Pydantic AI agent max_concurrency = agent-level max_concurrency - Compensating control
Avoid streaming runs through a concurrency-limited model until the affected pydantic-ai or pydantic-ai-slim component is upgraded.
Event History
Frequently Asked Questions
Which deployments are realistically exposed to service disruption?
Applications are exposed when they provide an affected streaming endpoint to network clients and use a long-lived shared limiter created with ConcurrencyLimitedModel or limit_model_concurrency. Agent-level max_concurrency and non-streaming model requests are not affected.
What must an attacker do to exhaust capacity?
An unauthenticated network client can repeatedly begin streamed requests and disconnect. Each affected request can retain a shared concurrency slot, eventually preventing later requests using that limiter from proceeding.
Does a stream have to be abandoned early for this to occur?
No. Early stream termination can trigger the issue, but a stream fully consumed through the default stream_text() debouncing behavior can also release its slot from a different task and leave capacity occupied.