CVE-2026-74714: bpf: tcp: Fix use-after-free in bpf_iter_tcp_established_batch()
In the Linux kernel, the following vulnerability has been resolved:
bpf: tcp: Fix use-after-free in bpfitertcpestablishedbatch()
reqskqueuehashreq() publishes a TCPNEWSYNRECV requestsock onto the ehash chain, drops the bucket lock, and only afterwards sets rskrefcnt to 3.
Lockless readers such as inetlookupestablished() handle this with refcountincnotzero(), but bpfitertcpestablishedbatch() uses plain sockhold() while holding the bucket lock, on the assumption that the lock guarantees skrefcnt > 0. That assumption does not hold for requestsock:
CPU 0 CPU 1 ----- ----- tcpconnrequest() reqskqueuehashreq() inetehashinsert(req) spinlock(bucket) sknullsaddnodercu(req) // rskrefcnt == 0 spinunlock(bucket) bpfitertcpestablishedbatch() spinlock(bucket) sockhold(req) <-- addition on 0 spinunlock(bucket) refcountset(&req->rskrefcnt, 3) // clobbers saturated value
which surfaces as:
refcountt: addition on 0; use-after-free. WARNING: lib/refcount.c:25 at refcountwarnsaturate+0x48/0x90, CPU#1 Call Trace: bpfitertcpestablishedbatch+0x14e/0x170 bpfitertcpbatch+0x53/0x200 bpfitertcpseqnext+0x27/0x70 bpfseqread+0x107/0x410 vfsread+0xb9/0x380
The iterator's stolen reference is lost when the publishing CPU's refcountset() overwrites the count, leaving the socket one reference short. When the last legitimate owner drops its reference the reqsk is freed while still reachable, leading to use-after-free.
This reproduces in seconds with tcpsyncookies=0, a handful of threads doing connect()/close() to a local listener while others read an iter/tcp link in a tight loop.
Use refcountincnotzero() and skip the socket on failure. A skipped socket is still part of the bucket, so keep counting it in expected. The reallocations are sized from expected, and a request sock whose refcount gets published while the lock is held across the last realloc must already have room.
A skipped socket is counted in expected but never batched, so endsk can be short of expected on a batch that is actually complete. Decide completeness by whether the walk left any socket behind instead. The WARN after the locked realloc checks the same, replacing an endsk == expected check that could not hold on that path since commit cdec67a489d4 ("bpf: tcp: Make sure iter->batch always contains a full bucket snapshot").
If every matching socket in a bucket is mid-init (refcount 0), endsk stays 0. Advance to the next bucket rather than returning a batch entry that was never filled this round.
Affected Software
Event History
Frequently Asked Questions
What conditions are required to trigger the issue?
A TCP_NEW_SYN_RECV request_sock must be published on the TCP established hash chain while its rsk_refcnt is still zero, and bpf_iter_tcp_established_batch() must inspect it during that interval. The race occurs when the iterator calls sock_hold() under the bucket lock before reqsk_queue_hash_req() sets rsk_refcnt to 3.
How can administrators recognize that the race has occurred?
The kernel can report a refcount warning such as "refcount_t: addition on 0; use-after-free." The reported call trace includes bpf_iter_tcp_established_batch(), bpf_iter_tcp_batch(), bpf_iter_tcp_seq_next(), bpf_seq_read(), and may include vfs_read().
Which code behavior is unsafe in the affected path?
bpf_iter_tcp_established_batch() uses plain sock_hold() for entries found while holding the ehash bucket lock. For a newly published request_sock, the bucket lock does not guarantee that sk_refcnt is nonzero, because the reference count is initialized after the entry is published.