In the Linux kernel, the following vulnerability has been resolved:
vfio/pci: Check BAR resources before exporting a DMABUF
A DMABUF exports access to BAR resources and, although they are requested at startup time, we need to ensure they really were reserved before exporting. Otherwise, it's possible to access unreserved resources through the export.
Add a check to the DMABUF-creation path.
In the Linux kernel, the following vulnerability has been resolved:
KVM: x86: Fix shadow paging use-after-free due to unexpected GFN
The shadow MMU computes GFNs for direct shadow pages using sp->gfn plus the SPTE index. This assumption breaks for shadow paging if the guest page tables are modified between VM entries (similar to commit aad885e77496, "KVM: x86/mmu: Drop/zap existing present SPTE even when creating an MMIO SPTE", 2026-03-27). The flow is as follows:
- a PDE is installed for a 2MB mapping, and a page in that area is accessed. KVM creates a kvmmmupage consisting of 512 4KB pages; the kvmmmupage is marked by FNAME(fetch) as direct-mapped because the guest's mapping is a huge page (and thus contiguous).
- the PDE mapping is changed from outside the guest.
- the guest accesses another page in the same 2MB area. KVM installs a new leaf SPTE and rmap entry; the SPTE uses the "correct" GFN (i.e. based on the new mapping, as changed in the previous step) but that GFN is outside of the [sp->gfn, sp->gfn + 511] range; therefore the rmap entry cannot be found and removed when the kvmmmupage is zapped.
- the memslot that covers the first 2MB mapping is deleted, and the kvmmmupage for the now-invalid GPA is zapped. However, rmapremove() only looks at the [sp->gfn, sp->gfn + 511] range established in step 1, and fails to find the rmap entry that was recorded by step 3.
- any operation that causes an rmap walk for the same page accessed by step 3 then walks a stale rmap and dereferences a freed kvmmmupage. This includes dirty logging or MMU notifier invalidations (e.g., from MADVDONTNEED).
The underlying issue is that KVM's walking of shadow PTEs assumes that if a SPTE is present when KVM wants to install a non-leaf SPTE, then the existing kvmmmupage must be for the correct gfn. Because the only way for the gfn to be wrong is if KVM messed up and failed to zap a SPTE... which shouldn't happen, but actually only happens in response to a guest write.
That bug dates back literally forever, as even the first version of KVM assumes that the GFN matches and walks into the "wrong" shadow page. However, that was only an imprecision until 2032a93d66fa ("KVM: MMU: Don't allocate gfns page for direct mmu pages") came along.
Fix it by checking for a target gfn mismatch and zapping the existing SPTE. That way the old SP and rmap entries are gone, KVM installs the rmap in the right location, and everyone is happy.
In the Linux kernel, the following vulnerability has been resolved:
ethtool: cmis: require exact CDB reply length
Malicious SFP module could respond with rpllen longer than what cmiscdbprocessreply() expected, leading to OOB writes. Malicious HW is a bit theoretical but some modules may just be buggy and/or the reads may occasionally get corrupted, so let's protect the kernel.
The existing check protects from short replies. We need to protect from long ones, too. All callers that pass a non-zero rplexplen cast the reply payload to a fixed-layout struct and read fields at fixed offsets, with no version negotiation or short-reply handling:
- cmiscdbvalidatepassword() - cmiscdbmodulefeaturesget() - cmisfwupdatefwmngfeaturesget()
so let's assume that responses longer than expected do not have to be handled gracefully here. Add a warning message to make the debug easier in case my understanding is wrong...
Note that pagedata->length (argument of kmalloc) comes from last arg to ethtoolcmispageinit() which is rplexplen.
Note2 that AIs also like to point out overflows in args->req.payload itself (which is a fixed-size 120 B buffer, on the stack), but callers should be reading structs defined by the standard, so protecting from requests for more data than max seem like defensive programming.
In the Linux kernel, the following vulnerability has been resolved:
ALSA: pcm: oss: Fix setup list UAF on proc write error
sndpcmossprocwrite() links a newly allocated setup entry into the OSS setup list before duplicating the task name. If the task-name allocation fails, the error path frees the already linked entry and leaves setuplist pointing at freed memory.
A later OSS device open can then walk the stale list entry in sndpcmosslookforsetup() and dereference freed memory.
Allocate the task name and initialize the setup entry before publishing the entry on setuplist. Also fetch the initial proc read iterator only after taking setupmutex, so all setuplist traversal follows the same list lifetime rules.
In the Linux kernel, the following vulnerability has been resolved:
accel/rocket: fix UAF via dangling GEM handle in createbo
rocketioctlcreatebo() inserts a GEM handle into the file's IDR via drmgemhandlecreate() early on, then performs several operations that can fail (sgt allocation, drmmm insert, iommumap). If any fail after the handle is live, the error path calls drmgemshmemobjectfree() which kfree's the object without removing the handle from the IDR.
This leaves a dangling handle pointing to freed slab memory. Any subsequent ioctl using that handle (PREPBO, FINIBO, SUBMIT) calls drmgemobjectlookup() and dereferences freed memory (UAF).
Fix by moving drmgemhandlecreate() to after all fallible operations succeed, matching the pattern used by panfrost, lima, and etnaviv.
Also fix drmmminsertnodegeneric() whose return value was silently overwritten by iommumapsgtable() on the next line. Add the missing error check.
[tomeu: Move handle creation to the very end]
In the Linux kernel, the following vulnerability has been resolved:
Bluetooth: L2CAP: Fix possible crash on l2capecredconnrsp
If dcid is received for an already-assigned destination CID the spec requires that both channels to be discarded, but calling l2capchandel may invalidate the tmp cursor created by listforeachentrysafe and in fact it is the wrong procedure as the chan->dcid may be assigned previously it really needs to be disconnected.
Calling l2capchanclone directly may still lead to l2capchandel so instead schedule l2capchantimeout with delay 0 to close the channel asynchronously.
In the Linux kernel, the following vulnerability has been resolved:
net: mana: Skip redundant detach on already-detached port
When manaperportqueueresetworkhandler() runs after a previous detach succeeded but attach failed, the port is left in a detached state with apc->txqp and apc->rxqs already freed. Calling manadetach() again unconditionally leads to NULL pointer dereferences during queue teardown.
Add an early exit in manadetach() when the port is already in detached state (!netifdevicepresent) for non-close callers, making it safe to call idempotently. This allows the queue reset handler and other recovery paths to simply retry manaattach() without redundant teardown.
In the Linux kernel, the following vulnerability has been resolved:
sctp: fix race between sctpwaitforconnect and peeloff
sctpwaitforconnect() drops and re-acquires the socket lock while waiting for the association to reach ESTABLISHED state. During this window, another thread can peeloff the association to a new socket via getsockopt(SCTPSOCKOPTPEELOFF), changing asoc->base.sk. After re-acquiring the old socket lock, sctpwaitforconnect() returns success without noticing the migration — the caller then accesses the association under the wrong lock in sctpdatamsgfromuser().
Add the same sk != asoc->base.sk check that sctpwaitforsndbuf() already has, returning an error if the association was migrated while we slept.
In the Linux kernel, the following vulnerability has been resolved:
vsock/virtio: bind uarg before filling zerocopy skb
virtiotransportsendpktinfo() allocates or reuses the zerocopy uarg before entering the send loop, but virtiotransportallocskb() still fills the skb before it inherits that uarg. When fixed-buffer vectored zerocopy hits MAXSKBFRAGS, iosgfromiter() may partially attach managed frags and return -EMSGSIZE. The rollback path call kfreeskb() to free an skb that carries SKBFLMANAGEDFRAGREFS but no uarg, so skbreleasedata() falls through to ordinary frag unref.
Pass the uarg into virtiotransportallocskb() and bind it immediately before virtiotransportfillskb(). This keeps control or no-payload skbs untouched while ensuring success and rollback share one lifetime rule.
In the Linux kernel, the following vulnerability has been resolved:
zram: fix use-after-free in zrambvecwritepartial()
zramreadpage() picks the sync or async backing device read path based on whether the parent bio is NULL. zrambvecwritepartial() passes its parent bio down, so for ZRAMWB slots the read is dispatched asynchronously and zramreadpage() returns 0 while the bio is still in flight. The caller then runs memcpyfrombvec(), zramwritepage() and freepage() on the buffer, leaving the async read to write into a freed page.
zrambvecreadpartial() was switched to NULL in commit 4e3c87b9421d ("zram: fix synchronous reads") for the same reason; the writepartial counterpart was missed.
In the Linux kernel, the following vulnerability has been resolved:
RDMA/mana: Validate rxhashkeylen
Sashiko points out that rxhashkeylen comes from a uAPI structure and is blindly passed to memcpy, allowing the userspace to trash kernel memory. Bounds check it so the memcpy cannot overflow.
In the Linux kernel, the following vulnerability has been resolved:
RDMA/vmwpvrdma: Fix double free on pvrdmaallocucontext() error path
Sashiko points out that pvrdmauarfree() is already called within pvrdmadeallocucontext(), so calling it before triggers a double free.
In the Linux kernel, the following vulnerability has been resolved:
net: gro: don't merge zcopy skbs
skbgroreceive() can currently copy frags between the source and GRO skb, without checking the zerocopy status, and in particular the SKBFLMANAGEDFRAGREFS flag.
When SKBFLMANAGEDFRAGREFS is set, the skb doesn't hold a reference on the pages in shinfo->frags. Appending those frags to another skb's frags without fixing up the page refcount can lead to UAF.
When either the last skb in the GRO chain (the one we would append frags to) or the source skb is zerocopy, don't merge the skbs.
btrfs: only release the dirty pages io tree after successful writes
In the Linux kernel, the following vulnerability has been resolved:
RDMA/mlx4: Fix mis-use of RCU in mlx4srqevent()
Sashiko points out the radixtree itself is RCU safe, but nothing ever frees the mlx4srq struct with RCU, and it isn't even accessed within the RCU critical section. It also will crash if an event is delivered before the srq object is finished initializing.
Use the spinlock since it isn't easy to make RCU work, use refcountincnotzero() to protect against partially initialized objects, and order the refcountset() to be after the srq is fully initialized.
In the Linux kernel, the following vulnerability has been resolved:
netfilter: bridge: ebtables: close module init race
sashiko reports for unrelated patch: Does the core ebtables initialization in ebtables.c suffer from a similar race? Once nfregistersockopt() completes, the sockopts are exposed globally.
sockopt has to be registered last, just like in ip/ip6/arptables.
In the Linux kernel, the following vulnerability has been resolved:
rbd: eliminate a race in lockdwork draining on unmap
Given how rbdlockaddrequest() and rbdimgexclusivelock() are written, lockdwork may be (re)queued more than it's actually needed: for example in case a new I/O request comes in while we are in the middle of rbdacquirelock() on behalf of another I/O request. This is expected and with rbdreleaselock() preemptively canceling lockdwork is benign under normal operation.
A more problematic example is maybekickacquire():
if (haverequests || delayedworkpending(&rbddev->lockdwork)) { dout("%s rbddev %p kicking lockdwork\n", func, rbddev); moddelayedwork(rbddev->taskwq, &rbddev->lockdwork, 0); }
It's not unrealistic for lockdwork to get canceled right after delayedworkpending() returns true and for moddelayedwork() to requeue it right there anyway. This is a classic TOCTOU race.
When it comes to unmapping the image, there is an implicit assumption of no self-initiated exclusive lock activity past the point of return from rbddevimageunlock() which unlocks the lock if it happens to be held. This unlock is assumed to be final and lockdwork (as well as all other exclusive lock tasks, really) isn't expected to get queued again. However, lockdwork is canceled only in canceltaskssync() (i.e. later in the unmap sequence) and on top of that the cancellation can get in effect nullified by maybekickacquire(). This may result in rbdacquirelock() executing after rbddevdevicerelease() and rbddevimagerelease() run and free and/or reset a bunch of things. One of the possible failure modes then is a violated
rbdassert(rbdimageformatvalid(rbddev->imageformat));
in rbddevheaderinfo() which is called via rbddevrefresh() from rbdpostacquireaction().
Redo exclusive lock task draining to provide saner semantics and try to meet the assumptions around rbddevimageunlock().
In the Linux kernel, the following vulnerability has been resolved:
netfilter: ebtables: move to two-stage removal scheme
Like previous patches for xtables, follow same pattern in ebtables. We can't reuse xt helpers: ebttable struct layout is incompatible.
table->ops assignment is now done while still holding the ebt mutex to make sure we never expose partially-filled table struct.
In the Linux kernel, the following vulnerability has been resolved:
netfilter: xtables: add and use xtablesunregistertableexit
Previous change added xtablesunregistertablepreexit to detach the table from the packetpath and to unlink it from the active table list. In case of rmmod, userspace that is doing set/getsockopt for this table will not be able to re-instantiate the table: 1. The larval table has been removed already 2. existing instantiated table is no longer on the xt pernet table list.
This adds the second stage helper:
unlink the table from the dying list, free the hook ops (if any) and do the audit notification. It replaces xtunregistertable().
In the Linux kernel, the following vulnerability has been resolved:
wifi: mac80211: capture fast-RX rate before mesh reuses skb->cb
ieee80211invokefastrx() reads RX status through IEEE80211SKBRXCB(skb), which aliases the same skb->cb storage that ieee80211rxmeshdata() reuses as IEEE80211TXINFO. In the unicast forward path, meshdata does:
info = IEEE80211SKBCB(fwdskb); memset(info, 0, sizeof(info));
on the same skb the caller still names via rx->skb, then either queues the skb for TX (success) or kfreeskb()'s it (no-route) before returning RXQUEUED. The caller's RXQUEUED arm then calls stastatsencoderate(status) on memory that is either zeroed (success path) or freed (no-route path). The latter is KASAN slab-use-after-free in ieee80211prepareandrxhandle.
Fix by encoding the rate from status before invoking ieee80211rxmeshdata(), so the RXQUEUED arm consumes a value captured while status was still backed by valid memory.
In the Linux kernel, the following vulnerability has been resolved:
net/mlx5e: xsk: Fix DMA and xdpframe leak on XDPTX xmit failure
In the XSK branch of mlx5exmitxdpbuff(), when sq->xmitxdpframe() returns false (e.g. XDPSQ is full), the function returns without unmapping the DMA address or freeing the xdpframe allocated by xdpconvertzctoxdpframe(). The xdpififo push only happens on success, so the completion path cannot recover these entries.
With CONFIGDMAAPIDEBUG=y, the leak surfaces on driver unbind:
DMA-API: pci 0000:08:00.0: device driver has pending DMA allocations while released from device [count=1116] One of leaked entries details: [device address=0x000000010ffd7028] [size=1534 bytes] [mapped with DMATODEVICE] [mapped as phy] WARNING: kernel/dma/debug.c:881 at dmadebugdevicechange+0x127/0x180 ... DMA-API: Mapped at: debugdmamapphys+0x4b/0xd0 dmamapphys+0xfd/0x2d0 mlx5exdphandle+0x5ae/0xac0 [mlx5core] mlx5exskskbfromcqempwrqlinear+0xc4/0x170 [mlx5core] mlx5ehandlerxcqempwrq+0xc1/0x290 [mlx5core]
Add the missing unmap + xdpreturnframe, matching the cleanup already done in mlx5exdpxmit(). hasfrags is rejected earlier in this branch, so no per-frag unmap is needed.
bpf: Free reuseport cBPF prog after RCU grace period.
In the Linux kernel, the following vulnerability has been resolved:
USB: serial: ioti: fix heap overflow in getmanufinfo()
getmanufinfo() reads le16tocpu(romdesc->Size) bytes from the device I2C EEPROM into a buffer allocated with kmallocobj(), which is sizeof(struct edgetimanufdescriptor) = 10 bytes.
The Size field comes from the device and is only validated (in checki2cimage()) to make sure the descriptor fits within TIMAXI2CSIZE (16384 bytes), not against the destination buffer size. A malicious USB device can therefore set Size to any value up to 16377, causing a heap overflow of up to 16367 bytes when plugged into a host running this driver.
validcsum() is called after readrom() and also iterates buffer[0..Size-1], compounding the out-of-bounds access.
Fix by rejecting descriptors with unexpected length before calling readrom().
[ johan: amend commit message; also check for short descriptors ]
In the Linux kernel, the following vulnerability has been resolved:
RDMA/mana: Remove user triggerable WARNON() in manaibcreateqprss()
Sashiko points out that the user can specify WQs sharing the same CQ as a part of the uAPI and this will trigger the WARNON() then go on to corrupt the kernel.
Just reject it outright and fail the QP creation.
A flaw in the Linux kernel's ebtables SNAT target allows writing to shared memory pages when rewriting ARP sender hardware addresses without ensuring writability, potentially causing file/memory corruption or denial of service.
In the Linux kernel, the following vulnerability has been resolved:
net/mlx5e: xsk: Fix unlocked writing to ICOSQ
During napi poll, when the affinity changes and there's still XSK work to be done, we trigger an ICOSQ interrupt on the new CPU. However, this triggering on the ICOSQ is done unprotected.
There are 2 such races:
A) mlx5etriggerirq() is called while mlx5exskallocrxmpwqe() is running from a different CPU due to affinity change. This can happen because IRQ triggering is done after napicompletedone(). At this point the NAPI can be scheduled on a different CPU. Like this:
CPU A (old affinity, NAPI tail) CPU B (new affinity, fresh NAPI) ------------------------------- -------------------------------- napicompletedone() clears SCHED mlx5ecqarm(...) napischeduleprep() sets SCHED mlx5enapipoll() mlx5exskallocrxmpwqe() mlx5eicosqsynclock() // noop memcpy 640 B UMR body advance sq->pc by 10 mlx5etriggerirq(&c->icosq) wqeinfo[pi] = {NOP, 1} mlx5epostnop() advances sq->pc
B) mlx5etriggerirq() is called on the ICOSQ when mlx5etriggernapiicosq() is running.
The obvious fix would be to lock the ICOSQ. But ICOSQ has an optimized locking scheme that doesn't work for this scenario. Kick the async ICOSQ instead which is always locked.
This issue was noticed in the wild with the following splat:
netdevice: ge-0-0-1: Bad OP in ICOSQ CQE: 0xd WARNING: drivers/net/ethernet/mellanox/mlx5/core/enrx.c:826 [...] [...] Call Trace: <IRQ> mlx5enapipoll+0x11d/0x7f0 [mlx5core] napipoll+0x30/0x200 ? skbdeferfreeflush+0x9c/0xc0 netrxaction+0x2fe/0x3f0 handlesoftirqs+0xd8/0x340 irqexitrcu+0xbc/0xe0 commoninterrupt+0x85/0xa0 </IRQ> <TASK> asmcommoninterrupt+0x26/0x40 [...] ---[ end trace 0000000000000000 ]--- mlx5core 0000:08:00.0 ge-0-0-1: Error cqe on cqn 0x548, ci 0x2022, qn 0x8f4, opcode 0xd, syndrome 0x2, vendor syndrome 0x68 00000000: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00000010: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00000020: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00000030: 00 00 00 00 01 00 68 02 01 00 08 f4 de 14 59 d2 WQE DUMP: WQ size 16384 WQ cur size 0, WQE index 0x1e14, len: 64 00000000: 00 00 00 01 d9 ed 80 02 00 00 00 01 d9 ed 90 02 00000010: 00 00 00 01 d9 ed a0 02 00 00 00 01 d9 ed b0 02 00000020: 00 00 00 01 d9 ed c0 02 00 00 00 01 d9 ed d0 02 00000030: 00 00 00 01 d9 ed e0 02 00 00 00 01 d9 ed f0 02 mlx5core 0000:08:00.0 ge-0-0-1: Error cqe on cqn 0x548, ci 0x2023, qn 0x8f4, opcode 0xd, syndrome 0x5, vendor syndrome 0xf9 00000000: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00000010: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00000020: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00000030: 00 00 00 00 01 00 f9 05 01 00 08 f4 de 15 cf d2
In the Linux kernel, the following vulnerability has been resolved:
staging: rtl8723bs: rtwmlme: add bounds checks before ielength subtraction
Add guards to ensure ielength is large enough before subtracting fixed IE offsets to prevent unsigned integer underflow.
In the Linux kernel, the following vulnerability has been resolved:
xfrm: defensively unhash xfrmstate lists in xfrmstatedelete
KASAN reproduces a slab-use-after-free in xfrmstatedelete()'s hlistdelrcu calls under syzkaller load on linux-6.12.y stable (reproduced on 6.12.47, also reachable via the same code path on torvalds/master and on the ipsec tree). Nine unique signatures cluster in the xfrmstate lifecycle, the load-bearing one being:
BUG: KASAN: slab-use-after-free in hlistdel include/linux/list.h:990 [inline] BUG: KASAN: slab-use-after-free in hlistdelrcu include/linux/rculist.h:516 [inline] BUG: KASAN: slab-use-after-free in xfrmstatedelete net/xfrm/xfrmstate.c Write of size 8 at addr ffff8881198bcb70 by task kworker/u8:9/435
Workqueue: netns cleanupnet Call Trace: hlistdel / hlistdelrcu xfrmstatedelete xfrmstatedelete xfrmstateflush xfrmstatefini opsexitlist cleanupnet
The other observed signatures hit the same slab object from xfrmstatelookup, xfrmallocspi, xfrmstateinsert and an OOB write variant of xfrmstatedelete, all on the byseq/byspi hash chains.
xfrmstatedelete() guards its byseq and byspi unhashes with value-based predicates:
if (x->km.seq) hlistdelrcu(&x->byseq); if (x->id.spi) hlistdelrcu(&x->byspi);
while everywhere else in the file (e.g. statecache, statecacheinput) the safer hlistunhashed() check is used. xfrmallocspi() sets x->id.spi = newspi inside xfrmstatelock and then immediately inserts into byspi, but a path that observes x->id.spi != 0 outside of xfrmstatelock can still skip-or-hit the byspi unhash inconsistently with whether x is actually on the list. The same holds for x->km.seq versus byseq, and the bydst/bysrc unhashes have no predicate at all, so a second xfrmstatedelete() on the same object writes through LISTPOISON pprev.
The defensive change here:
- Use hlistdelinitrcu() instead of hlistdelrcu() on bydst, bysrc, byseq and byspi so a second deletion is a no-op rather than a write through LISTPOISON pprev. The byseq/byspi nodes are already initialised in xfrmstatealloc(). - Test hlistunhashed() rather than the value predicate for byseq/byspi, so the unhash decision tracks list state rather than mutable scalar fields.
Empirical verification: applied this patch on top of v6.12.47, rebuilt, and re-ran the same syzkaller harness for 1h16m on a previously-crashy configuration that produced ~100 hits each of slab-use-after-free Read in xfrmallocspi / Read in xfrmstatelookup / Write in xfrmstatedelete. After the patch, 7.1M execs across 32 VMs at ~1550 exec/sec produced zero xfrmstate UAF/OOB hits. /proc/slabinfo confirms the xfrmstate slab is actively allocated and freed during the run (~143 KiB resident), so the fuzzer is still exercising those code paths -- they just no longer crash.
Reproduction:
- Linux 6.12.47 x8664 + KASANGENERIC + KASANINLINE + KCOV - syzkaller @ 746545b8b1e4c3a128db8652b340d3df90ce61db - 32 QEMU/KVM VMs x 2 vCPU on AWS c5.metal bare metal - 9 unique signatures collected in ~9h, all within xfrmstate lifecycle
In the Linux kernel, the following vulnerability has been resolved:
RDMA/mlx5: Fix error path fall-through in mlx5ibdevressrqinit()
mlx5ibdevressrqinit() allocates two SRQs, s0 and s1. When ibcreatesrq() fails for s1, the error branch destroys s0 but falls through and unconditionally assigns the freed s0 and the ERRPTR s1 to devr->s0 and devr->s1.
This leads to several problems: the lock-free fast path checks "if (devr->s1) return 0;" and treats the ERRPTR as already initialised; users in mlx5ibcreateqp() dereference the freed SRQ or ERRPTR via tomsrq(devr->s0)->msrq.srqn; and mlx5ibdevrescleanup() dereferences the ERRPTR and double-frees s0 on teardown.
Fix by adding the same goto unlock in the s1 failure path.
In the Linux kernel, the following vulnerability has been resolved: