CVE-2026-74357: drm/amdgpu: fix KASAN slab-out-of-bounds in amdgpu_coredump ring dump
In the Linux kernel, the following vulnerability has been resolved:
drm/amdgpu: fix KASAN slab-out-of-bounds in amdgpucoredump ring dump
The ring content dump in amdgpucoredump() uses two separate loops over adev->rings[]: the first counts rings with unsignalled fences to size the allocation, and the second copies ring data into the allocated buffers.
Both loops use the same condition to skip rings:
atomicread(&ring->fencedrv.lastseq) == ring->fencedrv.syncseq
Because lastseq is an atomic that is updated concurrently by the fence signalling path, additional rings may appear unsignalled in the second loop that were signalled during the first. When this happens, idx exceeds the allocated ringcount and the store to coredump->rings[idx] writes past the end of the kcalloc-ed buffer.
This was found during IGT stressful test amdqueuereset which triggers random GPU resets. The OVERSIZE subtest (CMDSTREAMEXECINVALIDPACKETLENGTHOVERSIZE on GFX ring) provokes a ring timeout and subsequent coredump, which hits the race between the counting and copying loops. The failure is non-deterministic and depends on fence signalling timing during the reset.
KASAN log:
BUG: KASAN: slab-out-of-bounds in amdgpucoredump+0x1274/0x12f0 [amdgpu] Write of size 4 at addr ffff888106154258 by task kworker/u128:5/23625 CPU: 16 UID: 0 PID: 23625 Comm: kworker/u128:5 Not tainted 6.19.0+ #35 Workqueue: amdgpu-reset-dev drmschedjobtimedout [gpusched] Call Trace: <TASK> dumpstacklvl+0xa5/0x110 printreport+0xd1/0x660 kasanreport+0xf3/0x130 asanreportstore4noabort+0x17/0x30 amdgpucoredump+0x1274/0x12f0 [amdgpu] amdgpujobtimedout+0xef0/0x16c0 [amdgpu] drmschedjobtimedout+0x194/0x5c0 [gpusched] processonework+0x84b/0x1990 workerthread+0x6b8/0x11b0 </TASK>
Allocated by task 23625: kasansavestack+0x39/0x70 kasankmalloc+0xc3/0xd0 kmallocnoprof+0x2ec/0x910 amdgpucoredump+0x5c5/0x12f0 [amdgpu] amdgpujobtimedout+0xef0/0x16c0 [amdgpu]
The buggy address belongs to the object at ffff888106154200 which belongs to the cache kmalloc-rnd-09-96 of size 96 The buggy address is located 16 bytes to the right of allocated 72-byte region [ffff888106154200, ffff888106154248)
72 bytes = 3 sizeof(struct amdgpucoredumpring), so ringcount was 3 but idx reached 3+, writing ringindex (at struct offset 16) 16 bytes past the allocation.
Fix by adding an idx < ringcount guard to the copy loop so it cannot exceed the allocated count even when the fence state changes between the two passes.
Affected Software
Remediation
Recommended actions to resolve this vulnerability, in priority order.
- Upgrade
Upgrade
amdgpu (Linux kernel)to a version that resolves this vulnerability.Patch drm/amdgpu: fix KASAN slab-out-of-bounds in amdgpu_coredump ring dump