CVE-2026-98256: signal: Prevent exec() race
In the Linux kernel, the following vulnerability has been resolved:
signal: Prevent exec() race
Hyunwoo debugged the following KASAN UAF splat:
BUG: KASAN: slab-use-after-free in sendsignallocked+0xb27/0xba0 Write of size 8 at addr ffff888007ed80c8 by task poc/79 ... Call Trace: sendsignallocked+0xb27/0xba0 dosendsiginfo+0xa7/0x160 dosendspecific+0x76/0xa0 x64systgkill+0x193/0x270 ... Allocated by task 80: dotimercreate+0x1a4/0x1030 x64systimercreate+0x145/0x190 ... Freed by task 12: kmemcachefreebulk+0x1f8/0x4a0 kvfreercubulk+0x14f/0x1c0 kfreercuwork+0x128/0x1a0 ... Last potentially related work creation: kvfreecallrcu+0x39/0x390 flushitimersignals+0x211/0x320 flushitimersignals+0x47/0x90 beginnewexec+0xa6b/0x28c0
It turned out that this happens with a non-leader exec() as Hyunwoo explained:
dethread() calls exchangetids() before releasetask(leader), so the struct pid held by a SIGEVTHREADID timer created against the leader's tid now points to the thread which called execve(). pidtask() returns that thread and locktasksighand() on it succeeds.
If the timer signal is blocked, its sigqueue stays queued on the leader's task::pending. The next expiry of that timer can then run while releasetask() flushes the queue.
posixtimersendsigqueue() checks whether the sigqueue is already queued with a plain listempty(), which only reads listhead::next. listdelinit() is not atomic and INITLISTHEAD() stores listhead::next before listhead::prev, so the check can pass in between. listaddtail() queues the entry on the task::pending of the live thread, and the listhead::prev store from the flush then overwrites the listhead::prev link that listaddtail() has just set.
flushitimersignals() does not undo that either. With listhead::prev pointing at the entry itself, its listdelinit() only stores the same values again, so the entry is not removed from the list. It is still there after the last reference is dropped and the timer is freed by RCU, and the listaddtail() of a later tgkill() follows that listhead::prev into the freed timer.
This problem surfaced with the recent commit which moved the sigqueue flush out of the sighand lock held region.
Hyonwoo proposed to fix this by using listdelinitcareful(), but that just papers over the problem. After some disucssions and various attempts to solve it, Eric pointed out that there is no reason to flush task::pending late in releasetask() and it should be done in exitsignals() already.
As nothing can collect and deliver signals which are queued in a dying task's pending queue, there is no reason to delay it further.
But it has to be ensured that no signals can be queued into it after that point. exitsignals() sets PFEXITING in task::flags, which can be used as an indicator for this.
Cure it by:
- Preventing signal queueing for task private signals (PIDTYPEPID) when the task has PFEXITING set in sendsignallocked() and in posixtimersendsigqueue().
- Protecting the unlocked setting of PFEXITING in exitsignals() for the task group empty and the group exit case with sighand lock
- Flushing task::pending signals right there.
Optimize that by moving the whole pending list to an on-stack list head under sighand lock and free the signals without the lock held.
There has been quite some discussion about the lockless flush and the non-leader exec case on weakly ordered systems. The problem is that a third party which tries to send a posix timer signal relies on the PID lookup to find the target task and that lookup might result in the new leader when the signal was originaly directed to the old leader. In case that the signal was queued on the old leader then the lockless flush raised a concern over the following situation:
oldleader newleader third party
A: flushlist() // listdelin ---truncated---
Affected Software
Event History
Frequently Asked Questions
What access does an attacker need to exploit this issue?
The CVSS vector indicates local access, low attack complexity, and low privileges are required. No user interaction is required.
What conditions trigger the use-after-free?
The reported scenario involves a non-leader thread calling execve() while a SIGEV_THREAD_ID timer was created for the thread-group leader. If the timer signal is blocked, its queued sigqueue can remain on the former leader's pending signal list and later be accessed after that task is freed.
What is the potential impact?
The issue is a kernel use-after-free in signal handling. The supplied CVSS rating assigns high impact to confidentiality, integrity, and availability.
How can I determine whether a system may be affected?
Review whether the running Linux kernel includes one of the referenced stable fixes. Systems running workloads that use SIGEV_THREAD_ID timers, blocked timer signals, and execve() from non-leader threads are the scenario described in the report.