CVE-2026-98220: sched_ext: Fix NULL sched deref in kfunc sub-sched error paths
In the Linux kernel, the following vulnerability has been resolved:
schedext: Fix NULL sched deref in kfunc sub-sched error paths
When the root scheduler has sub-scheds attached, the COMPAT kfunc wrappers scxbpfselectcpuand() and scxbpfdsqinsertvtime() refuse the call and report to @p's scheduler:
scxerror(scxtasksched(p), "... must be used");
The wrappers are reachable with tasks that have no scheduler. scxbpfselectcpuand() is in the selectcpu kfunc group, which scxkfunccontextfilter() opens to BPFPROGTYPESYSCALL programs; scxbpfdsqinsertvtime() is in the enqueuedispatch group, which ops.enqueue() and ops.dispatch() may call with any KFRCU task -- the group has no kftasks validation, and scxdsqinsertpreamble() checks task ownership with scxtaskonsched() precisely because @p may be an arbitrary task.
scxtasksched(p) is p->scx.sched, which is NULL for tasks past schedextdead() -- which clears it via scxdisableandexittask() on exit -- and for idle tasks, which the enable paths skip as they are never scheduled through SCX. It is also an rcudereferenceprotected() that expects @p's pilock or rq lock, which neither wrapper holds. Passing NULL to scxerror() reaches scxvexit(), which dereferences sch->exitinfo, oopsing the kernel.
One concrete trigger exercised while developing the fix: a BPFPROGTYPESYSCALL program calling the selectcpuand wrapper on an exited-but-not-reaped task while a sub-scheduler was attached (its pid stays findable while the zombie is unreaped; faulting instruction is the scxvexit() prologue "mov r15,[rdi+0x398]" with RDI=NULL and 0x398 the offset of sch->exitinfo):
schedext: BPF scheduler "kfuncsubschednull" enabled schedext: BPF sub-scheduler "kfuncsubschednull" enabled schedext: Unassociated program runselectcpu (id 76) BUG: kernel NULL pointer dereference, address: 0000000000000398 #PF: supervisor read access in kernel mode #PF: errorcode(0x0000) - not-present page Oops: Oops: 0000 [#1] SMP NOPTI CPU: 7 UID: 0 PID: 8201 Comm: kfunctestrunn Tainted: G W RIP: 0010:scxvexit+0x25/0xa0 Code: ... <4c> 8b bf 98 03 00 00 ... CR2: 0000000000000398 Call Trace: <TASK> scxexit+0x4f/0x70 scxbpfselectcpuand+0xab/0xb0 bpfprog430ed61a7b66e03arunselectcpuand+0x9c/0xe7 ? x64sysbpf+0x2c/0x40 bpfprogtestrunsyscall+0x130/0x2f0 sysbpf+0x930/0x10d0 ? x64sysbpf+0x2c/0x40 x64sysbpf+0x2c/0x40 dosyscall64+0xbc/0x460 entrySYSCALL64afterhwframe+0x76/0x7e </TASK>
Read @p's scheduler under RCU instead, which the wrappers can do from their guard(rcu)(): fault it when it can be determined, and when it can't be determined -- @p is a task past schedextdead() or an idle task -- there is nothing obviously wrong to report, so just refuse the call as before without faulting any scheduler.
These COMPAT wrappers are scheduled for eventual removal once the deprecation grace period elapses, but until then -- and regardless of their removal timeline -- they must not oops the kernel on a task they are handed.