From nobody Mon Sep 28 07:23:11 2026 Received: from out203-205-221-155.mail.qq.com (out203-205-221-155.mail.qq.com [203.205.221.155]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D74E03E0758 for ; Tue, 25 Aug 2026 09:13:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=203.205.221.155 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787649192; cv=none; b=lMzCPiPl4uyY+ftXqLkDpk5yrFOitqXhPQ+mWN6Q5R4zmfIq6gHBz6N7t302JHO6pLt8phJtam1AUQFaqc/gj58s1H5GbEq6BxLZJVKGLexBP3iHU9ey09k4Rcv1Y+wDKyE5Et8V/9Ll0kXs/SqKz7xgn8cI75knoL4BhcliPWc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787649192; c=relaxed/simple; bh=aLMwOWgS1UYjq1vewfAliYHENB1SJZ4WjwkN0VMmghE=; h=Message-ID:Date:MIME-Version:To:Cc:From:Subject:Content-Type; b=szVoQxRSjNuyjBWvfqqNkYj9gY5wCYc4lieAc/eZ17TQs5qE7YQEovcXlCGawtRIeBGTxWuJqJmUbG2RaWog1TKuXIMi+bGc2FGD/KIsdxk6vHCvcUhF6sBiZJQcHStLO8j5sjTwKrpgrgjNz/LiMCF7kVbAzqxEqfPgf2yuyf8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=qq.com; spf=pass smtp.mailfrom=qq.com; dkim=pass (1024-bit key) header.d=qq.com header.i=@qq.com header.b=hFy5ku7B; arc=none smtp.client-ip=203.205.221.155 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=qq.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=qq.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=qq.com header.i=@qq.com header.b="hFy5ku7B" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=qq.com; s=s201512; t=1787649181; bh=aLMwOWgS1UYjq1vewfAliYHENB1SJZ4WjwkN0VMmghE=; h=Date:To:Cc:From:Subject; b=hFy5ku7B7NZqFBu/CRSZBggRGV886ZaDZxjD5VTq9/CiC1c3J0R5MH1guuihusWgD dWkEqtrZEOW/obpo9Vp2bASXu1qrQWpeHx0yaDUXGaIaFI1KWOzq03YwQrk0o4zD2c rUHAr6kSFnTFlSKH+NUEs96P9kHuvTtqzzCQwwgU= Received: from [192.168.255.10] ([111.206.145.19]) by newxmesmtplogicsvrszc50-0.qq.com (NewEsmtp) with SMTP id 2ED898C4; Tue, 25 Aug 2026 17:11:45 +0800 X-QQ-mid: xmsmtpt1787649105te73z3zzb Message-ID: X-QQ-XMAILINFO: OIJV+wUmQOUA4oKfQabcsr+2XPn7yvy3PVwkBvdsfwqACLLeg9omL+TLv2Ozlw FsTnBkxtFNSeWA8T/jbWR8SD2j8dmSzk45h4/yytL0aeeDTiHeEz9/8NcCsLSlG78+ayBxwlxMZT /I8nlF13mCDXXrT8+eW6mgwrRRitNSe5bvJdi9r0XZ75JEs8AAV9hAg9YwDOadKh0wbLVt7MAp9F cb9FXbHist6vnGnkQBUJyN1u9lWku4qVQvC7AphBZJWNXQhim9s7CwgdMLz0uU5X1etkYQdQldyF +Z4LFdC2iw37hX9zyUyhiOtu+C6G2CjFGTLSx0IAfs+GUTAVAPc0XxNIAkzjHnXRTOc1yrY3SVG3 iQ+3bSuazXc0mAIEIZotgQqu8+IMfcfABjtXhZUEJVp6McmVzaJU/JpTgzIeqTgzn/9sBxrlMJQu Xms/fqFQ88Sr2E+dkbTybhf0fyof0qmRR3ph94rBfE72d0aKC7KzwxATldvPlHNwlz4LCx9m+q8f FZ2cefstaJN/Jz6lxez4NDH5evcHRIV6sBsVgoNLnnglYLAQ/6SNzBmFj3/YdhivyPdv2eezRvqv 27PHD61+AsOokqkxbNP0Lb83RSjbfNBIW7pDJ9yHLjTF/1I8y/ZZf12Q2bXbeyd63+w6UoIpzYBX C3JgS5hQhGpuMpwGNV/6b+hqNLH8IfsZke9a+zcRq8m16lFrVd4ta6F6k8eNtygEZ89qSo6SGSoD d2hEEjgdSmZaAFyfbn8xbx+OoFwrRVcfUKy+VshAK0VdzjqEcX6f4Se9V9DvY6RHQJagtYjE7CyT OsAQh2xgQbUgjDhmNvUKHCHmF5QhVTRoUZhfS55CoudFsFER9+fZf8zeMNl3orq7fEjfzcIqHejq sGXMICY9Wo4DuhL42lRxUuo7VIVbvDvp+aiJvzRLm8Zmlxq1VTLbyyOolZ61uRlOo3JakntW2QHK J55arvaURMsEJ2CUvgh1NRWOAUjpaf+XVfXySuANBXRYRKN2LKfzNWZFGoxfmqGKBY1EOqOnakOA gkir8jzoNNvgilcrxPPi8IYuBlWCqZtmf2RDCO2pWjokQun4LRmSrCxxkXTdugQlG3IvfWV+kEZO S7selFwt3wB8w6YGLDWeqWPdaAeg== X-QQ-XMRINFO: MPJ6Tf5t3I/ylTmHUqvI8+Wpn+Gzalws3A== X-OQ-MSGID: <563173dc-706c-4b5e-b37b-ae867211bc48@qq.com> Date: Tue, 25 Aug 2026 17:11:44 +0800 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird To: peterz@infradead.org, mingo@redhat.com, will@kernel.org, boqun@kernel.org, longman@redhat.com Cc: linux-kernel@vger.kernel.org From: Yang Zi <2959243019@qq.com> Subject: [PATCH] locking/mutex: Restore RCU read-side protection for optimistic spinning Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable The optimistic spinning paths in kernel/locking/mutex.c dereference the current lock owner's task_struct (via owner_on_cpu() -> owner->on_cpu and task_cpu(owner)) without any lifetime protection on the owner pointer. __mutex_owner() (kernel/locking/mutex.h) returns a bare task_struct pointer that is read directly from lock->owner. mutex_spin_on_owner() and mutex_can_spin_on_owner() then dereference it inside "while (__mutex_owner(lock) =3D=3D owner)" / immediately after the owner load, whi= le relying only on preempt_disable(). Since commit 6c2787f2a20c ("locking: Remove rcu_read_{,un}lock() for preempt_{dis,en}able()"), the code has assumed that preempt_disable() is equivalent to an RCU read-side critical section and therefore that the owner's task_struct "won't go away during the spinning period". That assumption does not hold on PREEMPT_LAZY kernels: preempt_disable() there no longer actually disables preemption, so the RCU callback that frees a just-exited owner's task_struct (put_task_struct_rcu_user() -> call_rcu(), run in rcu_do_batch()) can fire and free the object while a spinner still holds the stale pointer, leading to a slab use-after-free read of owner->on_cpu. This manifests in three independent syzkaller/KASAN reports with the same root cause: - BUG 123 (dw_edma_pcie, v7.1, PREEMPT(lazy)): finit_module -> =C2=A0 driver_register -> bus_add_driver -> bus_for_each_dev -> __driver_at= tach =C2=A0 -> device_lock -> __mutex_lock -> mutex_optimistic_spin -> =C2=A0 mutex_can_spin_on_owner -> owner_on_cpu. - BUG 132 (bna): read(2) -> uevent_show -> device_lock -> =C2=A0 mutex_optimistic_spin -> mutex_can_spin_on_owner -> owner_on_cpu. - BUG 208 (iavf, v7.1, PREEMPT(lazy)): delete_module -> =C2=A0 pci_unregister_driver -> ... -> iavf_remove -> netdev_lock -> =C2=A0 __mutex_lock -> mutex_optimistic_spin -> mutex_can_spin_on_owner -> =C2=A0 owner_on_cpu. Fix it by taking an explicit rcu_read_lock()/rcu_read_unlock() around the owner dereference in both mutex_spin_on_owner() and mutex_can_spin_on_owner= (), restoring the pre-5.16 protection, and update the now-inaccurate comments. Signed-off-by: Yang Zi <2959243019@qq.com> --- diff --git a/kernel/locking/mutex.c b/kernel/locking/mutex.c index 942a939cee95..cc2d7090bf36 100644 --- a/kernel/locking/mutex.c +++ b/kernel/locking/mutex.c @@ -389,14 +389,21 @@ bool mutex_spin_on_owner(struct mutex *lock, struct t= ask_struct *owner, =C2=A0 =C2=A0 =C2=A0 =C2=A0lockdep_assert_preemption_disabled(); =C2=A0 +=C2=A0 =C2=A0 /* +=C2=A0 =C2=A0 =C2=A0* Optimistic spinners run with preempt_disable() which= , on classic +=C2=A0 =C2=A0 =C2=A0* PREEMPT_RCU, was enough to act as an RCU read-side c= ritical +=C2=A0 =C2=A0 =C2=A0* section. On PREEMPT_LAZY kernels preempt_disable() n= o longer +=C2=A0 =C2=A0 =C2=A0* prevents RCU callbacks (e.g. the free of an exiting = owner's +=C2=A0 =C2=A0 =C2=A0* task_struct via call_rcu() in rcu_do_batch()) from r= unning, so take +=C2=A0 =C2=A0 =C2=A0* an explicit RCU read-side critical section to keep t= he owner alive. +=C2=A0 =C2=A0 =C2=A0*/ +=C2=A0 =C2=A0 rcu_read_lock(); =C2=A0 =C2=A0 =C2=A0while (__mutex_owner(lock) =3D=3D owner) { =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0/* -=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0* Ensure we emit the owner->on_cpu, dere= ference _after_ -=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0* checking lock->owner still matches own= er. And we already -=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0* disabled preemption which is equal to = the RCU read-side -=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0* crital section in optimistic spinning = code. Thus the -=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0* task_strcut structure won't go away du= ring the spinning -=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0* period +=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0* Ensure we emit the owner->on_cpu deref= erence _after_ +=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0* checking lock->owner still matches own= er. If that fails, +=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0* owner might point to freed memory. If = it still matches, +=C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0* the rcu_read_lock() ensures the memory= stays valid. =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 */ =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0barrier(); =C2=A0 @@ -415,6 +422,7 @@ bool mutex_spin_on_owner(struct mutex *lock, struct tas= k_struct *owner, =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0cpu_relax(); =C2=A0 =C2=A0 =C2=A0} +=C2=A0 =C2=A0 rcu_read_unlock(); =C2=A0 =C2=A0 =C2=A0 =C2=A0return ret; =C2=A0} @@ -433,13 +441,16 @@ static inline int mutex_can_spin_on_owner(struct mute= x *lock) =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0return 0; =C2=A0 =C2=A0 =C2=A0 =C2=A0/* -=C2=A0 =C2=A0 =C2=A0* We already disabled preemption which is equal to the= RCU read-side -=C2=A0 =C2=A0 =C2=A0* crital section in optimistic spinning code. Thus the= task_strcut -=C2=A0 =C2=A0 =C2=A0* structure won't go away during the spinning period. +=C2=A0 =C2=A0 =C2=A0* See the comment in mutex_spin_on_owner(): preempt_di= sable() is no +=C2=A0 =C2=A0 =C2=A0* longer an RCU read-side critical section on PREEMPT_= LAZY kernels, +=C2=A0 =C2=A0 =C2=A0* so protect the owner dereference with an explicit RC= U read-side +=C2=A0 =C2=A0 =C2=A0* critical section. =C2=A0 =C2=A0 =C2=A0 */ +=C2=A0 =C2=A0 rcu_read_lock(); =C2=A0 =C2=A0 =C2=A0owner =3D __mutex_owner(lock); =C2=A0 =C2=A0 =C2=A0if (owner) =C2=A0 =C2=A0 =C2=A0 =C2=A0 =C2=A0retval =3D owner_on_cpu(owner); +=C2=A0 =C2=A0 rcu_read_unlock(); =C2=A0 =C2=A0 =C2=A0 =C2=A0/* =C2=A0 =C2=A0 =C2=A0 * If lock->owner is not set, the mutex has been releas= ed. Return true