From nobody Thu Sep 24 19:04:36 2026 Delivered-To: importer@patchew.org Received-SPF: pass (zohomail.com: domain of lists.xenproject.org designates 192.237.175.120 as permitted sender) client-ip=192.237.175.120; envelope-from=xen-devel-bounces@lists.xenproject.org; helo=lists.xenproject.org; Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of lists.xenproject.org designates 192.237.175.120 as permitted sender) smtp.mailfrom=xen-devel-bounces@lists.xenproject.org ARC-Seal: i=1; a=rsa-sha256; t=1789478294; cv=none; d=zohomail.com; s=zohoarc; b=m56dIgHClZKrL/IkrPucRafsJz1x1wFtmtu/SjDvoPbNCxWLSHcWt3GdVAWsIJgo84FbJoZqtMttt0w7Kvs90IJyKkoMlOE820mu0pJkClHY5Vz/kF4a/bV1arxOr6m4R1EVrS1QrAKzLpVNVTEMSPFDTDZsTAQkhiq1AtCT5G8= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1789478294; h=Content-Type:Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=6aaaKNzMayGmjMwZUrTAjDsf1dO37tc7Q4XOVkfUDkI=; b=IxempiJlLZKfJb8BWVoRhJycXnletQLdjYT8rA///FZkvSGYM5Kh4bD30M5QS6lZ9BanY1rJG+/tnvdffe6BR5Io4DRaB8BXQc2dy1TEqe+rQE55dBaJdwyFYORq9nWLI1nm3Bl2Djs1UBYa/dKplZDhMe3gHb9eNIIvNTqLro4= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of lists.xenproject.org designates 192.237.175.120 as permitted sender) smtp.mailfrom=xen-devel-bounces@lists.xenproject.org Return-Path: Received: from lists.xenproject.org (lists.xenproject.org [192.237.175.120]) by mx.zohomail.com with SMTPS id 1789478294873866.9930870444733; Tue, 15 Sep 2026 06:18:14 -0700 (PDT) Received: from list by lists.xenproject.org with outflank-mailman.1421705.1647399 (Exim 4.92) (envelope-from ) id 1x6T2m-0004RN-IE; Tue, 15 Sep 2026 13:17:52 +0000 Received: by outflank-mailman (output) from mailman id 1421705.1647399; Tue, 15 Sep 2026 13:17:52 +0000 Received: from localhost ([127.0.0.1] helo=lists.xenproject.org) by lists.xenproject.org with esmtp (Exim 4.92) (envelope-from ) id 1x6T2m-0004RG-FL; Tue, 15 Sep 2026 13:17:52 +0000 Received: by outflank-mailman (input) for mailman id 1421705; Tue, 15 Sep 2026 13:17:51 +0000 Received: from mx.expurgate.net ([195.190.135.10]) by lists.xenproject.org with esmtp (Exim 4.92) (envelope-from ) id 1x6T2l-0004QS-G5 for xen-devel@lists.xenproject.org; Tue, 15 Sep 2026 13:17:51 +0000 Received: from mx.expurgate.net (helo=localhost) by mx.expurgate.net with esmtp id 1x6T2k-00EG5U-T6 for xen-devel@lists.xenproject.org; Tue, 15 Sep 2026 15:17:50 +0200 Received: from [10.42.69.4] (helo=localhost) by localhost with ESMTP (eXpurgate MTA 0.9.1) (envelope-from ) id 6aa94574-8faa-0a2a0a5109dd-0a2a45049d76-36 for ; Tue, 15 Sep 2026 15:17:50 +0200 Received: from [74.125.224.140] (helo=mail-yx2-f12.google.com) by tlsNG-ebf023.mxtls.expurgate.net with ESMTPS (eXpurgate 4.57.1) (envelope-from ) id 6aa9457d-b57f-0a2a45040019-4a7de08cbd16-3 for ; Tue, 15 Sep 2026 15:17:50 +0200 Received: by mail-yx2-f12.google.com with SMTP id 956f58d0204a3-66fabb1a601so278080d50.3 for ; Tue, 15 Sep 2026 06:17:49 -0700 (PDT) Received: from toxicpanda.com (ec2-34-228-114-98.compute-1.amazonaws.com. [34.228.114.98]) by smtp.gmail.com with ESMTPSA id 6a1803df08f44-9120f2032d4sm120309036d6.5.2026.09.15.06.17.46 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 15 Sep 2026 06:17:46 -0700 (PDT) X-Outflank-Mailman: Message body and most headers restored to incoming version X-BeenThere: xen-devel@lists.xenproject.org List-Id: Xen developer discussion List-Unsubscribe: , List-Post: List-Help: List-Subscribe: , Errors-To: xen-devel-bounces@lists.xenproject.org Precedence: list Sender: "Xen-devel" Authentication-Results: eu.smtp.expurgate.cloud; dkim=pass header.s=google header.d=toxicpanda.com header.i="@toxicpanda.com" header.h="Cc:To:In-Reply-To:References:Message-Id:Content-Transfer-Encoding:Content-Type:MIME-Version:Subject:Date:From" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=toxicpanda.com; s=google; t=1789478269; x=1790083069; darn=lists.xenproject.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=6aaaKNzMayGmjMwZUrTAjDsf1dO37tc7Q4XOVkfUDkI=; b=LI9pqqvYIoXl5Mr0o+z6D5mG+Q/pFzYWKGri+lONcUnSvWOmn6qKarwUffuXLS+Xuh ++/F5v1Y3GNQZDD5wefveGzjidioFc1rvbAFmBTnkE6x8DFhgP0OTOisgGAx+lQ+QHIX BXkllG3ooFi0yeD2NB/44MozbLubplBsFPuCyauMaHOY2gIEEBp882+WdWPIqrPYTYeo +FCdIl1YlQZNup5yIRgISJ5AMAKlSwflTvDN2YYxCXUWiDiYWH9h799tEd0kBUKtkQFE mcfhn+9oQNLLlsmpUDK88FII4rjoCKvZKAF9BUDJ96PAtSvX4QyOM8gt1rxp9txzHJxf BSCQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789478269; x=1790083069; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=6aaaKNzMayGmjMwZUrTAjDsf1dO37tc7Q4XOVkfUDkI=; b=zHguepcCqh0YQaLWp1hX8qEUBsJ//XjGvprqPOezVpeI6xtmKWjK98HuLzOuUrtfXZ 6HHko3ObiYoEPD4/PyD3WRHxNNaMGp5XlX/aAk3NwlQIbdIiUhu9gBjxOQbheGzcLbl9 YfhH3AHnSBuUr/uJoxCSQMqP7H4jLgiBJu2DSggmPvcak0+TS7wzlODWMhurAZJe1L5a xXLvBqYh48CHpTGNy3JfNO282xG8IBOBU/Xh4wTeEOSAfKwgEAsqjVKHFA74GpTPysg2 NxR4iEGkMwcIMrdsOippf7TYxvgLUN2A6YZUJ6D5babYC7hG+OHgWa93ffqyVhmlx9FT nXOg== X-Forwarded-Encrypted: i=1; AKwUvBx7ta9hS6feF3rqXJwbqdCPeBCLbt53UASyQrv9FTcbH6etIb6xcKpuh2MZo5q0x42+x9XdG8xQc+E=@lists.xenproject.org X-Gm-Message-State: AFuF++kPxHmlgdAFuG7mNO2uPoyoJ3W0qY6e5/HOO25kNkUW9pLjaAtZ AdCjCOC9Zs2zYmujtUgEyMMnwz6pkF5ZuoR8xFWU57ffu0uZ46A3cHE7O6TGdDF51yI= X-Gm-Gg: AYBFou18IYIGTf+6ZOEoKPfqViPwaOY+Emf6H2Z0pP3n95/QGdsMGc9uK9AvpMCnWFT OR1Jm4b6Lb7a4eLv0BSD+qC0hunvjLXkpLkraRUjnQ0vA9mH2Cs+0GrhfMK5eKiUqBPGHJ8KZ0L I/kxlENTtViAy9WDCzutCqtmtR/hvSdCwW6EPX987Y2sN8xZ2kcq6b1zyGN3MFNDRdrAh32HkmV KfXNbG/qBmx4sQ1e9962z2ROmnV9M7WpGIOLYlUHK17Jy2Ay753eBaZ2ZUVOgVgPpqtXJJnyqSt nglBsimGAKOrzZgc5dj3eyK/7ksP+wK6NuNdcJNONALYZmDY3iv9rM6Il+fZ13B1XJIxx+W8Hfu NtdYpysxZFMlhngT3i16tFGQ53SHDUxkcl6TpN3gaVeJqaCMBaPewRxX4mqi70p55d2CLahJc6V e1EfscCmSzZGGUEiaP+Ns5NmgZ9uPNLosEgPob936LdmZg4G5rXe1VMeAvrcyknqTTib398ozW6 7cjJFiVkg3nd/rSHQpPjK+uWoxkRZl8KXsWiym72l6gl/Icr3iYAVVy X-Received: by 2002:a53:ac8d:0:b0:66f:c1bc:c08e with SMTP id 956f58d0204a3-6715cbfa390mr343734d50.86.1789478267673; Tue, 15 Sep 2026 06:17:47 -0700 (PDT) From: Josef Bacik Date: Tue, 15 Sep 2026 13:17:28 +0000 Subject: [PATCH RFC v3 01/13] entry: Pass pt_regs to irqentry_exit_cond_resched() MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260915-b4-rcu-tasks-preempt-qs-v3-1-0ad30c4c5ee7@toxicpanda.com> References: <20260915-b4-rcu-tasks-preempt-qs-v3-0-0ad30c4c5ee7@toxicpanda.com> In-Reply-To: <20260915-b4-rcu-tasks-preempt-qs-v3-0-0ad30c4c5ee7@toxicpanda.com> To: "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Boqun Feng , Thomas Gleixner , Peter Zijlstra , Steven Rostedt , Masami Hiramatsu , Mark Rutland , Jiri Olsa , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , x86@kernel.org, Catalin Marinas , Will Deacon , Puranjay Mohan , Xu Kuohai , "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Boqun Feng , Thomas Gleixner , Peter Zijlstra , Steven Rostedt , Masami Hiramatsu , Mark Rutland , Jiri Olsa , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , x86@kernel.org, Catalin Marinas , Will Deacon , Puranjay Mohan , Xu Kuohai Cc: Andy Lutomirski , Josh Triplett , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Juergen Gross , Luis Chamberlain , Ihor Solodrai , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, xen-devel@lists.xenproject.org, Andy Lutomirski , Josh Triplett , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Juergen Gross , Luis Chamberlain , Ihor Solodrai , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, xen-devel@lists.xenproject.org, Josef Bacik X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openssh-sha256; t=1789478262; l=4149; i=josef@toxicpanda.com; h=from:subject:message-id; bh=4dQuHlI7KqFK9OlS3RCAgATHNrdOmuvuGUsCmQTn02w=; b=U1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgUBr36M/n0nWN0DNbnxwzIiCZez6MG JiruuNaSCI/zXsAAAAGcGF0YXR0AAAAAAAAAAZzaGE1MTIAAABTAAAAC3NzaC1lZDI1NTE5AAAA QEWkL7am6CPwY8WEkFAeUlQhZwIsKFKiG0LL4CTUEuVNlpBbsncWTVxiTecESmw1XMIxu5wivHh xkMDa7QjuUgk= X-Developer-Key: i=josef@toxicpanda.com; a=openssh; fpr=SHA256:C8kOX2QUJCMqnCX+KEeoqRAjLo9L+ELOSH2NSAJHqGA X-purgate-ID: tlsNG-ebf023/1789478270-50CDFB50-912E103E/0/0 X-purgate-type: clean X-purgate-size: 4151 X-ZohoMail-DKIM: pass (identity @toxicpanda.com) X-ZM-MESSAGEID: 1789478296837158500 The irq-exit preemption path is about to need the interrupted context's registers to decide whether the preemption may be reported to Tasks RCU as a quiescent state. irqentry_exit_to_kernel_mode_preempt() already has them; hand them down through irqentry_exit_cond_resched(), its PREEMPT_DYNAMIC static-call and static-key variants, and raw_irqentry_exit_cond_resched(). The only caller outside the generic entry code is Xen PV's upcall handler, which has regs as well. No functional change. Assisted-by: LLM Signed-off-by: Josef Bacik --- arch/x86/xen/enlighten_pv.c | 2 +- include/linux/irq-entry-common.h | 12 ++++++------ kernel/entry/common.c | 6 +++--- 3 files changed, 10 insertions(+), 10 deletions(-) diff --git a/arch/x86/xen/enlighten_pv.c b/arch/x86/xen/enlighten_pv.c index 2c64b388f616..3d85035f5624 100644 --- a/arch/x86/xen/enlighten_pv.c +++ b/arch/x86/xen/enlighten_pv.c @@ -739,7 +739,7 @@ __visible noinstr void xen_pv_evtchn_do_upcall(struct p= t_regs *regs) =20 inhcall =3D get_and_clear_inhcall(); if (inhcall && !WARN_ON_ONCE(state.exit_rcu)) { - irqentry_exit_cond_resched(); + irqentry_exit_cond_resched(regs); instrumentation_end(); restore_inhcall(inhcall); } else { diff --git a/include/linux/irq-entry-common.h b/include/linux/irq-entry-com= mon.h index 0bb6c03481fa..b811b469b0a7 100644 --- a/include/linux/irq-entry-common.h +++ b/include/linux/irq-entry-common.h @@ -346,21 +346,21 @@ typedef struct irqentry_state { * * Conditional reschedule with additional sanity checks. */ -void raw_irqentry_exit_cond_resched(void); +void raw_irqentry_exit_cond_resched(struct pt_regs *regs); =20 #ifdef CONFIG_PREEMPT_DYNAMIC #if defined(CONFIG_HAVE_PREEMPT_DYNAMIC_CALL) #define irqentry_exit_cond_resched_dynamic_enabled raw_irqentry_exit_cond_= resched #define irqentry_exit_cond_resched_dynamic_disabled NULL DECLARE_STATIC_CALL(irqentry_exit_cond_resched, raw_irqentry_exit_cond_res= ched); -#define irqentry_exit_cond_resched() static_call(irqentry_exit_cond_resche= d)() +#define irqentry_exit_cond_resched(regs) static_call(irqentry_exit_cond_re= sched)(regs) #elif defined(CONFIG_HAVE_PREEMPT_DYNAMIC_KEY) DECLARE_STATIC_KEY_TRUE(sk_dynamic_irqentry_exit_cond_resched); -void dynamic_irqentry_exit_cond_resched(void); -#define irqentry_exit_cond_resched() dynamic_irqentry_exit_cond_resched() +void dynamic_irqentry_exit_cond_resched(struct pt_regs *regs); +#define irqentry_exit_cond_resched(regs) dynamic_irqentry_exit_cond_resche= d(regs) #endif #else /* CONFIG_PREEMPT_DYNAMIC */ -#define irqentry_exit_cond_resched() raw_irqentry_exit_cond_resched() +#define irqentry_exit_cond_resched(regs) raw_irqentry_exit_cond_resched(re= gs) #endif /* CONFIG_PREEMPT_DYNAMIC */ =20 /** @@ -465,7 +465,7 @@ static inline void irqentry_exit_to_kernel_mode_preempt= (struct pt_regs *regs, return; =20 if (IS_ENABLED(CONFIG_PREEMPTION)) - irqentry_exit_cond_resched(); + irqentry_exit_cond_resched(regs); } =20 /** diff --git a/kernel/entry/common.c b/kernel/entry/common.c index e3d381fd3d25..e4acd50bd81a 100644 --- a/kernel/entry/common.c +++ b/kernel/entry/common.c @@ -134,7 +134,7 @@ static inline bool arch_irqentry_exit_need_resched(void= ); static inline bool arch_irqentry_exit_need_resched(void) { return true; } #endif =20 -void raw_irqentry_exit_cond_resched(void) +void raw_irqentry_exit_cond_resched(struct pt_regs *regs) { if (!preempt_count()) { /* Sanity check RCU and thread stack */ @@ -150,11 +150,11 @@ void raw_irqentry_exit_cond_resched(void) DEFINE_STATIC_CALL(irqentry_exit_cond_resched, raw_irqentry_exit_cond_resc= hed); #elif defined(CONFIG_HAVE_PREEMPT_DYNAMIC_KEY) DEFINE_STATIC_KEY_TRUE(sk_dynamic_irqentry_exit_cond_resched); -void dynamic_irqentry_exit_cond_resched(void) +void dynamic_irqentry_exit_cond_resched(struct pt_regs *regs) { if (!static_branch_unlikely(&sk_dynamic_irqentry_exit_cond_resched)) return; - raw_irqentry_exit_cond_resched(); + raw_irqentry_exit_cond_resched(regs); } #endif #endif --=20 2.55.0 From nobody Thu Sep 24 19:04:36 2026 Delivered-To: importer@patchew.org Received-SPF: pass (zohomail.com: domain of lists.xenproject.org designates 192.237.175.120 as permitted sender) client-ip=192.237.175.120; envelope-from=xen-devel-bounces@lists.xenproject.org; helo=lists.xenproject.org; Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of lists.xenproject.org designates 192.237.175.120 as permitted sender) smtp.mailfrom=xen-devel-bounces@lists.xenproject.org ARC-Seal: i=1; a=rsa-sha256; t=1789478289; cv=none; d=zohomail.com; s=zohoarc; b=LTclbCCk7kUclmzh36o+qxBDSnL26UxV/EUg8reLUL6hoxi12dw5zj8k11bYOGf8t0UxOR2uHVEaDxZ3kzoobIjEFiNwDVuQjFOBTyIKX/yESek+qa9P9cUZRgyLpVG/qJGESed5uKa5qgU0Jzr1SJgfoZE775RzfsCWeYKB+kI= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1789478289; h=Content-Type:Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=IVf/c/k7gzeyhXD1Q6M+2DLwLNI40n7fXEqvG2ZQN0k=; b=lm+/0S3//77iIDbUqGO3KOHtP0TFvMneLusSCrNKC2CZxU5diBIVH35TeZjZpqlVSPrTxFdI8mwnhAtT/jbpMCMI6iA+Cx+EB09n88kX3yfE2/BghfDP6hPfM84KJvmi8vA/eIIxC/r9G5DvxuxAt5ZX38f4VKc5F8F0QEN5zTc= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of lists.xenproject.org designates 192.237.175.120 as permitted sender) smtp.mailfrom=xen-devel-bounces@lists.xenproject.org Return-Path: Received: from lists.xenproject.org (lists.xenproject.org [192.237.175.120]) by mx.zohomail.com with SMTPS id 1789478289576486.4414089947212; Tue, 15 Sep 2026 06:18:09 -0700 (PDT) Received: from list by lists.xenproject.org with outflank-mailman.1421706.1647408 (Exim 4.92) (envelope-from ) id 1x6T2n-0004ez-OZ; Tue, 15 Sep 2026 13:17:53 +0000 Received: by outflank-mailman (output) from mailman id 1421706.1647408; Tue, 15 Sep 2026 13:17:53 +0000 Received: from localhost ([127.0.0.1] helo=lists.xenproject.org) by lists.xenproject.org with esmtp (Exim 4.92) (envelope-from ) id 1x6T2n-0004eq-Lt; Tue, 15 Sep 2026 13:17:53 +0000 Received: by outflank-mailman (input) for mailman id 1421706; Tue, 15 Sep 2026 13:17:52 +0000 Received: from mx.expurgate.net ([194.145.224.10]) by lists.xenproject.org with esmtp (Exim 4.92) (envelope-from ) id 1x6T2m-0004RF-KS for xen-devel@lists.xenproject.org; Tue, 15 Sep 2026 13:17:52 +0000 Received: from mx.expurgate.net (helo=localhost) by mx.expurgate.net with esmtp id 1x6T2m-00DQNP-1E for xen-devel@lists.xenproject.org; Tue, 15 Sep 2026 15:17:52 +0200 Received: from [10.42.69.11] (helo=localhost) by localhost with ESMTP (eXpurgate MTA 0.9.1) (envelope-from ) id 6aa94570-e002-0a2a0a5209dd-0a2a450babec-48 for ; Tue, 15 Sep 2026 15:17:51 +0200 Received: from [74.125.230.205] (helo=mail-qk2-f13.google.com) by tlsNG-42698a.mxtls.expurgate.net with ESMTPS (eXpurgate 4.57.1) (envelope-from ) id 6aa9457e-b7e8-0a2a450b0019-4a7de6cdd447-3 for ; Tue, 15 Sep 2026 15:17:51 +0200 Received: by mail-qk2-f13.google.com with SMTP id af79cd13be357-93a2218d7c2so67884885a.1 for ; Tue, 15 Sep 2026 06:17:51 -0700 (PDT) Received: from toxicpanda.com (ec2-34-228-114-98.compute-1.amazonaws.com. [34.228.114.98]) by smtp.gmail.com with ESMTPSA id af79cd13be357-939f19f94a1sm1137880185a.42.2026.09.15.06.17.48 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 15 Sep 2026 06:17:49 -0700 (PDT) X-Outflank-Mailman: Message body and most headers restored to incoming version X-BeenThere: xen-devel@lists.xenproject.org List-Id: Xen developer discussion List-Unsubscribe: , List-Post: List-Help: List-Subscribe: , Errors-To: xen-devel-bounces@lists.xenproject.org Precedence: list Sender: "Xen-devel" Authentication-Results: eu.smtp.expurgate.cloud; dkim=pass header.s=google header.d=toxicpanda.com header.i="@toxicpanda.com" header.h="Cc:To:In-Reply-To:References:Message-Id:Content-Transfer-Encoding:Content-Type:MIME-Version:Subject:Date:From" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=toxicpanda.com; s=google; t=1789478270; x=1790083070; darn=lists.xenproject.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=IVf/c/k7gzeyhXD1Q6M+2DLwLNI40n7fXEqvG2ZQN0k=; b=BDHNr5XS265fiyZv+nKnkdGvsCZz+AfecQhqLZ16Lbmq5KXvrTXGPDXjcpVb7K6HuX cuBWF4dZhFFusCnoB5VW9GuZ8tztdqGQ3V2FGK5SNKIrT5sWMf3358krpTKSB03uwP6c unP0FpBHUVAaTBVjYZaxdqnUFSEJMBKWcytpcXw2UKPhHY3oQ+50hLy5rZp6WFNv8gOR aJkEXy9N/wQj94ymlRYTJYo6IqJPQ2XcWp3T7BxrVgZT/HsZ1nbgGGB0fEd+2yKIAn3l qefageNO6GA9+qt+pehuW10h3JptI1mDaQQdIGSBDUx6jpxH7pWF71cOTLZkWAQHHrAz fSeg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789478270; x=1790083070; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=IVf/c/k7gzeyhXD1Q6M+2DLwLNI40n7fXEqvG2ZQN0k=; b=kN5yMI0+fccljiYTyYOb+IqRCpqE1S5tTseaIpHaPUhjFV8jOcbN4RbWc0hC75O8EZ 1/hOXL7Wf9lvcmfVnLJXQGYMUaue9qiL+n6rfrXdo//gysym4hnYQiuV0+lJvIE5XFTF LjDfGTdcpuLIsN/xl4sjG42Jsj1OudKS+P5qYYPUjVOLO008bT1Z0AnSrTajYeoEivTF iyV7Nj6wy1smYOCUhaMvklV52ajrx9mXfnY0kKTDSROJvUb8co7jGHIId+2wtGUZskhI MJ6KNWrMIH5C6mwwaoiVW9oEok3L1+UtKdXHchQbzYFMnc5yaq9nbdh3SoIVP3+2wldg /57g== X-Forwarded-Encrypted: i=1; AKwUvBwoI2WFvEQtvHwLBZGK5GyhxGBn4JPV7MuseUteg8Be6XtYaMr7Zg00DP94KJzMKxPafALzlU2X1Nc=@lists.xenproject.org X-Gm-Message-State: AFuF++ndQrZ3sdtlncP2arRnCETmFwrGKl93eFsWlYJQ7gZmTaB/dfdT uvViOcwHjPlAwGa0JL+ZEkw194a/z8V4gqwe1ORlTCpe096cYcl4GybEZS0IwL3s2hw= X-Gm-Gg: AYBFou0Vyz1GL0Z5QwKWbgwF/0rT4E9vB252vhwG1iFRl7RipISlMdXtjMm+Bt7X8oC yUrkJLoS/HBfy3Aactj6Ms3gtl7XY7KN0iJ8hj3r56WulBP0emZWTX3x9cAKk2mjvH7XVcx5pqV 4jR7XONNn/MB4zBzKV4yTQp6jXJWrN8l7fUeeZ4zVRgyJhq+Gf8sS61J0YJdTToiR2xKIG6kC8I kCid+AaRiUvjLOBWg+6uOMRalTIPEvCiKAT9W136oeVhkr6uhe0Fk6zyHEOCDl3cSYl3LASuUx0 julFLt927WzMnPPa2/BvdLLlRx8Un5Hs/g0WzusmEoXAN+I83HTbRBEyLsLR5UZsgTh5xNw0nVl 3qC+7I8/ImQcsNV4uY3rOdZKWoZXOKT8e8PXO9KZoFw9GBdQnpH/8u/roXyFzScGqwONGfpEs25 gIFqIbs8VDc2+eotmJug4zacSquL+hpqJxHL6YSEmHdLsuUlXEs+ifyWKbbq5ha5jYab8mT+Bym hPi6jkO0L/ws+ydwzQHrGPlfdjSqZIa5WKcXiwL6I3jwxXl8eoa7r6X X-Received: by 2002:a05:620a:2b8f:b0:936:d00a:4947 with SMTP id af79cd13be357-93a345745demr474274385a.2.1789478269641; Tue, 15 Sep 2026 06:17:49 -0700 (PDT) From: Josef Bacik Date: Tue, 15 Sep 2026 13:17:29 +0000 Subject: [PATCH RFC v3 02/13] rcu-tasks-trace: Inline rcu_read_lock_trace() and annotate inside the reader MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260915-b4-rcu-tasks-preempt-qs-v3-2-0ad30c4c5ee7@toxicpanda.com> References: <20260915-b4-rcu-tasks-preempt-qs-v3-0-0ad30c4c5ee7@toxicpanda.com> In-Reply-To: <20260915-b4-rcu-tasks-preempt-qs-v3-0-0ad30c4c5ee7@toxicpanda.com> To: "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Boqun Feng , Thomas Gleixner , Peter Zijlstra , Steven Rostedt , Masami Hiramatsu , Mark Rutland , Jiri Olsa , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , x86@kernel.org, Catalin Marinas , Will Deacon , Puranjay Mohan , Xu Kuohai , "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Boqun Feng , Thomas Gleixner , Peter Zijlstra , Steven Rostedt , Masami Hiramatsu , Mark Rutland , Jiri Olsa , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , x86@kernel.org, Catalin Marinas , Will Deacon , Puranjay Mohan , Xu Kuohai Cc: Andy Lutomirski , Josh Triplett , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Juergen Gross , Luis Chamberlain , Ihor Solodrai , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, xen-devel@lists.xenproject.org, Andy Lutomirski , Josh Triplett , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Juergen Gross , Luis Chamberlain , Ihor Solodrai , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, xen-devel@lists.xenproject.org, Josef Bacik X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openssh-sha256; t=1789478262; l=5472; i=josef@toxicpanda.com; h=from:subject:message-id; bh=WoSDn4ZEoocgxJI9vrvXQ3qPnrz/LPBQsbDWUONxsf4=; b=U1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgUBr36M/n0nWN0DNbnxwzIiCZez6MG JiruuNaSCI/zXsAAAAGcGF0YXR0AAAAAAAAAAZzaGE1MTIAAABTAAAAC3NzaC1lZDI1NTE5AAAA QOqgtWORMkFCOHo3WLOy6WMvyvf/Bd/PX1OTaW4H8LYu/zuiatDpxAZC81QUixDJwPnYFE7C1UL 57TTusaQZdwg= X-Developer-Key: i=josef@toxicpanda.com; a=openssh; fpr=SHA256:C8kOX2QUJCMqnCX+KEeoqRAjLo9L+ELOSH2NSAJHqGA X-purgate-ID: tlsNG-42698a/1789478271-1A4DB9EA-5F436B44/0/0 X-purgate-type: clean X-purgate-size: 5474 X-ZohoMail-DKIM: pass (identity @toxicpanda.com) X-ZM-MESSAGEID: 1789478291284158500 rcu_read_lock_trace() calls rcu_try_lock_acquire() before it has entered the SRCU-fast reader, and rcu_read_unlock_trace() calls srcu_lock_release() after it has left it. rcu_read_lock() and rcu_read_unlock() do it the other way around, annotating strictly inside the critical section, and rcu_read_lock_tasks_trace() already follows that order on the lock side. Make the trace variants match. Also make them, and the __srcu_read_lock_fast() and __srcu_read_unlock_fast() they are built on, __always_inline like rcu_read_lock() rather than leaving it to the compiler, which does outline all four in KASAN/KCOV builds. Besides consistency, this means the first thing a caller of rcu_read_lock_trace() does is enter the reader and the last thing rcu_read_unlock_trace() does is leave it, with no out-of-line call on the outside. A later patch relies on that for callers whose own text is protected by the reader they are about to take. Assisted-by: LLM Signed-off-by: Josef Bacik --- include/linux/rcupdate_trace.h | 22 ++++++++++------------ include/linux/srcutiny.h | 4 ++-- include/linux/srcutree.h | 5 +++-- 3 files changed, 15 insertions(+), 16 deletions(-) diff --git a/include/linux/rcupdate_trace.h b/include/linux/rcupdate_trace.h index 273c59a03251..4035054309d7 100644 --- a/include/linux/rcupdate_trace.h +++ b/include/linux/rcupdate_trace.h @@ -93,22 +93,20 @@ static inline void rcu_read_unlock_tasks_trace(struct s= rcu_ctr __percpu *scp) * * For more details, please see the documentation for rcu_read_lock(). */ -static inline void rcu_read_lock_trace(void) +static __always_inline void rcu_read_lock_trace(void) { int n; struct task_struct *t =3D current; =20 - rcu_try_lock_acquire(&rcu_tasks_trace_srcu_struct.dep_map); n =3D READ_ONCE(t->trc_reader_nesting); WRITE_ONCE(t->trc_reader_nesting, n + 1); - if (n) { - // In case we interrupted a Tasks Trace RCU reader. - return; - } - barrier(); // nesting before scp to protect against interrupt handler. - t->trc_reader_scp =3D __srcu_read_lock_fast(&rcu_tasks_trace_srcu_struct); - if (!IS_ENABLED(CONFIG_TASKS_TRACE_RCU_NO_MB)) - smp_mb(); // Placeholder for more selective ordering + if (!n) { + barrier(); // nesting before scp to protect against interrupt handler. + t->trc_reader_scp =3D __srcu_read_lock_fast(&rcu_tasks_trace_srcu_struct= ); + if (!IS_ENABLED(CONFIG_TASKS_TRACE_RCU_NO_MB)) + smp_mb(); // Placeholder for more selective ordering + } // Else we interrupted a Tasks Trace RCU reader. + rcu_try_lock_acquire(&rcu_tasks_trace_srcu_struct.dep_map); } =20 /** @@ -120,12 +118,13 @@ static inline void rcu_read_lock_trace(void) * * For more details, please see the documentation for rcu_read_unlock(). */ -static inline void rcu_read_unlock_trace(void) +static __always_inline void rcu_read_unlock_trace(void) { int n; struct srcu_ctr __percpu *scp; struct task_struct *t =3D current; =20 + srcu_lock_release(&rcu_tasks_trace_srcu_struct.dep_map); n =3D READ_ONCE(t->trc_reader_nesting) - 1; if (n) { WRITE_ONCE(t->trc_reader_nesting, n); @@ -137,7 +136,6 @@ static inline void rcu_read_unlock_trace(void) smp_mb(); // Placeholder for more selective ordering __srcu_read_unlock_fast(&rcu_tasks_trace_srcu_struct, scp); } - srcu_lock_release(&rcu_tasks_trace_srcu_struct.dep_map); } =20 /** diff --git a/include/linux/srcutiny.h b/include/linux/srcutiny.h index fbcf13bc12d1..a43bae11c81c 100644 --- a/include/linux/srcutiny.h +++ b/include/linux/srcutiny.h @@ -101,13 +101,13 @@ static inline struct srcu_ctr __percpu *__srcu_ctr_to= _ptr(struct srcu_struct *ss return (struct srcu_ctr __percpu *)(intptr_t)idx; } =20 -static inline struct srcu_ctr __percpu *__srcu_read_lock_fast(struct srcu_= struct *ssp) +static __always_inline struct srcu_ctr __percpu *__srcu_read_lock_fast(str= uct srcu_struct *ssp) __acquires_shared(ssp) { return __srcu_ctr_to_ptr(ssp, __srcu_read_lock(ssp)); } =20 -static inline void __srcu_read_unlock_fast(struct srcu_struct *ssp, struct= srcu_ctr __percpu *scp) +static __always_inline void __srcu_read_unlock_fast(struct srcu_struct *ss= p, struct srcu_ctr __percpu *scp) __releases_shared(ssp) { __srcu_read_unlock(ssp, __srcu_ptr_to_ctr(ssp, scp)); diff --git a/include/linux/srcutree.h b/include/linux/srcutree.h index 75e54e4f963f..fdb42ab50301 100644 --- a/include/linux/srcutree.h +++ b/include/linux/srcutree.h @@ -286,7 +286,8 @@ static inline struct srcu_ctr __percpu *__srcu_ctr_to_p= tr(struct srcu_struct *ss * on architectures that support NMIs but do not supply NMI-safe * implementations of this_cpu_inc(). */ -static inline struct srcu_ctr __percpu notrace *__srcu_read_lock_fast(stru= ct srcu_struct *ssp) +static __always_inline struct srcu_ctr __percpu notrace * +__srcu_read_lock_fast(struct srcu_struct *ssp) __acquires_shared(ssp) { struct srcu_ctr __percpu *scp =3D READ_ONCE(ssp->srcu_ctrp); @@ -309,7 +310,7 @@ static inline struct srcu_ctr __percpu notrace *__srcu_= read_lock_fast(struct src * Please see the __srcu_read_lock_fast() function's header comment for * information on implicit RCU readers and NMI safety. */ -static inline void notrace +static __always_inline void notrace __srcu_read_unlock_fast(struct srcu_struct *ssp, struct srcu_ctr __percpu = *scp) __releases_shared(ssp) { --=20 2.55.0 From nobody Thu Sep 24 19:04:36 2026 Delivered-To: importer@patchew.org Received-SPF: pass (zohomail.com: domain of lists.xenproject.org designates 192.237.175.120 as permitted sender) client-ip=192.237.175.120; envelope-from=xen-devel-bounces@lists.xenproject.org; helo=lists.xenproject.org; Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of lists.xenproject.org designates 192.237.175.120 as permitted sender) smtp.mailfrom=xen-devel-bounces@lists.xenproject.org ARC-Seal: i=1; a=rsa-sha256; t=1789478299; cv=none; d=zohomail.com; s=zohoarc; b=Vv84U3BtWgB0LJmBBa8u8bi6Tg2C295swlWDombNJxt9JJlWnUOiwJH5Fn1MEoUeUJU8r8B4C8pcK31IJQ4Msm8Sq7XFZ/sOUYoDMKX8dlxdGK4T4ZREMBm90CsJdI2vGHC60EYq3zDRPlIZvaNMxHOTBc8hrYymcrqHeCNccLk= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1789478299; h=Content-Type:Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=fYtErF6vBN+NnnD1JqJDCU7rpx142OZs1HR5OhHJnD0=; b=UwoEQcUXMnQ42qc9MFVdJDe8ioYz4G+odrd+SV5qNda2goPQXWjTZeXUTCtJZfaS/kyDcnHLlQbqWv/eK8lTwYrOLe05Bon2sTyWGT/JwG5hHUzr/9AB3sCI47d82ES2OfWT9uxzBkjCYG9JASVw5qVZuXWzW4YePf0Leej4ZiI= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of lists.xenproject.org designates 192.237.175.120 as permitted sender) smtp.mailfrom=xen-devel-bounces@lists.xenproject.org Return-Path: Received: from lists.xenproject.org (lists.xenproject.org [192.237.175.120]) by mx.zohomail.com with SMTPS id 178947829936112.693869687314077; Tue, 15 Sep 2026 06:18:19 -0700 (PDT) Received: from list by lists.xenproject.org with outflank-mailman.1421707.1647417 (Exim 4.92) (envelope-from ) id 1x6T2r-0004vD-25; Tue, 15 Sep 2026 13:17:57 +0000 Received: by outflank-mailman (output) from mailman id 1421707.1647417; Tue, 15 Sep 2026 13:17:57 +0000 Received: from localhost ([127.0.0.1] helo=lists.xenproject.org) by lists.xenproject.org with esmtp (Exim 4.92) (envelope-from ) id 1x6T2q-0004v6-Uz; Tue, 15 Sep 2026 13:17:56 +0000 Received: by outflank-mailman (input) for mailman id 1421707; Tue, 15 Sep 2026 13:17:55 +0000 Received: from mx.expurgate.net ([194.145.224.10]) by lists.xenproject.org with esmtp (Exim 4.92) (envelope-from ) id 1x6T2p-0004sv-8m for xen-devel@lists.xenproject.org; Tue, 15 Sep 2026 13:17:55 +0000 Received: from mx.expurgate.net (helo=localhost) by mx.expurgate.net with esmtp id 1x6T2o-00DQNP-Lg for xen-devel@lists.xenproject.org; Tue, 15 Sep 2026 15:17:54 +0200 Received: from [10.42.69.9] (helo=localhost) by localhost with ESMTP (eXpurgate MTA 0.9.1) (envelope-from ) id 6aa94581-e002-0a2a0a5209dd-0a2a45099780-4 for ; Tue, 15 Sep 2026 15:17:54 +0200 Received: from [209.85.219.54] (helo=mail-qv1-f54.google.com) by tlsNG-bad1c0.mxtls.expurgate.net with ESMTPS (eXpurgate 4.57.1) (envelope-from ) id 6aa94580-be1a-0a2a45090019-d155db36a875-3 for ; Tue, 15 Sep 2026 15:17:53 +0200 Received: by mail-qv1-f54.google.com with SMTP id 6a1803df08f44-90e9ad1a373so7400796d6.0 for ; Tue, 15 Sep 2026 06:17:53 -0700 (PDT) Received: from toxicpanda.com (ec2-34-228-114-98.compute-1.amazonaws.com. [34.228.114.98]) by smtp.gmail.com with ESMTPSA id 6a1803df08f44-912321108a5sm31206596d6.46.2026.09.15.06.17.50 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 15 Sep 2026 06:17:50 -0700 (PDT) X-Outflank-Mailman: Message body and most headers restored to incoming version X-BeenThere: xen-devel@lists.xenproject.org List-Id: Xen developer discussion List-Unsubscribe: , List-Post: List-Help: List-Subscribe: , Errors-To: xen-devel-bounces@lists.xenproject.org Precedence: list Sender: "Xen-devel" Authentication-Results: eu.smtp.expurgate.cloud; dkim=pass header.s=google header.d=toxicpanda.com header.i="@toxicpanda.com" header.h="Cc:To:In-Reply-To:References:Message-Id:Content-Transfer-Encoding:Content-Type:MIME-Version:Subject:Date:From" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=toxicpanda.com; s=google; t=1789478272; x=1790083072; darn=lists.xenproject.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=fYtErF6vBN+NnnD1JqJDCU7rpx142OZs1HR5OhHJnD0=; b=Vjm83Eu+tPVcteSMCRWXtdBwGLzYwSaSEdhiG55n9GZVTzotJYAc1HLQ6y6KBSf1iz GApzeln3B22k7qBhuAN8tqJazcFv5T8nZ0IWli+Pe5LDg7vI2nrQjPQowydDOjszlawi edRGeYa8FaqM7bG7lqL5LHDN7Lcq8Hin+BbkaJNL+EXyIPM/sxRxQ818EP/OBR+6/avU LZxigtEn1x/+bZxUjFdYL8/FBtNtUp8Q8N+CdDZ+XedhFmBvVZpegaF77D9YdOlLWVqg cAW87553JEu2RymSq+l1OosIBEtBPNIf9R7WMP12HyRXVolHusZPYiDDzBo9J5dmktUe U+BA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789478272; x=1790083072; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=fYtErF6vBN+NnnD1JqJDCU7rpx142OZs1HR5OhHJnD0=; b=qpE14rOYPRnDO+TWuLAzY+U/FXqY54C5s5UzwipCVSTQd/4LkW0bLXcmOBxIHnKClE gSjwCBgWg/gwrEf4Eyr7um0jSccTAMMZF+gH2BmH/nUXaVVdsEy2qK52M492+16MRxa8 PyWxX/m+2vnclkf4y8Sq+deit3bS5nHODS1Gst/NtF6uShboQDHuqqHjQpsQLr6fo12k 1EzVpvQv74hrAAbQ4AQsx+7U2qF7JNX+ORKwx39Y63nx3yIWHnpll9tj0h7w6LZ7n2sD O3f0fwLSIaqldAzdj5QrDtY0iAEzVe7McxJQ7pPRmQ+0BbfB0dc760FEvvpTflVzWa5g Oe4g== X-Forwarded-Encrypted: i=1; AKwUvBwKx9NwO2Emgm0c6am6Gxt5dcR4Ryi31LEdjdZG4fxYp+7DJRX9+6TglFevgTqCxmgeKAfLAuqixug=@lists.xenproject.org X-Gm-Message-State: AFuF++mS7e7fYRXlsQriPabgILhGRXHXvEk2fUZvDRntivjH7pimAx1K zWf3NBTJ6QkRj2LVfjgDyvn3qBPruB32k/F9MN1FQ1+kMT3yzO/xTzBkTlamA0qkzRE= X-Gm-Gg: AYBFou0gj9oUfvtqwhuT0SPHTnKyTfK5Z60bhFx545PSstsM4cdJrEvCtYn9JJ8pYwJ vAlhDfeQvdLmVvfq+LYB5zrCfX8ZDQQfl9drPIP0gdkyOvYqWdITQmDaUh6ejwZXizOe1KPaqad cvClcvcucz5GEo9qYgjtPGKW45A3nERGF5ZMHL/FlTwuLzmRvQoMxG/yK+XM4Z6LB/MeYyH+V6p k+3svzW2uAtparhlFRjxRkCCeTwSh4iELJHHUkB1Ejc0LkKeb04I9X9p4306lKXZMxb5hMZVZB9 LCDk3HC9zXWO0ovGaPQbEfC5CQttMUS77oRozM1LrHU0mwz6IwPEUbCjT0hC76l/A+EvQ0+Tc3U sbBujT5c3MFHYfZAr20T7q68QmnjUCUJJXnlafVilrI5MFPy+/enmgPAxDtznFW43La7Nr1Bkjw M6fWI8Z/KD3iYcm5omIJ4sVG3WFXJUuZwv9DENH/pR0AiwtMzX/E2KbJ/hVh6shA/s7rc8WdlLi rAox+mutTAJd9Ry+ikTO/xozvE1ml4g+cSQOxOFdOa05Nz3dlN2Svgg X-Received: by 2002:ad4:5aa4:0:b0:911:2a7c:65b9 with SMTP id 6a1803df08f44-912344347ccmr49325266d6.30.1789478271691; Tue, 15 Sep 2026 06:17:51 -0700 (PDT) From: Josef Bacik Date: Tue, 15 Sep 2026 13:17:30 +0000 Subject: [PATCH RFC v3 03/13] rcu-tasks: Add a Tasks RCU implementation for reader-marked trampolines MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260915-b4-rcu-tasks-preempt-qs-v3-3-0ad30c4c5ee7@toxicpanda.com> References: <20260915-b4-rcu-tasks-preempt-qs-v3-0-0ad30c4c5ee7@toxicpanda.com> In-Reply-To: <20260915-b4-rcu-tasks-preempt-qs-v3-0-0ad30c4c5ee7@toxicpanda.com> To: "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Boqun Feng , Thomas Gleixner , Peter Zijlstra , Steven Rostedt , Masami Hiramatsu , Mark Rutland , Jiri Olsa , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , x86@kernel.org, Catalin Marinas , Will Deacon , Puranjay Mohan , Xu Kuohai , "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Boqun Feng , Thomas Gleixner , Peter Zijlstra , Steven Rostedt , Masami Hiramatsu , Mark Rutland , Jiri Olsa , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , x86@kernel.org, Catalin Marinas , Will Deacon , Puranjay Mohan , Xu Kuohai Cc: Andy Lutomirski , Josh Triplett , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Juergen Gross , Luis Chamberlain , Ihor Solodrai , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, xen-devel@lists.xenproject.org, Andy Lutomirski , Josh Triplett , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Juergen Gross , Luis Chamberlain , Ihor Solodrai , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, xen-devel@lists.xenproject.org, Josef Bacik X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openssh-sha256; t=1789478262; l=32683; i=josef@toxicpanda.com; h=from:subject:message-id; bh=wQeEStQPLpkB0ZukCY5vknBzMvvV//pSfULH7QtsQGc=; b=U1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgUBr36M/n0nWN0DNbnxwzIiCZez6MG JiruuNaSCI/zXsAAAAGcGF0YXR0AAAAAAAAAAZzaGE1MTIAAABTAAAAC3NzaC1lZDI1NTE5AAAA QMM+dAGbEOlns74BRRoxfkCG0nJCZGoqB9FIgpGikpOfYMXm9qc7IWrJfu93xYxIrOxZ1yXrj33 aa/TFXf8Q9wU= X-Developer-Key: i=josef@toxicpanda.com; a=openssh; fpr=SHA256:C8kOX2QUJCMqnCX+KEeoqRAjLo9L+ELOSH2NSAJHqGA X-purgate-ID: tlsNG-bad1c0/1789478273-FD26A034-5D7126A2/0/0 X-purgate-type: clean X-purgate-size: 32685 X-ZohoMail-DKIM: pass (identity @toxicpanda.com) X-ZM-MESSAGEID: 1789478301069158500 Tasks RCU waits for every task to pass through a voluntary context switch, usermode or idle, because a preempted task might be sitting in a trampoline that is about to be freed and nothing marks it as such. With PREEMPT_LAZY that is a poor fit for servers: cond_resched() is a no-op, so a CPU-bound kthread only ever leaves the CPU by preemption, and one such kthread holds every synchronize_rcu_tasks() caller -- ftrace and BPF trampoline teardown under their mutexes, the kprobe jump optimizer under text_mutex and cpus_read_lock() -- hostage for as long as it runs. Following the discussion on v2, take the other road: let the architecture make its trampolines Tasks Trace RCU readers. When an architecture selects HAVE_RCU_TRAMPOLINE_READERS it promises that every trampoline whose lifetime Tasks RCU guards enters rcu_read_lock_trace() (or its assembly equivalent) before calling out and leaves it before returning, so a task anywhere inside such a call-out, preempted or not, is an ordinary Tasks Trace reader. That leaves the few instructions of trampoline text before the reader is entered and after it is left (plus, in a later patch, the bytes a kprobe jump optimization is about to overwrite). A task can only linger there by being interrupted there, and such text never calls anything that schedules, so instead of tracking tasks we track CPUs: every pass through __schedule() is a per-CPU quiescent event, except that the one context switch that can catch a task at an arbitrary instruction -- a preemption from irq exit -- first records the interrupted IP in the task and parks it on a per-CPU list for the duration (reusing the fields and lists the classic flavor keeps for its exit-path bookkeeping), and, if the IP is inside such "unmarked" text, puts the task on a short holdout list; the task takes itself off at its next context switch outside such a preemption or irq-exit check that finds it elsewhere. Usermode (the existing tick hook, or a nohz_full CPU in an RCU extended quiescent state) and idle count as well. rcu_tasks_trampoline_text() does the classification: anything outside core and module text, a new .text..rcu_tramp section for C glue that trampolines call before it has entered the reader (__rcu_trampoline), and an arch hook for things like static ftrace stubs and return thunks. The grace period, run by the existing rcu_tasks kthread so that call_rcu_tasks(), synchronize_rcu_tasks() and rcu_barrier_tasks() keep their names and callers, is: wait for every online non-idle CPU to context switch (nudging stragglers with resched_cpu() after a jiffy), drain the holdout list as it stood, synchronize_rcu_tasks_trace() for everything inside the readers, then one more CPU pass and drain for tasks that have since left the reader into the trailing instructions. That is bounded by a few jiffies, preempt-off latency and an SRCU grace period rather than by the longest stretch any task runs without sleeping, needs no per-task scan, and makes cond_resched_tasks_rcu_qs() unnecessary on such architectures. As before, idle tasks are not waited for. rcu_tasks_wait_irq_preempted() walks the parked lists for the one caller (the kprobe jump optimizer, later in the series) that makes ordinary text unsafe to be parked in and so has to wait out tasks that were preempted there before it said so. The classic implementation is untouched and remains the default; the new one is built only as CONFIG_TASKS_RCU_TRAMPOLINE_READERS when the architecture opts in and uses the generic irq entry code, whose reschedule check gains the rcu_tasks_irq_resched() call. Nothing selects it yet. Suggested-by: Paul E. McKenney Suggested-by: Alexei Starovoitov Assisted-by: LLM Signed-off-by: Josef Bacik --- include/asm-generic/vmlinux.lds.h | 11 + include/linux/rcupdate.h | 32 ++- include/linux/sched.h | 1 + kernel/entry/common.c | 8 +- kernel/fork.c | 1 + kernel/rcu/Kconfig | 22 ++ kernel/rcu/tasks.h | 460 ++++++++++++++++++++++++++++++++++= +++- kernel/rcu/update.c | 2 + 8 files changed, 528 insertions(+), 9 deletions(-) diff --git a/include/asm-generic/vmlinux.lds.h b/include/asm-generic/vmlinu= x.lds.h index b2988aa12f66..86e58c4fe370 100644 --- a/include/asm-generic/vmlinux.lds.h +++ b/include/asm-generic/vmlinux.lds.h @@ -571,6 +571,16 @@ __cpuidle_text_end =3D .; \ __noinstr_text_end =3D .; =20 +/* + * C glue called directly from Tasks-RCU-protected trampolines, bounded so + * that rcu_tasks_trampoline_text() can recognise it; see __rcu_trampoline. + */ +#define RCU_TRAMP_TEXT \ + ALIGN_FUNCTION(); \ + __rcu_tramp_text_start =3D .; \ + *(.text..rcu_tramp) \ + __rcu_tramp_text_end =3D .; + #define TEXT_SPLIT \ __split_text_start =3D .; \ *(.text.split .text.split.[0-9a-zA-Z_]*) \ @@ -607,6 +617,7 @@ TEXT_HOT \ *(TEXT_MAIN .text.fixup) \ NOINSTR_TEXT \ + RCU_TRAMP_TEXT \ *(.ref.text) =20 /* sched.text is aling to function alignment to secure we have same diff --git a/include/linux/rcupdate.h b/include/linux/rcupdate.h index 44c07a66edff..fb2a3889a696 100644 --- a/include/linux/rcupdate.h +++ b/include/linux/rcupdate.h @@ -50,6 +50,31 @@ token_context_lock_instance(RCU, RCU_BH); /* Exported common interfaces */ void call_rcu(struct rcu_head *head, rcu_callback_t func); void rcu_barrier_tasks(void); + +/* + * Trampoline-reader Tasks RCU (CONFIG_TASKS_RCU_TRAMPOLINE_READERS), see + * kernel/rcu/tasks.h. rcu_tasks_irq_resched_enter()/_exit() bracket the + * irq-exit preemption; rcu_tasks_trampoline_text() and the arch_ override + * classify an interrupted IP; rcu_tasks_wait_irq_preempted() lets a caller + * wait out tasks already preempted somewhere it is about to make unsafe. + * __rcu_trampoline places C code that such trampolines call directly, bef= ore + * it has entered its Tasks Trace reader, where that classification can se= e it. + */ +void rcu_tasks_irq_resched_enter(unsigned long ip); +void rcu_tasks_irq_resched_exit(void); +bool rcu_tasks_trampoline_text(unsigned long ip); +bool arch_rcu_tasks_trampoline_text(unsigned long ip); +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS +void rcu_tasks_wait_irq_preempted(bool (*inside)(unsigned long ip)); +#else +static inline void rcu_tasks_wait_irq_preempted(bool (*inside)(unsigned lo= ng ip)) { } +#endif +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS +/* Also keeps instrumentation calls out of the prologue, ahead of the read= er. */ +#define __rcu_trampoline __noinstr_section(".text..rcu_tramp") +#else +#define __rcu_trampoline +#endif void synchronize_rcu(void); =20 /* @@ -180,11 +205,16 @@ static inline void rcu_nocb_flush_deferred_wakeup(voi= d) { } #ifdef CONFIG_TASKS_RCU_GENERIC =20 # ifdef CONFIG_TASKS_RCU -# define rcu_tasks_classic_qs(t, preempt) \ +# ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS +void rcu_tasks_note_qs(struct task_struct *t, bool preempt); +# define rcu_tasks_classic_qs(t, preempt) rcu_tasks_note_qs((t), (preempt= )) +# else +# define rcu_tasks_classic_qs(t, preempt) \ do { \ if (!(preempt) && READ_ONCE((t)->rcu_tasks_holdout)) \ WRITE_ONCE((t)->rcu_tasks_holdout, false); \ } while (0) +# endif void call_rcu_tasks(struct rcu_head *head, rcu_callback_t func); void synchronize_rcu_tasks(void); void rcu_tasks_torture_stats_print(char *tt, char *tf); diff --git a/include/linux/sched.h b/include/linux/sched.h index 8b3d47a325cc..15beb44caa2c 100644 --- a/include/linux/sched.h +++ b/include/linux/sched.h @@ -957,6 +957,7 @@ struct task_struct { u8 rcu_tasks_holdout; u8 rcu_tasks_idx; int rcu_tasks_idle_cpu; + unsigned long rcu_tasks_irq_ip; struct list_head rcu_tasks_holdout_list; int rcu_tasks_exit_cpu; struct list_head rcu_tasks_exit_list; diff --git a/kernel/entry/common.c b/kernel/entry/common.c index e4acd50bd81a..94318519998c 100644 --- a/kernel/entry/common.c +++ b/kernel/entry/common.c @@ -6,6 +6,7 @@ #include #include #include +#include #include #include =20 @@ -141,8 +142,13 @@ void raw_irqentry_exit_cond_resched(struct pt_regs *re= gs) rcu_irq_exit_check_preempt(); if (IS_ENABLED(CONFIG_DEBUG_ENTRY)) WARN_ON_ONCE(!on_thread_stack()); - if (need_resched() && arch_irqentry_exit_need_resched()) + if (need_resched() && arch_irqentry_exit_need_resched()) { + if (IS_ENABLED(CONFIG_TASKS_RCU_TRAMPOLINE_READERS)) + rcu_tasks_irq_resched_enter(instruction_pointer(regs)); preempt_schedule_irq(); + if (IS_ENABLED(CONFIG_TASKS_RCU_TRAMPOLINE_READERS)) + rcu_tasks_irq_resched_exit(); + } } } #ifdef CONFIG_PREEMPT_DYNAMIC diff --git a/kernel/fork.c b/kernel/fork.c index 416758c8a3d4..8077336bb136 100644 --- a/kernel/fork.c +++ b/kernel/fork.c @@ -1871,6 +1871,7 @@ static inline void rcu_copy_process(struct task_struc= t *p) p->rcu_tasks_holdout =3D false; INIT_LIST_HEAD(&p->rcu_tasks_holdout_list); p->rcu_tasks_idle_cpu =3D -1; + p->rcu_tasks_irq_ip =3D 0; INIT_LIST_HEAD(&p->rcu_tasks_exit_list); #endif /* #ifdef CONFIG_TASKS_RCU */ #ifdef CONFIG_TASKS_TRACE_RCU diff --git a/kernel/rcu/Kconfig b/kernel/rcu/Kconfig index 332df7a7a634..bbab14bc14c3 100644 --- a/kernel/rcu/Kconfig +++ b/kernel/rcu/Kconfig @@ -107,6 +107,28 @@ config TASKS_RCU default NEED_TASKS_RCU && PREEMPTION select IRQ_WORK =20 +config HAVE_RCU_TRAMPOLINE_READERS + bool + help + Select this if the architecture uses the generic irq entry code and + every trampoline whose lifetime Tasks RCU guards on it (ftrace + trampolines, kprobe out-of-line and optimized-probe slots, BPF + trampolines, out-of-line ftrace direct-call trampolines) enters a + Tasks Trace RCU read-side critical section before calling out of + the trampoline and leaves it before returning, and any core text + that runs on behalf of such a trampoline outside that reader is + reported by arch_rcu_tasks_trampoline_text(). The assembly readers + use the this_cpu_inc() form of SRCU-fast, hence !NEED_SRCU_NMI_SAFE. + +config TASKS_RCU_TRAMPOLINE_READERS + def_bool TASKS_RCU && HAVE_RCU_TRAMPOLINE_READERS && GENERIC_IRQ_ENTRY &&= !NEED_SRCU_NMI_SAFE + select TASKS_TRACE_RCU + help + Implement the Tasks RCU grace period as a per-CPU pass over + context switches and irq-exit reschedules outside trampoline text + plus a Tasks Trace RCU grace period, instead of waiting for every + task to voluntarily context switch. See kernel/rcu/tasks.h. + config FORCE_TASKS_RUDE_RCU bool "Force selection of Tasks Rude RCU" depends on RCU_EXPERT diff --git a/kernel/rcu/tasks.h b/kernel/rcu/tasks.h index 627295396cd9..3a7c092361a6 100644 --- a/kernel/rcu/tasks.h +++ b/kernel/rcu/tasks.h @@ -152,7 +152,7 @@ static struct rcu_tasks rt_name =3D \ .kname =3D #rt_name, \ } =20 -#ifdef CONFIG_TASKS_RCU +#if defined(CONFIG_TASKS_RCU) && !defined(CONFIG_TASKS_RCU_TRAMPOLINE_READ= ERS) =20 /* Report delay of scan exiting tasklist in rcu_tasks_postscan(). */ static void tasks_rcu_exit_stall(struct timer_list *unused); @@ -802,7 +802,7 @@ static void rcu_tasks_torture_stats_print_generic(struc= t rcu_tasks *rtp, char *t =20 #endif // #ifndef CONFIG_TINY_RCU =20 -#if defined(CONFIG_TASKS_RCU) +#if defined(CONFIG_TASKS_RCU) && !defined(CONFIG_TASKS_RCU_TRAMPOLINE_READ= ERS) =20 //////////////////////////////////////////////////////////////////////// // @@ -897,10 +897,445 @@ static void rcu_tasks_wait_gp(struct rcu_tasks *rtp) rtp->postgp_func(rtp); } =20 -#endif /* #if defined(CONFIG_TASKS_RCU) */ +#endif /* #if defined(CONFIG_TASKS_RCU) && !defined(CONFIG_TASKS_RCU_TRAMP= OLINE_READERS) */ =20 #ifdef CONFIG_TASKS_RCU =20 +static int rcu_tasks_lazy_ms =3D -1; +module_param(rcu_tasks_lazy_ms, int, 0444); + +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS + +//////////////////////////////////////////////////////////////////////// +// +// Tasks RCU for architectures whose trampolines are Tasks Trace RCU +// readers (CONFIG_HAVE_RCU_TRAMPOLINE_READERS). +// +// On these architectures every piece of text whose lifetime Tasks RCU +// guards -- ftrace trampolines, kprobe optinsn slots, BPF trampoline +// images, out-of-line ftrace direct-call trampolines -- enters a Tasks +// Trace RCU read-side critical section before calling out of itself and +// leaves it before returning, so a task anywhere inside such a call-out, +// preempted or not, is an ordinary rcu_read_lock_trace() reader and +// synchronize_rcu_tasks_trace() waits for it. +// +// What that cannot cover is the handful of instructions in the trampoline +// before the reader is entered and after it is left, and the one user that +// has no trampoline at all: the bytes after a kprobe that the jump +// optimizer is about to overwrite. A task can only linger in such +// "unmarked" text by being interrupted there; unmarked text never calls +// anything that could schedule. So a context switch on a CPU tells us th= at +// whatever that CPU was running is out of unmarked text, with one +// exception: a preemption from the irq-exit path, which can happen at any +// instruction boundary. That path has the interrupted pt_regs in hand, so +// just before it preempts it records the IP in the task and checks it +// (rcu_tasks_trampoline_text()); if it is inside unmarked text the task +// goes on a short holdout list first, and takes itself off again at its +// next context switch outside such a preemption or its next irq-exit +// check that finds it elsewhere. With that, every pass through +// __schedule() is a per-CPU quiescent event, as are usermode and idle. +// +// A grace period is then: +// +// 1. Wait for every online, non-idle CPU to context switch, nudging +// stragglers with resched_cpu(). Afterwards no task is in the leading +// unmarked instructions of a dying trampoline unless it is on the +// holdout list. +// 2. Wait for the holdout list (as it stood) to drain. +// 3. synchronize_rcu_tasks_trace(), for everything inside the readers. +// 4. Repeat 1 and 2 for tasks that have since left the reader and are in +// the trailing unmarked instructions. +// +// which is bounded by a few jiffies plus preempt-off latency plus an SRCU +// grace period, independent of how long any task runs without sleeping. +// As with the classic implementation, the idle tasks are not waited for. + +static void rcu_tasks_tramp_wait_gp(struct rcu_tasks *rtp); +void call_rcu_tasks(struct rcu_head *rhp, rcu_callback_t func); +DEFINE_RCU_TASKS(rcu_tasks, rcu_tasks_tramp_wait_gp, call_rcu_tasks, "RCU = Tasks"); + +/* Per-CPU count of Tasks RCU quiescent events, and the GP kthread's snaps= hot. */ +static DEFINE_PER_CPU(unsigned long, rcu_tasks_qs_seq); +static DEFINE_PER_CPU(unsigned long, rcu_tasks_qs_snap); + +/* + * Tasks currently switched out by an irq-exit preemption are kept, with t= he + * interrupted IP, on the per-CPU rtp_exit_list of the CPU that preempted = them + * (reusing the list, lock and task_struct fields the classic flavor uses = for + * its exit-path bookkeeping, which this flavor does not need), so that + * rcu_tasks_wait_irq_preempted() can find them without a tasklist scan and + * regardless of where they are in exit. + */ + +/* Tasks last seen preempted inside unmarked trampoline text. */ +static LIST_HEAD(rcu_tasks_tramp_holdouts); +static DEFINE_RAW_SPINLOCK(rcu_tasks_tramp_lock); + +/* CPUs / holdouts the current grace period is still waiting for. */ +static struct cpumask rcu_tasks_pending_cpus; +static LIST_HEAD(rcu_tasks_gp_holdouts); + +extern char __rcu_tramp_text_start[], __rcu_tramp_text_end[]; + +/** + * arch_rcu_tasks_trampoline_text - Does the architecture treat @ip as unm= arked trampoline text? + * @ip: kernel text address inside core kernel text + * + * See rcu_tasks_trampoline_text(). Architectures override this to flag + * core text that runs on behalf of a trampoline outside its Tasks Trace + * reader, e.g. static ftrace entry stubs or return thunks that hold a + * trampoline address they are about to jump to. + */ +bool __weak arch_rcu_tasks_trampoline_text(unsigned long ip) +{ + return false; +} + +/** + * rcu_tasks_trampoline_text - Is @ip in text Tasks RCU protects but no re= ader marks? + * @ip: an interrupted instruction pointer + * + * True when a task interrupted at @ip may be executing, or about to enter + * or return into, text whose lifetime depends on synchronize_rcu_tasks() + * without being inside the Tasks Trace reader that text takes around its + * call-outs: + * + * - anything outside core kernel and module text (ftrace and BPF + * trampolines, kprobe slots and other dynamically allocated text; this + * deliberately does not ask is_ftrace_trampoline() and friends, since + * text being torn down may already be unregistered there); + * - the .text..rcu_tramp section, C glue called directly from such + * trampolines before it has entered the reader; + * - whatever the architecture adds via arch_rcu_tasks_trampoline_text(). + * + * A false positive only makes the task a holdout until its next quiescent + * event. Called with interrupts disabled from the irq-exit path. + */ +bool rcu_tasks_trampoline_text(unsigned long ip) +{ + if (core_kernel_text(ip)) { + if (ip >=3D (unsigned long)__rcu_tramp_text_start && + ip < (unsigned long)__rcu_tramp_text_end) + return true; + return arch_rcu_tasks_trampoline_text(ip); + } + return !is_module_text_address(ip); +} +NOKPROBE_SYMBOL(rcu_tasks_trampoline_text); + +/* Note a Tasks RCU quiescent event on this CPU. */ +static void rcu_tasks_qs_event(void) +{ + unsigned long *seq; + + guard(preempt_notrace)(); + seq =3D this_cpu_ptr(&rcu_tasks_qs_seq); + /* Order a preceding rcu_tasks_tramp_hold() before the count. */ + smp_store_release(seq, *seq + 1); +} + +static void rcu_tasks_tramp_hold(struct task_struct *t) +{ + unsigned long flags; + + if (t->rcu_tasks_holdout) + return; + raw_spin_lock_irqsave(&rcu_tasks_tramp_lock, flags); + list_add_tail(&t->rcu_tasks_holdout_list, &rcu_tasks_tramp_holdouts); + WRITE_ONCE(t->rcu_tasks_holdout, true); + raw_spin_unlock_irqrestore(&rcu_tasks_tramp_lock, flags); +} + +static void rcu_tasks_tramp_release(struct task_struct *t) +{ + unsigned long flags; + + if (likely(!t->rcu_tasks_holdout)) + return; + raw_spin_lock_irqsave(&rcu_tasks_tramp_lock, flags); + list_del_init(&t->rcu_tasks_holdout_list); + WRITE_ONCE(t->rcu_tasks_holdout, false); + raw_spin_unlock_irqrestore(&rcu_tasks_tramp_lock, flags); +} + +/** + * rcu_tasks_irq_resched_enter - Tasks RCU hook for the irq-exit reschedul= e check + * @ip: instruction pointer of the interrupted (task-level) context + * + * Called with interrupts disabled when an interrupt returning to kernel + * mode is about to preempt_schedule_irq(), the one context switch that can + * catch a task inside unmarked trampoline text. Record where the task is + * parked for as long as it is (rcu_tasks_wait_irq_preempted() looks at + * that), and if it is inside such text make it a holdout before + * __schedule() reports the quiescent event; if it is not, this is as good + * as a voluntary switch for ending an earlier hold. + */ +void rcu_tasks_irq_resched_enter(unsigned long ip) +{ + struct task_struct *t =3D current; + struct rcu_tasks_percpu *rtpcp =3D this_cpu_ptr(rcu_tasks.rtpcpu); + + lockdep_assert_irqs_disabled(); + WRITE_ONCE(t->rcu_tasks_irq_ip, ip); + t->rcu_tasks_exit_cpu =3D smp_processor_id(); + raw_spin_lock_rcu_node(rtpcp); + list_add(&t->rcu_tasks_exit_list, &rtpcp->rtp_exit_list); + raw_spin_unlock_rcu_node(rtpcp); + + if (unlikely(rcu_tasks_trampoline_text(ip))) + rcu_tasks_tramp_hold(t); + else + rcu_tasks_tramp_release(t); +} +NOKPROBE_SYMBOL(rcu_tasks_irq_resched_enter); + +/** + * rcu_tasks_irq_resched_exit - preempt_schedule_irq() has returned + * + * The task is running again (possibly elsewhere) and about to return to t= he + * interrupted context; it is no longer parked anywhere. + */ +void rcu_tasks_irq_resched_exit(void) +{ + struct task_struct *t =3D current; + struct rcu_tasks_percpu *rtpcp =3D per_cpu_ptr(rcu_tasks.rtpcpu, t->rcu_t= asks_exit_cpu); + + lockdep_assert_irqs_disabled(); + raw_spin_lock_rcu_node(rtpcp); + list_del_init(&t->rcu_tasks_exit_list); + raw_spin_unlock_rcu_node(rtpcp); + WRITE_ONCE(t->rcu_tasks_irq_ip, 0); +} +NOKPROBE_SYMBOL(rcu_tasks_irq_resched_exit); + +/** + * rcu_tasks_note_qs - Tasks RCU hook for a context switch or explicit QS + * @t: current + * @preempt: this is a preemption rather than a voluntary switch + * + * Every pass through __schedule() (and cond_resched_tasks_rcu_qs(), and a + * tick from userspace or idle) is a quiescent event for this CPU: unmarked + * trampoline text never calls anything that schedules, and the irq-exit + * path has already made @t a holdout if it is preempting inside such text. + * Any of these outside an irq-exit preemption also shows @t itself to be + * outside, ending an earlier hold -- including cond_resched() under + * PREEMPT_DYNAMIC's none/voluntary modes, where the irq-exit path is off. + */ +void rcu_tasks_note_qs(struct task_struct *t, bool preempt) +{ + WARN_ON_ONCE(t !=3D current); + if (!READ_ONCE(t->rcu_tasks_irq_ip)) + rcu_tasks_tramp_release(t); + rcu_tasks_qs_event(); +} +EXPORT_SYMBOL_GPL(rcu_tasks_note_qs); /* cond_resched_tasks_rcu_qs() */ + +/** + * rcu_tasks_wait_irq_preempted - wait for tasks irq-preempted inside @ins= ide + * @inside: predicate on a task's recorded irq-exit preemption IP + * + * For a caller about to make some ordinary text unsafe to be parked in + * (the kprobe jump optimizer): once the caller has arranged for + * rcu_tasks_trampoline_text() to cover that text, new irq-exit preemptions + * there become holdouts, but a task preempted there earlier is invisible + * to the grace period. Wait until no parked task's recorded preemption IP + * is inside; a following synchronize_rcu_tasks() then covers the rest. + * The leading synchronize_rcu() orders the caller's arrangement against + * preemptions in flight, which run with interrupts disabled. + */ +void rcu_tasks_wait_irq_preempted(bool (*inside)(unsigned long ip)) +{ + struct task_struct *t; + unsigned long flags; + int cpu, kick; + bool found; + + synchronize_rcu(); + for (;;) { + found =3D false; + for_each_possible_cpu(cpu) { + struct rcu_tasks_percpu *rtpcp =3D per_cpu_ptr(rcu_tasks.rtpcpu, cpu); + + kick =3D -1; + raw_spin_lock_irqsave_rcu_node(rtpcp, flags); + list_for_each_entry(t, &rtpcp->rtp_exit_list, rcu_tasks_exit_list) { + if (inside(READ_ONCE(t->rcu_tasks_irq_ip))) { + found =3D true; + if (task_curr(t)) + kick =3D task_cpu(t); + } + } + raw_spin_unlock_irqrestore_rcu_node(rtpcp, flags); + if (kick >=3D 0) + resched_cpu(kick); + } + if (!found) + return; + schedule_timeout_uninterruptible(1); + } +} + +/* Has @cpu passed a quiescent event since the snapshot, or need it not? */ +static bool rcu_tasks_cpu_quiescent(int cpu) +{ + if (!cpu_online(cpu)) + return true; + /* Pairs with the release in rcu_tasks_qs_event(). */ + if (smp_load_acquire(per_cpu_ptr(&rcu_tasks_qs_seq, cpu)) !=3D + per_cpu(rcu_tasks_qs_snap, cpu)) + return true; + /* + * Idle or nohz_full userspace (an RCU extended quiescent state): no + * task-level kernel frames there, and whatever ran before has switched + * out. As with the classic flavor, the idle task itself is not waited + * for. + */ + if (!(ct_rcu_watching_cpu(cpu) & CT_RCU_WATCHING)) + return true; + return idle_cpu(cpu); +} + +/* Rate-limited stall report; returns true if the caller should add detail= . */ +static bool rcu_tasks_tramp_stall(struct rcu_tasks *rtp, unsigned long *la= streport, + const char *what) +{ + int rtst =3D READ_ONCE(rcu_task_stall_timeout); + + if (rtst <=3D 0 || !time_after(jiffies, *lastreport + rtst)) + return false; + *lastreport =3D jiffies; + pr_err("INFO: %s: %s, grace period %lu is %lu jiffies old\n", rtp->kname, + what, rcu_seq_current(&rtp->tasks_gp_seq), jiffies - rtp->gp_start= ); + return true; +} + +/* + * Steps 1/4: wait until every online non-idle CPU has context switched. A + * CPU that has not after a jiffy is asked to with resched_cpu(), which ta= kes + * it through rcu_tasks_irq_resched_enter() and __schedule() (or, from + * userspace or a guest, straight to __schedule()). + */ +static void rcu_tasks_tramp_wait_cpus(struct rcu_tasks *rtp, unsigned long= *lastreport) +{ + struct cpumask *pending =3D &rcu_tasks_pending_cpus; + unsigned long start; + int cpu; + + /* + * The quiescent events run with preemption (in practice interrupts) + * disabled, so after this any event we go on to count began after the + * caller's updates -- the unpublished trampoline, and whatever + * rcu_tasks_trampoline_text() consults -- were visible to it. + */ + synchronize_rcu(); + + start =3D jiffies; + for_each_online_cpu(cpu) { + per_cpu(rcu_tasks_qs_snap, cpu) =3D READ_ONCE(per_cpu(rcu_tasks_qs_seq, = cpu)); + __cpumask_set_cpu(cpu, pending); + } + /* Snapshots before the checks below; pairs with rcu_tasks_qs_event(). */ + smp_mb(); + + for (;;) { + for_each_cpu(cpu, pending) + if (rcu_tasks_cpu_quiescent(cpu)) + __cpumask_clear_cpu(cpu, pending); + if (cpumask_empty(pending)) + break; + if (time_after(jiffies, start)) { + for_each_cpu(cpu, pending) + resched_cpu(cpu); + rtp->n_ipis +=3D cpumask_weight(pending); + } + schedule_timeout_idle(1); + if (rcu_tasks_tramp_stall(rtp, lastreport, "CPUs without a quiescent eve= nt")) + pr_err("\tCPUs: %*pbl\n", cpumask_pr_args(pending)); + } +} + +/* + * Steps 2/4: wait for the tasks that were holdouts when we looked to stop + * being holdouts. They are moved to a private list so that tasks becoming + * holdouts later (in live trampolines) cannot keep us here; each removes + * itself via rcu_tasks_tramp_release() wherever it is queued. + */ +static void rcu_tasks_tramp_wait_holdouts(struct rcu_tasks *rtp, unsigned = long *lastreport) +{ + struct task_struct *t; + unsigned long flags; + int cpu; + + raw_spin_lock_irqsave(&rcu_tasks_tramp_lock, flags); + list_splice_tail_init(&rcu_tasks_tramp_holdouts, &rcu_tasks_gp_holdouts); + raw_spin_unlock_irqrestore(&rcu_tasks_tramp_lock, flags); + + for (;;) { + struct cpumask *kick =3D &rcu_tasks_pending_cpus; + struct task_struct *show[8]; + int nshow =3D 0, i; + bool empty, report; + + report =3D rcu_tasks_tramp_stall(rtp, lastreport, + "tasks preempted in trampoline text"); + cpumask_clear(kick); + raw_spin_lock_irqsave(&rcu_tasks_tramp_lock, flags); + empty =3D list_empty(&rcu_tasks_gp_holdouts); + list_for_each_entry(t, &rcu_tasks_gp_holdouts, rcu_tasks_holdout_list) { + if (task_curr(t)) + __cpumask_set_cpu(task_cpu(t), kick); + if (report && nshow < ARRAY_SIZE(show)) + show[nshow++] =3D get_task_struct(t); + } + raw_spin_unlock_irqrestore(&rcu_tasks_tramp_lock, flags); + /* Never printk under the lock the irq-exit path takes. */ + for (i =3D 0; i < nshow; i++) { + sched_show_task(show[i]); + put_task_struct(show[i]); + } + if (empty) + break; + for_each_cpu(cpu, kick) + resched_cpu(cpu); + rtp->n_ipis +=3D cpumask_weight(kick); + schedule_timeout_idle(1); + } +} + +/* Wait for one trampoline-reader Tasks RCU grace period. */ +static void rcu_tasks_tramp_wait_gp(struct rcu_tasks *rtp) +{ + unsigned long lastreport =3D jiffies; + + set_tasks_gp_state(rtp, RTGS_WAIT_SCAN_HOLDOUTS); + rcu_tasks_tramp_wait_cpus(rtp, &lastreport); + rcu_tasks_tramp_wait_holdouts(rtp, &lastreport); + + set_tasks_gp_state(rtp, RTGS_WAIT_READERS); + synchronize_rcu_tasks_trace(); + + set_tasks_gp_state(rtp, RTGS_SCAN_HOLDOUTS); + rcu_tasks_tramp_wait_cpus(rtp, &lastreport); + rcu_tasks_tramp_wait_holdouts(rtp, &lastreport); + + set_tasks_gp_state(rtp, RTGS_POST_GP); +} + +static int __init rcu_spawn_tasks_kthread(void) +{ + rcu_tasks.gp_sleep =3D HZ / 10; + if (rcu_tasks_lazy_ms >=3D 0) + rcu_tasks.lazy_jiffies =3D msecs_to_jiffies(rcu_tasks_lazy_ms); + rcu_tasks.wait_state =3D TASK_IDLE; + rcu_spawn_tasks_kthread_generic(&rcu_tasks); + return 0; +} + +void exit_tasks_rcu_start(void) { } +void exit_tasks_rcu_finish(void) { } + +#else /* #ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS */ + //////////////////////////////////////////////////////////////////////// // // Simple variant of RCU whose quiescent states are voluntary context @@ -1173,6 +1608,8 @@ static void tasks_rcu_exit_stall(struct timer_list *u= nused) #endif // #ifndef CONFIG_TINY_RCU } =20 +#endif /* #else #ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS */ + /** * call_rcu_tasks() - Queue an RCU for invocation task-based grace period * @rhp: structure to be used for queueing the RCU updates. @@ -1187,6 +1624,12 @@ static void tasks_rcu_exit_stall(struct timer_list *= unused) * primitives analogous to rcu_read_lock() and rcu_read_unlock() because * this primitive is intended to determine that all tasks have passed * through a safe state, not so much for data-structure synchronization. + * On CONFIG_TASKS_RCU_TRAMPOLINE_READERS kernels a preemption outside + * trampoline text also ends one, and a reader whose protected window + * spans preemptible code must additionally be a Tasks Trace RCU reader + * (rcu_read_lock_trace(), as the trampolines there take around their + * call-outs); an arbitrary stretch of preemptible kernel code is not + * protected. * * See the description of call_rcu() for more detailed information on * memory ordering guarantees. @@ -1205,7 +1648,9 @@ EXPORT_SYMBOL_GPL(call_rcu_tasks); * executing rcu-tasks read-side critical sections have elapsed. These * read-side critical sections are delimited by calls to schedule(), * cond_resched_tasks_rcu_qs(), idle execution, userspace execution, calls - * to synchronize_rcu_tasks(), and (in theory, anyway) cond_resched(). + * to synchronize_rcu_tasks(), and (in theory, anyway) cond_resched(); + * on CONFIG_TASKS_RCU_TRAMPOLINE_READERS kernels also by preemption + * outside trampoline text, see call_rcu_tasks(). * * This is a very specialized primitive, intended only for a few uses in * tracing and other situations requiring manipulation of function @@ -1233,9 +1678,7 @@ void rcu_barrier_tasks(void) } EXPORT_SYMBOL_GPL(rcu_barrier_tasks); =20 -static int rcu_tasks_lazy_ms =3D -1; -module_param(rcu_tasks_lazy_ms, int, 0444); - +#ifndef CONFIG_TASKS_RCU_TRAMPOLINE_READERS static int __init rcu_spawn_tasks_kthread(void) { rcu_tasks.gp_sleep =3D HZ / 10; @@ -1251,6 +1694,7 @@ static int __init rcu_spawn_tasks_kthread(void) rcu_spawn_tasks_kthread_generic(&rcu_tasks); return 0; } +#endif /* #ifndef CONFIG_TASKS_RCU_TRAMPOLINE_READERS */ =20 #if !defined(CONFIG_TINY_RCU) void show_rcu_tasks_classic_gp_kthread(void) @@ -1279,6 +1723,7 @@ void rcu_tasks_get_gp_data(int *flags, unsigned long = *gp_seq) } EXPORT_SYMBOL_GPL(rcu_tasks_get_gp_data); =20 +#ifndef CONFIG_TASKS_RCU_TRAMPOLINE_READERS /* * Protect against tasklist scan blind spot while the task is exiting and * may be removed from the tasklist. Do this by adding the task to yet @@ -1322,6 +1767,7 @@ void exit_tasks_rcu_finish(void) list_del_init(&t->rcu_tasks_exit_list); raw_spin_unlock_irqrestore_rcu_node(rtpcp, flags); } +#endif /* #ifndef CONFIG_TASKS_RCU_TRAMPOLINE_READERS */ =20 #else /* #ifdef CONFIG_TASKS_RCU */ void exit_tasks_rcu_start(void) { } diff --git a/kernel/rcu/update.c b/kernel/rcu/update.c index b62735a67884..a122b8d1effb 100644 --- a/kernel/rcu/update.c +++ b/kernel/rcu/update.c @@ -40,7 +40,9 @@ #include #include #include +#include #include +#include #include #include #include --=20 2.55.0 From nobody Thu Sep 24 19:04:36 2026 Delivered-To: importer@patchew.org Received-SPF: pass (zohomail.com: domain of lists.xenproject.org designates 192.237.175.120 as permitted sender) client-ip=192.237.175.120; envelope-from=xen-devel-bounces@lists.xenproject.org; helo=lists.xenproject.org; Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of lists.xenproject.org designates 192.237.175.120 as permitted sender) smtp.mailfrom=xen-devel-bounces@lists.xenproject.org ARC-Seal: i=1; a=rsa-sha256; t=1789478296; cv=none; d=zohomail.com; s=zohoarc; b=gYJR5i4ruYr4HeOE84nulJlBkaOex9sShh3wMNgeuHHxmIETNy3uNsGEhKRJQQvI+gnZXUloqOAQqNZW08Q/7IvaIckpzZ0zDQUyM90LA8v1MDXWGqpknrZLxKYbyuSotLS88FyXP+GiBXF7h8cgRtY6BRjobwlFga/8IDYt23s= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1789478296; h=Content-Type:Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=zOBWKoOt8czuTeNdTxDxvLfMht4h54WO3jEdwibwnr4=; b=Bu5ezWxGf0IHGjjM371t7k7Fod5OV3EARXOYRD6SsZN3Iubr+tGkPq086hmWeZJNpxpKraBMAqMLTfKSpEb8edtrV8Hj2zZenKhGpenMkIywy3OHEy9qqpjqytajRSmRieL0f2IM0fG+yuPrKf8v1ptmiGdWnUMekVzDEWFGHMs= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of lists.xenproject.org designates 192.237.175.120 as permitted sender) smtp.mailfrom=xen-devel-bounces@lists.xenproject.org Return-Path: Received: from lists.xenproject.org (lists.xenproject.org [192.237.175.120]) by mx.zohomail.com with SMTPS id 178947829642235.84195442778241; Tue, 15 Sep 2026 06:18:16 -0700 (PDT) Received: from list by lists.xenproject.org with outflank-mailman.1421708.1647422 (Exim 4.92) (envelope-from ) id 1x6T2r-00050C-Ef; Tue, 15 Sep 2026 13:17:57 +0000 Received: by outflank-mailman (output) from mailman id 1421708.1647422; Tue, 15 Sep 2026 13:17:57 +0000 Received: from localhost ([127.0.0.1] helo=lists.xenproject.org) by lists.xenproject.org with esmtp (Exim 4.92) (envelope-from ) id 1x6T2r-0004yu-Ay; Tue, 15 Sep 2026 13:17:57 +0000 Received: by outflank-mailman (input) for mailman id 1421708; Tue, 15 Sep 2026 13:17:56 +0000 Received: from mx.expurgate.net ([195.190.135.10]) by lists.xenproject.org with esmtp (Exim 4.92) (envelope-from ) id 1x6T2q-0004tp-1c for xen-devel@lists.xenproject.org; Tue, 15 Sep 2026 13:17:56 +0000 Received: from mx.expurgate.net (helo=localhost) by mx.expurgate.net with esmtp id 1x6T2p-00EG9J-Dv for xen-devel@lists.xenproject.org; Tue, 15 Sep 2026 15:17:55 +0200 Received: from [10.42.69.3] (helo=localhost) by localhost with ESMTP (eXpurgate MTA 0.9.1) (envelope-from ) id 6aa9457b-bab6-0a2a0a5309dd-0a2a4503e980-14 for ; Tue, 15 Sep 2026 15:17:55 +0200 Received: from [74.125.230.204] (helo=mail-qk2-f12.google.com) by tlsNG-33051d.mxtls.expurgate.net with ESMTPS (eXpurgate 4.57.1) (envelope-from ) id 6aa94582-fae8-0a2a45030019-4a7de6ccce25-3 for ; Tue, 15 Sep 2026 15:17:55 +0200 Received: by mail-qk2-f12.google.com with SMTP id d75a77b69052e-52fb76906adso46626251cf.0 for ; Tue, 15 Sep 2026 06:17:55 -0700 (PDT) Received: from toxicpanda.com (ec2-34-228-114-98.compute-1.amazonaws.com. [34.228.114.98]) by smtp.gmail.com with ESMTPSA id d75a77b69052e-530ca4ba156sm126369461cf.19.2026.09.15.06.17.52 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 15 Sep 2026 06:17:52 -0700 (PDT) X-Outflank-Mailman: Message body and most headers restored to incoming version X-BeenThere: xen-devel@lists.xenproject.org List-Id: Xen developer discussion List-Unsubscribe: , List-Post: List-Help: List-Subscribe: , Errors-To: xen-devel-bounces@lists.xenproject.org Precedence: list Sender: "Xen-devel" Authentication-Results: eu.smtp.expurgate.cloud; dkim=pass header.s=google header.d=toxicpanda.com header.i="@toxicpanda.com" header.h="Cc:To:In-Reply-To:References:Message-Id:Content-Transfer-Encoding:Content-Type:MIME-Version:Subject:Date:From" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=toxicpanda.com; s=google; t=1789478274; x=1790083074; darn=lists.xenproject.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=zOBWKoOt8czuTeNdTxDxvLfMht4h54WO3jEdwibwnr4=; b=NqEVb7HrEnHrZA5W2qxk+nVH6z+qq08zjTm0pTrJJzvNRuBXta3K9r7d0mHCxGt+TV 2TV7DFHol2TC2OJt/KpByH24qiIEx3VyMpr6TZx0Si8mfDa2x5lU/pEp/UDs/vlzg/1d ObALVskR2C8G29X5BT8NQmzxpKu5zHt/v9GZTOI0U4ydA9upCBL9uyA1aWXv7HAwlgdA QSFZRF1fFT4SAmXKrXmdQtFcsEQ3kUWAr3oBeLT9OMrqZEpsjiYR8nUJ4Z0R2rcKh8TR kQGDw4jCp8crnP9YwLrOYuU9gStKk+ETPCfDNrfbZUTynhIlRBphoneCL8IT2CgX3MKC OzXA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789478274; x=1790083074; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=zOBWKoOt8czuTeNdTxDxvLfMht4h54WO3jEdwibwnr4=; b=EamyQS2PpCQXPpPQ/kj0eHsG8GTAa4AXpqU0+fr2W+yRG9XOE5cRjwppPlEbqKpRgx phx+KhjoutFHhL/z4bHS2WsUm6l6zDcRCA6TtA1icBMp02rP/g+YLHoH+ayFCQXp4buP rA4PSwg54nbiiZ8UoOCmu0gIbtxFEok8zy0GPHkl3bLbe7xxH1IDM1IzrEtm1Zg0mGUY 1aE/PEHbOeMo37ZMvNP0OWPi/ex8ETl6nas89i0sX1BMvHL4jmXJ2ORT6B1cQIQk8+M9 Aj+kAIU1mWDMwn8yIrxR1C7B02JVKcOi8t7l9mjGrCm27V3b0E8FJNWYq6QSc3cr7e0K FqNw== X-Forwarded-Encrypted: i=1; AKwUvBzAIqCXTH4APV9mDA0vamYLLwfsdVAFh4XjrbCVY4nhug0rSP1W9mtex7WCTpxTutHoMzm/Pn/FtWc=@lists.xenproject.org X-Gm-Message-State: AFuF++nPNT29Kp5WvjisUOvxQK0QfyJbMCM0fI6GKUyMJWafXnSmMoG4 F1oNW31qb/5700hXTK8nzvsykLhfiPy7wLFKgliNU+D6FHGkUVG0DleBxTbmQjrjIiM= X-Gm-Gg: AYBFou2Iczvmt9SQ4rsh/jknC9G1r0UB0gCiWFpCWvYQHecnrDerD1Y0e/HNAp7xtLD TfoeuKMIizMwbaXns+BjSKuH8HSB9O2mX+2jOS4XG/s9G6C4RRoxsXQYFvuhNRaibcTjOpBOFdP +O2iKzO3aLjIhmrepZgH6W0+WK3lF5cxTjGWGfJw75iaAwc4puVekDBT6+hOVaoSodh0Ytj/BPC 1z2KSKfJMmIfAF8IN7ama0U0b7pJsAvdf62XPD2qY8zKwQ4Q3y6c+nRVdU9R7WhPUhcNPlco/jw DdKkRYBM/KsvvWlRGNb9P5nMRku8M0ChD0mTNIwwUY8SVvh+ifZZW5kNLidXDwy553x7a+ZMDwq Je+cTXXj/VJB8s+hyXcO6S+nRMvJbD+YRRyHUDxEZW/Oxc+jxtasnK5MG2cQglHG6njhbyyBFf8 lPKAIj1AWTWRuo3vz04/sjBZE/cfv45BLpuZljnJSBhnXiQEyhK4l9nfwjPe64SISiILNZ18NZ6 IIPzrkzBkjsrNoKUfmEN74oMQ4IsNUs7bAKiDceWNb96Rzw8K4/tDd0 X-Received: by 2002:a05:622a:1646:b0:530:9bb:54ae with SMTP id d75a77b69052e-5310cf32770mr102643101cf.15.1789478273513; Tue, 15 Sep 2026 06:17:53 -0700 (PDT) From: Josef Bacik Date: Tue, 15 Sep 2026 13:17:31 +0000 Subject: [PATCH RFC v3 04/13] kprobes: Expose the optprobe jump window to Tasks RCU MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260915-b4-rcu-tasks-preempt-qs-v3-4-0ad30c4c5ee7@toxicpanda.com> References: <20260915-b4-rcu-tasks-preempt-qs-v3-0-0ad30c4c5ee7@toxicpanda.com> In-Reply-To: <20260915-b4-rcu-tasks-preempt-qs-v3-0-0ad30c4c5ee7@toxicpanda.com> To: "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Boqun Feng , Thomas Gleixner , Peter Zijlstra , Steven Rostedt , Masami Hiramatsu , Mark Rutland , Jiri Olsa , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , x86@kernel.org, Catalin Marinas , Will Deacon , Puranjay Mohan , Xu Kuohai , "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Boqun Feng , Thomas Gleixner , Peter Zijlstra , Steven Rostedt , Masami Hiramatsu , Mark Rutland , Jiri Olsa , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , x86@kernel.org, Catalin Marinas , Will Deacon , Puranjay Mohan , Xu Kuohai Cc: Andy Lutomirski , Josh Triplett , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Juergen Gross , Luis Chamberlain , Ihor Solodrai , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, xen-devel@lists.xenproject.org, Andy Lutomirski , Josh Triplett , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Juergen Gross , Luis Chamberlain , Ihor Solodrai , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, xen-devel@lists.xenproject.org, Josef Bacik X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openssh-sha256; t=1789478262; l=7478; i=josef@toxicpanda.com; h=from:subject:message-id; bh=2gqgm31o9vakOy6+NG5dFQrZfNnO16Cyr1QKmdk48dg=; b=U1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgUBr36M/n0nWN0DNbnxwzIiCZez6MG JiruuNaSCI/zXsAAAAGcGF0YXR0AAAAAAAAAAZzaGE1MTIAAABTAAAAC3NzaC1lZDI1NTE5AAAA QFs4wUTmA5uBGNzmizfs5h8i2jZ+BJEC+ZKMMB4U6RQcjTiVVDPqXAZ9evZNBFK/JcfouDSMfSj BhbWSxMKT7As= X-Developer-Key: i=josef@toxicpanda.com; a=openssh; fpr=SHA256:C8kOX2QUJCMqnCX+KEeoqRAjLo9L+ELOSH2NSAJHqGA X-purgate-ID: tlsNG-33051d/1789478275-6C8DD4E9-C8161ED8/0/0 X-purgate-type: clean X-purgate-size: 7480 X-ZohoMail-DKIM: pass (identity @toxicpanda.com) X-ZM-MESSAGEID: 1789478298869158500 kprobe_optimizer() is the one synchronize_rcu_tasks() user that is not about trampoline text: it waits for tasks that were interrupted on an instruction boundary inside the bytes it is about to overwrite with the optimized jump, so that none of them resumes into the middle of the new instruction. Those bytes are ordinary kernel or module text with no Tasks Trace reader around them, so on CONFIG_TASKS_RCU_TRAMPOLINE_READERS kernels the irq-exit quiescent-state check has to be told about them. Add kprobe_in_optimized_region(), a lockless and conservative form of get_optimized_kprobe() that reports whether any registered kprobe lies within MAX_OPTIMIZED_LENGTH before the given address regardless of its optimization state, and have rcu_tasks_trampoline_text() consult it for core and module text so that a task interrupted there becomes a holdout rather than a quiescent event. The hash walk only runs while kprobe_optimizer() is actually inside its synchronize_rcu_tasks(), tracked by a flag it sets around the call; otherwise the check is a single load. That check cannot see a task that was already preempted in the region before the flag went up (possibly before the kprobe even existed), and the new grace period does not otherwise wait for a preempted task to run again, so before synchronize_rcu_tasks() the optimizer calls rcu_tasks_wait_irq_preempted() to wait until no parked task's recorded irq-exit preemption IP is inside such a region; its leading synchronize_rcu() also publishes the flag to every (interrupts- disabled) check in flight. The kprobe hash is RCU-protected and every free path waits for a grace period after unhashing, so the lockless walk from the irq-exit path is safe. On other configurations the flag is set and cleared but nothing reads it and rcu_tasks_wait_irq_preempted() is a stub; the classic implementation already waits for such tasks. Assisted-by: LLM Signed-off-by: Josef Bacik --- include/linux/kprobes.h | 8 +++++++- kernel/kprobes.c | 50 +++++++++++++++++++++++++++++++++++++++++++++= ++++ kernel/rcu/tasks.h | 11 ++++++++--- 3 files changed, 65 insertions(+), 4 deletions(-) diff --git a/include/linux/kprobes.h b/include/linux/kprobes.h index e6de7ae55bda..74cc48c04417 100644 --- a/include/linux/kprobes.h +++ b/include/linux/kprobes.h @@ -530,11 +530,17 @@ static inline bool is_kprobe_insn_slot(unsigned long = addr) } #endif /* !CONFIG_KPROBES */ =20 -#ifndef CONFIG_OPTPROBES +#ifdef CONFIG_OPTPROBES +bool kprobe_in_optimized_region(unsigned long addr); +#else /* !CONFIG_OPTPROBES */ static inline bool is_kprobe_optinsn_slot(unsigned long addr) { return false; } +static inline bool kprobe_in_optimized_region(unsigned long addr) +{ + return false; +} #endif /* !CONFIG_OPTPROBES */ =20 #ifdef CONFIG_KRETPROBES diff --git a/kernel/kprobes.c b/kernel/kprobes.c index 6337da5cab9e..e460fba83e4a 100644 --- a/kernel/kprobes.c +++ b/kernel/kprobes.c @@ -511,6 +511,48 @@ static struct kprobe *get_optimized_kprobe(kprobe_opco= de_t *addr) return NULL; } =20 +/* + * True while kprobe_optimizer() is waiting for its Tasks RCU grace period. + * Only in that window can an interruption inside an optprobe's jump region + * matter to it, so kprobe_in_optimized_region() does no work otherwise. + */ +static bool kprobe_optimizer_waiting; + +/** + * kprobe_in_optimized_region - Could @addr be inside bytes a jump-optimiz= ed + * kprobe replaces? + * @addr: kernel text address, typically an interrupted instruction pointer + * + * kprobe_optimizer() relies on synchronize_rcu_tasks() to wait for tasks = that + * were interrupted on an instruction boundary inside the region about to = be + * overwritten by the optimized jump. Where Tasks RCU is built on + * reader-marked trampolines that region has no reader, so the irq-exit + * quiescent-state check asks this instead (see rcu_tasks_trampoline_text(= )). + * This is the lockless, conservative form of get_optimized_kprobe(): it d= oes + * not care whether the kprobe found is, or ever will be, optimized. May = be + * called from any context with preemption disabled; the kprobe hash is + * RCU-protected and every free path waits for a grace period after unhash= ing. + * + * The hash walk only runs while the optimizer is actually waiting. A task + * that was preempted in such a region before the flag went up is invisible + * to that check, so the optimizer first waits those out by their recorded + * preemption IP (rcu_tasks_wait_irq_preempted(), whose leading + * synchronize_rcu() also publishes the flag to every check in flight). + */ +bool kprobe_in_optimized_region(unsigned long addr) +{ + int i; + + if (!READ_ONCE(kprobe_optimizer_waiting)) + return false; + + for (i =3D 1; i < MAX_OPTIMIZED_LENGTH / sizeof(kprobe_opcode_t); i++) + if (get_kprobe((kprobe_opcode_t *)addr - i)) + return true; + return false; +} +NOKPROBE_SYMBOL(kprobe_in_optimized_region); + /* Optimization staging list, protected by 'kprobe_mutex' */ static LIST_HEAD(optimizing_list); static LIST_HEAD(unoptimizing_list); @@ -644,8 +686,16 @@ static void kprobe_optimizer(void) * to 2nd-Nth byte of jump instruction. This wait is for avoiding it. * Note that on non-preemptive kernel, this is transparently converted * to synchronoze_sched() to wait for all interrupts to have completed. + * kprobe_optimizer_waiting lets a reader-marked-trampoline Tasks RCU + * recognise tasks interrupted in such a region while we wait, and + * rcu_tasks_wait_irq_preempted() (a no-op elsewhere) first waits + * out any that were preempted there before we said so; see + * kprobe_in_optimized_region(). */ + WRITE_ONCE(kprobe_optimizer_waiting, true); + rcu_tasks_wait_irq_preempted(kprobe_in_optimized_region); synchronize_rcu_tasks(); + WRITE_ONCE(kprobe_optimizer_waiting, false); =20 /* Step 3: Optimize kprobes after quiesence period */ do_optimize_kprobes(); diff --git a/kernel/rcu/tasks.h b/kernel/rcu/tasks.h index 3a7c092361a6..866768462850 100644 --- a/kernel/rcu/tasks.h +++ b/kernel/rcu/tasks.h @@ -1006,7 +1006,9 @@ bool __weak arch_rcu_tasks_trampoline_text(unsigned l= ong ip) * text being torn down may already be unregistered there); * - the .text..rcu_tramp section, C glue called directly from such * trampolines before it has entered the reader; - * - whatever the architecture adds via arch_rcu_tasks_trampoline_text(). + * - whatever the architecture adds via arch_rcu_tasks_trampoline_text(); + * - the bytes after a kprobe that a pending jump optimization is about to + * overwrite, the one synchronize_rcu_tasks() user with no trampoline. * * A false positive only makes the task a holdout until its next quiescent * event. Called with interrupts disabled from the irq-exit path. @@ -1017,9 +1019,12 @@ bool rcu_tasks_trampoline_text(unsigned long ip) if (ip >=3D (unsigned long)__rcu_tramp_text_start && ip < (unsigned long)__rcu_tramp_text_end) return true; - return arch_rcu_tasks_trampoline_text(ip); + return arch_rcu_tasks_trampoline_text(ip) || + kprobe_in_optimized_region(ip); } - return !is_module_text_address(ip); + if (is_module_text_address(ip)) + return kprobe_in_optimized_region(ip); + return true; } NOKPROBE_SYMBOL(rcu_tasks_trampoline_text); =20 --=20 2.55.0 From nobody Thu Sep 24 19:04:36 2026 Delivered-To: importer@patchew.org Received-SPF: pass (zohomail.com: domain of lists.xenproject.org designates 192.237.175.120 as permitted sender) client-ip=192.237.175.120; envelope-from=xen-devel-bounces@lists.xenproject.org; helo=lists.xenproject.org; Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of lists.xenproject.org designates 192.237.175.120 as permitted sender) smtp.mailfrom=xen-devel-bounces@lists.xenproject.org ARC-Seal: i=1; a=rsa-sha256; t=1789478303; cv=none; d=zohomail.com; s=zohoarc; b=l25A6RqJu2O1oThboBWsDKzPFD3t4FqrojstDDDKM65ua9ERFeF8ECa3s+DlCFnJIeq/ACunTyMKJ9/aAiGOeCQ16W4gSvHUzZ4K9kjjwNQPnuK+OMS+dCHWAe4V1qDW3oVKS3g4IAqA6DQJ6snK9KxIacYxmhR4pLeeq3KxwPI= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1789478303; h=Content-Type:Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=N2aqKzOffwJC2uJ3O0Ky+yGqpaXXawvJtULMY2mEOoI=; b=I/CLqYdE7vHw4DzEmOqq9WxoXQ+zmidEPl1RUzR2JiN3rE4GL46Cf2qR5j5Gqq0muzLCgZw8Edsc0C87WeaQwQSLsr+wQrItuf5i4Jz7s+0NUFia4Ooz6XhhrHPcD9KImCfc4mway0qYdpQ+TZoc0g4PuEftADK7GjbgeFACulE= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of lists.xenproject.org designates 192.237.175.120 as permitted sender) smtp.mailfrom=xen-devel-bounces@lists.xenproject.org Return-Path: Received: from lists.xenproject.org (lists.xenproject.org [192.237.175.120]) by mx.zohomail.com with SMTPS id 1789478303895855.9924636566262; Tue, 15 Sep 2026 06:18:23 -0700 (PDT) Received: from list by lists.xenproject.org with outflank-mailman.1421709.1647436 (Exim 4.92) (envelope-from ) id 1x6T2u-0005QW-P8; Tue, 15 Sep 2026 13:18:00 +0000 Received: by outflank-mailman (output) from mailman id 1421709.1647436; Tue, 15 Sep 2026 13:18:00 +0000 Received: from localhost ([127.0.0.1] helo=lists.xenproject.org) by lists.xenproject.org with esmtp (Exim 4.92) (envelope-from ) id 1x6T2u-0005QI-Kt; Tue, 15 Sep 2026 13:18:00 +0000 Received: by outflank-mailman (input) for mailman id 1421709; Tue, 15 Sep 2026 13:17:59 +0000 Received: from mx.expurgate.net ([194.145.224.10]) by lists.xenproject.org with esmtp (Exim 4.92) (envelope-from ) id 1x6T2t-0005NY-8s for xen-devel@lists.xenproject.org; Tue, 15 Sep 2026 13:17:59 +0000 Received: from mx.expurgate.net (helo=localhost) by mx.expurgate.net with esmtp id 1x6T2s-00DQRf-Lp for xen-devel@lists.xenproject.org; Tue, 15 Sep 2026 15:17:58 +0200 Received: from [10.42.69.9] (helo=localhost) by localhost with ESMTP (eXpurgate MTA 0.9.1) (envelope-from ) id 6aa94581-e002-0a2a0a5209dd-0a2a45099780-36 for ; Tue, 15 Sep 2026 15:17:58 +0200 Received: from [74.125.230.140] (helo=mail-qv2-f12.google.com) by tlsNG-bad1c0.mxtls.expurgate.net with ESMTPS (eXpurgate 4.57.1) (envelope-from ) id 6aa94585-be1a-0a2a45090019-4a7de68ca160-3 for ; Tue, 15 Sep 2026 15:17:58 +0200 Received: by mail-qv2-f12.google.com with SMTP id 6a1803df08f44-90cdfc9b6eeso34079826d6.0 for ; Tue, 15 Sep 2026 06:17:58 -0700 (PDT) Received: from toxicpanda.com (ec2-34-228-114-98.compute-1.amazonaws.com. [34.228.114.98]) by smtp.gmail.com with ESMTPSA id 6a1803df08f44-9120f45a780sm120563676d6.10.2026.09.15.06.17.55 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 15 Sep 2026 06:17:55 -0700 (PDT) X-Outflank-Mailman: Message body and most headers restored to incoming version X-BeenThere: xen-devel@lists.xenproject.org List-Id: Xen developer discussion List-Unsubscribe: , List-Post: List-Help: List-Subscribe: , Errors-To: xen-devel-bounces@lists.xenproject.org Precedence: list Sender: "Xen-devel" Authentication-Results: eu.smtp.expurgate.cloud; dkim=pass header.s=google header.d=toxicpanda.com header.i="@toxicpanda.com" header.h="Cc:To:In-Reply-To:References:Message-Id:Content-Transfer-Encoding:Content-Type:MIME-Version:Subject:Date:From" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=toxicpanda.com; s=google; t=1789478277; x=1790083077; darn=lists.xenproject.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=N2aqKzOffwJC2uJ3O0Ky+yGqpaXXawvJtULMY2mEOoI=; b=LXtDBQP/6dEP0Lzvw94Zy3SApKxl1v9IYOvwpA7VnJ1VS3CoNbZTr6HNJCvkfkuUJJ 3Z0e7gV4mouR3AwZBTDMVY41b2qK2iu2SkiLxWKegxCzSHhPB04BaEm/kMqlB04DwEpG Hci7xMJ/MCjAfi77PpldnH906nfVov5UQbqt9tyQws/MqG7TXCSKQ8Hjft/gHV+TMAzb Q3Y1b7XeIuNHqOMzVsFRQ2G2TkbQYdhlbQLO1vKYx3HZ3eyByr9vY8MYvS2wvo3JrDPL 4jFSSD0efqxLz8uSfScs/Wn8GdIKakZ1fBSEcXXaGYwgGavw3RApehCW7ipd6EIZDrQF i6vQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789478277; x=1790083077; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=N2aqKzOffwJC2uJ3O0Ky+yGqpaXXawvJtULMY2mEOoI=; b=0yvdNvpw7AufpRQ/N9JLEH9PGRtQrikytyo5AeciLToyHA1bqu/eWc60sAsProC9yk fYtJtpI5db0WAKGydA1Kqq7JxJLEaMdVGBCmbHt3QF5OFAyRa2te7oIzLh2pPbqjlHKl 8tws7hdch15TIDpUnYhhT7lVmUB9JAw2LHPWMLHWBjJ9vRyjKJGbyn2I84abxFfZlH45 fs+2qGyBC2zB1H/k2foJUoM38yga5ORqXnAP/DAHx/vWF2ddx6nG/9wu8ME50iGP9qB5 FBuzV5++UE1xT7Vjmot/DoIhJdDvMffdK6QITKpt5Ckb7pyBSpoH7lQIc0UaYJSD4FCE jUIA== X-Forwarded-Encrypted: i=1; AKwUvBxKDND3YEPTWDWDlK4aPB5uLjex1meR97oDtym1NXGqmG1Bu3rVyeSlzhOlm6VMj19zI+H6DdOcILk=@lists.xenproject.org X-Gm-Message-State: AFuF++kC/6vq8vYRJbJ1O+WmM//LWVOaSlvZ+KVmMmnCkcBMgKo6C8TJ +zifRq/LlS4Nacw7IrWRxN2+P9PeipuaBDeapXfpPIu++dte348mMc07mQuoGtz6TZ4= X-Gm-Gg: AYBFou1tqz8Z2TQCIlgVcsqS1MHz4tmeetVDACerFLY4qDCH7BrLB4B5jJ74n51JSoA k5IGamoqdyYdGZ4CPwKbQdnLHUJV5I9+pQqCiWCN/J+MoLAQ9ioDxdqA4qTkKwn7vC+Izj7iJRL +v4FqjnRCp+lmv3Ryad8RHDlWiXX5Wp2+vk8stVXgUSEiO1De2UxYfmxklkpukMTPztwno953PR qWBv32KrM01QFUYGLlQutyaIDmEn/jQ0Vjdt2ZcGTaC7fGCO0dwbzDhzGIcNxb7v961lkQxaYWw +sjoXJp8ARuvyeiUUATIQkwoMLAy5/RWf7up/hfnWOlRG5R3ZPERBetWZ2x4a7O+Tbiv0xCj3ku MM+Z+OvDja8n73bhsgDH0T039WTvyHK1Ocz2w+bSJ8LXDKg9MvyS8dbZwjpAP4DeVb4+ZhjEucm vhaHOnFBueCcCqit6BnYkZdH6lsdCRGM9A4tUZLQv93en3l9qunvjY99xZczVvQot1yY6qupZcQ rW7drtxbklBIAH/+q46FUfR9aHSsrMVHXTIGMife9ewxu14dbt2VGzhHIqA7w3pxG4= X-Received: by 2002:a05:6214:4285:b0:90e:873f:112b with SMTP id 6a1803df08f44-9122e51912amr113387396d6.16.1789478276555; Tue, 15 Sep 2026 06:17:56 -0700 (PDT) From: Josef Bacik Date: Tue, 15 Sep 2026 13:17:32 +0000 Subject: [PATCH RFC v3 05/13] ftrace: Mark modules hosting direct-call trampolines for Tasks RCU MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260915-b4-rcu-tasks-preempt-qs-v3-5-0ad30c4c5ee7@toxicpanda.com> References: <20260915-b4-rcu-tasks-preempt-qs-v3-0-0ad30c4c5ee7@toxicpanda.com> In-Reply-To: <20260915-b4-rcu-tasks-preempt-qs-v3-0-0ad30c4c5ee7@toxicpanda.com> To: "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Boqun Feng , Thomas Gleixner , Peter Zijlstra , Steven Rostedt , Masami Hiramatsu , Mark Rutland , Jiri Olsa , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , x86@kernel.org, Catalin Marinas , Will Deacon , Puranjay Mohan , Xu Kuohai , "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Boqun Feng , Thomas Gleixner , Peter Zijlstra , Steven Rostedt , Masami Hiramatsu , Mark Rutland , Jiri Olsa , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , x86@kernel.org, Catalin Marinas , Will Deacon , Puranjay Mohan , Xu Kuohai Cc: Andy Lutomirski , Josh Triplett , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Juergen Gross , Luis Chamberlain , Ihor Solodrai , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, xen-devel@lists.xenproject.org, Andy Lutomirski , Josh Triplett , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Juergen Gross , Luis Chamberlain , Ihor Solodrai , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, xen-devel@lists.xenproject.org, Josef Bacik X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openssh-sha256; t=1789478262; l=7194; i=josef@toxicpanda.com; h=from:subject:message-id; bh=Fqdt5J86rRS/oaPi1G1Rp34bXnNDaywOKyZCOJUFY5s=; b=U1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgUBr36M/n0nWN0DNbnxwzIiCZez6MG JiruuNaSCI/zXsAAAAGcGF0YXR0AAAAAAAAAAZzaGE1MTIAAABTAAAAC3NzaC1lZDI1NTE5AAAA QFK/UccVcmZvo9+nJV8bUtDnpQLusVqacik5jCxxwwehteVs8AbBeSvZttSzdRt3+dA+Fv7F8p7 qiOCn2ZKvEwc= X-Developer-Key: i=josef@toxicpanda.com; a=openssh; fpr=SHA256:C8kOX2QUJCMqnCX+KEeoqRAjLo9L+ELOSH2NSAJHqGA X-purgate-ID: tlsNG-bad1c0/1789478278-3ACDF034-7F26B264/0/0 X-purgate-type: clean X-purgate-size: 7196 X-ZohoMail-DKIM: pass (identity @toxicpanda.com) X-ZM-MESSAGEID: 1789478304925158501 An out-of-line direct trampoline registered with register_ftrace_direct() is kept alive only by Tasks RCU while a task executes it or is preempted in something it called; ftrace_shutdown()'s synchronize_rcu_tasks() is what stops rmmod freeing it under such a task. Where Tasks RCU is built on reader-marked trampolines, such a trampoline must be a Tasks Trace reader across its call-out like the ftrace and BPF trampolines are, so document that in register_ftrace_direct(). That still leaves the few instructions before the reader is entered and after it is left. For BPF images those are in dynamically allocated text that rcu_tasks_trampoline_text() already treats as unmarked trampoline text, but the in-tree samples (and any similar user) place their trampolines in module .text. Add a sticky module::ftrace_direct_tramp flag, set by every register/modify path when the direct address is module text, and have rcu_tasks_trampoline_text() treat a task interrupted anywhere in such a module as a potential holdout. Other modules' text is unaffected. Assisted-by: LLM Signed-off-by: Josef Bacik --- include/linux/module.h | 7 +++++++ kernel/rcu/tasks.h | 21 ++++++++++++++++++--- kernel/trace/ftrace.c | 39 +++++++++++++++++++++++++++++++++++++++ 3 files changed, 64 insertions(+), 3 deletions(-) diff --git a/include/linux/module.h b/include/linux/module.h index 96cc98568eea..28488687cb01 100644 --- a/include/linux/module.h +++ b/include/linux/module.h @@ -521,6 +521,13 @@ struct module { unsigned int num_ftrace_callsites; unsigned long *ftrace_callsites; #endif +#ifdef CONFIG_DYNAMIC_FTRACE_WITH_DIRECT_CALLS + /* + * An ftrace direct-call trampoline lives in this module's text; see + * rcu_tasks_trampoline_text(). Sticky once set. + */ + bool ftrace_direct_tramp; +#endif #ifdef CONFIG_KPROBES void *kprobes_text_start; unsigned int kprobes_text_size; diff --git a/kernel/rcu/tasks.h b/kernel/rcu/tasks.h index 866768462850..ec54a27e47fa 100644 --- a/kernel/rcu/tasks.h +++ b/kernel/rcu/tasks.h @@ -1007,6 +1007,8 @@ bool __weak arch_rcu_tasks_trampoline_text(unsigned l= ong ip) * - the .text..rcu_tramp section, C glue called directly from such * trampolines before it has entered the reader; * - whatever the architecture adds via arch_rcu_tasks_trampoline_text(); + * - the text of a module that hosts an out-of-line ftrace direct-call + * trampoline (see ftrace_direct_mark_module()); * - the bytes after a kprobe that a pending jump optimization is about to * overwrite, the one synchronize_rcu_tasks() user with no trampoline. * @@ -1015,6 +1017,8 @@ bool __weak arch_rcu_tasks_trampoline_text(unsigned l= ong ip) */ bool rcu_tasks_trampoline_text(unsigned long ip) { + bool ret =3D true; + if (core_kernel_text(ip)) { if (ip >=3D (unsigned long)__rcu_tramp_text_start && ip < (unsigned long)__rcu_tramp_text_end) @@ -1022,9 +1026,20 @@ bool rcu_tasks_trampoline_text(unsigned long ip) return arch_rcu_tasks_trampoline_text(ip) || kprobe_in_optimized_region(ip); } - if (is_module_text_address(ip)) - return kprobe_in_optimized_region(ip); - return true; + +#ifdef CONFIG_MODULES + scoped_guard(rcu) { + struct module *mod =3D __module_text_address(ip); + + if (mod) { + ret =3D kprobe_in_optimized_region(ip); +#ifdef CONFIG_DYNAMIC_FTRACE_WITH_DIRECT_CALLS + ret =3D ret || READ_ONCE(mod->ftrace_direct_tramp); +#endif + } + } +#endif + return ret; } NOKPROBE_SYMBOL(rcu_tasks_trampoline_text); =20 diff --git a/kernel/trace/ftrace.c b/kernel/trace/ftrace.c index 53d5db60bfa5..efc4a518658a 100644 --- a/kernel/trace/ftrace.c +++ b/kernel/trace/ftrace.c @@ -6076,6 +6076,29 @@ static void reset_direct(struct ftrace_ops *ops, uns= igned long addr) ops->trampoline =3D 0; } =20 +/* + * A direct trampoline may live in module text rather than in dynamically + * allocated text that rcu_tasks_trampoline_text() recognises on its own (= see + * samples/ftrace/ftrace-direct*.c). The trampoline itself must be a Tasks + * Trace reader across its call-out (see register_ftrace_direct()); markin= g the + * owning module here covers the instructions before it enters that reader= and + * after it leaves it, where a task interrupted in the module's text must = not be + * counted as Tasks-RCU quiescent, so that ftrace_shutdown()'s + * synchronize_rcu_tasks() still keeps the module text from being freed un= der + * it. + */ +static void ftrace_direct_mark_module(unsigned long addr) +{ +#ifdef CONFIG_MODULES + struct module *mod; + + guard(rcu)(); + mod =3D __module_text_address(addr); + if (mod) + WRITE_ONCE(mod->ftrace_direct_tramp, true); +#endif +} + /** * register_ftrace_direct - Call a custom trampoline directly * for multiple functions registered in @ops @@ -6090,6 +6113,17 @@ static void reset_direct(struct ftrace_ops *ops, uns= igned long addr) * and save the parameters of the function being traced, and restore them * (or inject new ones if needed), before returning. * + * Nothing but Tasks RCU keeps the trampoline at @addr alive while a task = is + * executing it or is preempted in something it called. On architectures = that + * select HAVE_RCU_TRAMPOLINE_READERS, Tasks RCU only waits for such a tas= k if + * it is a Tasks Trace RCU reader, so the trampoline must enter one + * (rcu_read_lock_trace() or its assembly equivalent, see + * samples/ftrace/ftrace-direct.h) before calling out and leave it before + * returning, as the ftrace and BPF trampolines do. The few instructions + * before and after are covered by the irq-exit check: automatically for + * trampolines outside kernel and module text (e.g. BPF images), and via + * ftrace_direct_mark_module() for trampolines in module text. + * * Returns: * 0 on success * -EINVAL - The @ops object was already registered with this call or @@ -6169,6 +6203,7 @@ int register_ftrace_direct(struct ftrace_ops *ops, un= signed long addr) ops->flags |=3D MULTI_FLAGS; ops->trampoline =3D FTRACE_REGS_ADDR; ops->direct_call =3D addr; + ftrace_direct_mark_module(addr); =20 err =3D register_ftrace_function_nolock(ops); if (err) @@ -6237,6 +6272,8 @@ __modify_ftrace_direct(struct ftrace_ops *ops, unsign= ed long addr) =20 lockdep_assert_held_once(&direct_mutex); =20 + ftrace_direct_mark_module(addr); + /* Enable the tmp_ops to have the same functions as the direct ops */ ftrace_ops_init(&tmp_ops); tmp_ops.func_hash =3D ops->func_hash; @@ -6419,6 +6456,7 @@ int update_ftrace_direct_add(struct ftrace_ops *ops, = struct ftrace_hash *hash) hlist_for_each_entry(entry, &hash->buckets[i], hlist) { if (__ftrace_lookup_ip(direct_functions, entry->ip)) goto out_unlock; + ftrace_direct_mark_module(entry->direct); } } =20 @@ -6702,6 +6740,7 @@ int update_ftrace_direct_mod(struct ftrace_ops *ops, = struct ftrace_hash *hash, b tmp =3D __ftrace_lookup_ip(direct_hash, entry->ip); if (!tmp) continue; + ftrace_direct_mark_module(entry->direct); tmp->direct =3D entry->direct; } } --=20 2.55.0 From nobody Thu Sep 24 19:04:36 2026 Received: from mail-qk2-f13.google.com (mail-qk2-f13.google.com [74.125.230.205]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 023F149DBAA for ; Tue, 15 Sep 2026 13:17:59 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.205 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789478283; cv=none; b=AGjRKmTzcbbpoKBMjcw6gjm7fCLckqgdOEh+TaRoxp6ASyMMlXIqXEuEYSG1JA3lZn31a6xaeJjho79/q1Cp4UrOBvhV5k7iyxCyoGPy/zV2stUZ2eoG/MsHbqQVbgEBkTKEg8mVhr7AvbRk4KI/hR687ybaLM6bQaEjl7Ge2gk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789478283; c=relaxed/simple; bh=TXFGxaV6JtFDo/uJzdvTZ+O6U4MOxiEs7Jz7Lihg/aQ=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=r1nWaAvFQ8IzdtyWm4jUKfXxz4us+T4ROGGDYyEMafNoliqVCBgdSnzniHSSVgbayxe0AVroZPNIb6u6zPXYX8Q1lYJRIMnbPPhiFgknOvIrgSKEc6zr6XxeRkgltT0V/h/aHOePHHfkMLDPfH7xnFarUPs7gdrSU99Jz/T/GA4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com; spf=pass smtp.mailfrom=toxicpanda.com; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b=nG25Z0l9; arc=none smtp.client-ip=74.125.230.205 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b="nG25Z0l9" Received: by mail-qk2-f13.google.com with SMTP id af79cd13be357-939ca12ab70so408524985a.1 for ; Tue, 15 Sep 2026 06:17:59 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=toxicpanda.com; s=google; t=1789478279; x=1790083079; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=BZpG8a4K/YgKhJUrG5n9/0zVHbEqV2OSxLbwC0u3TSA=; b=nG25Z0l91ZM+th1tWAxHsJ7JGu0kvIPeeJhlvrfQRW8GLed5a50q4YpadNgQEYmCSR Ix72Nho79ujD+nPpWskArAs4OWdyANYjO8Aa3vKexrU+3F2lUqtSyeoOTqIz/YgW7IUz +t4rXbKzvkGXTMEIURhpcMNmQ3epqeJk59OI8OtbuPFffVddpQIVyBh1fMwUAZCJoJyF Yeuvliu38Kkjao4BBJ39UKO29jJOX91s9WLm46Degzty8+5TQIYisSMzh4HJXnaD7svT lxwwCh7Sf59jNYG8JG3Jjc+W4zVpxogObnB63PDL3fkLbJqs8atxVbL+Yjc9uxQJlkgj BocA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789478279; x=1790083079; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=BZpG8a4K/YgKhJUrG5n9/0zVHbEqV2OSxLbwC0u3TSA=; b=hFR6OuDte3DYq/R/W/d4AW2eADLx1/dW2oMr9XL63LV6GIvaBgNPvp/WYL1B6grgMu Pvhn269hyZsIL99sgWgOJlCa7eWqb/AHH40GMIZRZfw9Qj+fE/tkZf/oKcblVfZYMWP7 n5MTU4Qv88iRiNgxUPJUJuyIRXauobebadcUV8Z4U2ZOC4Qr9n6P68FsTEkA9RaTGi/O yeTIw1FG0dmk79pU99yOpcm5p5GIJaMfbTFMjwEs3Lik38Kigj5UZjzuVTbWljmCfmfk 9zW4ZHBCNFdI3SNwWAE+J2/dsMhW+fp3dtMqc58HubJQeGqvCwr3Kv63NIKA+ZOBgQ/B Xrcg== X-Forwarded-Encrypted: i=1; AKwUvBwYp/8BF3q4iHlMa9eIdy6QCrLMGgbf1uBgzHMzZQ+c6ZLTxrLW/1H5tqiot71qSpqpDa4HWmeDdaPNWS8=@vger.kernel.org X-Gm-Message-State: AFuF++kAL7JKUO9D+j8Eaio7AiTT6OwwWNlGlahVLEBBFDEeW0kKMOsF dpjumvW25xQod4XcbT0pXjM6o6oGcSd6ljQBBjLFysaf9ThG9Vcx/UuY580300GMAvY= X-Gm-Gg: AYBFou3GTwYXS4h4bDaMC6/HkEyEyIVyABwCCRxNWQQcBaWVW6eZq+40X76BeoWo1gc AOKKvkM8Gp2reXM3hpuojOMHBR02Dy0I8jfG0q9vJjgHjhFf2ksS/zzmu1cXN0Q78BDQ7TxKlQn IQRRRwEGelwdzi3woOpqMBsiJJoxzNY07Bgi3N1idgBwLxl+WkCWBEsvBJ9tlyEbo8kE9t8Q65d RjQd2BriR1+pRLR0OztOkT1kz7kypg+rICV/4sfiYpOL0EAkyBs62OBsoQqj1weDQB5eZFauTCj PirQ5x7YcTzzOrvYl13xv9dDWmwOVibkFIMwgojrOseMHwWq/mC8U4kmzQ/Vnoz+0NaKoVvqJVr MTEbocO0jDn38TaSA2Q83sgAaBaoho+PnNDyWPS+Xe0Pe837/29aPe4sfBhZGZZcMVPtALXyRAB 3sinXjBUFX/bfE3IBmBGE40yv2gE1wEb1OoJ4JvYHFoRnjjGNHi1sWZAHGf/TK+O0AmIsFD9mjb d2Xg0cT0ah2f9pfzl2RlcIy0SPjnRLFQtsnXEnfIqS0XMWZ/zPjqamw X-Received: by 2002:a05:620a:1b98:b0:93a:1092:48fa with SMTP id af79cd13be357-93a297e91bemr9631785a.4.1789478278679; Tue, 15 Sep 2026 06:17:58 -0700 (PDT) Received: from toxicpanda.com (ec2-34-228-114-98.compute-1.amazonaws.com. [34.228.114.98]) by smtp.gmail.com with ESMTPSA id af79cd13be357-939e86b7424sm1246377685a.32.2026.09.15.06.17.57 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 15 Sep 2026 06:17:58 -0700 (PDT) From: Josef Bacik Date: Tue, 15 Sep 2026 13:17:33 +0000 Subject: [PATCH RFC v3 06/13] bpf: Take a Tasks Trace reader in the trampoline glue Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260915-b4-rcu-tasks-preempt-qs-v3-6-0ad30c4c5ee7@toxicpanda.com> References: <20260915-b4-rcu-tasks-preempt-qs-v3-0-0ad30c4c5ee7@toxicpanda.com> In-Reply-To: <20260915-b4-rcu-tasks-preempt-qs-v3-0-0ad30c4c5ee7@toxicpanda.com> To: "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Boqun Feng , Thomas Gleixner , Peter Zijlstra , Steven Rostedt , Masami Hiramatsu , Mark Rutland , Jiri Olsa , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , x86@kernel.org, Catalin Marinas , Will Deacon , Puranjay Mohan , Xu Kuohai , "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Boqun Feng , Thomas Gleixner , Peter Zijlstra , Steven Rostedt , Masami Hiramatsu , Mark Rutland , Jiri Olsa , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , x86@kernel.org, Catalin Marinas , Will Deacon , Puranjay Mohan , Xu Kuohai Cc: Andy Lutomirski , Josh Triplett , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Juergen Gross , Luis Chamberlain , Ihor Solodrai , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, xen-devel@lists.xenproject.org, Andy Lutomirski , Josh Triplett , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Juergen Gross , Luis Chamberlain , Ihor Solodrai , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, xen-devel@lists.xenproject.org, Josef Bacik X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openssh-sha256; t=1789478262; l=11875; i=josef@toxicpanda.com; h=from:subject:message-id; bh=TXFGxaV6JtFDo/uJzdvTZ+O6U4MOxiEs7Jz7Lihg/aQ=; b=U1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgUBr36M/n0nWN0DNbnxwzIiCZez6MG JiruuNaSCI/zXsAAAAGcGF0YXR0AAAAAAAAAAZzaGE1MTIAAABTAAAAC3NzaC1lZDI1NTE5AAAA QMNOzSXH+MqwbhPyhl85UFDH60OjviU5by15ukzOGqWorm6a0N96E9GIqZviv+bWi33vuGmPHvW FnOzjv4m5SQI= X-Developer-Key: i=josef@toxicpanda.com; a=openssh; fpr=SHA256:C8kOX2QUJCMqnCX+KEeoqRAjLo9L+ELOSH2NSAJHqGA On CONFIG_TASKS_RCU_TRAMPOLINE_READERS kernels Tasks RCU keeps a BPF trampoline image allocated only while a task using it is a Tasks Trace RCU reader or is executing text that rcu_tasks_trampoline_text() recognises. The image itself is such text, but from it we call C glue in core kernel text -- __bpf_prog_enter*(), __bpf_prog_exit*(), __bpf_tramp_enter() and __bpf_tramp_exit() -- and today only the sleepable variants take rcu_read_lock_trace(). Rather than add anything to the JIT-emitted trampolines, close the gap in the glue: place all of it in .text..rcu_tramp via __rcu_trampoline so that a task interrupted in its prologue or epilogue is treated like one interrupted in the image, have every enter helper take rcu_read_lock_trace() before anything that could run out of line and every exit helper drop it last, and bracket the percpu_ref get and put in __bpf_tramp_enter()/__bpf_tramp_exit() the same way (the percpu_ref continues to cover the call to the original function). From the image's call to the glue's return the task is then always either in recognised text or a reader. The extra reader is compiled out on other configurations, and the sleepable paths are unchanged. bpf_tramp_image_put()'s call_rcu_tasks() stages are what now wait for those readers and for tasks interrupted in the image's own instructions. For images that call the original function nothing else changes: the percpu_ref pins the image from __bpf_tramp_enter() to __bpf_tramp_exit(), so only the instructions before and after need the grace periods they already get. A fentry-only image has no percpu_ref, and with one reader per prog a task walks reader, image, reader, image...; one grace period only guarantees such a task has left the reader or gap it was in when the grace period began, so on these kernels the fentry-only teardown requeues itself for one Tasks RCU grace period per prog (im->nr_progs, recorded when the image is built) before freeing. That costs nothing on the call path and a few more asynchronous grace periods on detach. Assisted-by: LLM Signed-off-by: Josef Bacik --- include/linux/bpf.h | 1 + kernel/bpf/trampoline.c | 107 +++++++++++++++++++++++++++++++++++++-------= ---- 2 files changed, 85 insertions(+), 23 deletions(-) diff --git a/include/linux/bpf.h b/include/linux/bpf.h index e57af902560c..b97ad80aacae 100644 --- a/include/linux/bpf.h +++ b/include/linux/bpf.h @@ -1366,6 +1366,7 @@ enum bpf_tramp_prog_type { struct bpf_tramp_image { void *image; int size; + int nr_progs; /* see bpf_tramp_image_put() */ struct bpf_ksym ksym; struct percpu_ref pcref; void *ip_after_call; diff --git a/kernel/bpf/trampoline.c b/kernel/bpf/trampoline.c index 90b70ea0d370..dbbc9bd7fd22 100644 --- a/kernel/bpf/trampoline.c +++ b/kernel/bpf/trampoline.c @@ -601,12 +601,25 @@ static void __bpf_tramp_image_put_rcu_tasks(struct rc= u_head *rcu) struct bpf_tramp_image *im; =20 im =3D container_of(rcu, struct bpf_tramp_image, rcu); - if (im->ip_after_call) + if (im->ip_after_call) { /* the case of fmod_ret/fexit trampoline and CONFIG_PREEMPTION=3Dy */ percpu_ref_kill(&im->pcref); - else + } else if (IS_ENABLED(CONFIG_TASKS_RCU_TRAMPOLINE_READERS) && + --im->nr_progs > 0) { + /* + * fentry-only trampoline on a reader-marked Tasks RCU: each prog + * runs in its own Tasks Trace reader with a few image + * instructions in between, and one rcu tasks grace period only + * guarantees that a task has moved on from the reader (or gap) it + * was in when the grace period started. A task walking the image + * therefore needs one grace period per prog before the image can + * go; keep requeueing until we have had that many. + */ + call_rcu_tasks(&im->rcu, __bpf_tramp_image_put_rcu_tasks); + } else { /* the case of fentry trampoline */ call_rcu_tasks(&im->rcu, __bpf_tramp_image_put_rcu); + } } =20 static void bpf_tramp_image_put(struct bpf_tramp_image *im) @@ -619,6 +632,14 @@ static void bpf_tramp_image_put(struct bpf_tramp_image= *im) * (which are few asm insns before __bpf_tramp_enter and * after __bpf_tramp_exit) * + * With CONFIG_TASKS_RCU_TRAMPOLINE_READERS, rcu tasks waits for a task + * in those asm insns because they are trampoline text, and for a task + * inside the glue or a prog because the glue makes it a + * rcu_read_lock_trace reader. The percpu_ref case is otherwise + * unchanged; the fentry-only case, having no percpu_ref across the + * whole image, takes one rcu tasks grace period per prog, see + * __bpf_tramp_image_put_rcu_tasks(). + * * The trampoline is unreachable before bpf_tramp_image_put(). * * First, patch the trampoline to avoid calling into fexit progs. @@ -776,6 +797,7 @@ static int bpf_trampoline_update(struct bpf_trampoline = *tr, bool lock_direct_mut err =3D PTR_ERR(im); goto out; } + im->nr_progs =3D total; =20 err =3D arch_prepare_bpf_trampoline(im, im->image, im->image + size, &tr->func.model, tr->flags, tnodes, @@ -1285,9 +1307,34 @@ static __always_inline u64 notrace bpf_prog_start_ti= me(void) * [2..MAX_U64] - execute bpf prog and record execution time. * This is start time. */ -static u64 notrace __bpf_prog_enter_recur(struct bpf_prog *prog, struct bp= f_tramp_run_ctx *run_ctx) +/* + * Where Tasks RCU is built on reader-marked trampolines + * (CONFIG_TASKS_RCU_TRAMPOLINE_READERS), the trampoline image that called= the + * glue below stays allocated only while the task is a Tasks Trace RCU rea= der + * or is executing text that rcu_tasks_trampoline_text() recognises: the i= mage + * itself, or this glue, which is therefore placed in .text..rcu_tramp + * (__rcu_trampoline). Each enter helper takes the reader before anything + * that could run out of line and each exit helper drops it last, so from = the + * image's call to the glue's return the task is always one or the other. = The + * sleepable variants already are such readers for their own reasons. + */ +static __always_inline void bpf_tramp_read_lock_trace(void) +{ + if (IS_ENABLED(CONFIG_TASKS_RCU_TRAMPOLINE_READERS)) + rcu_read_lock_trace(); +} + +static __always_inline void bpf_tramp_read_unlock_trace(void) +{ + if (IS_ENABLED(CONFIG_TASKS_RCU_TRAMPOLINE_READERS)) + rcu_read_unlock_trace(); +} + +static u64 notrace __rcu_trampoline +__bpf_prog_enter_recur(struct bpf_prog *prog, struct bpf_tramp_run_ctx *ru= n_ctx) __acquires(RCU) { + bpf_tramp_read_lock_trace(); rcu_read_lock_dont_migrate(); =20 run_ctx->saved_run_ctx =3D bpf_set_run_ctx(&run_ctx->run_ctx); @@ -1329,8 +1376,8 @@ static __always_inline void notrace update_prog_stats= (struct bpf_prog *prog, __update_prog_stats(prog, start); } =20 -static void notrace __bpf_prog_exit_recur(struct bpf_prog *prog, u64 start, - struct bpf_tramp_run_ctx *run_ctx) +static void notrace __rcu_trampoline +__bpf_prog_exit_recur(struct bpf_prog *prog, u64 start, struct bpf_tramp_r= un_ctx *run_ctx) __releases(RCU) { bpf_reset_run_ctx(run_ctx->saved_run_ctx); @@ -1338,15 +1385,17 @@ static void notrace __bpf_prog_exit_recur(struct bp= f_prog *prog, u64 start, update_prog_stats(prog, start); bpf_prog_put_recursion_context(prog); rcu_read_unlock_migrate(); + bpf_tramp_read_unlock_trace(); } =20 -static u64 notrace __bpf_prog_enter_lsm_cgroup(struct bpf_prog *prog, - struct bpf_tramp_run_ctx *run_ctx) +static u64 notrace __rcu_trampoline +__bpf_prog_enter_lsm_cgroup(struct bpf_prog *prog, struct bpf_tramp_run_ct= x *run_ctx) __acquires(RCU) { /* Runtime stats are exported via actual BPF_LSM_CGROUP * programs, not the shims. */ + bpf_tramp_read_lock_trace(); rcu_read_lock_dont_migrate(); =20 run_ctx->saved_run_ctx =3D bpf_set_run_ctx(&run_ctx->run_ctx); @@ -1354,17 +1403,18 @@ static u64 notrace __bpf_prog_enter_lsm_cgroup(stru= ct bpf_prog *prog, return NO_START_TIME; } =20 -static void notrace __bpf_prog_exit_lsm_cgroup(struct bpf_prog *prog, u64 = start, - struct bpf_tramp_run_ctx *run_ctx) +static void notrace __rcu_trampoline +__bpf_prog_exit_lsm_cgroup(struct bpf_prog *prog, u64 start, struct bpf_tr= amp_run_ctx *run_ctx) __releases(RCU) { bpf_reset_run_ctx(run_ctx->saved_run_ctx); =20 rcu_read_unlock_migrate(); + bpf_tramp_read_unlock_trace(); } =20 -u64 notrace __bpf_prog_enter_sleepable_recur(struct bpf_prog *prog, - struct bpf_tramp_run_ctx *run_ctx) +u64 notrace __rcu_trampoline +__bpf_prog_enter_sleepable_recur(struct bpf_prog *prog, struct bpf_tramp_r= un_ctx *run_ctx) { rcu_read_lock_trace(); migrate_disable(); @@ -1381,8 +1431,9 @@ u64 notrace __bpf_prog_enter_sleepable_recur(struct b= pf_prog *prog, return bpf_prog_start_time(); } =20 -void notrace __bpf_prog_exit_sleepable_recur(struct bpf_prog *prog, u64 st= art, - struct bpf_tramp_run_ctx *run_ctx) +void notrace __rcu_trampoline +__bpf_prog_exit_sleepable_recur(struct bpf_prog *prog, u64 start, + struct bpf_tramp_run_ctx *run_ctx) { bpf_reset_run_ctx(run_ctx->saved_run_ctx); =20 @@ -1392,8 +1443,8 @@ void notrace __bpf_prog_exit_sleepable_recur(struct b= pf_prog *prog, u64 start, rcu_read_unlock_trace(); } =20 -static u64 notrace __bpf_prog_enter_sleepable(struct bpf_prog *prog, - struct bpf_tramp_run_ctx *run_ctx) +static u64 notrace __rcu_trampoline +__bpf_prog_enter_sleepable(struct bpf_prog *prog, struct bpf_tramp_run_ctx= *run_ctx) { rcu_read_lock_trace(); migrate_disable(); @@ -1404,8 +1455,8 @@ static u64 notrace __bpf_prog_enter_sleepable(struct = bpf_prog *prog, return bpf_prog_start_time(); } =20 -static void notrace __bpf_prog_exit_sleepable(struct bpf_prog *prog, u64 s= tart, - struct bpf_tramp_run_ctx *run_ctx) +static void notrace __rcu_trampoline +__bpf_prog_exit_sleepable(struct bpf_prog *prog, u64 start, struct bpf_tra= mp_run_ctx *run_ctx) { bpf_reset_run_ctx(run_ctx->saved_run_ctx); =20 @@ -1414,10 +1465,11 @@ static void notrace __bpf_prog_exit_sleepable(struc= t bpf_prog *prog, u64 start, rcu_read_unlock_trace(); } =20 -static u64 notrace __bpf_prog_enter(struct bpf_prog *prog, - struct bpf_tramp_run_ctx *run_ctx) +static u64 notrace __rcu_trampoline +__bpf_prog_enter(struct bpf_prog *prog, struct bpf_tramp_run_ctx *run_ctx) __acquires(RCU) { + bpf_tramp_read_lock_trace(); rcu_read_lock_dont_migrate(); =20 run_ctx->saved_run_ctx =3D bpf_set_run_ctx(&run_ctx->run_ctx); @@ -1425,24 +1477,33 @@ static u64 notrace __bpf_prog_enter(struct bpf_prog= *prog, return bpf_prog_start_time(); } =20 -static void notrace __bpf_prog_exit(struct bpf_prog *prog, u64 start, - struct bpf_tramp_run_ctx *run_ctx) +static void notrace __rcu_trampoline +__bpf_prog_exit(struct bpf_prog *prog, u64 start, struct bpf_tramp_run_ctx= *run_ctx) __releases(RCU) { bpf_reset_run_ctx(run_ctx->saved_run_ctx); =20 update_prog_stats(prog, start); rcu_read_unlock_migrate(); + bpf_tramp_read_unlock_trace(); } =20 -void notrace __bpf_tramp_enter(struct bpf_tramp_image *tr) +/* + * The percpu_ref keeps the image alive across the call to the original + * function; the reader only has to cover getting and putting it, see abov= e. + */ +void notrace __rcu_trampoline __bpf_tramp_enter(struct bpf_tramp_image *tr) { + bpf_tramp_read_lock_trace(); percpu_ref_get(&tr->pcref); + bpf_tramp_read_unlock_trace(); } =20 -void notrace __bpf_tramp_exit(struct bpf_tramp_image *tr) +void notrace __rcu_trampoline __bpf_tramp_exit(struct bpf_tramp_image *tr) { + bpf_tramp_read_lock_trace(); percpu_ref_put(&tr->pcref); + bpf_tramp_read_unlock_trace(); } =20 bpf_trampoline_enter_t bpf_trampoline_enter(const struct bpf_prog *prog) --=20 2.55.0 From nobody Thu Sep 24 19:04:36 2026 Delivered-To: importer@patchew.org Received-SPF: pass (zohomail.com: domain of lists.xenproject.org designates 192.237.175.120 as permitted sender) client-ip=192.237.175.120; envelope-from=xen-devel-bounces@lists.xenproject.org; helo=lists.xenproject.org; Authentication-Results: mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of lists.xenproject.org designates 192.237.175.120 as permitted sender) smtp.mailfrom=xen-devel-bounces@lists.xenproject.org ARC-Seal: i=1; a=rsa-sha256; t=1789478300; cv=none; d=zohomail.com; s=zohoarc; b=gYG0vkgg/EKh7GhZ4ed7HjeT5fbpAHb/E2s2rMSYllifIeBwBJmkvBkMjpIAH6Mgva1nk9dZybX8QsgYDuWGyh3iXO4yDH1/HgPZFG2wjCofHCWeWo2bKSw0bUD8loS/b+v6a5LrfM9cdjJybETRWVE7eQe33he44G0fcmBXHSI= ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=zohomail.com; s=zohoarc; t=1789478300; h=Content-Type:Content-Transfer-Encoding:Cc:Cc:Date:Date:From:From:In-Reply-To:List-Subscribe:List-Post:List-Id:List-Help:List-Unsubscribe:MIME-Version:Message-ID:References:Sender:Subject:Subject:To:To:Message-Id:Reply-To; bh=pPH1+eiN4EhHtK6rDK3I01/qZvQ6oNbinGntBhruq3Q=; b=IBt+wikqmySyXNQ23Y9uAKQmkz0dhYiLT6J8jnTeBAMawSF7Z9EdThHiJF8OyiZZAVfkXiXFkeE0BeqXaNdNb2Neajpxd6gFEqK/FsPfYeL486KQMkin1JGnfeUo048GnLgu1mGsbXsHxB+13u1G2M7MIcpVoNbj9MNuuhSeP4E= ARC-Authentication-Results: i=1; mx.zohomail.com; dkim=pass; spf=pass (zohomail.com: domain of lists.xenproject.org designates 192.237.175.120 as permitted sender) smtp.mailfrom=xen-devel-bounces@lists.xenproject.org Return-Path: Received: from lists.xenproject.org (lists.xenproject.org [192.237.175.120]) by mx.zohomail.com with SMTPS id 1789478300006840.2505823077079; Tue, 15 Sep 2026 06:18:20 -0700 (PDT) Received: from list by lists.xenproject.org with outflank-mailman.1421712.1647445 (Exim 4.92) (envelope-from ) id 1x6T2z-0005m3-79; Tue, 15 Sep 2026 13:18:05 +0000 Received: by outflank-mailman (output) from mailman id 1421712.1647445; Tue, 15 Sep 2026 13:18:05 +0000 Received: from localhost ([127.0.0.1] helo=lists.xenproject.org) by lists.xenproject.org with esmtp (Exim 4.92) (envelope-from ) id 1x6T2z-0005lk-3M; Tue, 15 Sep 2026 13:18:05 +0000 Received: by outflank-mailman (input) for mailman id 1421712; Tue, 15 Sep 2026 13:18:03 +0000 Received: from mx.expurgate.net ([195.190.135.10]) by lists.xenproject.org with esmtp (Exim 4.92) (envelope-from ) id 1x6T2x-0005hh-By for xen-devel@lists.xenproject.org; Tue, 15 Sep 2026 13:18:03 +0000 Received: from mx.expurgate.net (helo=localhost) by mx.expurgate.net with esmtp id 1x6T2w-00EGDb-On for xen-devel@lists.xenproject.org; Tue, 15 Sep 2026 15:18:02 +0200 Received: from [10.42.69.11] (helo=localhost) by localhost with ESMTP (eXpurgate MTA 0.9.1) (envelope-from ) id 6aa9458a-bab6-0a2a0a5309dd-0a2a450bc220-0 for ; Tue, 15 Sep 2026 15:18:02 +0200 Received: from [74.125.230.140] (helo=mail-qv2-f12.google.com) by tlsNG-42698a.mxtls.expurgate.net with ESMTPS (eXpurgate 4.57.1) (envelope-from ) id 6aa94589-b7e8-0a2a450b0019-4a7de68cbe17-3 for ; Tue, 15 Sep 2026 15:18:02 +0200 Received: by mail-qv2-f12.google.com with SMTP id 6a1803df08f44-90cdfc9b6eeso34080786d6.0 for ; Tue, 15 Sep 2026 06:18:02 -0700 (PDT) Received: from toxicpanda.com (ec2-34-228-114-98.compute-1.amazonaws.com. [34.228.114.98]) by smtp.gmail.com with ESMTPSA id 6a1803df08f44-9120f47afc2sm120542886d6.20.2026.09.15.06.17.59 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 15 Sep 2026 06:17:59 -0700 (PDT) X-Outflank-Mailman: Message body and most headers restored to incoming version X-BeenThere: xen-devel@lists.xenproject.org List-Id: Xen developer discussion List-Unsubscribe: , List-Post: List-Help: List-Subscribe: , Errors-To: xen-devel-bounces@lists.xenproject.org Precedence: list Sender: "Xen-devel" Authentication-Results: eu.smtp.expurgate.cloud; dkim=pass header.s=google header.d=toxicpanda.com header.i="@toxicpanda.com" header.h="Cc:To:In-Reply-To:References:Message-Id:Content-Transfer-Encoding:Content-Type:MIME-Version:Subject:Date:From" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=toxicpanda.com; s=google; t=1789478281; x=1790083081; darn=lists.xenproject.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=pPH1+eiN4EhHtK6rDK3I01/qZvQ6oNbinGntBhruq3Q=; b=pdegXvcXrGOGWkHIqK5XopTSUhhctMzmxlyvokwlDr+7V/9rH3cWwRiT0YvCzX+OQD qcVOgZXRFaNcM5cvxvIMB0SJUKpJ2ccZ/CVT9q3uFgU3ghpVz8VxMaTZy36HxERmVncK 01Ub9upahL+7BN2TpuMj8QwxtcCQWhGUIVXobhJhjP2WeywYzhES4KzwVT8+56Rmpcn/ ISjJ73Ivg6Cq+HOzitMKswnptvE2vLukFyTybUt85TQVDtB7f6OKlllbIOC2xompU4VP diil+2SrocO2vSWOWXf/Kyyfg2Y+zAoDKILdSCRvLG0rY7RiT5c0YzgLyUdK3hwipIik J4Vg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789478281; x=1790083081; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=pPH1+eiN4EhHtK6rDK3I01/qZvQ6oNbinGntBhruq3Q=; b=hrffP7EnTbp4TMc6EBrOvskgssxmHqjjekIAmU2TuIEUu91u3maziXKUlAY1dCPsEd 7e5m8xTHxZjZwiMzLJUx+hK3G+Jzqp6es9s23yoNjU2aHItZzhQ59dlolJDAKqCvvmin a6itqTKRFAwoSoWywYjrIYG6+a5+BekKhsZhIBjLCsUwZwTOkRJ1VTTFBPSAggi3CFEY 17zeNri7x4ldWtiHnaoaLJVlLFNfIw9gAqiT/QbUFaDtilHn6yIaIEave3hvMgR7gfQZ 320nhpGD+bdsydHLdu7lMo6Dm3qSmTsUHnJ31/lXwjHJkn9JU0iz4oaXwTC260w+ypyy dIfQ== X-Forwarded-Encrypted: i=1; AKwUvBxkICOM4kvJIyMYcmETzPYoTturNX6Nvi2t5m8UcnoD3pE3KfVoAMm3ZMl/mlmgsPnDV+3rnpYvR20=@lists.xenproject.org X-Gm-Message-State: AFuF++lUSy0FVEP1MCXHFQWVlUn9rSRCUxkM6jYZA0DrTYFcvmtrdiKD YSmbvoLLRqSqgWfOaoB7yNBlBi2empzZ9v29xLiKvi6G+zqUPQwTbin5DUk2k1GjfskW4elFSSo vhgxeFk8PmA== X-Gm-Gg: AYBFou2eaHTTFmb7DbMAo09rh7DxT7GxXJNGkDM9MQfMQQ82suEtPna1xALFU3OxhAn JIySjHR7zusd5zTXnTFQ6fpNNr8KWPCIMWV2t8HGeuMDXDze3rei1wFPcxyThk6AYjnEn6E0wij ZtBkVoLNMUURkZtGaDzr3E8Lnt3rswh6MjK+04Oa9cYFP4kl5X9pvfkQbAMbiHwLz85IWU7k1PI Wx9pxkxF9ezsyNfetZ1BWhMWyMlXrDfG5ZFKPWmWdujArJ/EiXBT/klA/dwf7t8++KhgF6zxBFA XHH+Kqn/DWtGqBqIwfHcVdh2EVqJY0es6M1g+apDZFN8l3o+GNfJF3Us7aCbO5JVDk51wY97iih ZSnoY+I6LXo+xPatU6RFRUsJO6bZMzaOOQoTyuXRRO9IXp1zv4N9nyaVdferAtZ3LmdR5ub1LoB rTQOR+3aR4sQcM0KRPv0hfs4H0AtLcqhh8MoUNGurDscjRkhDkdp2FRzp5rjDjFjulG3nNR1Hhn mmimyi1FsrKsnAtXVZlqnUsyqVqfCAptXClhtQ2j9zRyOXEYoIlyCVC X-Received: by 2002:a0c:f40b:0:b0:907:dd19:fcb3 with SMTP id 6a1803df08f44-9122e55d2e0mr126180476d6.24.1789478280887; Tue, 15 Sep 2026 06:18:00 -0700 (PDT) From: Josef Bacik Date: Tue, 15 Sep 2026 13:17:34 +0000 Subject: [PATCH RFC v3 07/13] x86/ftrace: Take a Tasks Trace reader around ftrace_caller's call-out MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260915-b4-rcu-tasks-preempt-qs-v3-7-0ad30c4c5ee7@toxicpanda.com> References: <20260915-b4-rcu-tasks-preempt-qs-v3-0-0ad30c4c5ee7@toxicpanda.com> In-Reply-To: <20260915-b4-rcu-tasks-preempt-qs-v3-0-0ad30c4c5ee7@toxicpanda.com> To: "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Boqun Feng , Thomas Gleixner , Peter Zijlstra , Steven Rostedt , Masami Hiramatsu , Mark Rutland , Jiri Olsa , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , x86@kernel.org, Catalin Marinas , Will Deacon , Puranjay Mohan , Xu Kuohai , "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Boqun Feng , Thomas Gleixner , Peter Zijlstra , Steven Rostedt , Masami Hiramatsu , Mark Rutland , Jiri Olsa , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , x86@kernel.org, Catalin Marinas , Will Deacon , Puranjay Mohan , Xu Kuohai Cc: Andy Lutomirski , Josh Triplett , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Juergen Gross , Luis Chamberlain , Ihor Solodrai , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, xen-devel@lists.xenproject.org, Andy Lutomirski , Josh Triplett , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Juergen Gross , Luis Chamberlain , Ihor Solodrai , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, xen-devel@lists.xenproject.org, Josef Bacik X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openssh-sha256; t=1789478262; l=9964; i=josef@toxicpanda.com; h=from:subject:message-id; bh=EBgIXBiL5Z/I6XbFK2Nq9sAUVoT2dZ9k9OxjIjZChGA=; b=U1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgUBr36M/n0nWN0DNbnxwzIiCZez6MG JiruuNaSCI/zXsAAAAGcGF0YXR0AAAAAAAAAAZzaGE1MTIAAABTAAAAC3NzaC1lZDI1NTE5AAAA QJJZPzWCoHiGtBk3LJYNRCR9qGstVoSpG9sCCUThRLJ2AVjOv69KNmqUSOUhIKjcnlXv6dm2yvx k1FTR+gbuUgc= X-Developer-Key: i=josef@toxicpanda.com; a=openssh; fpr=SHA256:C8kOX2QUJCMqnCX+KEeoqRAjLo9L+ELOSH2NSAJHqGA X-purgate-ID: tlsNG-42698a/1789478282-194C39EA-1A4B1781/0/0 X-purgate-type: clean X-purgate-size: 9966 X-ZohoMail-DKIM: pass (identity @toxicpanda.com) X-ZM-MESSAGEID: 1789478301065158500 For HAVE_RCU_TRAMPOLINE_READERS the ftrace trampolines must be Tasks Trace RCU readers while they call out, since that -- and not the absence of a voluntary context switch -- is what synchronize_rcu_tasks() will wait for before ftrace_shutdown() frees a dynamic trampoline or its ops. Open-code rcu_read_lock_trace() and rcu_read_unlock_trace() in ftrace_caller and ftrace_regs_caller: bump current->trc_reader_nesting and, for the outermost reader, do the SRCU-fast per-CPU increment on rcu_tasks_trace_srcu_struct and stash the counter pointer in current->trc_reader_scp, exactly as the C inlines do (including the smp_mb() when CONFIG_TASKS_TRACE_RCU_NO_MB is not set). The lock sits before the function_trace_op load, because between that load and the call the ops pointer is protected only by Tasks RCU, and the unlock after the call returns. The sequences are inside the region that create_trampoline() copies for per-ops trampolines; their %rip-relative references are fixed up by text_poke_apply_relocation() like CALL_DEPTH_ACCOUNT's. %rax and %rcx are dead at both points. Two pieces of core text still run outside that reader while holding the address of a Tasks-RCU-protected trampoline they are about to enter: the static stubs themselves, whose direct-call tails keep a BPF trampoline address on the stack until the final RET, and, under CONFIG_MITIGATION_RETHUNK, the return thunk that RET expands to. Add an ftrace_static_tramp_end marker after ftrace_stub_direct_tramp and linker symbols around .text..__x86.return_thunk and .text..__x86.rethunk_safe, and provide arch_rcu_tasks_trampoline_text() covering [ftrace_caller, ftrace_static_tramp_end) and both thunk ranges so the irq-exit check treats a task interrupted there as a holdout. All of this is built only under CONFIG_TASKS_RCU_TRAMPOLINE_READERS, which x86 does not enable until a later patch. Assisted-by: LLM Signed-off-by: Josef Bacik --- arch/x86/kernel/asm-offsets.c | 8 +++++ arch/x86/kernel/ftrace.c | 43 +++++++++++++++++++++++++++ arch/x86/kernel/ftrace_64.S | 69 +++++++++++++++++++++++++++++++++++++++= ++++ arch/x86/kernel/vmlinux.lds.S | 4 +++ 4 files changed, 124 insertions(+) diff --git a/arch/x86/kernel/asm-offsets.c b/arch/x86/kernel/asm-offsets.c index 081816888f7a..876c3986419a 100644 --- a/arch/x86/kernel/asm-offsets.c +++ b/arch/x86/kernel/asm-offsets.c @@ -9,6 +9,7 @@ #include #include #include +#include #include #include #include @@ -46,6 +47,13 @@ static void __used common(void) #ifdef CONFIG_STACKPROTECTOR OFFSET(TASK_stack_canary, task_struct, stack_canary); #endif +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS + OFFSET(TASK_trc_reader_nesting, task_struct, trc_reader_nesting); + OFFSET(TASK_trc_reader_scp, task_struct, trc_reader_scp); + OFFSET(SRCU_srcu_ctrp, srcu_struct, srcu_ctrp); + OFFSET(SRCU_CTR_srcu_locks, srcu_ctr, srcu_locks); + OFFSET(SRCU_CTR_srcu_unlocks, srcu_ctr, srcu_unlocks); +#endif =20 BLANK(); OFFSET(pbe_address, pbe, address); diff --git a/arch/x86/kernel/ftrace.c b/arch/x86/kernel/ftrace.c index 17d6edfcb7e0..9babaed483eb 100644 --- a/arch/x86/kernel/ftrace.c +++ b/arch/x86/kernel/ftrace.c @@ -275,6 +275,49 @@ static inline void tramp_free(void *tramp) execmem_free(tramp); } =20 +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS +extern void ftrace_static_tramp_end(void); +extern char __return_thunk_start[], __return_thunk_end[]; +extern char __rethunk_safe_start[], __rethunk_safe_end[]; + +/* + * The SRCU-fast increments in TRACE_RCU_READ_LOCK/UNLOCK (ftrace_64.S) ar= e the + * this_cpu_inc() form. + */ +static_assert(!IS_ENABLED(CONFIG_NEED_SRCU_NMI_SAFE)); + +/* + * See rcu_tasks_trampoline_text(). Some core kernel text behaves like a + * trampoline for Tasks RCU purposes because a task executing there outside + * any Tasks Trace reader may still be about to enter a Tasks-RCU-protected + * trampoline whose address it already holds: + * + * - the static ftrace_caller / ftrace_regs_caller / ftrace_stub_direct_t= ramp + * stubs, which carry a direct-call target on the stack until their fin= al + * RET, and + * - the return thunks that RET expands to under CONFIG_MITIGATION_RETHUN= K, + * which run after leaving the stubs above and before landing in that + * target. + */ +bool arch_rcu_tasks_trampoline_text(unsigned long ip) +{ + if (ip >=3D (unsigned long)ftrace_caller && + ip < (unsigned long)ftrace_static_tramp_end) + return true; +#ifdef CONFIG_MITIGATION_RETPOLINE + if (ip >=3D (unsigned long)__return_thunk_start && + ip < (unsigned long)__return_thunk_end) + return true; +#endif +#ifdef CONFIG_MITIGATION_SRSO + if (ip >=3D (unsigned long)__rethunk_safe_start && + ip < (unsigned long)__rethunk_safe_end) + return true; +#endif + return false; +} +#endif /* CONFIG_TASKS_RCU_TRAMPOLINE_READERS */ + /* Defined as markers to the end of the ftrace default trampolines */ extern void ftrace_regs_caller_end(void); extern void ftrace_caller_end(void); diff --git a/arch/x86/kernel/ftrace_64.S b/arch/x86/kernel/ftrace_64.S index 62c1c93aa1c6..5d8cb3861978 100644 --- a/arch/x86/kernel/ftrace_64.S +++ b/arch/x86/kernel/ftrace_64.S @@ -7,6 +7,7 @@ #include #include #include +#include #include #include #include @@ -145,6 +146,53 @@ SYM_FUNC_END(ftrace_stub_graph) =20 #ifdef CONFIG_DYNAMIC_FTRACE =20 +/* + * Open-coded rcu_read_lock_trace() / rcu_read_unlock_trace(), see + * include/linux/rcupdate_trace.h and CONFIG_HAVE_RCU_TRAMPOLINE_READERS: = the + * trampoline and the ftrace_ops it is about to load are kept alive by Tas= ks + * RCU only while we are inside this reader, so the lock must precede the + * function_trace_op load and the unlock must follow the call. These live + * inside the region copied into dynamic trampolines; the %rip-relative + * references are fixed up by text_poke_apply_relocation() in + * create_trampoline(). Clobbers %rax, %rcx and flags. + */ +.macro TRACE_RCU_READ_LOCK +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS + movq PER_CPU_VAR(current_task), %rcx + movl TASK_trc_reader_nesting(%rcx), %eax + incl TASK_trc_reader_nesting(%rcx) + testl %eax, %eax + jnz .Ltrl_nested_\@ + movq rcu_tasks_trace_srcu_struct+SRCU_srcu_ctrp(%rip), %rax + incq %gs:SRCU_CTR_srcu_locks(%rax) + movq %rax, TASK_trc_reader_scp(%rcx) +#ifndef CONFIG_TASKS_TRACE_RCU_NO_MB + lock addl $0, -4(%rsp) /* smp_mb() */ +#endif +.Ltrl_nested_\@: +#endif +.endm + +.macro TRACE_RCU_READ_UNLOCK +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS + movq PER_CPU_VAR(current_task), %rcx + movl TASK_trc_reader_nesting(%rcx), %eax + subl $1, %eax + jnz .Ltru_nested_\@ + /* Outermost: pick up scp before an interrupt can see nesting =3D=3D 0. */ + movq TASK_trc_reader_scp(%rcx), %rax + movl $0, TASK_trc_reader_nesting(%rcx) +#ifndef CONFIG_TASKS_TRACE_RCU_NO_MB + lock addl $0, -4(%rsp) /* smp_mb() */ +#endif + incq %gs:SRCU_CTR_srcu_unlocks(%rax) + jmp .Ltru_done_\@ +.Ltru_nested_\@: + movl %eax, TASK_trc_reader_nesting(%rcx) +.Ltru_done_\@: +#endif +.endm + SYM_FUNC_START(__fentry__) ANNOTATE_NOENDBR CALL_DEPTH_ACCOUNT @@ -163,6 +211,8 @@ SYM_FUNC_START(ftrace_caller) leaq MCOUNT_REG_SIZE+8(%rsp), %rcx movq %rcx, RSP(%rsp) =20 + TRACE_RCU_READ_LOCK + SYM_INNER_LABEL(ftrace_caller_op_ptr, SYM_L_GLOBAL) ANNOTATE_NOENDBR /* Load the ftrace_ops into the 3rd parameter */ @@ -181,6 +231,8 @@ SYM_INNER_LABEL(ftrace_call, SYM_L_GLOBAL) ANNOTATE_NOENDBR call ftrace_stub =20 + TRACE_RCU_READ_UNLOCK + /* Handlers can change the RIP */ movq RIP(%rsp), %rax movq %rax, MCOUNT_REG_SIZE(%rsp) @@ -209,6 +261,8 @@ SYM_FUNC_START(ftrace_regs_caller) =20 CALL_DEPTH_ACCOUNT =20 + TRACE_RCU_READ_LOCK + SYM_INNER_LABEL(ftrace_regs_caller_op_ptr, SYM_L_GLOBAL) ANNOTATE_NOENDBR /* Load the ftrace_ops into the 3rd parameter */ @@ -246,6 +300,8 @@ SYM_INNER_LABEL(ftrace_regs_call, SYM_L_GLOBAL) ANNOTATE_NOENDBR call ftrace_stub =20 + TRACE_RCU_READ_UNLOCK + /* Copy flags back to SS, to restore them */ movq EFLAGS(%rsp), %rax movq %rax, MCOUNT_REG_SIZE(%rsp) @@ -328,6 +384,19 @@ SYM_FUNC_START(ftrace_stub_direct_tramp) RET SYM_FUNC_END(ftrace_stub_direct_tramp) =20 +/* + * [ftrace_caller, ftrace_static_tramp_end) is treated as trampoline text = by + * rcu_tasks_trampoline_text(): outside TRACE_RCU_READ_LOCK/UNLOCK the stu= bs + * may still hold a direct-call trampoline address (ORIG_RAX / the return + * address they RET to) that only Tasks RCU keeps alive. With return thun= ks + * the RET itself runs elsewhere; arch_rcu_tasks_trampoline_text() covers + * those too. + */ +SYM_CODE_START_NOALIGN(ftrace_static_tramp_end) + UNWIND_HINT_UNDEFINED + ANNOTATE_NOENDBR +SYM_CODE_END(ftrace_static_tramp_end) + #else /* ! CONFIG_DYNAMIC_FTRACE */ =20 SYM_FUNC_START(__fentry__) diff --git a/arch/x86/kernel/vmlinux.lds.S b/arch/x86/kernel/vmlinux.lds.S index 2438b89a4620..e546283dc267 100644 --- a/arch/x86/kernel/vmlinux.lds.S +++ b/arch/x86/kernel/vmlinux.lds.S @@ -151,7 +151,9 @@ SECTIONS * definition. */ . =3D srso_alias_untrain_ret | (1 << 2) | (1 << 8) | (1 << 14) | (1 << 2= 0); + __rethunk_safe_start =3D .; *(.text..__x86.rethunk_safe) + __rethunk_safe_end =3D .; #endif ALIGN_ENTRY_TEXT_END =20 @@ -162,7 +164,9 @@ SECTIONS SOFTIRQENTRY_TEXT #ifdef CONFIG_MITIGATION_RETPOLINE *(.text..__x86.indirect_thunk) + __return_thunk_start =3D .; *(.text..__x86.return_thunk) + __return_thunk_end =3D .; #endif STATIC_CALL_TEXT *(.gnu.warning) --=20 2.55.0 From nobody Thu Sep 24 19:04:36 2026 Received: from mail-qk2-f13.google.com (mail-qk2-f13.google.com [74.125.230.205]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 820DF4A1E0B for ; Tue, 15 Sep 2026 13:18:04 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.205 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789478288; cv=none; b=EeDlh+Rx4krcQryCo+bPszdcBcTJmaKM2rLiKLQa8O+ouZ4OchS9T4eJ1ubXe15PCzAic1iHFElXp/4ko6n2kzA0NDTcOvzdfym7Oppkex8OttHekY3td//EDwe9q4vJeBJKR0hObIHiTdmt9c/JKZdk+vtq2rifEG+EK/qdbis= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789478288; c=relaxed/simple; bh=cyOEpv7E+99KbsYBQF12i4qmiz/SB2W5CJ1v6BiG3CI=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=GPbkRd3b+gVXO/XbhGn7GIX/jtxm4SQJpVLidl0cqNA1RiHQ/ym3URBXCAa74L5A01vyTK31Se4unhnHlkUvrFW4fNCvs0ZWSSYsA/32G6u8eyFmo8wm0YKtiwXJsJ6ET/uXwjNB4GoFrJd5hK2mSmpMM6/v2IwJDLIH9b53ijc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com; spf=pass smtp.mailfrom=toxicpanda.com; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b=gYS2Bu1V; arc=none smtp.client-ip=74.125.230.205 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b="gYS2Bu1V" Received: by mail-qk2-f13.google.com with SMTP id d75a77b69052e-52fb76906adso46629811cf.0 for ; Tue, 15 Sep 2026 06:18:04 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=toxicpanda.com; s=google; t=1789478283; x=1790083083; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=9lDajmUPmaz8BTxhPEk6XaACKGtk0v9RwI05AN+Obys=; b=gYS2Bu1VPt21rSZr6oDMTEAKF0+UiLvIlptr2/gtMXk5FXlmg/u2TQksoBcSz6yRI9 +cw53Lz8Rgvomw5rm02teSz9gn9q1A7aCDekbNFumTcalcIWbTmmjQnOmMY/V69LWT6L DuwLk6kBVD0Ri2gt4Bml89x58yfI3lZ/eb1LPV4OTjJ0jeW2G6gGmzJ01ZcYHORj+4RV sFz8T4nKlFUrg7fTgLLeagu/JLWcNCOEfvLJd5e2GJMTExSnBhIue5Gl5yZuDwTFxwZs y/d6z4aG9nlnzM45zleX+3kkRwWCLD4s9ZzGaoUYOEs+LgqzETgkaJ6MLn/LwVOXy/wv uRCg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789478283; x=1790083083; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=9lDajmUPmaz8BTxhPEk6XaACKGtk0v9RwI05AN+Obys=; b=xtxs4QFAi/9W+D0evvdT7X+4/SuYIO7B3odN9FkvlaB6DRY1gMmAaZi2iCZ5yk9bfg GKWPGZrmtDdfmMvV8VWs5+a2bKf+y5HAUVpoodm4EyQuCiN1N3xXKMY7NWG51NuWmBEy a0Q2/hCPEPK//cEw/G5oiv+ptBnMJoqzxnhH1X1wbk3BWZoL4TfuMwhy1oDMw9PVaa/8 /LAjswCTHw/OPQRphBiF1VxKrJHQEUnCh/bWhKfB3JAeNqzX0NpAk42KaHNYseSBC6IJ JyL47JLpeNfiupm04aitU0qjuvdtpVzcGgyS0NM6bRKpuv+65Y/clGuZeFYLB+y0XTlj THXg== X-Forwarded-Encrypted: i=1; AKwUvBxXNWsFE+/sasHBFgzj7OO3Cu88vSYJz416ks2J03ORWkmfXUNNAv1MyUUUfNNZQKao9+4Qvg2s/AwSvdg=@vger.kernel.org X-Gm-Message-State: AFuF++kpI6Z8DTCTB6PW8eyP0kDWvDLsWUGWkckSJMviF2nVrDT76xf/ IV9ylKojUPIg3LXT3cOKu3t9reX2cp6VFuhUZfITqSe6noeEnBsuxYD7KHsq4Edr0gw= X-Gm-Gg: AYBFou1F1CCBlLocyBPod+GUBNrREaHcIePvKXcP1+ew3+DOj1YG7jPle097dLIK+Ee XNBbujZhufvO36KCSTB0mvmNa0jMIL8ImT+N14KtNcP5Ig/k/x0aiN4tYBt1EfN244NgzaFQsQ3 Nv8bHah6FRUZIC8dc/D5ZOPJB36ujV/Sbz+yMcsUpz5sGgJWre8DBWuoIPMAih3ZXMp3BlMs6EJ c1024NkKVJiOt8Ue1d8+5AbppJ4iM3c1weDas9teynnh3ab9i7z8XbH0mWiL1QCqdItV9WQGBtk zeP3SUZjstdM8sSa/c2Qe/x9R9dgJoq/vuGAsLD5RucvSZNRkxFQrCAkxiMvFZU8ibgvMBNNcoQ /LGwK3ZtwM+Xxn7i06tlEss604WU6a/uHYWF6ufu2HN2PRD530iuZGFbgHbAiEiB71q/H0oJbH3 EZzBTcjvTo++8HrvCWcJgRt+sElIwG2dt+NetUnaM/at6pUmtqaVGWDpZYWLXh85olUEPfrdQyY Sbyk+3+nebXwUDXjZzk4mNnTGBT4NWOd9ye9Kz0U72Z7Tsr81cDUqZP X-Received: by 2002:ac8:5882:0:b0:531:4d7:bbe1 with SMTP id d75a77b69052e-5310d0620aamr106343871cf.50.1789478282597; Tue, 15 Sep 2026 06:18:02 -0700 (PDT) Received: from toxicpanda.com (ec2-34-228-114-98.compute-1.amazonaws.com. [34.228.114.98]) by smtp.gmail.com with ESMTPSA id d75a77b69052e-530ca4899dasm125209321cf.12.2026.09.15.06.18.01 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 15 Sep 2026 06:18:02 -0700 (PDT) From: Josef Bacik Date: Tue, 15 Sep 2026 13:17:35 +0000 Subject: [PATCH RFC v3 08/13] x86/kprobes: Take a Tasks Trace reader in the optprobe template Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260915-b4-rcu-tasks-preempt-qs-v3-8-0ad30c4c5ee7@toxicpanda.com> References: <20260915-b4-rcu-tasks-preempt-qs-v3-0-0ad30c4c5ee7@toxicpanda.com> In-Reply-To: <20260915-b4-rcu-tasks-preempt-qs-v3-0-0ad30c4c5ee7@toxicpanda.com> To: "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Boqun Feng , Thomas Gleixner , Peter Zijlstra , Steven Rostedt , Masami Hiramatsu , Mark Rutland , Jiri Olsa , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , x86@kernel.org, Catalin Marinas , Will Deacon , Puranjay Mohan , Xu Kuohai , "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Boqun Feng , Thomas Gleixner , Peter Zijlstra , Steven Rostedt , Masami Hiramatsu , Mark Rutland , Jiri Olsa , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , x86@kernel.org, Catalin Marinas , Will Deacon , Puranjay Mohan , Xu Kuohai Cc: Andy Lutomirski , Josh Triplett , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Juergen Gross , Luis Chamberlain , Ihor Solodrai , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, xen-devel@lists.xenproject.org, Andy Lutomirski , Josh Triplett , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Juergen Gross , Luis Chamberlain , Ihor Solodrai , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, xen-devel@lists.xenproject.org, Josef Bacik X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openssh-sha256; t=1789478262; l=3977; i=josef@toxicpanda.com; h=from:subject:message-id; bh=cyOEpv7E+99KbsYBQF12i4qmiz/SB2W5CJ1v6BiG3CI=; b=U1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgUBr36M/n0nWN0DNbnxwzIiCZez6MG JiruuNaSCI/zXsAAAAGcGF0YXR0AAAAAAAAAAZzaGE1MTIAAABTAAAAC3NzaC1lZDI1NTE5AAAA QPs7HUne0IzQlm9wRQkt109QqoRgFF/pfpz8lJ+9BcfBer6XnRW2+MjSGUTLZ6JFUaWA9Uyg01K Y4lSDYJOfogI= X-Developer-Key: i=josef@toxicpanda.com; a=openssh; fpr=SHA256:C8kOX2QUJCMqnCX+KEeoqRAjLo9L+ELOSH2NSAJHqGA The jump-optimized kprobe template calls optimized_callback() from a dynamically allocated slot with preemption enabled, and only Tasks RCU keeps that slot alive under a task preempted in the callback. For HAVE_RCU_TRAMPOLINE_READERS that means the template must be a Tasks Trace reader across the call, so open-code rcu_read_lock_trace() and rcu_read_unlock_trace() around it as ftrace_64.S does. The template lives in .rodata and is memcpy()d into each slot without relocation processing, so the references to current_task and rcu_tasks_trace_srcu_struct are absolute (R_X86_64_32S, relocated for KASLR like any other) rather than %rip-relative. %rax and %rcx have already been saved by SAVE_REGS_STRING and are dead after the call. The slot itself is dynamically allocated text, so the instructions before the lock and after the unlock are covered by the irq-exit check. 64-bit only; 32-bit x86 does not take part. Assisted-by: LLM Signed-off-by: Josef Bacik --- arch/x86/kernel/kprobes/opt.c | 44 +++++++++++++++++++++++++++++++++++++++= ++++ 1 file changed, 44 insertions(+) diff --git a/arch/x86/kernel/kprobes/opt.c b/arch/x86/kernel/kprobes/opt.c index 3f8fea52619f..68a5de6cdabe 100644 --- a/arch/x86/kernel/kprobes/opt.c +++ b/arch/x86/kernel/kprobes/opt.c @@ -31,6 +31,7 @@ #include #include #include +#include =20 #include "common.h" =20 @@ -101,6 +102,47 @@ static void synthesize_set_arg1(kprobe_opcode_t *addr,= unsigned long val) *(unsigned long *)addr =3D val; } =20 +/* + * Open-coded rcu_read_lock_trace() / rcu_read_unlock_trace() around the c= all + * to optimized_callback(), see CONFIG_HAVE_RCU_TRAMPOLINE_READERS and the + * equivalent macros in ftrace_64.S. The template is memcpy()d into the s= lot + * without relocation processing, so memory references must be absolute + * rather than %rip-relative. %rax and %rcx are free at both points. + */ +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS +#ifndef CONFIG_TASKS_TRACE_RCU_NO_MB +#define OPTPROBE_TRACE_RCU_MB " lock addl $0, -4(%rsp)\n" +#else +#define OPTPROBE_TRACE_RCU_MB +#endif +#define OPTPROBE_TRACE_RCU_READ_LOCK \ + " movq %gs:current_task, %rcx\n" \ + " movl " __stringify(TASK_trc_reader_nesting) "(%rcx), %eax\n" \ + " incl " __stringify(TASK_trc_reader_nesting) "(%rcx)\n" \ + " testl %eax, %eax\n" \ + " jnz 1f\n" \ + " movq rcu_tasks_trace_srcu_struct+" __stringify(SRCU_srcu_ctrp) ", %rax= \n" \ + " incq %gs:" __stringify(SRCU_CTR_srcu_locks) "(%rax)\n" \ + " movq %rax, " __stringify(TASK_trc_reader_scp) "(%rcx)\n" \ + OPTPROBE_TRACE_RCU_MB \ + "1:\n" +#define OPTPROBE_TRACE_RCU_READ_UNLOCK \ + " movq %gs:current_task, %rcx\n" \ + " movl " __stringify(TASK_trc_reader_nesting) "(%rcx), %eax\n" \ + " subl $1, %eax\n" \ + " jnz 2f\n" \ + " movq " __stringify(TASK_trc_reader_scp) "(%rcx), %rax\n" \ + " movl $0, " __stringify(TASK_trc_reader_nesting) "(%rcx)\n" \ + OPTPROBE_TRACE_RCU_MB \ + " incq %gs:" __stringify(SRCU_CTR_srcu_unlocks) "(%rax)\n" \ + " jmp 3f\n" \ + "2: movl %eax, " __stringify(TASK_trc_reader_nesting) "(%rcx)\n" \ + "3:\n" +#else +#define OPTPROBE_TRACE_RCU_READ_LOCK +#define OPTPROBE_TRACE_RCU_READ_UNLOCK +#endif + asm ( ".pushsection .rodata\n" ".global optprobe_template_entry\n" @@ -114,6 +156,7 @@ asm ( "optprobe_template_clac:\n" ASM_NOP3 SAVE_REGS_STRING + OPTPROBE_TRACE_RCU_READ_LOCK " movq %rsp, %rsi\n" ".global optprobe_template_val\n" "optprobe_template_val:\n" @@ -122,6 +165,7 @@ asm ( ".global optprobe_template_call\n" "optprobe_template_call:\n" ASM_NOP5 + OPTPROBE_TRACE_RCU_READ_UNLOCK /* Copy 'regs->flags' into 'regs->ss'. */ " movq 18*8(%rsp), %rdx\n" " movq %rdx, 20*8(%rsp)\n" --=20 2.55.0 From nobody Thu Sep 24 19:04:36 2026 Received: from mail-qk2-f13.google.com (mail-qk2-f13.google.com [74.125.230.205]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AF80D4A2055 for ; Tue, 15 Sep 2026 13:18:06 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.205 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789478290; cv=none; b=LTmR4RjrGJGbfFhg2mBdEMuTOFwl9/oJ9PxgJ1KYXoycMdXDfKZdAKYu8YE5NJrxabYZwuGkqZXhtZfOCg9e/pU60DyAno3DdIGBR1aELOeVBcfMx75nlsCgDiIUUSbPwRqnOeJhwVeO2gDYv7hK9ZFU9IaOlsmRcosATWbIWBQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789478290; c=relaxed/simple; bh=Aqm0C5BBMBkqFusbG93kcm6leRi28F4VKStTPzg6dAE=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=dSyiY3nf4H159XE+PIdEl4iPD2JOBjgwnbPfHNb/TFihj4l/uXSk4YG1ZSK2krCCU5N2RH24GA/uNIjoXhNKluHJDrmUMjkrE4vNkVrQw6359PbdSRyBUQvU8AuQBZr2u8BWltyW9o6rR87lWxzb9+2XC5vBUfwoogCxs+hiNfM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com; spf=pass smtp.mailfrom=toxicpanda.com; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b=J6b+XA5m; arc=none smtp.client-ip=74.125.230.205 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b="J6b+XA5m" Received: by mail-qk2-f13.google.com with SMTP id d75a77b69052e-52fb76c9deeso50608481cf.1 for ; Tue, 15 Sep 2026 06:18:06 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=toxicpanda.com; s=google; t=1789478285; x=1790083085; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=ua8lq0Nx07F8Y+BgJ1JiUwgoDOin8IZmspGmQtj7ybE=; b=J6b+XA5mbLLkC1NUkEd3V9s5zF0cMihA22ycfERQDlxpyW3E09L4kuosySFa1KXRea 0F/j7c61EG2Ij0XN/0rSLSE9w5mQ1DbDnJIAokJY9SfzbvvkQHdoW2t3VLhXXvvBy0t1 lQwzaUzQ8xZQyds3fY5ZYTAV06JXJdWrFeTIG2rGYuGVApQTPuUCfuoaxFXfrMJ1NYuY HuED3nhFRk6ULfFWUPdTBQNJv4KHp2wOtneRezHUqR7LVldvkj//RZI3RbAd/+L9aOYj 2wlTU4bCsXviNzlmJzYB91jVy+RVbpnQaFSvpOe28HkqVdtWY+Y4UbBqRTvUH3wILJvQ tumA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789478285; x=1790083085; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=ua8lq0Nx07F8Y+BgJ1JiUwgoDOin8IZmspGmQtj7ybE=; b=1+4KxmJ23faIP/b7rpwg8pviGF2ooQreyUsbv3np4gPm8Y+B8kjBvme0jdMQ9l4Dy7 Ki+d1PNWbb6nE+fpeCVAx+l3Fdkhpgl5mszd/bB3pHjvR/SjSn1CG4fXLvAoPTigxX47 R8G1x+7BbZR11DXu+YfG2SEoXZD0uix8SL9R8ObR8L9Lzsvxd/OEQxRNLkTPVs0kW6yW 56jwJC2O5kcpbt5f0MOzmo2/ucLKJiUoRdeSFhsqAocQxai5A1bd14pF90kn9LX1DpZ0 IMVsjqRTnpCNVQSoATpeSY44ClWdO7/2pMJRLjN9UjA//hqytPp/mFQ/kSLTaKVrtWFN aVkw== X-Forwarded-Encrypted: i=1; AKwUvBwIAhxq+pd6RGyRdsZXllNN/lwG0JeOlg2ht21Xn069YwfXTCVt6zOLsi5GHsCAwt8lcGgCSX9nPjQGwLY=@vger.kernel.org X-Gm-Message-State: AFuF++nxTHCmpXUQ0qWVzigJ8Xhahu3WCJnqwO16C1QiMv0CwTNu87UK EAw0bN20HQUWix2KFHuPID2XWPcdwr8DX+nJjBAvKDIFUp8LotPbWmeeyqt3GIoVQtU= X-Gm-Gg: AYBFou3OhpMEjIkPJRjstI4cvnqDOp5owg5AtwIKN8nujhWbe3ISSzM9OeMbXggDOJD cb/ENi7dG6ZHHSopJ9+9XHSe5IGPW+qV+S3vizYMZXbyPb2t2wsNGES0iaaWRTW2PU3LrB7fvzR BmG1AoH5CFu8DCHgn5yPtwC/usSkHQjgV7VzpEdw84KpezoFG7rBeNoz93VwYp5Hn+/pD4Mz3Hp j+0+YFfgPrYKWRwSi17lu0GdNi41ienJLgcWVbgR6tMmZp42Z7QBQ07Pkb7p1MW7gaFu12P7x1d zj2Fu9k2nTlalzUO6NfoCjCbTmNP+HFz5XKssI53OZ7xOLepo2GECh7C7oZb9P/yWRT9axGGiIF jnyeO6S7q+N8LRl9io76ajzXlks8BrV1tvGocp50uM6HGOtOg9rBiciJ26iEnDyK9kf81EaTbt0 io5ILyfiLB4LrxOBBriTWsrrPUcGKKhNw+R3aO4RA4cqZOjqd2vfiv56HEOXaeK+n3uUOTJBOBF 0M+gy73ZxMseb1lU12sfN1fB2lL5jx4yGemaYxEHEh2tME4fh1Sm2Tb X-Received: by 2002:a05:622a:1919:b0:530:98bc:3974 with SMTP id d75a77b69052e-5310cf631f8mr109341941cf.22.1789478284669; Tue, 15 Sep 2026 06:18:04 -0700 (PDT) Received: from toxicpanda.com (ec2-34-228-114-98.compute-1.amazonaws.com. [34.228.114.98]) by smtp.gmail.com with ESMTPSA id d75a77b69052e-530f6849802sm80503791cf.10.2026.09.15.06.18.03 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 15 Sep 2026 06:18:03 -0700 (PDT) From: Josef Bacik Date: Tue, 15 Sep 2026 13:17:36 +0000 Subject: [PATCH RFC v3 09/13] arm64: ftrace: Take a Tasks Trace reader around ftrace_caller's call-out Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260915-b4-rcu-tasks-preempt-qs-v3-9-0ad30c4c5ee7@toxicpanda.com> References: <20260915-b4-rcu-tasks-preempt-qs-v3-0-0ad30c4c5ee7@toxicpanda.com> In-Reply-To: <20260915-b4-rcu-tasks-preempt-qs-v3-0-0ad30c4c5ee7@toxicpanda.com> To: "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Boqun Feng , Thomas Gleixner , Peter Zijlstra , Steven Rostedt , Masami Hiramatsu , Mark Rutland , Jiri Olsa , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , x86@kernel.org, Catalin Marinas , Will Deacon , Puranjay Mohan , Xu Kuohai , "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Boqun Feng , Thomas Gleixner , Peter Zijlstra , Steven Rostedt , Masami Hiramatsu , Mark Rutland , Jiri Olsa , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , x86@kernel.org, Catalin Marinas , Will Deacon , Puranjay Mohan , Xu Kuohai Cc: Andy Lutomirski , Josh Triplett , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Juergen Gross , Luis Chamberlain , Ihor Solodrai , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, xen-devel@lists.xenproject.org, Andy Lutomirski , Josh Triplett , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Juergen Gross , Luis Chamberlain , Ihor Solodrai , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, xen-devel@lists.xenproject.org, Josef Bacik X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openssh-sha256; t=1789478262; l=7839; i=josef@toxicpanda.com; h=from:subject:message-id; bh=Aqm0C5BBMBkqFusbG93kcm6leRi28F4VKStTPzg6dAE=; b=U1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgUBr36M/n0nWN0DNbnxwzIiCZez6MG JiruuNaSCI/zXsAAAAGcGF0YXR0AAAAAAAAAAZzaGE1MTIAAABTAAAAC3NzaC1lZDI1NTE5AAAA QGtKBbzIh64iNKja7atdfYbFNhvzop2423iTP8KKI74RhSORDHi3DpZgpvvVVV7stG2frrx5HDn 5ncEuv/+0CQs= X-Developer-Key: i=josef@toxicpanda.com; a=openssh; fpr=SHA256:C8kOX2QUJCMqnCX+KEeoqRAjLo9L+ELOSH2NSAJHqGA For HAVE_RCU_TRAMPOLINE_READERS the ftrace trampoline must be a Tasks Trace RCU reader while it calls out, since that is what synchronize_rcu_tasks() will wait for before ftrace_shutdown() frees an ftrace_ops (or, with CALL_OPS, lets its owner free it) under a task preempted in the callback. Open-code rcu_read_lock_trace() and rcu_read_unlock_trace() around the call to ops->func in ftrace_caller: bump current->trc_reader_nesting via sp_el0 and, for the outermost reader, do the SRCU-fast per-CPU increment on rcu_tasks_trace_srcu_struct and stash the counter pointer in current->trc_reader_scp, as the C inlines do (including the dmb when CONFIG_TASKS_TRACE_RCU_NO_MB is not set). The per-CPU increment is an LL/SC add on this CPU's counter; being migrated between reading the per-CPU offset and the store-exclusive only means another CPU's counter is incremented atomically instead, which SRCU sums over anyway. x12-x16 are free at both points. arm64 has no return thunks and, with CALL_OPS, no dynamic ftrace trampolines, but ftrace_caller itself carries the ops pointer in x11 from before the reader is entered and a direct-call BPF trampoline address in x17 until the final br/ret after it is left, so mark the end of the static trampoline text and provide arch_rcu_tasks_trampoline_text() covering [ftrace_caller, ftrace_static_tramp_end). Built only under CONFIG_TASKS_RCU_TRAMPOLINE_READERS, which arm64 does not enable until a later patch. Assisted-by: LLM Signed-off-by: Josef Bacik --- arch/arm64/kernel/asm-offsets.c | 8 +++++ arch/arm64/kernel/entry-ftrace.S | 74 ++++++++++++++++++++++++++++++++++++= ++++ arch/arm64/kernel/ftrace.c | 20 +++++++++++ 3 files changed, 102 insertions(+) diff --git a/arch/arm64/kernel/asm-offsets.c b/arch/arm64/kernel/asm-offset= s.c index 9c853ed3ceab..f6a8fb1f9b43 100644 --- a/arch/arm64/kernel/asm-offsets.c +++ b/arch/arm64/kernel/asm-offsets.c @@ -10,6 +10,7 @@ =20 #include #include +#include #include #include #include @@ -39,6 +40,13 @@ int main(void) DEFINE(TSK_STACK, offsetof(struct task_struct, stack)); #ifdef CONFIG_STACKPROTECTOR DEFINE(TSK_STACK_CANARY, offsetof(struct task_struct, stack_canary)); +#endif +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS + DEFINE(TSK_TRC_READER_NESTING, offsetof(struct task_struct, trc_reader_n= esting)); + DEFINE(TSK_TRC_READER_SCP, offsetof(struct task_struct, trc_reader_scp)); + DEFINE(SRCU_SRCU_CTRP, offsetof(struct srcu_struct, srcu_ctrp)); + DEFINE(SRCU_CTR_SRCU_LOCKS, offsetof(struct srcu_ctr, srcu_locks)); + DEFINE(SRCU_CTR_SRCU_UNLOCKS, offsetof(struct srcu_ctr, srcu_unlocks)); #endif BLANK(); DEFINE(THREAD_CPU_CONTEXT, offsetof(struct task_struct, thread.cpu_conte= xt)); diff --git a/arch/arm64/kernel/entry-ftrace.S b/arch/arm64/kernel/entry-ftr= ace.S index 025140caafe7..fc2805eb9e15 100644 --- a/arch/arm64/kernel/entry-ftrace.S +++ b/arch/arm64/kernel/entry-ftrace.S @@ -14,6 +14,72 @@ #include =20 #ifdef CONFIG_DYNAMIC_FTRACE_WITH_ARGS +/* + * Open-coded rcu_read_lock_trace() / rcu_read_unlock_trace(), see + * include/linux/rcupdate_trace.h and CONFIG_HAVE_RCU_TRAMPOLINE_READERS. = The + * whole of ftrace_caller is treated as trampoline text by the irq-exit ch= eck + * (see arch_rcu_tasks_trampoline_text()), so these only need to bracket t= he + * call out to ops->func; everything before the lock and after the unlock, + * including the direct-call tails that carry a BPF trampoline address in = x17, + * is covered by that. + * + * The SRCU-fast per-CPU increment is done LL/SC on this CPU's counter; be= ing + * migrated between reading the per-CPU offset and the store-exclusive only + * means another CPU's counter is (atomically) incremented, which SRCU sums + * over anyway. Ordering between the nesting count and the scp stash only + * matters against interrupts on this CPU, which observe program order. + * Clobbers x12-x16 and the flags. + */ + .macro trace_rcu_srcu_inc, addr:req, tmp:req, wtmp2:req +8888: ldxr \tmp, [\addr] + add \tmp, \tmp, #1 + stxr \wtmp2, \tmp, [\addr] + cbnz \wtmp2, 8888b + .endm + + .macro trace_rcu_read_lock +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS + mrs x12, sp_el0 // current + ldr w13, [x12, #TSK_TRC_READER_NESTING] + add w14, w13, #1 + str w14, [x12, #TSK_TRC_READER_NESTING] + cbnz w13, .Ltrl_nested\@ // interrupted a reader: done + ldr_l x13, rcu_tasks_trace_srcu_struct + SRCU_SRCU_CTRP + str x13, [x12, #TSK_TRC_READER_SCP] + get_this_cpu_offset x14 + add x14, x14, x13 + add x14, x14, #SRCU_CTR_SRCU_LOCKS + trace_rcu_srcu_inc x14, x15, w16 +#ifndef CONFIG_TASKS_TRACE_RCU_NO_MB + dmb ish +#endif +.Ltrl_nested\@: +#endif + .endm + + .macro trace_rcu_read_unlock +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS + mrs x12, sp_el0 // current + ldr w13, [x12, #TSK_TRC_READER_NESTING] + subs w13, w13, #1 + b.ne .Ltru_nested\@ + /* Outermost: pick up scp before an interrupt can see nesting =3D=3D 0. */ + ldr x14, [x12, #TSK_TRC_READER_SCP] + str wzr, [x12, #TSK_TRC_READER_NESTING] +#ifndef CONFIG_TASKS_TRACE_RCU_NO_MB + dmb ish +#endif + get_this_cpu_offset x15 + add x14, x14, x15 + add x14, x14, #SRCU_CTR_SRCU_UNLOCKS + trace_rcu_srcu_inc x14, x15, w16 + b .Ltru_done\@ +.Ltru_nested\@: + str w13, [x12, #TSK_TRC_READER_NESTING] +.Ltru_done\@: +#endif + .endm + /* * Due to -fpatchable-function-entry=3D2, the compiler has placed two NOPs= before * the regular function prologue. For an enabled callsite, ftrace_init_nop= () and @@ -94,6 +160,8 @@ SYM_CODE_START(ftrace_caller) stp x29, x30, [sp, #FREGS_SIZE] add x29, sp, #FREGS_SIZE =20 + trace_rcu_read_lock + /* Prepare arguments for the tracer func */ sub x0, x30, #AARCH64_INSN_SIZE // ip (callsite's BL insn) mov x1, x9 // parent_ip (callsite's LR) @@ -111,6 +179,8 @@ SYM_INNER_LABEL(ftrace_call, SYM_L_GLOBAL) bl ftrace_stub // func(ip, parent_ip, op, regs) #endif =20 + trace_rcu_read_unlock + /* * At the callsite x0-x8 and x19-x30 were live. Any C code will have prese= rved * x19-x29 per the AAPCS, and we created frame records upon entry, so we n= eed @@ -178,6 +248,10 @@ SYM_CODE_START(ftrace_stub_direct_tramp) SYM_CODE_END(ftrace_stub_direct_tramp) #endif /* CONFIG_DYNAMIC_FTRACE_WITH_DIRECT_CALLS */ =20 +/* End of [ftrace_caller, ...) for arch_rcu_tasks_trampoline_text(). */ +SYM_CODE_START(ftrace_static_tramp_end) +SYM_CODE_END(ftrace_static_tramp_end) + #else /* CONFIG_DYNAMIC_FTRACE_WITH_ARGS */ =20 /* diff --git a/arch/arm64/kernel/ftrace.c b/arch/arm64/kernel/ftrace.c index e1a3c0b3a051..5f4193f15cd9 100644 --- a/arch/arm64/kernel/ftrace.c +++ b/arch/arm64/kernel/ftrace.c @@ -17,6 +17,26 @@ #include #include =20 +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS +extern void ftrace_static_tramp_end(void); + +/* The SRCU-fast increments in entry-ftrace.S are the this_cpu_inc() form.= */ +static_assert(!IS_ENABLED(CONFIG_NEED_SRCU_NMI_SAFE)); + +/* + * See rcu_tasks_trampoline_text(). ftrace_caller and ftrace_stub_direct_= tramp + * are core kernel text but must be treated as trampolines: a task interru= pted + * in them outside the Tasks Trace reader may be carrying an ops pointer (= x11) + * or a direct-call BPF trampoline address (x17) whose lifetime is guarded= only + * by Tasks RCU. + */ +bool arch_rcu_tasks_trampoline_text(unsigned long ip) +{ + return ip >=3D (unsigned long)ftrace_caller && + ip < (unsigned long)ftrace_static_tramp_end; +} +#endif + #ifdef CONFIG_DYNAMIC_FTRACE_WITH_ARGS struct fregs_offset { const char *name; --=20 2.55.0 From nobody Thu Sep 24 19:04:36 2026 Received: from mail-vs2-f12.google.com (mail-vs2-f12.google.com [74.125.227.12]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 99F4C4921B1 for ; Tue, 15 Sep 2026 13:18:08 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.227.12 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789478292; cv=none; b=ufwrEwQ98qQ/rZNLBrZZHwbW5L4FJ6gvXTHnQkdtMQ7oIJtHfl0FBjq+sxq1yViiXpvZYqru+JGWDczx5QqtWa/HSdiVpLJsFWy/Qh1hr+3EWZs+sE7FoAZRtpLGPDVouLOFmhLGp3S1FUeM1cgx2lV4HVn/7adBii8sKY3XZcw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789478292; c=relaxed/simple; bh=I7rYeQZ4spsxQVfaE2Sayb+dFXGxTZqIS2lZCUAJvtU=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=pR5j2m6Q82LsIfax67vD4dYCJWz4Ri7VBcfKfcJs72wLdHE+15HrQC0WkTk1ty/vu425uOqH2hP0lAcAys2BHl5rFoLXoRjX6KUhKGsCrBt0XVW9paDbP1P39n6A+nejffp9s0+9V+ccdt9/eUyLJnwonp52MkN4wBz0qKRoD3s= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com; spf=pass smtp.mailfrom=toxicpanda.com; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b=ANeugNDz; arc=none smtp.client-ip=74.125.227.12 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b="ANeugNDz" Received: by mail-vs2-f12.google.com with SMTP id 71dfb90a1353d-5c67e512ee6so1017007e0c.1 for ; Tue, 15 Sep 2026 06:18:08 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=toxicpanda.com; s=google; t=1789478286; x=1790083086; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=Nk/uI/WquaXkRgHHd/dzhScVzrCkrWpp6IlwsBXJcUg=; b=ANeugNDzrXwRW1CDn4Wjk7U0DtcndRgd815Ok7tCSBRR4KPkAQN0FhlElQ8rq3OVv9 9Z9E9DDsT9dKlWvO78W3CImxRiZvRJv50irBt8O5uwabOwaLepptDRqqanlS594VWj9a /n4w79jr3Q5ZfFrjMJU819OnKMPXTITd43qNAIm6Vkwdf6sGE3HwIXAuJ9J2Ee/HODmJ 13B+JhW42r3cFV2paJ/SC+zbDkxrmunvz0lrIqz8Tt5GGkNo9Og9a2GtenFstOYn3aNB 221AoK/OU2VboSg3vg32S6JzZA8XdWyC2d9m3YRtLJd7bJnEMvGYNzuuSUSwoVgiCM/Q CXOg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789478286; x=1790083086; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=Nk/uI/WquaXkRgHHd/dzhScVzrCkrWpp6IlwsBXJcUg=; b=xY7cdgnIVv8xumaj2FrEkTuLYrLhvFkdG9Wq6Y2rTGv4N/iokSZb2NF/T7+8YadwOv gBmJdB7EhRibZ+xTvNtkOAys4yIg8318eCbvhD/eDjuSFxAuOGQyRX1HQaoj9Ru1b5cE 61FDCbQTO8jtnnWt4tawQr/m3RqIVCSAY41VJGHEziSU+RgCMo/A56vRQEDAcHXH49DZ hXHQBP+dTLrWaA2CptspPAZ8N74tLEzeD7zDh0GzbuJQe6Ab3FB2yhx5OXIU9YdoJ7d6 OkvdSm+W/dt0cRfEXSf+vAAEr7iJQlGTxfWvqFkhn7P7jJOb6CeC1Vc1053rtW6P3edJ OIrA== X-Forwarded-Encrypted: i=1; AKwUvBy+ve0YZZWYsnYQff9ebdYN7URz1ZC5zdVcgPWZvrnYwUlm0Oan6l+Ce4gNezdxySra6Xj+RpN9+mta9LI=@vger.kernel.org X-Gm-Message-State: AFuF++l5MF56Tz9G5pKRGd+7FEaP9oC2kwSBsAuR6+nwi7G7uDyWkcDt U/8H3l1zylrb2xR52BHc/tKUOtCNyd9Ri+ttMaq8kYwxui6nV+3eFGsGjr2LAp6CSeE= X-Gm-Gg: AYBFou1lVtSxIo8wYAo9txRwNMeKIhSNBUG2/7joJ3n92KdgrPFUKIsZjJkicr91cYe rYkcpIxXAm5dUu0DT2ElHq97/73Fhlyu+m6KYa3f0T/SynxW7EDCC5GvVFrgPzgnez1bsG24eU6 2cP+IE7uvmiE8OPV5BZqSP7An03Vxhp/qUpWU5oju+6uN4O76mMZWPqt9Ww1BhIIQrsJDTXypKi 2pr027ZSUeIaDE9ati4rgMbUcv1rPpcPQQLGF0MbwKnrI6S6/7aEQ3u9DAUeJTYxpc7ZSXkW4yq stGOitpxYsEt5qR1Vrc4Fagfdqo4FP7cZEJTGV769YbYtx8jc74PjPwI+Zdru9eFwhobPdAnzlk 5kstGJx/j7Pv6ie7rnBkVi35cFQxF3DuShUdEYgCFlSQA9rbZctKYCqenvZEpP1IEoJsoTDEBPw 4uvWwtmtxggPQDU6mii/1kp6mGdXtMo4J2BbbQHZt2FzCPu1wGEMz/apXJ4GwyAQWAO+0k2M+pr rifOCX/tD0tU46LCYEke7pg1lttdjK4LrMXHm7/6DrQbz4khYHncc7O X-Received: by 2002:a05:6122:547:b0:5c8:2ad8:1484 with SMTP id 71dfb90a1353d-5c981ad6d23mr8006048e0c.1.1789478286304; Tue, 15 Sep 2026 06:18:06 -0700 (PDT) Received: from toxicpanda.com (ec2-34-228-114-98.compute-1.amazonaws.com. [34.228.114.98]) by smtp.gmail.com with ESMTPSA id 6a1803df08f44-9120f4963fcsm120754396d6.33.2026.09.15.06.18.05 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 15 Sep 2026 06:18:05 -0700 (PDT) From: Josef Bacik Date: Tue, 15 Sep 2026 13:17:37 +0000 Subject: [PATCH RFC v3 10/13] samples: ftrace: Make the direct-call trampolines Tasks Trace readers Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260915-b4-rcu-tasks-preempt-qs-v3-10-0ad30c4c5ee7@toxicpanda.com> References: <20260915-b4-rcu-tasks-preempt-qs-v3-0-0ad30c4c5ee7@toxicpanda.com> In-Reply-To: <20260915-b4-rcu-tasks-preempt-qs-v3-0-0ad30c4c5ee7@toxicpanda.com> To: "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Boqun Feng , Thomas Gleixner , Peter Zijlstra , Steven Rostedt , Masami Hiramatsu , Mark Rutland , Jiri Olsa , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , x86@kernel.org, Catalin Marinas , Will Deacon , Puranjay Mohan , Xu Kuohai , "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Boqun Feng , Thomas Gleixner , Peter Zijlstra , Steven Rostedt , Masami Hiramatsu , Mark Rutland , Jiri Olsa , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , x86@kernel.org, Catalin Marinas , Will Deacon , Puranjay Mohan , Xu Kuohai Cc: Andy Lutomirski , Josh Triplett , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Juergen Gross , Luis Chamberlain , Ihor Solodrai , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, xen-devel@lists.xenproject.org, Andy Lutomirski , Josh Triplett , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Juergen Gross , Luis Chamberlain , Ihor Solodrai , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, xen-devel@lists.xenproject.org, Josef Bacik X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openssh-sha256; t=1789478262; l=13193; i=josef@toxicpanda.com; h=from:subject:message-id; bh=I7rYeQZ4spsxQVfaE2Sayb+dFXGxTZqIS2lZCUAJvtU=; b=U1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgUBr36M/n0nWN0DNbnxwzIiCZez6MG JiruuNaSCI/zXsAAAAGcGF0YXR0AAAAAAAAAAZzaGE1MTIAAABTAAAAC3NzaC1lZDI1NTE5AAAA QPOtf94yWU2d+WH5tRovPgYME0ctrPN2IFl8MP+VQNtbbXuFRpRyfsP4CCmH0G4Kr6G14TpI5yJ WZ3+go910yQc= X-Developer-Key: i=josef@toxicpanda.com; a=openssh; fpr=SHA256:C8kOX2QUJCMqnCX+KEeoqRAjLo9L+ELOSH2NSAJHqGA The sample direct trampolines are exactly the kind of out-of-line register_ftrace_direct() user whose lifetime depends on Tasks RCU waiting for a task inside them: nothing else stops rmmod while a task is preempted in my_direct_func(). On HAVE_RCU_TRAMPOLINE_READERS architectures that wait only covers Tasks Trace RCU readers, so give the samples a small shared header with rcu_read_lock_trace() and rcu_read_unlock_trace() open-coded as instruction strings for x86-64 and arm64 -- the same sequences as ftrace_64.S and entry-ftrace.S, using caller-saved non-argument scratch registers -- and bracket every call-out with them. The instructions outside the bracket are module text, covered by ftrace_direct_mark_module(). Other architectures get empty definitions. Assisted-by: LLM Signed-off-by: Josef Bacik --- samples/ftrace/ftrace-direct-modify.c | 9 ++ samples/ftrace/ftrace-direct-multi-modify.c | 9 ++ samples/ftrace/ftrace-direct-multi.c | 5 ++ samples/ftrace/ftrace-direct-too.c | 5 ++ samples/ftrace/ftrace-direct.c | 5 ++ samples/ftrace/ftrace-direct.h | 126 ++++++++++++++++++++++++= ++++ 6 files changed, 159 insertions(+) diff --git a/samples/ftrace/ftrace-direct-modify.c b/samples/ftrace/ftrace-= direct-modify.c index 164d9dd6fd92..937c8d8c2a1b 100644 --- a/samples/ftrace/ftrace-direct-modify.c +++ b/samples/ftrace/ftrace-direct-modify.c @@ -2,6 +2,7 @@ #include #include #include +#include "ftrace-direct.h" #if !defined(CONFIG_ARM64) && !defined(CONFIG_PPC32) #include #endif @@ -73,7 +74,9 @@ asm ( " pushq %rbp\n" " movq %rsp, %rbp\n" CALL_DEPTH_ACCOUNT + TRACE_RCU_READ_LOCK " call my_direct_func1\n" + TRACE_RCU_READ_UNLOCK " leave\n" ASM_RET " .size my_tramp1, .-my_tramp1\n" @@ -85,7 +88,9 @@ asm ( " pushq %rbp\n" " movq %rsp, %rbp\n" CALL_DEPTH_ACCOUNT + TRACE_RCU_READ_LOCK " call my_direct_func2\n" + TRACE_RCU_READ_UNLOCK " leave\n" ASM_RET " .size my_tramp2, .-my_tramp2\n" @@ -141,11 +146,13 @@ asm ( " .globl my_tramp1\n" " my_tramp1:" " hint 34\n" // bti c + TRACE_RCU_READ_LOCK " sub sp, sp, #16\n" " stp x9, x30, [sp]\n" " bl my_direct_func1\n" " ldp x30, x9, [sp]\n" " add sp, sp, #16\n" + TRACE_RCU_READ_UNLOCK " ret x9\n" " .size my_tramp1, .-my_tramp1\n" =20 @@ -153,11 +160,13 @@ asm ( " .globl my_tramp2\n" " my_tramp2:" " hint 34\n" // bti c + TRACE_RCU_READ_LOCK " sub sp, sp, #16\n" " stp x9, x30, [sp]\n" " bl my_direct_func2\n" " ldp x30, x9, [sp]\n" " add sp, sp, #16\n" + TRACE_RCU_READ_UNLOCK " ret x9\n" " .size my_tramp2, .-my_tramp2\n" " .popsection\n" diff --git a/samples/ftrace/ftrace-direct-multi-modify.c b/samples/ftrace/f= trace-direct-multi-modify.c index b03766c6217b..e12e5c8b83f0 100644 --- a/samples/ftrace/ftrace-direct-multi-modify.c +++ b/samples/ftrace/ftrace-direct-multi-modify.c @@ -2,6 +2,7 @@ #include #include #include +#include "ftrace-direct.h" #if !defined(CONFIG_ARM64) && !defined(CONFIG_PPC32) #include #endif @@ -77,10 +78,12 @@ asm ( " pushq %rbp\n" " movq %rsp, %rbp\n" CALL_DEPTH_ACCOUNT + TRACE_RCU_READ_LOCK " pushq %rdi\n" " movq 8(%rbp), %rdi\n" " call my_direct_func1\n" " popq %rdi\n" + TRACE_RCU_READ_UNLOCK " leave\n" ASM_RET " .size my_tramp1, .-my_tramp1\n" @@ -92,10 +95,12 @@ asm ( " pushq %rbp\n" " movq %rsp, %rbp\n" CALL_DEPTH_ACCOUNT + TRACE_RCU_READ_LOCK " pushq %rdi\n" " movq 8(%rbp), %rdi\n" " call my_direct_func2\n" " popq %rdi\n" + TRACE_RCU_READ_UNLOCK " leave\n" ASM_RET " .size my_tramp2, .-my_tramp2\n" @@ -154,6 +159,7 @@ asm ( " .globl my_tramp1\n" " my_tramp1:" " hint 34\n" // bti c + TRACE_RCU_READ_LOCK " sub sp, sp, #32\n" " stp x9, x30, [sp]\n" " str x0, [sp, #16]\n" @@ -162,6 +168,7 @@ asm ( " ldp x30, x9, [sp]\n" " ldr x0, [sp, #16]\n" " add sp, sp, #32\n" + TRACE_RCU_READ_UNLOCK " ret x9\n" " .size my_tramp1, .-my_tramp1\n" =20 @@ -169,6 +176,7 @@ asm ( " .globl my_tramp2\n" " my_tramp2:" " hint 34\n" // bti c + TRACE_RCU_READ_LOCK " sub sp, sp, #32\n" " stp x9, x30, [sp]\n" " str x0, [sp, #16]\n" @@ -177,6 +185,7 @@ asm ( " ldp x30, x9, [sp]\n" " ldr x0, [sp, #16]\n" " add sp, sp, #32\n" + TRACE_RCU_READ_UNLOCK " ret x9\n" " .size my_tramp2, .-my_tramp2\n" " .popsection\n" diff --git a/samples/ftrace/ftrace-direct-multi.c b/samples/ftrace/ftrace-d= irect-multi.c index 3fe6ddaf0b69..a970464ed378 100644 --- a/samples/ftrace/ftrace-direct-multi.c +++ b/samples/ftrace/ftrace-direct-multi.c @@ -3,6 +3,7 @@ =20 #include /* for handle_mm_fault() */ #include +#include "ftrace-direct.h" #include #if !defined(CONFIG_ARM64) && !defined(CONFIG_PPC32) #include @@ -56,10 +57,12 @@ asm ( " pushq %rbp\n" " movq %rsp, %rbp\n" CALL_DEPTH_ACCOUNT + TRACE_RCU_READ_LOCK " pushq %rdi\n" " movq 8(%rbp), %rdi\n" " call my_direct_func\n" " popq %rdi\n" + TRACE_RCU_READ_UNLOCK " leave\n" ASM_RET " .size my_tramp, .-my_tramp\n" @@ -101,6 +104,7 @@ asm ( " .globl my_tramp\n" " my_tramp:" " hint 34\n" // bti c + TRACE_RCU_READ_LOCK " sub sp, sp, #32\n" " stp x9, x30, [sp]\n" " str x0, [sp, #16]\n" @@ -109,6 +113,7 @@ asm ( " ldp x30, x9, [sp]\n" " ldr x0, [sp, #16]\n" " add sp, sp, #32\n" + TRACE_RCU_READ_UNLOCK " ret x9\n" " .size my_tramp, .-my_tramp\n" " .popsection\n" diff --git a/samples/ftrace/ftrace-direct-too.c b/samples/ftrace/ftrace-dir= ect-too.c index bf2411aa6fd7..abc098c2ab7a 100644 --- a/samples/ftrace/ftrace-direct-too.c +++ b/samples/ftrace/ftrace-direct-too.c @@ -3,6 +3,7 @@ =20 #include /* for handle_mm_fault() */ #include +#include "ftrace-direct.h" #if !defined(CONFIG_ARM64) && !defined(CONFIG_PPC32) #include #endif @@ -61,6 +62,7 @@ asm ( " pushq %rbp\n" " movq %rsp, %rbp\n" CALL_DEPTH_ACCOUNT + TRACE_RCU_READ_LOCK " pushq %rdi\n" " pushq %rsi\n" " pushq %rdx\n" @@ -70,6 +72,7 @@ asm ( " popq %rdx\n" " popq %rsi\n" " popq %rdi\n" + TRACE_RCU_READ_UNLOCK " leave\n" ASM_RET " .size my_tramp, .-my_tramp\n" @@ -110,6 +113,7 @@ asm ( " .globl my_tramp\n" " my_tramp:" " hint 34\n" // bti c + TRACE_RCU_READ_LOCK " sub sp, sp, #48\n" " stp x9, x30, [sp]\n" " stp x0, x1, [sp, #16]\n" @@ -119,6 +123,7 @@ asm ( " ldp x0, x1, [sp, #16]\n" " ldp x2, x3, [sp, #32]\n" " add sp, sp, #48\n" + TRACE_RCU_READ_UNLOCK " ret x9\n" " .size my_tramp, .-my_tramp\n" " .popsection\n" diff --git a/samples/ftrace/ftrace-direct.c b/samples/ftrace/ftrace-direct.c index 5368c8c39cbb..99b65ad2fccc 100644 --- a/samples/ftrace/ftrace-direct.c +++ b/samples/ftrace/ftrace-direct.c @@ -3,6 +3,7 @@ =20 #include /* for wake_up_process() */ #include +#include "ftrace-direct.h" #if !defined(CONFIG_ARM64) && !defined(CONFIG_PPC32) #include #endif @@ -54,9 +55,11 @@ asm ( " pushq %rbp\n" " movq %rsp, %rbp\n" CALL_DEPTH_ACCOUNT + TRACE_RCU_READ_LOCK " pushq %rdi\n" " call my_direct_func\n" " popq %rdi\n" + TRACE_RCU_READ_UNLOCK " leave\n" ASM_RET " .size my_tramp, .-my_tramp\n" @@ -97,6 +100,7 @@ asm ( " .globl my_tramp\n" " my_tramp:" " hint 34\n" // bti c + TRACE_RCU_READ_LOCK " sub sp, sp, #32\n" " stp x9, x30, [sp]\n" " str x0, [sp, #16]\n" @@ -104,6 +108,7 @@ asm ( " ldp x30, x9, [sp]\n" " ldr x0, [sp, #16]\n" " add sp, sp, #32\n" + TRACE_RCU_READ_UNLOCK " ret x9\n" " .size my_tramp, .-my_tramp\n" " .popsection\n" diff --git a/samples/ftrace/ftrace-direct.h b/samples/ftrace/ftrace-direct.h new file mode 100644 index 000000000000..726f67048ff5 --- /dev/null +++ b/samples/ftrace/ftrace-direct.h @@ -0,0 +1,126 @@ +/* SPDX-License-Identifier: GPL-2.0-only */ +#ifndef _SAMPLES_FTRACE_DIRECT_H +#define _SAMPLES_FTRACE_DIRECT_H + +#include + +/* + * A direct-call trampoline is entered with no lock, refcount or RCU marker + * held; only Tasks RCU keeps it (and, for a module, its text) alive while= a + * task is inside it or preempted in something it called. On architectures + * that select HAVE_RCU_TRAMPOLINE_READERS, Tasks RCU only waits for such a + * task while it is a Tasks Trace RCU reader, so the trampoline must enter= one + * before calling out and leave it afterwards, exactly like the ftrace and= BPF + * trampolines do. See register_ftrace_direct(). The instructions before= the + * lock and after the unlock are covered by ftrace_direct_mark_module(). + * + * These are rcu_read_lock_trace() / rcu_read_unlock_trace() open-coded as + * instruction strings for use inside the samples' asm() trampolines, after + * the versions in arch/x86/kernel/ftrace_64.S and + * arch/arm64/kernel/entry-ftrace.S. The scratch registers are caller-sav= ed + * and not argument registers, so they are dead on entry to and exit from = an + * fentry trampoline; the flags are clobbered. + * + * The generated asm-offsets.h is only pulled in on the architectures that= need + * it here: it is not generally safe to include from C (e.g. PPC32's TASK_= SIZE + * and arm64's TRAMP_VALIAS clash with the C definitions). + */ +#if defined(CONFIG_TASKS_RCU_TRAMPOLINE_READERS) && defined(CONFIG_X86_64) + +#include + +#ifndef CONFIG_TASKS_TRACE_RCU_NO_MB +#define TRACE_RCU_MB " lock addl $0, -4(%rsp)\n" +#else +#define TRACE_RCU_MB +#endif + +#define TRACE_RCU_READ_LOCK \ + " movq %gs:current_task(%rip), %r11\n" \ + " movl " __stringify(TASK_trc_reader_nesting) "(%r11), %r10d\n" \ + " incl " __stringify(TASK_trc_reader_nesting) "(%r11)\n" \ + " testl %r10d, %r10d\n" \ + " jnz 771f\n" \ + " movq rcu_tasks_trace_srcu_struct+" __stringify(SRCU_srcu_ctrp) "(%rip),= %r10\n" \ + " incq %gs:" __stringify(SRCU_CTR_srcu_locks) "(%r10)\n" \ + " movq %r10, " __stringify(TASK_trc_reader_scp) "(%r11)\n" \ + TRACE_RCU_MB \ + "771:\n" + +#define TRACE_RCU_READ_UNLOCK \ + " movq %gs:current_task(%rip), %r11\n" \ + " movl " __stringify(TASK_trc_reader_nesting) "(%r11), %r10d\n" \ + " subl $1, %r10d\n" \ + " jnz 772f\n" \ + " movq " __stringify(TASK_trc_reader_scp) "(%r11), %r10\n" \ + " movl $0, " __stringify(TASK_trc_reader_nesting) "(%r11)\n" \ + TRACE_RCU_MB \ + " incq %gs:" __stringify(SRCU_CTR_srcu_unlocks) "(%r10)\n" \ + " jmp 773f\n" \ + "772: movl %r10d, " __stringify(TASK_trc_reader_nesting) "(%r11)\n" \ + "773:\n" + +#elif defined(CONFIG_TASKS_RCU_TRAMPOLINE_READERS) && defined(CONFIG_ARM64) + +#include +#include +/* arm64's asm-offsets.h redefines TRAMP_VALIAS from . */ +#pragma push_macro("TRAMP_VALIAS") +#undef TRAMP_VALIAS +#include +#pragma pop_macro("TRAMP_VALIAS") + +#ifndef CONFIG_TASKS_TRACE_RCU_NO_MB +#define TRACE_RCU_MB " dmb ish\n" +#else +#define TRACE_RCU_MB +#endif + +#define TRACE_RCU_SRCU_CTRP "rcu_tasks_trace_srcu_struct+" __stringify(SRC= U_SRCU_CTRP) + +/* x14 =3D this CPU's offset; then atomically increment the long at x14 + = \areg */ +#define TRACE_RCU_PERCPU_INC(areg) \ + ALTERNATIVE(" mrs x14, tpidr_el1\n", " mrs x14, tpidr_el2\n", \ + ARM64_HAS_VIRT_HOST_EXTN) \ + " add x14, x14, " areg "\n" \ + "778: ldxr x15, [x14]\n" \ + " add x15, x15, #1\n" \ + " stxr w16, x15, [x14]\n" \ + " cbnz w16, 778b\n" + +#define TRACE_RCU_READ_LOCK \ + " mrs x12, sp_el0\n" \ + " ldr w13, [x12, #" __stringify(TSK_TRC_READER_NESTING) "]\n" \ + " add w14, w13, #1\n" \ + " str w14, [x12, #" __stringify(TSK_TRC_READER_NESTING) "]\n" \ + " cbnz w13, 771f\n" \ + " adrp x13, " TRACE_RCU_SRCU_CTRP "\n" \ + " ldr x13, [x13, #:lo12:" TRACE_RCU_SRCU_CTRP "]\n" \ + " str x13, [x12, #" __stringify(TSK_TRC_READER_SCP) "]\n" \ + " add x13, x13, #" __stringify(SRCU_CTR_SRCU_LOCKS) "\n" \ + TRACE_RCU_PERCPU_INC("x13") \ + TRACE_RCU_MB \ + "771:\n" + +#define TRACE_RCU_READ_UNLOCK \ + " mrs x12, sp_el0\n" \ + " ldr w13, [x12, #" __stringify(TSK_TRC_READER_NESTING) "]\n" \ + " subs w13, w13, #1\n" \ + " b.ne 772f\n" \ + " ldr x13, [x12, #" __stringify(TSK_TRC_READER_SCP) "]\n" \ + " str wzr, [x12, #" __stringify(TSK_TRC_READER_NESTING) "]\n" \ + TRACE_RCU_MB \ + " add x13, x13, #" __stringify(SRCU_CTR_SRCU_UNLOCKS) "\n" \ + TRACE_RCU_PERCPU_INC("x13") \ + " b 773f\n" \ + "772: str w13, [x12, #" __stringify(TSK_TRC_READER_NESTING) "]\n" \ + "773:\n" + +#else + +#define TRACE_RCU_READ_LOCK +#define TRACE_RCU_READ_UNLOCK + +#endif + +#endif /* _SAMPLES_FTRACE_DIRECT_H */ --=20 2.55.0 From nobody Thu Sep 24 19:04:36 2026 Received: from mail-qk2-f12.google.com (mail-qk2-f12.google.com [74.125.230.204]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6CA4D4A3865 for ; Tue, 15 Sep 2026 13:18:10 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.204 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789478293; cv=none; b=tqL6IKyTUzQFWUqPO1wIONYixnBu/iPgwcNEWQR69wu113p6wrqpXbyCR8zTPCpP+om9JDDM/MQN76BEq7X42RpnrzwScc5FRn7Os3Z0ZHXtTkuUDpl6RbNC1LgMNjU/3WuFROFQcYtjVY8m7cqtsmJh3g/EvTtLcmigCNsc5Kc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789478293; c=relaxed/simple; bh=O8YdW4YSHUYyF3GwEQN16WzJM7UkZZFfh8RzVwv7xJM=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=N+/xnYCQi/NXxnSLCokJqpXdUF+GqTBRENK6vr3mI82N8crCLb81VhjmNXWP+E/UeV3AaxAHkS5cziiqk3tkpO1A8vqoqz0lr2s3EpUis5RKfivgiWy9L7n54SQkHT3sEv7Yw0J97I0FAgTToDRhfweoqdMj0ruA9gH9vs3yAIw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com; spf=pass smtp.mailfrom=toxicpanda.com; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b=ZNRiEjkb; arc=none smtp.client-ip=74.125.230.204 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b="ZNRiEjkb" Received: by mail-qk2-f12.google.com with SMTP id af79cd13be357-93a222edc62so226650685a.0 for ; Tue, 15 Sep 2026 06:18:10 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=toxicpanda.com; s=google; t=1789478288; x=1790083088; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=MkP56mJgKixcgLXoBxrfzz0NyJo14L+hjKFrNfKE/2k=; b=ZNRiEjkbcxJzgkGQfsEmE4NquJ4MI7P5nVOndUM2EJRYj2e9cl8NvZt4ozVmun1SZa iGMGhKrQxzoLhhyYqPCsqrzDETI58u24FiNHJ4GiGgdSSnRkT9v583Vo9UJjwPEtYvt5 RqqJ5H8H+0JJBi2hxQfPZdPcJd639d0Kr78SlkxOQhBX67A0yrCOYH4LaZoMrgt1pHhw 3Or1HxFW59yV+SmxmiQBYV7sx8ur02+sgSxuxNAQzbC1tqqaQVRdOuK7m8H7IVRON2gh zcq+p9R0oeDIh8dbrHrRBZNnIn4AI9VCW3mYCjZUVoS1s7Lt7tFp9lPBdEoyVk1zHq9Q f7JQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789478288; x=1790083088; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=MkP56mJgKixcgLXoBxrfzz0NyJo14L+hjKFrNfKE/2k=; b=AVAMRrHkE1yRWY+9lIw8guXqGzXAVsg5r6qFxpiu+tI1qrf+5Dpkiyfn6eHb+lnXie PV0rhLLLJG+GPUQd9V8TJ4XRe0Cm5TVu6WMbNAfXxa+jJVqHq4+m2K0mG540dae8cKZ8 165xWURTrp5fPbX6hB0a5mKIGAtqWzHH6jCrY3WPMyUzE3YIzQxGHi9ZJz0WZrKFkfQd WZXbpulBvMAj7E5JW8w5QBaUTX32rLiT+225AA2vpVpQ6xy/QaW5rTDLD4t/pRjo5VKv 1EQohH6ZdfJCwbkYcOs0XFy96gmy1zpxK29wZd6rGuqg5SsphijucDRJAQeVUtXo0c07 Pk6Q== X-Forwarded-Encrypted: i=1; AKwUvBxMXTpgaZ3K8nmBAa3tyVdM3/dQeeBHLPcE0Zo008LeiivjTA0BfzEmbJ4pViu6Z3k2CdI681ISJAsVCF0=@vger.kernel.org X-Gm-Message-State: AFuF++kKKlde7ZnGCO/36s45Vnncj7fgCnwLI2vkBiX4uk3XxZKzJHNm L1nzY1Sh4uYMf68Vyr1jtyvD/wooWrloKuC03KKSlVfA6MZsnZYJqgSzkGwjIO+L8Q8= X-Gm-Gg: AYBFou2AvvViPimcE294OhBEN4rR7drR7n1vKs0UVtafgCh8sw7vMroVZgSlu0aKAzv l+6manDixwfXOHP0eMBkadYvLQWEXgX0vlQGjP8Y0A9MEd199nqPg5brj5MDR5XwZI/k2cjxGZr sBzGwnwLTPxA4BER7nBrXR7hnVyWy60Y5xpMoiUGils2stUJbvRqifoMUpDDVwcA+cEP/Fhl6Za BWccwWMsJSPv/DMYsmsHY03gLCc6bZu/EjIfYDAgdbLA5Sv/dCPJsWtp30ysu21B0lKBd3NTlm5 kTXvjrdAVIRnjusIZxrcVS+Jmpv5vdMLXcxSNElh3prYTBMANCQjETywrwcrL0MuElR0V1OgKUc iwaaSzqdHCDDZMHuRWrJllMv26gMHyEkp7+g0rXpchgv/M0VFndcz4AOhq8L7bnsM4axuyb5ATK tIi7594lgQ89KRnb2Cc41kl7hWuhGjFi2+KrQYTr9NxH8tLpBNe5RxMUicU+uWgUeBLnmIfzCxT 0AU039FsUpVkhicwGD4c+YT3K9RVi0FUh1Ay4rdFYuEzJLpKTpSRbCK X-Received: by 2002:a05:620a:1b87:b0:92e:76a1:96cd with SMTP id af79cd13be357-93a297e8f8cmr1159268485a.13.1789478288213; Tue, 15 Sep 2026 06:18:08 -0700 (PDT) Received: from toxicpanda.com (ec2-34-228-114-98.compute-1.amazonaws.com. [34.228.114.98]) by smtp.gmail.com with ESMTPSA id af79cd13be357-93a26921baesm489982785a.24.2026.09.15.06.18.07 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 15 Sep 2026 06:18:07 -0700 (PDT) From: Josef Bacik Date: Tue, 15 Sep 2026 13:17:38 +0000 Subject: [PATCH RFC v3 11/13] rcutorture: Make Tasks RCU readers Tasks Trace readers where required Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260915-b4-rcu-tasks-preempt-qs-v3-11-0ad30c4c5ee7@toxicpanda.com> References: <20260915-b4-rcu-tasks-preempt-qs-v3-0-0ad30c4c5ee7@toxicpanda.com> In-Reply-To: <20260915-b4-rcu-tasks-preempt-qs-v3-0-0ad30c4c5ee7@toxicpanda.com> To: "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Boqun Feng , Thomas Gleixner , Peter Zijlstra , Steven Rostedt , Masami Hiramatsu , Mark Rutland , Jiri Olsa , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , x86@kernel.org, Catalin Marinas , Will Deacon , Puranjay Mohan , Xu Kuohai , "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Boqun Feng , Thomas Gleixner , Peter Zijlstra , Steven Rostedt , Masami Hiramatsu , Mark Rutland , Jiri Olsa , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , x86@kernel.org, Catalin Marinas , Will Deacon , Puranjay Mohan , Xu Kuohai Cc: Andy Lutomirski , Josh Triplett , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Juergen Gross , Luis Chamberlain , Ihor Solodrai , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, xen-devel@lists.xenproject.org, Andy Lutomirski , Josh Triplett , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Juergen Gross , Luis Chamberlain , Ihor Solodrai , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, xen-devel@lists.xenproject.org, Josef Bacik X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openssh-sha256; t=1789478262; l=1716; i=josef@toxicpanda.com; h=from:subject:message-id; bh=O8YdW4YSHUYyF3GwEQN16WzJM7UkZZFfh8RzVwv7xJM=; b=U1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgUBr36M/n0nWN0DNbnxwzIiCZez6MG JiruuNaSCI/zXsAAAAGcGF0YXR0AAAAAAAAAAZzaGE1MTIAAABTAAAAC3NzaC1lZDI1NTE5AAAA QIccOQgqxisnPo/25JJmyx4Bx5hp1tfVVTZrdfU6SlgArQbvK/ARnj8AAoaQqqIY//mIhN1g1fK eVSGZVlcF5wA= X-Developer-Key: i=josef@toxicpanda.com; a=openssh; fpr=SHA256:C8kOX2QUJCMqnCX+KEeoqRAjLo9L+ELOSH2NSAJHqGA rcutorture's "tasks" flavor has empty readlock/readunlock hooks because a classic Tasks RCU reader is simply code that does not block. Under CONFIG_TASKS_RCU_TRAMPOLINE_READERS a preemption outside trampoline text is also a quiescent state, and the thing real readers (trampolines) do to stay protected across their call-outs is take rcu_read_lock_trace(), so have the torture readers do the same there. Otherwise a preempted torture reader would rightly be treated as quiescent and the test would report false-positive too-short grace periods. Assisted-by: LLM Signed-off-by: Josef Bacik --- kernel/rcu/rcutorture.c | 10 ++++++++++ 1 file changed, 10 insertions(+) diff --git a/kernel/rcu/rcutorture.c b/kernel/rcu/rcutorture.c index 794937e13e7c..ab870ef09af0 100644 --- a/kernel/rcu/rcutorture.c +++ b/kernel/rcu/rcutorture.c @@ -1142,13 +1142,23 @@ static struct rcu_torture_ops trivial_preempt_ops = =3D { * Definitions for RCU-tasks torture testing. */ =20 +/* + * A classic Tasks RCU reader is any stretch of kernel code that does not + * voluntarily block. With CONFIG_TASKS_RCU_TRAMPOLINE_READERS a preempti= on + * outside trampoline text also ends it, and what a trampoline does to stay + * protected across its call-out is take a Tasks Trace reader, so model th= at. + */ static int tasks_torture_read_lock(void) { + if (IS_ENABLED(CONFIG_TASKS_RCU_TRAMPOLINE_READERS)) + rcu_read_lock_trace(); return 0; } =20 static void tasks_torture_read_unlock(int idx) { + if (IS_ENABLED(CONFIG_TASKS_RCU_TRAMPOLINE_READERS)) + rcu_read_unlock_trace(); } =20 static void rcu_tasks_torture_deferred_free(struct rcu_torture *p) --=20 2.55.0 From nobody Thu Sep 24 19:04:36 2026 Received: from mail-qk2-f12.google.com (mail-qk2-f12.google.com [74.125.230.204]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BBA464A3F0F for ; Tue, 15 Sep 2026 13:18:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.204 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789478295; cv=none; b=T7CNeIVVp1oHUVxgHwsPlg17DlFsIY9mXGozdPbTTZuWz0YO3hFefGy/YJsfatBs6CyIh8s7Ae5OtJD86vrhEPS+Pzcl8cun5jJzkwtv0MYctvKZHyJrTVcSpRNkwQAZxQ2lZ7TLBteLNuEvXtHAR0rH7PjUQtcjN6i1XztJKvg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789478295; c=relaxed/simple; bh=a8x0095kiaPfODWOL6BE+0O7E5TXMNLRrf+m1qUM63A=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=gTlQyaiEVldp9LayZfnExhOLcngGFe+K3QVVbvNovW1r0PKJ3kOxNEP93sqMrQDR3/dG6riiovu/K4FpA1S+3a6nOqCcVhTK4Jul+oCZBVjC/dTwL4Lol/1KL2FG6d7UgWYrlRWWmflQbyJr/VvBLUEkhk2PiJH7KKnNHLJIGLI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com; spf=pass smtp.mailfrom=toxicpanda.com; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b=GcHM5kne; arc=none smtp.client-ip=74.125.230.204 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b="GcHM5kne" Received: by mail-qk2-f12.google.com with SMTP id d75a77b69052e-52fb76ef19aso34891111cf.2 for ; Tue, 15 Sep 2026 06:18:12 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=toxicpanda.com; s=google; t=1789478290; x=1790083090; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=4AdtdE/dNeJhNnC7k2tL4hzZr80UNkNMTy9M8Z5zSRM=; b=GcHM5knei46u64fvgeHMLZZx4CC/aABYu/dUG4fLhVq+Hn1pHHoaqKChILqst+JnJN QFv+pN4eoXxbWbi+yd75J4Fc4Lp00WxoeewVk7ba5YRVhHCVSSsQZhSCbyJvltk0ruC9 bzdeWyGSqWxVM3VY6BwDOQLyNbqZ7tBsAs3bx5RjawGGHkUx+Zo+EwHsgKa1yWCWiULQ sATFcDXQYBW0UHT+wz0aTjzmxw6UOhYRywOrto5iu92XbxS67ZrLuHb+vnaa4rmeqkIv gNOe4L7qpas5zIicX/Ty0gHI2mMIGfHFVK8OqCjUosGmtLdzfKPM76jbdHpQu0dDtCgk NpmA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789478290; x=1790083090; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=4AdtdE/dNeJhNnC7k2tL4hzZr80UNkNMTy9M8Z5zSRM=; b=X1pUmTlFsC3FEFKfn870xvPALH37eevEVeLElJ2aiJ/Zj1awZovetvDZ2AkzRl5950 DnuSbHdSg6u3NNofvfdr9Bp8D4Wuuq+H3pEILOXAlGgVRoJ6gq/OfofqApUDWfmEvV75 fZLFlcsvBQyQUB6m6/LhuoKC+RdsUi8hV0bOeO7KeugJZrFe/OV8uTF3LIQmbR0URq3h 1tVXa/pWWoaxGFgk78V0YLzt+sOgpuJ/wjJqmGs4Bxg6rKCVq4EadExP5VW//gtZq72/ 788gUWyCFb5tj6A/NWFMK5YHMuVW+7gJUhNKIIT4sU9yAr+vsDir0QW9t3ZM/2vIEbaF yJLg== X-Forwarded-Encrypted: i=1; AKwUvByYZR/TbRCSp38DFtF7JN6AG3zt9DKO6e5xN+YCsXD7qPfbHnyvbFFa+Md+mbXNJUcz2yC5hssNtz46EWc=@vger.kernel.org X-Gm-Message-State: AFuF++k//4U789vadBGht8cHopTDE5S+SkwODNduU/bg1Heu650Zrywu 6o2Xynut9MSLuj/mgn6L0n3xF9QDo46rnBUTeR/DOoeiFuFVuMr534EFxvDprR43hag= X-Gm-Gg: AYBFou18PXpl+yp2aAjYnx4laO/cl0Qg6wxztl00nh8lSUEZFkBnLt8AgGiXSsiQTh3 aw5xiuWlba6Qcy6td1qmK1v5FLQ5rGtLJnaZTRdz4Gq6fxCW45lpE6hFNqwOBQtYZgidnl9lBsk uc/vCIjqAEC7E8cqfgyWkVcfVkJ248OgYi6SvcSJu7Rj3Mh6CK+u/7961nHF+bjLslsgPo3QGoi iir4xxte0qoy74Wqruoioi6j7PSl9nIf+62C6txJ2NcSPvm8siBl7bvLqsfSkUJorLhWWLgcI9E 0fI9mzj6NlXc1NoYS0tkXY4nYj5uQtNh5Vy9rSUc2u/5FC1AVcZh84BY8e3TmwD2TAnscUin+a4 bS8fyFLzI+IF/AwnlOC3GLa2VWCNiEysJFMwu2kRuDO4i2TUZ8YlCOYbmkRi6qtuF25w6v5aF/g bZbYXYRNI7PTi2B+lJLy6UfoPWm9nk2AkMCCOdnWBeNiGSSioLusrEr4hL28I/ZuN4ggmh4qu4+ VGwZDcgoBifM4NCHbAmiOeBoGpQLzbOF4Zrmv8cBBZUWevIT9mlEFMs X-Received: by 2002:ac8:5705:0:b0:530:e6b9:f5d9 with SMTP id d75a77b69052e-5310cf60dcamr104610221cf.20.1789478290432; Tue, 15 Sep 2026 06:18:10 -0700 (PDT) Received: from toxicpanda.com (ec2-34-228-114-98.compute-1.amazonaws.com. [34.228.114.98]) by smtp.gmail.com with ESMTPSA id d75a77b69052e-53104165737sm56777241cf.27.2026.09.15.06.18.09 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 15 Sep 2026 06:18:09 -0700 (PDT) From: Josef Bacik Date: Tue, 15 Sep 2026 13:17:39 +0000 Subject: [PATCH RFC v3 12/13] rcu-tasks-trace: Assert no reader is held on return to userspace Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260915-b4-rcu-tasks-preempt-qs-v3-12-0ad30c4c5ee7@toxicpanda.com> References: <20260915-b4-rcu-tasks-preempt-qs-v3-0-0ad30c4c5ee7@toxicpanda.com> In-Reply-To: <20260915-b4-rcu-tasks-preempt-qs-v3-0-0ad30c4c5ee7@toxicpanda.com> To: "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Boqun Feng , Thomas Gleixner , Peter Zijlstra , Steven Rostedt , Masami Hiramatsu , Mark Rutland , Jiri Olsa , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , x86@kernel.org, Catalin Marinas , Will Deacon , Puranjay Mohan , Xu Kuohai , "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Boqun Feng , Thomas Gleixner , Peter Zijlstra , Steven Rostedt , Masami Hiramatsu , Mark Rutland , Jiri Olsa , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , x86@kernel.org, Catalin Marinas , Will Deacon , Puranjay Mohan , Xu Kuohai Cc: Andy Lutomirski , Josh Triplett , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Juergen Gross , Luis Chamberlain , Ihor Solodrai , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, xen-devel@lists.xenproject.org, Andy Lutomirski , Josh Triplett , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Juergen Gross , Luis Chamberlain , Ihor Solodrai , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, xen-devel@lists.xenproject.org, Josef Bacik X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openssh-sha256; t=1789478262; l=2589; i=josef@toxicpanda.com; h=from:subject:message-id; bh=a8x0095kiaPfODWOL6BE+0O7E5TXMNLRrf+m1qUM63A=; b=U1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgUBr36M/n0nWN0DNbnxwzIiCZez6MG JiruuNaSCI/zXsAAAAGcGF0YXR0AAAAAAAAAAZzaGE1MTIAAABTAAAAC3NzaC1lZDI1NTE5AAAA QNQ+h4Rgq+uBTWY42BKURCuzv1XjZ1QDbD5nWJycO3HnjSRSDGwSoVc9aADl5pV0kBRcRtS0O1v XzhUhrtoXcQU= X-Developer-Key: i=josef@toxicpanda.com; a=openssh; fpr=SHA256:C8kOX2QUJCMqnCX+KEeoqRAjLo9L+ELOSH2NSAJHqGA With trampolines and their glue now taking rcu_read_lock_trace() from assembly and from several C paths, an unbalanced reader would silently turn every later Tasks Trace grace period on that task into a stall. No task can legitimately reach userspace with current->trc_reader_nesting non-zero, so under CONFIG_PROVE_RCU check it in the generic entry code's return-to-user validation, next to the existing kmap and lockdep assertions. Compiles away otherwise. Assisted-by: LLM Signed-off-by: Josef Bacik --- include/linux/irq-entry-common.h | 2 ++ include/linux/rcupdate_trace.h | 7 +++++++ 2 files changed, 9 insertions(+) diff --git a/include/linux/irq-entry-common.h b/include/linux/irq-entry-com= mon.h index b811b469b0a7..58086cc1140c 100644 --- a/include/linux/irq-entry-common.h +++ b/include/linux/irq-entry-common.h @@ -5,6 +5,7 @@ #include #include #include +#include #include #include #include @@ -214,6 +215,7 @@ static __always_inline void __exit_to_user_mode_validat= e(void) { /* Ensure that kernel state is sane for a return to userspace */ kmap_assert_nomap(); + rcu_tasks_trace_assert_idle(); lockdep_assert_irqs_disabled(); lockdep_sys_exit(); } diff --git a/include/linux/rcupdate_trace.h b/include/linux/rcupdate_trace.h index 4035054309d7..f5a51de4a4ef 100644 --- a/include/linux/rcupdate_trace.h +++ b/include/linux/rcupdate_trace.h @@ -209,6 +209,12 @@ unsigned long rcu_tasks_trace_batches_completed(void); // Placeholders to enable stepwise transition. void __init rcu_tasks_trace_suppress_unused(void); =20 +/* A task must never reach userspace inside an rcu_read_lock_trace() reade= r. */ +static inline void rcu_tasks_trace_assert_idle(void) +{ + WARN_ON_ONCE(IS_ENABLED(CONFIG_PROVE_RCU) && READ_ONCE(current->trc_reade= r_nesting)); +} + #else static inline unsigned long rcu_tasks_trace_batches_completed(void) { retu= rn 0; } /* @@ -218,6 +224,7 @@ static inline unsigned long rcu_tasks_trace_batches_com= pleted(void) { return 0; static inline void call_rcu_tasks_trace(struct rcu_head *rhp, rcu_callback= _t func) { BUG(); } static inline void rcu_read_lock_trace(void) { BUG(); } static inline void rcu_read_unlock_trace(void) { BUG(); } +static inline void rcu_tasks_trace_assert_idle(void) { } #endif /* #ifdef CONFIG_TASKS_TRACE_RCU */ =20 DEFINE_LOCK_GUARD_0(rcu_tasks_trace, --=20 2.55.0 From nobody Thu Sep 24 19:04:36 2026 Received: from mail-qk2-f13.google.com (mail-qk2-f13.google.com [74.125.230.205]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 89BF74A482F for ; Tue, 15 Sep 2026 13:18:14 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.205 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789478297; cv=none; b=oNASqiHpGTZRted2TA8/HyPG8+ggJVpATM/EmROilS3PPEob12k4BFar+sxt8cPSaKthojiT2fSFSTbd0A6NO0p5/R5amJrCdb8+9OcxgB+w9/Ov2pIkRAgPDyopOvlQJx6UiO3KLTxrgz14S6XEtxnSsJY8UUl2/M/yaCrgrxs= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789478297; c=relaxed/simple; bh=R9t47tbzDpBlKabnqSyLf9h+YtCy+mgQn2kYE1SNKlg=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=TlOjE1T7X4c9N7GJlGFyiwhmjXDypaOqI/ZnVfI++c2URndZrFvfCfGoppwl06Z6/qYbDvfGfxEIX39AlrkCd3hTZ8fY3H5cL1CmzKerX8w6HBaya8bv7QX+5ayOyMIZsID8yKyWEB+2PVXG9ZViXiNirjf4ofnJCtngdFdYmLM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com; spf=pass smtp.mailfrom=toxicpanda.com; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b=kFiJqUSa; arc=none smtp.client-ip=74.125.230.205 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b="kFiJqUSa" Received: by mail-qk2-f13.google.com with SMTP id d75a77b69052e-52fb767cc7fso1726301cf.0 for ; Tue, 15 Sep 2026 06:18:14 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=toxicpanda.com; s=google; t=1789478293; x=1790083093; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=dyunHw12Wk9XaltfwN7bGL3PCRgKnyr3LfmugMg8Gj0=; b=kFiJqUSaNikHJYMEPzTggnxBXpK7rMLLDNVTmH1cC8JJgb50CIn5TUSeyjhJ1sVxbQ HlljZCvMm+WAQnYp/kmz7GlfD6ldCHEHR8bR4cXdw2m68qnh9Dh+am+IbmjhFXNGBy45 /z904sk2ET0G3cHwElQ9qgld2s0lHe3024NR1ChhyMYhEP5iDaUAgmPYO9RonbnyvK3Z FEN7dy+GKhQyPFRQphHubVi0XlNm9KSUBpf654zm5kdWKHuRvC1El5gA66hfbYNf9OFm 4dC1d7wCKm2msW5gsFkYkt5npCQ1lbQnQVD44iKvLoytjd4CSv1dn5I1zdxZKymWBd+P h/XA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789478293; x=1790083093; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=dyunHw12Wk9XaltfwN7bGL3PCRgKnyr3LfmugMg8Gj0=; b=QmYJjRn+d3JxepSzv8gJn3iPQP/C+750iHtQOpkexLm/xCAdamT0dalewLI9zc3uH8 kzLUPgOxjxqd+NdoyGEyWXSN0Tip7rNtdnn6nCfpnUR/4de8fkUDGN4byG6USvnFzv9N EhzLMseU+tWzRLWZjYESgfX1S6+NRNwCcAwZqL0fhc45NjSYQzOYrwftBV1usZizRDgi K7O/pF+NdutFWVoIvc+yanZY0B+SRCvBnG6vHqmHBmG5yHe63sHtWbUkrmIzsy1uCbkG N1+gyGfD2Atc7nIsfPXOl2YWau+kKliuJAK0c/gczTz/JKwaF3vprudK13pidlT36Q6g 1lQw== X-Forwarded-Encrypted: i=1; AKwUvBygRqxWy1uOQ0r4aUEfgNGbNwOSzc3T4GbCe5Ne4LAOCsMJdfiRDiOXrC97Cuo4NEo3xG4z57+KkgTXTdI=@vger.kernel.org X-Gm-Message-State: AFuF++k7XMuDaZ/Um/bXKprAxpfh4YAdF0wh8S/ub2rgQO1/mowUO7kD Smlj3lWAieA6zXzQH/EvB3elAB4mAVY9jwYcCcBq2XfRoqW6LW/Ctoo8m51NS+FfgwE= X-Gm-Gg: AYBFou3Lf5g0DG6OAgxfmoqSnpFqVfNqh4kqRwbse84cSL0i1P9tvmcxvYihcmntLqA HRAgdtu09FiUN+rN1qLwjnPWx5kPdn/LdNTIS5Z5KHsuUTo8i1IiMV3U060BOaJws2l9Hh483h7 5Me83vbnz47fG1rAdJKCjxGpd1Znz/go7oxw4xz+glj5F+QkD8REaniDtouz37FaK6J/IpsBruu hCvhJN2ebK7KlTFHQHEu7taxFYpM94SdQJNB7CHfqtmYOXF/XIH6WGMsVC5mDY6kf/cU4xic8r0 E1uSaZtZ9va690u+qhPtUxz1x3xC+pfckGxZf18ruwlmlkx46y1y5iF4KV0fN51WNiXdpSN/G2T ZfwqqcEkaMNbEB7oMjSpEQw3AwY0h6pFaWdjobzgmkcMQTW0rfGNPIb6bIt8lmUS+Zdr/Aps2kG lbs236PxI0MHpdZW6rvbwpiaOlLsAYkJyrS3ZRmQR67kV+PUwqLBtsAyg9O70A22aIPFmGN9lS2 4ImOwmZKjkLWWznodR8SPZ8e1v3/1m4AgBbCdSoT7c0SIBPIe9NU4OG X-Received: by 2002:ad4:594b:0:b0:910:4c56:76d8 with SMTP id 6a1803df08f44-9123a42ec8fmr14508336d6.1.1789478292499; Tue, 15 Sep 2026 06:18:12 -0700 (PDT) Received: from toxicpanda.com (ec2-34-228-114-98.compute-1.amazonaws.com. [34.228.114.98]) by smtp.gmail.com with ESMTPSA id 6a1803df08f44-9120f45cb06sm120943296d6.11.2026.09.15.06.18.11 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 15 Sep 2026 06:18:11 -0700 (PDT) From: Josef Bacik Date: Tue, 15 Sep 2026 13:17:40 +0000 Subject: [PATCH RFC v3 13/13] x86, arm64: Build Tasks RCU on Tasks Trace readers in trampolines Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260915-b4-rcu-tasks-preempt-qs-v3-13-0ad30c4c5ee7@toxicpanda.com> References: <20260915-b4-rcu-tasks-preempt-qs-v3-0-0ad30c4c5ee7@toxicpanda.com> In-Reply-To: <20260915-b4-rcu-tasks-preempt-qs-v3-0-0ad30c4c5ee7@toxicpanda.com> To: "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Boqun Feng , Thomas Gleixner , Peter Zijlstra , Steven Rostedt , Masami Hiramatsu , Mark Rutland , Jiri Olsa , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , x86@kernel.org, Catalin Marinas , Will Deacon , Puranjay Mohan , Xu Kuohai , "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Boqun Feng , Thomas Gleixner , Peter Zijlstra , Steven Rostedt , Masami Hiramatsu , Mark Rutland , Jiri Olsa , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , x86@kernel.org, Catalin Marinas , Will Deacon , Puranjay Mohan , Xu Kuohai Cc: Andy Lutomirski , Josh Triplett , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Juergen Gross , Luis Chamberlain , Ihor Solodrai , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, xen-devel@lists.xenproject.org, Andy Lutomirski , Josh Triplett , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Juergen Gross , Luis Chamberlain , Ihor Solodrai , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, xen-devel@lists.xenproject.org, Josef Bacik X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openssh-sha256; t=1789478262; l=5993; i=josef@toxicpanda.com; h=from:subject:message-id; bh=R9t47tbzDpBlKabnqSyLf9h+YtCy+mgQn2kYE1SNKlg=; b=U1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgUBr36M/n0nWN0DNbnxwzIiCZez6MG JiruuNaSCI/zXsAAAAGcGF0YXR0AAAAAAAAAAZzaGE1MTIAAABTAAAAC3NzaC1lZDI1NTE5AAAA QLrYYFTmT4yc2HnPNOm135tMG3u8qPYsi+9GRPNIPdwz0QfAjvOD+lsxTwptmj0sG65nDpJ4ULX TbSFHTaStyQI= X-Developer-Key: i=josef@toxicpanda.com; a=openssh; fpr=SHA256:C8kOX2QUJCMqnCX+KEeoqRAjLo9L+ELOSH2NSAJHqGA With the preceding patches every trampoline whose lifetime Tasks RCU guards on x86-64 and arm64 -- ftrace_caller and its copies, the optprobe template, BPF trampolines via their glue, and the sample direct-call trampolines -- is a Tasks Trace RCU reader around its call-out, and the text outside that reader is known to rcu_tasks_trampoline_text(). Select HAVE_RCU_TRAMPOLINE_READERS on both (x86-64 with SMP for Tree SRCU and DYNAMIC_FTRACE, which is where its ftrace_caller changes and arch_rcu_tasks_trampoline_text() live; arm64 with DYNAMIC_FTRACE_WITH_ARGS likewise), which switches CONFIG_TASKS_RCU to the implementation added earlier in the series: a Tasks RCU grace period becomes a per-CPU pass over context switches and irq-exit reschedules outside trampoline text plus a Tasks Trace grace period, bounded by a few jiffies and preempt-off latency instead of by the longest stretch any task runs without sleeping. Other architectures keep the classic implementation. Update Documentation/RCU and the FORCE_TASKS_RCU help text to describe the variant and the obligation it places on trampolines. Assisted-by: LLM Signed-off-by: Josef Bacik --- .../RCU/Design/Requirements/Requirements.rst | 20 ++++++++++++++++= ++++ Documentation/RCU/checklist.rst | 7 ++++++- arch/arm64/Kconfig | 1 + arch/x86/Kconfig | 1 + kernel/rcu/Kconfig | 6 ++++-- 5 files changed, 32 insertions(+), 3 deletions(-) diff --git a/Documentation/RCU/Design/Requirements/Requirements.rst b/Docum= entation/RCU/Design/Requirements/Requirements.rst index 8101fe6229d5..34b81512cc5f 100644 --- a/Documentation/RCU/Design/Requirements/Requirements.rst +++ b/Documentation/RCU/Design/Requirements/Requirements.rst @@ -2756,6 +2756,26 @@ synchronize_rcu(), and rcu_barrier(), respectively. = In three APIs are therefore implemented by separate functions that check for voluntary context switches. =20 +Architectures that select ``CONFIG_HAVE_RCU_TRAMPOLINE_READERS`` keep the +same three APIs but implement the grace period differently +(``CONFIG_TASKS_RCU_TRAMPOLINE_READERS``). There, every trampoline whose +lifetime Tasks RCU guards enters a Tasks Trace RCU read-side critical +section (rcu_read_lock_trace() or its assembly equivalent) before calling +out and leaves it before returning, so a task anywhere inside such a +call-out is an ordinary Tasks Trace reader whether or not it is +preempted. The few trampoline instructions outside that reader can only +be occupied by a task that was interrupted there, so the grace period +additionally waits for each CPU to pass through a context switch, and the +irq-exit preemption path, the only switch that can catch a task inside +such text (rcu_tasks_trampoline_text()), briefly makes such a task a +holdout until it is next seen elsewhere. On such kernels an involuntary +context switch outside trampoline text *is* a Tasks-RCU quiescent state, +a Tasks RCU grace period no longer depends on how long any task runs +without sleeping, cond_resched_tasks_rcu_qs() is unnecessary, and the +obligation moves to the trampolines: anything that relies on +synchronize_rcu_tasks() to protect code a task may be preempted in must +take the Tasks Trace reader (see register_ftrace_direct()). + Tasks Rude RCU ~~~~~~~~~~~~~~ =20 diff --git a/Documentation/RCU/checklist.rst b/Documentation/RCU/checklist.= rst index 4b30f701225f..7082686cbd66 100644 --- a/Documentation/RCU/checklist.rst +++ b/Documentation/RCU/checklist.rst @@ -252,7 +252,12 @@ over a rather long period of time, but improvements ar= e always welcome! a. If the updater uses synchronize_rcu_tasks() or call_rcu_tasks(), then the readers must refrain from executing voluntary context switches, that is, from - blocking. + blocking. On architectures that select + CONFIG_HAVE_RCU_TRAMPOLINE_READERS a reader must in + addition be a Tasks Trace RCU reader (that is what the + trampolines there do around their call-outs); an + arbitrary stretch of preemptible kernel code is not + protected. =20 b. If the updater uses call_rcu_tasks_trace() or synchronize_rcu_tasks_trace(), then the diff --git a/arch/arm64/Kconfig b/arch/arm64/Kconfig index b5a51b0ef944..bf0e56006863 100644 --- a/arch/arm64/Kconfig +++ b/arch/arm64/Kconfig @@ -218,6 +218,7 @@ config ARM64 select HAVE_PERF_REGS select HAVE_PERF_USER_STACK_DUMP select HAVE_PREEMPT_DYNAMIC_KEY + select HAVE_RCU_TRAMPOLINE_READERS if DYNAMIC_FTRACE_WITH_ARGS select HAVE_REGS_AND_STACK_ACCESS_API select HAVE_RELIABLE_STACKTRACE select HAVE_POSIX_CPU_TIMERS_TASK_WORK diff --git a/arch/x86/Kconfig b/arch/x86/Kconfig index 15fd9ec5ecac..64c3814eb745 100644 --- a/arch/x86/Kconfig +++ b/arch/x86/Kconfig @@ -288,6 +288,7 @@ config X86 select MMU_GATHER_RCU_TABLE_FREE select MMU_GATHER_MERGE_VMAS select HAVE_POSIX_CPU_TIMERS_TASK_WORK + select HAVE_RCU_TRAMPOLINE_READERS if X86_64 && SMP && DYNAMIC_FTRACE select HAVE_REGS_AND_STACK_ACCESS_API select HAVE_RELIABLE_STACKTRACE if UNWINDER_ORC || STACK_VALIDATION select HAVE_FUNCTION_ARG_ACCESS_API diff --git a/kernel/rcu/Kconfig b/kernel/rcu/Kconfig index bbab14bc14c3..341b68b972ef 100644 --- a/kernel/rcu/Kconfig +++ b/kernel/rcu/Kconfig @@ -95,8 +95,10 @@ config FORCE_TASKS_RCU help This option force-enables a task-based RCU implementation that uses only voluntary context switch (not preemption!), - idle, and user-mode execution as quiescent states. Not for - manual selection in most cases. + idle, and user-mode execution as quiescent states, or, on + HAVE_RCU_TRAMPOLINE_READERS architectures, the variant built on + Tasks Trace RCU readers in trampolines. Not for manual + selection in most cases. =20 config NEED_TASKS_RCU bool --=20 2.55.0