From nobody Thu Sep 24 17:09:51 2026 Received: from mail-qv2-f12.google.com (mail-qv2-f12.google.com [74.125.230.140]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 876B938B7D4 for ; Tue, 22 Sep 2026 02:25:31 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.140 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790043933; cv=none; b=InABm2cBcmMUMe//mr5fGIqyeKiDMZL6XoXumlO8FD0vVoq5eh9jv7VYkN/y8cXyYaex1SU0JNAfINm0YFegFvm5sqdI0VPTE5tz3frJn4fWn/7qJpK5gqN5AOcUyL8+0rquZZVsuA15mLKUc5u4dlqDFH3R8DVVvC+JIxoHies= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790043933; c=relaxed/simple; bh=vAWpKfoMcBcYFQiIgPpQZt2gAuw45VdrXu4vr32WnR0=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=WUlwOm+63V+XrgCSt6RRoQUBrzltUI8SBCA5sm5zyrBinu8/Msu1zcq/444COJg7fpuIGQYnIbb6XIn3BR6v7+t0PLVc9ii9Ik84meiYK5FL5Cjkl/+UGARMcuG8lZgXNJNV8+N2WblocY7wZMoX9MoSEARAQs7HNMrABc6/AZA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com; spf=pass smtp.mailfrom=toxicpanda.com; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b=LtyfqdJq; arc=none smtp.client-ip=74.125.230.140 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b="LtyfqdJq" Received: by mail-qv2-f12.google.com with SMTP id 6a1803df08f44-912355b1bf6so13259486d6.1 for ; Mon, 21 Sep 2026 19:25:31 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=toxicpanda.com; s=google; t=1790043930; x=1790648730; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=fr6kFLYx3hqPZHfOilkUN+oxGtJHBsVQ4xuLKdXd4To=; b=LtyfqdJqmY0e9ELL0ssGj0h7gUpGjOkkm70X85Gt+IkXbKkFbVpdREac17cbxynHk8 2Ko6UALOcS1AMVly8SqrbniSteNnkxoai6dGRYwrkTtn/kIr1nwGTTUfTVFRfrqa8W6q MAIFjID8tKfIwkBuw2kRGoV99HBHkxWZcR6GvhW/nIXErVd4o41M/81OKAWZITsGlB7I q4Mjce36Twhvl3JOjdQYHfU5nJO3D9slnugJATmT9ZgVVoAjpoAtWx5csMA4pGcIqt+0 yzCqcDPR9VNfBZyy1Q9bKmkVqkzRfy7JdywDXdr+upYi/mAO9Yyf68CWG4SYnkr8yDOX NZ0Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790043930; x=1790648730; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=fr6kFLYx3hqPZHfOilkUN+oxGtJHBsVQ4xuLKdXd4To=; b=L1XbpmXKhGjiWCI/02/aZZQACpCKCYFN+szmP85+JtX4jTQ+7ykJDX0OMMzy/05KQX w5ztOKicl+UO84bxaXpgrlxKwt5VqIN8xtcNXRYNlEoLubceLimApj+B9dgwYzhSqhss cgb2SXjIAbgujc6shh2Sb0w905Rb0WDlzKfBnU+OVAjjH7bOo648JChAXmYd1lI8CF5i GM27bYWZHrirZ0v5RQO+ZWhZUd39oiYTBaIWanlVO9v+b7INyID9ql5dTCukK3ureEim wba5o6cmFzYR5HEAg4PJmxpIZs5N6t5L6PzjRK3zATcFR3sIhsz1qLXfmuasfR8/4z5F AONw== X-Forwarded-Encrypted: i=1; AKwUvBx1n/WkC/1x6u1PZO/VvlvYQwNHSsCU162Md339lus2gR52TWg8lqGCXvsJTC4byjIiaY3usbO4A595lUQ=@vger.kernel.org X-Gm-Message-State: AFuF++nOMHEYtZ5GdjlWeBbh0FaAtoFmjsLs7gLX5p48Js65K2SJ8ILk 3f7JYZ48gEBN+JB7uLDg8J+bcbudd7PbPu/rr5NK07xFiFTaapcHgXTu1RwgC0IqB6I= X-Gm-Gg: AYBFou2OKeCqUnZ9ngMO6rzsLWQUjzXCDJWXdqY73B7X54Vydl2zRnuSW2mNXUDfr/7 k3JG23FzTlPHKZPE72il+WzdC88etZY/o2KsS3QRD0kHP6yoW0KekHgpuleNkzToV7KZl2JIUhX 1JtlWHO/7mQHAzddNbkDVfoCrDJ6tOGKovFdJvE1QeBHq0nV1JfWAFmNXMVDbrbHvGxkBi+3hFe Hg6QkXkzLNXwyPFiBbzSd4yA89ygikaOKENjruNFZitxTcCXA9uDju07l+ZU5P0RnOqLRUwurGX ItjJxyyLzlxwTtaqUKXXOMfp9uTLOwDJOSBZNx9peahOw6lGBq/SPedIbOgegPkRWLi+2tKKZH0 GxCWvvhomgY+2JOtfqS7RUvfJr35DlEl4J6jg5K5HvDfsDfFaO2KwXA3csUxFpm3in9rWqCR6Aj ap47sUqdD0Gxr95UJ+OrACz7Dj9wr5FkBaTd6/hTFIz4xOzgdBwUn3/rsWHZcZUOgkL6jSGkEa6 Uqtry080KgxDvo= X-Received: by 2002:ad4:5d44:0:b0:911:2a7e:abbb with SMTP id 6a1803df08f44-913ff001c87mr26718046d6.14.1790043930382; Mon, 21 Sep 2026 19:25:30 -0700 (PDT) Received: from toxicpanda.com ([153.61.196.242]) by smtp.gmail.com with ESMTPSA id 6a1803df08f44-914025165e2sm4420596d6.46.2026.09.21.19.25.29 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 21 Sep 2026 19:25:29 -0700 (PDT) From: Josef Bacik Date: Tue, 22 Sep 2026 02:23:20 +0000 Subject: [PATCH v5 01/13] entry: Pass pt_regs to irqentry_exit_cond_resched() Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260922-b4-rcu-tasks-preempt-qs-v5-1-410f57770bad@toxicpanda.com> References: <20260922-b4-rcu-tasks-preempt-qs-v5-0-410f57770bad@toxicpanda.com> In-Reply-To: <20260922-b4-rcu-tasks-preempt-qs-v5-0-410f57770bad@toxicpanda.com> To: "Paul E. McKenney" , Frederic Weisbecker , Alexei Starovoitov , Steven Rostedt Cc: Boqun Feng , Masami Hiramatsu , Mark Rutland , Peter Zijlstra , Thomas Gleixner , Daniel Borkmann , Andrii Nakryiko , Puranjay Mohan , rcu@vger.kernel.org, bpf@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, Josef Bacik X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openssh-sha256; t=1790043820; l=4297; i=josef@toxicpanda.com; h=from:subject:message-id; bh=vAWpKfoMcBcYFQiIgPpQZt2gAuw45VdrXu4vr32WnR0=; b=U1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgUBr36M/n0nWN0DNbnxwzIiCZez6MG JiruuNaSCI/zXsAAAAGcGF0YXR0AAAAAAAAAAZzaGE1MTIAAABTAAAAC3NzaC1lZDI1NTE5AAAA QCPeuFAp5uVxV2kH2vDDmkA4ssSDw+6kXcs/1v8hvtFa+fEUwj418WQRUH5jUvQqcwsIafENDDP 8QcAKkovgmQk= X-Developer-Key: i=josef@toxicpanda.com; a=openssh; fpr=SHA256:C8kOX2QUJCMqnCX+KEeoqRAjLo9L+ELOSH2NSAJHqGA The irq-exit preemption path is about to need the interrupted context's registers to decide whether the preemption may be reported to Tasks RCU as a quiescent state. irqentry_exit_to_kernel_mode_preempt() already has them; hand them down through irqentry_exit_cond_resched(), its PREEMPT_DYNAMIC static-call and static-key variants, and raw_irqentry_exit_cond_resched(). The only caller outside the generic entry code is Xen PV's upcall handler, which has regs as well. No functional change. Assisted-by: LLM Signed-off-by: Josef Bacik --- arch/x86/xen/enlighten_pv.c | 2 +- include/linux/irq-entry-common.h | 13 +++++++------ kernel/entry/common.c | 6 +++--- 3 files changed, 11 insertions(+), 10 deletions(-) diff --git a/arch/x86/xen/enlighten_pv.c b/arch/x86/xen/enlighten_pv.c index 2c64b388f616..3d85035f5624 100644 --- a/arch/x86/xen/enlighten_pv.c +++ b/arch/x86/xen/enlighten_pv.c @@ -739,7 +739,7 @@ __visible noinstr void xen_pv_evtchn_do_upcall(struct p= t_regs *regs) =20 inhcall =3D get_and_clear_inhcall(); if (inhcall && !WARN_ON_ONCE(state.exit_rcu)) { - irqentry_exit_cond_resched(); + irqentry_exit_cond_resched(regs); instrumentation_end(); restore_inhcall(inhcall); } else { diff --git a/include/linux/irq-entry-common.h b/include/linux/irq-entry-com= mon.h index 0bb6c03481fa..e19b41ee6b18 100644 --- a/include/linux/irq-entry-common.h +++ b/include/linux/irq-entry-common.h @@ -343,24 +343,25 @@ typedef struct irqentry_state { =20 /** * irqentry_exit_cond_resched - Conditionally reschedule on return from in= terrupt + * @regs: Pointer to pt_regs of interrupted context * * Conditional reschedule with additional sanity checks. */ -void raw_irqentry_exit_cond_resched(void); +void raw_irqentry_exit_cond_resched(struct pt_regs *regs); =20 #ifdef CONFIG_PREEMPT_DYNAMIC #if defined(CONFIG_HAVE_PREEMPT_DYNAMIC_CALL) #define irqentry_exit_cond_resched_dynamic_enabled raw_irqentry_exit_cond_= resched #define irqentry_exit_cond_resched_dynamic_disabled NULL DECLARE_STATIC_CALL(irqentry_exit_cond_resched, raw_irqentry_exit_cond_res= ched); -#define irqentry_exit_cond_resched() static_call(irqentry_exit_cond_resche= d)() +#define irqentry_exit_cond_resched(regs) static_call(irqentry_exit_cond_re= sched)(regs) #elif defined(CONFIG_HAVE_PREEMPT_DYNAMIC_KEY) DECLARE_STATIC_KEY_TRUE(sk_dynamic_irqentry_exit_cond_resched); -void dynamic_irqentry_exit_cond_resched(void); -#define irqentry_exit_cond_resched() dynamic_irqentry_exit_cond_resched() +void dynamic_irqentry_exit_cond_resched(struct pt_regs *regs); +#define irqentry_exit_cond_resched(regs) dynamic_irqentry_exit_cond_resche= d(regs) #endif #else /* CONFIG_PREEMPT_DYNAMIC */ -#define irqentry_exit_cond_resched() raw_irqentry_exit_cond_resched() +#define irqentry_exit_cond_resched(regs) raw_irqentry_exit_cond_resched(re= gs) #endif /* CONFIG_PREEMPT_DYNAMIC */ =20 /** @@ -465,7 +466,7 @@ static inline void irqentry_exit_to_kernel_mode_preempt= (struct pt_regs *regs, return; =20 if (IS_ENABLED(CONFIG_PREEMPTION)) - irqentry_exit_cond_resched(); + irqentry_exit_cond_resched(regs); } =20 /** diff --git a/kernel/entry/common.c b/kernel/entry/common.c index e3d381fd3d25..e4acd50bd81a 100644 --- a/kernel/entry/common.c +++ b/kernel/entry/common.c @@ -134,7 +134,7 @@ static inline bool arch_irqentry_exit_need_resched(void= ); static inline bool arch_irqentry_exit_need_resched(void) { return true; } #endif =20 -void raw_irqentry_exit_cond_resched(void) +void raw_irqentry_exit_cond_resched(struct pt_regs *regs) { if (!preempt_count()) { /* Sanity check RCU and thread stack */ @@ -150,11 +150,11 @@ void raw_irqentry_exit_cond_resched(void) DEFINE_STATIC_CALL(irqentry_exit_cond_resched, raw_irqentry_exit_cond_resc= hed); #elif defined(CONFIG_HAVE_PREEMPT_DYNAMIC_KEY) DEFINE_STATIC_KEY_TRUE(sk_dynamic_irqentry_exit_cond_resched); -void dynamic_irqentry_exit_cond_resched(void) +void dynamic_irqentry_exit_cond_resched(struct pt_regs *regs) { if (!static_branch_unlikely(&sk_dynamic_irqentry_exit_cond_resched)) return; - raw_irqentry_exit_cond_resched(); + raw_irqentry_exit_cond_resched(regs); } #endif #endif --=20 2.55.0 From nobody Thu Sep 24 17:09:51 2026 Received: from mail-qk2-f13.google.com (mail-qk2-f13.google.com [74.125.230.205]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EAA9938F25C for ; Tue, 22 Sep 2026 02:26:11 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.205 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790043974; cv=none; b=RysMwPP9zm3kOjEO0IekU2ayX9IvvZx8CvqSNsDvrjhLzJubvMhfzBpZRWcDou1B35MTD3AldpI1q2KP2dZkL98pRufgplix8oUp+uswri5O9V+kkqfDYKqQPfuQnt2gsrw+hT953GVy1+02Gk7Ph5ZseNeLGaF/jQveQ/yw9z0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790043974; c=relaxed/simple; bh=qhXNBlJJp01Ei07kqFZDAUfEsCu0gSp6eIBNqqgCukM=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=pwHw3MYMnXqiO5tg4G+RzNtncY4pmq/uMrUltDlAWhUEETvrtsOtaoWEd2G8gXSVThQ+VH2xhknJOCrwWcp/F5Ja8GJk2W4X/PcO4cL77qq3C6KFV83XvBN7nnvciDJkRMnHXylHr5OVynpA/OZqMz0O5xlDvzM+09sNHK6/89g= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com; spf=pass smtp.mailfrom=toxicpanda.com; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b=e1o7wX9U; arc=none smtp.client-ip=74.125.230.205 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b="e1o7wX9U" Received: by mail-qk2-f13.google.com with SMTP id af79cd13be357-939656ff6d9so384785085a.1 for ; Mon, 21 Sep 2026 19:26:11 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=toxicpanda.com; s=google; t=1790043970; x=1790648770; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=myll629Qfz4YrwGmTVALnxB11mMoUz/LpucXoYBTMJw=; b=e1o7wX9UWTZ7B6vRHp9rd8HOR0I07lXdFTREh5AkF+kr/RnnXALF+CVkcpTvyPhTbv 9hbqq9Nzx39usqOVcEDg0m0sgzqvwnxzbdS39OJVyf76sDKiVeMWsza95NVxyUtwmt4a bfoEMpzROv23uxc8d7p+01XQjgHLobw3o0ivcqxrRk/xJDfrNeS+mGXmm8NkcY3Jxr8c O+TVI48Su8l5tUcDDlGMqdnkXzv+aOi9VG1hhZvOwB22gUx5kHjnEnHmixbQJEBAgv+I 9nf4JG6wOSZszUZjAaGE1ZeBuP7JXMNrv63VQfmShCwQfCC7c70CJi90SBpfQEVzC5DI H/qg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790043970; x=1790648770; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=myll629Qfz4YrwGmTVALnxB11mMoUz/LpucXoYBTMJw=; b=CIt1MYZxE6gMAKzrhd0nbie64G33vExsGaq6bR2HX4t6jLh3aRgL4unfN4Tit56KTc 7d/SF1mBYBZZo+ChpngRWB0O+3Tz7rXho/kGw3PQCZAf3bh9henuzy31H5wXsY4a7uPp fSK78boCEl4JlxWwA3QJY1HYUw6/TJ8daShDqCi1Ru8ZHIPMNZbCP2RxnDk4vFIhDWaz PjA+DwRXo2V/BFO4U/w9FaOtOVGZHKVFuiiJglYUmRjhb2zOmow6089T+/v8o4uVOPMl 5KqWK8LXDuxfYRPbZg2i9x7q2seJGisUkVybXzR1v/IllO0vPkUZqcaD5BzjyYMnE86g logg== X-Forwarded-Encrypted: i=1; AKwUvBxWbrsQxLKbD+cEeIpnb9Q4bJUfTT1VQzzUVf/0YAWB+BlEkSc+NG055jfGAFG1hyrom6fSaMJkjItN9pE=@vger.kernel.org X-Gm-Message-State: AFuF++lkhziAUG0dy+lMjAjocvaJRznFhzygSv/SwGNjmc3ISPk5vrtX 5LJ/y9UuD01cm/WGXA1kqdAsAYcKz92tUYe81MrM4WbsDrXIpZ/EP0ptDMLtTtvTUic= X-Gm-Gg: AYBFou3JXjXLNpDdxHfy0jJAkyj/4UmEYNdTwbS04woWt6K1P+8oAkCn7Csrq5i3LGr oflhpMdsCzO7MPH/tH1ZSNocl6apt8OI/Kj0MjBFquA5Itx08v7R8bA4xu9nJvX+lcAU9ZTdWl+ BpZHpP8phr/zIZZ9MzfNn/y92jB6mAnzlKGrGITqUe22pnaEs5vxSXuz1+7TxkpvtBOmfHDJbkI 9D5avlvQ5ZNkuXX4KPSYXxP3rJv9jkyIaVoRd5hBpKWucj/zHOm9z0+KcsmpRtpKxmon/+g0yhg 7//gohLelWjBGJGyqAHWNgP9wR7qsYd1ejddSqYKKPb35hCSUR1mxS3qrvpxU2qEHZWkDmtNkfu uFw92mNHGteJ5O7ZBL8gv2UB/N8duohqR9YTXJRuA9dWZEEPfecVQ+76Hh8OwNl/stvBaW1M4W8 JuzWE8BONsp+J7LU1n8T9I1asJ0hnWOb07ql83Ml3ctxFh6+QWlDzVB3flWdOV6zOcxINXq/YdD +B/DrCOe7RNZAs= X-Received: by 2002:a05:620a:46aa:b0:93b:c45f:9024 with SMTP id af79cd13be357-93c15d9d1aamr363549185a.7.1790043970038; Mon, 21 Sep 2026 19:26:10 -0700 (PDT) Received: from toxicpanda.com ([153.61.196.250]) by smtp.gmail.com with ESMTPSA id af79cd13be357-93c1d1d513bsm21244885a.30.2026.09.21.19.26.08 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 21 Sep 2026 19:26:08 -0700 (PDT) From: Josef Bacik Date: Tue, 22 Sep 2026 02:23:21 +0000 Subject: [PATCH v5 02/13] rcu-tasks: Add a Tasks RCU implementation for reader-marked trampolines Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260922-b4-rcu-tasks-preempt-qs-v5-2-410f57770bad@toxicpanda.com> References: <20260922-b4-rcu-tasks-preempt-qs-v5-0-410f57770bad@toxicpanda.com> In-Reply-To: <20260922-b4-rcu-tasks-preempt-qs-v5-0-410f57770bad@toxicpanda.com> To: "Paul E. McKenney" , Frederic Weisbecker , Alexei Starovoitov , Steven Rostedt Cc: Boqun Feng , Masami Hiramatsu , Mark Rutland , Peter Zijlstra , Thomas Gleixner , Daniel Borkmann , Andrii Nakryiko , Puranjay Mohan , rcu@vger.kernel.org, bpf@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, Josef Bacik X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openssh-sha256; t=1790043820; l=31523; i=josef@toxicpanda.com; h=from:subject:message-id; bh=qhXNBlJJp01Ei07kqFZDAUfEsCu0gSp6eIBNqqgCukM=; b=U1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgUBr36M/n0nWN0DNbnxwzIiCZez6MG JiruuNaSCI/zXsAAAAGcGF0YXR0AAAAAAAAAAZzaGE1MTIAAABTAAAAC3NzaC1lZDI1NTE5AAAA QNU+11JcL4jIkRTQoAN/AarWuaUILZa7eBrbfGz4XpzyPKDByZwIFKdncoGOCLAxjUEpPdrqHoz Vr7fl3d/3Ww4= X-Developer-Key: i=josef@toxicpanda.com; a=openssh; fpr=SHA256:C8kOX2QUJCMqnCX+KEeoqRAjLo9L+ELOSH2NSAJHqGA Tasks RCU waits for every task to pass through a voluntary context switch, usermode or idle, because a preempted task might be sitting in a trampoline that is about to be freed and nothing marks it as such. With PREEMPT_LAZY that is a poor fit for servers: cond_resched() is a no-op, so a CPU-bound kthread only ever leaves the CPU by preemption, and one such kthread holds every synchronize_rcu_tasks() caller -- ftrace and BPF trampoline teardown under their mutexes, the kprobe jump optimizer under text_mutex and cpus_read_lock() -- hostage for as long as it runs. Following the discussion on v2, take the other road: let the architecture make its trampolines Tasks Trace RCU readers. When an architecture selects HAVE_RCU_TRAMPOLINE_READERS it promises that every trampoline whose lifetime Tasks RCU guards enters rcu_read_lock_trace() (or its assembly equivalent) before calling out and leaves it before returning, so a task anywhere inside such a call-out, preempted or not, is an ordinary Tasks Trace reader. That leaves the few instructions of trampoline text before the reader is entered and after it is left (plus, in a later patch, the bytes a kprobe jump optimization is about to overwrite). A task can only linger there by being interrupted there, and such text never calls anything that schedules, so instead of tracking tasks we track CPUs: every pass through __schedule() is a per-CPU quiescent event, except that the one context switch that can catch a task at an arbitrary instruction -- a preemption from irq exit -- first records the interrupted IP in the task and parks it on a per-CPU list for the duration (reusing the fields and lists the classic flavor keeps for its exit-path bookkeeping), and, if the IP is inside such "unmarked" text, puts the task on a short holdout list; the task takes itself off at its next context switch outside such a preemption or irq-exit check that finds it elsewhere. Usermode (the existing tick hook, or a nohz_full CPU in an RCU extended quiescent state) and idle count as well. rcu_tasks_trampoline_text() does the classification: anything outside core and module text, plus an arch hook for things like static ftrace stubs and return thunks. The grace period, run by the existing rcu_tasks kthread so that call_rcu_tasks(), synchronize_rcu_tasks() and rcu_barrier_tasks() keep their names and callers, is: wait for every online CPU to context switch or be seen in an RCU extended quiescent state (nudging stragglers with resched_cpu() after a jiffy), drain the holdout list as it stood, synchronize_rcu_tasks_trace() for everything inside the readers, then one more CPU pass and drain for tasks that have since left the reader into the trailing instructions. That is bounded by a few jiffies, preempt-off latency and an SRCU grace period rather than by the longest stretch any task runs without sleeping, needs no per-task scan, and makes cond_resched_tasks_rcu_qs() unnecessary on such architectures. Unlike the classic flavor it also waits for an idle task caught in a trampoline, since an idle CPU only counts while RCU is not watching it. rcu_tasks_wait_irq_preempted() walks the parked lists for the one caller (the kprobe jump optimizer, later in the series) that makes ordinary text unsafe to be parked in and so has to wait out tasks that were preempted there before it said so. The classic implementation is untouched and remains the default; the new one is built only as CONFIG_TASKS_RCU_TRAMPOLINE_READERS when the architecture opts in and uses the generic irq entry code, whose reschedule check gains the rcu_tasks_irq_resched() call. Nothing selects it yet. Suggested-by: Paul E. McKenney Suggested-by: Alexei Starovoitov Assisted-by: LLM Signed-off-by: Josef Bacik --- include/linux/rcupdate.h | 24 ++- include/linux/sched.h | 1 + kernel/entry/common.c | 8 +- kernel/fork.c | 1 + kernel/rcu/Kconfig | 22 +++ kernel/rcu/tasks.h | 459 +++++++++++++++++++++++++++++++++++++++++++= +++- kernel/rcu/update.c | 2 + 7 files changed, 508 insertions(+), 9 deletions(-) diff --git a/include/linux/rcupdate.h b/include/linux/rcupdate.h index 44c07a66edff..79e4b14f83b9 100644 --- a/include/linux/rcupdate.h +++ b/include/linux/rcupdate.h @@ -50,6 +50,23 @@ token_context_lock_instance(RCU, RCU_BH); /* Exported common interfaces */ void call_rcu(struct rcu_head *head, rcu_callback_t func); void rcu_barrier_tasks(void); + +/* + * Trampoline-reader Tasks RCU (CONFIG_TASKS_RCU_TRAMPOLINE_READERS), see + * kernel/rcu/tasks.h. rcu_tasks_irq_resched_enter()/_exit() bracket the + * irq-exit preemption; rcu_tasks_trampoline_text() and the arch_ override + * classify an interrupted IP; rcu_tasks_wait_irq_preempted() lets a caller + * wait out tasks already preempted somewhere it is about to make unsafe. + */ +void rcu_tasks_irq_resched_enter(unsigned long ip); +void rcu_tasks_irq_resched_exit(void); +bool rcu_tasks_trampoline_text(unsigned long ip); +bool arch_rcu_tasks_trampoline_text(unsigned long ip); +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS +void rcu_tasks_wait_irq_preempted(bool (*inside)(unsigned long ip)); +#else +static inline void rcu_tasks_wait_irq_preempted(bool (*inside)(unsigned lo= ng ip)) { } +#endif void synchronize_rcu(void); =20 /* @@ -180,11 +197,16 @@ static inline void rcu_nocb_flush_deferred_wakeup(voi= d) { } #ifdef CONFIG_TASKS_RCU_GENERIC =20 # ifdef CONFIG_TASKS_RCU -# define rcu_tasks_classic_qs(t, preempt) \ +# ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS +void rcu_tasks_note_qs(struct task_struct *t, bool preempt); +# define rcu_tasks_classic_qs(t, preempt) rcu_tasks_note_qs((t), (preempt= )) +# else +# define rcu_tasks_classic_qs(t, preempt) \ do { \ if (!(preempt) && READ_ONCE((t)->rcu_tasks_holdout)) \ WRITE_ONCE((t)->rcu_tasks_holdout, false); \ } while (0) +# endif void call_rcu_tasks(struct rcu_head *head, rcu_callback_t func); void synchronize_rcu_tasks(void); void rcu_tasks_torture_stats_print(char *tt, char *tf); diff --git a/include/linux/sched.h b/include/linux/sched.h index 8b3d47a325cc..15beb44caa2c 100644 --- a/include/linux/sched.h +++ b/include/linux/sched.h @@ -957,6 +957,7 @@ struct task_struct { u8 rcu_tasks_holdout; u8 rcu_tasks_idx; int rcu_tasks_idle_cpu; + unsigned long rcu_tasks_irq_ip; struct list_head rcu_tasks_holdout_list; int rcu_tasks_exit_cpu; struct list_head rcu_tasks_exit_list; diff --git a/kernel/entry/common.c b/kernel/entry/common.c index e4acd50bd81a..94318519998c 100644 --- a/kernel/entry/common.c +++ b/kernel/entry/common.c @@ -6,6 +6,7 @@ #include #include #include +#include #include #include =20 @@ -141,8 +142,13 @@ void raw_irqentry_exit_cond_resched(struct pt_regs *re= gs) rcu_irq_exit_check_preempt(); if (IS_ENABLED(CONFIG_DEBUG_ENTRY)) WARN_ON_ONCE(!on_thread_stack()); - if (need_resched() && arch_irqentry_exit_need_resched()) + if (need_resched() && arch_irqentry_exit_need_resched()) { + if (IS_ENABLED(CONFIG_TASKS_RCU_TRAMPOLINE_READERS)) + rcu_tasks_irq_resched_enter(instruction_pointer(regs)); preempt_schedule_irq(); + if (IS_ENABLED(CONFIG_TASKS_RCU_TRAMPOLINE_READERS)) + rcu_tasks_irq_resched_exit(); + } } } #ifdef CONFIG_PREEMPT_DYNAMIC diff --git a/kernel/fork.c b/kernel/fork.c index 416758c8a3d4..8077336bb136 100644 --- a/kernel/fork.c +++ b/kernel/fork.c @@ -1871,6 +1871,7 @@ static inline void rcu_copy_process(struct task_struc= t *p) p->rcu_tasks_holdout =3D false; INIT_LIST_HEAD(&p->rcu_tasks_holdout_list); p->rcu_tasks_idle_cpu =3D -1; + p->rcu_tasks_irq_ip =3D 0; INIT_LIST_HEAD(&p->rcu_tasks_exit_list); #endif /* #ifdef CONFIG_TASKS_RCU */ #ifdef CONFIG_TASKS_TRACE_RCU diff --git a/kernel/rcu/Kconfig b/kernel/rcu/Kconfig index 332df7a7a634..bbab14bc14c3 100644 --- a/kernel/rcu/Kconfig +++ b/kernel/rcu/Kconfig @@ -107,6 +107,28 @@ config TASKS_RCU default NEED_TASKS_RCU && PREEMPTION select IRQ_WORK =20 +config HAVE_RCU_TRAMPOLINE_READERS + bool + help + Select this if the architecture uses the generic irq entry code and + every trampoline whose lifetime Tasks RCU guards on it (ftrace + trampolines, kprobe out-of-line and optimized-probe slots, BPF + trampolines, out-of-line ftrace direct-call trampolines) enters a + Tasks Trace RCU read-side critical section before calling out of + the trampoline and leaves it before returning, and any core text + that runs on behalf of such a trampoline outside that reader is + reported by arch_rcu_tasks_trampoline_text(). The assembly readers + use the this_cpu_inc() form of SRCU-fast, hence !NEED_SRCU_NMI_SAFE. + +config TASKS_RCU_TRAMPOLINE_READERS + def_bool TASKS_RCU && HAVE_RCU_TRAMPOLINE_READERS && GENERIC_IRQ_ENTRY &&= !NEED_SRCU_NMI_SAFE + select TASKS_TRACE_RCU + help + Implement the Tasks RCU grace period as a per-CPU pass over + context switches and irq-exit reschedules outside trampoline text + plus a Tasks Trace RCU grace period, instead of waiting for every + task to voluntarily context switch. See kernel/rcu/tasks.h. + config FORCE_TASKS_RUDE_RCU bool "Force selection of Tasks Rude RCU" depends on RCU_EXPERT diff --git a/kernel/rcu/tasks.h b/kernel/rcu/tasks.h index 627295396cd9..f03be742be48 100644 --- a/kernel/rcu/tasks.h +++ b/kernel/rcu/tasks.h @@ -152,7 +152,7 @@ static struct rcu_tasks rt_name =3D \ .kname =3D #rt_name, \ } =20 -#ifdef CONFIG_TASKS_RCU +#if defined(CONFIG_TASKS_RCU) && !defined(CONFIG_TASKS_RCU_TRAMPOLINE_READ= ERS) =20 /* Report delay of scan exiting tasklist in rcu_tasks_postscan(). */ static void tasks_rcu_exit_stall(struct timer_list *unused); @@ -802,7 +802,7 @@ static void rcu_tasks_torture_stats_print_generic(struc= t rcu_tasks *rtp, char *t =20 #endif // #ifndef CONFIG_TINY_RCU =20 -#if defined(CONFIG_TASKS_RCU) +#if defined(CONFIG_TASKS_RCU) && !defined(CONFIG_TASKS_RCU_TRAMPOLINE_READ= ERS) =20 //////////////////////////////////////////////////////////////////////// // @@ -897,10 +897,444 @@ static void rcu_tasks_wait_gp(struct rcu_tasks *rtp) rtp->postgp_func(rtp); } =20 -#endif /* #if defined(CONFIG_TASKS_RCU) */ +#endif /* #if defined(CONFIG_TASKS_RCU) && !defined(CONFIG_TASKS_RCU_TRAMP= OLINE_READERS) */ =20 #ifdef CONFIG_TASKS_RCU =20 +static int rcu_tasks_lazy_ms =3D -1; +module_param(rcu_tasks_lazy_ms, int, 0444); + +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS + +//////////////////////////////////////////////////////////////////////// +// +// Tasks RCU for architectures whose trampolines are Tasks Trace RCU +// readers (CONFIG_HAVE_RCU_TRAMPOLINE_READERS). +// +// On these architectures every piece of text whose lifetime Tasks RCU +// guards -- ftrace trampolines, kprobe optinsn slots, BPF trampoline +// images, out-of-line ftrace direct-call trampolines -- enters a Tasks +// Trace RCU read-side critical section before calling out of itself and +// leaves it before returning, so a task anywhere inside such a call-out, +// preempted or not, is an ordinary rcu_read_lock_trace() reader and +// synchronize_rcu_tasks_trace() waits for it. +// +// What that cannot cover is the handful of instructions in the trampoline +// before the reader is entered and after it is left, and the one user that +// has no trampoline at all: the bytes after a kprobe that the jump +// optimizer is about to overwrite. A task can only linger in such +// "unmarked" text by being interrupted there; unmarked text never calls +// anything that could schedule. So a context switch on a CPU tells us th= at +// whatever that CPU was running is out of unmarked text, with one +// exception: a preemption from the irq-exit path, which can happen at any +// instruction boundary. That path has the interrupted pt_regs in hand, so +// just before it preempts it records the IP in the task and checks it +// (rcu_tasks_trampoline_text()); if it is inside unmarked text the task +// goes on a short holdout list first, and takes itself off again at its +// next context switch outside such a preemption or its next irq-exit +// check that finds it elsewhere. With that, every pass through +// __schedule() is a per-CPU quiescent event, as are usermode and idle. +// +// A grace period is then: +// +// 1. Wait for every online CPU to context switch or be found in an RCU +// extended quiescent state (deep idle, nohz_full userspace), nudging +// stragglers with resched_cpu(). Afterwards no task is in the leading +// unmarked instructions of a dying trampoline unless it is on the +// holdout list. +// 2. Wait for the holdout list (as it stood) to drain. +// 3. synchronize_rcu_tasks_trace(), for everything inside the readers. +// 4. Repeat 1 and 2 for tasks that have since left the reader and are in +// the trailing unmarked instructions. +// +// which is bounded by a few jiffies plus preempt-off latency plus an SRCU +// grace period, independent of how long any task runs without sleeping. +// Unlike the classic implementation this does wait for an idle task caught +// in a trampoline, since an idle CPU only counts while RCU is not watching +// it. + +static void rcu_tasks_tramp_wait_gp(struct rcu_tasks *rtp); +void call_rcu_tasks(struct rcu_head *rhp, rcu_callback_t func); +DEFINE_RCU_TASKS(rcu_tasks, rcu_tasks_tramp_wait_gp, call_rcu_tasks, "RCU = Tasks"); + +/* Per-CPU count of Tasks RCU quiescent events, and the GP kthread's snaps= hot. */ +static DEFINE_PER_CPU(unsigned long, rcu_tasks_qs_seq); +static DEFINE_PER_CPU(unsigned long, rcu_tasks_qs_snap); + +/* + * Tasks currently switched out by an irq-exit preemption are kept, with t= he + * interrupted IP, on the per-CPU rtp_exit_list of the CPU that preempted = them + * (reusing the list, lock and task_struct fields the classic flavor uses = for + * its exit-path bookkeeping, which this flavor does not need), so that + * rcu_tasks_wait_irq_preempted() can find them without a tasklist scan and + * regardless of where they are in exit. + */ + +/* Tasks last seen preempted inside unmarked trampoline text. */ +static LIST_HEAD(rcu_tasks_tramp_holdouts); +static DEFINE_RAW_SPINLOCK(rcu_tasks_tramp_lock); + +/* CPUs / holdouts the current grace period is still waiting for. */ +static struct cpumask rcu_tasks_pending_cpus; +static LIST_HEAD(rcu_tasks_gp_holdouts); + +/** + * arch_rcu_tasks_trampoline_text - Does the architecture treat @ip as unm= arked trampoline text? + * @ip: kernel text address inside core kernel text + * + * See rcu_tasks_trampoline_text(). Architectures override this to flag + * core text that runs on behalf of a trampoline outside its Tasks Trace + * reader, e.g. static ftrace entry stubs or return thunks that hold a + * trampoline address they are about to jump to. + */ +bool __weak arch_rcu_tasks_trampoline_text(unsigned long ip) +{ + return false; +} + +/** + * rcu_tasks_trampoline_text - Is @ip in text Tasks RCU protects but no re= ader marks? + * @ip: an interrupted instruction pointer + * + * True when a task interrupted at @ip may be executing, or about to enter + * or return into, text whose lifetime depends on synchronize_rcu_tasks() + * without being inside the Tasks Trace reader that text takes around its + * call-outs: + * + * - anything outside core kernel and module text (ftrace and BPF + * trampolines, kprobe slots and other dynamically allocated text; this + * deliberately does not ask is_ftrace_trampoline() and friends, since + * text being torn down may already be unregistered there); + * - whatever the architecture adds via arch_rcu_tasks_trampoline_text(). + * + * A false positive only makes the task a holdout until its next quiescent + * event. Called with interrupts disabled from the irq-exit path. + */ +bool rcu_tasks_trampoline_text(unsigned long ip) +{ + if (core_kernel_text(ip)) + return arch_rcu_tasks_trampoline_text(ip); + return !is_module_text_address(ip); +} +NOKPROBE_SYMBOL(rcu_tasks_trampoline_text); + +/* Note a Tasks RCU quiescent event on this CPU. */ +static void rcu_tasks_qs_event(void) +{ + unsigned long *seq; + + guard(preempt_notrace)(); + seq =3D this_cpu_ptr(&rcu_tasks_qs_seq); + /* Order a preceding rcu_tasks_tramp_hold() before the count. */ + smp_store_release(seq, *seq + 1); +} + +static void rcu_tasks_tramp_hold(struct task_struct *t) +{ + unsigned long flags; + + if (t->rcu_tasks_holdout) + return; + raw_spin_lock_irqsave(&rcu_tasks_tramp_lock, flags); + list_add_tail(&t->rcu_tasks_holdout_list, &rcu_tasks_tramp_holdouts); + WRITE_ONCE(t->rcu_tasks_holdout, true); + raw_spin_unlock_irqrestore(&rcu_tasks_tramp_lock, flags); +} + +static void rcu_tasks_tramp_release(struct task_struct *t) +{ + unsigned long flags; + + if (likely(!t->rcu_tasks_holdout)) + return; + raw_spin_lock_irqsave(&rcu_tasks_tramp_lock, flags); + list_del_init(&t->rcu_tasks_holdout_list); + WRITE_ONCE(t->rcu_tasks_holdout, false); + raw_spin_unlock_irqrestore(&rcu_tasks_tramp_lock, flags); +} + +/** + * rcu_tasks_irq_resched_enter - Tasks RCU hook for the irq-exit reschedul= e check + * @ip: instruction pointer of the interrupted (task-level) context + * + * Called with interrupts disabled when an interrupt returning to kernel + * mode is about to preempt_schedule_irq(), the one context switch that can + * catch a task inside unmarked trampoline text. Record where the task is + * parked for as long as it is (rcu_tasks_wait_irq_preempted() looks at + * that), and if it is inside such text make it a holdout before + * __schedule() reports the quiescent event; if it is not, this is as good + * as a voluntary switch for ending an earlier hold. + */ +void rcu_tasks_irq_resched_enter(unsigned long ip) +{ + struct task_struct *t =3D current; + struct rcu_tasks_percpu *rtpcp =3D this_cpu_ptr(rcu_tasks.rtpcpu); + + lockdep_assert_irqs_disabled(); + WRITE_ONCE(t->rcu_tasks_irq_ip, ip); + t->rcu_tasks_exit_cpu =3D smp_processor_id(); + raw_spin_lock_rcu_node(rtpcp); + list_add(&t->rcu_tasks_exit_list, &rtpcp->rtp_exit_list); + raw_spin_unlock_rcu_node(rtpcp); + + if (unlikely(rcu_tasks_trampoline_text(ip))) + rcu_tasks_tramp_hold(t); + else + rcu_tasks_tramp_release(t); +} +NOKPROBE_SYMBOL(rcu_tasks_irq_resched_enter); + +/** + * rcu_tasks_irq_resched_exit - preempt_schedule_irq() has returned + * + * The task is running again (possibly elsewhere) and about to return to t= he + * interrupted context; it is no longer parked anywhere. + */ +void rcu_tasks_irq_resched_exit(void) +{ + struct task_struct *t =3D current; + struct rcu_tasks_percpu *rtpcp =3D per_cpu_ptr(rcu_tasks.rtpcpu, t->rcu_t= asks_exit_cpu); + + lockdep_assert_irqs_disabled(); + raw_spin_lock_rcu_node(rtpcp); + list_del_init(&t->rcu_tasks_exit_list); + raw_spin_unlock_rcu_node(rtpcp); + WRITE_ONCE(t->rcu_tasks_irq_ip, 0); +} +NOKPROBE_SYMBOL(rcu_tasks_irq_resched_exit); + +/** + * rcu_tasks_note_qs - Tasks RCU hook for a context switch or explicit QS + * @t: current + * @preempt: this is a preemption rather than a voluntary switch + * + * Every pass through __schedule() (and cond_resched_tasks_rcu_qs(), and a + * tick from userspace or idle) is a quiescent event for this CPU: unmarked + * trampoline text never calls anything that schedules, and the irq-exit + * path has already made @t a holdout if it is preempting inside such text. + * Any of these outside an irq-exit preemption also shows @t itself to be + * outside, ending an earlier hold -- including cond_resched() under + * PREEMPT_DYNAMIC's none/voluntary modes, where the irq-exit path is off. + */ +void rcu_tasks_note_qs(struct task_struct *t, bool preempt) +{ + WARN_ON_ONCE(t !=3D current); + if (!READ_ONCE(t->rcu_tasks_irq_ip)) + rcu_tasks_tramp_release(t); + rcu_tasks_qs_event(); +} +EXPORT_SYMBOL_GPL(rcu_tasks_note_qs); /* cond_resched_tasks_rcu_qs() */ + +/** + * rcu_tasks_wait_irq_preempted - wait for tasks irq-preempted inside @ins= ide + * @inside: predicate on a task's recorded irq-exit preemption IP + * + * For a caller about to make some ordinary text unsafe to be parked in + * (the kprobe jump optimizer): once the caller has arranged for + * rcu_tasks_trampoline_text() to cover that text, new irq-exit preemptions + * there become holdouts, but a task preempted there earlier is invisible + * to the grace period. Wait until no parked task's recorded preemption IP + * is inside; a following synchronize_rcu_tasks() then covers the rest. + * The leading synchronize_rcu() orders the caller's arrangement against + * preemptions in flight, which run with interrupts disabled. + */ +void rcu_tasks_wait_irq_preempted(bool (*inside)(unsigned long ip)) +{ + struct task_struct *t; + unsigned long flags; + int cpu, kick; + bool found; + + synchronize_rcu(); + for (;;) { + found =3D false; + for_each_possible_cpu(cpu) { + struct rcu_tasks_percpu *rtpcp =3D per_cpu_ptr(rcu_tasks.rtpcpu, cpu); + + kick =3D -1; + raw_spin_lock_irqsave_rcu_node(rtpcp, flags); + list_for_each_entry(t, &rtpcp->rtp_exit_list, rcu_tasks_exit_list) { + if (inside(READ_ONCE(t->rcu_tasks_irq_ip))) { + found =3D true; + if (task_curr(t)) + kick =3D task_cpu(t); + } + } + raw_spin_unlock_irqrestore_rcu_node(rtpcp, flags); + if (kick >=3D 0) + resched_cpu(kick); + } + if (!found) + return; + schedule_timeout_uninterruptible(1); + } +} + +/* Has @cpu passed a quiescent event since the snapshot, or need it not? */ +static bool rcu_tasks_cpu_quiescent(int cpu) +{ + if (!cpu_online(cpu)) + return true; + /* Pairs with the release in rcu_tasks_qs_event(). */ + if (smp_load_acquire(per_cpu_ptr(&rcu_tasks_qs_seq, cpu)) !=3D + per_cpu(rcu_tasks_qs_snap, cpu)) + return true; + /* + * Idle or nohz_full userspace in an RCU extended quiescent state: no + * task-level kernel frames can be live in a trampoline there, and + * whatever ran before has switched out. An idle CPU that RCU is + * watching (an interrupt from idle, or the traceable part of the idle + * loop) is deliberately not let through: the idle task may be in a + * trampoline with that interrupt on top, and since it is never + * preempted from irq exit nothing else would catch it. It gets the + * resched_cpu() like anyone else and counts once the idle loop itself + * schedules, which it cannot do from inside a trampoline. + */ + return !(ct_rcu_watching_cpu(cpu) & CT_RCU_WATCHING); +} + +/* Rate-limited stall report; returns true if the caller should add detail= . */ +static bool rcu_tasks_tramp_stall(struct rcu_tasks *rtp, unsigned long *la= streport, + const char *what) +{ + int rtst =3D READ_ONCE(rcu_task_stall_timeout); + + if (rtst <=3D 0 || !time_after(jiffies, *lastreport + rtst)) + return false; + *lastreport =3D jiffies; + pr_err("INFO: %s: %s, grace period %lu is %lu jiffies old\n", rtp->kname, + what, rcu_seq_current(&rtp->tasks_gp_seq), jiffies - rtp->gp_start= ); + return true; +} + +/* + * Steps 1/4: wait until every online CPU has context switched or is in an + * RCU extended quiescent state. A CPU that has neither after a jiffy is + * asked to switch with resched_cpu(), which takes it through + * rcu_tasks_irq_resched_enter() and __schedule() (or, from userspace, a + * guest or the idle loop, straight to __schedule()). + */ +static void rcu_tasks_tramp_wait_cpus(struct rcu_tasks *rtp, unsigned long= *lastreport) +{ + struct cpumask *pending =3D &rcu_tasks_pending_cpus; + unsigned long start; + int cpu; + + /* + * The quiescent events run with preemption (in practice interrupts) + * disabled, so after this any event we go on to count began after the + * caller's updates -- the unpublished trampoline, and whatever + * rcu_tasks_trampoline_text() consults -- were visible to it. + */ + synchronize_rcu(); + + start =3D jiffies; + for_each_online_cpu(cpu) { + per_cpu(rcu_tasks_qs_snap, cpu) =3D READ_ONCE(per_cpu(rcu_tasks_qs_seq, = cpu)); + __cpumask_set_cpu(cpu, pending); + } + /* Snapshots before the checks below; pairs with rcu_tasks_qs_event(). */ + smp_mb(); + + for (;;) { + for_each_cpu(cpu, pending) + if (rcu_tasks_cpu_quiescent(cpu)) + __cpumask_clear_cpu(cpu, pending); + if (cpumask_empty(pending)) + break; + if (time_after(jiffies, start)) { + for_each_cpu(cpu, pending) + resched_cpu(cpu); + rtp->n_ipis +=3D cpumask_weight(pending); + } + schedule_timeout_idle(1); + if (rcu_tasks_tramp_stall(rtp, lastreport, "CPUs without a quiescent eve= nt")) + pr_err("\tCPUs: %*pbl\n", cpumask_pr_args(pending)); + } +} + +/* + * Steps 2/4: wait for the tasks that were holdouts when we looked to stop + * being holdouts. They are moved to a private list so that tasks becoming + * holdouts later (in live trampolines) cannot keep us here; each removes + * itself via rcu_tasks_tramp_release() wherever it is queued. + */ +static void rcu_tasks_tramp_wait_holdouts(struct rcu_tasks *rtp, unsigned = long *lastreport) +{ + struct task_struct *t; + unsigned long flags; + int cpu; + + raw_spin_lock_irqsave(&rcu_tasks_tramp_lock, flags); + list_splice_tail_init(&rcu_tasks_tramp_holdouts, &rcu_tasks_gp_holdouts); + raw_spin_unlock_irqrestore(&rcu_tasks_tramp_lock, flags); + + for (;;) { + struct cpumask *kick =3D &rcu_tasks_pending_cpus; + struct task_struct *show[8]; + int nshow =3D 0, i; + bool empty, report; + + report =3D rcu_tasks_tramp_stall(rtp, lastreport, + "tasks preempted in trampoline text"); + cpumask_clear(kick); + raw_spin_lock_irqsave(&rcu_tasks_tramp_lock, flags); + empty =3D list_empty(&rcu_tasks_gp_holdouts); + list_for_each_entry(t, &rcu_tasks_gp_holdouts, rcu_tasks_holdout_list) { + if (task_curr(t)) + __cpumask_set_cpu(task_cpu(t), kick); + if (report && nshow < ARRAY_SIZE(show)) + show[nshow++] =3D get_task_struct(t); + } + raw_spin_unlock_irqrestore(&rcu_tasks_tramp_lock, flags); + /* Never printk under the lock the irq-exit path takes. */ + for (i =3D 0; i < nshow; i++) { + sched_show_task(show[i]); + put_task_struct(show[i]); + } + if (empty) + break; + for_each_cpu(cpu, kick) + resched_cpu(cpu); + rtp->n_ipis +=3D cpumask_weight(kick); + schedule_timeout_idle(1); + } +} + +/* Wait for one trampoline-reader Tasks RCU grace period. */ +static void rcu_tasks_tramp_wait_gp(struct rcu_tasks *rtp) +{ + unsigned long lastreport =3D jiffies; + + set_tasks_gp_state(rtp, RTGS_WAIT_SCAN_HOLDOUTS); + rcu_tasks_tramp_wait_cpus(rtp, &lastreport); + rcu_tasks_tramp_wait_holdouts(rtp, &lastreport); + + set_tasks_gp_state(rtp, RTGS_WAIT_READERS); + synchronize_rcu_tasks_trace(); + + set_tasks_gp_state(rtp, RTGS_SCAN_HOLDOUTS); + rcu_tasks_tramp_wait_cpus(rtp, &lastreport); + rcu_tasks_tramp_wait_holdouts(rtp, &lastreport); + + set_tasks_gp_state(rtp, RTGS_POST_GP); +} + +static int __init rcu_spawn_tasks_kthread(void) +{ + rcu_tasks.gp_sleep =3D HZ / 10; + if (rcu_tasks_lazy_ms >=3D 0) + rcu_tasks.lazy_jiffies =3D msecs_to_jiffies(rcu_tasks_lazy_ms); + rcu_tasks.wait_state =3D TASK_IDLE; + rcu_spawn_tasks_kthread_generic(&rcu_tasks); + return 0; +} + +void exit_tasks_rcu_start(void) { } +void exit_tasks_rcu_finish(void) { } + +#else /* #ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS */ + //////////////////////////////////////////////////////////////////////// // // Simple variant of RCU whose quiescent states are voluntary context @@ -1173,6 +1607,8 @@ static void tasks_rcu_exit_stall(struct timer_list *u= nused) #endif // #ifndef CONFIG_TINY_RCU } =20 +#endif /* #else #ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS */ + /** * call_rcu_tasks() - Queue an RCU for invocation task-based grace period * @rhp: structure to be used for queueing the RCU updates. @@ -1187,6 +1623,12 @@ static void tasks_rcu_exit_stall(struct timer_list *= unused) * primitives analogous to rcu_read_lock() and rcu_read_unlock() because * this primitive is intended to determine that all tasks have passed * through a safe state, not so much for data-structure synchronization. + * On CONFIG_TASKS_RCU_TRAMPOLINE_READERS kernels a preemption outside + * trampoline text also ends one, and a reader whose protected window + * spans preemptible code must additionally be a Tasks Trace RCU reader + * (rcu_read_lock_trace(), as the trampolines there take around their + * call-outs); an arbitrary stretch of preemptible kernel code is not + * protected. * * See the description of call_rcu() for more detailed information on * memory ordering guarantees. @@ -1205,7 +1647,9 @@ EXPORT_SYMBOL_GPL(call_rcu_tasks); * executing rcu-tasks read-side critical sections have elapsed. These * read-side critical sections are delimited by calls to schedule(), * cond_resched_tasks_rcu_qs(), idle execution, userspace execution, calls - * to synchronize_rcu_tasks(), and (in theory, anyway) cond_resched(). + * to synchronize_rcu_tasks(), and (in theory, anyway) cond_resched(); + * on CONFIG_TASKS_RCU_TRAMPOLINE_READERS kernels also by preemption + * outside trampoline text, see call_rcu_tasks(). * * This is a very specialized primitive, intended only for a few uses in * tracing and other situations requiring manipulation of function @@ -1233,9 +1677,7 @@ void rcu_barrier_tasks(void) } EXPORT_SYMBOL_GPL(rcu_barrier_tasks); =20 -static int rcu_tasks_lazy_ms =3D -1; -module_param(rcu_tasks_lazy_ms, int, 0444); - +#ifndef CONFIG_TASKS_RCU_TRAMPOLINE_READERS static int __init rcu_spawn_tasks_kthread(void) { rcu_tasks.gp_sleep =3D HZ / 10; @@ -1251,6 +1693,7 @@ static int __init rcu_spawn_tasks_kthread(void) rcu_spawn_tasks_kthread_generic(&rcu_tasks); return 0; } +#endif /* #ifndef CONFIG_TASKS_RCU_TRAMPOLINE_READERS */ =20 #if !defined(CONFIG_TINY_RCU) void show_rcu_tasks_classic_gp_kthread(void) @@ -1279,6 +1722,7 @@ void rcu_tasks_get_gp_data(int *flags, unsigned long = *gp_seq) } EXPORT_SYMBOL_GPL(rcu_tasks_get_gp_data); =20 +#ifndef CONFIG_TASKS_RCU_TRAMPOLINE_READERS /* * Protect against tasklist scan blind spot while the task is exiting and * may be removed from the tasklist. Do this by adding the task to yet @@ -1322,6 +1766,7 @@ void exit_tasks_rcu_finish(void) list_del_init(&t->rcu_tasks_exit_list); raw_spin_unlock_irqrestore_rcu_node(rtpcp, flags); } +#endif /* #ifndef CONFIG_TASKS_RCU_TRAMPOLINE_READERS */ =20 #else /* #ifdef CONFIG_TASKS_RCU */ void exit_tasks_rcu_start(void) { } diff --git a/kernel/rcu/update.c b/kernel/rcu/update.c index b62735a67884..a122b8d1effb 100644 --- a/kernel/rcu/update.c +++ b/kernel/rcu/update.c @@ -40,7 +40,9 @@ #include #include #include +#include #include +#include #include #include #include --=20 2.55.0 From nobody Thu Sep 24 17:09:51 2026 Received: from mail-qk2-f12.google.com (mail-qk2-f12.google.com [74.125.230.204]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 434A238C410 for ; Tue, 22 Sep 2026 02:26:46 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.204 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790044009; cv=none; b=hJ6VuC+OPpyOEvsNxV8h/fau65qk2ImicnPmzay1dCE2L1nvgqIgO4Q7Gcvykz36tr+m9GgJ1AVbMAa8StTaWRhMVZ6YxAYxpSIPpis1dNAc7irgDy6yqHxBYawN6dOMDzlLyF9VkhACIOAGnInx73SDOKiSJ1wRimV3jp5aXFA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790044009; c=relaxed/simple; bh=97rXNEsO0S02JFiFJq38moNkWCht7E9TlyxYtLTVeYI=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=RZPovecKHd7Vu8j3WjP6xqHsRmxo9suxwaRMQDvxnajt2yXLNAC0VJoSEVS8OHd6PsfgUhftNmshPe5sYRsA42D69e4hEKoyKTghCzY6Z2MMxAubnHbU1Cq2OIC8H8Pt84TOpgABDO8jnABV6qEb/Y2eam/o4QPgPZ6mjf4Mgek= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com; spf=pass smtp.mailfrom=toxicpanda.com; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b=XSRleOQg; arc=none smtp.client-ip=74.125.230.204 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b="XSRleOQg" Received: by mail-qk2-f12.google.com with SMTP id af79cd13be357-93910cadeb7so369118785a.0 for ; Mon, 21 Sep 2026 19:26:46 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=toxicpanda.com; s=google; t=1790044006; x=1790648806; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=n57jpnbqR3xEpFvVL4pj+x/9TynCnpwMD9+qM+zc540=; b=XSRleOQgjUQzCCbGEqntpF1vr5U/KVWqAt+XGZTZyL7fS/CGUQWvFYjOql/DtI5/ub Nmg95FYAnqGSAU1R/dqqGGRjGJMTw12mQNR+E++vzI+/sbf8ai7wIXNQ+Lm0KJNIjGFj H9bd+G2Bltwo/mRIntAlADHD8ps41xJql+bya1+EyOFPjE6skdbc/Chdbp0DUN3rgl2l d+Zkv5KxY27DLAV4SoE8EhalcodDzL+OsYpn+N8YNQobrt7sFVqNBQPn/l4IKKv0P43O 5O1XnoEd8Potdb0IatO8OmwzK0n4U13eXoXdtpodtRQUJcaoWuiKyygELPsQcC/1d06D jgZg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790044006; x=1790648806; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=n57jpnbqR3xEpFvVL4pj+x/9TynCnpwMD9+qM+zc540=; b=ESskqiC9fdllwcg4wZz4UddRi1IApqvFHpDi2MVEi49Nyvswqn15cpDVSTgNWqx/ht n87W85ISYopEVgkm2rDe4skDNQh4W/i/x46+GrxZPo6/My2LRRhZuGHLVENx4qUA9r5m xHM8bEMQeQ3F75jZLrerxNw01MdHgTSiY3suToIbBDoaIpvEb4MuBr+QEf7kyxxnXEeX j4zXoCaZ8zjOVUFVb4MD8vilZIdthNlkSoXBZNyHvQNb1v3drxHzyabMFWjzck2hfmoB E8wOijFV0mqnJi2m9Ft3Dq9MSAvOXuVtSA6cb8uN4CW+JF+r1Xype1V5spLy8UX9aj2y FY1Q== X-Forwarded-Encrypted: i=1; AKwUvBw3BT5VLlywodXsS5CbrML4zQTV2JFqlnCCQnYqdyGR/ekQI7Dq0kG18lkhEKyZ+iuAJw+pXctUFE4dKu8=@vger.kernel.org X-Gm-Message-State: AFuF++myzXiskIWLSGz2RHR38MWkmU9yqCI7WtVl4DMTHnJDp18nVjmf y2j8RgWH8MgN6rz1NMZviodFDNxHXZCgQXMda6aG4PYaH65H2EeiBvPGkTSQ1wj8s8M= X-Gm-Gg: AYBFou0nZlwHOY25ha/SHLSkf4eM4ehFHbqKHORPSIHFcMmpjQTm7YAutjngdGug/Tl IjwUq+AUDz05BVnN4cEPOlctVHh/m8sS+xUe/eJzmEkDMIJJYARrg54gNQfS/tgDb6wlwa6PCY6 c6Gf3BBCKNLEmid8xPgt9ndgebclyNVsnNATw4PcPNatQ/8/IiDcBlDX+qCYI0aqFm0ckxybvz8 qMtesbVzL6qBbXZen85P/L1695ex/96jBj5uF0XVLuIRbWu7BphDZvum1I624kKmdksIbAaRo4O hs+hY/xjPUWdmxzio+iHxJDCR9rjVVugylQcpc7gUIDyBHBFPBUIOk1hJsXxEP5krYj2sdtv1Gz 31YGAWeTSUj3PNvdq6dg0fmTFHFvCpZxHeqjiCnyTyoqy5OkyhlLThcs3Dq7NGE9zWOZu0gbU2e cr5K0FaqcSvZN50gAYsRZJxep3NZGJYraoVXAE4gLrDQCoSUv+iBEYxugK8UL5ihWEXmpDn21dd XU+4YI8mLyL3xU= X-Received: by 2002:a05:620a:458f:b0:939:bc7c:15c8 with SMTP id af79cd13be357-93c15e783eamr375359285a.45.1790044005713; Mon, 21 Sep 2026 19:26:45 -0700 (PDT) Received: from toxicpanda.com ([153.61.196.251]) by smtp.gmail.com with ESMTPSA id af79cd13be357-93c1d131c4dsm21683685a.15.2026.09.21.19.26.44 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 21 Sep 2026 19:26:45 -0700 (PDT) From: Josef Bacik Date: Tue, 22 Sep 2026 02:23:22 +0000 Subject: [PATCH v5 03/13] kprobes: Expose the optprobe jump window to Tasks RCU Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260922-b4-rcu-tasks-preempt-qs-v5-3-410f57770bad@toxicpanda.com> References: <20260922-b4-rcu-tasks-preempt-qs-v5-0-410f57770bad@toxicpanda.com> In-Reply-To: <20260922-b4-rcu-tasks-preempt-qs-v5-0-410f57770bad@toxicpanda.com> To: "Paul E. McKenney" , Frederic Weisbecker , Alexei Starovoitov , Steven Rostedt Cc: Boqun Feng , Masami Hiramatsu , Mark Rutland , Peter Zijlstra , Thomas Gleixner , Daniel Borkmann , Andrii Nakryiko , Puranjay Mohan , rcu@vger.kernel.org, bpf@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, Josef Bacik X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openssh-sha256; t=1790043820; l=7473; i=josef@toxicpanda.com; h=from:subject:message-id; bh=97rXNEsO0S02JFiFJq38moNkWCht7E9TlyxYtLTVeYI=; b=U1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgUBr36M/n0nWN0DNbnxwzIiCZez6MG JiruuNaSCI/zXsAAAAGcGF0YXR0AAAAAAAAAAZzaGE1MTIAAABTAAAAC3NzaC1lZDI1NTE5AAAA QDgns7oV6kZusWmA4JvFz3vXN+lrAHSlz7TH5kJB91LDL9uOpEUIYQaSIenbGihoBc2IuMs6vNX 0WqfawsK3hAU= X-Developer-Key: i=josef@toxicpanda.com; a=openssh; fpr=SHA256:C8kOX2QUJCMqnCX+KEeoqRAjLo9L+ELOSH2NSAJHqGA kprobe_optimizer() is the one synchronize_rcu_tasks() user that is not about trampoline text: it waits for tasks that were interrupted on an instruction boundary inside the bytes it is about to overwrite with the optimized jump, so that none of them resumes into the middle of the new instruction. Those bytes are ordinary kernel or module text with no Tasks Trace reader around them, so on CONFIG_TASKS_RCU_TRAMPOLINE_READERS kernels the irq-exit quiescent-state check has to be told about them. Add kprobe_in_optimized_region(), a lockless and conservative form of get_optimized_kprobe() that reports whether any registered kprobe lies within MAX_OPTIMIZED_LENGTH before the given address regardless of its optimization state, and have rcu_tasks_trampoline_text() consult it for core and module text so that a task interrupted there becomes a holdout rather than a quiescent event. The hash walk only runs while kprobe_optimizer() is actually inside its synchronize_rcu_tasks(), tracked by a flag it sets around the call; otherwise the check is a single load. That check cannot see a task that was already preempted in the region before the flag went up (possibly before the kprobe even existed), and the new grace period does not otherwise wait for a preempted task to run again, so before synchronize_rcu_tasks() the optimizer calls rcu_tasks_wait_irq_preempted() to wait until no parked task's recorded irq-exit preemption IP is inside such a region; its leading synchronize_rcu() also publishes the flag to every (interrupts- disabled) check in flight. The kprobe hash is RCU-protected and every free path waits for a grace period after unhashing, so the lockless walk from the irq-exit path is safe. On other configurations the flag is set and cleared but nothing reads it and rcu_tasks_wait_irq_preempted() is a stub; the classic implementation already waits for such tasks. Assisted-by: LLM Signed-off-by: Josef Bacik --- include/linux/kprobes.h | 8 +++++++- kernel/kprobes.c | 50 +++++++++++++++++++++++++++++++++++++++++++++= ++++ kernel/rcu/tasks.h | 11 ++++++++--- 3 files changed, 65 insertions(+), 4 deletions(-) diff --git a/include/linux/kprobes.h b/include/linux/kprobes.h index e6de7ae55bda..74cc48c04417 100644 --- a/include/linux/kprobes.h +++ b/include/linux/kprobes.h @@ -530,11 +530,17 @@ static inline bool is_kprobe_insn_slot(unsigned long = addr) } #endif /* !CONFIG_KPROBES */ =20 -#ifndef CONFIG_OPTPROBES +#ifdef CONFIG_OPTPROBES +bool kprobe_in_optimized_region(unsigned long addr); +#else /* !CONFIG_OPTPROBES */ static inline bool is_kprobe_optinsn_slot(unsigned long addr) { return false; } +static inline bool kprobe_in_optimized_region(unsigned long addr) +{ + return false; +} #endif /* !CONFIG_OPTPROBES */ =20 #ifdef CONFIG_KRETPROBES diff --git a/kernel/kprobes.c b/kernel/kprobes.c index 6337da5cab9e..e460fba83e4a 100644 --- a/kernel/kprobes.c +++ b/kernel/kprobes.c @@ -511,6 +511,48 @@ static struct kprobe *get_optimized_kprobe(kprobe_opco= de_t *addr) return NULL; } =20 +/* + * True while kprobe_optimizer() is waiting for its Tasks RCU grace period. + * Only in that window can an interruption inside an optprobe's jump region + * matter to it, so kprobe_in_optimized_region() does no work otherwise. + */ +static bool kprobe_optimizer_waiting; + +/** + * kprobe_in_optimized_region - Could @addr be inside bytes a jump-optimiz= ed + * kprobe replaces? + * @addr: kernel text address, typically an interrupted instruction pointer + * + * kprobe_optimizer() relies on synchronize_rcu_tasks() to wait for tasks = that + * were interrupted on an instruction boundary inside the region about to = be + * overwritten by the optimized jump. Where Tasks RCU is built on + * reader-marked trampolines that region has no reader, so the irq-exit + * quiescent-state check asks this instead (see rcu_tasks_trampoline_text(= )). + * This is the lockless, conservative form of get_optimized_kprobe(): it d= oes + * not care whether the kprobe found is, or ever will be, optimized. May = be + * called from any context with preemption disabled; the kprobe hash is + * RCU-protected and every free path waits for a grace period after unhash= ing. + * + * The hash walk only runs while the optimizer is actually waiting. A task + * that was preempted in such a region before the flag went up is invisible + * to that check, so the optimizer first waits those out by their recorded + * preemption IP (rcu_tasks_wait_irq_preempted(), whose leading + * synchronize_rcu() also publishes the flag to every check in flight). + */ +bool kprobe_in_optimized_region(unsigned long addr) +{ + int i; + + if (!READ_ONCE(kprobe_optimizer_waiting)) + return false; + + for (i =3D 1; i < MAX_OPTIMIZED_LENGTH / sizeof(kprobe_opcode_t); i++) + if (get_kprobe((kprobe_opcode_t *)addr - i)) + return true; + return false; +} +NOKPROBE_SYMBOL(kprobe_in_optimized_region); + /* Optimization staging list, protected by 'kprobe_mutex' */ static LIST_HEAD(optimizing_list); static LIST_HEAD(unoptimizing_list); @@ -644,8 +686,16 @@ static void kprobe_optimizer(void) * to 2nd-Nth byte of jump instruction. This wait is for avoiding it. * Note that on non-preemptive kernel, this is transparently converted * to synchronoze_sched() to wait for all interrupts to have completed. + * kprobe_optimizer_waiting lets a reader-marked-trampoline Tasks RCU + * recognise tasks interrupted in such a region while we wait, and + * rcu_tasks_wait_irq_preempted() (a no-op elsewhere) first waits + * out any that were preempted there before we said so; see + * kprobe_in_optimized_region(). */ + WRITE_ONCE(kprobe_optimizer_waiting, true); + rcu_tasks_wait_irq_preempted(kprobe_in_optimized_region); synchronize_rcu_tasks(); + WRITE_ONCE(kprobe_optimizer_waiting, false); =20 /* Step 3: Optimize kprobes after quiesence period */ do_optimize_kprobes(); diff --git a/kernel/rcu/tasks.h b/kernel/rcu/tasks.h index f03be742be48..eb1388dd8a61 100644 --- a/kernel/rcu/tasks.h +++ b/kernel/rcu/tasks.h @@ -1005,7 +1005,9 @@ bool __weak arch_rcu_tasks_trampoline_text(unsigned l= ong ip) * trampolines, kprobe slots and other dynamically allocated text; this * deliberately does not ask is_ftrace_trampoline() and friends, since * text being torn down may already be unregistered there); - * - whatever the architecture adds via arch_rcu_tasks_trampoline_text(). + * - whatever the architecture adds via arch_rcu_tasks_trampoline_text(); + * - the bytes after a kprobe that a pending jump optimization is about to + * overwrite, the one synchronize_rcu_tasks() user with no trampoline. * * A false positive only makes the task a holdout until its next quiescent * event. Called with interrupts disabled from the irq-exit path. @@ -1013,8 +1015,11 @@ bool __weak arch_rcu_tasks_trampoline_text(unsigned = long ip) bool rcu_tasks_trampoline_text(unsigned long ip) { if (core_kernel_text(ip)) - return arch_rcu_tasks_trampoline_text(ip); - return !is_module_text_address(ip); + return arch_rcu_tasks_trampoline_text(ip) || + kprobe_in_optimized_region(ip); + if (is_module_text_address(ip)) + return kprobe_in_optimized_region(ip); + return true; } NOKPROBE_SYMBOL(rcu_tasks_trampoline_text); =20 --=20 2.55.0 From nobody Thu Sep 24 17:09:51 2026 Received: from mail-qk2-f42.google.com (mail-qk2-f42.google.com [74.125.230.234]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2F1F538A715 for ; Tue, 22 Sep 2026 02:27:54 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.234 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790044077; cv=none; b=H0FxkxYm+uG4TIdKcxw65Vtxgjb8yzyPKUFZV/6D+It3YeucQljBJ8pp2V1nsUl3izJKS1c46qx+DX3o3PMuLVVAtDvs+LV03C1lQm5zJ6FTOcbnCN1zCgz2JCzRuc6ZwE8l2PGwMVmVQ9GrG6joYbQ/361d71Nvv+ADeSO3xkA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790044077; c=relaxed/simple; bh=WxuURWfyPfVrbwoJYzsme+dJrO0g2fOexfelnDzg3io=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=EqQLkJgPsg2Pu9SK9AVqWUaOQqKb+fG731pQapkH0uCQj04Z+IKq3lkzpfot03qdI43UxeWCF2sk9JFix3uSvoJnRSQQeDK9gyZLI6p1YM5EB84eaJAjUWcTZTJsbqQoCvO6tJCq1Ocnc8Q1Jyof0Ta2Ag8yl7WHRFUibkje6pU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com; spf=pass smtp.mailfrom=toxicpanda.com; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b=I6lHTrEY; arc=none smtp.client-ip=74.125.230.234 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b="I6lHTrEY" Received: by mail-qk2-f42.google.com with SMTP id af79cd13be357-93be29bb454so412727385a.1 for ; Mon, 21 Sep 2026 19:27:54 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=toxicpanda.com; s=google; t=1790044073; x=1790648873; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=cjEABplA81VfGcXS6Z1Q3Y7Xt5Unb50ofxErtTl6HPY=; b=I6lHTrEYx3XIEou9LiLDPR92sFcFypxuNY4FR11g5RVmVFmEJXnA2bKGJ3S1IoKiR6 rmOgW7sNNl4wqFHPc0yQDeKQtMF9U7+ywjV3lOzf7T1gpR4Y4laOsQ0egVCyGyXuLV1P LqQoGUW05qT6lN6UTmJKjG/T24ScCPmKBqunbRX4ZWESiyp0gUaZCBmxibLwdEFZbXvq 8DlU3zinQ9xOlM012tgY9qR1n46KQ0bsGJGiRcI3YpAF6gQoSQgR8IMJnKcBWk+qLf7O HYxNr7MbyyPfHCjuvTZHA+ns8o/vjiMcb2e6ZSlRyP7BADorS/YI+lD5hQ/KuFpXUB+J 5IeA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790044073; x=1790648873; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=cjEABplA81VfGcXS6Z1Q3Y7Xt5Unb50ofxErtTl6HPY=; b=dJLIMyfc2ZBRsdKI8Cmc+oAgv67FDaFVBtXp6S5dhBAggcuaygI/erY4emZjzci1mF Vk98U0u9dOEKB9xrZzjDTc0CClDgECY+6wygwDAK74KpNXg9qxnnTkWqkNBH3JwbTOMx j4QFFnUXqpZK9Y8vP2rBNTK0hHrnYHuUI3k0iX344zKsGpc9B1YtynF+qH2oWg2KiQBi 1Q6usXscmyxE8GPKowS5YQdbXqAChg/ykJMsZte82Ftc85eKSaLRioTbaGR9Brpv0dEn 7X/du3F5XceSjkO5bg/o6Hs14OMQmMOhCneuVcM9y/CG/ZlNSCUO+BMwnl1O3k7jV5xx Ae0w== X-Forwarded-Encrypted: i=1; AKwUvBxs9UwiyqeONRJ2q1ZEPZO5QLjTENWXElHYz0qAeSltk1kn7b08CHeOta7grBnyXHWbI7DHUJBvKG22ChY=@vger.kernel.org X-Gm-Message-State: AFuF++nX8bpi6vEyj3J/m3macVJ2EmJU7Ixi7EugJJPt3eOactumy249 VwlxTeQ0gQzCOvAse3mQCt31WJKWN1+gaJuJYW9MIGa6VFRAL4Mz96TKRD8muogv29I= X-Gm-Gg: AYBFou3WFaVcwsGiTIfanFtZ0PuC2n4vtIzYc1JF5awOaRIEXZclwL57SksYWV8oLiD Tk/KKieWfm2hUOlxaA8MWefrhRw/HZ1V3ba57GlNv/8HZ/eeTToRh20vs7AKQumghKfa6d0uExX FKXHShYZOVUHK89V+lPQC3ZuRNeOkjSTYkyopN0zkWOF4avQHBm7P8sD7aRLu61w4Sl5ApcqROa vXCNMVUHM5HFOxXt8Uxbo+wOvbnTlJ0wqJcTMMFKpuDz15S56lu0Hgwtu4rf33mMhJ/7pbm6ibm yRCQ6B+xWv9CpoCRYvqMfqCLCB1Y1xXPf9HSIppKc36s9RyNNPbKJLoRxY0J6RVpijSYTl95p9a zdJ91v6nlqxiwCV5PXw9aH5wNoPtj+hdBEGI3c5/VxLuYYpUpa94qq5fOLjnSLONa0q4HxoidVc VzsubXA9jAQ7f7DnN19tGef6xK1fLYxXhVphJGLngxG21ouaTSfLgwmFPvkrR48olBLkdCASjpx WNJBUdlaSNlqL8= X-Received: by 2002:a05:620a:2949:b0:93b:d7a0:d9dd with SMTP id af79cd13be357-93c19a67345mr151109685a.55.1790044073015; Mon, 21 Sep 2026 19:27:53 -0700 (PDT) Received: from toxicpanda.com ([153.61.196.251]) by smtp.gmail.com with ESMTPSA id af79cd13be357-93c1d1e9c78sm21071585a.32.2026.09.21.19.27.52 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 21 Sep 2026 19:27:52 -0700 (PDT) From: Josef Bacik Date: Tue, 22 Sep 2026 02:23:23 +0000 Subject: [PATCH v5 04/13] ftrace: Mark modules hosting direct-call trampolines for Tasks RCU Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260922-b4-rcu-tasks-preempt-qs-v5-4-410f57770bad@toxicpanda.com> References: <20260922-b4-rcu-tasks-preempt-qs-v5-0-410f57770bad@toxicpanda.com> In-Reply-To: <20260922-b4-rcu-tasks-preempt-qs-v5-0-410f57770bad@toxicpanda.com> To: "Paul E. McKenney" , Frederic Weisbecker , Alexei Starovoitov , Steven Rostedt Cc: Boqun Feng , Masami Hiramatsu , Mark Rutland , Peter Zijlstra , Thomas Gleixner , Daniel Borkmann , Andrii Nakryiko , Puranjay Mohan , rcu@vger.kernel.org, bpf@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, Josef Bacik X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openssh-sha256; t=1790043820; l=7065; i=josef@toxicpanda.com; h=from:subject:message-id; bh=WxuURWfyPfVrbwoJYzsme+dJrO0g2fOexfelnDzg3io=; b=U1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgUBr36M/n0nWN0DNbnxwzIiCZez6MG JiruuNaSCI/zXsAAAAGcGF0YXR0AAAAAAAAAAZzaGE1MTIAAABTAAAAC3NzaC1lZDI1NTE5AAAA QIG8cGFVLUxRckCJD2FMFcjxlH7HHUuPFKzc+YD5qvFJCI6wuoTYH6RLHGPqY8+XII71ZLcy6t+ Vmn6AcyiciAU= X-Developer-Key: i=josef@toxicpanda.com; a=openssh; fpr=SHA256:C8kOX2QUJCMqnCX+KEeoqRAjLo9L+ELOSH2NSAJHqGA An out-of-line direct trampoline registered with register_ftrace_direct() is kept alive only by Tasks RCU while a task executes it or is preempted in something it called; ftrace_shutdown()'s synchronize_rcu_tasks() is what stops rmmod freeing it under such a task. Where Tasks RCU is built on reader-marked trampolines, such a trampoline must be a Tasks Trace reader across its call-out like the ftrace and BPF trampolines are, so document that in register_ftrace_direct(). That still leaves the few instructions before the reader is entered and after it is left. For BPF images those are in dynamically allocated text that rcu_tasks_trampoline_text() already treats as unmarked trampoline text, but the in-tree samples (and any similar user) place their trampolines in module .text. Add a sticky module::ftrace_direct_tramp flag (under CONFIG_TASKS_RCU_TRAMPOLINE_READERS, its only consumer), set by every register/modify path when the direct address is module text, and have rcu_tasks_trampoline_text() treat a task interrupted anywhere in such a module as a potential holdout. Other modules' text is unaffected. Assisted-by: LLM Signed-off-by: Josef Bacik --- include/linux/module.h | 7 +++++++ kernel/rcu/tasks.h | 18 +++++++++++++++--- kernel/trace/ftrace.c | 39 +++++++++++++++++++++++++++++++++++++++ 3 files changed, 61 insertions(+), 3 deletions(-) diff --git a/include/linux/module.h b/include/linux/module.h index 96cc98568eea..82ca4f774725 100644 --- a/include/linux/module.h +++ b/include/linux/module.h @@ -521,6 +521,13 @@ struct module { unsigned int num_ftrace_callsites; unsigned long *ftrace_callsites; #endif +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS + /* + * An ftrace direct-call trampoline lives in this module's text; see + * rcu_tasks_trampoline_text(). Sticky once set. + */ + bool ftrace_direct_tramp; +#endif #ifdef CONFIG_KPROBES void *kprobes_text_start; unsigned int kprobes_text_size; diff --git a/kernel/rcu/tasks.h b/kernel/rcu/tasks.h index eb1388dd8a61..42ea6e0e61cb 100644 --- a/kernel/rcu/tasks.h +++ b/kernel/rcu/tasks.h @@ -1006,6 +1006,8 @@ bool __weak arch_rcu_tasks_trampoline_text(unsigned l= ong ip) * deliberately does not ask is_ftrace_trampoline() and friends, since * text being torn down may already be unregistered there); * - whatever the architecture adds via arch_rcu_tasks_trampoline_text(); + * - the text of a module that hosts an out-of-line ftrace direct-call + * trampoline (see ftrace_direct_mark_module()); * - the bytes after a kprobe that a pending jump optimization is about to * overwrite, the one synchronize_rcu_tasks() user with no trampoline. * @@ -1014,12 +1016,22 @@ bool __weak arch_rcu_tasks_trampoline_text(unsigned= long ip) */ bool rcu_tasks_trampoline_text(unsigned long ip) { + bool ret =3D true; + if (core_kernel_text(ip)) return arch_rcu_tasks_trampoline_text(ip) || kprobe_in_optimized_region(ip); - if (is_module_text_address(ip)) - return kprobe_in_optimized_region(ip); - return true; + +#ifdef CONFIG_MODULES + scoped_guard(rcu) { + struct module *mod =3D __module_text_address(ip); + + if (mod) + ret =3D READ_ONCE(mod->ftrace_direct_tramp) || + kprobe_in_optimized_region(ip); + } +#endif + return ret; } NOKPROBE_SYMBOL(rcu_tasks_trampoline_text); =20 diff --git a/kernel/trace/ftrace.c b/kernel/trace/ftrace.c index 53d5db60bfa5..f69f71591358 100644 --- a/kernel/trace/ftrace.c +++ b/kernel/trace/ftrace.c @@ -6076,6 +6076,29 @@ static void reset_direct(struct ftrace_ops *ops, uns= igned long addr) ops->trampoline =3D 0; } =20 +/* + * A direct trampoline may live in module text rather than in dynamically + * allocated text that rcu_tasks_trampoline_text() recognises on its own (= see + * samples/ftrace/ftrace-direct*.c). The trampoline itself must be a Tasks + * Trace reader across its call-out (see register_ftrace_direct()); markin= g the + * owning module here covers the instructions before it enters that reader= and + * after it leaves it, where a task interrupted in the module's text must = not be + * counted as Tasks-RCU quiescent, so that ftrace_shutdown()'s + * synchronize_rcu_tasks() still keeps the module text from being freed un= der + * it. + */ +static void ftrace_direct_mark_module(unsigned long addr) +{ +#if defined(CONFIG_MODULES) && defined(CONFIG_TASKS_RCU_TRAMPOLINE_READERS) + struct module *mod; + + guard(rcu)(); + mod =3D __module_text_address(addr); + if (mod) + WRITE_ONCE(mod->ftrace_direct_tramp, true); +#endif +} + /** * register_ftrace_direct - Call a custom trampoline directly * for multiple functions registered in @ops @@ -6090,6 +6113,17 @@ static void reset_direct(struct ftrace_ops *ops, uns= igned long addr) * and save the parameters of the function being traced, and restore them * (or inject new ones if needed), before returning. * + * Nothing but Tasks RCU keeps the trampoline at @addr alive while a task = is + * executing it or is preempted in something it called. On architectures = that + * select HAVE_RCU_TRAMPOLINE_READERS, Tasks RCU only waits for such a tas= k if + * it is a Tasks Trace RCU reader, so the trampoline must enter one + * (rcu_read_lock_trace() or an assembly equivalent) before calling out and + * leave it before returning, just as that option requires of the in-kernel + * ftrace and BPF trampolines. The few instructions before and after are + * covered by the irq-exit check: automatically for trampolines outside ke= rnel + * and module text (e.g. BPF images), and via ftrace_direct_mark_module() = for + * trampolines in module text. + * * Returns: * 0 on success * -EINVAL - The @ops object was already registered with this call or @@ -6169,6 +6203,7 @@ int register_ftrace_direct(struct ftrace_ops *ops, un= signed long addr) ops->flags |=3D MULTI_FLAGS; ops->trampoline =3D FTRACE_REGS_ADDR; ops->direct_call =3D addr; + ftrace_direct_mark_module(addr); =20 err =3D register_ftrace_function_nolock(ops); if (err) @@ -6237,6 +6272,8 @@ __modify_ftrace_direct(struct ftrace_ops *ops, unsign= ed long addr) =20 lockdep_assert_held_once(&direct_mutex); =20 + ftrace_direct_mark_module(addr); + /* Enable the tmp_ops to have the same functions as the direct ops */ ftrace_ops_init(&tmp_ops); tmp_ops.func_hash =3D ops->func_hash; @@ -6419,6 +6456,7 @@ int update_ftrace_direct_add(struct ftrace_ops *ops, = struct ftrace_hash *hash) hlist_for_each_entry(entry, &hash->buckets[i], hlist) { if (__ftrace_lookup_ip(direct_functions, entry->ip)) goto out_unlock; + ftrace_direct_mark_module(entry->direct); } } =20 @@ -6702,6 +6740,7 @@ int update_ftrace_direct_mod(struct ftrace_ops *ops, = struct ftrace_hash *hash, b tmp =3D __ftrace_lookup_ip(direct_hash, entry->ip); if (!tmp) continue; + ftrace_direct_mark_module(entry->direct); tmp->direct =3D entry->direct; } } --=20 2.55.0 From nobody Thu Sep 24 17:09:51 2026 Received: from mail-qv2-f12.google.com (mail-qv2-f12.google.com [74.125.230.140]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0C05038F64D for ; Tue, 22 Sep 2026 02:28:33 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.140 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790044116; cv=none; b=AyV5gxE4FUGDOyeJOgswdHZ3+OuKl0GtRMD8M93UAzpTf9d4qv41KqknOHfW/grJ6y03w+/AGL7jH7AXoLM+BpeTCJk8CL99HnZmpMIZPcCUnTTs54s6usZDwAFyYjXNAyUsircQXEnRkhXX6FXcsY4/hqrUuHnIOFK2l6WIzDI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790044116; c=relaxed/simple; bh=1mhltqxNnCEQ1bcCikw9zgimeks9K7E+TczS0kwNUMw=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=ef4gmKOmed+JPBsbMp4A9HNifiXfAhlUS51ede3ttBRAX68xf5B6uoaqDfYoahMbb6UzeeI0mNQO5MOYQ1e4Hmksp61L5tGqvD7xQYF7FXKi/19BgNiVGiVGZAMKbNiKKEr7+AkrmD9t53qDxgBUZimRHBc8iDroVPfsXdGOzpw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com; spf=pass smtp.mailfrom=toxicpanda.com; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b=gOS4p/MZ; arc=none smtp.client-ip=74.125.230.140 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b="gOS4p/MZ" Received: by mail-qv2-f12.google.com with SMTP id 6a1803df08f44-90cdfcb5cb3so39429256d6.3 for ; Mon, 21 Sep 2026 19:28:33 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=toxicpanda.com; s=google; t=1790044112; x=1790648912; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=HscWMACQBaXdyi2sA4ZdN71wnPrvQNOWSxMsTvhV6Ow=; b=gOS4p/MZxN1xotHEZgcAjYlBtny/eJM1fu/EfC3N+5dfz44x8oYkflUTC56UG8BhOe iRqhZ3r2v1DKXiSj02Hc56rqiQ4otMzMDvq71roZDyEQwXRpsjaUwZ+Xc6nCovjS73rw CxpOHZlMAxETLLUq5c0sk10OVpHBsErS2pWblEkNbSHmwoPVi5ZiO7m+r1MEM71Q6wwI 8gd414eq04o96iajy0l2mQ7qgO3hzRWrKShkVNOR6jqkYaPcfYwrv3ZYf3oRs5jcUHYX 8kMWT4ytye4vN2AWqiLb7SbJWAV8bc9uW0OabkQGx8TbJNBS99dHCq4nREw0jl5kvbtE QxPQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790044112; x=1790648912; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=HscWMACQBaXdyi2sA4ZdN71wnPrvQNOWSxMsTvhV6Ow=; b=dzm1rK/5afx41WSD89R67DfdmtWcUhJen6QWwTwAxVJLKNCXUjeAsEtl2+GwhBfUfb 9Kpn8SFX3FDPhf/ijVkq5qVMI0ZCgFHnZTcq8Qt7Oau7E3HVYpyqiA6LJ+6cYZnVwV7t G2rTySou/ZEJ8tH6BUDf74+vcizRs4g+/k8x0acg2Bi/Zjg0Ie5jPelXcJUEzsItsoyM mV5Fq+IHdQ3SMvXBzOSove1ZHVJD+I5kvcYs85fVIf6tZAd26oZ01Q8LnV1av1S7wRsa 4du06Hcdli6WjVDFiE/nn1sOwZW7N6mNnT8k+luFQ0A2qLhrY0BiYu7y+mJodB4FU2p2 mIUg== X-Forwarded-Encrypted: i=1; AKwUvByw3WtFPJ0OFrhn0JyPTGoJp9WF4MfsaTaiCJMSxqWxW3+vNxYOrt5Xxk6Ggzc/SASoKnLZyN7ezZtpKU8=@vger.kernel.org X-Gm-Message-State: AFuF++nolaaA1swrmKJoZdurKHwpLVK+AywH35Rqk/JxEU2quWrevn1x dT8UVFMBZz0HWGBr8N7OwOjb27rTSmoe3b4g0A1iEG2DJGsSv1sifhMKdVVQBJvBK14= X-Gm-Gg: AYBFou0FZJCjOnEIzC7AG1KXwGfhmAMtjHoY2jr5cqkYBcFecLYxXAPoNkoOAqz4cgS /9ONkd+FcXs4dzKtyLCorBjPcctPtiOIKr7y/ZQOdGhUHzDZsDFo9afBeQQq5zFklBrbAeaVJMj tNfcMLEY5Yc2uUyEyLNdIfLi1EZvJXv8jgyNo2HT5OVIzoqZRMWte4GMD/10EHMPq/ygGNo60Yr e3q8HxoiYXG2kc7UU+LrqyskoscYHgLVCiPQoTGCYN5VwQo1x3ZZC8pk/iuJqbiR+zm8buOLUHM 6vtPBjcLrGFc3CAP1Z7A/Vy/mh5TLKIK1y1X5ri17F8hT9zSWFIeF3TEoqUNXNUBnuoN0Na2rL6 GT2qt2JqtG8Wh8EF/nOwIhIioFKQkfyOLORbuk6p31X39CPNEiGdJGswe4CMl4ouh7DMCcjKvpa ggha2sgabLmwyrdc/hXLI0F/hJhvpjKlIIE+2GyR/weaa9iz2xAuUx+i/DPo3mezeHMgkuKGt4x weuln0D6/3cZ2k= X-Received: by 2002:a05:620a:a414:20b0:93b:d7a1:ba0e with SMTP id af79cd13be357-93c166802c5mr232637985a.57.1790044112434; Mon, 21 Sep 2026 19:28:32 -0700 (PDT) Received: from toxicpanda.com ([153.61.196.243]) by smtp.gmail.com with ESMTPSA id af79cd13be357-93c1d19ff89sm22030785a.24.2026.09.21.19.28.31 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 21 Sep 2026 19:28:31 -0700 (PDT) From: Josef Bacik Date: Tue, 22 Sep 2026 02:23:24 +0000 Subject: [PATCH v5 05/13] x86/ftrace: Take a Tasks Trace reader around ftrace_caller's call-out Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260922-b4-rcu-tasks-preempt-qs-v5-5-410f57770bad@toxicpanda.com> References: <20260922-b4-rcu-tasks-preempt-qs-v5-0-410f57770bad@toxicpanda.com> In-Reply-To: <20260922-b4-rcu-tasks-preempt-qs-v5-0-410f57770bad@toxicpanda.com> To: "Paul E. McKenney" , Frederic Weisbecker , Alexei Starovoitov , Steven Rostedt Cc: Boqun Feng , Masami Hiramatsu , Mark Rutland , Peter Zijlstra , Thomas Gleixner , Daniel Borkmann , Andrii Nakryiko , Puranjay Mohan , rcu@vger.kernel.org, bpf@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, Josef Bacik X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openssh-sha256; t=1790043820; l=10955; i=josef@toxicpanda.com; h=from:subject:message-id; bh=1mhltqxNnCEQ1bcCikw9zgimeks9K7E+TczS0kwNUMw=; b=U1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgUBr36M/n0nWN0DNbnxwzIiCZez6MG JiruuNaSCI/zXsAAAAGcGF0YXR0AAAAAAAAAAZzaGE1MTIAAABTAAAAC3NzaC1lZDI1NTE5AAAA QF9CdIkGApa945J1vk2W+WTwCDI23VixB6KL4xohl05jh4SjouoFpLKPU7RYPX/mzDNzARG4WqR 0vsgs6UviVQA= X-Developer-Key: i=josef@toxicpanda.com; a=openssh; fpr=SHA256:C8kOX2QUJCMqnCX+KEeoqRAjLo9L+ELOSH2NSAJHqGA For HAVE_RCU_TRAMPOLINE_READERS the ftrace trampolines must be Tasks Trace RCU readers while they call out, since that -- and not the absence of a voluntary context switch -- is what synchronize_rcu_tasks() will wait for before ftrace_shutdown() frees a dynamic trampoline or its ops. Open-code rcu_read_lock_trace() and rcu_read_unlock_trace() in ftrace_caller and ftrace_regs_caller: bump current->trc_reader_nesting and, for the outermost reader, do the SRCU-fast per-CPU increment on rcu_tasks_trace_srcu_struct and stash the counter pointer in current->trc_reader_scp, exactly as the C inlines do (including the smp_mb() when CONFIG_TASKS_TRACE_RCU_NO_MB is not set). The lock sits before the function_trace_op load, because between that load and the call the ops pointer is protected only by Tasks RCU, and the unlock after the call returns. Note at rcu_read_lock_trace() that it now has open-coded copies that have to be kept in step. The sequences are inside t= he region that create_trampoline() copies for per-ops trampolines; their %rip-relative references are fixed up by text_poke_apply_relocation() like CALL_DEPTH_ACCOUNT's. %rax and %rcx are dead at both points. Two pieces of core text still run outside that reader while holding the address of a Tasks-RCU-protected trampoline they are about to enter: the static stubs themselves, whose direct-call tails keep a BPF trampoline address on the stack until the final RET, and, under CONFIG_MITIGATION_RETHUNK, the return thunk that RET expands to. Add an ftrace_static_tramp_end marker after ftrace_stub_direct_tramp and linker symbols around .text..__x86.return_thunk and .text..__x86.rethunk_safe, and provide arch_rcu_tasks_trampoline_text() covering [ftrace_caller, ftrace_static_tramp_end) and both thunk ranges so the irq-exit check treats a task interrupted there as a holdout. All of this is built only under CONFIG_TASKS_RCU_TRAMPOLINE_READERS, which x86 does not enable until a later patch. Assisted-by: LLM Signed-off-by: Josef Bacik --- arch/x86/kernel/asm-offsets.c | 8 +++++ arch/x86/kernel/ftrace.c | 43 ++++++++++++++++++++++++++ arch/x86/kernel/ftrace_64.S | 69 ++++++++++++++++++++++++++++++++++++++= ++++ arch/x86/kernel/vmlinux.lds.S | 4 +++ include/linux/rcupdate_trace.h | 6 ++++ 5 files changed, 130 insertions(+) diff --git a/arch/x86/kernel/asm-offsets.c b/arch/x86/kernel/asm-offsets.c index 081816888f7a..876c3986419a 100644 --- a/arch/x86/kernel/asm-offsets.c +++ b/arch/x86/kernel/asm-offsets.c @@ -9,6 +9,7 @@ #include #include #include +#include #include #include #include @@ -46,6 +47,13 @@ static void __used common(void) #ifdef CONFIG_STACKPROTECTOR OFFSET(TASK_stack_canary, task_struct, stack_canary); #endif +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS + OFFSET(TASK_trc_reader_nesting, task_struct, trc_reader_nesting); + OFFSET(TASK_trc_reader_scp, task_struct, trc_reader_scp); + OFFSET(SRCU_srcu_ctrp, srcu_struct, srcu_ctrp); + OFFSET(SRCU_CTR_srcu_locks, srcu_ctr, srcu_locks); + OFFSET(SRCU_CTR_srcu_unlocks, srcu_ctr, srcu_unlocks); +#endif =20 BLANK(); OFFSET(pbe_address, pbe, address); diff --git a/arch/x86/kernel/ftrace.c b/arch/x86/kernel/ftrace.c index 17d6edfcb7e0..9babaed483eb 100644 --- a/arch/x86/kernel/ftrace.c +++ b/arch/x86/kernel/ftrace.c @@ -275,6 +275,49 @@ static inline void tramp_free(void *tramp) execmem_free(tramp); } =20 +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS +extern void ftrace_static_tramp_end(void); +extern char __return_thunk_start[], __return_thunk_end[]; +extern char __rethunk_safe_start[], __rethunk_safe_end[]; + +/* + * The SRCU-fast increments in TRACE_RCU_READ_LOCK/UNLOCK (ftrace_64.S) ar= e the + * this_cpu_inc() form. + */ +static_assert(!IS_ENABLED(CONFIG_NEED_SRCU_NMI_SAFE)); + +/* + * See rcu_tasks_trampoline_text(). Some core kernel text behaves like a + * trampoline for Tasks RCU purposes because a task executing there outside + * any Tasks Trace reader may still be about to enter a Tasks-RCU-protected + * trampoline whose address it already holds: + * + * - the static ftrace_caller / ftrace_regs_caller / ftrace_stub_direct_t= ramp + * stubs, which carry a direct-call target on the stack until their fin= al + * RET, and + * - the return thunks that RET expands to under CONFIG_MITIGATION_RETHUN= K, + * which run after leaving the stubs above and before landing in that + * target. + */ +bool arch_rcu_tasks_trampoline_text(unsigned long ip) +{ + if (ip >=3D (unsigned long)ftrace_caller && + ip < (unsigned long)ftrace_static_tramp_end) + return true; +#ifdef CONFIG_MITIGATION_RETPOLINE + if (ip >=3D (unsigned long)__return_thunk_start && + ip < (unsigned long)__return_thunk_end) + return true; +#endif +#ifdef CONFIG_MITIGATION_SRSO + if (ip >=3D (unsigned long)__rethunk_safe_start && + ip < (unsigned long)__rethunk_safe_end) + return true; +#endif + return false; +} +#endif /* CONFIG_TASKS_RCU_TRAMPOLINE_READERS */ + /* Defined as markers to the end of the ftrace default trampolines */ extern void ftrace_regs_caller_end(void); extern void ftrace_caller_end(void); diff --git a/arch/x86/kernel/ftrace_64.S b/arch/x86/kernel/ftrace_64.S index 62c1c93aa1c6..5d8cb3861978 100644 --- a/arch/x86/kernel/ftrace_64.S +++ b/arch/x86/kernel/ftrace_64.S @@ -7,6 +7,7 @@ #include #include #include +#include #include #include #include @@ -145,6 +146,53 @@ SYM_FUNC_END(ftrace_stub_graph) =20 #ifdef CONFIG_DYNAMIC_FTRACE =20 +/* + * Open-coded rcu_read_lock_trace() / rcu_read_unlock_trace(), see + * include/linux/rcupdate_trace.h and CONFIG_HAVE_RCU_TRAMPOLINE_READERS: = the + * trampoline and the ftrace_ops it is about to load are kept alive by Tas= ks + * RCU only while we are inside this reader, so the lock must precede the + * function_trace_op load and the unlock must follow the call. These live + * inside the region copied into dynamic trampolines; the %rip-relative + * references are fixed up by text_poke_apply_relocation() in + * create_trampoline(). Clobbers %rax, %rcx and flags. + */ +.macro TRACE_RCU_READ_LOCK +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS + movq PER_CPU_VAR(current_task), %rcx + movl TASK_trc_reader_nesting(%rcx), %eax + incl TASK_trc_reader_nesting(%rcx) + testl %eax, %eax + jnz .Ltrl_nested_\@ + movq rcu_tasks_trace_srcu_struct+SRCU_srcu_ctrp(%rip), %rax + incq %gs:SRCU_CTR_srcu_locks(%rax) + movq %rax, TASK_trc_reader_scp(%rcx) +#ifndef CONFIG_TASKS_TRACE_RCU_NO_MB + lock addl $0, -4(%rsp) /* smp_mb() */ +#endif +.Ltrl_nested_\@: +#endif +.endm + +.macro TRACE_RCU_READ_UNLOCK +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS + movq PER_CPU_VAR(current_task), %rcx + movl TASK_trc_reader_nesting(%rcx), %eax + subl $1, %eax + jnz .Ltru_nested_\@ + /* Outermost: pick up scp before an interrupt can see nesting =3D=3D 0. */ + movq TASK_trc_reader_scp(%rcx), %rax + movl $0, TASK_trc_reader_nesting(%rcx) +#ifndef CONFIG_TASKS_TRACE_RCU_NO_MB + lock addl $0, -4(%rsp) /* smp_mb() */ +#endif + incq %gs:SRCU_CTR_srcu_unlocks(%rax) + jmp .Ltru_done_\@ +.Ltru_nested_\@: + movl %eax, TASK_trc_reader_nesting(%rcx) +.Ltru_done_\@: +#endif +.endm + SYM_FUNC_START(__fentry__) ANNOTATE_NOENDBR CALL_DEPTH_ACCOUNT @@ -163,6 +211,8 @@ SYM_FUNC_START(ftrace_caller) leaq MCOUNT_REG_SIZE+8(%rsp), %rcx movq %rcx, RSP(%rsp) =20 + TRACE_RCU_READ_LOCK + SYM_INNER_LABEL(ftrace_caller_op_ptr, SYM_L_GLOBAL) ANNOTATE_NOENDBR /* Load the ftrace_ops into the 3rd parameter */ @@ -181,6 +231,8 @@ SYM_INNER_LABEL(ftrace_call, SYM_L_GLOBAL) ANNOTATE_NOENDBR call ftrace_stub =20 + TRACE_RCU_READ_UNLOCK + /* Handlers can change the RIP */ movq RIP(%rsp), %rax movq %rax, MCOUNT_REG_SIZE(%rsp) @@ -209,6 +261,8 @@ SYM_FUNC_START(ftrace_regs_caller) =20 CALL_DEPTH_ACCOUNT =20 + TRACE_RCU_READ_LOCK + SYM_INNER_LABEL(ftrace_regs_caller_op_ptr, SYM_L_GLOBAL) ANNOTATE_NOENDBR /* Load the ftrace_ops into the 3rd parameter */ @@ -246,6 +300,8 @@ SYM_INNER_LABEL(ftrace_regs_call, SYM_L_GLOBAL) ANNOTATE_NOENDBR call ftrace_stub =20 + TRACE_RCU_READ_UNLOCK + /* Copy flags back to SS, to restore them */ movq EFLAGS(%rsp), %rax movq %rax, MCOUNT_REG_SIZE(%rsp) @@ -328,6 +384,19 @@ SYM_FUNC_START(ftrace_stub_direct_tramp) RET SYM_FUNC_END(ftrace_stub_direct_tramp) =20 +/* + * [ftrace_caller, ftrace_static_tramp_end) is treated as trampoline text = by + * rcu_tasks_trampoline_text(): outside TRACE_RCU_READ_LOCK/UNLOCK the stu= bs + * may still hold a direct-call trampoline address (ORIG_RAX / the return + * address they RET to) that only Tasks RCU keeps alive. With return thun= ks + * the RET itself runs elsewhere; arch_rcu_tasks_trampoline_text() covers + * those too. + */ +SYM_CODE_START_NOALIGN(ftrace_static_tramp_end) + UNWIND_HINT_UNDEFINED + ANNOTATE_NOENDBR +SYM_CODE_END(ftrace_static_tramp_end) + #else /* ! CONFIG_DYNAMIC_FTRACE */ =20 SYM_FUNC_START(__fentry__) diff --git a/arch/x86/kernel/vmlinux.lds.S b/arch/x86/kernel/vmlinux.lds.S index 2438b89a4620..e546283dc267 100644 --- a/arch/x86/kernel/vmlinux.lds.S +++ b/arch/x86/kernel/vmlinux.lds.S @@ -151,7 +151,9 @@ SECTIONS * definition. */ . =3D srso_alias_untrain_ret | (1 << 2) | (1 << 8) | (1 << 14) | (1 << 2= 0); + __rethunk_safe_start =3D .; *(.text..__x86.rethunk_safe) + __rethunk_safe_end =3D .; #endif ALIGN_ENTRY_TEXT_END =20 @@ -162,7 +164,9 @@ SECTIONS SOFTIRQENTRY_TEXT #ifdef CONFIG_MITIGATION_RETPOLINE *(.text..__x86.indirect_thunk) + __return_thunk_start =3D .; *(.text..__x86.return_thunk) + __return_thunk_end =3D .; #endif STATIC_CALL_TEXT *(.gnu.warning) diff --git a/include/linux/rcupdate_trace.h b/include/linux/rcupdate_trace.h index 273c59a03251..dcdb11643496 100644 --- a/include/linux/rcupdate_trace.h +++ b/include/linux/rcupdate_trace.h @@ -92,6 +92,12 @@ static inline void rcu_read_unlock_tasks_trace(struct sr= cu_ctr __percpu *scp) * the all the other tasks exit their critical sections. * * For more details, please see the documentation for rcu_read_lock(). + * + * CONFIG_HAVE_RCU_TRAMPOLINE_READERS architectures open-code this pair in + * their ftrace and BPF trampolines (arch/x86/kernel/ftrace_64.S, + * arch/x86/kernel/kprobes/opt.c, arch/x86/net/bpf_jit_comp.c, + * arch/arm64/kernel/entry-ftrace.S, arch/arm64/net/bpf_jit_comp.c, + * samples/ftrace/ftrace-direct.h); changes here need to be mirrored there. */ static inline void rcu_read_lock_trace(void) { --=20 2.55.0 From nobody Thu Sep 24 17:09:51 2026 Received: from mail-qk2-f43.google.com (mail-qk2-f43.google.com [74.125.230.235]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 777E22AD03 for ; Tue, 22 Sep 2026 02:29:47 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.235 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790044189; cv=none; b=PnEfWmo3LXwNMm+Fly+DR7n+d74sgWB68mxjQgaKeKKGwyZNBgEJqupo8FUnMOXxqc/1QwsFanw8lzlHjbfvWAOvk1wP5wV2KHl1nzq0pG0ub7FT1FDJBY8396TJsntqpw6J104ke+CMRvTCWIY8JlZ4+b+4LRciXbjKaC6txfA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790044189; c=relaxed/simple; bh=cyOEpv7E+99KbsYBQF12i4qmiz/SB2W5CJ1v6BiG3CI=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=BBp5byJKHdn1A1inTZ0vD90Aiesl7QGovkrwDO48iBD2d2MXBZPE/acg4t2s+yBHA/xj6JWFXWCQjsLkYYOKRCeXVyuJ+b4ouyxEwX2Of4qz8dLwrs2HhyQjRN6oJ/MOl8i1dd0hCwP1uKyA8T1hWlAvDzloIF8yj8weehEKYt4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com; spf=pass smtp.mailfrom=toxicpanda.com; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b=A6Iq5M5b; arc=none smtp.client-ip=74.125.230.235 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b="A6Iq5M5b" Received: by mail-qk2-f43.google.com with SMTP id d75a77b69052e-52fb76906adso54783431cf.0 for ; Mon, 21 Sep 2026 19:29:47 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=toxicpanda.com; s=google; t=1790044186; x=1790648986; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=9lDajmUPmaz8BTxhPEk6XaACKGtk0v9RwI05AN+Obys=; b=A6Iq5M5bPgVhobuemnvHhFbQG6jm+RTffV5wIx7w00uz1L4YTq8GGbBw7igz3LiCes wCcbvC4cyYIp5FoZR9GNMoyL1DJ/1Q+QzdVJ22UXl9EvhoOpV5An4KQ5eOT7+ejxkyIF vGNVKA7o4FVI0yS3caZAq8TGsIjW22efro13CXYtX42w8j0O7XqlZGCsVIrhoUcb/wb0 d8W5c6JJLWia/fftkiI/pv30B8QXOaZIds5aYcKN7SKeGCwSmu3OELlQVW8gg1C87jOC b0DIPyFEmMptCu7Ty1ICn14lAGrOLUjNztbZTt0kXDF3INLgfGF6k7bhGx5XTXsSZ+pd a5CQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790044186; x=1790648986; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=9lDajmUPmaz8BTxhPEk6XaACKGtk0v9RwI05AN+Obys=; b=M9WLL80PFlQ+PksX6/882EaDXAaWS0gMR0a0Y+QjEFoJWKQvUpXgY1galkSr46Ne8b LFLM0gXwl7bwCOzHldiTGewILQ1UHn1fXuidpjoINjoW3qpHQFUDDl373w33aJVozvxh cr8agxyhoE0QRaEKwmPLUW+MLEKypGWYdRn5zMJy8loeoztKsYFCGsf8h+romPyDzglx oBdCcMLGoPhiJInqkYMV/MtE38TtoqmlnSByc60Hnyu5XTsIKNITJyKQchtvKzw7tgL1 hu6clRsZNp1XWMfVHmh5F+whBoFW3ANpvFEMTs2RFct2eQxkDiTQhHbMOv/g7NX+MTSv Z7rw== X-Forwarded-Encrypted: i=1; AKwUvByrkl6dol6sWX4oOn3Min/8RoVq6vgMQzeP4ozFCyi61dj33JSLHv7nISniTcYHWGjxM0EUlOgtTveU1YE=@vger.kernel.org X-Gm-Message-State: AFuF++naAu3JfM1tk7xAwbDnnCaQ0W2ZtDtIJ0IIHHmwh7X1ktdxoUB8 qUiP/tXffLywFuxwpXcKMtGQ4gQ7JtighU+ctLuEvs2xlVZEOmaTHWHXRDmKb3Ub5i4= X-Gm-Gg: AYBFou0BjzzEsVwyUMnG7A4/plAGguyudwfMKWZ5qo8RtGMiEyOuCyL/s3Mxuunz038 79wBtFAl0JdoRjpX1BAZACoUG4oLmcPPedD8uybLNu/uFO7sfxR2WnoLIzXPjpyDwKKNMTrUhRE CH1HX6PJwIsMFijTgjZQFEYz0X5XYPadBrisAQXuCk6XvD3G+oT2KQHCHAitRw9zSejD5ZE+R40 7zWGhFnGIpB0g04qWTwFGqYw8Kjlnbo2Yb33F1MVHKxa5jr6JcKe6nkXV85j009IjXXrpkmn5VA s9lHG86RoxUSFQtCzEj5EOQc12L6zaKkP8xzvB0SuAD7kjnPN5I+51BHv7ihv13qblgftIR4d22 unKyaJyBtPblGMZmRzEjfpT4RT8ezW/vq52Vym8YUcDY0n+lOP6SxRY6OCabm0IQFm4breVrq0w r+gPPUoks6bmynDCgGRKTMXRspcfpUpVX+PWgPXISJXkSgHvGQofmj4HWemum66sXLc+LpWBqCx My7UBaai3Fmk2Y= X-Received: by 2002:a05:622a:5a19:b0:532:b3ce:f38 with SMTP id d75a77b69052e-532dd3cc9a4mr16978661cf.4.1790044186247; Mon, 21 Sep 2026 19:29:46 -0700 (PDT) Received: from toxicpanda.com ([153.61.196.243]) by smtp.gmail.com with ESMTPSA id d75a77b69052e-532e192c463sm1639311cf.21.2026.09.21.19.29.45 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 21 Sep 2026 19:29:45 -0700 (PDT) From: Josef Bacik Date: Tue, 22 Sep 2026 02:23:25 +0000 Subject: [PATCH v5 06/13] x86/kprobes: Take a Tasks Trace reader in the optprobe template Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260922-b4-rcu-tasks-preempt-qs-v5-6-410f57770bad@toxicpanda.com> References: <20260922-b4-rcu-tasks-preempt-qs-v5-0-410f57770bad@toxicpanda.com> In-Reply-To: <20260922-b4-rcu-tasks-preempt-qs-v5-0-410f57770bad@toxicpanda.com> To: "Paul E. McKenney" , Frederic Weisbecker , Alexei Starovoitov , Steven Rostedt Cc: Boqun Feng , Masami Hiramatsu , Mark Rutland , Peter Zijlstra , Thomas Gleixner , Daniel Borkmann , Andrii Nakryiko , Puranjay Mohan , rcu@vger.kernel.org, bpf@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, Josef Bacik X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openssh-sha256; t=1790043820; l=3977; i=josef@toxicpanda.com; h=from:subject:message-id; bh=cyOEpv7E+99KbsYBQF12i4qmiz/SB2W5CJ1v6BiG3CI=; b=U1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgUBr36M/n0nWN0DNbnxwzIiCZez6MG JiruuNaSCI/zXsAAAAGcGF0YXR0AAAAAAAAAAZzaGE1MTIAAABTAAAAC3NzaC1lZDI1NTE5AAAA QGWEkZzmBIn08GZX0Y8n1aYaLj9DKrBu1cjgNBvrnGmt+51rXj4Tis5qLbxNWpv8gzG2Md0fp7c HpXWdXp6OPQc= X-Developer-Key: i=josef@toxicpanda.com; a=openssh; fpr=SHA256:C8kOX2QUJCMqnCX+KEeoqRAjLo9L+ELOSH2NSAJHqGA The jump-optimized kprobe template calls optimized_callback() from a dynamically allocated slot with preemption enabled, and only Tasks RCU keeps that slot alive under a task preempted in the callback. For HAVE_RCU_TRAMPOLINE_READERS that means the template must be a Tasks Trace reader across the call, so open-code rcu_read_lock_trace() and rcu_read_unlock_trace() around it as ftrace_64.S does. The template lives in .rodata and is memcpy()d into each slot without relocation processing, so the references to current_task and rcu_tasks_trace_srcu_struct are absolute (R_X86_64_32S, relocated for KASLR like any other) rather than %rip-relative. %rax and %rcx have already been saved by SAVE_REGS_STRING and are dead after the call. The slot itself is dynamically allocated text, so the instructions before the lock and after the unlock are covered by the irq-exit check. 64-bit only; 32-bit x86 does not take part. Assisted-by: LLM Signed-off-by: Josef Bacik --- arch/x86/kernel/kprobes/opt.c | 44 +++++++++++++++++++++++++++++++++++++++= ++++ 1 file changed, 44 insertions(+) diff --git a/arch/x86/kernel/kprobes/opt.c b/arch/x86/kernel/kprobes/opt.c index 3f8fea52619f..68a5de6cdabe 100644 --- a/arch/x86/kernel/kprobes/opt.c +++ b/arch/x86/kernel/kprobes/opt.c @@ -31,6 +31,7 @@ #include #include #include +#include =20 #include "common.h" =20 @@ -101,6 +102,47 @@ static void synthesize_set_arg1(kprobe_opcode_t *addr,= unsigned long val) *(unsigned long *)addr =3D val; } =20 +/* + * Open-coded rcu_read_lock_trace() / rcu_read_unlock_trace() around the c= all + * to optimized_callback(), see CONFIG_HAVE_RCU_TRAMPOLINE_READERS and the + * equivalent macros in ftrace_64.S. The template is memcpy()d into the s= lot + * without relocation processing, so memory references must be absolute + * rather than %rip-relative. %rax and %rcx are free at both points. + */ +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS +#ifndef CONFIG_TASKS_TRACE_RCU_NO_MB +#define OPTPROBE_TRACE_RCU_MB " lock addl $0, -4(%rsp)\n" +#else +#define OPTPROBE_TRACE_RCU_MB +#endif +#define OPTPROBE_TRACE_RCU_READ_LOCK \ + " movq %gs:current_task, %rcx\n" \ + " movl " __stringify(TASK_trc_reader_nesting) "(%rcx), %eax\n" \ + " incl " __stringify(TASK_trc_reader_nesting) "(%rcx)\n" \ + " testl %eax, %eax\n" \ + " jnz 1f\n" \ + " movq rcu_tasks_trace_srcu_struct+" __stringify(SRCU_srcu_ctrp) ", %rax= \n" \ + " incq %gs:" __stringify(SRCU_CTR_srcu_locks) "(%rax)\n" \ + " movq %rax, " __stringify(TASK_trc_reader_scp) "(%rcx)\n" \ + OPTPROBE_TRACE_RCU_MB \ + "1:\n" +#define OPTPROBE_TRACE_RCU_READ_UNLOCK \ + " movq %gs:current_task, %rcx\n" \ + " movl " __stringify(TASK_trc_reader_nesting) "(%rcx), %eax\n" \ + " subl $1, %eax\n" \ + " jnz 2f\n" \ + " movq " __stringify(TASK_trc_reader_scp) "(%rcx), %rax\n" \ + " movl $0, " __stringify(TASK_trc_reader_nesting) "(%rcx)\n" \ + OPTPROBE_TRACE_RCU_MB \ + " incq %gs:" __stringify(SRCU_CTR_srcu_unlocks) "(%rax)\n" \ + " jmp 3f\n" \ + "2: movl %eax, " __stringify(TASK_trc_reader_nesting) "(%rcx)\n" \ + "3:\n" +#else +#define OPTPROBE_TRACE_RCU_READ_LOCK +#define OPTPROBE_TRACE_RCU_READ_UNLOCK +#endif + asm ( ".pushsection .rodata\n" ".global optprobe_template_entry\n" @@ -114,6 +156,7 @@ asm ( "optprobe_template_clac:\n" ASM_NOP3 SAVE_REGS_STRING + OPTPROBE_TRACE_RCU_READ_LOCK " movq %rsp, %rsi\n" ".global optprobe_template_val\n" "optprobe_template_val:\n" @@ -122,6 +165,7 @@ asm ( ".global optprobe_template_call\n" "optprobe_template_call:\n" ASM_NOP5 + OPTPROBE_TRACE_RCU_READ_UNLOCK /* Copy 'regs->flags' into 'regs->ss'. */ " movq 18*8(%rsp), %rdx\n" " movq %rdx, 20*8(%rsp)\n" --=20 2.55.0 From nobody Thu Sep 24 17:09:51 2026 Received: from mail-qv2-f43.google.com (mail-qv2-f43.google.com [74.125.230.171]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4EB273793A8 for ; Tue, 22 Sep 2026 02:30:26 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.171 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790044228; cv=none; b=q5z8WZoTd0VuHUmqqVi9xAqQXDW86Hn/AlZTh/2YdLN1+uXuRx//V0wTA1eEX1JE1upPYVLwstmYzA5mFP2vA7qoYMP8WxIv+jZkr/GWT1G4U/uwHhlbR/f5dRVxL+13XuN2LklMPFDgXOFW7bUsxutX0mmGT4jENU0QTxWWa0U= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790044228; c=relaxed/simple; bh=EPAPVj1FhQ7/+pop5mUZQszLO5ulnMl6jVPcEL6t5Gk=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=eEGHXh6YO3+BAMZRYa9vecan1kykdVEDtencjdxpnXiAJ9abVLsPuU5DXeBNBT/gPl71l/9hulW6Rzxk2QZjOzH2iewYmWL3hskhLGUYLYUSoAqqPd1IOABtqjJY7BY2X303RtZnFZgM+NAgcXy8xdoVZAvkR910+O7QK+2KZQI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com; spf=pass smtp.mailfrom=toxicpanda.com; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b=prV/DsyD; arc=none smtp.client-ip=74.125.230.171 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b="prV/DsyD" Received: by mail-qv2-f43.google.com with SMTP id 6a1803df08f44-91400f92ae8so1896556d6.0 for ; Mon, 21 Sep 2026 19:30:26 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=toxicpanda.com; s=google; t=1790044225; x=1790649025; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=EMtH/e2BOZ+5Q/fPJ0djJ3i8PvS8/cEMBigMOOsU40U=; b=prV/DsyDw/ndMPYWkg7lHH+Q2B149NYvbSo39qG+pRUIcK/3aF/VNkhTvJ0Ajj079C jIVJFgGWmvnccqBOLrHdm7QPFleFXTdkeEldYtJEZvG7mlNpSNYc4oM94PigsAcqpW+Q 7Fi9wDmeRs3VdHcPbWTV2r/9me39NoN9SZzfgnOywZGThii9ILgiqGaZ3Oz+V/A1uKKB FDTjVUVyS7f4lq+et1c6ZL7+8eoyxx2UkbgwgxELdm1VahZR6BXFzVe03uJSW/e532sF ymNKMLqefuplISr+GDYiePrdy9tlzAgFtxRlM0T1JxNMldZxyL4X7NnRSBjAgRK5KR+/ C7jw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790044225; x=1790649025; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=EMtH/e2BOZ+5Q/fPJ0djJ3i8PvS8/cEMBigMOOsU40U=; b=2cgseGsQwOGWAatE2AWxvxKQ/10gE4apj6pTKGE3WqK8QAmur3fsNLV2izWfl64d4L N/wFZXIDizDoMEequlfy+tq6VEiwuDjZugBwOVatz+dqOVs1NPrzqlxkxBSOXY2teXBv E2YxXMjIf5zzqS76A9yXMkVOQYVWvh1ViCQsZ6TylN11A/Y394fgpC1+SIyLkf19kZcc lJ9eGx7CnhagDWPUr+g0OdFpCL/93vo0k9d0AptOyVy+iREXRs2cBQi51k/64hoESHO8 HseFhVZl0P+rYqH4xZ0Uaj7RkaTzqDgnZSu0JK4IjignliQKFCXilReXCOFEO62rNSVV 7Z2g== X-Forwarded-Encrypted: i=1; AKwUvBynOXIw1GW+ekSh5mrgcRhLnOXqFvdLNCHl83HZy0vzMOFSgjToknfkxXPHK/fNjzdiWM1ydkML67BwL2o=@vger.kernel.org X-Gm-Message-State: AFuF++lD5GTtCcPdRO+Pq8FvVuC6ZST8ltBrRPnB7JPnVZ9LHGV5/5cS YirgqeY+3LC85UUtEDoLLbxLrvOgsSsuH9wI8i2QoKqi95VSYCP5MdHBEsO1rwVT+RQ= X-Gm-Gg: AYBFou3n+PbZW1/eflJd+g3Yv7oVNNr5KBMX7m0ohD8sdN348Cl0UnhUx5zSiWNmeHd YfRWhJctXqcdis6cWYiPxyJ28jKAp0opWXmeQTh0JwUjmYAenkfWNOdrQOlNQphBctNBkTnlPfn H78zb3OBoJFp2GaC1CSJXfsq8uTj6uWLxGY/85zWYsS+ZZKdqpoqehCyag2u0ed6gcQRIRIDGoz 2/LVEqslWjXynCG23+nl8k/o20r7WO+vk2KgDWgX3I68R9Kwi1P2/OutN5VgLQpVT2nChpW4VTx ULfuXTFLHz7XNhGt0+qYyZJHih0F3x6FWC1ITPFBLKxfbSx1TzwnkAzrzFEV/Ue8sdxCeU6cWx5 82q1c/psI+ZakMNHsJ/j1oAMRK9PWPxyriAHkp/K2XBZTwhJkBprReW7MBWQUKMOzL6XOIgJMGR ZIRF60R7JNbtaofeoOmxDFpGXQXKpZ99i0i1OvV1t2Cynkp3M39Hd/ItgpGPihGMt4oe2T/S8ve IPn3pitmKJV3XI= X-Received: by 2002:a05:6214:3ca0:b0:912:435f:1a8b with SMTP id 6a1803df08f44-913fc8d3895mr33550386d6.22.1790044225028; Mon, 21 Sep 2026 19:30:25 -0700 (PDT) Received: from toxicpanda.com ([153.61.196.243]) by smtp.gmail.com with ESMTPSA id 6a1803df08f44-914024a1b21sm4504326d6.15.2026.09.21.19.30.24 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 21 Sep 2026 19:30:24 -0700 (PDT) From: Josef Bacik Date: Tue, 22 Sep 2026 02:23:26 +0000 Subject: [PATCH v5 07/13] bpf, x86: Take a Tasks Trace reader in the trampoline around its call-outs Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260922-b4-rcu-tasks-preempt-qs-v5-7-410f57770bad@toxicpanda.com> References: <20260922-b4-rcu-tasks-preempt-qs-v5-0-410f57770bad@toxicpanda.com> In-Reply-To: <20260922-b4-rcu-tasks-preempt-qs-v5-0-410f57770bad@toxicpanda.com> To: "Paul E. McKenney" , Frederic Weisbecker , Alexei Starovoitov , Steven Rostedt Cc: Boqun Feng , Masami Hiramatsu , Mark Rutland , Peter Zijlstra , Thomas Gleixner , Daniel Borkmann , Andrii Nakryiko , Puranjay Mohan , rcu@vger.kernel.org, bpf@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, Josef Bacik X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openssh-sha256; t=1790043820; l=8237; i=josef@toxicpanda.com; h=from:subject:message-id; bh=EPAPVj1FhQ7/+pop5mUZQszLO5ulnMl6jVPcEL6t5Gk=; b=U1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgUBr36M/n0nWN0DNbnxwzIiCZez6MG JiruuNaSCI/zXsAAAAGcGF0YXR0AAAAAAAAAAZzaGE1MTIAAABTAAAAC3NzaC1lZDI1NTE5AAAA QKeNfl/NvsHWWZiu8ABGjBm0Ik1LIwwY7PXte5+OSYt8vF7GS12LYmo7jIhnhxlZSNoBmV+d/6B f4uZu/duViwk= X-Developer-Key: i=josef@toxicpanda.com; a=openssh; fpr=SHA256:C8kOX2QUJCMqnCX+KEeoqRAjLo9L+ELOSH2NSAJHqGA On HAVE_RCU_TRAMPOLINE_READERS kernels Tasks RCU keeps a BPF trampoline image allocated only while a task using it is a Tasks Trace RCU reader or is executing text that rcu_tasks_trampoline_text() recognises. The image itself is such text, but the C glue and the programs it calls are not, and only sleepable programs take rcu_read_lock_trace() today. Have the x86-64 JIT open-code rcu_read_lock_trace() and rcu_read_unlock_trace() in the trampoline, as ftrace_64.S does for ftrace_caller: one reader from just after the frame is set up to just before the original function is called, covering __bpf_tramp_enter() and the fentry and fmod_ret programs, and a second one from just after the original function returns to just before the final register restore, covering the fexit programs and __bpf_tramp_exit(). The original function itself runs outside both, since it may run for a long time and the image is pinned by im->pcref across it. Trampolines that do not call the original function get a single reader around all their programs. The second reader is entered before ip_after_call, so the ip_after_call -> ip_epilogue jump that bpf_tramp_image_put() patches in is inside it, and the fmod_ret early-exit branch lands after that point still holding the first reader, so exactly one is held on every path. The sequence uses r10 and r11, which are scratch at each emission point, and references current_task and rcu_tasks_trace_srcu_struct by absolute sign-extended address, the form the JIT already relies on for this_cpu_off. Sleepable programs' own rcu_read_lock_trace() simply nests. Nothing is emitted on other configurations. Suggested-by: Alexei Starovoitov Assisted-by: LLM Signed-off-by: Josef Bacik --- arch/x86/net/bpf_jit_comp.c | 107 ++++++++++++++++++++++++++++++++++++++++= ++++ 1 file changed, 107 insertions(+) diff --git a/arch/x86/net/bpf_jit_comp.c b/arch/x86/net/bpf_jit_comp.c index 2853e87797a7..b77e29c9599d 100644 --- a/arch/x86/net/bpf_jit_comp.c +++ b/arch/x86/net/bpf_jit_comp.c @@ -14,6 +14,7 @@ #include #include #include +#include #include #include #include @@ -722,6 +723,91 @@ static void emit_indirect_jump(u8 **pprog, int bpf_reg= , u8 *ip) *pprog =3D prog; } =20 +/* + * Open-coded rcu_read_lock_trace() / rcu_read_unlock_trace() for the + * trampoline, see CONFIG_HAVE_RCU_TRAMPOLINE_READERS and the equivalent + * macros in arch/x86/kernel/ftrace_64.S. The image is not relocated, so + * current_task and rcu_tasks_trace_srcu_struct are referenced by absolute + * (sign-extended 32-bit) address, the form the JIT already relies on for + * this_cpu_off. Uses r10 and r11, which are scratch at every emission + * point, and clobbers flags. + * + * lock: unlock: + * mov r11, gs:[current_task] mov r11, gs:[current_task] + * mov r10d, [r11+nesting_off] mov r10d, [r11+nesting_off] + * inc dword ptr [r11+nesting_off] sub r10d, 1 + * test r10d, r10d jnz 2f + * jnz 1f mov r10, [r11+scp_off] + * mov r10, [&srcu.srcu_ctrp] mov dword ptr [r11+nesting_of= f], 0 + * inc qword ptr gs:[r10+locks_off] (smp_mb) + * mov [r11+scp_off], r10 inc qword ptr gs:[r10+unlocks= _off] + * (smp_mb) jmp 3f + * 1: 2: mov [r11+nesting_off], r10d + * 3: + */ +static void emit_trace_rcu_reader(u8 **pprog, bool lock) +{ +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS + const u32 nesting_off =3D offsetof(struct task_struct, trc_reader_nesting= ); + const u32 scp_off =3D offsetof(struct task_struct, trc_reader_scp); + const u32 locks_off =3D offsetof(struct srcu_ctr, srcu_locks); + const u32 unlocks_off =3D offsetof(struct srcu_ctr, srcu_unlocks); + const bool mb =3D !IS_ENABLED(CONFIG_TASKS_TRACE_RCU_NO_MB); + u8 *prog =3D *pprog; + + /* The plain this_cpu_inc() form of __srcu_read_lock_fast(). */ + BUILD_BUG_ON(IS_ENABLED(CONFIG_NEED_SRCU_NMI_SAFE)); + + /* mov r11, gs:[abs32 current_task] */ + EMIT2(0x65, 0x4C); + EMIT3(0x8B, 0x1C, 0x25); + EMIT((u32)(unsigned long)¤t_task, 4); + /* mov r10d, dword ptr [r11 + nesting_off] */ + EMIT3_off32(0x45, 0x8B, 0x93, nesting_off); + + if (lock) { + /* inc dword ptr [r11 + nesting_off] */ + EMIT3_off32(0x41, 0xFF, 0x83, nesting_off); + /* test r10d, r10d */ + EMIT3(0x45, 0x85, 0xD2); + /* jnz 1f */ + EMIT2(X86_JNE, 8 + 8 + 7 + (mb ? 6 : 0)); + /* mov r10, qword ptr [abs32 &rcu_tasks_trace_srcu_struct.srcu_ctrp] */ + EMIT4(0x4C, 0x8B, 0x14, 0x25); + EMIT((u32)(unsigned long)&rcu_tasks_trace_srcu_struct.srcu_ctrp, 4); + /* inc qword ptr gs:[r10 + locks_off] */ + EMIT4_off32(0x65, 0x49, 0xFF, 0x82, locks_off); + /* mov qword ptr [r11 + scp_off], r10 */ + EMIT3_off32(0x4D, 0x89, 0x93, scp_off); + /* smp_mb(): lock add dword ptr [rsp - 4], 0 */ + if (mb) + EMIT2_off32(0xF0, 0x83, 0x00FC2444); + /* 1: */ + } else { + /* sub r10d, 1 */ + EMIT4(0x41, 0x83, 0xEA, 0x01); + /* jnz 2f */ + EMIT2(X86_JNE, 7 + 11 + (mb ? 6 : 0) + 8 + 2); + /* mov r10, qword ptr [r11 + scp_off] */ + EMIT3_off32(0x4D, 0x8B, 0x93, scp_off); + /* mov dword ptr [r11 + nesting_off], 0 */ + EMIT3_off32(0x41, 0xC7, 0x83, nesting_off); + EMIT(0, 4); + if (mb) + EMIT2_off32(0xF0, 0x83, 0x00FC2444); + /* inc qword ptr gs:[r10 + unlocks_off] */ + EMIT4_off32(0x65, 0x49, 0xFF, 0x82, unlocks_off); + /* jmp 3f */ + EMIT2(0xEB, 7); + /* 2: mov dword ptr [r11 + nesting_off], r10d */ + EMIT3_off32(0x45, 0x89, 0x93, nesting_off); + /* 3: */ + } + + *pprog =3D prog; +#endif +} + static void emit_return(u8 **pprog, u8 *ip) { u8 *prog =3D *pprog; @@ -3610,6 +3696,16 @@ static int __arch_prepare_bpf_trampoline(struct bpf_= tramp_image *im, void *rw_im /* mov QWORD PTR [rbp - rbx_off], rbx */ emit_stx(&prog, BPF_DW, BPF_REG_FP, BPF_REG_6, -rbx_off); =20 + /* + * Tasks RCU keeps this image alive only while we are a Tasks Trace + * reader; the instructions before this point (and after the final + * unlock) are covered by the irq-exit IP check. One reader spans + * __bpf_tramp_enter() and the fentry/fmod_ret progs, a second one + * the fexit progs and __bpf_tramp_exit(); the original function runs + * outside both, with the image pinned by im->pcref instead. + */ + emit_trace_rcu_reader(&prog, true); + func_meta =3D nr_regs; /* Store number of argument registers of the traced function */ emit_store_stack_imm64(&prog, BPF_REG_0, -func_meta_off, func_meta); @@ -3660,6 +3756,7 @@ static int __arch_prepare_bpf_trampoline(struct bpf_t= ramp_image *im, void *rw_im } =20 if (flags & BPF_TRAMP_F_CALL_ORIG) { + emit_trace_rcu_reader(&prog, false); restore_regs(m, &prog, regs_off); save_args(m, &prog, arg_stack_off, true, flags, 0); =20 @@ -3682,6 +3779,13 @@ static int __arch_prepare_bpf_trampoline(struct bpf_= tramp_image *im, void *rw_im } /* remember return value in a stack for bpf prog to access */ emit_stx(&prog, BPF_DW, BPF_REG_FP, BPF_REG_0, -8); + /* + * Second reader. Taken before ip_after_call so that the + * ip_after_call -> ip_epilogue jump patched in at teardown is + * inside it too; the fmod_ret early exit jumps past this still + * holding the first reader, so either way exactly one is held. + */ + emit_trace_rcu_reader(&prog, true); im->ip_after_call =3D image + (prog - (u8 *)rw_image); emit_nops(&prog, X86_PATCH_SIZE); } @@ -3737,6 +3841,9 @@ static int __arch_prepare_bpf_trampoline(struct bpf_t= ramp_image *im, void *rw_im LOAD_TRAMP_TAIL_CALL_CNT_PTR(stack_size); } =20 + /* Remaining instructions are covered by the irq-exit IP check. */ + emit_trace_rcu_reader(&prog, false); + /* restore return value of orig_call or fentry prog back into RAX */ if (save_ret) emit_ldx(&prog, BPF_DW, BPF_REG_0, BPF_REG_FP, -8); --=20 2.55.0 From nobody Thu Sep 24 17:09:51 2026 Received: from mail-qk2-f12.google.com (mail-qk2-f12.google.com [74.125.230.204]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CA901345CDD for ; Tue, 22 Sep 2026 02:31:05 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.204 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790044268; cv=none; b=tEhIKfQBJeBxrQPPs6HnKLeO6aYqC6cjcuw4ZiFqlpLwS6CWLhBXMUAw3YqgNmsuyxENE+LoBX/8KmbXOTZ9JOvVYIxNHnbDS7YYtM7XO7D0fEIy1dUcu1xINTbfDghlhMY8aPF1Q6jhYqr+D2ud6Md7LGebwKLaTJaq4aqHWhc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790044268; c=relaxed/simple; bh=Aqm0C5BBMBkqFusbG93kcm6leRi28F4VKStTPzg6dAE=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=nuHH4I83fwway/ZEmApQ0ZL1L5HiqPq6dNPy+chIkER89d/ulAszwBEWVP6C/t9HcRn0Lewv2dzScdx0SlEgKKHBtEPXk91IGN1oo5olScVIadaVozqJggxz5/Q+lWVSlSJD4m2L0D8z0Mwqwzvc4vdV5CrBvg7evElNRHroQdI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com; spf=pass smtp.mailfrom=toxicpanda.com; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b=JJF7HRNU; arc=none smtp.client-ip=74.125.230.204 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b="JJF7HRNU" Received: by mail-qk2-f12.google.com with SMTP id d75a77b69052e-52fb769ca02so40725711cf.0 for ; Mon, 21 Sep 2026 19:31:05 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=toxicpanda.com; s=google; t=1790044264; x=1790649064; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=ua8lq0Nx07F8Y+BgJ1JiUwgoDOin8IZmspGmQtj7ybE=; b=JJF7HRNUj4Rg+u6UUnt5B/1+Cg9nERrChWYvF9PRZpo17wYkti0nxO4fiuYGiK/2+w vim9WVpcTdwvRlI+hsJ4EW7IuPQVS4wYM3vwK9UG20WVAsa9risAu7/vYvBgvLHu0qtY 0DUpbd0omYtFH8oM4rV76gR93YsZkZAZ0LB3mW2u9SPgecxfi+P6f1JUqZHdRKfJGY3S ZFjF2Ezp86/ruhdAT1LBjSjr3SEEd5DR+Pn3PMWTm53QBnerKFWilA6tN9vigEA50ZYg X7Wf5/fGEsGQLLHEO8HQe1fkY/kLmwWofn0PzJqw3Xd83p4l5BOz5ImDh3Pm10Zn2VU6 YliQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790044264; x=1790649064; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=ua8lq0Nx07F8Y+BgJ1JiUwgoDOin8IZmspGmQtj7ybE=; b=tr52fKLTmHdr2354dEf9oa0Q5WXlxTfOs2bIg4NLsYjNWpjgDMvNVXYjLLRtaIdtxu cvE0M3Lor9ROXuOB3Ddbq0QUoOacnotU2ZBvfGbq32vfCZn2+/M2lQkis2XSb9gGitA/ DKGwf61SVp2RO9WoCtawm8ymaBEmYq1REAwLOAWGVDyYb4gnC6GoOS2MNznBlHti2XGX VrZCe8LLyigKNQJR2Q2FgWwxMLiWE64vS6+T1m+xlmtturp0Ne/t1/IZ8ArkPDffmXDg kHU1zbBfXAiPvPgHtv3oR5+btY8bBK5OVDPFxmo4m7OBVQDdNfE1YhTQBmJxkeQX+s2H EDWg== X-Forwarded-Encrypted: i=1; AKwUvBwu+/ujJN+NyuZuQZFzj/QkDvZPBStTEr/is2KM5rFhhvDU43J3EhdI2knzMIf9rD8KIGlMR71EXIlmwKc=@vger.kernel.org X-Gm-Message-State: AFuF++melPcBT27ZlhlGtrJOHwrlP6r7WeBnEqf2KnvFdtZqKfR+ZZmJ D2ZQYnNIZ+RjlgWV2jI0OTpuv9NWFIbgXSq6crrFGfAmxMhSC2SF/C7Vxfg3ZI5UaYQ= X-Gm-Gg: AYBFou0tHmTFFrg7uljpgKQVP0VNwY2ZJlPFDevlgn98PVQts+T+eNhdwb17auXsfKN Fc4jg7fK3sty6pVqfqxn0kWlKGaqEKz/uQ5xYvm6/Qr19JLlsZZZgI7IDO2AXbdyTY+fbPurxIy D65LdUPVAyqdV+K2+a7gg7WBJ2QF2rpcBqaa5cpagIPLagxAvmMR8en2x82foB9MRIQnccMe96X h6ftQCZU3vem+1myOBvh4HO5uZCm9eZKr18dr1WUOxkVqUXkLCYaXBh51R5zhD9VwLPtqyKJXY0 Cr7fCw6inlTv8mnf+7OjXx4ISBRsMo8AZO5s6mQsNmIxoqxbJc7yhLUwnwz4GjzfVNlMXTS19Sr RXEXkLyc0pAZMMIDeBRljnOKNJMjnMrfU8RB+WnEgybpfH5Neg/hw6Una9UdlSENS/UTrLzE3n4 V/9jl9ayzUzY8TnzAUVxanA9nEdU/hgtiJ9Ldk57ipSF8pBeFlo7nlK6wmD7ofNSCSNCix+CUT+ DaFW8Vdkp6QDQudZXK4nUYJsw== X-Received: by 2002:a05:620a:3199:b0:936:b923:4832 with SMTP id af79cd13be357-93c15db3eecmr402787485a.28.1790044264382; Mon, 21 Sep 2026 19:31:04 -0700 (PDT) Received: from toxicpanda.com ([153.61.196.250]) by smtp.gmail.com with ESMTPSA id af79cd13be357-93c1d18725esm22523185a.22.2026.09.21.19.31.03 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 21 Sep 2026 19:31:03 -0700 (PDT) From: Josef Bacik Date: Tue, 22 Sep 2026 02:23:27 +0000 Subject: [PATCH v5 08/13] arm64: ftrace: Take a Tasks Trace reader around ftrace_caller's call-out Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260922-b4-rcu-tasks-preempt-qs-v5-8-410f57770bad@toxicpanda.com> References: <20260922-b4-rcu-tasks-preempt-qs-v5-0-410f57770bad@toxicpanda.com> In-Reply-To: <20260922-b4-rcu-tasks-preempt-qs-v5-0-410f57770bad@toxicpanda.com> To: "Paul E. McKenney" , Frederic Weisbecker , Alexei Starovoitov , Steven Rostedt Cc: Boqun Feng , Masami Hiramatsu , Mark Rutland , Peter Zijlstra , Thomas Gleixner , Daniel Borkmann , Andrii Nakryiko , Puranjay Mohan , rcu@vger.kernel.org, bpf@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, Josef Bacik X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openssh-sha256; t=1790043820; l=7839; i=josef@toxicpanda.com; h=from:subject:message-id; bh=Aqm0C5BBMBkqFusbG93kcm6leRi28F4VKStTPzg6dAE=; b=U1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgUBr36M/n0nWN0DNbnxwzIiCZez6MG JiruuNaSCI/zXsAAAAGcGF0YXR0AAAAAAAAAAZzaGE1MTIAAABTAAAAC3NzaC1lZDI1NTE5AAAA QLzF5n3GUkTWLii3NiHsvpD1TinWMCM9IV4Llmu+a7yPt78yVhVhsGdhEjVxKNXB2QbrNgDy62p d2pOMkWs6WwQ= X-Developer-Key: i=josef@toxicpanda.com; a=openssh; fpr=SHA256:C8kOX2QUJCMqnCX+KEeoqRAjLo9L+ELOSH2NSAJHqGA For HAVE_RCU_TRAMPOLINE_READERS the ftrace trampoline must be a Tasks Trace RCU reader while it calls out, since that is what synchronize_rcu_tasks() will wait for before ftrace_shutdown() frees an ftrace_ops (or, with CALL_OPS, lets its owner free it) under a task preempted in the callback. Open-code rcu_read_lock_trace() and rcu_read_unlock_trace() around the call to ops->func in ftrace_caller: bump current->trc_reader_nesting via sp_el0 and, for the outermost reader, do the SRCU-fast per-CPU increment on rcu_tasks_trace_srcu_struct and stash the counter pointer in current->trc_reader_scp, as the C inlines do (including the dmb when CONFIG_TASKS_TRACE_RCU_NO_MB is not set). The per-CPU increment is an LL/SC add on this CPU's counter; being migrated between reading the per-CPU offset and the store-exclusive only means another CPU's counter is incremented atomically instead, which SRCU sums over anyway. x12-x16 are free at both points. arm64 has no return thunks and, with CALL_OPS, no dynamic ftrace trampolines, but ftrace_caller itself carries the ops pointer in x11 from before the reader is entered and a direct-call BPF trampoline address in x17 until the final br/ret after it is left, so mark the end of the static trampoline text and provide arch_rcu_tasks_trampoline_text() covering [ftrace_caller, ftrace_static_tramp_end). Built only under CONFIG_TASKS_RCU_TRAMPOLINE_READERS, which arm64 does not enable until a later patch. Assisted-by: LLM Signed-off-by: Josef Bacik --- arch/arm64/kernel/asm-offsets.c | 8 +++++ arch/arm64/kernel/entry-ftrace.S | 74 ++++++++++++++++++++++++++++++++++++= ++++ arch/arm64/kernel/ftrace.c | 20 +++++++++++ 3 files changed, 102 insertions(+) diff --git a/arch/arm64/kernel/asm-offsets.c b/arch/arm64/kernel/asm-offset= s.c index 9c853ed3ceab..f6a8fb1f9b43 100644 --- a/arch/arm64/kernel/asm-offsets.c +++ b/arch/arm64/kernel/asm-offsets.c @@ -10,6 +10,7 @@ =20 #include #include +#include #include #include #include @@ -39,6 +40,13 @@ int main(void) DEFINE(TSK_STACK, offsetof(struct task_struct, stack)); #ifdef CONFIG_STACKPROTECTOR DEFINE(TSK_STACK_CANARY, offsetof(struct task_struct, stack_canary)); +#endif +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS + DEFINE(TSK_TRC_READER_NESTING, offsetof(struct task_struct, trc_reader_n= esting)); + DEFINE(TSK_TRC_READER_SCP, offsetof(struct task_struct, trc_reader_scp)); + DEFINE(SRCU_SRCU_CTRP, offsetof(struct srcu_struct, srcu_ctrp)); + DEFINE(SRCU_CTR_SRCU_LOCKS, offsetof(struct srcu_ctr, srcu_locks)); + DEFINE(SRCU_CTR_SRCU_UNLOCKS, offsetof(struct srcu_ctr, srcu_unlocks)); #endif BLANK(); DEFINE(THREAD_CPU_CONTEXT, offsetof(struct task_struct, thread.cpu_conte= xt)); diff --git a/arch/arm64/kernel/entry-ftrace.S b/arch/arm64/kernel/entry-ftr= ace.S index 025140caafe7..fc2805eb9e15 100644 --- a/arch/arm64/kernel/entry-ftrace.S +++ b/arch/arm64/kernel/entry-ftrace.S @@ -14,6 +14,72 @@ #include =20 #ifdef CONFIG_DYNAMIC_FTRACE_WITH_ARGS +/* + * Open-coded rcu_read_lock_trace() / rcu_read_unlock_trace(), see + * include/linux/rcupdate_trace.h and CONFIG_HAVE_RCU_TRAMPOLINE_READERS. = The + * whole of ftrace_caller is treated as trampoline text by the irq-exit ch= eck + * (see arch_rcu_tasks_trampoline_text()), so these only need to bracket t= he + * call out to ops->func; everything before the lock and after the unlock, + * including the direct-call tails that carry a BPF trampoline address in = x17, + * is covered by that. + * + * The SRCU-fast per-CPU increment is done LL/SC on this CPU's counter; be= ing + * migrated between reading the per-CPU offset and the store-exclusive only + * means another CPU's counter is (atomically) incremented, which SRCU sums + * over anyway. Ordering between the nesting count and the scp stash only + * matters against interrupts on this CPU, which observe program order. + * Clobbers x12-x16 and the flags. + */ + .macro trace_rcu_srcu_inc, addr:req, tmp:req, wtmp2:req +8888: ldxr \tmp, [\addr] + add \tmp, \tmp, #1 + stxr \wtmp2, \tmp, [\addr] + cbnz \wtmp2, 8888b + .endm + + .macro trace_rcu_read_lock +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS + mrs x12, sp_el0 // current + ldr w13, [x12, #TSK_TRC_READER_NESTING] + add w14, w13, #1 + str w14, [x12, #TSK_TRC_READER_NESTING] + cbnz w13, .Ltrl_nested\@ // interrupted a reader: done + ldr_l x13, rcu_tasks_trace_srcu_struct + SRCU_SRCU_CTRP + str x13, [x12, #TSK_TRC_READER_SCP] + get_this_cpu_offset x14 + add x14, x14, x13 + add x14, x14, #SRCU_CTR_SRCU_LOCKS + trace_rcu_srcu_inc x14, x15, w16 +#ifndef CONFIG_TASKS_TRACE_RCU_NO_MB + dmb ish +#endif +.Ltrl_nested\@: +#endif + .endm + + .macro trace_rcu_read_unlock +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS + mrs x12, sp_el0 // current + ldr w13, [x12, #TSK_TRC_READER_NESTING] + subs w13, w13, #1 + b.ne .Ltru_nested\@ + /* Outermost: pick up scp before an interrupt can see nesting =3D=3D 0. */ + ldr x14, [x12, #TSK_TRC_READER_SCP] + str wzr, [x12, #TSK_TRC_READER_NESTING] +#ifndef CONFIG_TASKS_TRACE_RCU_NO_MB + dmb ish +#endif + get_this_cpu_offset x15 + add x14, x14, x15 + add x14, x14, #SRCU_CTR_SRCU_UNLOCKS + trace_rcu_srcu_inc x14, x15, w16 + b .Ltru_done\@ +.Ltru_nested\@: + str w13, [x12, #TSK_TRC_READER_NESTING] +.Ltru_done\@: +#endif + .endm + /* * Due to -fpatchable-function-entry=3D2, the compiler has placed two NOPs= before * the regular function prologue. For an enabled callsite, ftrace_init_nop= () and @@ -94,6 +160,8 @@ SYM_CODE_START(ftrace_caller) stp x29, x30, [sp, #FREGS_SIZE] add x29, sp, #FREGS_SIZE =20 + trace_rcu_read_lock + /* Prepare arguments for the tracer func */ sub x0, x30, #AARCH64_INSN_SIZE // ip (callsite's BL insn) mov x1, x9 // parent_ip (callsite's LR) @@ -111,6 +179,8 @@ SYM_INNER_LABEL(ftrace_call, SYM_L_GLOBAL) bl ftrace_stub // func(ip, parent_ip, op, regs) #endif =20 + trace_rcu_read_unlock + /* * At the callsite x0-x8 and x19-x30 were live. Any C code will have prese= rved * x19-x29 per the AAPCS, and we created frame records upon entry, so we n= eed @@ -178,6 +248,10 @@ SYM_CODE_START(ftrace_stub_direct_tramp) SYM_CODE_END(ftrace_stub_direct_tramp) #endif /* CONFIG_DYNAMIC_FTRACE_WITH_DIRECT_CALLS */ =20 +/* End of [ftrace_caller, ...) for arch_rcu_tasks_trampoline_text(). */ +SYM_CODE_START(ftrace_static_tramp_end) +SYM_CODE_END(ftrace_static_tramp_end) + #else /* CONFIG_DYNAMIC_FTRACE_WITH_ARGS */ =20 /* diff --git a/arch/arm64/kernel/ftrace.c b/arch/arm64/kernel/ftrace.c index e1a3c0b3a051..5f4193f15cd9 100644 --- a/arch/arm64/kernel/ftrace.c +++ b/arch/arm64/kernel/ftrace.c @@ -17,6 +17,26 @@ #include #include =20 +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS +extern void ftrace_static_tramp_end(void); + +/* The SRCU-fast increments in entry-ftrace.S are the this_cpu_inc() form.= */ +static_assert(!IS_ENABLED(CONFIG_NEED_SRCU_NMI_SAFE)); + +/* + * See rcu_tasks_trampoline_text(). ftrace_caller and ftrace_stub_direct_= tramp + * are core kernel text but must be treated as trampolines: a task interru= pted + * in them outside the Tasks Trace reader may be carrying an ops pointer (= x11) + * or a direct-call BPF trampoline address (x17) whose lifetime is guarded= only + * by Tasks RCU. + */ +bool arch_rcu_tasks_trampoline_text(unsigned long ip) +{ + return ip >=3D (unsigned long)ftrace_caller && + ip < (unsigned long)ftrace_static_tramp_end; +} +#endif + #ifdef CONFIG_DYNAMIC_FTRACE_WITH_ARGS struct fregs_offset { const char *name; --=20 2.55.0 From nobody Thu Sep 24 17:09:51 2026 Received: from mail-qk2-f12.google.com (mail-qk2-f12.google.com [74.125.230.204]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 375C5399004 for ; Tue, 22 Sep 2026 02:31:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.204 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790044319; cv=none; b=tGlV5oES4Sn/Wc4fJUA5sLoyw2qIVty+qqJFHEVWgCr6mSp28gE0I9ZdBAD3bPnMFYAZpRWlpDQ3t5KbwPAxPLBpYoe2Bd2ByWRO4KKlzhoWXhmOR8+BXizbqtrppErBXeR9t4fMp4qSc0a6D4jQRyp+sTv5aOtbNn0uZmZcLSY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790044319; c=relaxed/simple; bh=DWFFSokO8NG0r0yBeEFehUgLH0GBe2lFI780YAKeFAM=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=r2F0rPcMEtSnkRrCZKaALyPJDhWHDahoSox8axEzED/CipbcAO67o53mAx+076pixNAwJgOv5chG6toO1brOYDyCnMje0BUr9wjjdSFrT7qUbIH4INwvhDaZBx16m9s7XZSFle58RvLquJnqkTOq3NSA2tUNxJjf7evCF6tjAG4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com; spf=pass smtp.mailfrom=toxicpanda.com; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b=dTBh52DT; arc=none smtp.client-ip=74.125.230.204 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b="dTBh52DT" Received: by mail-qk2-f12.google.com with SMTP id af79cd13be357-93910cc46c8so464552085a.0 for ; Mon, 21 Sep 2026 19:31:56 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=toxicpanda.com; s=google; t=1790044316; x=1790649116; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=OcQuspsT/gTElVk43OpaGWxo4+ZF80NPyngxmwv0iAI=; b=dTBh52DTgoCaf3XekPuztKNni2H6uYOfvrOGHWVo7ZZ0S4r2xd2lCmtcHxYOu0MKXp xhC+3Wt45K9vSOOlub868shQWVXtfvPLR2jmSdYxaSKGw/jx3og0hdO50/KbFTRyZyo2 Y18inx+7kwqzK/xo2TQnv9fF4vTAyb6N5/k+ubcU1zyWjElty0Qktew/j7prk950mFGy ZIh8r62kfHgPBj6b5T3z/+8sHIbqcKcMlTZOzVhl6gnH9lG/iyvo7ZkBUCR+EwFnhC22 iYnEbJruN9WZewY+kSSXYrKNPFNsgXoWkQ+YVWcuMf0MoFFlhfPulnbXzSmZ/TTE5tJu EBSQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790044316; x=1790649116; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=OcQuspsT/gTElVk43OpaGWxo4+ZF80NPyngxmwv0iAI=; b=EDiN+mdZB0+CQB5u/GEncfkmZ9f5KDAzWZkFnXzTTVcupRBLrqg9ARhKtrsxEn54UX 8onSaR2RsnuB93RH7JGDQ10oV/aZMrbaCcf0lpDi2BZjXGfTt8JX+E3PSk4BbQg7MCqf qa5kXk3xRXjIp+thVzLmstqmD+EU9Kj+Z70cOHXo0T8KB4cnOut0wiGCj2nzlKH1jFjY ro7u41aPKZJ11DubjOQDVdQCJVqPybWGQRY7McZX384acCShhKxFbe7X6z65ndAeweIW 733D2bUIp5H+1Vxd0MMPW0IyT5aoD0mooromsLdF8LbQhubolGRbnizjI7p8050HZh1d VscA== X-Forwarded-Encrypted: i=1; AKwUvBxOeP06s5w9bGlT0aZ31vzltefFa9aaiifWbIHATxtiFWtwxkUgGrZ75eEnOghnUe4f2RF+KT8/+shGrd0=@vger.kernel.org X-Gm-Message-State: AFuF++ndM+6ezQ+npzcWy9BlbgmCr7mC25ezaa62LHKPGnLI5z9Jlr1T qampq+XyDwBxzgknb4d/2F7qNiuXCRxBh0c46yBHeAwcGtD62RoWIZdO8z8G5QxpL7k= X-Gm-Gg: AYBFou2Kzf6jyOLn1cmgCnFeI/qVm7fzjMJSsZBS9Bm6MLYnp1wqNCW+6U7uDYVUjKE Ukv2mvZ5/Wj+HgOSDWI5xxvjSPbuEnteElPXUcDTgy5DscDPxbL2xvmUemiUI0VdByUW6mu2tgX HsdEgrTN58co72WbCIuERwlq0O0IzUEb7xM9x9CBLUH4n2iTZJMHqXLUBOtxOBI0jkBqlQwHV1G Zwol9RPhNKyz18+0kGinztRTXg6pWSLwzQvo+GNhZaGJ6L8Z0PkzcjGSgE4FUSCs+iEsK9WLBdS RIyZpULrcy73hcpnMHMHIvwrjfJqssii2dnKLfINwrvLLhmdVx/+9v9bd0BoeDqj+wuWvfeEvUB l1re+WxpNjqES4VKjEAU1KsNAEAXjjJhwwc3FbS3QBJooM4A4op7hPjFje83CvtzmmwQxvB0n3s jFjgYy80cVS0KVpQb25sfZFD6gHVQzjXQRi9SBQFefhHtjOM5dBPOMjVLxL45A411ZiOdZ0/w5A QM5ZBCONhNLGuY= X-Received: by 2002:a05:620a:4614:b0:936:dc4c:3bf4 with SMTP id af79cd13be357-93c198dc1c9mr159925485a.28.1790044315913; Mon, 21 Sep 2026 19:31:55 -0700 (PDT) Received: from toxicpanda.com ([153.61.196.243]) by smtp.gmail.com with ESMTPSA id af79cd13be357-93c1d03b170sm22998785a.16.2026.09.21.19.31.55 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 21 Sep 2026 19:31:55 -0700 (PDT) From: Josef Bacik Date: Tue, 22 Sep 2026 02:23:28 +0000 Subject: [PATCH v5 09/13] bpf, arm64: Take a Tasks Trace reader in the trampoline around its call-outs Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260922-b4-rcu-tasks-preempt-qs-v5-9-410f57770bad@toxicpanda.com> References: <20260922-b4-rcu-tasks-preempt-qs-v5-0-410f57770bad@toxicpanda.com> In-Reply-To: <20260922-b4-rcu-tasks-preempt-qs-v5-0-410f57770bad@toxicpanda.com> To: "Paul E. McKenney" , Frederic Weisbecker , Alexei Starovoitov , Steven Rostedt Cc: Boqun Feng , Masami Hiramatsu , Mark Rutland , Peter Zijlstra , Thomas Gleixner , Daniel Borkmann , Andrii Nakryiko , Puranjay Mohan , rcu@vger.kernel.org, bpf@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, Josef Bacik X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openssh-sha256; t=1790043820; l=7373; i=josef@toxicpanda.com; h=from:subject:message-id; bh=DWFFSokO8NG0r0yBeEFehUgLH0GBe2lFI780YAKeFAM=; b=U1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgUBr36M/n0nWN0DNbnxwzIiCZez6MG JiruuNaSCI/zXsAAAAGcGF0YXR0AAAAAAAAAAZzaGE1MTIAAABTAAAAC3NzaC1lZDI1NTE5AAAA QPVoMRhiz6gAdvKlENOzZ1c8jxAfTwJ9sVEsY25YoCvyFY/4f6gXn6G4z9PxCL6u0E2EccNLDn3 +KDAp6wx4zwk= X-Developer-Key: i=josef@toxicpanda.com; a=openssh; fpr=SHA256:C8kOX2QUJCMqnCX+KEeoqRAjLo9L+ELOSH2NSAJHqGA Same scheme as x86: have the arm64 BPF JIT open-code rcu_read_lock_trace() and rcu_read_unlock_trace() in the trampoline, one reader from after the callee-saved registers are stored to just before the original function is called (covering __bpf_tramp_enter() and the fentry/fmod_ret programs) and a second from just after it returns to just before those registers are restored (covering the fexit programs and __bpf_tramp_exit()), with the original function itself outside both and pinned by im->pcref. The second reader is entered before ip_after_call and the fmod_ret cbnz lands past it still holding the first, so exactly one is held on every path; trampolines without an original call get a single reader. The sequence mirrors entry-ftrace.S: current via sp_el0, trc_reader_nesting bumped, and for the outermost reader the SRCU-fast per-CPU counter incremented LL/SC and the counter pointer stashed in trc_reader_scp (plus the dmb when CONFIG_TASKS_TRACE_RCU_NO_MB is not set). x10-x15 are scratch at every emission point. Nothing is emitted on other configurations. Suggested-by: Alexei Starovoitov Assisted-by: LLM Signed-off-by: Josef Bacik --- arch/arm64/net/bpf_jit_comp.c | 93 +++++++++++++++++++++++++++++++++++++++= ++++ 1 file changed, 93 insertions(+) diff --git a/arch/arm64/net/bpf_jit_comp.c b/arch/arm64/net/bpf_jit_comp.c index c18e005a41db..87bbcca0298f 100644 --- a/arch/arm64/net/bpf_jit_comp.c +++ b/arch/arm64/net/bpf_jit_comp.c @@ -14,6 +14,7 @@ #include #include #include +#include #include =20 #include @@ -2591,6 +2592,77 @@ static void emit_arena_arg_conv(struct jit_ctx *ctx,= u8 dst, u8 src, bool nullab emit(A64_SUB(0, dst, src, base_lo), ctx); } =20 +/* + * Open-coded rcu_read_lock_trace() / rcu_read_unlock_trace() for the + * trampoline, see CONFIG_HAVE_RCU_TRAMPOLINE_READERS and the equivalent m= acros + * in arch/arm64/kernel/entry-ftrace.S. The SRCU-fast per-CPU increment i= s an + * LL/SC add on this CPU's counter; being migrated between reading the per= -CPU + * offset and the store-exclusive only means another CPU's counter is + * incremented atomically instead, which SRCU sums over anyway. Uses x10-= x15, + * which are scratch at every emission point, and the flags. + */ +static void emit_trace_rcu_reader(struct jit_ctx *ctx, bool lock) +{ +#ifdef CONFIG_TASKS_RCU_TRAMPOLINE_READERS + const int nesting_off =3D offsetof(struct task_struct, trc_reader_nesting= ); + const int scp_off =3D offsetof(struct task_struct, trc_reader_scp); + const int ctr_off =3D lock ? offsetof(struct srcu_ctr, srcu_locks) + : offsetof(struct srcu_ctr, srcu_unlocks); + const bool mb =3D !IS_ENABLED(CONFIG_TASKS_TRACE_RCU_NO_MB); + const u8 tsk =3D A64_R(10), n =3D A64_R(11), tmp =3D A64_R(12); + const u8 ctr =3D A64_R(13), addr =3D A64_R(14), val =3D A64_R(15); + + /* The plain this_cpu_inc() form of __srcu_read_lock_fast(). */ + BUILD_BUG_ON(IS_ENABLED(CONFIG_NEED_SRCU_NMI_SAFE)); + /* LDR/STR (immediate, unsigned offset) ranges */ + BUILD_BUG_ON((nesting_off & 3) || nesting_off >=3D SZ_16K || + (scp_off & 7) || scp_off >=3D SZ_32K); + + emit(A64_MRS_SP_EL0(tsk), ctx); /* current */ + emit(A64_LDR32I(n, tsk, nesting_off), ctx); + if (lock) { + emit(A64_ADD_I(0, tmp, n, 1), ctx); + emit(A64_STR32I(tmp, tsk, nesting_off), ctx); + /* interrupted a reader: done (branch to 1:) */ + emit(A64_CBNZ(0, n, 13 + mb), ctx); + /* scp =3D rcu_tasks_trace_srcu_struct.srcu_ctrp; current->trc_reader_sc= p =3D scp */ + emit_addr_mov_i64(ctr, (u64)&rcu_tasks_trace_srcu_struct.srcu_ctrp, ctx); + emit(A64_LDR64I(ctr, ctr, 0), ctx); + emit(A64_STR64I(ctr, tsk, scp_off), ctx); + } else { + emit(A64_SUB_I(0, n, n, 1), ctx); + /* still nested: just store the count (branch to 2:) */ + emit(A64_CBNZ(0, n, 11 + mb), ctx); + /* outermost: pick up scp before an interrupt can see nesting =3D=3D 0 */ + emit(A64_LDR64I(ctr, tsk, scp_off), ctx); + emit(A64_STR32I(A64_ZR, tsk, nesting_off), ctx); + if (mb) + emit(A64_DMB_ISH, ctx); + } + /* this_cpu_inc(scp->srcu_locks / scp->srcu_unlocks) */ + if (cpus_have_cap(ARM64_HAS_VIRT_HOST_EXTN)) + emit(A64_MRS_TPIDR_EL2(addr), ctx); + else + emit(A64_MRS_TPIDR_EL1(addr), ctx); + emit(A64_ADD(1, addr, addr, ctr), ctx); + emit(A64_ADD_I(1, addr, addr, ctr_off), ctx); + emit(A64_LDXR(1, val, addr), ctx); + emit(A64_ADD_I(1, val, val, 1), ctx); + emit(A64_STXR(1, val, addr, tmp), ctx); + emit(A64_CBNZ(0, tmp, -3), ctx); + if (lock) { + if (mb) + emit(A64_DMB_ISH, ctx); + /* 1: */ + } else { + emit(A64_B(2), ctx); + /* 2: */ + emit(A64_STR32I(n, tsk, nesting_off), ctx); + /* 3: */ + } +#endif +} + static void save_args(struct jit_ctx *ctx, int bargs_off, int oargs_off, const struct btf_func_model *m, const struct arg_aux *a, bool for_call_origin, bool is_struct_ops, u64 arena_base) @@ -2854,6 +2926,16 @@ static int prepare_trampoline(struct jit_ctx *ctx, s= truct bpf_tramp_image *im, emit(A64_STR64I(A64_R(19), A64_SP, regs_off), ctx); emit(A64_STR64I(A64_R(20), A64_SP, regs_off + 8), ctx); =20 + /* + * Tasks RCU keeps this image alive only while we are a Tasks Trace + * reader; the instructions before this point (and after the final + * unlock) are covered by the irq-exit IP check. One reader spans + * __bpf_tramp_enter() and the fentry/fmod_ret progs, a second one the + * fexit progs and __bpf_tramp_exit(); the original function runs + * outside both, with the image pinned by im->pcref instead. + */ + emit_trace_rcu_reader(ctx, true); + if (flags & BPF_TRAMP_F_CALL_ORIG) { /* for the first pass, assume the worst case */ if (!ctx->image) @@ -2898,12 +2980,20 @@ static int prepare_trampoline(struct jit_ctx *ctx, = struct bpf_tramp_image *im, if (flags & BPF_TRAMP_F_CALL_ORIG) { /* the original func takes kernel addresses, never converted ones */ save_args(ctx, bargs_off, oargs_off, m, a, true, is_struct_ops, 0); + emit_trace_rcu_reader(ctx, false); /* call original func */ emit(A64_LDR64I(A64_R(10), A64_SP, retaddr_off), ctx); emit(A64_ADR(A64_LR, AARCH64_INSN_SIZE * 2), ctx); emit(A64_RET(A64_R(10)), ctx); /* store return value */ emit(A64_STR64I(A64_R(0), A64_SP, retval_off), ctx); + /* + * Second reader. Taken before ip_after_call so that the branch + * to the epilogue patched in at teardown is inside it too; the + * fmod_ret early exit lands past this still holding the first + * reader, so either way exactly one is held. + */ + emit_trace_rcu_reader(ctx, true); /* reserve a nop for bpf_tramp_image_put */ im->ip_after_call =3D ctx->ro_image + ctx->idx; emit(A64_NOP, ctx); @@ -2945,6 +3035,9 @@ static int prepare_trampoline(struct jit_ctx *ctx, st= ruct bpf_tramp_image *im, if (flags & BPF_TRAMP_F_RESTORE_REGS) restore_args(ctx, bargs_off, a->regs_for_args); =20 + /* Remaining instructions are covered by the irq-exit IP check. */ + emit_trace_rcu_reader(ctx, false); + /* restore callee saved register x19 and x20 */ emit(A64_LDR64I(A64_R(19), A64_SP, regs_off), ctx); emit(A64_LDR64I(A64_R(20), A64_SP, regs_off + 8), ctx); --=20 2.55.0 From nobody Thu Sep 24 17:09:51 2026 Received: from mail-qv2-f42.google.com (mail-qv2-f42.google.com [74.125.230.170]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 01FD6382292 for ; Tue, 22 Sep 2026 02:32:46 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.170 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790044368; cv=none; b=GONqII9D8eh5+DcT/R4xNpM9wM81hJFyKf7VLXxMzvqF9VOigLcY+TWS0U95jj/6IQmE05/3D6SzM43Pheo9wDRNDmkmVGnUWw0MmyCXzUm9xaFtDtaWrFMDkPwv9yDrWSqSyGveTJs+/1CNCEalvE0iqgTuq1Dy5D+wYGfw7ZE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790044368; c=relaxed/simple; bh=I7rYeQZ4spsxQVfaE2Sayb+dFXGxTZqIS2lZCUAJvtU=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=dDObhWAta8YY0noxaiaNm9ayJUNWpgD0Z2UB9yK3iKUAZMfNU77E+nidkFVwTKSdFiHlV3AS0XVjoqDPsmSxarizAlEh8kK3ik1jzyTIkY54phzOTh6BkJIiRrV5o0bQvJm4gROMvyM1F9jG3vXD2F+a4RKfgq73YyNpw4bNi7s= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com; spf=pass smtp.mailfrom=toxicpanda.com; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b=stsEQMTS; arc=none smtp.client-ip=74.125.230.170 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b="stsEQMTS" Received: by mail-qv2-f42.google.com with SMTP id 6a1803df08f44-90cdfc6db0aso32274326d6.1 for ; Mon, 21 Sep 2026 19:32:46 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=toxicpanda.com; s=google; t=1790044366; x=1790649166; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=Nk/uI/WquaXkRgHHd/dzhScVzrCkrWpp6IlwsBXJcUg=; b=stsEQMTSXsH6TSaXgLPanfDG+ZB9lpdKb8Qdd/MQSQhEiAZ9Tsq4CzuncjAQMrBqY8 haGPpmbT9v4BGjvCjknX1P1frm+r/f9N5QmtFRhmp4bFyqebWatGWSm/Kkh3tTOjiKA9 J6gtUAz1SKrCgD+AANX5IbOlKlbDf1y0Ri53op/zGkR96yi8Jcyh6NhxeY9rMtIV5odH dMTH401e7p802G6mRqrJB7+Uswuv7vZsEi/gB4twsXD91CTzEQC3wvBpFtyCnOxGPh99 Q58SwoKnQ2IppVz0UnrJQ3igez3msApyIouP5HCaBJNZWmsyXaoZlIp0S+jLYT07C+Qd PA5A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790044366; x=1790649166; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=Nk/uI/WquaXkRgHHd/dzhScVzrCkrWpp6IlwsBXJcUg=; b=GhPDEqhCikC+NFGqcJXF+NjESIiKrRDVIC5oW6AyIgdYDQ8wwYgrsBPq3ro7ZA3SXC nqRUjUcf0MHOW+87lZYo51XL1LlmuYWRtX8pdWAm/bBTARYjYh/ZloIFISTmuOYraYgh bAbHvms1B7i3TvVu7dLgipHzgkgmfPTRPMGJgf0g6SEg0uaemzbWQQiRmgCXjRw+jkhd cevihyB1LpdS7tMzsirg4smTJ8AwJB9/ijwdMcvjvKlAe6Hjbjs+aJyxGpH5Z57pLlcg FfFAzzDq6Uq/KQI7qMj1hhEkME7gom3x6jHitqQGoajg7RHk1fkWbp2YJ6vNHSS/b5ve n85w== X-Forwarded-Encrypted: i=1; AKwUvBxM7WfS5jWzBclreE21jNYg7zOHp6hmO50STfqz8G7xiwWvZOFnk0DmEBSW0WfDOgCWz+kphQk1fzj05dM=@vger.kernel.org X-Gm-Message-State: AFuF++l1PU9S0Bh3Rza4HYeNNcIYOpM4CuGC73foUEbQNkXRe+WoKpwa pN74SZqRyBDw5QH832ZZidGT6dSJ1axb5OBAT7+f4/891UplUTsW6SpALLpBgwHyu3k= X-Gm-Gg: AYBFou1oXoTpQfDtfhdJijBPLONorQlKtUrxUFZQFropA2ZD2ofP/ntY6xz2idDxuix PfxvhWihtldl53ii7hT4YumgkYpGjNNJwQ2nH33YQ8t+ftRS53T0oSti5QHWt/uLiwdkSEII7W6 5HpyhqQdgYz4h4LgGDqDCEhUPd8i2nzfCF2MEFEZS+qFnWNjaJyAE/yMFtS9pqXumwh2h5eGwWG nWRrHujVxTVDuJQj+ByxWRzws4nH4FMkw/k4WHz6y2VpN/zvwslbDmoTfxAABKU02U3hWSfKn8I evAGDZH6+nAvIzQp8RnEwThknVGBDiJuiTkfn0vsXa+urk0V8ui93CowdZassrcjilyxGdrl3vH UGkKVyxIYvHE7BYFWiqJMoKtfEudp4HpyNhyFe9PA+W40Q4Et7PybiEQg6XUSs59V+BTpq+B78L YTzYgXkZRL3t5byM6cWJcizsY/kua0kB3pSrk3cPt5qcoWxDtHaK72izP6nb6Rj/x1wx5uxvS2l cTEvH6gNdlTQjw= X-Received: by 2002:a05:6214:5017:b0:911:9385:8581 with SMTP id 6a1803df08f44-913fc948279mr35279286d6.28.1790044365561; Mon, 21 Sep 2026 19:32:45 -0700 (PDT) Received: from toxicpanda.com ([153.61.196.250]) by smtp.gmail.com with ESMTPSA id 6a1803df08f44-914024c20e0sm4740526d6.21.2026.09.21.19.32.44 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 21 Sep 2026 19:32:44 -0700 (PDT) From: Josef Bacik Date: Tue, 22 Sep 2026 02:23:29 +0000 Subject: [PATCH v5 10/13] samples: ftrace: Make the direct-call trampolines Tasks Trace readers Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260922-b4-rcu-tasks-preempt-qs-v5-10-410f57770bad@toxicpanda.com> References: <20260922-b4-rcu-tasks-preempt-qs-v5-0-410f57770bad@toxicpanda.com> In-Reply-To: <20260922-b4-rcu-tasks-preempt-qs-v5-0-410f57770bad@toxicpanda.com> To: "Paul E. McKenney" , Frederic Weisbecker , Alexei Starovoitov , Steven Rostedt Cc: Boqun Feng , Masami Hiramatsu , Mark Rutland , Peter Zijlstra , Thomas Gleixner , Daniel Borkmann , Andrii Nakryiko , Puranjay Mohan , rcu@vger.kernel.org, bpf@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, Josef Bacik X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openssh-sha256; t=1790043820; l=13193; i=josef@toxicpanda.com; h=from:subject:message-id; bh=I7rYeQZ4spsxQVfaE2Sayb+dFXGxTZqIS2lZCUAJvtU=; b=U1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgUBr36M/n0nWN0DNbnxwzIiCZez6MG JiruuNaSCI/zXsAAAAGcGF0YXR0AAAAAAAAAAZzaGE1MTIAAABTAAAAC3NzaC1lZDI1NTE5AAAA QM+2t5V2gfw8IUSy1qDoDGQkdXuWXVx8UOA8l4idX6itSYIr+3cJvizuGAn0kC3lS3jW3yYg3Kb 0Ft51r6OGEA0= X-Developer-Key: i=josef@toxicpanda.com; a=openssh; fpr=SHA256:C8kOX2QUJCMqnCX+KEeoqRAjLo9L+ELOSH2NSAJHqGA The sample direct trampolines are exactly the kind of out-of-line register_ftrace_direct() user whose lifetime depends on Tasks RCU waiting for a task inside them: nothing else stops rmmod while a task is preempted in my_direct_func(). On HAVE_RCU_TRAMPOLINE_READERS architectures that wait only covers Tasks Trace RCU readers, so give the samples a small shared header with rcu_read_lock_trace() and rcu_read_unlock_trace() open-coded as instruction strings for x86-64 and arm64 -- the same sequences as ftrace_64.S and entry-ftrace.S, using caller-saved non-argument scratch registers -- and bracket every call-out with them. The instructions outside the bracket are module text, covered by ftrace_direct_mark_module(). Other architectures get empty definitions. Assisted-by: LLM Signed-off-by: Josef Bacik --- samples/ftrace/ftrace-direct-modify.c | 9 ++ samples/ftrace/ftrace-direct-multi-modify.c | 9 ++ samples/ftrace/ftrace-direct-multi.c | 5 ++ samples/ftrace/ftrace-direct-too.c | 5 ++ samples/ftrace/ftrace-direct.c | 5 ++ samples/ftrace/ftrace-direct.h | 126 ++++++++++++++++++++++++= ++++ 6 files changed, 159 insertions(+) diff --git a/samples/ftrace/ftrace-direct-modify.c b/samples/ftrace/ftrace-= direct-modify.c index 164d9dd6fd92..937c8d8c2a1b 100644 --- a/samples/ftrace/ftrace-direct-modify.c +++ b/samples/ftrace/ftrace-direct-modify.c @@ -2,6 +2,7 @@ #include #include #include +#include "ftrace-direct.h" #if !defined(CONFIG_ARM64) && !defined(CONFIG_PPC32) #include #endif @@ -73,7 +74,9 @@ asm ( " pushq %rbp\n" " movq %rsp, %rbp\n" CALL_DEPTH_ACCOUNT + TRACE_RCU_READ_LOCK " call my_direct_func1\n" + TRACE_RCU_READ_UNLOCK " leave\n" ASM_RET " .size my_tramp1, .-my_tramp1\n" @@ -85,7 +88,9 @@ asm ( " pushq %rbp\n" " movq %rsp, %rbp\n" CALL_DEPTH_ACCOUNT + TRACE_RCU_READ_LOCK " call my_direct_func2\n" + TRACE_RCU_READ_UNLOCK " leave\n" ASM_RET " .size my_tramp2, .-my_tramp2\n" @@ -141,11 +146,13 @@ asm ( " .globl my_tramp1\n" " my_tramp1:" " hint 34\n" // bti c + TRACE_RCU_READ_LOCK " sub sp, sp, #16\n" " stp x9, x30, [sp]\n" " bl my_direct_func1\n" " ldp x30, x9, [sp]\n" " add sp, sp, #16\n" + TRACE_RCU_READ_UNLOCK " ret x9\n" " .size my_tramp1, .-my_tramp1\n" =20 @@ -153,11 +160,13 @@ asm ( " .globl my_tramp2\n" " my_tramp2:" " hint 34\n" // bti c + TRACE_RCU_READ_LOCK " sub sp, sp, #16\n" " stp x9, x30, [sp]\n" " bl my_direct_func2\n" " ldp x30, x9, [sp]\n" " add sp, sp, #16\n" + TRACE_RCU_READ_UNLOCK " ret x9\n" " .size my_tramp2, .-my_tramp2\n" " .popsection\n" diff --git a/samples/ftrace/ftrace-direct-multi-modify.c b/samples/ftrace/f= trace-direct-multi-modify.c index b03766c6217b..e12e5c8b83f0 100644 --- a/samples/ftrace/ftrace-direct-multi-modify.c +++ b/samples/ftrace/ftrace-direct-multi-modify.c @@ -2,6 +2,7 @@ #include #include #include +#include "ftrace-direct.h" #if !defined(CONFIG_ARM64) && !defined(CONFIG_PPC32) #include #endif @@ -77,10 +78,12 @@ asm ( " pushq %rbp\n" " movq %rsp, %rbp\n" CALL_DEPTH_ACCOUNT + TRACE_RCU_READ_LOCK " pushq %rdi\n" " movq 8(%rbp), %rdi\n" " call my_direct_func1\n" " popq %rdi\n" + TRACE_RCU_READ_UNLOCK " leave\n" ASM_RET " .size my_tramp1, .-my_tramp1\n" @@ -92,10 +95,12 @@ asm ( " pushq %rbp\n" " movq %rsp, %rbp\n" CALL_DEPTH_ACCOUNT + TRACE_RCU_READ_LOCK " pushq %rdi\n" " movq 8(%rbp), %rdi\n" " call my_direct_func2\n" " popq %rdi\n" + TRACE_RCU_READ_UNLOCK " leave\n" ASM_RET " .size my_tramp2, .-my_tramp2\n" @@ -154,6 +159,7 @@ asm ( " .globl my_tramp1\n" " my_tramp1:" " hint 34\n" // bti c + TRACE_RCU_READ_LOCK " sub sp, sp, #32\n" " stp x9, x30, [sp]\n" " str x0, [sp, #16]\n" @@ -162,6 +168,7 @@ asm ( " ldp x30, x9, [sp]\n" " ldr x0, [sp, #16]\n" " add sp, sp, #32\n" + TRACE_RCU_READ_UNLOCK " ret x9\n" " .size my_tramp1, .-my_tramp1\n" =20 @@ -169,6 +176,7 @@ asm ( " .globl my_tramp2\n" " my_tramp2:" " hint 34\n" // bti c + TRACE_RCU_READ_LOCK " sub sp, sp, #32\n" " stp x9, x30, [sp]\n" " str x0, [sp, #16]\n" @@ -177,6 +185,7 @@ asm ( " ldp x30, x9, [sp]\n" " ldr x0, [sp, #16]\n" " add sp, sp, #32\n" + TRACE_RCU_READ_UNLOCK " ret x9\n" " .size my_tramp2, .-my_tramp2\n" " .popsection\n" diff --git a/samples/ftrace/ftrace-direct-multi.c b/samples/ftrace/ftrace-d= irect-multi.c index 3fe6ddaf0b69..a970464ed378 100644 --- a/samples/ftrace/ftrace-direct-multi.c +++ b/samples/ftrace/ftrace-direct-multi.c @@ -3,6 +3,7 @@ =20 #include /* for handle_mm_fault() */ #include +#include "ftrace-direct.h" #include #if !defined(CONFIG_ARM64) && !defined(CONFIG_PPC32) #include @@ -56,10 +57,12 @@ asm ( " pushq %rbp\n" " movq %rsp, %rbp\n" CALL_DEPTH_ACCOUNT + TRACE_RCU_READ_LOCK " pushq %rdi\n" " movq 8(%rbp), %rdi\n" " call my_direct_func\n" " popq %rdi\n" + TRACE_RCU_READ_UNLOCK " leave\n" ASM_RET " .size my_tramp, .-my_tramp\n" @@ -101,6 +104,7 @@ asm ( " .globl my_tramp\n" " my_tramp:" " hint 34\n" // bti c + TRACE_RCU_READ_LOCK " sub sp, sp, #32\n" " stp x9, x30, [sp]\n" " str x0, [sp, #16]\n" @@ -109,6 +113,7 @@ asm ( " ldp x30, x9, [sp]\n" " ldr x0, [sp, #16]\n" " add sp, sp, #32\n" + TRACE_RCU_READ_UNLOCK " ret x9\n" " .size my_tramp, .-my_tramp\n" " .popsection\n" diff --git a/samples/ftrace/ftrace-direct-too.c b/samples/ftrace/ftrace-dir= ect-too.c index bf2411aa6fd7..abc098c2ab7a 100644 --- a/samples/ftrace/ftrace-direct-too.c +++ b/samples/ftrace/ftrace-direct-too.c @@ -3,6 +3,7 @@ =20 #include /* for handle_mm_fault() */ #include +#include "ftrace-direct.h" #if !defined(CONFIG_ARM64) && !defined(CONFIG_PPC32) #include #endif @@ -61,6 +62,7 @@ asm ( " pushq %rbp\n" " movq %rsp, %rbp\n" CALL_DEPTH_ACCOUNT + TRACE_RCU_READ_LOCK " pushq %rdi\n" " pushq %rsi\n" " pushq %rdx\n" @@ -70,6 +72,7 @@ asm ( " popq %rdx\n" " popq %rsi\n" " popq %rdi\n" + TRACE_RCU_READ_UNLOCK " leave\n" ASM_RET " .size my_tramp, .-my_tramp\n" @@ -110,6 +113,7 @@ asm ( " .globl my_tramp\n" " my_tramp:" " hint 34\n" // bti c + TRACE_RCU_READ_LOCK " sub sp, sp, #48\n" " stp x9, x30, [sp]\n" " stp x0, x1, [sp, #16]\n" @@ -119,6 +123,7 @@ asm ( " ldp x0, x1, [sp, #16]\n" " ldp x2, x3, [sp, #32]\n" " add sp, sp, #48\n" + TRACE_RCU_READ_UNLOCK " ret x9\n" " .size my_tramp, .-my_tramp\n" " .popsection\n" diff --git a/samples/ftrace/ftrace-direct.c b/samples/ftrace/ftrace-direct.c index 5368c8c39cbb..99b65ad2fccc 100644 --- a/samples/ftrace/ftrace-direct.c +++ b/samples/ftrace/ftrace-direct.c @@ -3,6 +3,7 @@ =20 #include /* for wake_up_process() */ #include +#include "ftrace-direct.h" #if !defined(CONFIG_ARM64) && !defined(CONFIG_PPC32) #include #endif @@ -54,9 +55,11 @@ asm ( " pushq %rbp\n" " movq %rsp, %rbp\n" CALL_DEPTH_ACCOUNT + TRACE_RCU_READ_LOCK " pushq %rdi\n" " call my_direct_func\n" " popq %rdi\n" + TRACE_RCU_READ_UNLOCK " leave\n" ASM_RET " .size my_tramp, .-my_tramp\n" @@ -97,6 +100,7 @@ asm ( " .globl my_tramp\n" " my_tramp:" " hint 34\n" // bti c + TRACE_RCU_READ_LOCK " sub sp, sp, #32\n" " stp x9, x30, [sp]\n" " str x0, [sp, #16]\n" @@ -104,6 +108,7 @@ asm ( " ldp x30, x9, [sp]\n" " ldr x0, [sp, #16]\n" " add sp, sp, #32\n" + TRACE_RCU_READ_UNLOCK " ret x9\n" " .size my_tramp, .-my_tramp\n" " .popsection\n" diff --git a/samples/ftrace/ftrace-direct.h b/samples/ftrace/ftrace-direct.h new file mode 100644 index 000000000000..726f67048ff5 --- /dev/null +++ b/samples/ftrace/ftrace-direct.h @@ -0,0 +1,126 @@ +/* SPDX-License-Identifier: GPL-2.0-only */ +#ifndef _SAMPLES_FTRACE_DIRECT_H +#define _SAMPLES_FTRACE_DIRECT_H + +#include + +/* + * A direct-call trampoline is entered with no lock, refcount or RCU marker + * held; only Tasks RCU keeps it (and, for a module, its text) alive while= a + * task is inside it or preempted in something it called. On architectures + * that select HAVE_RCU_TRAMPOLINE_READERS, Tasks RCU only waits for such a + * task while it is a Tasks Trace RCU reader, so the trampoline must enter= one + * before calling out and leave it afterwards, exactly like the ftrace and= BPF + * trampolines do. See register_ftrace_direct(). The instructions before= the + * lock and after the unlock are covered by ftrace_direct_mark_module(). + * + * These are rcu_read_lock_trace() / rcu_read_unlock_trace() open-coded as + * instruction strings for use inside the samples' asm() trampolines, after + * the versions in arch/x86/kernel/ftrace_64.S and + * arch/arm64/kernel/entry-ftrace.S. The scratch registers are caller-sav= ed + * and not argument registers, so they are dead on entry to and exit from = an + * fentry trampoline; the flags are clobbered. + * + * The generated asm-offsets.h is only pulled in on the architectures that= need + * it here: it is not generally safe to include from C (e.g. PPC32's TASK_= SIZE + * and arm64's TRAMP_VALIAS clash with the C definitions). + */ +#if defined(CONFIG_TASKS_RCU_TRAMPOLINE_READERS) && defined(CONFIG_X86_64) + +#include + +#ifndef CONFIG_TASKS_TRACE_RCU_NO_MB +#define TRACE_RCU_MB " lock addl $0, -4(%rsp)\n" +#else +#define TRACE_RCU_MB +#endif + +#define TRACE_RCU_READ_LOCK \ + " movq %gs:current_task(%rip), %r11\n" \ + " movl " __stringify(TASK_trc_reader_nesting) "(%r11), %r10d\n" \ + " incl " __stringify(TASK_trc_reader_nesting) "(%r11)\n" \ + " testl %r10d, %r10d\n" \ + " jnz 771f\n" \ + " movq rcu_tasks_trace_srcu_struct+" __stringify(SRCU_srcu_ctrp) "(%rip),= %r10\n" \ + " incq %gs:" __stringify(SRCU_CTR_srcu_locks) "(%r10)\n" \ + " movq %r10, " __stringify(TASK_trc_reader_scp) "(%r11)\n" \ + TRACE_RCU_MB \ + "771:\n" + +#define TRACE_RCU_READ_UNLOCK \ + " movq %gs:current_task(%rip), %r11\n" \ + " movl " __stringify(TASK_trc_reader_nesting) "(%r11), %r10d\n" \ + " subl $1, %r10d\n" \ + " jnz 772f\n" \ + " movq " __stringify(TASK_trc_reader_scp) "(%r11), %r10\n" \ + " movl $0, " __stringify(TASK_trc_reader_nesting) "(%r11)\n" \ + TRACE_RCU_MB \ + " incq %gs:" __stringify(SRCU_CTR_srcu_unlocks) "(%r10)\n" \ + " jmp 773f\n" \ + "772: movl %r10d, " __stringify(TASK_trc_reader_nesting) "(%r11)\n" \ + "773:\n" + +#elif defined(CONFIG_TASKS_RCU_TRAMPOLINE_READERS) && defined(CONFIG_ARM64) + +#include +#include +/* arm64's asm-offsets.h redefines TRAMP_VALIAS from . */ +#pragma push_macro("TRAMP_VALIAS") +#undef TRAMP_VALIAS +#include +#pragma pop_macro("TRAMP_VALIAS") + +#ifndef CONFIG_TASKS_TRACE_RCU_NO_MB +#define TRACE_RCU_MB " dmb ish\n" +#else +#define TRACE_RCU_MB +#endif + +#define TRACE_RCU_SRCU_CTRP "rcu_tasks_trace_srcu_struct+" __stringify(SRC= U_SRCU_CTRP) + +/* x14 =3D this CPU's offset; then atomically increment the long at x14 + = \areg */ +#define TRACE_RCU_PERCPU_INC(areg) \ + ALTERNATIVE(" mrs x14, tpidr_el1\n", " mrs x14, tpidr_el2\n", \ + ARM64_HAS_VIRT_HOST_EXTN) \ + " add x14, x14, " areg "\n" \ + "778: ldxr x15, [x14]\n" \ + " add x15, x15, #1\n" \ + " stxr w16, x15, [x14]\n" \ + " cbnz w16, 778b\n" + +#define TRACE_RCU_READ_LOCK \ + " mrs x12, sp_el0\n" \ + " ldr w13, [x12, #" __stringify(TSK_TRC_READER_NESTING) "]\n" \ + " add w14, w13, #1\n" \ + " str w14, [x12, #" __stringify(TSK_TRC_READER_NESTING) "]\n" \ + " cbnz w13, 771f\n" \ + " adrp x13, " TRACE_RCU_SRCU_CTRP "\n" \ + " ldr x13, [x13, #:lo12:" TRACE_RCU_SRCU_CTRP "]\n" \ + " str x13, [x12, #" __stringify(TSK_TRC_READER_SCP) "]\n" \ + " add x13, x13, #" __stringify(SRCU_CTR_SRCU_LOCKS) "\n" \ + TRACE_RCU_PERCPU_INC("x13") \ + TRACE_RCU_MB \ + "771:\n" + +#define TRACE_RCU_READ_UNLOCK \ + " mrs x12, sp_el0\n" \ + " ldr w13, [x12, #" __stringify(TSK_TRC_READER_NESTING) "]\n" \ + " subs w13, w13, #1\n" \ + " b.ne 772f\n" \ + " ldr x13, [x12, #" __stringify(TSK_TRC_READER_SCP) "]\n" \ + " str wzr, [x12, #" __stringify(TSK_TRC_READER_NESTING) "]\n" \ + TRACE_RCU_MB \ + " add x13, x13, #" __stringify(SRCU_CTR_SRCU_UNLOCKS) "\n" \ + TRACE_RCU_PERCPU_INC("x13") \ + " b 773f\n" \ + "772: str w13, [x12, #" __stringify(TSK_TRC_READER_NESTING) "]\n" \ + "773:\n" + +#else + +#define TRACE_RCU_READ_LOCK +#define TRACE_RCU_READ_UNLOCK + +#endif + +#endif /* _SAMPLES_FTRACE_DIRECT_H */ --=20 2.55.0 From nobody Thu Sep 24 17:09:51 2026 Received: from mail-qk2-f12.google.com (mail-qk2-f12.google.com [74.125.230.204]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 577C038737A for ; Tue, 22 Sep 2026 02:33:34 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.204 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790044416; cv=none; b=cWNrwBe2kWz+DyQ8V8/NdKRU5c/hkmiT0XOJN44bXLqBbfqHHdJ53AYiSlhuWYZ6iPstmPHGTD/6sHNuOi/botSPB6U+vH+g1eeN8pxKLRa1EO6DaF1my8VXUagozx+4sR/NA97F6W6HI6OIuy0h/MeqLt2Kid00HUxXkkKAvWk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790044416; c=relaxed/simple; bh=O8YdW4YSHUYyF3GwEQN16WzJM7UkZZFfh8RzVwv7xJM=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=dWy/pEwKOxtHeh0l47q295bjVxhEvvhnbU2NZVs/mVJouNZ2Qf0Kjyp7n8ocvpYiZGWYicxnEgIyXi1WfiMaeWcZfM0LhtD+EVw15qWTLCL43I5BsVT7EjKb+4VHjRbDpX5YvPN1xaIbUQCwls3wHzE+AJUyNKuJKy6mLoHX2tY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com; spf=pass smtp.mailfrom=toxicpanda.com; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b=qnop8KF3; arc=none smtp.client-ip=74.125.230.204 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b="qnop8KF3" Received: by mail-qk2-f12.google.com with SMTP id d75a77b69052e-530d8a00bcdso23239981cf.2 for ; Mon, 21 Sep 2026 19:33:34 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=toxicpanda.com; s=google; t=1790044413; x=1790649213; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=MkP56mJgKixcgLXoBxrfzz0NyJo14L+hjKFrNfKE/2k=; b=qnop8KF3/xW7MwL1cIJXYaz6P00OjjBFrfHjtnUWJtITYukvpJlNeH67b5oRueMoeN +9iNeNuROHtNBwVOTMxFGCYqdea5MgnfuzTmuV7+Uyo0h97iJRgpdNvB1Zl3SXORxe/6 UFXPHDHz2UTjAK3LnoIlPD0HJW460Ck1CYDtKPikLXiurZkm7bRXd9XLl7PMJSVJGqPu Vm/c99t9vVQgjsuKmh7g9c+oi5FWSwaSdqjES/dhgofRGDi6Iu9GPbDAPpR1Q44b/oMv a2JxoQUU+ci+gBPzSJpVIHq/QWMAcxnJYplsZ3xVsgDXF/VJ5YBP+po7IIbHbnom3PE8 iIUA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790044413; x=1790649213; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=MkP56mJgKixcgLXoBxrfzz0NyJo14L+hjKFrNfKE/2k=; b=UR/GEMaPYKo3r6UKU/0/4cAFS/xIAYvQdC4Dm+9Vxx6D02TMZfmUpx0iY1MOScnJXJ QhkSnF+NIBIss2IBrVgLHVneOacn51jQ5mWJv1uqXev9OdMwLFeoc50gnQnRekWBaVOo Jm/KECAMtrD9zRuFLAnZXULd4mGztXdySZnvCdTkmIRJQ2QIUOtij77qeblXvRjopfCr Wa94Xo6aP9dEuuHydNv+GrS2B9tQFnfIrZNoiFyB0WK2dVxMiXSsbrz8ILeAL3X36Tah isAZPSJCovYXOGbAuoOPYDa4s+Wk6suZNZBYdBTBrlq6xUeHMJmw9Z4KQq38KvXdUSHO vVOw== X-Forwarded-Encrypted: i=1; AKwUvBxmG7ylCFW62Ygk4KxElWBmqjFAfm+YhlLtMMLAJO+wfzDKHAthKwp9/csZolD4btp0HXwD4bS8PWvRD5Q=@vger.kernel.org X-Gm-Message-State: AFuF++khKZrKfVG5NiLVhY2n4bc+YkwQ5/M12ghkQrweLsOQ7uX5j6BT KU2epspgIIBdkNXbQl6mt+/zXi5C3bDN9rVyO2SME54xxEz8maGyf0IgGi7ouIztZew= X-Gm-Gg: AYBFou1iyqjll7Dn7ZhrVLTaCnj/dQFRNPxMdvawAx27Lo5u4TUbtclhPzi30uFJuc5 1r5Gy+zJ95oRL+XoXO2+hSu6g+TJZvOHMtHUFn+7j4QrvPozTgPI78F3SRjR/H56GY9rNAP9BDQ pxZ5LdPz2uRNYGhLRs+bCNm7HyrLGO6269fIh5ON9YJT2a0SfKUUKbmu7blfhkKfZhTu+FA1pvj UozI43aioJ953GRSCoLPgjAzki7vzB8+9KVURghsjHCKBJk0Sx9yoeWaFgkFEXWBqPawKEbQP+r Pp3VdLJYJTjNVihK2xMVL5Ovx4LkX6MMAh3gos3BFJL0jkoQcTQQ3t95jdyJ44KErmz2tBU3Vkf Tu/+C9UN/1R0w4iWftej7MK1Bkmiwk0sbofv+2OxIGMUnL/cNZIAsq6TYYB8anz+vr6Pg6ljVn0 9SlAp5kqBmJnhLW3SNmqU/akeF/zSUz0ErEIcB7fAGFkqsAgPpt+BYXEvef17GwcLlVk3gjtC0G gnq57D/px8qIncs X-Received: by 2002:ac8:5a12:0:b0:52d:cda5:abfd with SMTP id d75a77b69052e-532d8e10f9amr37728211cf.38.1790044412692; Mon, 21 Sep 2026 19:33:32 -0700 (PDT) Received: from toxicpanda.com ([153.61.196.250]) by smtp.gmail.com with ESMTPSA id 6a1803df08f44-914024ee211sm4656006d6.34.2026.09.21.19.33.32 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 21 Sep 2026 19:33:32 -0700 (PDT) From: Josef Bacik Date: Tue, 22 Sep 2026 02:23:30 +0000 Subject: [PATCH v5 11/13] rcutorture: Make Tasks RCU readers Tasks Trace readers where required Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260922-b4-rcu-tasks-preempt-qs-v5-11-410f57770bad@toxicpanda.com> References: <20260922-b4-rcu-tasks-preempt-qs-v5-0-410f57770bad@toxicpanda.com> In-Reply-To: <20260922-b4-rcu-tasks-preempt-qs-v5-0-410f57770bad@toxicpanda.com> To: "Paul E. McKenney" , Frederic Weisbecker , Alexei Starovoitov , Steven Rostedt Cc: Boqun Feng , Masami Hiramatsu , Mark Rutland , Peter Zijlstra , Thomas Gleixner , Daniel Borkmann , Andrii Nakryiko , Puranjay Mohan , rcu@vger.kernel.org, bpf@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, Josef Bacik X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openssh-sha256; t=1790043820; l=1716; i=josef@toxicpanda.com; h=from:subject:message-id; bh=O8YdW4YSHUYyF3GwEQN16WzJM7UkZZFfh8RzVwv7xJM=; b=U1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgUBr36M/n0nWN0DNbnxwzIiCZez6MG JiruuNaSCI/zXsAAAAGcGF0YXR0AAAAAAAAAAZzaGE1MTIAAABTAAAAC3NzaC1lZDI1NTE5AAAA QN+SSmbqaUDY715i7pHjObQRKDl/n4FuMBV449sPKr77qotFd+q5BWjod9YbtBMN8iSeWuFGbqu 3faoKTTHOvw4= X-Developer-Key: i=josef@toxicpanda.com; a=openssh; fpr=SHA256:C8kOX2QUJCMqnCX+KEeoqRAjLo9L+ELOSH2NSAJHqGA rcutorture's "tasks" flavor has empty readlock/readunlock hooks because a classic Tasks RCU reader is simply code that does not block. Under CONFIG_TASKS_RCU_TRAMPOLINE_READERS a preemption outside trampoline text is also a quiescent state, and the thing real readers (trampolines) do to stay protected across their call-outs is take rcu_read_lock_trace(), so have the torture readers do the same there. Otherwise a preempted torture reader would rightly be treated as quiescent and the test would report false-positive too-short grace periods. Assisted-by: LLM Signed-off-by: Josef Bacik --- kernel/rcu/rcutorture.c | 10 ++++++++++ 1 file changed, 10 insertions(+) diff --git a/kernel/rcu/rcutorture.c b/kernel/rcu/rcutorture.c index 794937e13e7c..ab870ef09af0 100644 --- a/kernel/rcu/rcutorture.c +++ b/kernel/rcu/rcutorture.c @@ -1142,13 +1142,23 @@ static struct rcu_torture_ops trivial_preempt_ops = =3D { * Definitions for RCU-tasks torture testing. */ =20 +/* + * A classic Tasks RCU reader is any stretch of kernel code that does not + * voluntarily block. With CONFIG_TASKS_RCU_TRAMPOLINE_READERS a preempti= on + * outside trampoline text also ends it, and what a trampoline does to stay + * protected across its call-out is take a Tasks Trace reader, so model th= at. + */ static int tasks_torture_read_lock(void) { + if (IS_ENABLED(CONFIG_TASKS_RCU_TRAMPOLINE_READERS)) + rcu_read_lock_trace(); return 0; } =20 static void tasks_torture_read_unlock(int idx) { + if (IS_ENABLED(CONFIG_TASKS_RCU_TRAMPOLINE_READERS)) + rcu_read_unlock_trace(); } =20 static void rcu_tasks_torture_deferred_free(struct rcu_torture *p) --=20 2.55.0 From nobody Thu Sep 24 17:09:51 2026 Received: from mail-qv2-f35.google.com (mail-qv2-f35.google.com [74.125.230.163]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 839A737F8BA for ; Tue, 22 Sep 2026 02:34:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.163 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790044451; cv=none; b=LKt289Q4rri+72yLNrHs722ClFjOIONDX27F1xxqDp+kB8QOWGdpEpGQ67CoMGyymdIHpJRGeE9gWKer+jwpy9bYqz91t1UNlyrwzr1KYCT9NnOq0Up3oJOP6EjGv1wrDJ4/szTLLWBQ7KiYcQAtyllXv0RppoxifP4UeUYQfrQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790044451; c=relaxed/simple; bh=EMsOGg9VRxBcocei3P3dGUkU8foVu6cO35v23va9IOs=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=LuSScAWyskN8ChugKjXnok6D8dVDbM+BLvs2z0wcz6I4H1Ur67gPvSZ2hxH8sIEEJLesi+amXBXCk84n9/812HTRApQHfYKrb4we/sGo3ZXlzUF7U9YkTZnxOoLVeoenXRURNqffC0GB3Uaj+AM3lC+i5kEKQ4Yl8Mo5EE28WXY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com; spf=pass smtp.mailfrom=toxicpanda.com; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b=gxu6iCps; arc=none smtp.client-ip=74.125.230.163 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b="gxu6iCps" Received: by mail-qv2-f35.google.com with SMTP id 6a1803df08f44-9104e54b1abso49337406d6.1 for ; Mon, 21 Sep 2026 19:34:09 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=toxicpanda.com; s=google; t=1790044448; x=1790649248; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=vixDe7y/8fdtCmwAQ+Tlt/gsuGiEiA9j8JVU0h22FWI=; b=gxu6iCpsUuLPXgYUiCTsOhdCO0oUcbJMBjMPpP+jv0eOHhhCXhGwUeIs2HbNQ20DHx 74qKmviwa3RDrCJ7/9SzC3LYu0B9qZiq+5geSmAmd1lXncRixovrhMmqKpR0VohQ1W0Q jTaNOe2nZxjJVrveqr8GHCCcHdZWmjGs5eqXGSVA08vZ/Cy4PVwRnkqczI8d4de7SEBh PIl1D3KJ8rVTur1Y2QIIEAXJZviOyS0YuCtPLU/DFmzblAiFHYxUvd3ejTUgCcd4nLFr qGMy8pp96s6aeeS5/W0b4ATIkBgLliiR2U1+mMJ7ruU1y5jxhOgCgcRf219wL53YetDD dafw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790044448; x=1790649248; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=vixDe7y/8fdtCmwAQ+Tlt/gsuGiEiA9j8JVU0h22FWI=; b=gBsPf1sebBYqUu+LioKwoE3dDepuB2SAxV+Sysz/xnOCEXcQvHzpV2fU/1ba4sGjYG YHwrcuxzM6z96Oiy0tPfvPe0vOz9mDS/jBFMxYItAOD6aR8nNgVhHyP0nsrpox2bZAsz sddNS1ftvfMXb2Ddag7LJpQ8CbuNO7jb6xBiz7WjeKjP7XECdPOFSaefTkGiLC6VG9nc AtcKPUEuK7GUJ6dX5VpIkpCuqrpREx8lB/F8v6PYElCUduUUYJl9DiQWhRGm0sZEkHAV 5aObsFNe8TXY97JpbqghmLbA5nuBRfHV0vugV2051bVg4WMxx/z3R8RIGdeZigfRFS5r Ggpw== X-Forwarded-Encrypted: i=1; AKwUvBx8FSchPgsOGC9Jg1nSSJFGB3JKBt9+AlSJsZHiIHJhgLO7382o8UpXYpi9eOqyQwC9eNZVpAyPp7JPaUQ=@vger.kernel.org X-Gm-Message-State: AFuF++l3yTMfzYiUhBNjCGE/TfIxhZO70Fau1I5aL4LzxUgmQbs3+fzu sBxxHMrqGc06EZnsWEUI5TTmgLX2E0DzVH20X8JEuM0F3o9EiQrzqemlI3Pjue7gM+8= X-Gm-Gg: AYBFou2ZXAEAFfPHMW40gtDTGfCf98yu5xw7YxQv84sqaR0/pa7/BBu/eI+ucYkd0Mk 2PIPbbpPEjfTz1jhlmpU0t7OwrttR6KQ2AnDrXD9OU+1anpr2myAb/lxBfJ2x8xIg2u7WU8zsWv LZyk0GbpB6m9R4iL3cjHTNfQLavEYNEuWssSxI74knK6iWbh7ohP025q+GGzQQoMIzEHHetgDLl xV4u0nyxs3xJ0Xt0zDCuKdXS53J5he0oqT42gauMvvRYOMd1htxiLyQ4W42kyO7VlhHzogmUU88 Zrn4d31YznZRFOWouvAWYaQL5/MCbViRkJs42ejM3Aw3qJn8IuVTnEfHPmHATIvLJb993bokXBb M3ONhAT/fNx67kjUKBvfUkc92xS0Se7NWbglt9Y1FELcydtFQXAJLMFbA1rrDgToulwFIoIr2aD 7bS1dgTUMQC2JX9jvUFuXjLRK0yq1kExe2XI18mic9h62q6jWXOGKwglt9V+snVT9onTiLmWD57 E6BSkE1JnweZj0= X-Received: by 2002:a05:6214:1304:b0:90e:9875:b48c with SMTP id 6a1803df08f44-913fc8b1f2dmr34711036d6.12.1790044448184; Mon, 21 Sep 2026 19:34:08 -0700 (PDT) Received: from toxicpanda.com ([153.61.196.250]) by smtp.gmail.com with ESMTPSA id 6a1803df08f44-9140250b82asm4582396d6.39.2026.09.21.19.34.07 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 21 Sep 2026 19:34:07 -0700 (PDT) From: Josef Bacik Date: Tue, 22 Sep 2026 02:23:31 +0000 Subject: [PATCH v5 12/13] rcu-tasks-trace: Assert no reader is held on return to userspace Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260922-b4-rcu-tasks-preempt-qs-v5-12-410f57770bad@toxicpanda.com> References: <20260922-b4-rcu-tasks-preempt-qs-v5-0-410f57770bad@toxicpanda.com> In-Reply-To: <20260922-b4-rcu-tasks-preempt-qs-v5-0-410f57770bad@toxicpanda.com> To: "Paul E. McKenney" , Frederic Weisbecker , Alexei Starovoitov , Steven Rostedt Cc: Boqun Feng , Masami Hiramatsu , Mark Rutland , Peter Zijlstra , Thomas Gleixner , Daniel Borkmann , Andrii Nakryiko , Puranjay Mohan , rcu@vger.kernel.org, bpf@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, Josef Bacik X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openssh-sha256; t=1790043820; l=2581; i=josef@toxicpanda.com; h=from:subject:message-id; bh=EMsOGg9VRxBcocei3P3dGUkU8foVu6cO35v23va9IOs=; b=U1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgUBr36M/n0nWN0DNbnxwzIiCZez6MG JiruuNaSCI/zXsAAAAGcGF0YXR0AAAAAAAAAAZzaGE1MTIAAABTAAAAC3NzaC1lZDI1NTE5AAAA QLzeLb/2gtbRDPx3ssZtnGVRsANVTMmhDuPYTamLR1rD3jjMOpYIakVJmblm2P6SqGgYyithDIa ERiWg9/irxQs= X-Developer-Key: i=josef@toxicpanda.com; a=openssh; fpr=SHA256:C8kOX2QUJCMqnCX+KEeoqRAjLo9L+ELOSH2NSAJHqGA With trampolines now taking rcu_read_lock_trace() open-coded from assembly and JIT-emitted code, an unbalanced reader would silently turn every later Tasks Trace grace period on that task into a stall. No task can legitimately reach userspace with current->trc_reader_nesting non-zero, so under CONFIG_PROVE_RCU check it in the generic entry code's return-to-user validation, next to the existing kmap and lockdep assertions. Compiles away otherwise. Assisted-by: LLM Signed-off-by: Josef Bacik --- include/linux/irq-entry-common.h | 2 ++ include/linux/rcupdate_trace.h | 7 +++++++ 2 files changed, 9 insertions(+) diff --git a/include/linux/irq-entry-common.h b/include/linux/irq-entry-com= mon.h index e19b41ee6b18..d42568a488a7 100644 --- a/include/linux/irq-entry-common.h +++ b/include/linux/irq-entry-common.h @@ -5,6 +5,7 @@ #include #include #include +#include #include #include #include @@ -214,6 +215,7 @@ static __always_inline void __exit_to_user_mode_validat= e(void) { /* Ensure that kernel state is sane for a return to userspace */ kmap_assert_nomap(); + rcu_tasks_trace_assert_idle(); lockdep_assert_irqs_disabled(); lockdep_sys_exit(); } diff --git a/include/linux/rcupdate_trace.h b/include/linux/rcupdate_trace.h index dcdb11643496..9eb7f92c710b 100644 --- a/include/linux/rcupdate_trace.h +++ b/include/linux/rcupdate_trace.h @@ -217,6 +217,12 @@ unsigned long rcu_tasks_trace_batches_completed(void); // Placeholders to enable stepwise transition. void __init rcu_tasks_trace_suppress_unused(void); =20 +/* A task must never reach userspace inside an rcu_read_lock_trace() reade= r. */ +static inline void rcu_tasks_trace_assert_idle(void) +{ + WARN_ON_ONCE(IS_ENABLED(CONFIG_PROVE_RCU) && READ_ONCE(current->trc_reade= r_nesting)); +} + #else static inline unsigned long rcu_tasks_trace_batches_completed(void) { retu= rn 0; } /* @@ -226,6 +232,7 @@ static inline unsigned long rcu_tasks_trace_batches_com= pleted(void) { return 0; static inline void call_rcu_tasks_trace(struct rcu_head *rhp, rcu_callback= _t func) { BUG(); } static inline void rcu_read_lock_trace(void) { BUG(); } static inline void rcu_read_unlock_trace(void) { BUG(); } +static inline void rcu_tasks_trace_assert_idle(void) { } #endif /* #ifdef CONFIG_TASKS_TRACE_RCU */ =20 DEFINE_LOCK_GUARD_0(rcu_tasks_trace, --=20 2.55.0 From nobody Thu Sep 24 17:09:51 2026 Received: from mail-qk2-f12.google.com (mail-qk2-f12.google.com [74.125.230.204]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EE44D398902 for ; Tue, 22 Sep 2026 02:34:42 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.230.204 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790044486; cv=none; b=TXutzW8sqoCOn+Vm229FUmRT6BM3l/W2g0HK25D7+x2cfnyI4/JvP8oT9VJo5MsSFurms60xkU/5rKwy7GXTbDbputkloqEqf+heth9i4wyVbuU8h9Cqt+hcWWWI53YX+qTuMTRRxm/fzlCIoNla7kfD6T6Zc/wdU+50hy+7sMw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790044486; c=relaxed/simple; bh=VsJyzZYJ6ksShlnC7Yiz29R4LafGY+0/mhkU2yBL4O8=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=boRDA6R0gVfT4HiJEkKVjRnAR0yvcHt3YLLiTUppBoDW2nLsfAW/EpijcoueGiiewnTfDWgFR6dUT8pyPVhSmTTMT/GNATrQaYAU/HwgzX567rId8whMHLES9PUbn39ANluVIcNzEj44cSftBuKMClxnLE0iJTHyxpa/YZzDMEw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com; spf=pass smtp.mailfrom=toxicpanda.com; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b=kIQNNrcu; arc=none smtp.client-ip=74.125.230.204 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=toxicpanda.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=toxicpanda.com header.i=@toxicpanda.com header.b="kIQNNrcu" Received: by mail-qk2-f12.google.com with SMTP id d75a77b69052e-530de452c33so30842621cf.1 for ; Mon, 21 Sep 2026 19:34:42 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=toxicpanda.com; s=google; t=1790044482; x=1790649282; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=A8c5rUXaTXu9B0UqDL2lNbjmJqabE8+fQOLsAHyr+Co=; b=kIQNNrcubijMLU9kXufwCnfuMATpJB/6pd2klffOZfFFiOw2Hu3fE5NyJXnf5RpQet M0pRQLpDQBo1YMaOAf8dHLgkk5rryrp30J/zo4i59zq/nUnsrr9OuUpEhkoZXThOE8Dv TzSWpKgTtSaXlxosklQ9aiBMkp96te8d9wdI8x9JkUACON8sEyuwynTMT9NHupATxuXM va7O0CdKCHNbKTmWER+Mmwb+JyrMg5AQUP6TXoPG9ow3k7xRophnT1i2bzk2BfHUZ4sD R9B7FrQk6w5cjPkZJSE7UObwQB4ManZn/TRRwhj8m1PuJhyGAX77xLqC6n7OOz2XOC5R pwfg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790044482; x=1790649282; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=A8c5rUXaTXu9B0UqDL2lNbjmJqabE8+fQOLsAHyr+Co=; b=G2mQxXhVCUi0OzwE9WZt2n6BxW1gyqiCK5ybt2xbsBVN/pFBtBL00y8N5EA4JEFrh5 IwcFd7Ly0Z226pW1TtEvBiD812U9FQdkqHAQv/dcn9x4ASN2i6Zbi0+rO7AyuHJ4fDhs FOx8lHNAtIuxNbL39h36Cp2fTnDK0SWVLIBgCFA/JThuVKEZGCeauPRl8HEHx6GVnO9T OLZhrcpZmuRKKHSfgXVTSDALOTj4QAFUUOXrmx7kamqxQXHZNhZuhbuBkD0wHEUVanq/ v4j953VF7hxZ759aD78VwOYF+1upcKHLd/4S8nJrR7VuOyTHINSUFIkxBi6XdVCudRT0 pw4Q== X-Forwarded-Encrypted: i=1; AKwUvBxgVK2ZjOHdpvOk5aqfj9d7XZqi49mc5xFjDLB4PYWX0BS9j0ZCqqo8XWNgbAFmZwcFAhK+Xf+J4cGSrws=@vger.kernel.org X-Gm-Message-State: AFuF++mpMK+c0CukDJBaDn2hLrMb5w8C02R5nw6ipilcFsWAARAJsNWF bHQJv4nHQdWKvPkbb3wue8zsrKrr0KdpG0+bwjqnbgssqpMZtC+vRFg6m0K7deiKi20= X-Gm-Gg: AYBFou0J5Y1JNjLp6KqSwMQbP1IMHHKdNsoy9m0iotBRonLsGJP8f08MjKrok7boXFW ys+jFfU1h3/Zzp/M0KizrVTAueqnhTCqgCTWcgg+zpuuqThCvR/6/CZFVGY7GDilzYXlcAmgpbo +gWXvkaaXfP4cQ4EJJ+st1kVNFYO6+BTF36i6jR8C5SzEra2rq+WN5dYLPoqc3+7DsNCU3dhszq 8MJmuH9VPYNgvo81x1+01DNb1AsqVv/m+hUm+r4vblNxfr+QJUpPLHkqjBPCyAl/OZahzwG1+Xq B4RHEZ1EKrZRkjP2AuRtv9ky0qkZwza1tnCqyzuObiSZmn0c7fm6dmQ4jEUn6E7cUiWg8THKcds taiRgnvCpvxfVLbKTiG+w2V899eyN3LrVWwCTp0oCHXZX5R+Lp7aSkeVb9V58yut1t8ULqQMefR 9NvXkKpdY6ZTLxyeNk1UfGq6yLTtRdB/2JTSq6vnP8KMZ4nBSxOyLFVxGmh1x067vJydspCp92S iOC80RUrEox5H7D X-Received: by 2002:ac8:5981:0:b0:531:341:685b with SMTP id d75a77b69052e-532d8e6db17mr35559661cf.57.1790044481601; Mon, 21 Sep 2026 19:34:41 -0700 (PDT) Received: from toxicpanda.com ([153.61.196.242]) by smtp.gmail.com with ESMTPSA id d75a77b69052e-532e19042f2sm1747801cf.17.2026.09.21.19.34.40 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 21 Sep 2026 19:34:41 -0700 (PDT) From: Josef Bacik Date: Tue, 22 Sep 2026 02:23:32 +0000 Subject: [PATCH v5 13/13] x86, arm64: Build Tasks RCU on Tasks Trace readers in trampolines Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260922-b4-rcu-tasks-preempt-qs-v5-13-410f57770bad@toxicpanda.com> References: <20260922-b4-rcu-tasks-preempt-qs-v5-0-410f57770bad@toxicpanda.com> In-Reply-To: <20260922-b4-rcu-tasks-preempt-qs-v5-0-410f57770bad@toxicpanda.com> To: "Paul E. McKenney" , Frederic Weisbecker , Alexei Starovoitov , Steven Rostedt Cc: Boqun Feng , Masami Hiramatsu , Mark Rutland , Peter Zijlstra , Thomas Gleixner , Daniel Borkmann , Andrii Nakryiko , Puranjay Mohan , rcu@vger.kernel.org, bpf@vger.kernel.org, linux-trace-kernel@vger.kernel.org, linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, Josef Bacik X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openssh-sha256; t=1790043821; l=6284; i=josef@toxicpanda.com; h=from:subject:message-id; bh=VsJyzZYJ6ksShlnC7Yiz29R4LafGY+0/mhkU2yBL4O8=; b=U1NIU0lHAAAAAQAAADMAAAALc3NoLWVkMjU1MTkAAAAgUBr36M/n0nWN0DNbnxwzIiCZez6MG JiruuNaSCI/zXsAAAAGcGF0YXR0AAAAAAAAAAZzaGE1MTIAAABTAAAAC3NzaC1lZDI1NTE5AAAA QAul8WLe0tDTHb9SU5nk2yovv9PSDoMlv/2xo0UkHrqLg8PKI6TfwzvMu0u8RzKLGszaLneeE+P K/bP4/WyaqA8= X-Developer-Key: i=josef@toxicpanda.com; a=openssh; fpr=SHA256:C8kOX2QUJCMqnCX+KEeoqRAjLo9L+ELOSH2NSAJHqGA With the preceding patches every trampoline whose lifetime Tasks RCU guards on x86-64 and arm64 -- ftrace_caller and its copies, the optprobe template, BPF trampolines, and the sample direct-call trampolines -- is a Tasks Trace RCU reader around its call-out, and the text outside that reader is known to rcu_tasks_trampoline_text(). Select HAVE_RCU_TRAMPOLINE_READERS on both (x86-64 with SMP for Tree SRCU and DYNAMIC_FTRACE, which is where its ftrace_caller changes and arch_rcu_tasks_trampoline_text() live; arm64 with DYNAMIC_FTRACE_WITH_ARGS likewise), which switches CONFIG_TASKS_RCU to the implementation added earlier in the series: a Tasks RCU grace period becomes a per-CPU pass over context switches and irq-exit reschedules outside trampoline text plus a Tasks Trace grace period, bounded by a few jiffies and preempt-off latency instead of by the longest stretch any task runs without sleeping. Other architectures keep the classic implementation. Update Documentation/RCU and the FORCE_TASKS_RCU help text to describe the variant and the obligation it places on trampolines. Assisted-by: LLM Signed-off-by: Josef Bacik --- .../RCU/Design/Requirements/Requirements.rst | 24 ++++++++++++++++++= ++++ Documentation/RCU/checklist.rst | 7 ++++++- arch/arm64/Kconfig | 1 + arch/x86/Kconfig | 1 + kernel/rcu/Kconfig | 6 ++++-- 5 files changed, 36 insertions(+), 3 deletions(-) diff --git a/Documentation/RCU/Design/Requirements/Requirements.rst b/Docum= entation/RCU/Design/Requirements/Requirements.rst index 8101fe6229d5..c059c463e083 100644 --- a/Documentation/RCU/Design/Requirements/Requirements.rst +++ b/Documentation/RCU/Design/Requirements/Requirements.rst @@ -2756,6 +2756,30 @@ synchronize_rcu(), and rcu_barrier(), respectively. = In three APIs are therefore implemented by separate functions that check for voluntary context switches. =20 +Architectures that select ``CONFIG_HAVE_RCU_TRAMPOLINE_READERS`` keep the +same three APIs but implement the grace period differently +(``CONFIG_TASKS_RCU_TRAMPOLINE_READERS``). There, every trampoline whose +lifetime Tasks RCU guards enters a Tasks Trace RCU read-side critical +section (rcu_read_lock_trace() or its assembly equivalent) before calling +out and leaves it before returning, so a task anywhere inside such a +call-out is an ordinary Tasks Trace reader whether or not it is +preempted. The few trampoline instructions outside that reader can only +be occupied by a task that was interrupted there, so the grace period +additionally waits for each CPU to pass through a context switch, and the +irq-exit preemption path, the only switch that can catch a task inside +such text (rcu_tasks_trampoline_text()), briefly makes such a task a +holdout until it is next seen elsewhere. On such kernels an involuntary +context switch outside trampoline text *is* a Tasks-RCU quiescent state, +a Tasks RCU grace period no longer depends on how long any task runs +without sleeping, cond_resched_tasks_rcu_qs() is unnecessary, and the +obligation moves to the trampolines: anything that relies on +synchronize_rcu_tasks() to protect code a task may be preempted in must +take the Tasks Trace reader (see register_ftrace_direct()), or, where +that is impossible because the code is ordinary text with no trampoline +of its own, make it known to rcu_tasks_trampoline_text() and wait out +tasks already parked there with rcu_tasks_wait_irq_preempted(), as the +kprobe jump optimizer does. + Tasks Rude RCU ~~~~~~~~~~~~~~ =20 diff --git a/Documentation/RCU/checklist.rst b/Documentation/RCU/checklist.= rst index 4b30f701225f..87dd506bffef 100644 --- a/Documentation/RCU/checklist.rst +++ b/Documentation/RCU/checklist.rst @@ -252,7 +252,12 @@ over a rather long period of time, but improvements ar= e always welcome! a. If the updater uses synchronize_rcu_tasks() or call_rcu_tasks(), then the readers must refrain from executing voluntary context switches, that is, from - blocking. + blocking. On architectures that select + CONFIG_HAVE_RCU_TRAMPOLINE_READERS a reader must in + addition be a Tasks Trace RCU reader (that is what the + trampolines there do around their call-outs) or be text + that rcu_tasks_trampoline_text() recognises; an arbitrary + stretch of preemptible kernel code is not protected. =20 b. If the updater uses call_rcu_tasks_trace() or synchronize_rcu_tasks_trace(), then the diff --git a/arch/arm64/Kconfig b/arch/arm64/Kconfig index b5a51b0ef944..bf0e56006863 100644 --- a/arch/arm64/Kconfig +++ b/arch/arm64/Kconfig @@ -218,6 +218,7 @@ config ARM64 select HAVE_PERF_REGS select HAVE_PERF_USER_STACK_DUMP select HAVE_PREEMPT_DYNAMIC_KEY + select HAVE_RCU_TRAMPOLINE_READERS if DYNAMIC_FTRACE_WITH_ARGS select HAVE_REGS_AND_STACK_ACCESS_API select HAVE_RELIABLE_STACKTRACE select HAVE_POSIX_CPU_TIMERS_TASK_WORK diff --git a/arch/x86/Kconfig b/arch/x86/Kconfig index 15fd9ec5ecac..64c3814eb745 100644 --- a/arch/x86/Kconfig +++ b/arch/x86/Kconfig @@ -288,6 +288,7 @@ config X86 select MMU_GATHER_RCU_TABLE_FREE select MMU_GATHER_MERGE_VMAS select HAVE_POSIX_CPU_TIMERS_TASK_WORK + select HAVE_RCU_TRAMPOLINE_READERS if X86_64 && SMP && DYNAMIC_FTRACE select HAVE_REGS_AND_STACK_ACCESS_API select HAVE_RELIABLE_STACKTRACE if UNWINDER_ORC || STACK_VALIDATION select HAVE_FUNCTION_ARG_ACCESS_API diff --git a/kernel/rcu/Kconfig b/kernel/rcu/Kconfig index bbab14bc14c3..341b68b972ef 100644 --- a/kernel/rcu/Kconfig +++ b/kernel/rcu/Kconfig @@ -95,8 +95,10 @@ config FORCE_TASKS_RCU help This option force-enables a task-based RCU implementation that uses only voluntary context switch (not preemption!), - idle, and user-mode execution as quiescent states. Not for - manual selection in most cases. + idle, and user-mode execution as quiescent states, or, on + HAVE_RCU_TRAMPOLINE_READERS architectures, the variant built on + Tasks Trace RCU readers in trampolines. Not for manual + selection in most cases. =20 config NEED_TASKS_RCU bool --=20 2.55.0