From nobody Fri Oct 2 03:53:34 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 58983459AC4; Wed, 5 Aug 2026 12:24:01 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785932642; cv=none; b=IWbgIRTIaUArRKa52y4cUE9BKK1j/SIQaNKKb0vVe6CZwyosrrReTHag2qWW24l6jfMIJ9nQ1m2zb1K9fTbF8oI/g87zLj9Mo2zu6Nn+VV6fDY7TwIUK/VK+t0Sn6tAVVq/kdbKWl1W1lNS2FNiDCLEUdzD6iK9DPty+IGnzS8M= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785932642; c=relaxed/simple; bh=EReJDLhEGlZPErokzaY75PeN749eccArMOHc/45VCa0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=O79JafiiACi05DMwiWegKh3Mare15Ksq3L9oq+E+LS9ZwxFqycgBePfHQ+r2UJYIqKErHvEUhcYIdrLaRBUjqSa7du4EwrvxLnnq8Knxdtfjsx+vs5pLJWeatcENQpBEgIXW4sTLQZhT/EiV0n6gCQa3k/2o3qzK8UiAWY8BV5I= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=GaImOvrT; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="GaImOvrT" Received: by smtp.kernel.org (Postfix) with ESMTPSA id E8AA31F000E9; Wed, 5 Aug 2026 12:24:00 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785932641; bh=kuNIjXo4zSvHj6D7JE/Gr+QwVuRbAcjkxyeYD0d8ipU=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=GaImOvrTEaIAcDaaRuVDqv4/esokMKc+5d/Uo42OP+gH9Qjo4rGH0M2IeTohD/RDs yXLwWG5shVefHuvRMny2EHNCLOEcE4AsLvy0EjRDDb9xwfqmgvSzBhHyljvTuoN4vW 5+FZe8D+bwCo6MTiEiEUlrucAvcdCYt47vhJVCOKLxcb1/fMr6Zqnoh9o9YBmY86ru uXp3+//H+1FIwP+QLeO/m3wOh8JO3PMSkK9hcTQDn1nh7SqMuYwWNoSGzhEh0FOfRM SnJgqIa403+AI3ItZjNyLOBbS/6C2594lCTDXCNL8SNogUouqdA/EmwTmJYngMKEkW ta3YCJlLYTQQQ== From: Puranjay Mohan To: "Lai Jiangshan" , "Paul E. McKenney" , "Josh Triplett" , =?UTF-8?q?Onur=20=C3=96zkan?= , "Frederic Weisbecker" , "Neeraj Upadhyay" , "Joel Fernandes" , "Boqun Feng" , "Uladzislau Rezki" , "Davidlohr Bueso" , "Andrii Nakryiko" , "Eduard Zingerman" , "Alexei Starovoitov" , "Daniel Borkmann" , "Kumar Kartikeya Dwivedi" Cc: Puranjay Mohan , Steven Rostedt , Mathieu Desnoyers , Zqiang , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Matt Fleming , "Harry Yoo (Oracle)" , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, bpf@vger.kernel.org, linux-rt-devel@lists.linux.dev Subject: [PATCH v3 1/6] rcu: Make call_rcu() safe to call from any context Date: Wed, 5 Aug 2026 05:23:38 -0700 Message-ID: <20260805122346.269445-2-puranjay@kernel.org> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260805122346.269445-1-puranjay@kernel.org> References: <20260805122346.269445-1-puranjay@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" RCU's per-CPU callback list is only touched with interrupts disabled: the enqueue runs under local_irq_save() (and the nocb locks when offloaded), as do callback invocation and grace-period work. A call_rcu() that arrives with interrupts already disabled, whether from an NMI or from instrumentation that re-enters RCU, can interrupt one of those and corrupt the list or deadlock. Handle it by deferring: stage the callback on a per-CPU llist and raise an irq_work that re-issues it once interrupts are on. The re-issue goes straight to the enqueue so it cannot defer again. Callers that only hold interrupts off are deferred too, which is harmless. Skip the gate while the scheduler is down (RCU_SCHEDULER_INACTIVE), since irq_work is not usable that early and rcu_init() already calls call_rcu(). rcu_barrier() flushes deferred callbacks before it scans the lists by draining every CPU's ->defer_head itself, and rcutree_migrate_callbacks() drains an outgoing CPU's. ->defer_lock is held across llist_del_all() and the re-issue, so these drainers serialize: a drain that finds the list empty knows any re-issue in flight has already reached a callback list. A deferred callback is re-issued without the lazy hint: it has already waited for the drain, so batching it further would only add latency. The re-issue runs with interrupts disabled, so instrumentation on the enqueue path can re-enter call_rcu(), stage another callback, and re-raise the irq_work, livelocking the drain. A per-CPU flag guards it: a call_rcu() that tries to defer while the drain is running, and is not from an NMI, is dropped with WARN_ONCE() instead of re-queued. Dropping leaks the callback, but the alternative is an unbounded loop. The irq_work is IRQ_WORK_INIT_HARD so the re-issue stays prompt on PREEMPT_RT, where a non-HARD irq_work runs in a kthread that can be delayed under load. A hidden CONFIG_RCU_DEFER gates the code and its IRQ_WORK dependency; without it call_rcu() enqueues directly as before. Under CONFIG_PROVE_RCU, warn if the direct path is reached from an NMI. Suggested-by: Paul E. McKenney Signed-off-by: Puranjay Mohan --- kernel/rcu/Kconfig | 6 ++ kernel/rcu/rcu.h | 15 +++++ kernel/rcu/tree.c | 139 +++++++++++++++++++++++++++++++++++++++++---- kernel/rcu/tree.h | 6 ++ 4 files changed, 156 insertions(+), 10 deletions(-) diff --git a/kernel/rcu/Kconfig b/kernel/rcu/Kconfig index f15da8038d0ba..1a5fb3156c062 100644 --- a/kernel/rcu/Kconfig +++ b/kernel/rcu/Kconfig @@ -175,6 +175,12 @@ config RCU_STALL_COMMON config RCU_NEED_SEGCBLIST def_bool ( TREE_RCU || TREE_SRCU || TASKS_RCU_GENERIC ) =20 +# The deferral (and the IRQ_WORK it uses) is only needed where call_rcu() / +# call_srcu() can be invoked while a callback-list operation is in flight. +config RCU_DEFER + def_bool HAVE_NMI || KPROBES || FUNCTION_TRACER || TRACEPOINTS + select IRQ_WORK + config RCU_FANOUT int "Tree-based hierarchical RCU fanout value" range 2 64 if 64BIT diff --git a/kernel/rcu/rcu.h b/kernel/rcu/rcu.h index 39a9f6fa9a7b2..fd075d91b80cf 100644 --- a/kernel/rcu/rcu.h +++ b/kernel/rcu/rcu.h @@ -572,6 +572,21 @@ static inline void tasks_cblist_init_generic(void) { } #define RCU_SCHEDULER_INIT 1 #define RCU_SCHEDULER_RUNNING 2 =20 +/* + * Defer a call_rcu()/call_srcu() callback rather than enqueue it now? De= fer + * whenever interrupts are disabled, since a callback-list operation may b= e in + * flight on this CPU. Not before the scheduler is up, though: that covers + * early boot, where irq_work is not yet usable and rcu_init() itself alre= ady + * calls call_rcu(). + */ +static inline bool should_rcu_defer(void) +{ + if (!IS_ENABLED(CONFIG_RCU_DEFER)) + return false; + + return irqs_disabled() && rcu_scheduler_active !=3D RCU_SCHEDULER_INACTIV= E; +} + enum rcutorture_type { RCU_FLAVOR, RCU_TASKS_FLAVOR, diff --git a/kernel/rcu/tree.c b/kernel/rcu/tree.c index 21b6ce1dffb63..6a83408974c2b 100644 --- a/kernel/rcu/tree.c +++ b/kernel/rcu/tree.c @@ -24,6 +24,7 @@ #include #include #include +#include #include #include #include @@ -3148,21 +3149,25 @@ static void check_cb_ovld(struct rcu_data *rdp) raw_spin_unlock_rcu_node(rnp); } =20 -static void -__call_rcu_common(struct rcu_head *head, rcu_callback_t func, bool lazy_in) +/* + * The callback list is only accessed with interrupts disabled, so a call_= rcu() + * that arrives with interrupts off (see should_rcu_defer()) stages the ca= llback + * on a per-CPU llist that an irq_work re-issues once interrupts are on. + */ +static void rcu_defer_drain(struct irq_work *iw); + +/* + * Enqueue @head on this CPU's rcu_segcblist. Also called by rcu_defer_dr= ain() + * to re-issue a deferred callback, so it must not re-check the deferral + * condition. Either caller may have interrupts already disabled. + */ +static void rcu_do_enqueue(struct rcu_head *head, rcu_callback_t func, boo= l lazy_in) { static atomic_t doublefrees; unsigned long flags; bool lazy; struct rcu_data *rdp; =20 - /* Misaligned rcu_head! */ - WARN_ON_ONCE((unsigned long)head & (sizeof(void *) - 1)); - - /* Avoid NULL dereference if callback is NULL. */ - if (WARN_ON_ONCE(!func)) - return; - if (debug_rcu_head_queue(head)) { /* * Probable double call_rcu(), so leak the callback. @@ -3206,6 +3211,104 @@ __call_rcu_common(struct rcu_head *head, rcu_callba= ck_t func, bool lazy_in) local_irq_restore(flags); } =20 +/* + * Re-issue deferred callbacks straight to the enqueue so they cannot defer + * again. ->defer_lock serializes the drainers: this CPU's irq_work, + * rcu_defer_flush() and rcutree_migrate_callbacks(). + */ +static void __rcu_defer_drain(struct rcu_data *rdp, bool guard) +{ + struct llist_node *node, *next; + unsigned long flags; + + if (!IS_ENABLED(CONFIG_RCU_DEFER)) + return; + + raw_spin_lock_irqsave(&rdp->defer_lock, flags); + if (guard) + WRITE_ONCE(rdp->defer_draining, true); + llist_for_each_safe(node, next, llist_del_all(&rdp->defer_head)) { + struct rcu_head *head =3D (struct rcu_head *)node; + + head->next =3D NULL; + rcu_do_enqueue(head, head->func, false); + } + if (guard) + WRITE_ONCE(rdp->defer_draining, false); + raw_spin_unlock_irqrestore(&rdp->defer_lock, flags); +} + +/* + * Only the irq_work drain can be re-fed by its own re-issue, so only it s= ets + * ->defer_draining. A direct drain re-issues onto this CPU, and anything + * staged during it is picked up by that CPU's own irq_work. + */ +static void rcu_defer_drain(struct irq_work *iw) +{ + __rcu_defer_drain(container_of(iw, struct rcu_data, defer_work), true); +} + +/* Stage @head for this CPU's irq_work when call_rcu() cannot enqueue now.= */ +static void call_rcu_defer(struct rcu_head *head, rcu_callback_t func) +{ + struct rcu_data *rdp =3D this_cpu_ptr(&rcu_data); + + /* + * Instrumentation on the enqueue path can re-enter here from inside + * rcu_defer_drain(). Re-queuing would livelock the drain, so drop the + * callback; an NMI is one-shot and cannot loop, so let it through. + */ + if (READ_ONCE(rdp->defer_draining) && !in_nmi()) { + WARN_ONCE(IS_ENABLED(CONFIG_PROVE_RCU), + "call_rcu() re-entered during callback drain; leaking callback\n"); + return; + } + head->func =3D func; + if (llist_add((struct llist_node *)head, &rdp->defer_head)) + irq_work_queue(&rdp->defer_work); +} + +/* + * Register pending deferred callbacks into the callback lists so a follow= ing + * rcu_barrier() waits for them. This runs before rcu_barrier() scans the + * lists. + */ +static void rcu_defer_flush(void) +{ + int cpu; + + if (!IS_ENABLED(CONFIG_RCU_DEFER)) + return; + + for_each_possible_cpu(cpu) { + struct rcu_data *rdp =3D per_cpu_ptr(&rcu_data, cpu); + + if (!llist_empty(&rdp->defer_head)) + __rcu_defer_drain(rdp, false); + } +} + +static void +__call_rcu_common(struct rcu_head *head, rcu_callback_t func, bool lazy_in) +{ + /* Misaligned rcu_head! */ + WARN_ON_ONCE((unsigned long)head & (sizeof(void *) - 1)); + + /* Avoid NULL dereference if callback is NULL. */ + if (WARN_ON_ONCE(!func)) + return; + + if (should_rcu_defer()) { + call_rcu_defer(head, func); + return; + } + + /* An NMI reaching here entered with irqs enabled, so the enqueue can rac= e. */ + WARN_ON_ONCE(IS_ENABLED(CONFIG_PROVE_RCU) && in_nmi()); + + rcu_do_enqueue(head, func, lazy_in); +} + #ifdef CONFIG_RCU_LAZY static bool enable_rcu_lazy __read_mostly =3D !IS_ENABLED(CONFIG_RCU_LAZY_= DEFAULT_OFF); module_param(enable_rcu_lazy, bool, 0444); @@ -3896,8 +3999,12 @@ void rcu_barrier(void) unsigned long flags; unsigned long gseq; struct rcu_data *rdp; - unsigned long s =3D rcu_seq_snap(&rcu_state.barrier_sequence); + unsigned long s; =20 + /* Register any deferred callbacks before snapshotting the sequence. */ + rcu_defer_flush(); + + s =3D rcu_seq_snap(&rcu_state.barrier_sequence); rcu_barrier_trace(TPS("Begin"), -1, s); =20 /* Take mutex to serialize concurrent rcu_barrier() requests. */ @@ -4231,6 +4338,10 @@ rcu_boot_init_percpu_data(int cpu) rdp->rcu_onl_gp_state =3D RCU_GP_CLEANED; rdp->last_sched_clock =3D jiffies; rdp->cpu =3D cpu; + init_llist_head(&rdp->defer_head); + raw_spin_lock_init(&rdp->defer_lock); + /* Hard irq_work so the re-issue runs promptly. */ + rdp->defer_work =3D IRQ_WORK_INIT_HARD(rcu_defer_drain); rcu_boot_init_nocb_percpu_data(rdp); } =20 @@ -4528,6 +4639,14 @@ void rcutree_migrate_callbacks(int cpu) struct rcu_data *rdp =3D per_cpu_ptr(&rcu_data, cpu); bool needwake; =20 + /* + * Callbacks the outgoing CPU deferred late in the offline path (past the + * point its irq_work can run) sit on ->defer_head, which the ->cblist + * migration below does not cover. Drain them here, before the early + * returns; the re-issue lands on this CPU. + */ + __rcu_defer_drain(rdp, false); + if (rcu_rdp_is_offloaded(rdp)) return; =20 diff --git a/kernel/rcu/tree.h b/kernel/rcu/tree.h index eedfa43059e80..b7cac7a13b4f2 100644 --- a/kernel/rcu/tree.h +++ b/kernel/rcu/tree.h @@ -229,6 +229,12 @@ struct rcu_data { struct rcu_head barrier_head; int exp_watching_snap; /* Double-check need for IPI. */ =20 + /* Deferral of an NMI/reentrant call_rcu(); see __call_rcu_common(). */ + struct llist_head defer_head; + struct irq_work defer_work; + raw_spinlock_t defer_lock; + bool defer_draining; + /* 5) Callback offloading. */ #ifdef CONFIG_RCU_NOCB_CPU struct swait_queue_head nocb_cb_wq; /* For nocb kthreads to sleep on. */ --=20 2.53.0-Meta From nobody Fri Oct 2 03:53:34 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0983145D5E4; Wed, 5 Aug 2026 12:24:04 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785932646; cv=none; b=Nv6DbFOCZKXfsk9s19RmC+ysCLXUiJ0HbLSGe3z57MT6CC4FXB7m5VBWpyo9h/O01WNJ7MZtuvkjksA6Xj02VtpDNyZg8eKCF3kL1E2YqxuN2kCBY0s3QJHiXLVuutCTmKwyfLGEXFu0oKcpAy2bZvbV83mV+gLa8X46Ep4/PV4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785932646; c=relaxed/simple; bh=KfqAmejNRotNqn4BSRDfTw6HG1fKLGZx2qlRvQWCHbM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=pSsAKFvvhsaEqV6q/Me6Shpc2E/5+mWW7Lqc49LF9Hc7Y1IIDvLMOPFuuMnARTTRsV5Pzf5ew6amrYsli+FCYGzlOqmD+O53M6GH9B/i3u2Kn/I/kj2Jvg9H+AuJ5bmDZaPjxeQemsiYzdA9Rd+zx1kt8TBJljLBr9vAedEC5cY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=c9Rle6RE; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="c9Rle6RE" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 7DE641F00A3A; Wed, 5 Aug 2026 12:24:04 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785932644; bh=F4fKMwHkoXgcYwZ+GDVXYH9cJRDr3Ob52RPoP9ggB8A=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=c9Rle6RERUaulViFjb6THs/Hdl2wKywsQYxNtbSFUewhkIHqZCo+GoeF7/cIPLVnR kAJaqNKa+9m5DRG1oPbPkmaqUOS3EhAwf+FMhNfqAKkTI+8OH7vMi4L0MtgaxQluNK 5By1QBBtMP7J0ev7y2K1Q23JvXLe7CJ8ZLXiBf72Yu9PqAa8z2hT1uIMYGO7eoYnwp LfMx5Ovf0beY5vjNxqJa992hJGLobfm3Be7VBnPv7HtxPiCUiuKRsqLfHbbjEipIhV b7ZqtjCtWzRMlKjDX1cgTlGQyed6/OBr20E3FNoNaWWEJ+U4LCQhNycpKbRtetsHg8 ZEJF+CJqMIUwQ== From: Puranjay Mohan To: "Lai Jiangshan" , "Paul E. McKenney" , "Josh Triplett" , =?UTF-8?q?Onur=20=C3=96zkan?= , "Frederic Weisbecker" , "Neeraj Upadhyay" , "Joel Fernandes" , "Boqun Feng" , "Uladzislau Rezki" , "Davidlohr Bueso" , "Andrii Nakryiko" , "Eduard Zingerman" , "Alexei Starovoitov" , "Daniel Borkmann" , "Kumar Kartikeya Dwivedi" Cc: Puranjay Mohan , Steven Rostedt , Mathieu Desnoyers , Zqiang , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Matt Fleming , "Harry Yoo (Oracle)" , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, bpf@vger.kernel.org, linux-rt-devel@lists.linux.dev Subject: [PATCH v3 2/6] rcu: Make Tiny call_rcu() safe to call from any context Date: Wed, 5 Aug 2026 05:23:39 -0700 Message-ID: <20260805122346.269445-3-puranjay@kernel.org> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260805122346.269445-1-puranjay@kernel.org> References: <20260805122346.269445-1-puranjay@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Give Tiny call_rcu() the same treatment as Tree RCU. When interrupts are disabled and the scheduler is up, stage the callback on a lockless list that an irq_work re-issues later. One global list and irq_work suffice since Tiny RCU is uniprocessor, and there is no CPU-offline drain. The re-issue runs with interrupts disabled and can be re-entered by instrumentation, so a draining flag drops a deferring call_rcu() seen mid-drain (unless from an NMI), as in Tree RCU. Gated by CONFIG_RCU_DEFER. Suggested-by: Paul E. McKenney Signed-off-by: Puranjay Mohan --- kernel/rcu/tiny.c | 117 ++++++++++++++++++++++++++++++++++++++-------- 1 file changed, 98 insertions(+), 19 deletions(-) diff --git a/kernel/rcu/tiny.c b/kernel/rcu/tiny.c index dccccd6be9411..baffe660043f5 100644 --- a/kernel/rcu/tiny.c +++ b/kernel/rcu/tiny.c @@ -11,6 +11,8 @@ */ #include #include +#include +#include #include #include #include @@ -42,8 +44,99 @@ static struct rcu_ctrlblk rcu_ctrlblk =3D { .gp_seq =3D 0 - 300UL, }; =20 +/* + * The callback list is only accessed with interrupts disabled, so a call_= rcu() + * that arrives with interrupts off stages the callback on a lockless list= that + * an irq_work re-issues later. One global list and irq_work suffice, as = Tiny + * RCU is uniprocessor. + */ +static void rcu_defer_drain(struct irq_work *iw); +static LLIST_HEAD(rcu_defer_list); +static struct irq_work rcu_defer_iw =3D IRQ_WORK_INIT_HARD(rcu_defer_drain= ); +static bool rcu_defer_draining; + +/* + * Enqueue @head on the callback list. Also called by rcu_defer_drain() to + * re-issue a deferred callback, so it must not re-check the deferral cond= ition. + */ +static void rcu_do_enqueue(struct rcu_head *head, rcu_callback_t func) +{ + static atomic_t doublefrees; + unsigned long flags; + + if (debug_rcu_head_queue(head)) { + if (atomic_inc_return(&doublefrees) < 4) { + pr_err("%s(): Double-freed CB %p->%pS()!!! ", __func__, head, head->fu= nc); + mem_dump_obj(head); + } + return; + } + + head->func =3D func; + head->next =3D NULL; + + local_irq_save(flags); + *rcu_ctrlblk.curtail =3D head; + rcu_ctrlblk.curtail =3D &head->next; + local_irq_restore(flags); + + if (unlikely(is_idle_task(current))) { + /* force scheduling for rcu_qs() */ + resched_cpu(0); + } +} + +static void __rcu_defer_drain(bool guard) +{ + struct llist_node *node, *next; + unsigned long flags; + + /* Callbacks are unordered, so drain in llist order without reversing. */ + local_irq_save(flags); + if (guard) + WRITE_ONCE(rcu_defer_draining, true); + llist_for_each_safe(node, next, llist_del_all(&rcu_defer_list)) { + struct rcu_head *head =3D (struct rcu_head *)node; + + head->next =3D NULL; + rcu_do_enqueue(head, head->func); + } + if (guard) + WRITE_ONCE(rcu_defer_draining, false); + local_irq_restore(flags); +} + +/* Only the irq_work drain can be re-fed by its own re-issue; see Tree RCU= . */ +static void rcu_defer_drain(struct irq_work *iw) +{ + __rcu_defer_drain(true); +} + +static void call_rcu_defer(struct rcu_head *head, rcu_callback_t func) +{ + /* A re-entrant call_rcu() during the drain would livelock it; drop it. */ + if (READ_ONCE(rcu_defer_draining) && !in_nmi()) { + WARN_ONCE(IS_ENABLED(CONFIG_PROVE_RCU), + "call_rcu() re-entered during callback drain; leaking callback\n"); + return; + } + head->func =3D func; + if (llist_add((struct llist_node *)head, &rcu_defer_list)) + irq_work_queue(&rcu_defer_iw); +} + +/* Register any deferred callbacks so a following rcu_barrier() waits for = them. */ +static void rcu_defer_flush(void) +{ + if (!IS_ENABLED(CONFIG_RCU_DEFER)) + return; + __rcu_defer_drain(false); +} + void rcu_barrier(void) { + /* Register any deferred callbacks first. */ + rcu_defer_flush(); wait_rcu_gp(call_rcu_hurry); } EXPORT_SYMBOL(rcu_barrier); @@ -157,29 +250,15 @@ EXPORT_SYMBOL_GPL(synchronize_rcu); */ void call_rcu(struct rcu_head *head, rcu_callback_t func) { - static atomic_t doublefrees; - unsigned long flags; - - if (debug_rcu_head_queue(head)) { - if (atomic_inc_return(&doublefrees) < 4) { - pr_err("%s(): Double-freed CB %p->%pS()!!! ", __func__, head, head->fu= nc); - mem_dump_obj(head); - } + if (should_rcu_defer()) { + call_rcu_defer(head, func); return; } =20 - head->func =3D func; - head->next =3D NULL; - - local_irq_save(flags); - *rcu_ctrlblk.curtail =3D head; - rcu_ctrlblk.curtail =3D &head->next; - local_irq_restore(flags); + /* An NMI reaching here entered with irqs enabled, so the enqueue can rac= e. */ + WARN_ON_ONCE(IS_ENABLED(CONFIG_PROVE_RCU) && in_nmi()); =20 - if (unlikely(is_idle_task(current))) { - /* force scheduling for rcu_qs() */ - resched_cpu(0); - } + rcu_do_enqueue(head, func); } EXPORT_SYMBOL_GPL(call_rcu); =20 --=20 2.53.0-Meta From nobody Fri Oct 2 03:53:34 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A2218469828; Wed, 5 Aug 2026 12:24:08 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785932650; cv=none; b=ROYqllprjxwCdGd0zN3P4pTIfAIA2guZRDJ+jrdari0RbB4hyPKGpVgTME+5Qh4f03sa9q5HRCHVCoQ60McvfRiUrfgIZdzX51kTEWr9/MFx49FM0uai9PcmwcMlAPN8hDJqA1N2xDJFyPdM3bzljZrkAIySV4xMzlhb9S9oFLk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785932650; c=relaxed/simple; bh=2HzwJqqykpeRV+fRENdx3Rl435thQUPWJjcF5t8VHQ4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=OZENupGhmUvXRnqt3cl2OdkmMOL5KAjHTCLlv6ShqCS9FyxevUosnMYDwmFiqppyJTcfLZIXQkzRulpZVatigmMEUcj54En0p/VmOS+8avxW7aB3HINSNQzWijxrZWPXCSeUINeBHntk7mMiV0suM2O+DKAhkgWBTF1Y5iygFo0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=g5VdbbW2; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="g5VdbbW2" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 3B3371F00A3A; Wed, 5 Aug 2026 12:24:08 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785932648; bh=Zuf9N90kRGuzCtffHw96HWj2lxYV2drR4iLhb0zurcU=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=g5VdbbW2QAlQPVDqvZngUXIhSinQCYtAw3H8tKMGcfoLiM5gDtq27681M7i2fbFVq QnSRv7vrJ2xzpvvpGLMKuqs+WcHNKraJhJv+LQdrYEmDJVLMGC5/rTTAiWB6IzjHvA QdMgjptztOJTJ9b0JaLtYWL55nZwkaJS6UOsnMkDXCFsud/94rndMZW9+gyrh45i/S cp+i0Or7EA2KZ7V/MP7yxuGxQRb3osvwAAsSiNIANBKOi6LRI46sT+NfFg7/gRQUis aouXQUsmgR8NiC1HE8+mdtnUwK5jMp7HHTrHuKylwxIe+M1VRJougTBb6n7mrfbcGi cm52NRdVQY0nw== From: Puranjay Mohan To: "Lai Jiangshan" , "Paul E. McKenney" , "Josh Triplett" , =?UTF-8?q?Onur=20=C3=96zkan?= , "Frederic Weisbecker" , "Neeraj Upadhyay" , "Joel Fernandes" , "Boqun Feng" , "Uladzislau Rezki" , "Davidlohr Bueso" , "Andrii Nakryiko" , "Eduard Zingerman" , "Alexei Starovoitov" , "Daniel Borkmann" , "Kumar Kartikeya Dwivedi" Cc: Puranjay Mohan , Steven Rostedt , Mathieu Desnoyers , Zqiang , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Matt Fleming , "Harry Yoo (Oracle)" , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, bpf@vger.kernel.org, linux-rt-devel@lists.linux.dev Subject: [PATCH v3 3/6] srcu: Make call_srcu() safe to call from any context Date: Wed, 5 Aug 2026 05:23:40 -0700 Message-ID: <20260805122346.269445-4-puranjay@kernel.org> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260805122346.269445-1-puranjay@kernel.org> References: <20260805122346.269445-1-puranjay@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" call_srcu() has the same constraint as call_rcu(): its callback list and locks are only touched with interrupts disabled. srcu_gp_start_if_needed() enqueues under raw_spin_lock_irqsave() and may walk the srcu_node tree, as do callback invocation and grace-period work. A call_srcu() with interrupts disabled can race a list operation in flight on this CPU and corrupt the list or deadlock. call_rcu_tasks_trace() is call_srcu() under the hood, so a sleepable BPF program freeing an object can reach this. Defer as call_rcu() does: stage the callback on the srcu_data's ->defer_cbs, chain that srcu_data onto a per-CPU list, and raise a per-CPU irq_work that re-issues it straight to the enqueue helper, never back through __call_srcu(). The common path is unchanged and keeps interrupts enabled across srcu_gp_start_if_needed(). The irq_work is per-CPU rather than per-srcu_struct and statically initialized, so deferral never runs check_init_srcu_struct(); it is IRQ_WORK_INIT_HARD as for call_rcu(). srcu_barrier() and cleanup_srcu_struct() flush it first, and rcutree_migrate_callbacks() calls srcu_offline_drain() for an outgoing CPU. ->lock is held across the drain so the drainers serialize. As in call_rcu(), the re-issue runs with interrupts disabled and can be re-entered by instrumentation, so a flag on the srcu_data being drained drops a deferring call_srcu() seen mid-drain (unless from an NMI). Staging records only the callback, so an expedited request is remembered per srcu_data in ->defer_exp and the whole batch is re-issued expedited rather than silently downgraded to a normal grace period. A dropped callback can also strand state its caller associated with it, not just the callback itself. Gated by CONFIG_RCU_DEFER. Under CONFIG_PROVE_RCU, warn if the direct path is reached from an NMI. Suggested-by: Paul E. McKenney Signed-off-by: Puranjay Mohan --- include/linux/srcutree.h | 5 ++ kernel/rcu/rcu.h | 3 + kernel/rcu/srcutree.c | 171 ++++++++++++++++++++++++++++++++++++++- kernel/rcu/tree.c | 2 + 4 files changed, 177 insertions(+), 4 deletions(-) diff --git a/include/linux/srcutree.h b/include/linux/srcutree.h index 75e54e4f963fa..09a9c8f4a6d24 100644 --- a/include/linux/srcutree.h +++ b/include/linux/srcutree.h @@ -13,6 +13,8 @@ =20 #include #include +#include +#include =20 struct srcu_node; struct srcu_struct; @@ -41,6 +43,9 @@ struct srcu_data { bool srcu_cblist_invoking; /* Invoking these CBs? */ struct timer_list delay_work; /* Delay for CB invoking */ struct work_struct work; /* Context for CB invoking. */ + struct llist_head defer_cbs; /* Callbacks deferred on re-entry. */ + struct llist_node defer_link; /* Links onto the per-CPU deferral drain l= ist */ + bool defer_exp; /* A deferred callback asked to expedite. */ struct rcu_head srcu_barrier_head; /* For srcu_barrier() use. */ struct rcu_head srcu_ec_head; /* For srcu_expedite_current() use. */ int srcu_ec_state; /* State for srcu_expedite_current(). */ diff --git a/kernel/rcu/rcu.h b/kernel/rcu/rcu.h index fd075d91b80cf..84d74cd5a351c 100644 --- a/kernel/rcu/rcu.h +++ b/kernel/rcu/rcu.h @@ -587,6 +587,9 @@ static inline bool should_rcu_defer(void) return irqs_disabled() && rcu_scheduler_active !=3D RCU_SCHEDULER_INACTIV= E; } =20 +/* Drain an outgoing CPU's deferred SRCU callbacks; see rcutree_migrate_ca= llbacks(). */ +void srcu_offline_drain(int cpu); + enum rcutorture_type { RCU_FLAVOR, RCU_TASKS_FLAVOR, diff --git a/kernel/rcu/srcutree.c b/kernel/rcu/srcutree.c index 304112674e8a2..35fface51d50b 100644 --- a/kernel/rcu/srcutree.c +++ b/kernel/rcu/srcutree.c @@ -20,6 +20,7 @@ #include #include #include +#include #include #include #include @@ -79,6 +80,45 @@ static void process_srcu(struct work_struct *work); static void srcu_irq_work(struct irq_work *work); static void srcu_delay_timer(struct timer_list *t); =20 +struct srcu_defer; +static void srcu_defer_drain(struct irq_work *iw); +static void __srcu_defer_drain(struct srcu_defer *sndp, bool guard); + +/* + * Per-CPU call_srcu() deferral state, shared by every srcu_struct. A def= erred + * callback is staged on its srcu_data's ->defer_cbs; that srcu_data is ch= ained + * via ->defer_link onto ->list, which the irq_work walks. + */ +struct srcu_defer { + struct llist_head list; + struct irq_work iw; + raw_spinlock_t lock; + bool draining; +}; + +static DEFINE_PER_CPU(struct srcu_defer, srcu_defer) =3D { + .lock =3D __RAW_SPIN_LOCK_UNLOCKED(srcu_defer.lock), + .iw =3D IRQ_WORK_INIT_HARD(srcu_defer_drain), +}; + +/* + * Flush pending deferred callbacks so a following srcu_barrier() waits fo= r them. + */ +static void srcu_defer_flush(void) +{ + int cpu; + + if (!IS_ENABLED(CONFIG_RCU_DEFER)) + return; + + for_each_possible_cpu(cpu) { + struct srcu_defer *sndp =3D &per_cpu(srcu_defer, cpu); + + if (!llist_empty(&sndp->list)) + __srcu_defer_drain(sndp, false); + } +} + /* * Initialize SRCU per-CPU data. Note that statically allocated * srcu_struct structures might already have srcu_read_lock() and @@ -107,6 +147,11 @@ static void init_srcu_struct_data(struct srcu_struct *= ssp) sdp->cpu =3D cpu; INIT_WORK(&sdp->work, srcu_invoke_callbacks); timer_setup(&sdp->delay_work, srcu_delay_timer, 0); + /* + * ->defer_cbs, ->defer_link and ->defer_exp are valid when zeroed + * and are not reinitialized here, lest we clobber callbacks a + * reentrant call_srcu() already staged. See __call_srcu(). + */ sdp->ssp =3D ssp; } } @@ -695,7 +740,12 @@ void cleanup_srcu_struct(struct srcu_struct *ssp) return; /* Just leak it! */ if (WARN_ON(srcu_readers_active(ssp))) return; /* Just leak it! */ - /* Wait for irq_work to finish first as it may queue a new work. */ + /* + * Drain deferred callbacks before syncing ->irq_work: re-issuing one can + * start a grace period and re-queue ->irq_work, which then schedules + * ->work, so both must be waited out after the drain. + */ + srcu_defer_flush(); irq_work_sync(&sup->irq_work); flush_delayed_work(&sup->work); for_each_possible_cpu(cpu) { @@ -1410,8 +1460,8 @@ static unsigned long srcu_gp_start_if_needed(struct s= rcu_struct *ssp, * srcu_read_lock(), and srcu_read_unlock() that are all passed the same * srcu_struct structure. */ -static void __call_srcu(struct srcu_struct *ssp, struct rcu_head *rhp, - rcu_callback_t func, bool do_norm) +static void srcu_do_enqueue(struct srcu_struct *ssp, struct rcu_head *rhp, + rcu_callback_t func, bool do_norm) { if (debug_rcu_head_queue(rhp)) { /* Probable double call_srcu(), so leak the callback. */ @@ -1423,6 +1473,111 @@ static void __call_srcu(struct srcu_struct *ssp, st= ruct rcu_head *rhp, (void)srcu_gp_start_if_needed(ssp, rhp, do_norm); } =20 +/* + * The srcu_cblist and srcu_node tree are only accessed with interrupts di= sabled + * (srcu_gp_start_if_needed() enqueues under raw_spin_lock_irqsave() and m= ay walk + * the tree). Like call_rcu(), __call_srcu() defers when interrupts are a= lready + * disabled, so a re-entrant call_srcu() -- e.g. call_rcu_tasks_trace() fr= om a + * BPF program -- cannot corrupt the list or deadlock. + */ +static void __call_srcu(struct srcu_struct *ssp, struct rcu_head *rhp, + rcu_callback_t func, bool do_norm) +{ + if (should_rcu_defer()) { + struct srcu_defer *sndp =3D this_cpu_ptr(&srcu_defer); + struct srcu_data *sdp; + + /* + * Instrumentation on the enqueue path can re-enter here from + * inside srcu_defer_drain(). Re-queuing would livelock the + * drain, so drop the callback; an NMI cannot loop, so let it in. + */ + if (READ_ONCE(sndp->draining) && !in_nmi()) { + WARN_ONCE(IS_ENABLED(CONFIG_PROVE_RCU), + "call_srcu() re-entered during callback drain; leaking callback\n"); + return; + } + sdp =3D this_cpu_ptr(ssp->sda); + rhp->func =3D func; + if (!do_norm) + WRITE_ONCE(sdp->defer_exp, true); + if (llist_add((struct llist_node *)rhp, &sdp->defer_cbs)) { + /* + * Chain this srcu_data for the drain. ->ssp must be + * published here: deferral skips check_init_srcu_struct(), + * so on a never-initialized static srcu_struct the + * statically zeroed ->sda still has a NULL ->ssp. + */ + sdp->ssp =3D ssp; + if (llist_add(&sdp->defer_link, &sndp->list)) + irq_work_queue(&sndp->iw); + } + return; + } + + /* An NMI reaching here entered with irqs enabled, so the enqueue can rac= e. */ + WARN_ON_ONCE(IS_ENABLED(CONFIG_PROVE_RCU) && in_nmi()); + + srcu_do_enqueue(ssp, rhp, func, do_norm); +} + +/* + * Re-issue deferred callbacks straight to srcu_do_enqueue() so they canno= t defer + * again. ->lock serializes the drainers: the irq_work, srcu_defer_flush(= ) and + * srcu_offline_drain(). + */ +static void __srcu_defer_drain(struct srcu_defer *sndp, bool guard) +{ + struct llist_node *snode, *snext; + unsigned long flags; + + raw_spin_lock_irqsave(&sndp->lock, flags); + if (guard) + WRITE_ONCE(sndp->draining, true); + llist_for_each_safe(snode, snext, llist_del_all(&sndp->list)) { + struct srcu_data *sdp =3D container_of(snode, struct srcu_data, defer_li= nk); + struct srcu_struct *ssp =3D sdp->ssp; + struct llist_node *cnode, *cnext; + bool do_norm; + + cnode =3D llist_del_all(&sdp->defer_cbs); + do_norm =3D !READ_ONCE(sdp->defer_exp); + if (!do_norm) + WRITE_ONCE(sdp->defer_exp, false); + llist_for_each_safe(cnode, cnext, cnode) { + struct rcu_head *rhp =3D (struct rcu_head *)cnode; + + rhp->next =3D NULL; + srcu_do_enqueue(ssp, rhp, rhp->func, do_norm); + } + } + if (guard) + WRITE_ONCE(sndp->draining, false); + raw_spin_unlock_irqrestore(&sndp->lock, flags); +} + +/* + * Only the irq_work drain can be re-fed by its own re-issue, so only it s= ets + * ->draining. A direct drain re-issues onto this CPU, and anything staged + * during it is picked up by that CPU's own irq_work. + */ +static void srcu_defer_drain(struct irq_work *iw) +{ + __srcu_defer_drain(container_of(iw, struct srcu_defer, iw), true); +} + +/* + * Drain @cpu's deferred call_srcu() callbacks from rcutree_migrate_callba= cks() + * once @cpu is dead. One pass covers every srcu_struct, and the re-issue= lands + * on the current CPU. + */ +void srcu_offline_drain(int cpu) +{ + if (!IS_ENABLED(CONFIG_RCU_DEFER)) + return; + __srcu_defer_drain(&per_cpu(srcu_defer, cpu), false); +} + /** * call_srcu() - Queue a callback for invocation after an SRCU grace period * @ssp: srcu_struct in queue the callback @@ -1677,9 +1832,17 @@ void srcu_barrier(struct srcu_struct *ssp) { int cpu; int idx; - unsigned long s =3D rcu_seq_snap(&ssp->srcu_sup->srcu_barrier_seq); + unsigned long s; =20 check_init_srcu_struct(ssp); + + /* + * Register any deferred callbacks before snapshotting the sequence. The + * shared irq_work may also drain other srcu_structs', which is harmless. + */ + srcu_defer_flush(); + + s =3D rcu_seq_snap(&ssp->srcu_sup->srcu_barrier_seq); mutex_lock(&ssp->srcu_sup->srcu_barrier_mutex); if (rcu_seq_done(&ssp->srcu_sup->srcu_barrier_seq, s)) { smp_mb(); /* Force ordering following return. */ diff --git a/kernel/rcu/tree.c b/kernel/rcu/tree.c index 6a83408974c2b..227790a87b02f 100644 --- a/kernel/rcu/tree.c +++ b/kernel/rcu/tree.c @@ -4646,6 +4646,8 @@ void rcutree_migrate_callbacks(int cpu) * returns; the re-issue lands on this CPU. */ __rcu_defer_drain(rdp, false); + /* Likewise for the outgoing CPU's deferred call_srcu() callbacks. */ + srcu_offline_drain(cpu); =20 if (rcu_rdp_is_offloaded(rdp)) return; --=20 2.53.0-Meta From nobody Fri Oct 2 03:53:34 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B0D7D46985D; Wed, 5 Aug 2026 12:24:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785932654; cv=none; b=sbZ6h1IqZxHXnXWVVxNJNNsDVO4rzBwLVYBAAVVeC70q8rmfhZdrSRIKu8OBipZ+qyrCRjRbmiA6xsSpMurUhwbyr+Pg0d1eIYozEIqzHtIZJO6KrluMr9Aa7/CWevXzHSH4cI0/OFae2HO7pUfGU3fAyPZeKnpm18/PcA/7fGk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785932654; c=relaxed/simple; bh=2ekOa4qmCh1B2oLM5xZ5ukA5EFOV/uXlpohEZL9VNRY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=o/jvla/7dfzgpDODpE3w2Ju/uOmZEsMvkizzW5fYS4QLswUWZ99/6nhrmu5bRg3AlA9I6jyc0Dolb8Ixrgt9kEANfk/qfDw4tub4ZDWr6xv3Gncgid45CS0VdZcNNQlUGx4b6T0NWKYwiJaBfTTFnTYtyAxw2IAxPPwiLLoC/Jk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=Y0p3Rhpl; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="Y0p3Rhpl" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 48CB81F000E9; Wed, 5 Aug 2026 12:24:12 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785932652; bh=l+7jmh53O2Wea598EFhmmOEa2DQP55RrLDb9VbBnMb0=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=Y0p3Rhpl/oq+N7qp8FOcnZJuZJ1aOc8kA/ThYLoDOMxRIyMuZGKN+pnJGfw2XpvWd 3Yj9ifB9oGW4L4MkhvKetzILcuCF95AoM5HGgGXWJ94GyM+CpgosfxSItUUzO4on/C bVOYE9d26pm1cKo+hUVghEQFzbuAwg85P/I/6H+ntIh7fATG9kwT/dvLlO42oKOseO BbWDlWX4LunBbcN9AqCI1TdXPeOMW+lqr7iBL8RMl7jJrkLtalbi4wPAUksD+tJ23y vDfQ4MdDeuFMKmS46CxEEWTtOhInG2M2ZkpEENq+nvlNafShRuVZaRuSiSxht7Zmeg K74xvo9Q9YokQ== From: Puranjay Mohan To: "Lai Jiangshan" , "Paul E. McKenney" , "Josh Triplett" , =?UTF-8?q?Onur=20=C3=96zkan?= , "Frederic Weisbecker" , "Neeraj Upadhyay" , "Joel Fernandes" , "Boqun Feng" , "Uladzislau Rezki" , "Davidlohr Bueso" , "Andrii Nakryiko" , "Eduard Zingerman" , "Alexei Starovoitov" , "Daniel Borkmann" , "Kumar Kartikeya Dwivedi" Cc: Puranjay Mohan , Steven Rostedt , Mathieu Desnoyers , Zqiang , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Matt Fleming , "Harry Yoo (Oracle)" , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, bpf@vger.kernel.org, linux-rt-devel@lists.linux.dev Subject: [PATCH v3 4/6] srcu: Make Tiny call_srcu() safe to call from any context Date: Wed, 5 Aug 2026 05:23:41 -0700 Message-ID: <20260805122346.269445-5-puranjay@kernel.org> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260805122346.269445-1-puranjay@kernel.org> References: <20260805122346.269445-1-puranjay@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Give Tiny call_srcu() the same treatment as Tree SRCU. When interrupts are disabled and the scheduler is up, stage the callback on the srcu_struct's lockless list for an irq_work to re-issue later. Tiny SRCU is uniprocessor, so there is no CPU-offline drain. A draining flag drops a deferring call_srcu() that re-enters mid-drain (unless from an NMI), as in Tree SRCU. srcu_barrier() (now out of line) and cleanup_srcu_struct() drain the deferred list first, so a deferred callback is re-issued onto the callback list and invoked by the grace-period work that cleanup_srcu_struct() flushes, rather than stranded on a soon-to-be-freed srcu_struct. cleanup also syncs ->defer_iw, since that irq_work is embedded in the srcu_struct the caller is about to free. Gated by CONFIG_RCU_DEFER, like Tree SRCU. Suggested-by: Paul E. McKenney Signed-off-by: Puranjay Mohan --- include/linux/srcutiny.h | 12 ++++-- kernel/rcu/srcutiny.c | 79 ++++++++++++++++++++++++++++++++++++++-- 2 files changed, 83 insertions(+), 8 deletions(-) diff --git a/include/linux/srcutiny.h b/include/linux/srcutiny.h index fbcf13bc12d15..a0e54d9182baa 100644 --- a/include/linux/srcutiny.h +++ b/include/linux/srcutiny.h @@ -12,6 +12,7 @@ #define _LINUX_SRCU_TINY_H =20 #include +#include #include =20 struct srcu_struct { @@ -26,6 +27,8 @@ struct srcu_struct { struct rcu_head **srcu_cb_tail; /* Pending callbacks: Tail. */ struct work_struct srcu_work; /* For driving grace periods. */ struct irq_work srcu_irq_work; /* Defer schedule_work() to irq work. */ + struct llist_head defer_cbs; /* Callbacks deferred on re-entry. */ + struct irq_work defer_iw; /* Registers defer_cbs later. */ #ifdef CONFIG_DEBUG_LOCK_ALLOC struct lockdep_map dep_map; #endif /* #ifdef CONFIG_DEBUG_LOCK_ALLOC */ @@ -33,6 +36,7 @@ struct srcu_struct { =20 void srcu_drive_gp(struct work_struct *wp); void srcu_tiny_irq_work(struct irq_work *irq_work); +void srcu_defer_drain(struct irq_work *irq_work); =20 #define __SRCU_STRUCT_INIT(name, __ignored, ___ignored, ____ignored) \ { \ @@ -40,6 +44,9 @@ void srcu_tiny_irq_work(struct irq_work *irq_work); .srcu_cb_tail =3D &name.srcu_cb_head, \ .srcu_work =3D __WORK_INITIALIZER(name.srcu_work, srcu_drive_gp), \ .srcu_irq_work =3D { .func =3D srcu_tiny_irq_work }, \ + .defer_cbs =3D LLIST_HEAD_INIT(name.defer_cbs), \ + .defer_iw =3D { .node =3D { .u_flags =3D IRQ_WORK_HARD_IRQ }, \ + .func =3D srcu_defer_drain }, \ __SRCU_DEP_MAP_INIT(name) \ } =20 @@ -131,10 +138,7 @@ static inline void synchronize_srcu_expedited(struct s= rcu_struct *ssp) synchronize_srcu(ssp); } =20 -static inline void srcu_barrier(struct srcu_struct *ssp) -{ - synchronize_srcu(ssp); -} +void srcu_barrier(struct srcu_struct *ssp); =20 static inline void srcu_expedite_current(struct srcu_struct *ssp) { } #define srcu_check_read_flavor(ssp, read_flavor) do { } while (0) diff --git a/kernel/rcu/srcutiny.c b/kernel/rcu/srcutiny.c index f9c498ae75df2..ba36067909fd0 100644 --- a/kernel/rcu/srcutiny.c +++ b/kernel/rcu/srcutiny.c @@ -10,6 +10,7 @@ =20 #include #include +#include #include #include #include @@ -29,6 +30,8 @@ extern int rcu_scheduler_active; static LIST_HEAD(srcu_boot_list); static bool srcu_init_done; =20 +static void __srcu_defer_drain(struct srcu_struct *ssp, bool guard); + static int init_srcu_struct_fields(struct srcu_struct *ssp) { ssp->srcu_lock_nesting[0] =3D 0; @@ -43,6 +46,8 @@ static int init_srcu_struct_fields(struct srcu_struct *ss= p) INIT_WORK(&ssp->srcu_work, srcu_drive_gp); INIT_LIST_HEAD(&ssp->srcu_work.entry); init_irq_work(&ssp->srcu_irq_work, srcu_tiny_irq_work); + init_llist_head(&ssp->defer_cbs); + ssp->defer_iw =3D IRQ_WORK_INIT_HARD(srcu_defer_drain); return 0; } =20 @@ -86,6 +91,11 @@ EXPORT_SYMBOL_GPL(init_srcu_struct_generic); void cleanup_srcu_struct(struct srcu_struct *ssp) { WARN_ON(srcu_readers_active(ssp)); + /* Re-issue any deferred callbacks, then wait out ->defer_iw before it is= freed. */ + if (IS_ENABLED(CONFIG_RCU_DEFER)) { + __srcu_defer_drain(ssp, false); + irq_work_sync(&ssp->defer_iw); + } irq_work_sync(&ssp->srcu_irq_work); flush_work(&ssp->srcu_work); WARN_ON(ssp->srcu_gp_running); @@ -215,11 +225,11 @@ static void srcu_gp_start_if_needed(struct srcu_struc= t *ssp) } =20 /* - * Enqueue an SRCU callback on the specified srcu_struct structure, - * initiating grace-period processing if it is not already running. + * Enqueue @rhp on the callback list. Also called by srcu_defer_drain() to + * re-issue a deferred callback, so it must not re-check the deferral cond= ition. */ -void call_srcu(struct srcu_struct *ssp, struct rcu_head *rhp, - rcu_callback_t func) +static void srcu_do_enqueue(struct srcu_struct *ssp, struct rcu_head *rhp, + rcu_callback_t func) { unsigned long flags; =20 @@ -233,6 +243,58 @@ void call_srcu(struct srcu_struct *ssp, struct rcu_hea= d *rhp, srcu_gp_start_if_needed(ssp); preempt_enable(); } + +/* Set while srcu_defer_drain() re-issues, to catch a re-entrant call_srcu= (). */ +static bool srcu_defer_draining; + +static void __srcu_defer_drain(struct srcu_struct *ssp, bool guard) +{ + struct llist_node *node, *next; + unsigned long flags; + + /* Callbacks are unordered, so drain in llist order without reversing. */ + local_irq_save(flags); + if (guard) + WRITE_ONCE(srcu_defer_draining, true); + llist_for_each_safe(node, next, llist_del_all(&ssp->defer_cbs)) { + struct rcu_head *rhp =3D (struct rcu_head *)node; + + rhp->next =3D NULL; + srcu_do_enqueue(ssp, rhp, rhp->func); + } + if (guard) + WRITE_ONCE(srcu_defer_draining, false); + local_irq_restore(flags); +} + +/* Only the irq_work drain can be re-fed by its own re-issue; see Tree SRC= U. */ +void srcu_defer_drain(struct irq_work *iw) +{ + __srcu_defer_drain(container_of(iw, struct srcu_struct, defer_iw), true); +} +EXPORT_SYMBOL_GPL(srcu_defer_drain); + +void call_srcu(struct srcu_struct *ssp, struct rcu_head *rhp, + rcu_callback_t func) +{ + if (should_rcu_defer()) { + /* A re-entrant call_srcu() during the drain would livelock it. */ + if (READ_ONCE(srcu_defer_draining) && !in_nmi()) { + WARN_ONCE(IS_ENABLED(CONFIG_PROVE_RCU), + "call_srcu() re-entered during callback drain; leaking callback\n"); + return; + } + rhp->func =3D func; + if (llist_add((struct llist_node *)rhp, &ssp->defer_cbs)) + irq_work_queue(&ssp->defer_iw); + return; + } + + /* An NMI reaching here entered with irqs enabled, so the enqueue can rac= e. */ + WARN_ON_ONCE(IS_ENABLED(CONFIG_PROVE_RCU) && in_nmi()); + + srcu_do_enqueue(ssp, rhp, func); +} EXPORT_SYMBOL_GPL(call_srcu); =20 /* @@ -262,6 +324,15 @@ void synchronize_srcu(struct srcu_struct *ssp) } EXPORT_SYMBOL_GPL(synchronize_srcu); =20 +/* Register any deferred callbacks, then wait for all in-flight ones. */ +void srcu_barrier(struct srcu_struct *ssp) +{ + if (IS_ENABLED(CONFIG_RCU_DEFER)) + __srcu_defer_drain(ssp, false); + synchronize_srcu(ssp); +} +EXPORT_SYMBOL_GPL(srcu_barrier); + /* * get_state_synchronize_srcu - Provide an end-of-grace-period cookie */ --=20 2.53.0-Meta From nobody Fri Oct 2 03:53:34 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 338DB46AA8D; Wed, 5 Aug 2026 12:24:17 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785932658; cv=none; b=hsrWX0K3xR7S0gKln3gTUsri9/aAn1HdLGsVQz46y91RNQkA/veFbGEUS1IbLGaQlLUalYaim446IMQdhJ6XwdSZzwV7OOOQYOy0z+m8cQY4l7NO0CD0UakbSZ5+ktx/YICwDSV9MIzw7ZZWFYRPS9jf7mpNHvfZslKgVgtWCgo= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785932658; c=relaxed/simple; bh=AHXnQVK2u8e2iwCU1KLipfIlXvpMOke0bbFculNVqkU=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=ElnRAerRdqDhaIDT9cyNas6GZjfIqZUg9CrELyETMmo01eLs8we1YCOPCnmukJ0Hs8rTPKVJbkAUKQj6DaiLx63USa2PZTcOEV4roqwDqm6GiSA+Luj0bAONWv+kI5wXMqN1erVPXZf18AMRdjw7fqNw2JEDJxoPGvVyebfdIJ8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=af+CYDeS; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="af+CYDeS" Received: by smtp.kernel.org (Postfix) with ESMTPSA id D3C2B1F000E9; Wed, 5 Aug 2026 12:24:16 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785932657; bh=gLRQFYdOJ7aPkEIiatEnmo9QdkwxnK0Nw4TUVjsGhJ0=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=af+CYDeS445OxjQX4XrX2XwNLJHIoll+r3UOi8B6Gl0Ehj7ISD7xtqrQaKx7j/+z5 nUs6JkH0KdfUfAk21wuGx/xwbXUTvA/daalLzPpA3ANVNTC5vXd5oZMrCg2fbK1uyw BBj/7EmwAATGJNeXevo7y+Pi/hggw2n7wdus97ypd6MuoTa/ZAGhEG42nzUSWghtJ0 dDEAaJKgNzwHff06S5cgZo52YSBHUgRoMYi19jKpdkZOaT5Hx3pNsg2hs97OUQi09W BolWeGrI3jdYY+KYIXa/4bJZAXQH/XV5BWPdr871ES1GAXe87Cb9+NDt/096Sji0DM 0AMHEMLzoLe/w== From: Puranjay Mohan To: "Lai Jiangshan" , "Paul E. McKenney" , "Josh Triplett" , =?UTF-8?q?Onur=20=C3=96zkan?= , "Frederic Weisbecker" , "Neeraj Upadhyay" , "Joel Fernandes" , "Boqun Feng" , "Uladzislau Rezki" , "Davidlohr Bueso" , "Andrii Nakryiko" , "Eduard Zingerman" , "Alexei Starovoitov" , "Daniel Borkmann" , "Kumar Kartikeya Dwivedi" Cc: Puranjay Mohan , Steven Rostedt , Mathieu Desnoyers , Zqiang , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Matt Fleming , "Harry Yoo (Oracle)" , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, bpf@vger.kernel.org, linux-rt-devel@lists.linux.dev Subject: [PATCH v3 5/6] rcutorture: Exercise ->call() from NMI context Date: Wed, 5 Aug 2026 05:23:42 -0700 Message-ID: <20260805122346.269445-6-puranjay@kernel.org> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260805122346.269445-1-puranjay@kernel.org> References: <20260805122346.269445-1-puranjay@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" call_rcu() and call_srcu() are now safe to invoke from NMI, but rcutorture never does, leaving the deferral path untested. Add an ->nmi_capable flag to rcu_torture_ops. For flavors that set it, arm a per-CPU hardware perf counter whose overflow handler submits a callback via ->call(). The handler acts only when in_nmi(), so only a genuine NMI exercises the deferral path. One preallocated callback is kept in flight (guarded by an atomic) to avoid allocating in NMI. Report the count issued from NMI ("nmi-calls:") and the count invoked ("nmi-cbs:"). rcu_torture_cleanup() disables the counters and then calls cb_barrier(), which drains every deferred callback, so the two counts must then match; a mismatch means a lost callback and fails the test. This relies on srcu_barrier()/rcu_barrier() flushing deferred callbacks, as added earlier in the series. Set ->nmi_capable on the NMI-safe flavors: rcu, srcu, srcud, and tasks-tracing (call_srcu() under the hood). Tasks and Tasks Rude are left alone, as call_rcu_tasks_generic() is not yet NMI-safe. Enabled by default; the nmi_calls parameter disables it, which helps rule NMI handling in or out when triaging a failure. Requires CONFIG_PERF_EVENTS and a hardware PMU, else silently skipped. Signed-off-by: Puranjay Mohan --- kernel/rcu/rcutorture.c | 154 +++++++++++++++++++++++++++++++++++++++- 1 file changed, 152 insertions(+), 2 deletions(-) diff --git a/kernel/rcu/rcutorture.c b/kernel/rcu/rcutorture.c index 39426a8718fe9..715c4c6c51b90 100644 --- a/kernel/rcu/rcutorture.c +++ b/kernel/rcu/rcutorture.c @@ -48,6 +48,7 @@ #include #include #include +#include =20 #include "rcu.h" =20 @@ -115,6 +116,7 @@ torture_param(int, leakpointer, 0, "Leak pointer derefe= rences from readers"); torture_param(int, n_barrier_cbs, 0, "# of callbacks/kthreads for barrier = testing"); torture_param(int, n_up_down, 32, "# of concurrent up/down hrtimer-based R= CU readers"); torture_param(int, nfakewriters, 4, "Number of RCU fake writer threads"); +torture_param(bool, nmi_calls, true, "Exercise ->call() from NMI on nmi_ca= pable flavors"); torture_param(int, nreaders, -1, "Number of RCU reader threads"); torture_param(bool, nwriters, 1, "Number of RCU writer threads (0 or 1)"); torture_param(int, object_debug, 0, "Enable debug-object double call_rcu()= testing"); @@ -216,6 +218,8 @@ static long n_rcu_torture_boost_failure; static long n_rcu_torture_boosts; static atomic_long_t n_rcu_torture_timers; static atomic_long_t n_rcu_torture_irqs; +static atomic_long_t n_rcu_torture_nmi_call; +static atomic_long_t n_rcu_torture_nmi_cb; static long n_barrier_attempts; static long n_barrier_successes; /* did rcu_barrier test succeed? */ static unsigned long n_read_exits; @@ -433,6 +437,7 @@ struct rcu_torture_ops { bool (*is_task_rcu_boosted)(void); long cbflood_max; int irq_capable; + int nmi_capable; int can_boost; int extendables; int slow_gps; @@ -648,6 +653,7 @@ static struct rcu_torture_ops rcu_ops =3D { .extendables =3D RCUTORTURE_MAX_EXTEND, .debug_objects =3D 1, .start_poll_irqsoff =3D 1, + .nmi_capable =3D 1, .name =3D "rcu" }; =20 @@ -942,6 +948,7 @@ static struct rcu_torture_ops srcu_ops =3D { .debug_objects =3D 1, .have_up_down =3D IS_ENABLED(CONFIG_TINY_SRCU) ? 0 : SRCU_READ_FLAVOR_NORMAL | SRCU_READ_FLAVOR_FAST_UPDOWN, + .nmi_capable =3D 1, .name =3D "srcu" }; =20 @@ -1005,6 +1012,7 @@ static struct rcu_torture_ops srcud_ops =3D { .debug_objects =3D 1, .have_up_down =3D IS_ENABLED(CONFIG_TINY_SRCU) ? 0 : SRCU_READ_FLAVOR_NORMAL | SRCU_READ_FLAVOR_FAST_UPDOWN, + .nmi_capable =3D 1, .name =3D "srcud" }; =20 @@ -1269,6 +1277,7 @@ static struct rcu_torture_ops tasks_tracing_ops =3D { .cbflood_max =3D 50000, .irq_capable =3D 1, .slow_gps =3D 1, + .nmi_capable =3D 1, .name =3D "tasks-tracing" }; =20 @@ -2659,12 +2668,131 @@ static bool rcu_torture_one_read(struct torture_ra= ndom_state *trsp, long myid) =20 static DEFINE_TORTURE_RANDOM_PERCPU(rcu_torture_timer_rand); =20 +/* + * Exercise ->call() from NMI context for flavors that set ->nmi_capable. = A + * per-CPU hardware perf counter overflows into an NMI, and its handler su= bmits + * one preallocated callback via ->call(). One callback is in flight at a= time + * (guarded by an atomic) to avoid allocating in NMI. This mirrors how a = BPF + * program reaches ->call() from NMI. + */ +#ifdef CONFIG_PERF_EVENTS +static struct perf_event_attr rcu_torture_nmi_attr =3D { + .type =3D PERF_TYPE_HARDWARE, + .config =3D PERF_COUNT_HW_CPU_CYCLES, + .size =3D sizeof(struct perf_event_attr), + .pinned =3D 1, + .disabled =3D 1, + .freq =3D 1, + .sample_freq =3D 100, +}; + +/* One in-flight callback per CPU; ->inuse is released by the callback. */ +struct rcu_torture_nmi_cbs { + struct rcu_head rh; + atomic_t inuse; +}; + +static struct perf_event **rcu_torture_nmi_events; +static int rcu_torture_nmi_hp_state; +static DEFINE_PER_CPU(struct rcu_torture_nmi_cbs, rcu_torture_nmi_cbs); + +static void rcu_torture_nmi_cb(struct rcu_head *rhp) +{ + struct rcu_torture_nmi_cbs *rtncp =3D container_of(rhp, struct rcu_tortur= e_nmi_cbs, rh); + + atomic_long_inc(&n_rcu_torture_nmi_cb); + atomic_set(&rtncp->inuse, 0); +} + +static void rcu_torture_nmi_overflow(struct perf_event *event, + struct perf_sample_data *data, + struct pt_regs *regs) +{ + struct rcu_torture_nmi_cbs *rtncp =3D this_cpu_ptr(&rcu_torture_nmi_cbs); + + if (!in_nmi()) + return; + if (cur_ops->call && !atomic_xchg(&rtncp->inuse, 1)) { + atomic_long_inc(&n_rcu_torture_nmi_call); + cur_ops->call(&rtncp->rh, rcu_torture_nmi_cb); + } +} + +static int rcu_torture_nmi_online(unsigned int cpu) +{ + struct perf_event *event; + + event =3D perf_event_create_kernel_counter(&rcu_torture_nmi_attr, cpu, NU= LL, + rcu_torture_nmi_overflow, NULL); + if (IS_ERR(event)) + return 0; + rcu_torture_nmi_events[cpu] =3D event; + perf_event_enable(event); + return 0; +} + +static int rcu_torture_nmi_offline(unsigned int cpu) +{ + struct perf_event *event =3D rcu_torture_nmi_events[cpu]; + + if (event) { + rcu_torture_nmi_events[cpu] =3D NULL; + perf_event_disable(event); + perf_event_release_kernel(event); + } + return 0; +} + +/* + * Drive the counters from CPU-hotplug callbacks so that coverage survives= the + * onoff testing that most scenarios run. + */ +static void rcu_torture_nmi_init(void) +{ + int ret; + + if (!nmi_calls || !cur_ops->nmi_capable || !cur_ops->call) + return; + rcu_torture_nmi_events =3D kcalloc(nr_cpu_ids, sizeof(*rcu_torture_nmi_ev= ents), + GFP_KERNEL); + if (!rcu_torture_nmi_events) + return; + ret =3D cpuhp_setup_state(CPUHP_AP_ONLINE_DYN, "rcutorture/nmi:online", + rcu_torture_nmi_online, rcu_torture_nmi_offline); + if (ret < 0) { + kfree(rcu_torture_nmi_events); + rcu_torture_nmi_events =3D NULL; + return; + } + rcu_torture_nmi_hp_state =3D ret; +} + +static void rcu_torture_nmi_cleanup(void) +{ + if (!rcu_torture_nmi_events) + return; + if (rcu_torture_nmi_hp_state > 0) { + cpuhp_remove_state(rcu_torture_nmi_hp_state); + rcu_torture_nmi_hp_state =3D 0; + } + kfree(rcu_torture_nmi_events); + rcu_torture_nmi_events =3D NULL; + if (!atomic_long_read(&n_rcu_torture_nmi_call)) + pr_alert("%s: nmi_calls set but no ->call() issued from NMI: not support= ed here.\n", + __func__); +} +#else /* #ifdef CONFIG_PERF_EVENTS */ +static void rcu_torture_nmi_init(void) { } +static void rcu_torture_nmi_cleanup(void) { } +#endif /* #else #ifdef CONFIG_PERF_EVENTS */ + /* * RCU torture reader from timer handler. Dereferences rcu_torture_curren= t, * incrementing the corresponding element of the pipeline array. The * counter in the element should never be greater than 1, otherwise, the * RCU implementation is broken. */ + static void rcu_torture_timer(struct timer_list *unused) { WARN_ON_ONCE(!in_serving_softirq()); @@ -3047,6 +3175,9 @@ rcu_torture_stats_print(void) data_race(n_barrier_attempts), data_race(n_rcu_torture_barrier_error)); pr_cont("read-exits: %ld ", data_race(n_read_exits)); // Statistic. + pr_cont("nmi-calls: %ld nmi-cbs: %ld ", + atomic_long_read(&n_rcu_torture_nmi_call), + atomic_long_read(&n_rcu_torture_nmi_cb)); pr_cont("nocb-toggles: %ld:%ld ", atomic_long_read(&n_nocb_offload), atomic_long_read(&n_nocb_deoffload)); pr_cont("gpwraps: %ld\n", n_gpwraps); @@ -3195,7 +3326,7 @@ rcu_torture_print_module_parms(struct rcu_torture_ops= *cur_ops, const char *tag) "read_exit_delay=3D%d read_exit_burst=3D%d " "reader_flavor=3D%x " "nocbs_nthreads=3D%d nocbs_toggle=3D%d " - "test_nmis=3D%d " + "test_nmis=3D%d nmi_calls=3D%d " "preempt_duration=3D%d preempt_interval=3D%d n_up_down=3D%d\n", torture_type, tag, nrealreaders, nwriters, nrealfakewriters, stat_interval, verbose, test_no_idle_hz, shuffle_interval, @@ -3209,7 +3340,7 @@ rcu_torture_print_module_parms(struct rcu_torture_ops= *cur_ops, const char *tag) read_exit_delay, read_exit_burst, reader_flavor, nocbs_nthreads, nocbs_toggle, - test_nmis, + test_nmis, nmi_calls, preempt_duration, preempt_interval, n_up_down); } =20 @@ -4283,6 +4414,7 @@ rcu_torture_cleanup(void) int i; =20 if (torture_cleanup_begin()) { + rcu_torture_nmi_cleanup(); if (cur_ops->cb_barrier !=3D NULL) { pr_info("%s: Invoking %pS().\n", __func__, cur_ops->cb_barrier); cur_ops->cb_barrier(); @@ -4325,6 +4457,8 @@ rcu_torture_cleanup(void) kfree(reader_tasks); reader_tasks =3D NULL; } + /* Disable the perf counters (and thus the NMI ->call() firing) now. */ + rcu_torture_nmi_cleanup(); kfree(rcu_torture_reader_mbchk); rcu_torture_reader_mbchk =3D NULL; =20 @@ -4354,6 +4488,20 @@ rcu_torture_cleanup(void) pr_info("%s: Invoking %pS().\n", __func__, cur_ops->cb_barrier); cur_ops->cb_barrier(); } + + /* + * cb_barrier() above drained every deferred NMI ->call() callback, so the + * count issued from NMI must equal the count invoked; a mismatch means a + * callback was lost. + */ + if (atomic_long_read(&n_rcu_torture_nmi_call) !=3D + atomic_long_read(&n_rcu_torture_nmi_cb)) { + pr_alert("%s: NMI ->call() lost a callback: issued %ld invoked %ld\n", + __func__, atomic_long_read(&n_rcu_torture_nmi_call), + atomic_long_read(&n_rcu_torture_nmi_cb)); + atomic_inc(&n_rcu_torture_error); + } + if (cur_ops->cleanup !=3D NULL) cur_ops->cleanup(); =20 @@ -4786,6 +4934,8 @@ rcu_torture_init(void) firsterr =3D -ENOMEM; goto unwind; } + /* Arm the per-CPU perf counters that drive ->call() from NMI. */ + rcu_torture_nmi_init(); for (i =3D 0; i < nrealreaders; i++) { rcu_torture_reader_mbchk[i].rtc_chkrdr =3D -1; firsterr =3D torture_create_kthread(rcu_torture_reader, (void *)i, --=20 2.53.0-Meta From nobody Fri Oct 2 03:53:34 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 70BC84657FD; Wed, 5 Aug 2026 12:24:21 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785932662; cv=none; b=BcFOrVJb+BEJDRrI3Zz6WFCXvSoSBDMn4D1kNlH2+fQTlVjPQBItfErZhfB5NzcUMnKFVNgEnEE263pBHAJldYPgG1rHKCkfLwMTog2RBiQW8naYjj3pxLSak9sSjHPo2vYH/t0QFPf53PIHVRABp780EZ3R5BcnTJWKCVZaHwg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785932662; c=relaxed/simple; bh=u43ECnkxfmdGYW8Bf4pwUp69IBSkp9X0YymN34IHNp4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=nrr1/zAkbOlB67ao68y7x8wYWtPvRULfsdrVkGMgrjKxp3FdFcxWheRrbm/xIbe6oseC/uNjy1Ronuse8CspQQ2hCGa5FhaXSvWAV/yWJM7d1QqlZM74keDfIvtQ2QziQuQ/PFT/Lo/vAG7GJKHqa+pALAWmwAy7M/skdXUAfek= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=dSSC+WFQ; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="dSSC+WFQ" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 1EFAD1F000E9; Wed, 5 Aug 2026 12:24:21 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785932661; bh=wZC1Z+3adPrHkC8Qv8oyMha0PitI7krXLmp3Adhm/jI=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=dSSC+WFQ00umcfx+tqieACzSJnSWVkJCWdn9wqPvWedAh/OxhGwB+Manb2VPnhQsz WCmuPJtBVldoxTSbpDiXLrYwhRZcqI/C4I/WABG2yIsGxBSEfa6fBP7TuiSCCKUdf9 2TEmjqFSlQs7G8SH9HiM+unVa6sqciokFlbYml7Z16f4ToVSHv/oR84Zuq+m0/Qk2x TLxMZOzxvAA4CH6h+g/3jAyYyaascOhSdMHjw7C6nLjlWwUY3zJ11pg8FgikNTLhB8 2si6lGIhhsAMwvcQvfWYs84TCvzVj3NHhQAiqyiJY8JTtgTqf3KkG8kgUx9ux9okgd HdTQ4mmwsUERg== From: Puranjay Mohan To: "Lai Jiangshan" , "Paul E. McKenney" , "Josh Triplett" , =?UTF-8?q?Onur=20=C3=96zkan?= , "Frederic Weisbecker" , "Neeraj Upadhyay" , "Joel Fernandes" , "Boqun Feng" , "Uladzislau Rezki" , "Davidlohr Bueso" , "Andrii Nakryiko" , "Eduard Zingerman" , "Alexei Starovoitov" , "Daniel Borkmann" , "Kumar Kartikeya Dwivedi" Cc: Puranjay Mohan , Steven Rostedt , Mathieu Desnoyers , Zqiang , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Matt Fleming , "Harry Yoo (Oracle)" , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, bpf@vger.kernel.org, linux-rt-devel@lists.linux.dev Subject: [PATCH v3 6/6] selftests/bpf: Add a call_srcu() re-entry reproducer Date: Wed, 5 Aug 2026 05:23:43 -0700 Message-ID: <20260805122346.269445-7-puranjay@kernel.org> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260805122346.269445-1-puranjay@kernel.org> References: <20260805122346.269445-1-puranjay@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Re-enter call_srcu() from a BPF program to exercise its any-context safety (and thus call_rcu_tasks_trace(), which is call_srcu() on rcu_tasks_trace_srcu_struct). An fentry program on rcu_segcblist_enqueue() fires mid-enqueue: that function is reached from srcu_gp_start_if_needed() with the srcu_data ->lock held. The program does a task-storage delete, whose only deferred work is call_rcu_tasks_trace(), re-entering the enqueue on the same CPU. The triggering thread is pinned to one CPU and matched by TID, so the program fires only for the test's own delete. Without the fix the nested call re-takes the same sdp lock and self-deadlocks; with it the nested __call_srcu() sees interrupts disabled and defers via irq_work, so the delete returns and the test passes. Since it can hang an unfixed kernel, run it only against a kernel carrying the fix. Signed-off-by: Puranjay Mohan Acked-by: Kumar Kartikeya Dwivedi --- .../selftests/bpf/prog_tests/rcu_reentry.c | 95 +++++++++++++++++++ .../testing/selftests/bpf/progs/rcu_reentry.c | 51 ++++++++++ 2 files changed, 146 insertions(+) create mode 100644 tools/testing/selftests/bpf/prog_tests/rcu_reentry.c create mode 100644 tools/testing/selftests/bpf/progs/rcu_reentry.c diff --git a/tools/testing/selftests/bpf/prog_tests/rcu_reentry.c b/tools/t= esting/selftests/bpf/prog_tests/rcu_reentry.c new file mode 100644 index 0000000000000..fa813d1d492b6 --- /dev/null +++ b/tools/testing/selftests/bpf/prog_tests/rcu_reentry.c @@ -0,0 +1,95 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Exercise re-entry into call_srcu() from BPF; see progs/rcu_reentry.c. + * + * On a kernel without the call_srcu() any-context fix the nested call + * self-deadlocks on the srcu_data lock, so this hangs rather than fails. + */ +#define _GNU_SOURCE +#include +#include +#include +#include "rcu_reentry.skel.h" + +static int sys_pidfd_open(pid_t pid, unsigned int flags) +{ + return syscall(__NR_pidfd_open, pid, flags); +} + +/* Tiny RCU builds have no rcu_segcblist_enqueue() to attach to. */ +static bool have_attach_target(void) +{ + char buf[256]; + bool found =3D false; + FILE *f; + + f =3D fopen("/proc/kallsyms", "r"); + if (!f) + return true; /* cannot tell; let the attach decide */ + while (fgets(buf, sizeof(buf), f)) { + if (strstr(buf, " rcu_segcblist_enqueue\n")) { + found =3D true; + break; + } + } + fclose(f); + return found; +} + +void test_rcu_reentry(void) +{ + struct rcu_reentry *skel; + int err, pidfd =3D -1, map_fd; + cpu_set_t set, old_set; + bool affinity_saved; + __u64 val =3D 1; + + if (!have_attach_target()) { + test__skip(); + return; + } + + skel =3D rcu_reentry__open_and_load(); + if (!ASSERT_OK_PTR(skel, "skel_open_and_load")) + return; + + err =3D rcu_reentry__attach(skel); + if (!ASSERT_OK(err, "skel_attach")) + goto out; + + /* Keep the re-entry on a single CPU. */ + affinity_saved =3D !sched_getaffinity(0, sizeof(old_set), &old_set); + CPU_ZERO(&set); + CPU_SET(0, &set); + if (!ASSERT_OK(sched_setaffinity(0, sizeof(set), &set), "setaffinity")) + goto out; + + pidfd =3D sys_pidfd_open(getpid(), 0); + if (!ASSERT_GE(pidfd, 0, "pidfd_open")) + goto restore; + map_fd =3D bpf_map__fd(skel->maps.task_stg); + err =3D bpf_map_update_elem(map_fd, &pidfd, &val, BPF_NOEXIST); + if (!ASSERT_OK(err, "boot_create")) + goto restore; + + /* Arm the handler for this thread, then trigger call_rcu_tasks_trace(). = */ + skel->bss->target_pid =3D syscall(__NR_gettid); + err =3D bpf_map_delete_elem(map_fd, &pidfd); + if (!ASSERT_OK(err, "boot_delete")) + goto restore; + + /* Only Tree SRCU enqueues via rcu_segcblist_enqueue(); skip elsewhere. */ + if (!skel->bss->hits) { + test__skip(); + goto restore; + } + ASSERT_EQ(skel->bss->get_errs, 0, "nested_storage_get"); + ASSERT_EQ(skel->bss->del_errs, 0, "nested_storage_delete"); +restore: + if (affinity_saved) + sched_setaffinity(0, sizeof(old_set), &old_set); +out: + if (pidfd >=3D 0) + close(pidfd); + rcu_reentry__destroy(skel); +} diff --git a/tools/testing/selftests/bpf/progs/rcu_reentry.c b/tools/testin= g/selftests/bpf/progs/rcu_reentry.c new file mode 100644 index 0000000000000..47a36f704cf3e --- /dev/null +++ b/tools/testing/selftests/bpf/progs/rcu_reentry.c @@ -0,0 +1,51 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Re-enter call_srcu() from a BPF program. fentry on rcu_segcblist_enque= ue() + * fires inside call_srcu()'s enqueue (reached from srcu_gp_start_if_neede= d() + * with the srcu_data ->lock held); the handler then calls call_rcu_tasks_= trace() + * -- itself call_srcu() on rcu_tasks_trace_srcu_struct -- re-entering the= same + * srcu_data on the same CPU. + */ +#include "vmlinux.h" +#include +#include + +char _license[] SEC("license") =3D "GPL"; + +struct { + __uint(type, BPF_MAP_TYPE_TASK_STORAGE); + __uint(map_flags, BPF_F_NO_PREALLOC); + __type(key, int); + __type(value, __u64); +} task_stg SEC(".maps"); + +int target_pid; +int hits; +int get_errs; +int del_errs; +int done; + +SEC("fentry/rcu_segcblist_enqueue") +int BPF_PROG(reenter) +{ + struct task_struct *cur; + + if (done || !target_pid) + return 0; + + cur =3D bpf_get_current_task_btf(); + if (cur->pid !=3D target_pid) + return 0; + + /* Issue the nested call exactly once, so the test is deterministic. */ + done =3D 1; + __sync_fetch_and_add(&hits, 1); + + /* Re-enter via a task-storage delete, which calls call_rcu_tasks_trace()= . */ + if (!bpf_task_storage_get(&task_stg, cur, 0, BPF_LOCAL_STORAGE_GET_F_CREA= TE)) + __sync_fetch_and_add(&get_errs, 1); + else if (bpf_task_storage_delete(&task_stg, cur)) + __sync_fetch_and_add(&del_errs, 1); + + return 0; +} --=20 2.53.0-Meta