From nobody Fri Oct 2 08:25:15 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 58F6C40F8ED; Mon, 3 Aug 2026 13:53:45 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785765226; cv=none; b=mSx4AG8Fff+ujjUazIeKjBcvByxzbKeaPg/midQLA4jS9xP1UfwYcZET8mYPca4OmzdsqfUtJhoUybsC/I8MDRtfZIcDEMOz24pdyjJdkTfadwOJvUXAGexScfH63Y44kb6OQGP5eioQf0DUbjT9sb8n1YV6QOSa5b+2/JxT2pM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785765226; c=relaxed/simple; bh=SuY+2cmy5+frrBrBWY9FtmrnhhFO7LK0X/x0Q9jIuaQ=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=PujEuvpqEfjUmiadcNtpzEmKlh3EXG3d7WTyA/5AShl1N9Ca8uuLEo6TV/jy5F5HjLMoprpXMmebDLPm3bOCDhAllC6D2N21KE3qbJitTMimP1TvLwKRPVJNo/ubBcUc/jdC0Z7OvtvUthXrC6VQzD2GzNSKJYqR2xaVUxhwilE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=gRean5fo; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="gRean5fo" Received: by smtp.kernel.org (Postfix) with ESMTPSA id ED8141F000E9; Mon, 3 Aug 2026 13:53:44 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785765225; bh=E7FsjnOMScN9KjMOsRPgBmpBN+1GiwYAOe/GbQx/MUA=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=gRean5foXhST/WWlfSKa7djrsDmHzboTv3pG/ikjsBj7KEQ5bf3Dne0Q6xJnng2Rk 9wwBxiUNjmABCBCqCCRwE0sJLbu851QmDiMtUCVqcdYFSoO+f91Mn1Bb6Ae2y3BDE1 BrpChm1PZm5aOW7IZ1sOxeBjK/UZ+BonTqRiVJtl6Rtg7BzAoLOl8GcaAlBrFah06h LpBPceqEoDVGiXiVmtCQ0CVeQjtc2qBEqu9itKdo8z7ytbtrbI4NjwC6hKhTpjV3lt O7qgtXCQ3V0RIm+8pGI3Gm8p6IQDxxPtis92XQboozLUdvGx8jDnTfCaU03BDsLWKF v7/nv0s42Z0aw== From: Puranjay Mohan To: "Lai Jiangshan" , "Paul E. McKenney" , "Josh Triplett" , =?UTF-8?q?Onur=20=C3=96zkan?= , "Frederic Weisbecker" , "Neeraj Upadhyay" , "Joel Fernandes" , "Boqun Feng" , "Uladzislau Rezki" , "Davidlohr Bueso" , "Andrii Nakryiko" , "Eduard Zingerman" , "Alexei Starovoitov" , "Daniel Borkmann" , "Kumar Kartikeya Dwivedi" Cc: Puranjay Mohan , Steven Rostedt , Mathieu Desnoyers , Zqiang , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Matt Fleming , "Harry Yoo (Oracle)" , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, bpf@vger.kernel.org, linux-rt-devel@lists.linux.dev Subject: [PATCH v2 1/6] rcu: Make call_rcu() safe to call from any context Date: Mon, 3 Aug 2026 06:53:24 -0700 Message-ID: <20260803135329.2327280-1-puranjay@kernel.org> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260803134839.2103051-1-puranjay@kernel.org> References: <20260803134839.2103051-1-puranjay@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" RCU's per-CPU callback list is only touched with interrupts disabled: the enqueue runs under local_irq_save() (and the nocb locks when offloaded), as do callback invocation and grace-period work. A call_rcu() that arrives with interrupts already disabled, whether from an NMI or from instrumentation that re-enters RCU, can interrupt one of those and corrupt the list or deadlock. Handle it by deferring: stage the callback on a per-CPU llist and raise an irq_work that re-issues it once interrupts are on. The re-issue goes straight to the enqueue so it cannot defer again. Callers that only hold interrupts off are deferred too, which is harmless. Skip the gate while the scheduler is down (RCU_SCHEDULER_INACTIVE), since irq_work is not usable that early and rcu_init() already calls call_rcu(). rcu_barrier() flushes deferred callbacks before it scans the lists: it waits out each online CPU's irq_work and drains an offline CPU's list directly, since that irq_work may never run again. rcutree_migrate_callbacks() drains an outgoing CPU's ->defer_head for the same reason. ->defer_lock is held across llist_del_all() and the re-issue so these drainers serialize. The re-issue runs with interrupts disabled, so instrumentation on the enqueue path can re-enter call_rcu(), stage another callback, and re-raise the irq_work, livelocking the drain. A per-CPU flag guards it: a call_rcu() that tries to defer while the drain is running, and is not from an NMI, is dropped with WARN_ONCE() instead of re-queued. Dropping leaks the callback, but the alternative is an unbounded loop. The irq_work is IRQ_WORK_INIT_HARD so the re-issue stays prompt on PREEMPT_RT, where a non-HARD irq_work runs in a kthread that can be delayed under load. A hidden CONFIG_RCU_DEFER gates the code and its IRQ_WORK dependency; without it call_rcu() enqueues directly as before. Under CONFIG_PROVE_RCU, warn if the direct path is reached from an NMI. Suggested-by: Paul E. McKenney Signed-off-by: Puranjay Mohan --- kernel/rcu/Kconfig | 6 +++ kernel/rcu/rcu.h | 14 +++++ kernel/rcu/tree.c | 130 +++++++++++++++++++++++++++++++++++++++++---- kernel/rcu/tree.h | 5 ++ 4 files changed, 145 insertions(+), 10 deletions(-) diff --git a/kernel/rcu/Kconfig b/kernel/rcu/Kconfig index f15da8038d0ba..1a5fb3156c062 100644 --- a/kernel/rcu/Kconfig +++ b/kernel/rcu/Kconfig @@ -175,6 +175,12 @@ config RCU_STALL_COMMON config RCU_NEED_SEGCBLIST def_bool ( TREE_RCU || TREE_SRCU || TASKS_RCU_GENERIC ) =20 +# The deferral (and the IRQ_WORK it uses) is only needed where call_rcu() / +# call_srcu() can be invoked while a callback-list operation is in flight. +config RCU_DEFER + def_bool HAVE_NMI || KPROBES || FUNCTION_TRACER || TRACEPOINTS + select IRQ_WORK + config RCU_FANOUT int "Tree-based hierarchical RCU fanout value" range 2 64 if 64BIT diff --git a/kernel/rcu/rcu.h b/kernel/rcu/rcu.h index 39a9f6fa9a7b2..f8add8f8eae15 100644 --- a/kernel/rcu/rcu.h +++ b/kernel/rcu/rcu.h @@ -572,6 +572,20 @@ static inline void tasks_cblist_init_generic(void) { } #define RCU_SCHEDULER_INIT 1 #define RCU_SCHEDULER_RUNNING 2 =20 +/* + * Defer a call_rcu()/call_srcu() callback rather than enqueue it now? De= fer + * whenever interrupts are disabled, since a callback-list operation may b= e in + * flight on this CPU. Not while the scheduler is down, though: irq_work = is + * unusable before init_IRQ(), yet rcu_init() already calls call_rcu(). + */ +static inline bool should_rcu_defer(void) +{ + if (!IS_ENABLED(CONFIG_RCU_DEFER)) + return false; + + return irqs_disabled() && rcu_scheduler_active !=3D RCU_SCHEDULER_INACTIV= E; +} + enum rcutorture_type { RCU_FLAVOR, RCU_TASKS_FLAVOR, diff --git a/kernel/rcu/tree.c b/kernel/rcu/tree.c index 21b6ce1dffb63..744a7cb60db4a 100644 --- a/kernel/rcu/tree.c +++ b/kernel/rcu/tree.c @@ -24,6 +24,7 @@ #include #include #include +#include #include #include #include @@ -3148,21 +3149,28 @@ static void check_cb_ovld(struct rcu_data *rdp) raw_spin_unlock_rcu_node(rnp); } =20 -static void -__call_rcu_common(struct rcu_head *head, rcu_callback_t func, bool lazy_in) +/* + * The callback list is only accessed with interrupts disabled, so a call_= rcu() + * that arrives with interrupts off (see should_rcu_defer()) stages the ca= llback + * on a per-CPU llist that an irq_work re-issues once interrupts are on. + */ +static void rcu_defer_drain(struct irq_work *iw); + +/* Set while rcu_defer_drain() re-issues, to catch a re-entrant call_rcu()= . */ +static DEFINE_PER_CPU(bool, rcu_defer_draining); + +/* + * Enqueue @head on this CPU's rcu_segcblist. Also called by rcu_defer_dr= ain() + * to re-issue a deferred callback, so it must not re-check the deferral + * condition. Either caller may have interrupts already disabled. + */ +static void rcu_do_enqueue(struct rcu_head *head, rcu_callback_t func, boo= l lazy_in) { static atomic_t doublefrees; unsigned long flags; bool lazy; struct rcu_data *rdp; =20 - /* Misaligned rcu_head! */ - WARN_ON_ONCE((unsigned long)head & (sizeof(void *) - 1)); - - /* Avoid NULL dereference if callback is NULL. */ - if (WARN_ON_ONCE(!func)) - return; - if (debug_rcu_head_queue(head)) { /* * Probable double call_rcu(), so leak the callback. @@ -3206,6 +3214,92 @@ __call_rcu_common(struct rcu_head *head, rcu_callbac= k_t func, bool lazy_in) local_irq_restore(flags); } =20 +/* + * Re-issue deferred callbacks straight to the enqueue so they cannot defer + * again. ->defer_lock serializes the drainers: this CPU's irq_work, + * rcu_defer_flush() and rcutree_migrate_callbacks(). + */ +static void rcu_defer_drain(struct irq_work *iw) +{ + struct rcu_data *rdp =3D container_of(iw, struct rcu_data, defer_work); + struct llist_node *node, *next; + unsigned long flags; + + raw_spin_lock_irqsave(&rdp->defer_lock, flags); + this_cpu_write(rcu_defer_draining, true); + llist_for_each_safe(node, next, llist_del_all(&rdp->defer_head)) { + struct rcu_head *head =3D (struct rcu_head *)node; + + rcu_do_enqueue(head, head->func, false); + } + this_cpu_write(rcu_defer_draining, false); + raw_spin_unlock_irqrestore(&rdp->defer_lock, flags); +} + +/* Stage @head for this CPU's irq_work when call_rcu() cannot enqueue now.= */ +static void call_rcu_defer(struct rcu_head *head, rcu_callback_t func) +{ + struct rcu_data *rdp =3D this_cpu_ptr(&rcu_data); + + /* + * Instrumentation on the enqueue path can re-enter here from inside + * rcu_defer_drain(). Re-queuing would livelock the drain, so drop the + * callback; an NMI is one-shot and cannot loop, so let it through. + */ + if (this_cpu_read(rcu_defer_draining) && !in_nmi()) { + WARN_ONCE(1, "call_rcu() re-entered during callback drain; leaking callb= ack\n"); + return; + } + head->func =3D func; + if (llist_add((struct llist_node *)head, &rdp->defer_head)) + irq_work_queue(&rdp->defer_work); +} + +/* + * Register pending deferred callbacks into the callback lists so a follow= ing + * rcu_barrier() waits for them. This runs before rcu_barrier() scans the + * lists. An online CPU's own irq_work re-issues its callbacks, so wait i= t out; + * an offline CPU's irq_work may never run again, so drain its list direct= ly + * onto this CPU instead. + */ +static void rcu_defer_flush(void) +{ + int cpu; + + if (!IS_ENABLED(CONFIG_RCU_DEFER)) + return; + + for_each_possible_cpu(cpu) { + struct rcu_data *rdp =3D per_cpu_ptr(&rcu_data, cpu); + + if (cpu_online(cpu)) + irq_work_sync(&rdp->defer_work); + else + rcu_defer_drain(&rdp->defer_work); + } +} + +static void +__call_rcu_common(struct rcu_head *head, rcu_callback_t func, bool lazy_in) +{ + /* Misaligned rcu_head! */ + WARN_ON_ONCE((unsigned long)head & (sizeof(void *) - 1)); + + /* Avoid NULL dereference if callback is NULL. */ + if (WARN_ON_ONCE(!func)) + return; + + if (should_rcu_defer()) { + call_rcu_defer(head, func); + return; + } + + /* An NMI reaching here entered with irqs enabled, so the enqueue can rac= e. */ + WARN_ON_ONCE(IS_ENABLED(CONFIG_PROVE_RCU) && in_nmi()); + + rcu_do_enqueue(head, func, lazy_in); +} + #ifdef CONFIG_RCU_LAZY static bool enable_rcu_lazy __read_mostly =3D !IS_ENABLED(CONFIG_RCU_LAZY_= DEFAULT_OFF); module_param(enable_rcu_lazy, bool, 0444); @@ -3896,8 +3990,12 @@ void rcu_barrier(void) unsigned long flags; unsigned long gseq; struct rcu_data *rdp; - unsigned long s =3D rcu_seq_snap(&rcu_state.barrier_sequence); + unsigned long s; =20 + /* Register any deferred callbacks before snapshotting the sequence. */ + rcu_defer_flush(); + + s =3D rcu_seq_snap(&rcu_state.barrier_sequence); rcu_barrier_trace(TPS("Begin"), -1, s); =20 /* Take mutex to serialize concurrent rcu_barrier() requests. */ @@ -4231,6 +4329,10 @@ rcu_boot_init_percpu_data(int cpu) rdp->rcu_onl_gp_state =3D RCU_GP_CLEANED; rdp->last_sched_clock =3D jiffies; rdp->cpu =3D cpu; + init_llist_head(&rdp->defer_head); + raw_spin_lock_init(&rdp->defer_lock); + /* Hard irq_work so the re-issue runs promptly. */ + rdp->defer_work =3D IRQ_WORK_INIT_HARD(rcu_defer_drain); rcu_boot_init_nocb_percpu_data(rdp); } =20 @@ -4528,6 +4630,14 @@ void rcutree_migrate_callbacks(int cpu) struct rcu_data *rdp =3D per_cpu_ptr(&rcu_data, cpu); bool needwake; =20 + /* + * Callbacks the outgoing CPU deferred late in the offline path (past the + * point its irq_work can run) sit on ->defer_head, which the ->cblist + * migration below does not cover. Drain them here, before the early + * returns; the re-issue lands on this CPU. + */ + rcu_defer_drain(&rdp->defer_work); + if (rcu_rdp_is_offloaded(rdp)) return; =20 diff --git a/kernel/rcu/tree.h b/kernel/rcu/tree.h index eedfa43059e80..3a8e17136c5a7 100644 --- a/kernel/rcu/tree.h +++ b/kernel/rcu/tree.h @@ -229,6 +229,11 @@ struct rcu_data { struct rcu_head barrier_head; int exp_watching_snap; /* Double-check need for IPI. */ =20 + /* Deferral of an NMI/reentrant call_rcu(); see __call_rcu_common(). */ + struct llist_head defer_head; + struct irq_work defer_work; + raw_spinlock_t defer_lock; + /* 5) Callback offloading. */ #ifdef CONFIG_RCU_NOCB_CPU struct swait_queue_head nocb_cb_wq; /* For nocb kthreads to sleep on. */ --=20 2.53.0-Meta From nobody Fri Oct 2 08:25:15 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9F90B430CD7; Mon, 3 Aug 2026 13:53:49 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785765230; cv=none; b=bIroINwwgL+NQldj9Y5J0F0eoty2+Ug/hSKxgi9oXCw4RMNwTJl6yEdWXUlqkfFuZtYkvSn+LzQATxl/jWAEzsBXXP4S/j7MEi636VwK1FuSTdk4hLDYF57jEHMDDfKQbT6OPUQcRFce8k7uWMCOYstO6mMieM1JyZH4+TO2k6U= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785765230; c=relaxed/simple; bh=3JdWddTI48wnb/fe2oe4n46d/M3HF6MGQ3GXZ7fKgos=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=gAb5q4HBKWsLh/Pc/ofpLZfWYwrk3bdLBnS8d/wrcYGJ9CYJU2e3S+tPWIPuspioMxlM0d55nIXk3lR15ldRbgCNNB3BokKL9izEdEYvXU9riw9AhjZ2ddgAMvc8f8ZvbUFg+V6P5Llyd8LPZoVOoUT11gT6BX272j+ePeHNxrs= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=OqQOAh0s; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="OqQOAh0s" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 14A1F1F000E9; Mon, 3 Aug 2026 13:53:49 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785765229; bh=sM/jD8c1DZVPZ/ENWd+7nK7/Ih+2gZL3XN/QDhd+77o=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=OqQOAh0sbQkjiesz7kcnZrq7CVO5HLK/Z/KSMgCLh/+jIZ0nsiKKymUSMkM5VayT+ JvdRDd9VbVN/qWvggvOQG55zvrrQ6ZQdbaufH1vc8Zz+7BnoKOB7MoI5IXNWdLZUYw JXyINrOf1dgIMglPhGhkdZrNgeYMGd8N9dcMRnBZSjTVx3TPtWR/mMqkKlt6MQUf4W diBWRS60smfy24m2eYaXvUyedbYbLHEhHfxCidfnMa2EYkXYVN+6BwXsgzbhq9aGsa rhDXBZsYhbAbpFMEhoNk1ZxMp7WZErgwj0YuI8mxORcxtd5Riz/TNUJNYCD4mzKtcb D4RfieUGePksQ== From: Puranjay Mohan To: "Lai Jiangshan" , "Paul E. McKenney" , "Josh Triplett" , =?UTF-8?q?Onur=20=C3=96zkan?= , "Frederic Weisbecker" , "Neeraj Upadhyay" , "Joel Fernandes" , "Boqun Feng" , "Uladzislau Rezki" , "Davidlohr Bueso" , "Andrii Nakryiko" , "Eduard Zingerman" , "Alexei Starovoitov" , "Daniel Borkmann" , "Kumar Kartikeya Dwivedi" Cc: Puranjay Mohan , Steven Rostedt , Mathieu Desnoyers , Zqiang , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Matt Fleming , "Harry Yoo (Oracle)" , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, bpf@vger.kernel.org, linux-rt-devel@lists.linux.dev Subject: [PATCH v2 2/6] rcu: Make Tiny call_rcu() safe to call from any context Date: Mon, 3 Aug 2026 06:53:25 -0700 Message-ID: <20260803135329.2327280-2-puranjay@kernel.org> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260803134839.2103051-1-puranjay@kernel.org> References: <20260803134839.2103051-1-puranjay@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Give Tiny call_rcu() the same treatment as Tree RCU. When interrupts are disabled and the scheduler is up, stage the callback on a lockless list that an irq_work re-issues later. One global list and irq_work suffice since Tiny RCU is uniprocessor, and there is no CPU-offline drain. The re-issue runs with interrupts disabled and can be re-entered by instrumentation, so a draining flag drops a deferring call_rcu() seen mid-drain (unless from an NMI), as in Tree RCU. Gated by CONFIG_RCU_DEFER. Suggested-by: Paul E. McKenney Signed-off-by: Puranjay Mohan --- kernel/rcu/tiny.c | 104 +++++++++++++++++++++++++++++++++++++--------- 1 file changed, 85 insertions(+), 19 deletions(-) diff --git a/kernel/rcu/tiny.c b/kernel/rcu/tiny.c index dccccd6be9411..5736b964d8ee1 100644 --- a/kernel/rcu/tiny.c +++ b/kernel/rcu/tiny.c @@ -11,6 +11,8 @@ */ #include #include +#include +#include #include #include #include @@ -42,8 +44,86 @@ static struct rcu_ctrlblk rcu_ctrlblk =3D { .gp_seq =3D 0 - 300UL, }; =20 +/* + * The callback list is only accessed with interrupts disabled, so a call_= rcu() + * that arrives with interrupts off stages the callback on a lockless list= that + * an irq_work re-issues later. One global list and irq_work suffice, as = Tiny + * RCU is uniprocessor. + */ +static void rcu_defer_drain(struct irq_work *iw); +static LLIST_HEAD(rcu_defer_list); +static DEFINE_IRQ_WORK(rcu_defer_iw, rcu_defer_drain); +static bool rcu_defer_draining; + +/* + * Enqueue @head on the callback list. Also called by rcu_defer_drain() to + * re-issue a deferred callback, so it must not re-check the deferral cond= ition. + */ +static void rcu_do_enqueue(struct rcu_head *head, rcu_callback_t func) +{ + static atomic_t doublefrees; + unsigned long flags; + + if (debug_rcu_head_queue(head)) { + if (atomic_inc_return(&doublefrees) < 4) { + pr_err("%s(): Double-freed CB %p->%pS()!!! ", __func__, head, head->fu= nc); + mem_dump_obj(head); + } + return; + } + + head->func =3D func; + head->next =3D NULL; + + local_irq_save(flags); + *rcu_ctrlblk.curtail =3D head; + rcu_ctrlblk.curtail =3D &head->next; + local_irq_restore(flags); + + if (unlikely(is_idle_task(current))) { + /* force scheduling for rcu_qs() */ + resched_cpu(0); + } +} + +static void rcu_defer_drain(struct irq_work *iw) +{ + struct llist_node *node, *next; + + /* Callbacks are unordered, so drain in llist order without reversing. */ + rcu_defer_draining =3D true; + llist_for_each_safe(node, next, llist_del_all(&rcu_defer_list)) { + struct rcu_head *head =3D (struct rcu_head *)node; + + rcu_do_enqueue(head, head->func); + } + rcu_defer_draining =3D false; +} + +static void call_rcu_defer(struct rcu_head *head, rcu_callback_t func) +{ + /* A re-entrant call_rcu() during the drain would livelock it; drop it. */ + if (rcu_defer_draining && !in_nmi()) { + WARN_ONCE(1, "call_rcu() re-entered during callback drain; leaking callb= ack\n"); + return; + } + head->func =3D func; + if (llist_add((struct llist_node *)head, &rcu_defer_list)) + irq_work_queue(&rcu_defer_iw); +} + +/* Register any deferred callbacks so a following rcu_barrier() waits for = them. */ +static void rcu_defer_flush(void) +{ + if (!IS_ENABLED(CONFIG_RCU_DEFER)) + return; + irq_work_sync(&rcu_defer_iw); +} + void rcu_barrier(void) { + /* Register any deferred callbacks first. */ + rcu_defer_flush(); wait_rcu_gp(call_rcu_hurry); } EXPORT_SYMBOL(rcu_barrier); @@ -157,29 +237,15 @@ EXPORT_SYMBOL_GPL(synchronize_rcu); */ void call_rcu(struct rcu_head *head, rcu_callback_t func) { - static atomic_t doublefrees; - unsigned long flags; - - if (debug_rcu_head_queue(head)) { - if (atomic_inc_return(&doublefrees) < 4) { - pr_err("%s(): Double-freed CB %p->%pS()!!! ", __func__, head, head->fu= nc); - mem_dump_obj(head); - } + if (should_rcu_defer()) { + call_rcu_defer(head, func); return; } =20 - head->func =3D func; - head->next =3D NULL; - - local_irq_save(flags); - *rcu_ctrlblk.curtail =3D head; - rcu_ctrlblk.curtail =3D &head->next; - local_irq_restore(flags); + /* An NMI reaching here entered with irqs enabled, so the enqueue can rac= e. */ + WARN_ON_ONCE(IS_ENABLED(CONFIG_PROVE_RCU) && in_nmi()); =20 - if (unlikely(is_idle_task(current))) { - /* force scheduling for rcu_qs() */ - resched_cpu(0); - } + rcu_do_enqueue(head, func); } EXPORT_SYMBOL_GPL(call_rcu); =20 --=20 2.53.0-Meta From nobody Fri Oct 2 08:25:15 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id F23D4430CD7; Mon, 3 Aug 2026 13:53:53 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785765235; cv=none; b=IU9vQeSBiEQOvZ7rfG+AOkQSRCCf6TS0q/CMDaT/YDPnXuEoNqw6//OoYMBujhMjT2BGA2mAOdKW0wgx+Z22IlLMNSuPcM92T+Br0AJ89Znv/p/T0gHvSJTVqpr7OpfdYSdnPmKlidt2ct36LXARFu8E+HqGP/DBVML0VQOoHw4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785765235; c=relaxed/simple; bh=sgX5v03Uf3SdzrUd6juElSLoELKtnTfEvlitg3o4GHM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Gd+O+0tynUMAhy554cv+FNh1e/JZXD06canFk5bU0QvdU02lYqdKs4HFbPrxUBDxfCTjobu3cN9HBSH9Rn9Oc/zS8xQNP4PJk0hYz1dsx1SySsqjtBxdoc+r7ArgQf0xkEBQ7DVEdDoiDo1gXl5wjfjYd5iu1vzXDQN0B12BPr8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=n2Wm0XeT; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="n2Wm0XeT" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 5704D1F000E9; Mon, 3 Aug 2026 13:53:53 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785765233; bh=haeBS9YY2oto0H157hsY6qsK+0st9zbTF5qicBAxnag=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=n2Wm0XeT23wGtqzBPkwe72czVazqbJhmBqL60blQN6M5jXROKFO5WQLl23HKXaJeg bVHMCcMsBHUyKyP4bxMduIRSimT6aHRL/krXCTkICdvcsBp8zLHkQTIlvTZa2Eq9dN qq+Rho0BC/MbX8phBH59YTAGCqQ1oxY8A3RKf6zQCi0zx15lZZieCUsgVAEtROcCDC Ls7Gacx1OA1DxluaYC98iODlS7fZpWyh7RorUtGdO+nDw8aHqrob+f0OFLdujy/95i YpK28iBK/8Vv7AyduQwoXLWea93KK1iO8SdN/9qFkKy0vFYNvxCBl6sXAd3WHUWyw+ PCcw1XhhogZNg== From: Puranjay Mohan To: "Lai Jiangshan" , "Paul E. McKenney" , "Josh Triplett" , =?UTF-8?q?Onur=20=C3=96zkan?= , "Frederic Weisbecker" , "Neeraj Upadhyay" , "Joel Fernandes" , "Boqun Feng" , "Uladzislau Rezki" , "Davidlohr Bueso" , "Andrii Nakryiko" , "Eduard Zingerman" , "Alexei Starovoitov" , "Daniel Borkmann" , "Kumar Kartikeya Dwivedi" Cc: Puranjay Mohan , Steven Rostedt , Mathieu Desnoyers , Zqiang , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Matt Fleming , "Harry Yoo (Oracle)" , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, bpf@vger.kernel.org, linux-rt-devel@lists.linux.dev Subject: [PATCH v2 3/6] srcu: Make call_srcu() safe to call from any context Date: Mon, 3 Aug 2026 06:53:26 -0700 Message-ID: <20260803135329.2327280-3-puranjay@kernel.org> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260803134839.2103051-1-puranjay@kernel.org> References: <20260803134839.2103051-1-puranjay@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" call_srcu() has the same constraint as call_rcu(): its callback list and locks are only touched with interrupts disabled. srcu_gp_start_if_needed() enqueues under raw_spin_lock_irqsave() and may walk the srcu_node tree, as do callback invocation and grace-period work. A call_srcu() with interrupts disabled can race a list operation in flight on this CPU and corrupt the list or deadlock. call_rcu_tasks_trace() is call_srcu() under the hood, so a sleepable BPF program freeing an object can reach this. Defer as call_rcu() does: stage the callback on the srcu_data's ->defer_cbs, chain that srcu_data onto a per-CPU list, and raise a per-CPU irq_work that re-issues it straight to the enqueue helper, never back through __call_srcu(). The common path is unchanged and keeps interrupts enabled across srcu_gp_start_if_needed(). The irq_work is per-CPU rather than per-srcu_struct and statically initialized, so deferral never runs check_init_srcu_struct(); it is IRQ_WORK_INIT_HARD as for call_rcu(). srcu_barrier() and cleanup_srcu_struct() flush it first, and rcutree_migrate_callbacks() calls srcu_offline_drain() for an outgoing CPU. ->lock is held across the drain so the drainers serialize. As in call_rcu(), the re-issue runs with interrupts disabled and can be re-entered by instrumentation, so a per-CPU flag drops a deferring call_srcu() seen mid-drain (unless from an NMI). Gated by CONFIG_RCU_DEFER. Under CONFIG_PROVE_RCU, warn if the direct path is reached from an NMI. Suggested-by: Paul E. McKenney Signed-off-by: Puranjay Mohan --- include/linux/srcutree.h | 4 ++ kernel/rcu/rcu.h | 3 + kernel/rcu/srcutree.c | 150 +++++++++++++++++++++++++++++++++++++-- kernel/rcu/tree.c | 2 + 4 files changed, 155 insertions(+), 4 deletions(-) diff --git a/include/linux/srcutree.h b/include/linux/srcutree.h index 75e54e4f963fa..1ce759fb70948 100644 --- a/include/linux/srcutree.h +++ b/include/linux/srcutree.h @@ -13,6 +13,8 @@ =20 #include #include +#include +#include =20 struct srcu_node; struct srcu_struct; @@ -41,6 +43,8 @@ struct srcu_data { bool srcu_cblist_invoking; /* Invoking these CBs? */ struct timer_list delay_work; /* Delay for CB invoking */ struct work_struct work; /* Context for CB invoking. */ + struct llist_head defer_cbs; /* Callbacks deferred on re-entry. */ + struct llist_node defer_link; /* Links onto the per-CPU deferral drain l= ist */ struct rcu_head srcu_barrier_head; /* For srcu_barrier() use. */ struct rcu_head srcu_ec_head; /* For srcu_expedite_current() use. */ int srcu_ec_state; /* State for srcu_expedite_current(). */ diff --git a/kernel/rcu/rcu.h b/kernel/rcu/rcu.h index f8add8f8eae15..ca05d48773c79 100644 --- a/kernel/rcu/rcu.h +++ b/kernel/rcu/rcu.h @@ -586,6 +586,9 @@ static inline bool should_rcu_defer(void) return irqs_disabled() && rcu_scheduler_active !=3D RCU_SCHEDULER_INACTIV= E; } =20 +/* Drain an outgoing CPU's deferred SRCU callbacks; see rcutree_migrate_ca= llbacks(). */ +void srcu_offline_drain(int cpu); + enum rcutorture_type { RCU_FLAVOR, RCU_TASKS_FLAVOR, diff --git a/kernel/rcu/srcutree.c b/kernel/rcu/srcutree.c index 304112674e8a2..2669594a6402f 100644 --- a/kernel/rcu/srcutree.c +++ b/kernel/rcu/srcutree.c @@ -20,6 +20,7 @@ #include #include #include +#include #include #include #include @@ -79,6 +80,47 @@ static void process_srcu(struct work_struct *work); static void srcu_irq_work(struct irq_work *work); static void srcu_delay_timer(struct timer_list *t); =20 +static void srcu_defer_drain(struct irq_work *iw); + +/* + * Per-CPU call_srcu() deferral state, shared by every srcu_struct. A def= erred + * callback is staged on its srcu_data's ->defer_cbs; that srcu_data is ch= ained + * via ->defer_link onto ->list, which the irq_work walks. + */ +struct srcu_defer { + struct llist_head list; + struct irq_work iw; + raw_spinlock_t lock; +}; + +static DEFINE_PER_CPU(struct srcu_defer, srcu_defer) =3D { + .lock =3D __RAW_SPIN_LOCK_UNLOCKED(srcu_defer.lock), + .iw =3D IRQ_WORK_INIT_HARD(srcu_defer_drain), +}; + +/* Set while srcu_defer_drain() re-issues, to catch a re-entrant call_srcu= (). */ +static DEFINE_PER_CPU(bool, srcu_defer_draining); + +/* + * Flush pending deferred callbacks so a following srcu_barrier() waits fo= r them. + * Wait out an online CPU's irq_work; drain an offline CPU's list directly= , as + * its irq_work may never run again. + */ +static void srcu_defer_flush(void) +{ + int cpu; + + if (!IS_ENABLED(CONFIG_RCU_DEFER)) + return; + + for_each_possible_cpu(cpu) { + if (cpu_online(cpu)) + irq_work_sync(&per_cpu(srcu_defer, cpu).iw); + else + srcu_defer_drain(&per_cpu(srcu_defer, cpu).iw); + } +} + /* * Initialize SRCU per-CPU data. Note that statically allocated * srcu_struct structures might already have srcu_read_lock() and @@ -107,6 +149,11 @@ static void init_srcu_struct_data(struct srcu_struct *= ssp) sdp->cpu =3D cpu; INIT_WORK(&sdp->work, srcu_invoke_callbacks); timer_setup(&sdp->delay_work, srcu_delay_timer, 0); + /* + * ->defer_cbs and ->defer_link are valid when zeroed and are not + * reinitialized here, lest we clobber callbacks a reentrant + * call_srcu() already staged. See __call_srcu(). + */ sdp->ssp =3D ssp; } } @@ -695,7 +742,12 @@ void cleanup_srcu_struct(struct srcu_struct *ssp) return; /* Just leak it! */ if (WARN_ON(srcu_readers_active(ssp))) return; /* Just leak it! */ - /* Wait for irq_work to finish first as it may queue a new work. */ + /* + * Drain deferred callbacks before syncing ->irq_work: re-issuing one can + * start a grace period and re-queue ->irq_work, which then schedules + * ->work, so both must be waited out after the drain. + */ + srcu_defer_flush(); irq_work_sync(&sup->irq_work); flush_delayed_work(&sup->work); for_each_possible_cpu(cpu) { @@ -1410,8 +1462,15 @@ static unsigned long srcu_gp_start_if_needed(struct = srcu_struct *ssp, * srcu_read_lock(), and srcu_read_unlock() that are all passed the same * srcu_struct structure. */ -static void __call_srcu(struct srcu_struct *ssp, struct rcu_head *rhp, - rcu_callback_t func, bool do_norm) +/* + * The srcu_cblist and srcu_node tree are only accessed with interrupts di= sabled + * (srcu_gp_start_if_needed() enqueues under raw_spin_lock_irqsave() and m= ay walk + * the tree). Like call_rcu(), __call_srcu() defers when interrupts are a= lready + * disabled, so a re-entrant call_srcu() -- e.g. call_rcu_tasks_trace() fr= om a + * BPF program -- cannot corrupt the list or deadlock. + */ +static void srcu_do_enqueue(struct srcu_struct *ssp, struct rcu_head *rhp, + rcu_callback_t func, bool do_norm) { if (debug_rcu_head_queue(rhp)) { /* Probable double call_srcu(), so leak the callback. */ @@ -1423,6 +1482,81 @@ static void __call_srcu(struct srcu_struct *ssp, str= uct rcu_head *rhp, (void)srcu_gp_start_if_needed(ssp, rhp, do_norm); } =20 +static void __call_srcu(struct srcu_struct *ssp, struct rcu_head *rhp, + rcu_callback_t func, bool do_norm) +{ + if (should_rcu_defer()) { + struct srcu_data *sdp; + + /* + * Instrumentation on the enqueue path can re-enter here from + * inside srcu_defer_drain(). Re-queuing would livelock the + * drain, so drop the callback; an NMI cannot loop, so let it in. + */ + if (this_cpu_read(srcu_defer_draining) && !in_nmi()) { + WARN_ONCE(1, "call_srcu() re-entered during callback drain; leaking cal= lback\n"); + return; + } + sdp =3D this_cpu_ptr(ssp->sda); + rhp->func =3D func; + if (llist_add((struct llist_node *)rhp, &sdp->defer_cbs)) { + /* First deferral on this srcu_data: chain it for the drain. */ + struct srcu_defer *sndp =3D this_cpu_ptr(&srcu_defer); + + sdp->ssp =3D ssp; + if (llist_add(&sdp->defer_link, &sndp->list)) + irq_work_queue(&sndp->iw); + } + return; + } + + /* An NMI reaching here entered with irqs enabled, so the enqueue can rac= e. */ + WARN_ON_ONCE(IS_ENABLED(CONFIG_PROVE_RCU) && in_nmi()); + + srcu_do_enqueue(ssp, rhp, func, do_norm); +} + +/* + * Re-issue deferred callbacks straight to srcu_do_enqueue() so they canno= t defer + * again. ->lock serializes the drainers: the irq_work, srcu_defer_flush(= ) and + * srcu_offline_drain(). + */ +static void srcu_defer_drain(struct irq_work *iw) +{ + struct srcu_defer *sndp =3D container_of(iw, struct srcu_defer, iw); + struct llist_node *snode, *snext; + unsigned long flags; + + raw_spin_lock_irqsave(&sndp->lock, flags); + this_cpu_write(srcu_defer_draining, true); + llist_for_each_safe(snode, snext, llist_del_all(&sndp->list)) { + struct srcu_data *sdp =3D container_of(snode, struct srcu_data, defer_li= nk); + struct srcu_struct *ssp =3D sdp->ssp; + struct llist_node *cnode, *cnext; + + cnode =3D llist_del_all(&sdp->defer_cbs); + llist_for_each_safe(cnode, cnext, cnode) { + struct rcu_head *rhp =3D (struct rcu_head *)cnode; + + srcu_do_enqueue(ssp, rhp, rhp->func, true); + } + } + this_cpu_write(srcu_defer_draining, false); + raw_spin_unlock_irqrestore(&sndp->lock, flags); +} + +/* + * Drain @cpu's deferred call_srcu() callbacks from rcutree_migrate_callba= cks() + * once @cpu is dead. One pass covers every srcu_struct, and the re-issue= lands + * on the current CPU. + */ +void srcu_offline_drain(int cpu) +{ + if (!IS_ENABLED(CONFIG_RCU_DEFER)) + return; + srcu_defer_drain(&per_cpu(srcu_defer, cpu).iw); +} + /** * call_srcu() - Queue a callback for invocation after an SRCU grace period * @ssp: srcu_struct in queue the callback @@ -1677,9 +1811,17 @@ void srcu_barrier(struct srcu_struct *ssp) { int cpu; int idx; - unsigned long s =3D rcu_seq_snap(&ssp->srcu_sup->srcu_barrier_seq); + unsigned long s; =20 check_init_srcu_struct(ssp); + + /* + * Register any deferred callbacks before snapshotting the sequence. The + * shared irq_work may also drain other srcu_structs', which is harmless. + */ + srcu_defer_flush(); + + s =3D rcu_seq_snap(&ssp->srcu_sup->srcu_barrier_seq); mutex_lock(&ssp->srcu_sup->srcu_barrier_mutex); if (rcu_seq_done(&ssp->srcu_sup->srcu_barrier_seq, s)) { smp_mb(); /* Force ordering following return. */ diff --git a/kernel/rcu/tree.c b/kernel/rcu/tree.c index 744a7cb60db4a..3f95c941ed977 100644 --- a/kernel/rcu/tree.c +++ b/kernel/rcu/tree.c @@ -4637,6 +4637,8 @@ void rcutree_migrate_callbacks(int cpu) * returns; the re-issue lands on this CPU. */ rcu_defer_drain(&rdp->defer_work); + /* Likewise for the outgoing CPU's deferred call_srcu() callbacks. */ + srcu_offline_drain(cpu); =20 if (rcu_rdp_is_offloaded(rdp)) return; --=20 2.53.0-Meta From nobody Fri Oct 2 08:25:15 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B8DB24322FA; Mon, 3 Aug 2026 13:53:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785765239; cv=none; b=ga3v3DdsdspZntPy9E4SUHB5JGfj/BHGGvfYrfkML1Yy8oMLjCdkcKzL9a/r2w6c3h2Kl0gtv79T1VDNHxRu3ZSvs86BYENwc+GQ7JVcml4fTarAItOPFfea/H9YpRi34lwmE0mx3eG6PglkMu8LHfCGz6+Y6HDWAxezoxjcjUE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785765239; c=relaxed/simple; bh=fJkdKO3QR4xIxABIq31W7cUWz9W4+q87lI0ykMlCOfk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=czbtXGUqdwvbbRGKmP4n/J1QrGS/w0QMwWHFmJSWFE+VWEVUkjjiYubmm9erGTnflzLrHnt3qto0gNJsUMrVUwHWQleGjfs2JecxK58Od32TGHDwAgqTzHoOOU72fLx+uGRfCC6LWKQei6XRExXQa90e/YeU0sEbVknnEXd5m6E= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=TrOFvIGr; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="TrOFvIGr" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 2CAE21F000E9; Mon, 3 Aug 2026 13:53:57 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785765237; bh=UW/yu4v8sY9etcTcyqgcC/GqOw7GQXjCCRdpg+TEMmI=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=TrOFvIGrQOmvNXyJ8LzUwztA9t8h/TrkvfTxbqRDoWMau29NuEhngB4XMJU/jLv4K WsNFGpMHw8zfpK2vxMfkWifkrh/zTx3EdBhA3Nvfe4xTItIQbZMh/rz1U/184jWVLJ InPhRzZkA38B9DBsWNi56xeS75r8QU35AT1q2jgAbceYr4FNLLOzrSGt0LThki6Ryu 56ZhM2myQBjYZ1VMAsCuJoclvnfKdiXg+SJ0qu2v0PAP5UAVdR44rLYeLLjP2osOcT MJWMqJ9Rd5AhJIZDlB8HRe3/B5d35P4IqmzReVIPq4E0wFFo28Mw3dc+0aSh9x1xvI Z3JsOYNIsiAdg== From: Puranjay Mohan To: "Lai Jiangshan" , "Paul E. McKenney" , "Josh Triplett" , =?UTF-8?q?Onur=20=C3=96zkan?= , "Frederic Weisbecker" , "Neeraj Upadhyay" , "Joel Fernandes" , "Boqun Feng" , "Uladzislau Rezki" , "Davidlohr Bueso" , "Andrii Nakryiko" , "Eduard Zingerman" , "Alexei Starovoitov" , "Daniel Borkmann" , "Kumar Kartikeya Dwivedi" Cc: Puranjay Mohan , Steven Rostedt , Mathieu Desnoyers , Zqiang , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Matt Fleming , "Harry Yoo (Oracle)" , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, bpf@vger.kernel.org, linux-rt-devel@lists.linux.dev Subject: [PATCH v2 4/6] srcu: Make Tiny call_srcu() safe to call from any context Date: Mon, 3 Aug 2026 06:53:27 -0700 Message-ID: <20260803135329.2327280-4-puranjay@kernel.org> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260803134839.2103051-1-puranjay@kernel.org> References: <20260803134839.2103051-1-puranjay@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Give Tiny call_srcu() the same treatment as Tree SRCU. When interrupts are disabled and the scheduler is up, stage the callback on the srcu_struct's lockless list for an irq_work to re-issue later. Tiny SRCU is uniprocessor, so there is no CPU-offline drain. A draining flag drops a deferring call_srcu() that re-enters mid-drain (unless from an NMI), as in Tree SRCU. srcu_barrier() (now out of line) and cleanup_srcu_struct() sync the irq_work before checking for outstanding callbacks, so a deferred callback is re-issued onto the callback list, where the leak checks can see it, rather than stranded on a soon-to-be-freed srcu_struct. Gated by CONFIG_RCU_DEFER, like Tree SRCU. Suggested-by: Paul E. McKenney Signed-off-by: Puranjay Mohan --- include/linux/srcutiny.h | 11 ++++--- kernel/rcu/srcutiny.c | 63 +++++++++++++++++++++++++++++++++++++--- 2 files changed, 66 insertions(+), 8 deletions(-) diff --git a/include/linux/srcutiny.h b/include/linux/srcutiny.h index fbcf13bc12d15..47275d182966c 100644 --- a/include/linux/srcutiny.h +++ b/include/linux/srcutiny.h @@ -12,6 +12,7 @@ #define _LINUX_SRCU_TINY_H =20 #include +#include #include =20 struct srcu_struct { @@ -26,6 +27,8 @@ struct srcu_struct { struct rcu_head **srcu_cb_tail; /* Pending callbacks: Tail. */ struct work_struct srcu_work; /* For driving grace periods. */ struct irq_work srcu_irq_work; /* Defer schedule_work() to irq work. */ + struct llist_head defer_cbs; /* Callbacks deferred on re-entry. */ + struct irq_work defer_iw; /* Registers defer_cbs later. */ #ifdef CONFIG_DEBUG_LOCK_ALLOC struct lockdep_map dep_map; #endif /* #ifdef CONFIG_DEBUG_LOCK_ALLOC */ @@ -33,6 +36,7 @@ struct srcu_struct { =20 void srcu_drive_gp(struct work_struct *wp); void srcu_tiny_irq_work(struct irq_work *irq_work); +void srcu_defer_drain(struct irq_work *irq_work); =20 #define __SRCU_STRUCT_INIT(name, __ignored, ___ignored, ____ignored) \ { \ @@ -40,6 +44,8 @@ void srcu_tiny_irq_work(struct irq_work *irq_work); .srcu_cb_tail =3D &name.srcu_cb_head, \ .srcu_work =3D __WORK_INITIALIZER(name.srcu_work, srcu_drive_gp), \ .srcu_irq_work =3D { .func =3D srcu_tiny_irq_work }, \ + .defer_cbs =3D LLIST_HEAD_INIT(name.defer_cbs), \ + .defer_iw =3D { .func =3D srcu_defer_drain }, \ __SRCU_DEP_MAP_INIT(name) \ } =20 @@ -131,10 +137,7 @@ static inline void synchronize_srcu_expedited(struct s= rcu_struct *ssp) synchronize_srcu(ssp); } =20 -static inline void srcu_barrier(struct srcu_struct *ssp) -{ - synchronize_srcu(ssp); -} +void srcu_barrier(struct srcu_struct *ssp); =20 static inline void srcu_expedite_current(struct srcu_struct *ssp) { } #define srcu_check_read_flavor(ssp, read_flavor) do { } while (0) diff --git a/kernel/rcu/srcutiny.c b/kernel/rcu/srcutiny.c index f9c498ae75df2..988819e6ddd4b 100644 --- a/kernel/rcu/srcutiny.c +++ b/kernel/rcu/srcutiny.c @@ -10,6 +10,7 @@ =20 #include #include +#include #include #include #include @@ -43,6 +44,8 @@ static int init_srcu_struct_fields(struct srcu_struct *ss= p) INIT_WORK(&ssp->srcu_work, srcu_drive_gp); INIT_LIST_HEAD(&ssp->srcu_work.entry); init_irq_work(&ssp->srcu_irq_work, srcu_tiny_irq_work); + init_llist_head(&ssp->defer_cbs); + init_irq_work(&ssp->defer_iw, srcu_defer_drain); return 0; } =20 @@ -86,6 +89,9 @@ EXPORT_SYMBOL_GPL(init_srcu_struct_generic); void cleanup_srcu_struct(struct srcu_struct *ssp) { WARN_ON(srcu_readers_active(ssp)); + /* Re-issue any deferred callbacks so ->srcu_cb_head sees them below. */ + if (IS_ENABLED(CONFIG_RCU_DEFER)) + irq_work_sync(&ssp->defer_iw); irq_work_sync(&ssp->srcu_irq_work); flush_work(&ssp->srcu_work); WARN_ON(ssp->srcu_gp_running); @@ -215,11 +221,11 @@ static void srcu_gp_start_if_needed(struct srcu_struc= t *ssp) } =20 /* - * Enqueue an SRCU callback on the specified srcu_struct structure, - * initiating grace-period processing if it is not already running. + * Enqueue @rhp on the callback list. Also called by srcu_defer_drain() to + * re-issue a deferred callback, so it must not re-check the deferral cond= ition. */ -void call_srcu(struct srcu_struct *ssp, struct rcu_head *rhp, - rcu_callback_t func) +static void srcu_do_enqueue(struct srcu_struct *ssp, struct rcu_head *rhp, + rcu_callback_t func) { unsigned long flags; =20 @@ -233,6 +239,46 @@ void call_srcu(struct srcu_struct *ssp, struct rcu_hea= d *rhp, srcu_gp_start_if_needed(ssp); preempt_enable(); } + +/* Set while srcu_defer_drain() re-issues, to catch a re-entrant call_srcu= (). */ +static bool srcu_defer_draining; + +void srcu_defer_drain(struct irq_work *iw) +{ + struct srcu_struct *ssp =3D container_of(iw, struct srcu_struct, defer_iw= ); + struct llist_node *node, *next; + + /* Callbacks are unordered, so drain in llist order without reversing. */ + srcu_defer_draining =3D true; + llist_for_each_safe(node, next, llist_del_all(&ssp->defer_cbs)) { + struct rcu_head *rhp =3D (struct rcu_head *)node; + + srcu_do_enqueue(ssp, rhp, rhp->func); + } + srcu_defer_draining =3D false; +} +EXPORT_SYMBOL_GPL(srcu_defer_drain); + +void call_srcu(struct srcu_struct *ssp, struct rcu_head *rhp, + rcu_callback_t func) +{ + if (should_rcu_defer()) { + /* A re-entrant call_srcu() during the drain would livelock it. */ + if (srcu_defer_draining && !in_nmi()) { + WARN_ONCE(1, "call_srcu() re-entered during callback drain; leaking cal= lback\n"); + return; + } + rhp->func =3D func; + if (llist_add((struct llist_node *)rhp, &ssp->defer_cbs)) + irq_work_queue(&ssp->defer_iw); + return; + } + + /* An NMI reaching here entered with irqs enabled, so the enqueue can rac= e. */ + WARN_ON_ONCE(IS_ENABLED(CONFIG_PROVE_RCU) && in_nmi()); + + srcu_do_enqueue(ssp, rhp, func); +} EXPORT_SYMBOL_GPL(call_srcu); =20 /* @@ -262,6 +308,15 @@ void synchronize_srcu(struct srcu_struct *ssp) } EXPORT_SYMBOL_GPL(synchronize_srcu); =20 +/* Register any deferred callbacks, then wait for all in-flight ones. */ +void srcu_barrier(struct srcu_struct *ssp) +{ + if (IS_ENABLED(CONFIG_RCU_DEFER)) + irq_work_sync(&ssp->defer_iw); + synchronize_srcu(ssp); +} +EXPORT_SYMBOL_GPL(srcu_barrier); + /* * get_state_synchronize_srcu - Provide an end-of-grace-period cookie */ --=20 2.53.0-Meta From nobody Fri Oct 2 08:25:15 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 511D4432E8D; Mon, 3 Aug 2026 13:54:02 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785765243; cv=none; b=Y5FEOrd29fB/6AOxeyeqT3c+cNLxPFmii+mnjQuUhTSpeie9X28XUxhhoos8QUw0zLkCV//LWe4b3XVNletP+5rhV5R7tyM9Svz8sLKB7wyiutSjR/8WTCDu5ykZj/6gVCNiM6c7CpHViJlQ8PPCTWVgimk640fvkP+r3lk6SPg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785765243; c=relaxed/simple; bh=LwqhIkjuquQWAUBMVGKetFHEeDxPNFlLvPVxDwO4Okk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=ckR7Y9npTsQM4qG3n1veT0DtDRXGl+7eRwng7vG7yiTqYE2mUvTmeoNjd9zCFZCfs7l8E2JEyonOdDG5VCmizy3guOL5xOf23e360wVScrRZWndjozNVJ6y84wAOLM/IHSuZ/WTgdiRq8iXLSC5BvkLdrqlySyNzzM7cccFyPxw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=aBzy2hZz; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="aBzy2hZz" Received: by smtp.kernel.org (Postfix) with ESMTPSA id C758E1F000E9; Mon, 3 Aug 2026 13:54:01 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785765242; bh=Zmv3bjU4WtAFGwErteowPPW79rmJUtxdZMl4u0lwb30=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=aBzy2hZzX3qTzh0sP4PdUT9Z8+IVeR2A+mhEt9nXiztAkugJngSp/F4O3hnqMdIUx 1Y15Pi5htgRcKM849PHk+SUdZegtm2NNPHX7Sr2Yh3oD3KE5GK6bLUqcV4atLU3kvh PpHPbLtTrHGTDpBBacbMueDmsPTrGEi5laZrWty88y4R1WEviqRXwinfp4uvlzYBK2 jO1BMjCTj/Diy/WJmTb0mkvOpnor9c2VBMQjFw4FBdOw5DWBBoFjwCngOr9OBHYuwc iTVEM/0ugph3Fkp6JAAbAN8/XQ3yjtNc1a7dY6o/rxJiHFCE2gTB5UMyNyNh3uVrDy NpqVGgsIc2zzQ== From: Puranjay Mohan To: "Lai Jiangshan" , "Paul E. McKenney" , "Josh Triplett" , =?UTF-8?q?Onur=20=C3=96zkan?= , "Frederic Weisbecker" , "Neeraj Upadhyay" , "Joel Fernandes" , "Boqun Feng" , "Uladzislau Rezki" , "Davidlohr Bueso" , "Andrii Nakryiko" , "Eduard Zingerman" , "Alexei Starovoitov" , "Daniel Borkmann" , "Kumar Kartikeya Dwivedi" Cc: Puranjay Mohan , Steven Rostedt , Mathieu Desnoyers , Zqiang , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Matt Fleming , "Harry Yoo (Oracle)" , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, bpf@vger.kernel.org, linux-rt-devel@lists.linux.dev Subject: [PATCH v2 5/6] rcutorture: Exercise ->call() from NMI context Date: Mon, 3 Aug 2026 06:53:28 -0700 Message-ID: <20260803135329.2327280-5-puranjay@kernel.org> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260803134839.2103051-1-puranjay@kernel.org> References: <20260803134839.2103051-1-puranjay@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" call_rcu() and call_srcu() are now safe to invoke from NMI, but rcutorture never does, leaving the deferral path untested. Add an ->nmi_capable flag to rcu_torture_ops. For flavors that set it, arm a per-CPU hardware perf counter whose overflow handler submits a callback via ->call(). The handler acts only when in_nmi(), so only a genuine NMI exercises the deferral path. One preallocated callback is kept in flight (guarded by an atomic) to avoid allocating in NMI. Report the count issued from NMI ("nmi-calls:") and the count invoked ("nmi-cbs:"). rcu_torture_cleanup() disables the counters and then calls cb_barrier(), which drains every deferred callback, so the two counts must then match; a mismatch means a lost callback and fails the test. This relies on srcu_barrier()/rcu_barrier() flushing deferred callbacks, as added earlier in the series. Set ->nmi_capable on the NMI-safe flavors: rcu, srcu, srcud, and tasks-tracing (call_srcu() under the hood). Tasks and Tasks Rude are left alone, as call_rcu_tasks_generic() is not yet NMI-safe. Enabled by default; the nmi_calls parameter disables it, which helps rule NMI handling in or out when triaging a failure. Requires CONFIG_PERF_EVENTS and a hardware PMU, else silently skipped. Signed-off-by: Puranjay Mohan --- kernel/rcu/rcutorture.c | 112 ++++++++++++++++++++++++++++++++++++++++ 1 file changed, 112 insertions(+) diff --git a/kernel/rcu/rcutorture.c b/kernel/rcu/rcutorture.c index 39426a8718fe9..ae3bbe34809e1 100644 --- a/kernel/rcu/rcutorture.c +++ b/kernel/rcu/rcutorture.c @@ -48,6 +48,7 @@ #include #include #include +#include =20 #include "rcu.h" =20 @@ -115,6 +116,7 @@ torture_param(int, leakpointer, 0, "Leak pointer derefe= rences from readers"); torture_param(int, n_barrier_cbs, 0, "# of callbacks/kthreads for barrier = testing"); torture_param(int, n_up_down, 32, "# of concurrent up/down hrtimer-based R= CU readers"); torture_param(int, nfakewriters, 4, "Number of RCU fake writer threads"); +torture_param(bool, nmi_calls, true, "Exercise ->call() from NMI on nmi_ca= pable flavors"); torture_param(int, nreaders, -1, "Number of RCU reader threads"); torture_param(bool, nwriters, 1, "Number of RCU writer threads (0 or 1)"); torture_param(int, object_debug, 0, "Enable debug-object double call_rcu()= testing"); @@ -216,6 +218,8 @@ static long n_rcu_torture_boost_failure; static long n_rcu_torture_boosts; static atomic_long_t n_rcu_torture_timers; static atomic_long_t n_rcu_torture_irqs; +static atomic_long_t n_rcu_torture_nmi_call; +static atomic_long_t n_rcu_torture_nmi_cb; static long n_barrier_attempts; static long n_barrier_successes; /* did rcu_barrier test succeed? */ static unsigned long n_read_exits; @@ -433,6 +437,7 @@ struct rcu_torture_ops { bool (*is_task_rcu_boosted)(void); long cbflood_max; int irq_capable; + int nmi_capable; int can_boost; int extendables; int slow_gps; @@ -648,6 +653,7 @@ static struct rcu_torture_ops rcu_ops =3D { .extendables =3D RCUTORTURE_MAX_EXTEND, .debug_objects =3D 1, .start_poll_irqsoff =3D 1, + .nmi_capable =3D 1, .name =3D "rcu" }; =20 @@ -942,6 +948,7 @@ static struct rcu_torture_ops srcu_ops =3D { .debug_objects =3D 1, .have_up_down =3D IS_ENABLED(CONFIG_TINY_SRCU) ? 0 : SRCU_READ_FLAVOR_NORMAL | SRCU_READ_FLAVOR_FAST_UPDOWN, + .nmi_capable =3D 1, .name =3D "srcu" }; =20 @@ -1005,6 +1012,7 @@ static struct rcu_torture_ops srcud_ops =3D { .debug_objects =3D 1, .have_up_down =3D IS_ENABLED(CONFIG_TINY_SRCU) ? 0 : SRCU_READ_FLAVOR_NORMAL | SRCU_READ_FLAVOR_FAST_UPDOWN, + .nmi_capable =3D 1, .name =3D "srcud" }; =20 @@ -1269,6 +1277,7 @@ static struct rcu_torture_ops tasks_tracing_ops =3D { .cbflood_max =3D 50000, .irq_capable =3D 1, .slow_gps =3D 1, + .nmi_capable =3D 1, .name =3D "tasks-tracing" }; =20 @@ -2659,12 +2668,94 @@ static bool rcu_torture_one_read(struct torture_ran= dom_state *trsp, long myid) =20 static DEFINE_TORTURE_RANDOM_PERCPU(rcu_torture_timer_rand); =20 +/* + * Exercise ->call() from NMI context for flavors that set ->nmi_capable. = A + * per-CPU hardware perf counter overflows into an NMI, and its handler su= bmits + * one preallocated callback via ->call(). One callback is in flight at a= time + * (guarded by an atomic) to avoid allocating in NMI. This mirrors how a = BPF + * program reaches ->call() from NMI. + */ +#ifdef CONFIG_PERF_EVENTS +static struct perf_event_attr rcu_torture_nmi_attr =3D { + .type =3D PERF_TYPE_HARDWARE, + .config =3D PERF_COUNT_HW_CPU_CYCLES, + .size =3D sizeof(struct perf_event_attr), + .pinned =3D 1, + .disabled =3D 1, + .freq =3D 1, + .sample_freq =3D 1000, +}; + +static struct perf_event **rcu_torture_nmi_events; +static struct rcu_head rcu_torture_nmi_rh; +static atomic_t rcu_torture_nmi_rh_inuse; + +static void rcu_torture_nmi_cb(struct rcu_head *rhp) +{ + atomic_long_inc(&n_rcu_torture_nmi_cb); + atomic_set(&rcu_torture_nmi_rh_inuse, 0); +} + +static void rcu_torture_nmi_overflow(struct perf_event *event, + struct perf_sample_data *data, + struct pt_regs *regs) +{ + if (!in_nmi()) + return; + if (cur_ops->call && !atomic_xchg(&rcu_torture_nmi_rh_inuse, 1)) { + cur_ops->call(&rcu_torture_nmi_rh, rcu_torture_nmi_cb); + atomic_long_inc(&n_rcu_torture_nmi_call); + } +} + +static void rcu_torture_nmi_init(void) +{ + struct perf_event *event; + int cpu; + + if (!nmi_calls || !cur_ops->nmi_capable || !cur_ops->call) + return; + rcu_torture_nmi_events =3D kcalloc(nr_cpu_ids, sizeof(*rcu_torture_nmi_ev= ents), + GFP_KERNEL); + if (!rcu_torture_nmi_events) + return; + for_each_online_cpu(cpu) { + event =3D perf_event_create_kernel_counter(&rcu_torture_nmi_attr, cpu, + NULL, rcu_torture_nmi_overflow, NULL); + if (IS_ERR(event)) + continue; + rcu_torture_nmi_events[cpu] =3D event; + perf_event_enable(event); + } +} + +static void rcu_torture_nmi_cleanup(void) +{ + int cpu; + + if (!rcu_torture_nmi_events) + return; + for_each_possible_cpu(cpu) { + if (!rcu_torture_nmi_events[cpu]) + continue; + perf_event_disable(rcu_torture_nmi_events[cpu]); + perf_event_release_kernel(rcu_torture_nmi_events[cpu]); + } + kfree(rcu_torture_nmi_events); + rcu_torture_nmi_events =3D NULL; +} +#else /* #ifdef CONFIG_PERF_EVENTS */ +static void rcu_torture_nmi_init(void) { } +static void rcu_torture_nmi_cleanup(void) { } +#endif /* #else #ifdef CONFIG_PERF_EVENTS */ + /* * RCU torture reader from timer handler. Dereferences rcu_torture_curren= t, * incrementing the corresponding element of the pipeline array. The * counter in the element should never be greater than 1, otherwise, the * RCU implementation is broken. */ + static void rcu_torture_timer(struct timer_list *unused) { WARN_ON_ONCE(!in_serving_softirq()); @@ -3047,6 +3138,9 @@ rcu_torture_stats_print(void) data_race(n_barrier_attempts), data_race(n_rcu_torture_barrier_error)); pr_cont("read-exits: %ld ", data_race(n_read_exits)); // Statistic. + pr_cont("nmi-calls: %ld nmi-cbs: %ld ", + atomic_long_read(&n_rcu_torture_nmi_call), + atomic_long_read(&n_rcu_torture_nmi_cb)); pr_cont("nocb-toggles: %ld:%ld ", atomic_long_read(&n_nocb_offload), atomic_long_read(&n_nocb_deoffload)); pr_cont("gpwraps: %ld\n", n_gpwraps); @@ -4325,6 +4419,8 @@ rcu_torture_cleanup(void) kfree(reader_tasks); reader_tasks =3D NULL; } + /* Disable the perf counters (and thus the NMI ->call() firing) now. */ + rcu_torture_nmi_cleanup(); kfree(rcu_torture_reader_mbchk); rcu_torture_reader_mbchk =3D NULL; =20 @@ -4354,6 +4450,20 @@ rcu_torture_cleanup(void) pr_info("%s: Invoking %pS().\n", __func__, cur_ops->cb_barrier); cur_ops->cb_barrier(); } + + /* + * cb_barrier() above drained every deferred NMI ->call() callback, so the + * count issued from NMI must equal the count invoked; a mismatch means a + * callback was lost. + */ + if (atomic_long_read(&n_rcu_torture_nmi_call) !=3D + atomic_long_read(&n_rcu_torture_nmi_cb)) { + pr_alert("%s: NMI ->call() lost a callback: issued %ld invoked %ld\n", + __func__, atomic_long_read(&n_rcu_torture_nmi_call), + atomic_long_read(&n_rcu_torture_nmi_cb)); + atomic_inc(&n_rcu_torture_error); + } + if (cur_ops->cleanup !=3D NULL) cur_ops->cleanup(); =20 @@ -4786,6 +4896,8 @@ rcu_torture_init(void) firsterr =3D -ENOMEM; goto unwind; } + /* Arm the per-CPU perf counters that drive ->call() from NMI. */ + rcu_torture_nmi_init(); for (i =3D 0; i < nrealreaders; i++) { rcu_torture_reader_mbchk[i].rtc_chkrdr =3D -1; firsterr =3D torture_create_kthread(rcu_torture_reader, (void *)i, --=20 2.53.0-Meta From nobody Fri Oct 2 08:25:15 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8BD0841D11E; Mon, 3 Aug 2026 13:54:05 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785765246; cv=none; b=SYF6ztfDvpoJ8ghGHGuP5AQVXcOd7xQ6RUySaRc/nOOFXYnkfNulSnFVMUGmQBwTDd630Wn2CfKNGiL7BX+YNSBpfpmw/Lk51pmw3qqLsK3HszIpxaznynwP6pBpCg924a95uAvWrTl2MODpF0NKxqc30+S2TGuVjCcvWbdGW04= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785765246; c=relaxed/simple; bh=wp013QQZyotJrLjBrIasqIfJoCCnycwD9Rfdk91JbKg=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=stw0bKfRkbdBwhLE7bpO1NVgSNC9I0vnaZyElfKLrWziqpUxiTu9D4laSWuCPP+GGoA6jg1QHL8ttA49lYzVv8flBvZVYBddX/b/mDDaNPEzxdq29KmCrByoirvKIuvn4Ctd0e9gSl8AWrUvtaiQ2pUVHbuq0fzcwjpt68AkPVM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=HxocRJ7A; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="HxocRJ7A" Received: by smtp.kernel.org (Postfix) with ESMTPSA id EC5741F000E9; Mon, 3 Aug 2026 13:54:04 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1785765245; bh=qeYRh45l20SmIGanD5yoc+u5gkMxwHBGK5QfYK5yQdU=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=HxocRJ7AMEVpGOfIuBEa4Y61N0xsjHcLE1kmTy6GW40/nz4fBMYxYrDyeVUMhEfxr JDDWidJBVHH2zB3nCRWx8JazDCj19bW044m/cqjKTRZV1MhR6j/al9vLwBNhetyJfh 0BMcUIDDs2tIN8zbvCBy3+ZWPROM+VAcHZcli0enwHIuSTfRAaCWS5zh0iPAQXQzxF uom5jAWj56WhnWG8t584YuNCswpT13AskTNlrS32UhAUVeZhBTo7uskvkIJUUV27Af fWlnyQWMgwzr6fFnguNI3UV8tZxxjz4g6SjJRTuw9pZ0/nAKRci94yGvwmkHXuzx4+ p3phlcimf6Jbg== From: Puranjay Mohan To: "Lai Jiangshan" , "Paul E. McKenney" , "Josh Triplett" , =?UTF-8?q?Onur=20=C3=96zkan?= , "Frederic Weisbecker" , "Neeraj Upadhyay" , "Joel Fernandes" , "Boqun Feng" , "Uladzislau Rezki" , "Davidlohr Bueso" , "Andrii Nakryiko" , "Eduard Zingerman" , "Alexei Starovoitov" , "Daniel Borkmann" , "Kumar Kartikeya Dwivedi" Cc: Puranjay Mohan , Steven Rostedt , Mathieu Desnoyers , Zqiang , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Matt Fleming , "Harry Yoo (Oracle)" , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, bpf@vger.kernel.org, linux-rt-devel@lists.linux.dev Subject: [PATCH v2 6/6] selftests/bpf: Add a call_srcu() re-entry reproducer Date: Mon, 3 Aug 2026 06:53:29 -0700 Message-ID: <20260803135329.2327280-6-puranjay@kernel.org> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260803134839.2103051-1-puranjay@kernel.org> References: <20260803134839.2103051-1-puranjay@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Re-enter call_srcu() from a BPF program to exercise its any-context safety (and thus call_rcu_tasks_trace(), which is call_srcu() on rcu_tasks_trace_srcu_struct). An fentry program on rcu_segcblist_enqueue() fires mid-enqueue: that function is reached from srcu_gp_start_if_needed() with the srcu_data ->lock held. The program does a task-storage delete, whose only deferred work is call_rcu_tasks_trace(), re-entering the enqueue on the same CPU. The triggering thread is pinned to one CPU and matched by TID, so the program fires only for the test's own delete. Without the fix the nested call re-takes the same sdp lock and self-deadlocks; with it the nested __call_srcu() sees interrupts disabled and defers via irq_work, so the delete returns and the test passes. Since it can hang an unfixed kernel, run it only against a kernel carrying the fix. Signed-off-by: Puranjay Mohan Acked-by: Kumar Kartikeya Dwivedi --- .../selftests/bpf/prog_tests/rcu_reentry.c | 58 +++++++++++++++++++ .../testing/selftests/bpf/progs/rcu_reentry.c | 45 ++++++++++++++ 2 files changed, 103 insertions(+) create mode 100644 tools/testing/selftests/bpf/prog_tests/rcu_reentry.c create mode 100644 tools/testing/selftests/bpf/progs/rcu_reentry.c diff --git a/tools/testing/selftests/bpf/prog_tests/rcu_reentry.c b/tools/t= esting/selftests/bpf/prog_tests/rcu_reentry.c new file mode 100644 index 0000000000000..f6ecd93be30f4 --- /dev/null +++ b/tools/testing/selftests/bpf/prog_tests/rcu_reentry.c @@ -0,0 +1,58 @@ +// SPDX-License-Identifier: GPL-2.0 +/* Exercise re-entry into call_srcu() from BPF; see progs/rcu_reentry.c. */ +#define _GNU_SOURCE +#include +#include +#include +#include "rcu_reentry.skel.h" + +static int sys_pidfd_open(pid_t pid, unsigned int flags) +{ + return syscall(__NR_pidfd_open, pid, flags); +} + +void test_rcu_reentry(void) +{ + struct rcu_reentry *skel; + int err, pidfd =3D -1, map_fd; + __u64 val =3D 1; + cpu_set_t set; + + skel =3D rcu_reentry__open_and_load(); + if (!ASSERT_OK_PTR(skel, "skel_open_and_load")) + return; + + err =3D rcu_reentry__attach(skel); + if (!ASSERT_OK(err, "skel_attach")) + goto out; + + /* Keep the re-entry on a single CPU. */ + CPU_ZERO(&set); + CPU_SET(0, &set); + if (sched_setaffinity(0, sizeof(set), &set)) + perror("sched_setaffinity"); + + pidfd =3D sys_pidfd_open(getpid(), 0); + if (!ASSERT_GE(pidfd, 0, "pidfd_open")) + goto out; + map_fd =3D bpf_map__fd(skel->maps.task_stg); + err =3D bpf_map_update_elem(map_fd, &pidfd, &val, BPF_NOEXIST); + if (!ASSERT_OK(err, "boot_create")) + goto out; + + /* Arm the handler for this thread, then trigger call_rcu_tasks_trace(). = */ + skel->bss->target_pid =3D syscall(__NR_gettid); + err =3D bpf_map_delete_elem(map_fd, &pidfd); + ASSERT_OK(err, "boot_delete"); + + /* Only Tree SRCU enqueues via rcu_segcblist_enqueue(); skip elsewhere. */ + if (!skel->bss->hits) { + test__skip(); + goto out; + } + ASSERT_EQ(skel->bss->reentered, 1, "reentry_deferred"); +out: + if (pidfd >=3D 0) + close(pidfd); + rcu_reentry__destroy(skel); +} diff --git a/tools/testing/selftests/bpf/progs/rcu_reentry.c b/tools/testin= g/selftests/bpf/progs/rcu_reentry.c new file mode 100644 index 0000000000000..d92a927ff51c0 --- /dev/null +++ b/tools/testing/selftests/bpf/progs/rcu_reentry.c @@ -0,0 +1,45 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Re-enter call_srcu() from a BPF program. fentry on rcu_segcblist_enque= ue() + * fires inside call_srcu()'s enqueue (reached from srcu_gp_start_if_neede= d() + * with the srcu_data ->lock held); the handler then calls call_rcu_tasks_= trace() + * -- itself call_srcu() on rcu_tasks_trace_srcu_struct -- re-entering the= same + * srcu_data on the same CPU. + */ +#include "vmlinux.h" +#include +#include + +char _license[] SEC("license") =3D "GPL"; + +struct { + __uint(type, BPF_MAP_TYPE_TASK_STORAGE); + __uint(map_flags, BPF_F_NO_PREALLOC); + __type(key, int); + __type(value, __u64); +} task_stg SEC(".maps"); + +int target_pid; +int hits; +int reentered; + +SEC("fentry/rcu_segcblist_enqueue") +int BPF_PROG(reenter) +{ + struct task_struct *cur; + + if (reentered || !target_pid) + return 0; + + cur =3D bpf_get_current_task_btf(); + if (!cur || cur->pid !=3D target_pid) + return 0; + + /* Re-enter via a task-storage delete, which calls call_rcu_tasks_trace()= . */ + __sync_fetch_and_add(&hits, 1); + bpf_task_storage_get(&task_stg, cur, 0, BPF_LOCAL_STORAGE_GET_F_CREATE); + bpf_task_storage_delete(&task_stg, cur); + + reentered =3D 1; + return 0; +} --=20 2.53.0-Meta