From nobody Tue Sep 29 08:25:35 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5043B3D75A1; Mon, 10 Aug 2026 12:28:17 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786364898; cv=none; b=TIz/rh8KM37wvCe8SGO+zk5uJg7ulxp2kh00LGr0hibLmMUkoP5pwqWk7ZdU+6cAoa1CFchrIJWihRpLwl8rrBZcsZ+W6WX6Mb9qWmEmhsscPJWWRSyxXsLCODo3FjpNfy/nXIBXMoHaPL9SlfRIQrmQFP+XYHwFVB8uV3HUcv4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786364898; c=relaxed/simple; bh=T9XDYwpQqs/O8dKn05DDy/qIwIvaH2o+0z7ppRdGyl4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=eS8plx6rE79sm+pfJYAcdByFYyq+VSqN+rVaKU1xcpEWHIa8xs+pwwAoyw519sfuo+ey9y5Qs0SaFxltDp2MF40PpPaxOC+PYKSUqt8ZAf08RErJzLFmMgvW+KhASqhHoLtCtSv/SjXdNzVcJTXIZj/WC0XlvRndEomtcvWsf04= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=OoPP1ipJ; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="OoPP1ipJ" Received: by smtp.kernel.org (Postfix) with ESMTPSA id AF0F51F000E9; Mon, 10 Aug 2026 12:28:16 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786364896; bh=PpX8WmDOIXfXSuF7Dzly1j9xo/oluC1UGMrpWDF2+X8=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=OoPP1ipJZmbTLPkzwBWD1VvFXq6RiD3t/HbiGlRrZsARph8FVY0nA7TXSKdgeiv6J Liugltr6CFCTBNmX/cVY19iMc/wfray+EbCTMJLY/e9cjOXg70x8hRYKwn9v71RG8j oHZQRp2pcyed6kd0K1JGFvs1F5rX0MKrjZWYO0unRaxlog8tqOK08RSwzIpYtaX/hs NNdYrLWP3gwoFHOkwCTX3K4qTI7vWe7lEWF6rdUGmdhmC9coqMyJrKGkyL6FAvNFSv HdKAzbvwXaRis+aq33wxSSPAibL35GOoAeV66RoipwTEmljhNvLbToP7dpAvfZwZB+ 7Fb5fKRWoMpHw== From: Puranjay Mohan To: "Lai Jiangshan" , "Paul E. McKenney" , "Josh Triplett" , =?UTF-8?q?Onur=20=C3=96zkan?= , "Frederic Weisbecker" , "Neeraj Upadhyay" , "Joel Fernandes" , "Boqun Feng" , "Uladzislau Rezki" , "Davidlohr Bueso" , "Andrii Nakryiko" , "Eduard Zingerman" , "Alexei Starovoitov" , "Daniel Borkmann" , "Kumar Kartikeya Dwivedi" Cc: Puranjay Mohan , Steven Rostedt , Mathieu Desnoyers , Zqiang , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Matt Fleming , "Harry Yoo (Oracle)" , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, bpf@vger.kernel.org, linux-rt-devel@lists.linux.dev Subject: [PATCH v4 1/6] rcu: Make call_rcu() safe to call from any context Date: Mon, 10 Aug 2026 05:27:50 -0700 Message-ID: <20260810122758.183765-2-puranjay@kernel.org> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260810122758.183765-1-puranjay@kernel.org> References: <20260810122758.183765-1-puranjay@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" RCU's per-CPU callback list is only touched with interrupts disabled: the enqueue runs under local_irq_save() (and the nocb locks when offloaded), as do callback invocation and grace-period work. A call_rcu() that arrives with interrupts already disabled, whether from an NMI or from instrumentation that re-enters RCU, can interrupt one of those and corrupt the list or deadlock. Defer instead: stage the callback on a per-CPU llist and raise an irq_work that re-issues it once interrupts are on, straight to the enqueue so it cannot defer again. The gate is bare irqs_disabled(), so callers that merely hold interrupts off are deferred too and pay one irq_work hop. Skip it while the scheduler is down (RCU_SCHEDULER_INACTIVE): irq_work is not usable that early, rcu_init() already calls call_rcu(), and the per-CPU deferral state is not initialised until rcu_init_one() runs later in it. rcu_barrier() drains every CPU's ->defer_head before it scans the lists, and rcutree_migrate_callbacks() drains an outgoing CPU's. A drain re-issues onto the draining CPU, so a barrier moves other CPUs' staged callbacks onto its own ->cblist; call_rcu() promises no CPU affinity for invocation. ->defer_lock is held across llist_del_all() and the whole re-issue so the drainers serialize: one that finds the list empty can conclude that everything staged before it is already on a callback list. Interrupts stay off for the batch. Where the arch has an irq_work self-IPI that is what one interrupts-disabled region could stage, normally a single callback; where arch_irq_work_has_interrupt() is false the drain waits for the tick, so several regions can accumulate first. The drain clears ->next before re-issuing. A double call_rcu() on a head that is already debug-object-active self-links the staged node, and rcu_do_enqueue()'s duplicate path returns without clearing it, so the drain would spin. A re-add behind other staged callbacks makes a longer cycle, which that does not bound; a double call_rcu() stays undefined. llist_del_all() yields newest-first, so a batch is re-issued in reverse call order; nothing depends on call_rcu() ordering. The re-issue drops the lazy hint, since staging records only ->func, so a deferred callback loses its batching on CONFIG_RCU_LAZY. kasan_record_aux_stack() moves to __call_rcu_common() so a use-after-free report names the caller rather than the irq_work. The re-issue runs with interrupts disabled, so instrumentation on the enqueue path can re-enter call_rcu(), stage another callback and re-raise the irq_work, livelocking the drain. A per-CPU flag guards it: a deferral that arrives while this CPU is draining, and is not from an NMI, is dropped. The WARN_ONCE() is under CONFIG_PROVE_RCU, so a production kernel drops it silently. That leaks the callback and can strand state the caller tied to it, since a one-shot flag only the callback clears never resets, but the alternative is an unbounded loop. A callback deferred past the CPUHP_AP_SMPCFD_DYING irq_work flush leaves ->defer_work claimed with its self-IPI lost. rcutree_migrate_callbacks() still re-issues the callback, but the first deferral after that CPU comes back raises no IPI and waits for the next irq_work there, or for rcu_barrier(). Unqueueing an irq_work is not something the API offers. The irq_work is IRQ_WORK_INIT_HARD so the re-issue stays prompt on PREEMPT_RT, where a non-HARD irq_work runs in a kthread that can be delayed under load. A hidden CONFIG_RCU_DEFER gates the deferral code and its IRQ_WORK dependency, though the rcu_data members are unconditional; without it call_rcu() enqueues directly as before. Under CONFIG_PROVE_RCU, warn if the direct path is reached from an NMI. Suggested-by: Paul E. McKenney Signed-off-by: Puranjay Mohan --- kernel/rcu/Kconfig | 6 +++ kernel/rcu/rcu.h | 11 ++++ kernel/rcu/tree.c | 131 +++++++++++++++++++++++++++++++++++++++++---- kernel/rcu/tree.h | 6 +++ 4 files changed, 143 insertions(+), 11 deletions(-) diff --git a/kernel/rcu/Kconfig b/kernel/rcu/Kconfig index f15da8038d0ba..1a5fb3156c062 100644 --- a/kernel/rcu/Kconfig +++ b/kernel/rcu/Kconfig @@ -175,6 +175,12 @@ config RCU_STALL_COMMON config RCU_NEED_SEGCBLIST def_bool ( TREE_RCU || TREE_SRCU || TASKS_RCU_GENERIC ) =20 +# The deferral (and the IRQ_WORK it uses) is only needed where call_rcu() / +# call_srcu() can be invoked while a callback-list operation is in flight. +config RCU_DEFER + def_bool HAVE_NMI || KPROBES || FUNCTION_TRACER || TRACEPOINTS + select IRQ_WORK + config RCU_FANOUT int "Tree-based hierarchical RCU fanout value" range 2 64 if 64BIT diff --git a/kernel/rcu/rcu.h b/kernel/rcu/rcu.h index 39a9f6fa9a7b2..91e33571a554d 100644 --- a/kernel/rcu/rcu.h +++ b/kernel/rcu/rcu.h @@ -572,6 +572,17 @@ static inline void tasks_cblist_init_generic(void) { } #define RCU_SCHEDULER_INIT 1 #define RCU_SCHEDULER_RUNNING 2 =20 +/* + * Defer whenever interrupts are disabled, since a callback-list operation= may + * be in flight on this CPU. Not before the scheduler is up: irq_work is = not + * usable that early, and rcu_init() itself calls call_rcu(). + */ +static inline bool should_rcu_defer(void) +{ + return IS_ENABLED(CONFIG_RCU_DEFER) && irqs_disabled() && + rcu_scheduler_active !=3D RCU_SCHEDULER_INACTIVE; +} + enum rcutorture_type { RCU_FLAVOR, RCU_TASKS_FLAVOR, diff --git a/kernel/rcu/tree.c b/kernel/rcu/tree.c index 21b6ce1dffb63..3bf3a250f9de8 100644 --- a/kernel/rcu/tree.c +++ b/kernel/rcu/tree.c @@ -24,6 +24,7 @@ #include #include #include +#include #include #include #include @@ -3148,21 +3149,19 @@ static void check_cb_ovld(struct rcu_data *rdp) raw_spin_unlock_rcu_node(rnp); } =20 -static void -__call_rcu_common(struct rcu_head *head, rcu_callback_t func, bool lazy_in) +/* + * Also called by __rcu_defer_drain() to re-issue a deferred callback, so = it + * must not re-check the deferral condition. Either caller may have inter= rupts + * already disabled, and a drain of a remote CPU re-issues onto the draini= ng + * CPU. + */ +static void rcu_do_enqueue(struct rcu_head *head, rcu_callback_t func, boo= l lazy_in) { static atomic_t doublefrees; unsigned long flags; bool lazy; struct rcu_data *rdp; =20 - /* Misaligned rcu_head! */ - WARN_ON_ONCE((unsigned long)head & (sizeof(void *) - 1)); - - /* Avoid NULL dereference if callback is NULL. */ - if (WARN_ON_ONCE(!func)) - return; - if (debug_rcu_head_queue(head)) { /* * Probable double call_rcu(), so leak the callback. @@ -3178,7 +3177,6 @@ __call_rcu_common(struct rcu_head *head, rcu_callback= _t func, bool lazy_in) } head->func =3D func; head->next =3D NULL; - kasan_record_aux_stack(head); =20 local_irq_save(flags); rdp =3D this_cpu_ptr(&rcu_data); @@ -3206,6 +3204,103 @@ __call_rcu_common(struct rcu_head *head, rcu_callba= ck_t func, bool lazy_in) local_irq_restore(flags); } =20 +/* + * Re-issue deferred callbacks straight to the enqueue so they cannot defer + * again. ->defer_lock serializes the drainers: this CPU's irq_work, + * rcu_defer_flush() and rcutree_migrate_callbacks(). + */ +static void __rcu_defer_drain(struct rcu_data *rdp) +{ + struct llist_node *node, *next; + unsigned long flags; + + if (!IS_ENABLED(CONFIG_RCU_DEFER)) + return; + + raw_spin_lock_irqsave(&rdp->defer_lock, flags); + llist_for_each_safe(node, next, llist_del_all(&rdp->defer_head)) { + struct rcu_head *head =3D (struct rcu_head *)node; + + /* Bounds a node self-linked by a double call_rcu(). */ + head->next =3D NULL; + rcu_do_enqueue(head, head->func, false); + } + raw_spin_unlock_irqrestore(&rdp->defer_lock, flags); +} + +/* + * Only the irq_work drain can be re-fed by its own re-issue, so only it s= ets + * ->defer_draining. Anything staged during a direct drain is picked up b= y the + * staging CPU's own irq_work. Every caller of irq_work_run_list() has + * interrupts disabled, so the flag is never visible with them enabled. + */ +static void rcu_defer_drain(struct irq_work *iw) +{ + struct rcu_data *rdp =3D container_of(iw, struct rcu_data, defer_work); + + WRITE_ONCE(rdp->defer_draining, true); + __rcu_defer_drain(rdp); + WRITE_ONCE(rdp->defer_draining, false); +} + +/* + * Stage @head for this CPU's irq_work to re-issue once interrupts are on.= Only + * the drain side takes a lock, so this stays safe from NMI. + */ +static void call_rcu_defer(struct rcu_head *head, rcu_callback_t func) +{ + struct rcu_data *rdp =3D this_cpu_ptr(&rcu_data); + + /* + * Instrumentation on the enqueue path can re-enter here from inside the + * drain. Re-queuing would livelock it, so drop the callback; an NMI + * cannot loop, so let it through. + */ + if (READ_ONCE(rdp->defer_draining) && !in_nmi()) { + WARN_ONCE(IS_ENABLED(CONFIG_PROVE_RCU), + "call_rcu() re-entered during callback drain; leaking callback\n"); + return; + } + head->func =3D func; + if (llist_add((struct llist_node *)head, &rdp->defer_head)) + irq_work_queue(&rdp->defer_work); +} + +static void rcu_defer_flush(void) +{ + int cpu; + + for_each_possible_cpu(cpu) + __rcu_defer_drain(per_cpu_ptr(&rcu_data, cpu)); +} + +static void +__call_rcu_common(struct rcu_head *head, rcu_callback_t func, bool lazy_in) +{ + /* Misaligned rcu_head! */ + WARN_ON_ONCE((unsigned long)head & (sizeof(void *) - 1)); + + /* Avoid NULL dereference if callback is NULL. */ + if (WARN_ON_ONCE(!func)) + return; + + /* Record the caller: the irq_work's stack says nothing about it. */ + kasan_record_aux_stack(head); + + if (should_rcu_defer()) { + call_rcu_defer(head, func); + return; + } + + /* + * Only reachable from an NMI when deferral is off: before the scheduler + * is up, or with CONFIG_RCU_DEFER=3Dn. The enqueue can then race. + */ + WARN_ON_ONCE(IS_ENABLED(CONFIG_PROVE_RCU) && in_nmi()); + + rcu_do_enqueue(head, func, lazy_in); +} + #ifdef CONFIG_RCU_LAZY static bool enable_rcu_lazy __read_mostly =3D !IS_ENABLED(CONFIG_RCU_LAZY_= DEFAULT_OFF); module_param(enable_rcu_lazy, bool, 0444); @@ -3896,8 +3991,12 @@ void rcu_barrier(void) unsigned long flags; unsigned long gseq; struct rcu_data *rdp; - unsigned long s =3D rcu_seq_snap(&rcu_state.barrier_sequence); + unsigned long s; =20 + /* Register any deferred callbacks before snapshotting the sequence. */ + rcu_defer_flush(); + + s =3D rcu_seq_snap(&rcu_state.barrier_sequence); rcu_barrier_trace(TPS("Begin"), -1, s); =20 /* Take mutex to serialize concurrent rcu_barrier() requests. */ @@ -4231,6 +4330,9 @@ rcu_boot_init_percpu_data(int cpu) rdp->rcu_onl_gp_state =3D RCU_GP_CLEANED; rdp->last_sched_clock =3D jiffies; rdp->cpu =3D cpu; + init_llist_head(&rdp->defer_head); + raw_spin_lock_init(&rdp->defer_lock); + rdp->defer_work =3D IRQ_WORK_INIT_HARD(rcu_defer_drain); rcu_boot_init_nocb_percpu_data(rdp); } =20 @@ -4528,6 +4630,13 @@ void rcutree_migrate_callbacks(int cpu) struct rcu_data *rdp =3D per_cpu_ptr(&rcu_data, cpu); bool needwake; =20 + /* + * Callbacks deferred past the point the outgoing CPU's irq_work can run + * sit on ->defer_head, which the ->cblist migration below does not + * cover. Drain them here, before the early returns. + */ + __rcu_defer_drain(rdp); + if (rcu_rdp_is_offloaded(rdp)) return; =20 diff --git a/kernel/rcu/tree.h b/kernel/rcu/tree.h index eedfa43059e80..b7cac7a13b4f2 100644 --- a/kernel/rcu/tree.h +++ b/kernel/rcu/tree.h @@ -229,6 +229,12 @@ struct rcu_data { struct rcu_head barrier_head; int exp_watching_snap; /* Double-check need for IPI. */ =20 + /* Deferral of an NMI/reentrant call_rcu(); see __call_rcu_common(). */ + struct llist_head defer_head; + struct irq_work defer_work; + raw_spinlock_t defer_lock; + bool defer_draining; + /* 5) Callback offloading. */ #ifdef CONFIG_RCU_NOCB_CPU struct swait_queue_head nocb_cb_wq; /* For nocb kthreads to sleep on. */ --=20 2.53.0-Meta From nobody Tue Sep 29 08:25:35 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EBE7D3D566E; Mon, 10 Aug 2026 12:28:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786364902; cv=none; b=MGv1fzODT+Sr9C6VnSYA84/Q3uZn42qKiRZn+9SBcTDfJ5RFwj9i1LUYgso1ttnG/gl4AnGQU9tS+iiDgeeZOM/nB7UZq0MeFMqWpFPPcxACHvdSlXizbfc9NfCiEiknf9c8fQLPrvkwhHb5xREmijIiQtK3funSiHf4FPkKuOw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786364902; c=relaxed/simple; bh=IDZo4ZqB0SNp2RdCYkNGm42x/R3vAXot0IbrrluCG8M=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=PHXdqZlcaqc4OQbhEF/MKybUt7tTAQWMqg+sSHavkm9RjohKJfOxODnzfK7sbm+0qq+XI9AjoJ1vbBY+a3vfTbnUOn6Eo0poCiSKVrjzSesHkZH5xdC545UkyTjlluNJ+ApRwMUWBMrenDfRmAht747ygguUd0M+eO9HtVlG5Ao= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=RrUmJWHz; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="RrUmJWHz" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 5CA7F1F000E9; Mon, 10 Aug 2026 12:28:20 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786364900; bh=zRx2cUHJLwtDrbQpC2tDv9j1PG1ZjnTzGstPAq7DiUw=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=RrUmJWHzkGPFPWSlQuZnSypETXpQF31sBXz9FzLi7H/Ij3DvAybwYeU3iNJg5RNUg NGfQy4XoeF7TWps404hjbBx+bssP0DXnL2ySj4s+LRkvfH58AAhk0NO/gSEbSiBeAM 9pa6gTMtRN44XCJ0qHgICUBYCFk0B84MWZ2eIxZN9M9eKQr7kFBAOZ86DOV6mXz9yv o/9OdWv4ji/gWtz3xHtvy/JxCzjkXdkIfJ7hEbFRInk7S6saPZz95aL823zvJ9ze65 NCd3HYSH0K6U5fo1I5zVbdKyemTKkLUIEuBmWUut3bwXa2zE522apjSmvScBBhTP4r TymQp9K0XkB7g== From: Puranjay Mohan To: "Lai Jiangshan" , "Paul E. McKenney" , "Josh Triplett" , =?UTF-8?q?Onur=20=C3=96zkan?= , "Frederic Weisbecker" , "Neeraj Upadhyay" , "Joel Fernandes" , "Boqun Feng" , "Uladzislau Rezki" , "Davidlohr Bueso" , "Andrii Nakryiko" , "Eduard Zingerman" , "Alexei Starovoitov" , "Daniel Borkmann" , "Kumar Kartikeya Dwivedi" Cc: Puranjay Mohan , Steven Rostedt , Mathieu Desnoyers , Zqiang , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Matt Fleming , "Harry Yoo (Oracle)" , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, bpf@vger.kernel.org, linux-rt-devel@lists.linux.dev Subject: [PATCH v4 2/6] rcu: Make Tiny call_rcu() safe to call from any context Date: Mon, 10 Aug 2026 05:27:51 -0700 Message-ID: <20260810122758.183765-3-puranjay@kernel.org> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260810122758.183765-1-puranjay@kernel.org> References: <20260810122758.183765-1-puranjay@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Give Tiny call_rcu() the same treatment as Tree RCU. When interrupts are disabled and the scheduler is up, stage the callback on a lockless list that an irq_work re-issues later. One global list and irq_work suffice since Tiny RCU is uniprocessor, and there is no CPU-offline drain. The re-issue runs with interrupts disabled and can be re-entered by instrumentation, so a draining flag drops a deferring call_rcu() seen mid-drain (unless from an NMI), as in Tree RCU. Gated by CONFIG_RCU_DEFER, though the deferral state is unconditional. Interrupts stay off for the whole batch, but the re-issue is a tail append with no locks. TINY_RCU implies !SMP, where arch_irq_work_has_interrupt() is false, so the drain always waits for the tick and a batch is whatever one tick's worth of interrupts-disabled call_rcu()s staged. As in Tree RCU the drain clears ->next before re-issuing, which bounds a node self-linked by a double call_rcu(): rcu_do_enqueue()'s duplicate path returns without clearing it. A longer cycle is not bounded; a double call_rcu() stays undefined. The idle-task reschedule moves out of the enqueue helper so that a drain does it once for the batch rather than once per callback, which would otherwise take the runqueue lock N times with interrupts disabled. Suggested-by: Paul E. McKenney Signed-off-by: Puranjay Mohan --- kernel/rcu/tiny.c | 127 +++++++++++++++++++++++++++++++++++++--------- 1 file changed, 104 insertions(+), 23 deletions(-) diff --git a/kernel/rcu/tiny.c b/kernel/rcu/tiny.c index dccccd6be9411..656b6a682e31a 100644 --- a/kernel/rcu/tiny.c +++ b/kernel/rcu/tiny.c @@ -11,6 +11,8 @@ */ #include #include +#include +#include #include #include #include @@ -42,8 +44,100 @@ static struct rcu_ctrlblk rcu_ctrlblk =3D { .gp_seq =3D 0 - 300UL, }; =20 +/* + * The callback list is only accessed with interrupts disabled, so a call_= rcu() + * that arrives with interrupts off stages the callback on a lockless list= that + * an irq_work re-issues later. One global list and irq_work suffice, as = Tiny + * RCU is uniprocessor. + */ +static void rcu_defer_drain(struct irq_work *iw); +static LLIST_HEAD(rcu_defer_list); +static struct irq_work rcu_defer_iw =3D IRQ_WORK_INIT_HARD(rcu_defer_drain= ); +static bool rcu_defer_draining; + +/* + * Also called by __rcu_defer_drain() to re-issue a deferred callback, so = it + * must not re-check the deferral condition. + */ +static void rcu_do_enqueue(struct rcu_head *head, rcu_callback_t func) +{ + static atomic_t doublefrees; + unsigned long flags; + + if (debug_rcu_head_queue(head)) { + if (atomic_inc_return(&doublefrees) < 4) { + pr_err("%s(): Double-freed CB %p->%pS()!!! ", __func__, head, head->fu= nc); + mem_dump_obj(head); + } + return; + } + + head->func =3D func; + head->next =3D NULL; + + local_irq_save(flags); + *rcu_ctrlblk.curtail =3D head; + rcu_ctrlblk.curtail =3D &head->next; + local_irq_restore(flags); +} + +/* Force scheduling for rcu_qs() when enqueuing from the idle task. */ +static void rcu_resched_if_idle(void) +{ + if (unlikely(is_idle_task(current))) + resched_cpu(0); +} + +static void __rcu_defer_drain(void) +{ + struct llist_node *node, *next; + bool drained =3D false; + unsigned long flags; + + if (!IS_ENABLED(CONFIG_RCU_DEFER)) + return; + + /* Re-issued newest-first; nothing depends on call_rcu() ordering. */ + local_irq_save(flags); + llist_for_each_safe(node, next, llist_del_all(&rcu_defer_list)) { + struct rcu_head *head =3D (struct rcu_head *)node; + + /* Bounds a node self-linked by a double call_rcu(). */ + head->next =3D NULL; + rcu_do_enqueue(head, head->func); + drained =3D true; + } + local_irq_restore(flags); + + if (drained) + rcu_resched_if_idle(); +} + +/* Only the irq_work drain can be re-fed by its own re-issue; see Tree RCU= . */ +static void rcu_defer_drain(struct irq_work *iw) +{ + WRITE_ONCE(rcu_defer_draining, true); + __rcu_defer_drain(); + WRITE_ONCE(rcu_defer_draining, false); +} + +static void call_rcu_defer(struct rcu_head *head, rcu_callback_t func) +{ + /* A re-entrant call_rcu() during the drain would livelock it; drop it. */ + if (READ_ONCE(rcu_defer_draining) && !in_nmi()) { + WARN_ONCE(IS_ENABLED(CONFIG_PROVE_RCU), + "call_rcu() re-entered during callback drain; leaking callback\n"); + return; + } + head->func =3D func; + if (llist_add((struct llist_node *)head, &rcu_defer_list)) + irq_work_queue(&rcu_defer_iw); +} + void rcu_barrier(void) { + /* Register any deferred callbacks so the wait below covers them. */ + __rcu_defer_drain(); wait_rcu_gp(call_rcu_hurry); } EXPORT_SYMBOL(rcu_barrier); @@ -157,29 +251,19 @@ EXPORT_SYMBOL_GPL(synchronize_rcu); */ void call_rcu(struct rcu_head *head, rcu_callback_t func) { - static atomic_t doublefrees; - unsigned long flags; - - if (debug_rcu_head_queue(head)) { - if (atomic_inc_return(&doublefrees) < 4) { - pr_err("%s(): Double-freed CB %p->%pS()!!! ", __func__, head, head->fu= nc); - mem_dump_obj(head); - } + if (should_rcu_defer()) { + call_rcu_defer(head, func); return; } =20 - head->func =3D func; - head->next =3D NULL; + /* + * Only reachable from an NMI when deferral is off: before the scheduler + * is up, or with CONFIG_RCU_DEFER=3Dn. The enqueue can then race. + */ + WARN_ON_ONCE(IS_ENABLED(CONFIG_PROVE_RCU) && in_nmi()); =20 - local_irq_save(flags); - *rcu_ctrlblk.curtail =3D head; - rcu_ctrlblk.curtail =3D &head->next; - local_irq_restore(flags); - - if (unlikely(is_idle_task(current))) { - /* force scheduling for rcu_qs() */ - resched_cpu(0); - } + rcu_do_enqueue(head, func); + rcu_resched_if_idle(); } EXPORT_SYMBOL_GPL(call_rcu); =20 @@ -211,10 +295,7 @@ unsigned long start_poll_synchronize_rcu(void) { unsigned long gp_seq =3D get_state_synchronize_rcu(); =20 - if (unlikely(is_idle_task(current))) { - /* force scheduling for rcu_qs() */ - resched_cpu(0); - } + rcu_resched_if_idle(); return gp_seq; } EXPORT_SYMBOL_GPL(start_poll_synchronize_rcu); --=20 2.53.0-Meta From nobody Tue Sep 29 08:25:35 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C24973C1D54; Mon, 10 Aug 2026 12:28:26 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786364908; cv=none; b=U9VOctxRMHXQCDWEIdKmeRyRfbQBOsygeUenaAK85vOp0lIGdsUsLlj/JhyFA8KSViNpgT4w6CZnW2XwL+kfTGNcmEoZdL05OeVww++l9DjRYp/8GNQMcwXkiVRXMXpfDvRqCa9+KqGcFDanTjhdOtRiHKvzDgJ5kjj7pIqb8q0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786364908; c=relaxed/simple; bh=8IZkTjWxB/mCH6zxRPg2wXo/r2AOxyyVU2JLi8n4lpI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=mPSf5wgpVSMS/GAAtN5mnwsgYpV12HFQA48nSCIMaN4GHCekwpXUS7OA9+LCkgiBg5ztx6qhr0fd+oM/9M9fmzGtE50qSevyboUG9WGC8d7C+5AUhMOVuEhZhbDmI6wfeSQ5vJCGUGrqvVVWqREG99ojDjJIx5SaUP7JoEEuvRU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=Af1lWU6b; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="Af1lWU6b" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 1C8AC1F000E9; Mon, 10 Aug 2026 12:28:26 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786364906; bh=4q2JdJfCqnB/x9SGqg8GRyZ5t0vijKgS1SaGHiWaOyg=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=Af1lWU6bwIPMp64uVrYCQvo7DK4a91++7zNa9ptL78PZzXt3c7e0zqnFwm1MZsYrv ZasTQ5RZoOhGcn1dJ4yewohhCzPUcLyxRzFW8gysdA6jmIwSi4u9O10mT/dFedqaTE 0ut55EWURwqQgdSTj+GT5I/fjDxAZP+Ckoadj9N6rRqns8mvoiNVmwswpJmpfajCkh 3mGhj6jy59ZFp7N4Dp2UPBAUAetUG6fuE9Q9zZ4oj0knXa4lCk74LoNFsPJUt3m9p9 uJYV6mNqUL/x+UYT0+G2OnGC4bkbljutjJxJSJrLUnqZujq0N28T21kOE+nBR6okWU zRC8UCWmPxCRA== From: Puranjay Mohan To: "Lai Jiangshan" , "Paul E. McKenney" , "Josh Triplett" , =?UTF-8?q?Onur=20=C3=96zkan?= , "Frederic Weisbecker" , "Neeraj Upadhyay" , "Joel Fernandes" , "Boqun Feng" , "Uladzislau Rezki" , "Davidlohr Bueso" , "Andrii Nakryiko" , "Eduard Zingerman" , "Alexei Starovoitov" , "Daniel Borkmann" , "Kumar Kartikeya Dwivedi" Cc: Puranjay Mohan , Steven Rostedt , Mathieu Desnoyers , Zqiang , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Matt Fleming , "Harry Yoo (Oracle)" , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, bpf@vger.kernel.org, linux-rt-devel@lists.linux.dev Subject: [PATCH v4 3/6] srcu: Make call_srcu() safe to call from any context Date: Mon, 10 Aug 2026 05:27:52 -0700 Message-ID: <20260810122758.183765-4-puranjay@kernel.org> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260810122758.183765-1-puranjay@kernel.org> References: <20260810122758.183765-1-puranjay@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" call_srcu() has the same constraint as call_rcu(): srcu_gp_start_if_needed() enqueues under raw_spin_lock_irqsave() and may walk the srcu_node tree, as do callback invocation and grace-period work, so a call_srcu() with interrupts already disabled can race an operation in flight on this CPU. call_rcu_tasks_trace() is call_srcu() under the hood, so a sleepable BPF program freeing an object can reach this. Defer as call_rcu() does: stage the callback on the srcu_data's ->defer_cbs, chain that srcu_data onto a per-CPU list, and raise a per-CPU irq_work that re-issues it straight to the enqueue helper, never back through __call_srcu(). The irqs-enabled path is unchanged; as for call_rcu() the gate is bare irqs_disabled(), so callers that merely hold interrupts off are deferred too and pay one irq_work hop, including call_rcu_tasks_trace() from the BPF memalloc irq_work. The irq_work is per-CPU rather than per-srcu_struct and statically initialized, so deferral never runs check_init_srcu_struct(); it is IRQ_WORK_INIT_HARD as for call_rcu(). srcu_barrier() flushes it first, and rcutree_migrate_callbacks() calls srcu_offline_drain() for an outgoing CPU. cleanup_srcu_struct() drains before its "just leak it" early returns, and srcu_module_going() before freeing any ->sda, since a staged srcu_data left chained on a per-CPU list would dangle. Staging is two steps, the callback onto ->defer_cbs and then the srcu_data onto the per-CPU list, so a flusher can find the per-CPU list empty while a callback whose call_srcu() has not returned sits on ->defer_cbs; the staging CPU's own irq_work takes that one. The per-CPU srcu_defer ->lock, not any srcu_data's, is held with interrupts off across the whole nested drain: the chain of srcu_datas staged on that CPU and, for each, its callbacks, with srcu_do_enqueue() taking that srcu_data's ->lock and possibly starting a grace period for every one. That is what serializes the drainers. The bound is as for call_rcu(): what one interrupts-disabled region could stage, normally a single callback, or whatever accumulates before the tick where arch_irq_work_has_interrupt() is false. Both lists are drained newest-first; nothing depends on call_srcu() ordering. The drain clears ->next before re-issuing, which bounds a node self-linked by a double call_srcu(); a longer cycle is not bounded, and a double call_srcu() stays undefined, as for call_rcu(). A callback deferred past the CPUHP_AP_SMPCFD_DYING irq_work flush leaves that CPU's srcu_defer ->iw claimed with its self-IPI lost, as for call_rcu(). srcu_offline_drain() still re-issues the callback, but the irq_work cannot be un-queued, and here the claim is shared by every srcu_struct on the CPU. As in call_rcu(), the re-issue runs with interrupts disabled and can be re-entered by instrumentation, so a per-CPU flag, set only while that CPU is inside its own irq_work drain, drops a deferring call_srcu() seen mid-drain unless it comes from an NMI. Such a drop can strand state the caller associated with the callback, not just the callback itself. Staging records only the callback, so a deferred expedited call_srcu() completes as a normal grace period. Only srcu_expedite_current() can hit that, and only when invoked with interrupts already disabled. Gated by CONFIG_RCU_DEFER, though the srcu_data members and the per-CPU srcu_defer are unconditional. Under CONFIG_PROVE_RCU, warn if the direct path is reached from an NMI. Suggested-by: Paul E. McKenney Signed-off-by: Puranjay Mohan --- include/linux/srcutree.h | 4 + kernel/rcu/rcu.h | 3 + kernel/rcu/srcutree.c | 170 ++++++++++++++++++++++++++++++++++++++- kernel/rcu/tree.c | 2 + 4 files changed, 175 insertions(+), 4 deletions(-) diff --git a/include/linux/srcutree.h b/include/linux/srcutree.h index 75e54e4f963fa..1ce759fb70948 100644 --- a/include/linux/srcutree.h +++ b/include/linux/srcutree.h @@ -13,6 +13,8 @@ =20 #include #include +#include +#include =20 struct srcu_node; struct srcu_struct; @@ -41,6 +43,8 @@ struct srcu_data { bool srcu_cblist_invoking; /* Invoking these CBs? */ struct timer_list delay_work; /* Delay for CB invoking */ struct work_struct work; /* Context for CB invoking. */ + struct llist_head defer_cbs; /* Callbacks deferred on re-entry. */ + struct llist_node defer_link; /* Links onto the per-CPU deferral drain l= ist */ struct rcu_head srcu_barrier_head; /* For srcu_barrier() use. */ struct rcu_head srcu_ec_head; /* For srcu_expedite_current() use. */ int srcu_ec_state; /* State for srcu_expedite_current(). */ diff --git a/kernel/rcu/rcu.h b/kernel/rcu/rcu.h index 91e33571a554d..d60444bf3a02d 100644 --- a/kernel/rcu/rcu.h +++ b/kernel/rcu/rcu.h @@ -583,6 +583,9 @@ static inline bool should_rcu_defer(void) rcu_scheduler_active !=3D RCU_SCHEDULER_INACTIVE; } =20 +/* Drain an outgoing CPU's deferred SRCU callbacks; see rcutree_migrate_ca= llbacks(). */ +void srcu_offline_drain(int cpu); + enum rcutorture_type { RCU_FLAVOR, RCU_TASKS_FLAVOR, diff --git a/kernel/rcu/srcutree.c b/kernel/rcu/srcutree.c index 304112674e8a2..10c89f34920c8 100644 --- a/kernel/rcu/srcutree.c +++ b/kernel/rcu/srcutree.c @@ -20,6 +20,7 @@ #include #include #include +#include #include #include #include @@ -79,6 +80,38 @@ static void process_srcu(struct work_struct *work); static void srcu_irq_work(struct irq_work *work); static void srcu_delay_timer(struct timer_list *t); =20 +struct srcu_defer; +static void srcu_defer_drain(struct irq_work *iw); +static void __srcu_defer_drain(struct srcu_defer *sndp); + +/* + * Per-CPU call_srcu() deferral state, shared by every srcu_struct. A def= erred + * callback is staged on its srcu_data's ->defer_cbs; that srcu_data is ch= ained + * via ->defer_link onto ->list, which the irq_work walks. + */ +struct srcu_defer { + struct llist_head list; + struct irq_work iw; + raw_spinlock_t lock; + bool draining; +}; + +static DEFINE_PER_CPU(struct srcu_defer, srcu_defer) =3D { + .lock =3D __RAW_SPIN_LOCK_UNLOCKED(srcu_defer.lock), + .iw =3D IRQ_WORK_INIT_HARD(srcu_defer_drain), +}; + +/* + * Flush pending deferred callbacks so a following srcu_barrier() waits fo= r them. + */ +static void srcu_defer_flush(void) +{ + int cpu; + + for_each_possible_cpu(cpu) + __srcu_defer_drain(&per_cpu(srcu_defer, cpu)); +} + /* * Initialize SRCU per-CPU data. Note that statically allocated * srcu_struct structures might already have srcu_read_lock() and @@ -107,6 +140,11 @@ static void init_srcu_struct_data(struct srcu_struct *= ssp) sdp->cpu =3D cpu; INIT_WORK(&sdp->work, srcu_invoke_callbacks); timer_setup(&sdp->delay_work, srcu_delay_timer, 0); + /* + * ->defer_cbs and ->defer_link are valid when zeroed and are not + * reinitialized here: that would clobber callbacks a reentrant + * call_srcu() already staged. See __call_srcu(). + */ sdp->ssp =3D ssp; } } @@ -688,6 +726,14 @@ void cleanup_srcu_struct(struct srcu_struct *ssp) unsigned long delay; struct srcu_usage *sup =3D ssp->srcu_sup; =20 + /* + * Drain before the early returns below: they leak the srcu_struct, but + * srcu_module_going() frees ->sda regardless, and a staged srcu_data + * left chained on a per-CPU list would then dangle. Draining first also + * has to precede the ->irq_work sync, since re-issuing a callback can + * start a grace period and re-queue ->irq_work, which schedules ->work. + */ + srcu_defer_flush(); raw_spin_lock_irq_rcu_node(ssp->srcu_sup); delay =3D srcu_get_delay(ssp); raw_spin_unlock_irq_rcu_node(ssp->srcu_sup); @@ -695,7 +741,6 @@ void cleanup_srcu_struct(struct srcu_struct *ssp) return; /* Just leak it! */ if (WARN_ON(srcu_readers_active(ssp))) return; /* Just leak it! */ - /* Wait for irq_work to finish first as it may queue a new work. */ irq_work_sync(&sup->irq_work); flush_delayed_work(&sup->work); for_each_possible_cpu(cpu) { @@ -1410,8 +1455,8 @@ static unsigned long srcu_gp_start_if_needed(struct s= rcu_struct *ssp, * srcu_read_lock(), and srcu_read_unlock() that are all passed the same * srcu_struct structure. */ -static void __call_srcu(struct srcu_struct *ssp, struct rcu_head *rhp, - rcu_callback_t func, bool do_norm) +static void srcu_do_enqueue(struct srcu_struct *ssp, struct rcu_head *rhp, + rcu_callback_t func, bool do_norm) { if (debug_rcu_head_queue(rhp)) { /* Probable double call_srcu(), so leak the callback. */ @@ -1423,6 +1468,108 @@ static void __call_srcu(struct srcu_struct *ssp, st= ruct rcu_head *rhp, (void)srcu_gp_start_if_needed(ssp, rhp, do_norm); } =20 +/* + * The srcu_cblist and srcu_node tree are only accessed with interrupts + * disabled, so defer when interrupts are already off rather than enqueue = into + * an operation that may be in flight on this CPU. + */ +static void __call_srcu(struct srcu_struct *ssp, struct rcu_head *rhp, + rcu_callback_t func, bool do_norm) +{ + if (should_rcu_defer()) { + struct srcu_defer *sndp =3D this_cpu_ptr(&srcu_defer); + struct srcu_data *sdp; + + /* + * Instrumentation on the enqueue path can re-enter here from + * inside the drain. Re-queuing would livelock it, so drop the + * callback; an NMI cannot loop, so let it through. + */ + if (READ_ONCE(sndp->draining) && !in_nmi()) { + WARN_ONCE(IS_ENABLED(CONFIG_PROVE_RCU), + "call_srcu() re-entered during callback drain; leaking callback\n"); + return; + } + sdp =3D this_cpu_ptr(ssp->sda); + rhp->func =3D func; + if (llist_add((struct llist_node *)rhp, &sdp->defer_cbs)) { + /* + * Chain this srcu_data for the drain. ->ssp must be + * published here: deferral skips + * check_init_srcu_struct(), so on a never-initialized + * static srcu_struct the srcu_data are still zeroed and + * the drain would read a NULL ->ssp. + */ + sdp->ssp =3D ssp; + if (llist_add(&sdp->defer_link, &sndp->list)) + irq_work_queue(&sndp->iw); + } + return; + } + + /* + * Only reachable from an NMI when deferral is off: before the scheduler + * is up, or with CONFIG_RCU_DEFER=3Dn. The enqueue can then race. + */ + WARN_ON_ONCE(IS_ENABLED(CONFIG_PROVE_RCU) && in_nmi()); + + srcu_do_enqueue(ssp, rhp, func, do_norm); +} + +/* + * Re-issue deferred callbacks straight to srcu_do_enqueue() so they canno= t defer + * again. ->lock serializes the drainers: the irq_work, srcu_defer_flush(= ) and + * srcu_offline_drain(). + */ +static void __srcu_defer_drain(struct srcu_defer *sndp) +{ + struct llist_node *snode, *snext; + unsigned long flags; + + if (!IS_ENABLED(CONFIG_RCU_DEFER)) + return; + + raw_spin_lock_irqsave(&sndp->lock, flags); + llist_for_each_safe(snode, snext, llist_del_all(&sndp->list)) { + struct srcu_data *sdp =3D container_of(snode, struct srcu_data, defer_li= nk); + struct srcu_struct *ssp =3D sdp->ssp; + struct llist_node *cnode, *cnext; + + cnode =3D llist_del_all(&sdp->defer_cbs); + llist_for_each_safe(cnode, cnext, cnode) { + struct rcu_head *rhp =3D (struct rcu_head *)cnode; + + /* Bounds a node self-linked by a double call_srcu(). */ + rhp->next =3D NULL; + srcu_do_enqueue(ssp, rhp, rhp->func, true); + } + } + raw_spin_unlock_irqrestore(&sndp->lock, flags); +} + +/* + * Only the irq_work drain can be re-fed by its own re-issue, so only it s= ets + * ->draining. A direct drain re-issues onto this CPU, and anything staged + * during it is picked up by that CPU's own irq_work. + */ +static void srcu_defer_drain(struct irq_work *iw) +{ + struct srcu_defer *sndp =3D container_of(iw, struct srcu_defer, iw); + + WRITE_ONCE(sndp->draining, true); + __srcu_defer_drain(sndp); + WRITE_ONCE(sndp->draining, false); +} + +/* + * Drain @cpu's deferred call_srcu() callbacks once @cpu is dead. One pass + * covers every srcu_struct; the re-issue lands on the current CPU. + */ +void srcu_offline_drain(int cpu) +{ + __srcu_defer_drain(&per_cpu(srcu_defer, cpu)); +} + /** * call_srcu() - Queue a callback for invocation after an SRCU grace period * @ssp: srcu_struct in queue the callback @@ -1677,9 +1824,18 @@ void srcu_barrier(struct srcu_struct *ssp) { int cpu; int idx; - unsigned long s =3D rcu_seq_snap(&ssp->srcu_sup->srcu_barrier_seq); + unsigned long s; =20 check_init_srcu_struct(ssp); + + /* + * Register any deferred callbacks before snapshotting the sequence. The + * staging list is per-CPU, not per-srcu_struct, so this also drains + * other srcu_structs'. + */ + srcu_defer_flush(); + + s =3D rcu_seq_snap(&ssp->srcu_sup->srcu_barrier_seq); mutex_lock(&ssp->srcu_sup->srcu_barrier_mutex); if (rcu_seq_done(&ssp->srcu_sup->srcu_barrier_seq, s)) { smp_mb(); /* Force ordering following return. */ @@ -2134,6 +2290,12 @@ static void srcu_module_going(struct module *mod) struct srcu_struct *ssp; struct srcu_struct **sspp =3D mod->srcu_struct_ptrs; =20 + /* + * Deferral skips check_init_srcu_struct(), so cleanup_srcu_struct() + * below can be skipped for an srcu_struct that has staged callbacks. + * Drain them before any ->sda is freed. + */ + srcu_defer_flush(); for (i =3D 0; i < mod->num_srcu_structs; i++) { ssp =3D *(sspp++); if (!rcu_seq_state(smp_load_acquire(&ssp->srcu_sup->srcu_gp_seq_needed))= && diff --git a/kernel/rcu/tree.c b/kernel/rcu/tree.c index 3bf3a250f9de8..7a9b18825e029 100644 --- a/kernel/rcu/tree.c +++ b/kernel/rcu/tree.c @@ -4636,6 +4636,8 @@ void rcutree_migrate_callbacks(int cpu) * cover. Drain them here, before the early returns. */ __rcu_defer_drain(rdp); + /* Likewise for the outgoing CPU's deferred call_srcu() callbacks. */ + srcu_offline_drain(cpu); =20 if (rcu_rdp_is_offloaded(rdp)) return; --=20 2.53.0-Meta From nobody Tue Sep 29 08:25:35 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6E9E73D891F; Mon, 10 Aug 2026 12:28:31 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786364912; cv=none; b=T+2lZ9U1IJcoGUwZKKi6zosutTiqJmyqq4wlfTuEATLFWherUMX5Gli0hsQjUZee5xsL1TI7E6b1+uA1dD+0efHAL91Mp8OEtsB+XNTfucUdAkXUMZ6eNtpKR0cGGmfBSImRdHEHSjTUmrA2DEu/ySB/+8/H7rebQxUzeAjve5g= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786364912; c=relaxed/simple; bh=sWLF8Lu85GKRJJUDw2IixZ09zb8cNHc9glLQxA1l1zo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=jdHAhI0yMbdzPjK8xU4+Y4qOzcHk8SFBLQplAGvlfCDqVuKRmmT0zVDMMHTBPeFKTGG8pNZ+GXYv3OoaAs1vb4SmovpDQG9labyNf0te93CEo2bLYrLqX/9GbuHU5XxHu5w2v2K70DuUfxYA/uxLfsUw6FoeZmcVMP2CXUmJP54= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=NHxtv9kL; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="NHxtv9kL" Received: by smtp.kernel.org (Postfix) with ESMTPSA id D08AC1F000E9; Mon, 10 Aug 2026 12:28:30 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786364911; bh=Rr6LlzJz94VLUxbndZCJiHNS6lr28WG/jTb07f20srE=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=NHxtv9kLjeZDw4kHWkb8jNqa8rqRc4ZITD3HwIiaXA8ePPHbv+CR1PaxSm3YdWrJR Y068NKKDkvjjsECTdmwDr51yFl0wn+hZzbneMT77jsNaUoOY2jxr50qYR3wUG5jXcl soyZk2Uf9YbMa2DpAyAaqN8kQ/8sePcMHGNIvI8tPOiqetqfi4/UorW6HQB5rfSICi XKGNfoONUwbTAivkrp6URvYcrA5Z/h7+mUqPC0r7OzYLsG4xqrtdXGJbF+KqDfaK31 rpeHJlSOWyn3c45rhjUAbRWlucUQBvNpAhf3G6kaijSDWGWceFI9A0cbkHCyGlCRsn NvSWbuDFRajHA== From: Puranjay Mohan To: "Lai Jiangshan" , "Paul E. McKenney" , "Josh Triplett" , =?UTF-8?q?Onur=20=C3=96zkan?= , "Frederic Weisbecker" , "Neeraj Upadhyay" , "Joel Fernandes" , "Boqun Feng" , "Uladzislau Rezki" , "Davidlohr Bueso" , "Andrii Nakryiko" , "Eduard Zingerman" , "Alexei Starovoitov" , "Daniel Borkmann" , "Kumar Kartikeya Dwivedi" Cc: Puranjay Mohan , Steven Rostedt , Mathieu Desnoyers , Zqiang , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Matt Fleming , "Harry Yoo (Oracle)" , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, bpf@vger.kernel.org, linux-rt-devel@lists.linux.dev Subject: [PATCH v4 4/6] srcu: Make Tiny call_srcu() safe to call from any context Date: Mon, 10 Aug 2026 05:27:53 -0700 Message-ID: <20260810122758.183765-5-puranjay@kernel.org> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260810122758.183765-1-puranjay@kernel.org> References: <20260810122758.183765-1-puranjay@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Give Tiny call_srcu() the same treatment as Tree SRCU. When interrupts are disabled and the scheduler is up, stage the callback on the srcu_struct's lockless list for an irq_work to re-issue later. Tiny SRCU is uniprocessor, so there is no CPU-offline drain. A draining flag drops a deferring call_srcu() that re-enters mid-drain (unless from an NMI), as in Tree SRCU; such a drop can strand state the caller tied to the callback, not just the callback itself. Interrupts stay off for the whole batch. TINY_SRCU implies !SMP, where arch_irq_work_has_interrupt() is false, so the drain always waits for the tick and a batch is whatever one tick's worth of interrupts-disabled call_srcu()s staged. Unlike the other three flavors srcu_do_enqueue() here has no debug_rcu_head_queue(), so nothing reports a double call_srcu(); termination of the drain rests on srcu_do_enqueue() clearing ->next, and the callback list self-links at the tail exactly as a double call_srcu() made it before. srcu_barrier() (now out of line) and cleanup_srcu_struct() drain the deferred list first, so a deferred callback is re-issued onto the callback list and invoked by the grace-period work that cleanup_srcu_struct() flushes, rather than stranded on a soon-to-be-freed srcu_struct. cleanup_srcu_struct() also syncs ->defer_iw, since that irq_work is embedded in the srcu_struct the caller is about to free. The draining flag is global rather than per-srcu_struct: a re-entrant call_srcu(B) inside a drain of A raises B's own ->defer_iw, whose drain can stage back onto A, so a per-srcu_struct flag would not break the chain. The cost is that a drain of A also drops a non-NMI call_srcu() to any other srcu_struct for its duration. Gated by CONFIG_RCU_DEFER like Tree SRCU, though the srcu_struct members are unconditional. Suggested-by: Paul E. McKenney Signed-off-by: Puranjay Mohan --- include/linux/srcutiny.h | 12 ++++-- kernel/rcu/srcutiny.c | 93 ++++++++++++++++++++++++++++++++++++++-- 2 files changed, 97 insertions(+), 8 deletions(-) diff --git a/include/linux/srcutiny.h b/include/linux/srcutiny.h index fbcf13bc12d15..85b5de438450b 100644 --- a/include/linux/srcutiny.h +++ b/include/linux/srcutiny.h @@ -12,6 +12,7 @@ #define _LINUX_SRCU_TINY_H =20 #include +#include #include =20 struct srcu_struct { @@ -26,6 +27,8 @@ struct srcu_struct { struct rcu_head **srcu_cb_tail; /* Pending callbacks: Tail. */ struct work_struct srcu_work; /* For driving grace periods. */ struct irq_work srcu_irq_work; /* Defer schedule_work() to irq work. */ + struct llist_head defer_cbs; /* Callbacks deferred on re-entry. */ + struct irq_work defer_iw; /* Re-issues defer_cbs later. */ #ifdef CONFIG_DEBUG_LOCK_ALLOC struct lockdep_map dep_map; #endif /* #ifdef CONFIG_DEBUG_LOCK_ALLOC */ @@ -33,6 +36,7 @@ struct srcu_struct { =20 void srcu_drive_gp(struct work_struct *wp); void srcu_tiny_irq_work(struct irq_work *irq_work); +void srcu_defer_drain(struct irq_work *irq_work); =20 #define __SRCU_STRUCT_INIT(name, __ignored, ___ignored, ____ignored) \ { \ @@ -40,6 +44,9 @@ void srcu_tiny_irq_work(struct irq_work *irq_work); .srcu_cb_tail =3D &name.srcu_cb_head, \ .srcu_work =3D __WORK_INITIALIZER(name.srcu_work, srcu_drive_gp), \ .srcu_irq_work =3D { .func =3D srcu_tiny_irq_work }, \ + .defer_cbs =3D LLIST_HEAD_INIT(name.defer_cbs), \ + .defer_iw =3D { .node =3D { .u_flags =3D IRQ_WORK_HARD_IRQ }, \ + .func =3D srcu_defer_drain }, \ __SRCU_DEP_MAP_INIT(name) \ } =20 @@ -131,10 +138,7 @@ static inline void synchronize_srcu_expedited(struct s= rcu_struct *ssp) synchronize_srcu(ssp); } =20 -static inline void srcu_barrier(struct srcu_struct *ssp) -{ - synchronize_srcu(ssp); -} +void srcu_barrier(struct srcu_struct *ssp); =20 static inline void srcu_expedite_current(struct srcu_struct *ssp) { } #define srcu_check_read_flavor(ssp, read_flavor) do { } while (0) diff --git a/kernel/rcu/srcutiny.c b/kernel/rcu/srcutiny.c index f9c498ae75df2..32b37d63d58aa 100644 --- a/kernel/rcu/srcutiny.c +++ b/kernel/rcu/srcutiny.c @@ -10,6 +10,7 @@ =20 #include #include +#include #include #include #include @@ -29,6 +30,8 @@ extern int rcu_scheduler_active; static LIST_HEAD(srcu_boot_list); static bool srcu_init_done; =20 +static void __srcu_defer_drain(struct srcu_struct *ssp); + static int init_srcu_struct_fields(struct srcu_struct *ssp) { ssp->srcu_lock_nesting[0] =3D 0; @@ -43,6 +46,8 @@ static int init_srcu_struct_fields(struct srcu_struct *ss= p) INIT_WORK(&ssp->srcu_work, srcu_drive_gp); INIT_LIST_HEAD(&ssp->srcu_work.entry); init_irq_work(&ssp->srcu_irq_work, srcu_tiny_irq_work); + init_llist_head(&ssp->defer_cbs); + ssp->defer_iw =3D IRQ_WORK_INIT_HARD(srcu_defer_drain); return 0; } =20 @@ -86,6 +91,16 @@ EXPORT_SYMBOL_GPL(init_srcu_struct_generic); void cleanup_srcu_struct(struct srcu_struct *ssp) { WARN_ON(srcu_readers_active(ssp)); + /* + * Re-issue any deferred callbacks, then wait out ->defer_iw before it is + * freed. Skipped entirely with CONFIG_RCU_DEFER=3Dn: irq_work_sync() en= ds + * in an unconditional synchronize_rcu() wherever + * arch_irq_work_has_interrupt() is false, which is every !SMP target. + */ + if (IS_ENABLED(CONFIG_RCU_DEFER)) { + __srcu_defer_drain(ssp); + irq_work_sync(&ssp->defer_iw); + } irq_work_sync(&ssp->srcu_irq_work); flush_work(&ssp->srcu_work); WARN_ON(ssp->srcu_gp_running); @@ -215,11 +230,11 @@ static void srcu_gp_start_if_needed(struct srcu_struc= t *ssp) } =20 /* - * Enqueue an SRCU callback on the specified srcu_struct structure, - * initiating grace-period processing if it is not already running. + * Also called by __srcu_defer_drain() to re-issue a deferred callback, so= it + * must not re-check the deferral condition. */ -void call_srcu(struct srcu_struct *ssp, struct rcu_head *rhp, - rcu_callback_t func) +static void srcu_do_enqueue(struct srcu_struct *ssp, struct rcu_head *rhp, + rcu_callback_t func) { unsigned long flags; =20 @@ -233,6 +248,68 @@ void call_srcu(struct srcu_struct *ssp, struct rcu_hea= d *rhp, srcu_gp_start_if_needed(ssp); preempt_enable(); } + +/* + * Set only by the irq_work drain, the one drain its own re-issue can re-f= eed; + * a callback staged during a direct drain is taken by ->defer_iw afterwar= ds. + * Global rather than per-srcu_struct: a re-entrant call_srcu(B) inside a = drain + * of A raises B's own ->defer_iw, whose drain can stage back onto A. + */ +static bool srcu_defer_draining; + +static void __srcu_defer_drain(struct srcu_struct *ssp) +{ + struct llist_node *node, *next; + unsigned long flags; + + if (!IS_ENABLED(CONFIG_RCU_DEFER)) + return; + + /* Re-issued newest-first; nothing depends on call_srcu() ordering. */ + local_irq_save(flags); + llist_for_each_safe(node, next, llist_del_all(&ssp->defer_cbs)) { + struct rcu_head *rhp =3D (struct rcu_head *)node; + + srcu_do_enqueue(ssp, rhp, rhp->func); + } + local_irq_restore(flags); +} + +/* Only the irq_work drain can be re-fed by its own re-issue; see Tree SRC= U. */ +void srcu_defer_drain(struct irq_work *iw) +{ + struct srcu_struct *ssp =3D container_of(iw, struct srcu_struct, defer_iw= ); + + WRITE_ONCE(srcu_defer_draining, true); + __srcu_defer_drain(ssp); + WRITE_ONCE(srcu_defer_draining, false); +} +EXPORT_SYMBOL_GPL(srcu_defer_drain); + +void call_srcu(struct srcu_struct *ssp, struct rcu_head *rhp, + rcu_callback_t func) +{ + if (should_rcu_defer()) { + /* A re-entrant call_srcu() during the drain would livelock it. */ + if (READ_ONCE(srcu_defer_draining) && !in_nmi()) { + WARN_ONCE(IS_ENABLED(CONFIG_PROVE_RCU), + "call_srcu() re-entered during callback drain; leaking callback\n"); + return; + } + rhp->func =3D func; + if (llist_add((struct llist_node *)rhp, &ssp->defer_cbs)) + irq_work_queue(&ssp->defer_iw); + return; + } + + /* + * Only reachable from an NMI when deferral is off: before the scheduler + * is up, or with CONFIG_RCU_DEFER=3Dn. The enqueue can then race. + */ + WARN_ON_ONCE(IS_ENABLED(CONFIG_PROVE_RCU) && in_nmi()); + + srcu_do_enqueue(ssp, rhp, func); +} EXPORT_SYMBOL_GPL(call_srcu); =20 /* @@ -262,6 +339,14 @@ void synchronize_srcu(struct srcu_struct *ssp) } EXPORT_SYMBOL_GPL(synchronize_srcu); =20 +/* Register any deferred callbacks, then wait for all in-flight ones. */ +void srcu_barrier(struct srcu_struct *ssp) +{ + __srcu_defer_drain(ssp); + synchronize_srcu(ssp); +} +EXPORT_SYMBOL_GPL(srcu_barrier); + /* * get_state_synchronize_srcu - Provide an end-of-grace-period cookie */ --=20 2.53.0-Meta From nobody Tue Sep 29 08:25:35 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 011F53DA5B2; Mon, 10 Aug 2026 12:28:34 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786364916; cv=none; b=DVReWrwoKINv8EwuSBc0n0djjtmrZ7aScKs+p24AOh4gKO38wmszLt6XkEOYHg2MCafQHABD3JvZAai83/nAjVUj4gUqY7q0XliL2/RQJX0Q5+Cv3uV5jR43GZImtssbG6IxtnbJrK/vRNhZd6a7CYX41hjkVf9FZqSM4q2ePcA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786364916; c=relaxed/simple; bh=f+TKqUJk+3u/FWIYXTSUrkYH1ZB4Fx4QrtB9D036iTw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=BZSRDXtTrroIACp34q6U92GmQnbW+p/icdVeVqZWnvL70FSUgarL8ot8hidxY36Vu8T/nSMaiTQarE5PXu7SUIDxNKmLst666NiCG5BeH3PmR0mwU745ER4kF1HsuP9WQhb1FS1h8LieQOIHZxTXG0ui6t1dGu6PLNkdZyyJpfc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=egfD1DaB; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="egfD1DaB" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 579941F000E9; Mon, 10 Aug 2026 12:28:34 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786364914; bh=hT108n9EELsgMbbDZwstA3an3hbd454cZbSMEBHg/x0=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=egfD1DaB7equ9mQqn1d/DqstY8k/6sSQdylIjBDjtNvhUZbgARgOzREOF1+gRKLsn I4zv/HABlaqWCFNnuo1PpUFw856skACePLe7BeOtosa4YqQdd6vULC3CTfuNZecYqY XMtddtkwq7CKm33ccEV1QzlQQg42/B1QHW4hN2EhWAEbx1cc2oaR9w+cmkoXBhqG0R bY9dd8dpJwMK79o349582sBBBNK2YCvuCfnoSs3aTWh5iEKTOjSu1QAk4u94FNh6l/ QPTpZ/1ol3rmRC+SHmXBow01996UvmAu3LmUgHyh5C99foeJo+1NPBnINm9cBXBt5p kKUqPRT9dvsbQ== From: Puranjay Mohan To: "Lai Jiangshan" , "Paul E. McKenney" , "Josh Triplett" , =?UTF-8?q?Onur=20=C3=96zkan?= , "Frederic Weisbecker" , "Neeraj Upadhyay" , "Joel Fernandes" , "Boqun Feng" , "Uladzislau Rezki" , "Davidlohr Bueso" , "Andrii Nakryiko" , "Eduard Zingerman" , "Alexei Starovoitov" , "Daniel Borkmann" , "Kumar Kartikeya Dwivedi" Cc: Puranjay Mohan , Steven Rostedt , Mathieu Desnoyers , Zqiang , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Matt Fleming , "Harry Yoo (Oracle)" , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, bpf@vger.kernel.org, linux-rt-devel@lists.linux.dev Subject: [PATCH v4 5/6] rcutorture: Exercise ->call() from NMI context Date: Mon, 10 Aug 2026 05:27:54 -0700 Message-ID: <20260810122758.183765-6-puranjay@kernel.org> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260810122758.183765-1-puranjay@kernel.org> References: <20260810122758.183765-1-puranjay@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" call_rcu() and call_srcu() are now safe to invoke from NMI, but rcutorture never does, leaving the deferral path untested. Add an ->nmi_capable flag to rcu_torture_ops. For flavors that set it, arm a per-CPU hardware perf counter whose overflow handler submits a callback via ->call(). The handler acts only when in_nmi(), so only a genuine NMI exercises the deferral path. One preallocated callback per CPU is kept in flight, guarded by an atomic, to avoid allocating in NMI. The counter uses a fixed sample period rather than a frequency: a frequency-based event sets TICK_DEP_BIT_PERF_EVENTS and would pin the tick for the whole run on NO_HZ_FULL kernels. Report the count issued from NMI ("nmi-calls:") and the count invoked ("nmi-cbs:"). rcu_torture_cleanup() disables the counters and then calls cb_barrier(), which drains every deferred callback, so the two counts must then match; a mismatch fails the test. This relies on rcu_barrier()/srcu_barrier() flushing deferred callbacks, as added earlier in the series. Set ->nmi_capable on the NMI-safe flavors: rcu, srcu, srcud, and tasks-tracing (call_srcu() under the hood). Tasks and Tasks Rude are left alone, as call_rcu_tasks_generic() is not yet NMI-safe. Enabled by default; the nmi_calls parameter disables it, which helps rule NMI handling in or out when triaging a failure. Requires CONFIG_PERF_EVENTS and a hardware PMU: without one nothing is issued from NMI and the end-of-test check compares zero against zero, so a pass does not by itself mean the path ran. That is the case under kvm.sh, which boots qemu with -cpu kvm64 and no vPMU; rcu_torture_nmi_cleanup() says so on the console. Signed-off-by: Puranjay Mohan --- .../admin-guide/kernel-parameters.txt | 7 + kernel/rcu/rcutorture.c | 150 +++++++++++++++++- 2 files changed, 155 insertions(+), 2 deletions(-) diff --git a/Documentation/admin-guide/kernel-parameters.txt b/Documentatio= n/admin-guide/kernel-parameters.txt index 6ebc1cd26c7c8..a3d7a4f8bca33 100644 --- a/Documentation/admin-guide/kernel-parameters.txt +++ b/Documentation/admin-guide/kernel-parameters.txt @@ -6198,6 +6198,13 @@ Kernel parameters stress RCU, they don't participate in the actual test, hence the "fake". =20 + rcutorture.nmi_calls=3D [KNL] + Enable issuing RCU callbacks from an NMI, on the + RCU flavors that support it, to exercise the + any-context callback path. Requires + CONFIG_PERF_EVENTS and a hardware PMU; without + both, nothing is issued. Defaults to enabled. + rcutorture.nocbs_nthreads=3D [KNL] Set number of RCU callback-offload togglers. Zero (the default) disables toggling. diff --git a/kernel/rcu/rcutorture.c b/kernel/rcu/rcutorture.c index 39426a8718fe9..0e8b88368bbfc 100644 --- a/kernel/rcu/rcutorture.c +++ b/kernel/rcu/rcutorture.c @@ -48,6 +48,7 @@ #include #include #include +#include =20 #include "rcu.h" =20 @@ -115,6 +116,7 @@ torture_param(int, leakpointer, 0, "Leak pointer derefe= rences from readers"); torture_param(int, n_barrier_cbs, 0, "# of callbacks/kthreads for barrier = testing"); torture_param(int, n_up_down, 32, "# of concurrent up/down hrtimer-based R= CU readers"); torture_param(int, nfakewriters, 4, "Number of RCU fake writer threads"); +torture_param(bool, nmi_calls, true, "Exercise ->call() from NMI on nmi_ca= pable flavors"); torture_param(int, nreaders, -1, "Number of RCU reader threads"); torture_param(bool, nwriters, 1, "Number of RCU writer threads (0 or 1)"); torture_param(int, object_debug, 0, "Enable debug-object double call_rcu()= testing"); @@ -216,6 +218,8 @@ static long n_rcu_torture_boost_failure; static long n_rcu_torture_boosts; static atomic_long_t n_rcu_torture_timers; static atomic_long_t n_rcu_torture_irqs; +static atomic_long_t n_rcu_torture_nmi_call; +static atomic_long_t n_rcu_torture_nmi_cb; static long n_barrier_attempts; static long n_barrier_successes; /* did rcu_barrier test succeed? */ static unsigned long n_read_exits; @@ -433,6 +437,7 @@ struct rcu_torture_ops { bool (*is_task_rcu_boosted)(void); long cbflood_max; int irq_capable; + int nmi_capable; int can_boost; int extendables; int slow_gps; @@ -648,6 +653,7 @@ static struct rcu_torture_ops rcu_ops =3D { .extendables =3D RCUTORTURE_MAX_EXTEND, .debug_objects =3D 1, .start_poll_irqsoff =3D 1, + .nmi_capable =3D 1, .name =3D "rcu" }; =20 @@ -942,6 +948,7 @@ static struct rcu_torture_ops srcu_ops =3D { .debug_objects =3D 1, .have_up_down =3D IS_ENABLED(CONFIG_TINY_SRCU) ? 0 : SRCU_READ_FLAVOR_NORMAL | SRCU_READ_FLAVOR_FAST_UPDOWN, + .nmi_capable =3D 1, .name =3D "srcu" }; =20 @@ -1005,6 +1012,7 @@ static struct rcu_torture_ops srcud_ops =3D { .debug_objects =3D 1, .have_up_down =3D IS_ENABLED(CONFIG_TINY_SRCU) ? 0 : SRCU_READ_FLAVOR_NORMAL | SRCU_READ_FLAVOR_FAST_UPDOWN, + .nmi_capable =3D 1, .name =3D "srcud" }; =20 @@ -1269,6 +1277,7 @@ static struct rcu_torture_ops tasks_tracing_ops =3D { .cbflood_max =3D 50000, .irq_capable =3D 1, .slow_gps =3D 1, + .nmi_capable =3D 1, .name =3D "tasks-tracing" }; =20 @@ -2659,6 +2668,124 @@ static bool rcu_torture_one_read(struct torture_ran= dom_state *trsp, long myid) =20 static DEFINE_TORTURE_RANDOM_PERCPU(rcu_torture_timer_rand); =20 +/* + * Exercise ->call() from NMI context for flavors that set ->nmi_capable. = A + * per-CPU hardware perf counter overflows into an NMI, and its handler su= bmits + * a preallocated callback via ->call(). One callback per CPU is in fligh= t at a + * time (guarded by an atomic) to avoid allocating in NMI. + */ +#ifdef CONFIG_PERF_EVENTS +static struct perf_event_attr rcu_torture_nmi_attr =3D { + .type =3D PERF_TYPE_HARDWARE, + .config =3D PERF_COUNT_HW_CPU_CYCLES, + .size =3D sizeof(struct perf_event_attr), + .pinned =3D 1, + .disabled =3D 1, + /* + * A fixed period rather than .freq: a frequency-based event bumps + * nr_freq_events, which sets TICK_DEP_BIT_PERF_EVENTS and would pin the + * tick for the whole run on NO_HZ_FULL kernels. + */ + .sample_period =3D 20 * 1000 * 1000, +}; + +/* One in-flight callback per CPU; ->inuse is released by the callback. */ +struct rcu_torture_nmi_cb { + struct rcu_head rh; + atomic_t inuse; +}; + +static struct perf_event **rcu_torture_nmi_events; +static int rcu_torture_nmi_hp_state; +static DEFINE_PER_CPU(struct rcu_torture_nmi_cb, rcu_torture_nmi_cb); + +static void rcu_torture_nmi_invoked(struct rcu_head *rhp) +{ + struct rcu_torture_nmi_cb *rtncp =3D container_of(rhp, struct rcu_torture= _nmi_cb, rh); + + atomic_long_inc(&n_rcu_torture_nmi_cb); + atomic_set(&rtncp->inuse, 0); +} + +static void rcu_torture_nmi_overflow(struct perf_event *event, + struct perf_sample_data *data, + struct pt_regs *regs) +{ + struct rcu_torture_nmi_cb *rtncp =3D this_cpu_ptr(&rcu_torture_nmi_cb); + + if (!in_nmi()) + return; + if (cur_ops->call && !atomic_xchg(&rtncp->inuse, 1)) { + atomic_long_inc(&n_rcu_torture_nmi_call); + cur_ops->call(&rtncp->rh, rcu_torture_nmi_invoked); + } +} + +static int rcu_torture_nmi_online(unsigned int cpu) +{ + struct perf_event *event; + + event =3D perf_event_create_kernel_counter(&rcu_torture_nmi_attr, cpu, NU= LL, + rcu_torture_nmi_overflow, NULL); + if (IS_ERR(event)) + return 0; + rcu_torture_nmi_events[cpu] =3D event; + perf_event_enable(event); + return 0; +} + +static int rcu_torture_nmi_offline(unsigned int cpu) +{ + struct perf_event *event =3D rcu_torture_nmi_events[cpu]; + + if (event) { + rcu_torture_nmi_events[cpu] =3D NULL; + perf_event_disable(event); + perf_event_release_kernel(event); + } + return 0; +} + +/* Drive the counters from hotplug callbacks so coverage survives onoff. */ +static void rcu_torture_nmi_init(void) +{ + int ret; + + if (!nmi_calls || !cur_ops->nmi_capable || !cur_ops->call) + return; + rcu_torture_nmi_events =3D kcalloc(nr_cpu_ids, sizeof(*rcu_torture_nmi_ev= ents), + GFP_KERNEL); + if (!rcu_torture_nmi_events) + return; + ret =3D cpuhp_setup_state(CPUHP_AP_ONLINE_DYN, "rcutorture/nmi:online", + rcu_torture_nmi_online, rcu_torture_nmi_offline); + if (ret < 0) { + kfree(rcu_torture_nmi_events); + rcu_torture_nmi_events =3D NULL; + return; + } + rcu_torture_nmi_hp_state =3D ret; +} + +static void rcu_torture_nmi_cleanup(void) +{ + if (!rcu_torture_nmi_events) + return; + if (rcu_torture_nmi_hp_state > 0) { + cpuhp_remove_state(rcu_torture_nmi_hp_state); + rcu_torture_nmi_hp_state =3D 0; + } + kfree(rcu_torture_nmi_events); + rcu_torture_nmi_events =3D NULL; + if (!atomic_long_read(&n_rcu_torture_nmi_call)) + pr_alert("%s: nmi_calls set but no ->call() ever issued from NMI, so NMI= ->call() went untested (no PMU, or NMIs unavailable here).\n", + __func__); +} +#else /* #ifdef CONFIG_PERF_EVENTS */ +static void rcu_torture_nmi_init(void) { } +static void rcu_torture_nmi_cleanup(void) { } +#endif /* #else #ifdef CONFIG_PERF_EVENTS */ + /* * RCU torture reader from timer handler. Dereferences rcu_torture_curren= t, * incrementing the corresponding element of the pipeline array. The @@ -3047,6 +3174,9 @@ rcu_torture_stats_print(void) data_race(n_barrier_attempts), data_race(n_rcu_torture_barrier_error)); pr_cont("read-exits: %ld ", data_race(n_read_exits)); // Statistic. + pr_cont("nmi-calls: %ld nmi-cbs: %ld ", + atomic_long_read(&n_rcu_torture_nmi_call), + atomic_long_read(&n_rcu_torture_nmi_cb)); pr_cont("nocb-toggles: %ld:%ld ", atomic_long_read(&n_nocb_offload), atomic_long_read(&n_nocb_deoffload)); pr_cont("gpwraps: %ld\n", n_gpwraps); @@ -3195,7 +3325,7 @@ rcu_torture_print_module_parms(struct rcu_torture_ops= *cur_ops, const char *tag) "read_exit_delay=3D%d read_exit_burst=3D%d " "reader_flavor=3D%x " "nocbs_nthreads=3D%d nocbs_toggle=3D%d " - "test_nmis=3D%d " + "test_nmis=3D%d nmi_calls=3D%d " "preempt_duration=3D%d preempt_interval=3D%d n_up_down=3D%d\n", torture_type, tag, nrealreaders, nwriters, nrealfakewriters, stat_interval, verbose, test_no_idle_hz, shuffle_interval, @@ -3209,7 +3339,7 @@ rcu_torture_print_module_parms(struct rcu_torture_ops= *cur_ops, const char *tag) read_exit_delay, read_exit_burst, reader_flavor, nocbs_nthreads, nocbs_toggle, - test_nmis, + test_nmis, nmi_calls, preempt_duration, preempt_interval, n_up_down); } =20 @@ -4283,6 +4413,7 @@ rcu_torture_cleanup(void) int i; =20 if (torture_cleanup_begin()) { + rcu_torture_nmi_cleanup(); if (cur_ops->cb_barrier !=3D NULL) { pr_info("%s: Invoking %pS().\n", __func__, cur_ops->cb_barrier); cur_ops->cb_barrier(); @@ -4325,6 +4456,7 @@ rcu_torture_cleanup(void) kfree(reader_tasks); reader_tasks =3D NULL; } + rcu_torture_nmi_cleanup(); kfree(rcu_torture_reader_mbchk); rcu_torture_reader_mbchk =3D NULL; =20 @@ -4354,6 +4486,19 @@ rcu_torture_cleanup(void) pr_info("%s: Invoking %pS().\n", __func__, cur_ops->cb_barrier); cur_ops->cb_barrier(); } + + /* + * cb_barrier() above drained every deferred callback, so the count + * issued from NMI must equal the count invoked. + */ + if (atomic_long_read(&n_rcu_torture_nmi_call) !=3D + atomic_long_read(&n_rcu_torture_nmi_cb)) { + pr_alert("%s: NMI ->call() lost a callback: issued %ld invoked %ld\n", + __func__, atomic_long_read(&n_rcu_torture_nmi_call), + atomic_long_read(&n_rcu_torture_nmi_cb)); + atomic_inc(&n_rcu_torture_error); + } + if (cur_ops->cleanup !=3D NULL) cur_ops->cleanup(); =20 @@ -4786,6 +4931,7 @@ rcu_torture_init(void) firsterr =3D -ENOMEM; goto unwind; } + rcu_torture_nmi_init(); for (i =3D 0; i < nrealreaders; i++) { rcu_torture_reader_mbchk[i].rtc_chkrdr =3D -1; firsterr =3D torture_create_kthread(rcu_torture_reader, (void *)i, --=20 2.53.0-Meta From nobody Tue Sep 29 08:25:35 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 529973D7A18; Mon, 10 Aug 2026 12:28:37 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786364919; cv=none; b=Vj1O569Hm3jIJMS6uai3pEOV9XAykevz+JuRwsz6pOorFU0IWf6ZIhp54Jmead3asrBq3S4pYTti1EudkQRYkut9o64m6Kg6VJPjnobEzbbkHhFFSI0eEUNcrGEl+K7D9SRKeskN9TNNE9sjhFc4wekJjZ0pWc/Isa5dLTxuPAw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786364919; c=relaxed/simple; bh=YBqj1Odqlr5vJhp8U6pwcvWqvXBL4GqH9aoBir/rry8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=dc/9WRD0O5gyiPRr7OLKbZf1zFo9tQgplP1KWhEsx6R928UkTy57tLQnt+5xOPtyQwiKu3OwuAkucPvwvU2twBRmJ3G7QVvqNt43LDvSmlcy4b88zBYIz+F1htz93zRi01uF1/F8cCHRVH5kTrSjtA9dIUIEqFYSiLGeSckOxVs= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=Ew+IznzZ; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="Ew+IznzZ" Received: by smtp.kernel.org (Postfix) with ESMTPSA id A033C1F000E9; Mon, 10 Aug 2026 12:28:37 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786364917; bh=iEwmQpckeIO3ur7ByQuHq0/2DDHm8BL4KuY4DhULyRA=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=Ew+IznzZhDjEDaEVeztJWpzfJ3zFSWUACyqgd831U+jl0xnKiSCoLMuDu0XMKhzRJ JS3kcsk67D5lqb965hqFmJ4ZqhTISJhLy7W1JFBK+6Z7S4H1IrccRFkyITdfuJaUlK PtXGl1rA1SPd2vAtbWDBK2NMn47//3aLZgmIbvqXslHtBSIVyIY2m33UkiVQ0YpMwE /u6F/EQnzzuunqtWIe1uTCwR1vBr0I+hZHd3txboz0rAQeRD4LNROklEYdd2Rhrts+ V0SqM/sAMmgbzDCX6dlcNNCs8bnymhKrGlIui/6bBHb0yUzZgiT/D2sQel9wBmLHuJ Vzo7hhrjpdNAg== From: Puranjay Mohan To: "Lai Jiangshan" , "Paul E. McKenney" , "Josh Triplett" , =?UTF-8?q?Onur=20=C3=96zkan?= , "Frederic Weisbecker" , "Neeraj Upadhyay" , "Joel Fernandes" , "Boqun Feng" , "Uladzislau Rezki" , "Davidlohr Bueso" , "Andrii Nakryiko" , "Eduard Zingerman" , "Alexei Starovoitov" , "Daniel Borkmann" , "Kumar Kartikeya Dwivedi" Cc: Puranjay Mohan , Steven Rostedt , Mathieu Desnoyers , Zqiang , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Matt Fleming , "Harry Yoo (Oracle)" , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, bpf@vger.kernel.org, linux-rt-devel@lists.linux.dev Subject: [PATCH v4 6/6] selftests/bpf: Add a call_srcu() re-entry reproducer Date: Mon, 10 Aug 2026 05:27:55 -0700 Message-ID: <20260810122758.183765-7-puranjay@kernel.org> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260810122758.183765-1-puranjay@kernel.org> References: <20260810122758.183765-1-puranjay@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Re-enter call_srcu() from a BPF program to exercise its any-context safety, via call_rcu_tasks_trace(), which is call_srcu() on rcu_tasks_trace_srcu_struct. An fentry program on rcu_segcblist_enqueue() fires mid-enqueue: that function is reached from srcu_gp_start_if_needed() with the srcu_data ->lock held. The program does a task-storage delete, whose only deferred work is call_rcu_tasks_trace(), re-entering the enqueue on the same CPU. The handler matches on TID and fires once; pinning the thread removes the migration window between picking the srcu_data and taking its lock. Without the fix the nested call re-takes the same sdp lock and self-deadlocks; with it the nested __call_srcu() sees interrupts disabled and defers via irq_work, so the delete returns and the test passes. The test skips where it does not apply: Tiny RCU has no rcu_segcblist_enqueue() to attach to, and a UP+PREEMPT kernel pairs Tree RCU with Tiny SRCU, so the attach succeeds but call_srcu() never reaches the enqueue. Tiny SRCU is told apart by srcu_expedite_current(), which it stubs out, so on Tree SRCU a zero hit count fails rather than skips and the reproducer cannot quietly stop reproducing. Signed-off-by: Puranjay Mohan Acked-by: Kumar Kartikeya Dwivedi --- .../selftests/bpf/prog_tests/rcu_reentry.c | 93 +++++++++++++++++++ .../testing/selftests/bpf/progs/rcu_reentry.c | 51 ++++++++++ 2 files changed, 144 insertions(+) create mode 100644 tools/testing/selftests/bpf/prog_tests/rcu_reentry.c create mode 100644 tools/testing/selftests/bpf/progs/rcu_reentry.c diff --git a/tools/testing/selftests/bpf/prog_tests/rcu_reentry.c b/tools/t= esting/selftests/bpf/prog_tests/rcu_reentry.c new file mode 100644 index 0000000000000..de23a14b3d408 --- /dev/null +++ b/tools/testing/selftests/bpf/prog_tests/rcu_reentry.c @@ -0,0 +1,93 @@ +// SPDX-License-Identifier: GPL-2.0 +/* Exercise re-entry into call_srcu() from BPF; see progs/rcu_reentry.c. */ +#define _GNU_SOURCE +#include +#include +#include "task_local_storage_helpers.h" +#include "trace_helpers.h" +#include "rcu_reentry.skel.h" + +/* Tiny RCU has no rcu_segcblist_enqueue() to attach to. */ +static bool have_attach_target(void) +{ + unsigned long long addr; + + return kallsyms_find("rcu_segcblist_enqueue", &addr) =3D=3D 0; +} + +/* Tiny SRCU stubs out srcu_expedite_current(); Tree SRCU exports it. */ +static bool have_tree_srcu(void) +{ + unsigned long long addr; + + return kallsyms_find("srcu_expedite_current", &addr) =3D=3D 0; +} + +void test_rcu_reentry(void) +{ + struct rcu_reentry *skel; + int err, pidfd =3D -1, map_fd; + cpu_set_t set, old_set; + bool affinity_saved; + __u64 val =3D 1; + int cpu; + + if (!have_attach_target()) { + test__skip(); + return; + } + + skel =3D rcu_reentry__open_and_load(); + if (!ASSERT_OK_PTR(skel, "skel_open_and_load")) + return; + + err =3D rcu_reentry__attach(skel); + if (!ASSERT_OK(err, "skel_attach")) + goto out; + + /* Keep the re-entry on a single CPU; a cpuset may exclude CPU 0. */ + affinity_saved =3D !sched_getaffinity(0, sizeof(old_set), &old_set); + cpu =3D sched_getcpu(); + if (!ASSERT_GE(cpu, 0, "getcpu")) + goto out; + CPU_ZERO(&set); + CPU_SET(cpu, &set); + if (!ASSERT_OK(sched_setaffinity(0, sizeof(set), &set), "setaffinity")) + goto out; + + pidfd =3D sys_pidfd_open(getpid(), 0); + if (!ASSERT_GE(pidfd, 0, "pidfd_open")) + goto restore; + map_fd =3D bpf_map__fd(skel->maps.task_stg); + err =3D bpf_map_update_elem(map_fd, &pidfd, &val, BPF_NOEXIST); + if (!ASSERT_OK(err, "boot_create")) + goto restore; + + /* Arm the handler for this thread, then trigger call_rcu_tasks_trace(). = */ + skel->bss->target_pid =3D syscall(__NR_gettid); + err =3D bpf_map_delete_elem(map_fd, &pidfd); + if (!ASSERT_OK(err, "boot_delete")) + goto restore; + + /* + * Only Tree SRCU reaches rcu_segcblist_enqueue() from call_srcu(); a + * UP+PREEMPT kernel pairs Tree RCU with Tiny SRCU, so the attach + * succeeds but nothing fires. On Tree SRCU it must fire. + */ + if (!skel->bss->hits) { + if (have_tree_srcu()) + ASSERT_GT(skel->bss->hits, 0, "prog_fired"); + else + test__skip(); + goto restore; + } + ASSERT_EQ(skel->bss->get_errs, 0, "nested_storage_get"); + ASSERT_EQ(skel->bss->del_errs, 0, "nested_storage_delete"); +restore: + if (affinity_saved) + sched_setaffinity(0, sizeof(old_set), &old_set); +out: + if (pidfd >=3D 0) + close(pidfd); + rcu_reentry__destroy(skel); +} diff --git a/tools/testing/selftests/bpf/progs/rcu_reentry.c b/tools/testin= g/selftests/bpf/progs/rcu_reentry.c new file mode 100644 index 0000000000000..47a36f704cf3e --- /dev/null +++ b/tools/testing/selftests/bpf/progs/rcu_reentry.c @@ -0,0 +1,51 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Re-enter call_srcu() from a BPF program. fentry on rcu_segcblist_enque= ue() + * fires inside call_srcu()'s enqueue (reached from srcu_gp_start_if_neede= d() + * with the srcu_data ->lock held); the handler then calls call_rcu_tasks_= trace() + * -- itself call_srcu() on rcu_tasks_trace_srcu_struct -- re-entering the= same + * srcu_data on the same CPU. + */ +#include "vmlinux.h" +#include +#include + +char _license[] SEC("license") =3D "GPL"; + +struct { + __uint(type, BPF_MAP_TYPE_TASK_STORAGE); + __uint(map_flags, BPF_F_NO_PREALLOC); + __type(key, int); + __type(value, __u64); +} task_stg SEC(".maps"); + +int target_pid; +int hits; +int get_errs; +int del_errs; +int done; + +SEC("fentry/rcu_segcblist_enqueue") +int BPF_PROG(reenter) +{ + struct task_struct *cur; + + if (done || !target_pid) + return 0; + + cur =3D bpf_get_current_task_btf(); + if (cur->pid !=3D target_pid) + return 0; + + /* Issue the nested call exactly once, so the test is deterministic. */ + done =3D 1; + __sync_fetch_and_add(&hits, 1); + + /* Re-enter via a task-storage delete, which calls call_rcu_tasks_trace()= . */ + if (!bpf_task_storage_get(&task_stg, cur, 0, BPF_LOCAL_STORAGE_GET_F_CREA= TE)) + __sync_fetch_and_add(&get_errs, 1); + else if (bpf_task_storage_delete(&task_stg, cur)) + __sync_fetch_and_add(&del_errs, 1); + + return 0; +} --=20 2.53.0-Meta