mm/vmscan.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-)
I am seeing some rcu_tasks stalls in the Meta fleet during reclaim.
INFO: rcu_tasks detected stalls on tasks:
0000000088620d09: .. nvcsw: 6735/6735 holdout: 1 idle_cpu: -1/8
task:GlobalCPUThread state:R running task pid:2552016 tgid:2524552
Call Trace:
shrink_lruvec
mem_cgroup_iter
shrink_node
do_try_to_free_pages
try_to_free_pages
__alloc_frozen_pages_noprof
alloc_pages_noprof
pte_alloc_one
__pte_alloc
handle_mm_fault
Nothing promises direct reclaim returns in bounded time, and the scan
loop in shrink_lruvec() only calls cond_resched(), which is a no-op on
PREEMPTION kernels. Involuntary preemption is not a Tasks-RCU
quiescent state, so the reclaiming task never reports one and becomes a
holdout.
Upgrade it to cond_resched_tasks_rcu_qs(), which reports a quiescent
state even when cond_resched() does nothing.
PS: This has been discussed in [1]
Link: https://lore.kernel.org/all/amdWVTs0WKOxguxP@gmail.com/ [1]
Cc: stable@vger.kernel.org
Signed-off-by: Breno Leitao <leitao@debian.org>
---
mm/vmscan.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/mm/vmscan.c b/mm/vmscan.c
index 26436059ea394..6ac2fde137b89 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -6023,7 +6023,7 @@ static void shrink_lruvec(struct lruvec *lruvec, struct scan_control *sc)
}
}
- cond_resched();
+ cond_resched_tasks_rcu_qs();
if (nr_reclaimed < nr_to_reclaim || proportional_reclaim)
continue;
---
base-commit: 6b8c8af514d739d0335f5579b585e02babe8a727
change-id: 20260810-rcu_task_shrink_lruvec-711112e87de0
Best regards,
--
Breno Leitao <leitao@debian.org>
On Mon, Aug 10, 2026 at 02:57:36AM -0700, Breno Leitao wrote: > I am seeing some rcu_tasks stalls in the Meta fleet during reclaim. > > INFO: rcu_tasks detected stalls on tasks: > 0000000088620d09: .. nvcsw: 6735/6735 holdout: 1 idle_cpu: -1/8 > task:GlobalCPUThread state:R running task pid:2552016 tgid:2524552 > Call Trace: > shrink_lruvec > mem_cgroup_iter > shrink_node > do_try_to_free_pages > try_to_free_pages > __alloc_frozen_pages_noprof > alloc_pages_noprof > pte_alloc_one > __pte_alloc > handle_mm_fault > > Nothing promises direct reclaim returns in bounded time, and the scan > loop in shrink_lruvec() only calls cond_resched(), which is a no-op on > PREEMPTION kernels. Involuntary preemption is not a Tasks-RCU > quiescent state, so the reclaiming task never reports one and becomes a > holdout. I still don't understand why cond_resched() is being treated as involuntary preemption but that is orthogonal to this patch. > > Upgrade it to cond_resched_tasks_rcu_qs(), which reports a quiescent > state even when cond_resched() does nothing. > > PS: This has been discussed in [1] > > Link: https://lore.kernel.org/all/amdWVTs0WKOxguxP@gmail.com/ [1] > Cc: stable@vger.kernel.org > Signed-off-by: Breno Leitao <leitao@debian.org> Acked-by: Shakeel Butt <shakeel.butt@linux.dev>
On Mon, Aug 10, 2026 at 09:43:03PM -0700, Shakeel Butt wrote: > On Mon, Aug 10, 2026 at 02:57:36AM -0700, Breno Leitao wrote: > > I am seeing some rcu_tasks stalls in the Meta fleet during reclaim. > > > > INFO: rcu_tasks detected stalls on tasks: > > 0000000088620d09: .. nvcsw: 6735/6735 holdout: 1 idle_cpu: -1/8 > > task:GlobalCPUThread state:R running task pid:2552016 tgid:2524552 > > Call Trace: > > shrink_lruvec > > mem_cgroup_iter > > shrink_node > > do_try_to_free_pages > > try_to_free_pages > > __alloc_frozen_pages_noprof > > alloc_pages_noprof > > pte_alloc_one > > __pte_alloc > > handle_mm_fault > > > > Nothing promises direct reclaim returns in bounded time, and the scan > > loop in shrink_lruvec() only calls cond_resched(), which is a no-op on > > PREEMPTION kernels. Involuntary preemption is not a Tasks-RCU > > quiescent state, so the reclaiming task never reports one and becomes a > > holdout. > > I still don't understand why cond_resched() is being treated as involuntary > preemption but that is orthogonal to this patch. The history is that cond_resched() was originally intended to be a preemption point in any otherwise non-preemptible kernel. Therefore, because it is a preemption point, it counts as an involuntary context switch. Thanx, Paul > > Upgrade it to cond_resched_tasks_rcu_qs(), which reports a quiescent > > state even when cond_resched() does nothing. > > > > PS: This has been discussed in [1] > > > > Link: https://lore.kernel.org/all/amdWVTs0WKOxguxP@gmail.com/ [1] > > Cc: stable@vger.kernel.org > > Signed-off-by: Breno Leitao <leitao@debian.org> > > Acked-by: Shakeel Butt <shakeel.butt@linux.dev> >
On Tue, Aug 11, 2026 at 10:38:40AM -0700, Paul E. McKenney wrote: > On Mon, Aug 10, 2026 at 09:43:03PM -0700, Shakeel Butt wrote: > > On Mon, Aug 10, 2026 at 02:57:36AM -0700, Breno Leitao wrote: > > > I am seeing some rcu_tasks stalls in the Meta fleet during reclaim. > > > > > > INFO: rcu_tasks detected stalls on tasks: > > > 0000000088620d09: .. nvcsw: 6735/6735 holdout: 1 idle_cpu: -1/8 > > > task:GlobalCPUThread state:R running task pid:2552016 tgid:2524552 > > > Call Trace: > > > shrink_lruvec > > > mem_cgroup_iter > > > shrink_node > > > do_try_to_free_pages > > > try_to_free_pages > > > __alloc_frozen_pages_noprof > > > alloc_pages_noprof > > > pte_alloc_one > > > __pte_alloc > > > handle_mm_fault > > > > > > Nothing promises direct reclaim returns in bounded time, and the scan > > > loop in shrink_lruvec() only calls cond_resched(), which is a no-op on > > > PREEMPTION kernels. Involuntary preemption is not a Tasks-RCU > > > quiescent state, so the reclaiming task never reports one and becomes a > > > holdout. > > > > I still don't understand why cond_resched() is being treated as involuntary > > preemption but that is orthogonal to this patch. > > The history is that cond_resched() was originally intended to be a > preemption point in any otherwise non-preemptible kernel. Therefore, > because it is a preemption point, it counts as an involuntary context > switch. Thanks for the explanation. Is cond_resched_tasks_rcu_qs() voluntary or involuntary context switch?
On Tue, Aug 11, 2026 at 10:54:23AM -0700, Shakeel Butt wrote: > On Tue, Aug 11, 2026 at 10:38:40AM -0700, Paul E. McKenney wrote: > > On Mon, Aug 10, 2026 at 09:43:03PM -0700, Shakeel Butt wrote: > > > On Mon, Aug 10, 2026 at 02:57:36AM -0700, Breno Leitao wrote: > > > > I am seeing some rcu_tasks stalls in the Meta fleet during reclaim. > > > > > > > > INFO: rcu_tasks detected stalls on tasks: > > > > 0000000088620d09: .. nvcsw: 6735/6735 holdout: 1 idle_cpu: -1/8 > > > > task:GlobalCPUThread state:R running task pid:2552016 tgid:2524552 > > > > Call Trace: > > > > shrink_lruvec > > > > mem_cgroup_iter > > > > shrink_node > > > > do_try_to_free_pages > > > > try_to_free_pages > > > > __alloc_frozen_pages_noprof > > > > alloc_pages_noprof > > > > pte_alloc_one > > > > __pte_alloc > > > > handle_mm_fault > > > > > > > > Nothing promises direct reclaim returns in bounded time, and the scan > > > > loop in shrink_lruvec() only calls cond_resched(), which is a no-op on > > > > PREEMPTION kernels. Involuntary preemption is not a Tasks-RCU > > > > quiescent state, so the reclaiming task never reports one and becomes a > > > > holdout. > > > > > > I still don't understand why cond_resched() is being treated as involuntary > > > preemption but that is orthogonal to this patch. > > > > The history is that cond_resched() was originally intended to be a > > preemption point in any otherwise non-preemptible kernel. Therefore, > > because it is a preemption point, it counts as an involuntary context > > switch. > > Thanks for the explanation. Is cond_resched_tasks_rcu_qs() voluntary or > involuntary context switch? That one is voluntary in order to take care of Tasks RCU's need for a voluntary context switch. Except that it doesn't bother involving the scheduler at all, but instead just updates the current task's state with respect to Tasks RCU. Much faster that way. ;-) Thanx, Paul
On Mon, Aug 10, 2026 at 02:57:36AM -0700, Breno Leitao wrote: > I am seeing some rcu_tasks stalls in the Meta fleet during reclaim. > > INFO: rcu_tasks detected stalls on tasks: > 0000000088620d09: .. nvcsw: 6735/6735 holdout: 1 idle_cpu: -1/8 > task:GlobalCPUThread state:R running task pid:2552016 tgid:2524552 > Call Trace: > shrink_lruvec > mem_cgroup_iter > shrink_node > do_try_to_free_pages > try_to_free_pages > __alloc_frozen_pages_noprof > alloc_pages_noprof > pte_alloc_one > __pte_alloc > handle_mm_fault > > Nothing promises direct reclaim returns in bounded time, and the scan > loop in shrink_lruvec() only calls cond_resched(), which is a no-op on > PREEMPTION kernels. Involuntary preemption is not a Tasks-RCU > quiescent state, so the reclaiming task never reports one and becomes a > holdout. > > Upgrade it to cond_resched_tasks_rcu_qs(), which reports a quiescent > state even when cond_resched() does nothing. > > PS: This has been discussed in [1] > > Link: https://lore.kernel.org/all/amdWVTs0WKOxguxP@gmail.com/ [1] > Cc: stable@vger.kernel.org > Signed-off-by: Breno Leitao <leitao@debian.org> Acked-by: Johannes Weiner <hannes@cmpxchg.org>
On Mon, Aug 10, 2026 at 02:57:36AM -0700, Breno Leitao wrote: > I am seeing some rcu_tasks stalls in the Meta fleet during reclaim. > > INFO: rcu_tasks detected stalls on tasks: > 0000000088620d09: .. nvcsw: 6735/6735 holdout: 1 idle_cpu: -1/8 > task:GlobalCPUThread state:R running task pid:2552016 tgid:2524552 > Call Trace: > shrink_lruvec > mem_cgroup_iter > shrink_node > do_try_to_free_pages > try_to_free_pages > __alloc_frozen_pages_noprof > alloc_pages_noprof > pte_alloc_one > __pte_alloc > handle_mm_fault > > Nothing promises direct reclaim returns in bounded time, and the scan > loop in shrink_lruvec() only calls cond_resched(), which is a no-op on > PREEMPTION kernels. Involuntary preemption is not a Tasks-RCU > quiescent state, so the reclaiming task never reports one and becomes a > holdout. > > Upgrade it to cond_resched_tasks_rcu_qs(), which reports a quiescent > state even when cond_resched() does nothing. > > PS: This has been discussed in [1] > > Link: https://lore.kernel.org/all/amdWVTs0WKOxguxP@gmail.com/ [1] > Cc: stable@vger.kernel.org > Signed-off-by: Breno Leitao <leitao@debian.org> Reviewed-by: Paul E. McKenney <paulmck@kernel.org> > --- > mm/vmscan.c | 2 +- > 1 file changed, 1 insertion(+), 1 deletion(-) > > diff --git a/mm/vmscan.c b/mm/vmscan.c > index 26436059ea394..6ac2fde137b89 100644 > --- a/mm/vmscan.c > +++ b/mm/vmscan.c > @@ -6023,7 +6023,7 @@ static void shrink_lruvec(struct lruvec *lruvec, struct scan_control *sc) > } > } > > - cond_resched(); > + cond_resched_tasks_rcu_qs(); > > if (nr_reclaimed < nr_to_reclaim || proportional_reclaim) > continue; > > --- > base-commit: 6b8c8af514d739d0335f5579b585e02babe8a727 > change-id: 20260810-rcu_task_shrink_lruvec-711112e87de0 > > Best regards, > -- > Breno Leitao <leitao@debian.org> >
© 2016 - 2026 Red Hat, Inc.