[PATCH] mm/vmscan: report RCU-tasks quiescent states in shrink_lruvec()

Breno Leitao posted 1 patch 1 month, 2 weeks ago
mm/vmscan.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
[PATCH] mm/vmscan: report RCU-tasks quiescent states in shrink_lruvec()
Posted by Breno Leitao 1 month, 2 weeks ago
I am seeing some rcu_tasks stalls in the Meta fleet during reclaim.

  INFO: rcu_tasks detected stalls on tasks:
	0000000088620d09: .. nvcsw: 6735/6735 holdout: 1 idle_cpu: -1/8
	task:GlobalCPUThread state:R  running task  pid:2552016 tgid:2524552
  Call Trace:
   shrink_lruvec
   mem_cgroup_iter
   shrink_node
   do_try_to_free_pages
   try_to_free_pages
   __alloc_frozen_pages_noprof
   alloc_pages_noprof
   pte_alloc_one
   __pte_alloc
   handle_mm_fault

Nothing promises direct reclaim returns in bounded time, and the scan
loop in shrink_lruvec() only calls cond_resched(), which is a no-op on
PREEMPTION kernels.  Involuntary preemption is not a Tasks-RCU
quiescent state, so the reclaiming task never reports one and becomes a
holdout.

Upgrade it to cond_resched_tasks_rcu_qs(), which reports a quiescent
state even when cond_resched() does nothing.

PS: This has been discussed in [1]

Link: https://lore.kernel.org/all/amdWVTs0WKOxguxP@gmail.com/ [1]
Cc: stable@vger.kernel.org
Signed-off-by: Breno Leitao <leitao@debian.org>
---
 mm/vmscan.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/mm/vmscan.c b/mm/vmscan.c
index 26436059ea394..6ac2fde137b89 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -6023,7 +6023,7 @@ static void shrink_lruvec(struct lruvec *lruvec, struct scan_control *sc)
 			}
 		}
 
-		cond_resched();
+		cond_resched_tasks_rcu_qs();
 
 		if (nr_reclaimed < nr_to_reclaim || proportional_reclaim)
 			continue;

---
base-commit: 6b8c8af514d739d0335f5579b585e02babe8a727
change-id: 20260810-rcu_task_shrink_lruvec-711112e87de0

Best regards,
--  
Breno Leitao <leitao@debian.org>
Re: [PATCH] mm/vmscan: report RCU-tasks quiescent states in shrink_lruvec()
Posted by Shakeel Butt 1 month, 2 weeks ago
On Mon, Aug 10, 2026 at 02:57:36AM -0700, Breno Leitao wrote:
> I am seeing some rcu_tasks stalls in the Meta fleet during reclaim.
> 
>   INFO: rcu_tasks detected stalls on tasks:
> 	0000000088620d09: .. nvcsw: 6735/6735 holdout: 1 idle_cpu: -1/8
> 	task:GlobalCPUThread state:R  running task  pid:2552016 tgid:2524552
>   Call Trace:
>    shrink_lruvec
>    mem_cgroup_iter
>    shrink_node
>    do_try_to_free_pages
>    try_to_free_pages
>    __alloc_frozen_pages_noprof
>    alloc_pages_noprof
>    pte_alloc_one
>    __pte_alloc
>    handle_mm_fault
> 
> Nothing promises direct reclaim returns in bounded time, and the scan
> loop in shrink_lruvec() only calls cond_resched(), which is a no-op on
> PREEMPTION kernels.  Involuntary preemption is not a Tasks-RCU
> quiescent state, so the reclaiming task never reports one and becomes a
> holdout.

I still don't understand why cond_resched() is being treated as involuntary
preemption but that is orthogonal to this patch.

> 
> Upgrade it to cond_resched_tasks_rcu_qs(), which reports a quiescent
> state even when cond_resched() does nothing.
> 
> PS: This has been discussed in [1]
> 
> Link: https://lore.kernel.org/all/amdWVTs0WKOxguxP@gmail.com/ [1]
> Cc: stable@vger.kernel.org
> Signed-off-by: Breno Leitao <leitao@debian.org>

Acked-by: Shakeel Butt <shakeel.butt@linux.dev>
Re: [PATCH] mm/vmscan: report RCU-tasks quiescent states in shrink_lruvec()
Posted by Paul E. McKenney 1 month, 2 weeks ago
On Mon, Aug 10, 2026 at 09:43:03PM -0700, Shakeel Butt wrote:
> On Mon, Aug 10, 2026 at 02:57:36AM -0700, Breno Leitao wrote:
> > I am seeing some rcu_tasks stalls in the Meta fleet during reclaim.
> > 
> >   INFO: rcu_tasks detected stalls on tasks:
> > 	0000000088620d09: .. nvcsw: 6735/6735 holdout: 1 idle_cpu: -1/8
> > 	task:GlobalCPUThread state:R  running task  pid:2552016 tgid:2524552
> >   Call Trace:
> >    shrink_lruvec
> >    mem_cgroup_iter
> >    shrink_node
> >    do_try_to_free_pages
> >    try_to_free_pages
> >    __alloc_frozen_pages_noprof
> >    alloc_pages_noprof
> >    pte_alloc_one
> >    __pte_alloc
> >    handle_mm_fault
> > 
> > Nothing promises direct reclaim returns in bounded time, and the scan
> > loop in shrink_lruvec() only calls cond_resched(), which is a no-op on
> > PREEMPTION kernels.  Involuntary preemption is not a Tasks-RCU
> > quiescent state, so the reclaiming task never reports one and becomes a
> > holdout.
> 
> I still don't understand why cond_resched() is being treated as involuntary
> preemption but that is orthogonal to this patch.

The history is that cond_resched() was originally intended to be a
preemption point in any otherwise non-preemptible kernel.  Therefore,
because it is a preemption point, it counts as an involuntary context
switch.

							Thanx, Paul

> > Upgrade it to cond_resched_tasks_rcu_qs(), which reports a quiescent
> > state even when cond_resched() does nothing.
> > 
> > PS: This has been discussed in [1]
> > 
> > Link: https://lore.kernel.org/all/amdWVTs0WKOxguxP@gmail.com/ [1]
> > Cc: stable@vger.kernel.org
> > Signed-off-by: Breno Leitao <leitao@debian.org>
> 
> Acked-by: Shakeel Butt <shakeel.butt@linux.dev>
>
Re: [PATCH] mm/vmscan: report RCU-tasks quiescent states in shrink_lruvec()
Posted by Shakeel Butt 1 month, 2 weeks ago
On Tue, Aug 11, 2026 at 10:38:40AM -0700, Paul E. McKenney wrote:
> On Mon, Aug 10, 2026 at 09:43:03PM -0700, Shakeel Butt wrote:
> > On Mon, Aug 10, 2026 at 02:57:36AM -0700, Breno Leitao wrote:
> > > I am seeing some rcu_tasks stalls in the Meta fleet during reclaim.
> > > 
> > >   INFO: rcu_tasks detected stalls on tasks:
> > > 	0000000088620d09: .. nvcsw: 6735/6735 holdout: 1 idle_cpu: -1/8
> > > 	task:GlobalCPUThread state:R  running task  pid:2552016 tgid:2524552
> > >   Call Trace:
> > >    shrink_lruvec
> > >    mem_cgroup_iter
> > >    shrink_node
> > >    do_try_to_free_pages
> > >    try_to_free_pages
> > >    __alloc_frozen_pages_noprof
> > >    alloc_pages_noprof
> > >    pte_alloc_one
> > >    __pte_alloc
> > >    handle_mm_fault
> > > 
> > > Nothing promises direct reclaim returns in bounded time, and the scan
> > > loop in shrink_lruvec() only calls cond_resched(), which is a no-op on
> > > PREEMPTION kernels.  Involuntary preemption is not a Tasks-RCU
> > > quiescent state, so the reclaiming task never reports one and becomes a
> > > holdout.
> > 
> > I still don't understand why cond_resched() is being treated as involuntary
> > preemption but that is orthogonal to this patch.
> 
> The history is that cond_resched() was originally intended to be a
> preemption point in any otherwise non-preemptible kernel.  Therefore,
> because it is a preemption point, it counts as an involuntary context
> switch.

Thanks for the explanation. Is cond_resched_tasks_rcu_qs() voluntary or
involuntary context switch?
Re: [PATCH] mm/vmscan: report RCU-tasks quiescent states in shrink_lruvec()
Posted by Paul E. McKenney 1 month, 2 weeks ago
On Tue, Aug 11, 2026 at 10:54:23AM -0700, Shakeel Butt wrote:
> On Tue, Aug 11, 2026 at 10:38:40AM -0700, Paul E. McKenney wrote:
> > On Mon, Aug 10, 2026 at 09:43:03PM -0700, Shakeel Butt wrote:
> > > On Mon, Aug 10, 2026 at 02:57:36AM -0700, Breno Leitao wrote:
> > > > I am seeing some rcu_tasks stalls in the Meta fleet during reclaim.
> > > > 
> > > >   INFO: rcu_tasks detected stalls on tasks:
> > > > 	0000000088620d09: .. nvcsw: 6735/6735 holdout: 1 idle_cpu: -1/8
> > > > 	task:GlobalCPUThread state:R  running task  pid:2552016 tgid:2524552
> > > >   Call Trace:
> > > >    shrink_lruvec
> > > >    mem_cgroup_iter
> > > >    shrink_node
> > > >    do_try_to_free_pages
> > > >    try_to_free_pages
> > > >    __alloc_frozen_pages_noprof
> > > >    alloc_pages_noprof
> > > >    pte_alloc_one
> > > >    __pte_alloc
> > > >    handle_mm_fault
> > > > 
> > > > Nothing promises direct reclaim returns in bounded time, and the scan
> > > > loop in shrink_lruvec() only calls cond_resched(), which is a no-op on
> > > > PREEMPTION kernels.  Involuntary preemption is not a Tasks-RCU
> > > > quiescent state, so the reclaiming task never reports one and becomes a
> > > > holdout.
> > > 
> > > I still don't understand why cond_resched() is being treated as involuntary
> > > preemption but that is orthogonal to this patch.
> > 
> > The history is that cond_resched() was originally intended to be a
> > preemption point in any otherwise non-preemptible kernel.  Therefore,
> > because it is a preemption point, it counts as an involuntary context
> > switch.
> 
> Thanks for the explanation. Is cond_resched_tasks_rcu_qs() voluntary or
> involuntary context switch?

That one is voluntary in order to take care of Tasks RCU's need for a
voluntary context switch.

Except that it doesn't bother involving the scheduler at all, but
instead just updates the current task's state with respect to Tasks RCU.
Much faster that way.  ;-)

							Thanx, Paul
Re: [PATCH] mm/vmscan: report RCU-tasks quiescent states in shrink_lruvec()
Posted by Johannes Weiner 1 month, 2 weeks ago
On Mon, Aug 10, 2026 at 02:57:36AM -0700, Breno Leitao wrote:
> I am seeing some rcu_tasks stalls in the Meta fleet during reclaim.
> 
>   INFO: rcu_tasks detected stalls on tasks:
> 	0000000088620d09: .. nvcsw: 6735/6735 holdout: 1 idle_cpu: -1/8
> 	task:GlobalCPUThread state:R  running task  pid:2552016 tgid:2524552
>   Call Trace:
>    shrink_lruvec
>    mem_cgroup_iter
>    shrink_node
>    do_try_to_free_pages
>    try_to_free_pages
>    __alloc_frozen_pages_noprof
>    alloc_pages_noprof
>    pte_alloc_one
>    __pte_alloc
>    handle_mm_fault
> 
> Nothing promises direct reclaim returns in bounded time, and the scan
> loop in shrink_lruvec() only calls cond_resched(), which is a no-op on
> PREEMPTION kernels.  Involuntary preemption is not a Tasks-RCU
> quiescent state, so the reclaiming task never reports one and becomes a
> holdout.
> 
> Upgrade it to cond_resched_tasks_rcu_qs(), which reports a quiescent
> state even when cond_resched() does nothing.
> 
> PS: This has been discussed in [1]
> 
> Link: https://lore.kernel.org/all/amdWVTs0WKOxguxP@gmail.com/ [1]
> Cc: stable@vger.kernel.org
> Signed-off-by: Breno Leitao <leitao@debian.org>

Acked-by: Johannes Weiner <hannes@cmpxchg.org>
Re: [PATCH] mm/vmscan: report RCU-tasks quiescent states in shrink_lruvec()
Posted by Paul E. McKenney 1 month, 2 weeks ago
On Mon, Aug 10, 2026 at 02:57:36AM -0700, Breno Leitao wrote:
> I am seeing some rcu_tasks stalls in the Meta fleet during reclaim.
> 
>   INFO: rcu_tasks detected stalls on tasks:
> 	0000000088620d09: .. nvcsw: 6735/6735 holdout: 1 idle_cpu: -1/8
> 	task:GlobalCPUThread state:R  running task  pid:2552016 tgid:2524552
>   Call Trace:
>    shrink_lruvec
>    mem_cgroup_iter
>    shrink_node
>    do_try_to_free_pages
>    try_to_free_pages
>    __alloc_frozen_pages_noprof
>    alloc_pages_noprof
>    pte_alloc_one
>    __pte_alloc
>    handle_mm_fault
> 
> Nothing promises direct reclaim returns in bounded time, and the scan
> loop in shrink_lruvec() only calls cond_resched(), which is a no-op on
> PREEMPTION kernels.  Involuntary preemption is not a Tasks-RCU
> quiescent state, so the reclaiming task never reports one and becomes a
> holdout.
> 
> Upgrade it to cond_resched_tasks_rcu_qs(), which reports a quiescent
> state even when cond_resched() does nothing.
> 
> PS: This has been discussed in [1]
> 
> Link: https://lore.kernel.org/all/amdWVTs0WKOxguxP@gmail.com/ [1]
> Cc: stable@vger.kernel.org
> Signed-off-by: Breno Leitao <leitao@debian.org>

Reviewed-by: Paul E. McKenney <paulmck@kernel.org>

> ---
>  mm/vmscan.c | 2 +-
>  1 file changed, 1 insertion(+), 1 deletion(-)
> 
> diff --git a/mm/vmscan.c b/mm/vmscan.c
> index 26436059ea394..6ac2fde137b89 100644
> --- a/mm/vmscan.c
> +++ b/mm/vmscan.c
> @@ -6023,7 +6023,7 @@ static void shrink_lruvec(struct lruvec *lruvec, struct scan_control *sc)
>  			}
>  		}
>  
> -		cond_resched();
> +		cond_resched_tasks_rcu_qs();
>  
>  		if (nr_reclaimed < nr_to_reclaim || proportional_reclaim)
>  			continue;
> 
> ---
> base-commit: 6b8c8af514d739d0335f5579b585e02babe8a727
> change-id: 20260810-rcu_task_shrink_lruvec-711112e87de0
> 
> Best regards,
> --  
> Breno Leitao <leitao@debian.org>
>