[PATCH AUTOSEL 6.18-6.1] sched/fair: Reject misfit pulls onto busy SMT siblings on asym-capacity

Sasha Levin posted 1 patch 3 weeks, 5 days ago
kernel/sched/fair.c | 11 ++++++++++-
1 file changed, 10 insertions(+), 1 deletion(-)
[PATCH AUTOSEL 6.18-6.1] sched/fair: Reject misfit pulls onto busy SMT siblings on asym-capacity
Posted by Sasha Levin 3 weeks, 5 days ago
From: Andrea Righi <arighi@nvidia.com>

[ Upstream commit bf6aa722198d3c06e4236e8c5a480f30a64e1513 ]

When SD_ASYM_CPUCAPACITY load balancing considers pulling a misfit task,
capacity_of(dst_cpu) can overstate available compute if the SMT sibling is
busy: the core does not deliver its full nominal capacity.

If SMT is active and dst_cpu is not on a fully idle core, skip this
destination so we do not migrate a misfit expecting a capacity upgrade we
cannot actually provide.

Reported-by: Felix Abecassis <fabecassis@nvidia.com>
Signed-off-by: Andrea Righi <arighi@nvidia.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Reviewed-by: Vincent Guittot <vincent.guittot@linaro.org>
Link: https://patch.msgid.link/20260509180955.1840064-5-arighi@nvidia.com
Signed-off-by: Sasha Levin <sashal@kernel.org>
---

LLM Generated explanations, may be completely bogus:

Connection lost, reconnecting to https://agentn.us.api5.cursor.sh
(attempt 1)...
Retry attempt 1...
The background `git log --grep` search finished without finding this
commit in the 6.18.43 tree — which matches what we saw in code
(`dst_core_idle` is absent from `kernel/sched/fair.c`).

**Verdict for 6.18.y: YES.** It's a small, maintainer-reviewed
scheduling correctness fix for misfit task migration on asymmetric-
capacity + SMT systems. The buggy logic is present in this tree and the
patch should apply cleanly.

 kernel/sched/fair.c | 11 ++++++++++-
 1 file changed, 10 insertions(+), 1 deletion(-)

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index ac5f08cd01a83..fee58d01ae0f9 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -9395,6 +9395,7 @@ struct lb_env {
 
 	int			dst_cpu;
 	struct rq		*dst_rq;
+	bool			dst_core_idle;
 
 	struct cpumask		*dst_grpmask;
 	int			new_dst_cpu;
@@ -10636,10 +10637,16 @@ static bool update_sd_pick_busiest(struct lb_env *env,
 	 * We can use max_capacity here as reduction in capacity on some
 	 * CPUs in the group should either be possible to resolve
 	 * internally or be covered by avg_load imbalance (eventually).
+	 *
+	 * When SMT is active, only pull a misfit to dst_cpu if it is on a
+	 * fully idle core; otherwise the effective capacity of the core is
+	 * reduced and we may not actually provide more capacity than the
+	 * source.
 	 */
 	if ((env->sd->flags & SD_ASYM_CPUCAPACITY) &&
 	    (sgs->group_type == group_misfit_task) &&
-	    (!capacity_greater(capacity_of(env->dst_cpu), sg->sgc->max_capacity) ||
+	    (!env->dst_core_idle ||
+	     !capacity_greater(capacity_of(env->dst_cpu), sg->sgc->max_capacity) ||
 	     sds->local_stat.group_type != group_has_spare))
 		return false;
 
@@ -11206,6 +11213,8 @@ static inline void update_sd_lb_stats(struct lb_env *env, struct sd_lb_stats *sd
 	unsigned long sum_util = 0;
 	bool sg_overloaded = 0, sg_overutilized = 0;
 
+	env->dst_core_idle = !sched_smt_active() || is_core_idle(env->dst_cpu);
+
 	do {
 		struct sg_lb_stats *sgs = &tmp_sgs;
 		int local_group;
-- 
2.53.0

Re: [PATCH AUTOSEL 6.18-6.1] sched/fair: Reject misfit pulls onto busy SMT siblings on asym-capacity
Posted by Matthieu Baerts 1 week, 4 days ago
Hi Sasha,

On 31/08/2026 15:23, Sasha Levin wrote:
> From: Andrea Righi <arighi@nvidia.com>
> 
> [ Upstream commit bf6aa722198d3c06e4236e8c5a480f30a64e1513 ]
> 
> When SD_ASYM_CPUCAPACITY load balancing considers pulling a misfit task,
> capacity_of(dst_cpu) can overstate available compute if the SMT sibling is
> busy: the core does not deliver its full nominal capacity.
> 
> If SMT is active and dst_cpu is not on a fully idle core, skip this
> destination so we do not migrate a misfit expecting a capacity upgrade we
> cannot actually provide.
> 
> Reported-by: Felix Abecassis <fabecassis@nvidia.com>
> Signed-off-by: Andrea Righi <arighi@nvidia.com>
> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
> Reviewed-by: Vincent Guittot <vincent.guittot@linaro.org>
> Link: https://patch.msgid.link/20260509180955.1840064-5-arighi@nvidia.com
> Signed-off-by: Sasha Levin <sashal@kernel.org>
> ---
> 
> LLM Generated explanations, may be completely bogus:
> 
> Connection lost, reconnecting to https://agentn.us.api5.cursor.sh
> (attempt 1)...
> Retry attempt 1...
> The background `git log --grep` search finished without finding this
> commit in the 6.18.43 tree — which matches what we saw in code
> (`dst_core_idle` is absent from `kernel/sched/fair.c`).

Mmh, you might need to check this. Is it why it didn't detect issues on v6.1,
see below.

> **Verdict for 6.18.y: YES.** It's a small, maintainer-reviewed
> scheduling correctness fix for misfit task migration on asymmetric-
> capacity + SMT systems. The buggy logic is present in this tree and the
> patch should apply cleanly.
> 
>  kernel/sched/fair.c | 11 ++++++++++-
>  1 file changed, 10 insertions(+), 1 deletion(-)
> 
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index ac5f08cd01a83..fee58d01ae0f9 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -9395,6 +9395,7 @@ struct lb_env {
>  
>  	int			dst_cpu;
>  	struct rq		*dst_rq;
> +	bool			dst_core_idle;
>  
>  	struct cpumask		*dst_grpmask;
>  	int			new_dst_cpu;
> @@ -10636,10 +10637,16 @@ static bool update_sd_pick_busiest(struct lb_env *env,
>  	 * We can use max_capacity here as reduction in capacity on some
>  	 * CPUs in the group should either be possible to resolve
>  	 * internally or be covered by avg_load imbalance (eventually).
> +	 *
> +	 * When SMT is active, only pull a misfit to dst_cpu if it is on a
> +	 * fully idle core; otherwise the effective capacity of the core is
> +	 * reduced and we may not actually provide more capacity than the
> +	 * source.
>  	 */
>  	if ((env->sd->flags & SD_ASYM_CPUCAPACITY) &&
>  	    (sgs->group_type == group_misfit_task) &&
> -	    (!capacity_greater(capacity_of(env->dst_cpu), sg->sgc->max_capacity) ||
> +	    (!env->dst_core_idle ||
> +	     !capacity_greater(capacity_of(env->dst_cpu), sg->sgc->max_capacity) ||
>  	     sds->local_stat.group_type != group_has_spare))
>  		return false;
>  
> @@ -11206,6 +11213,8 @@ static inline void update_sd_lb_stats(struct lb_env *env, struct sd_lb_stats *sd
>  	unsigned long sum_util = 0;
>  	bool sg_overloaded = 0, sg_overutilized = 0;
>  
> +	env->dst_core_idle = !sched_smt_active() || is_core_idle(env->dst_cpu);

On v6.1, I got this:

kernel/sched/fair.c: In function 'update_sd_lb_stats':
kernel/sched/fair.c:9912:53: error: implicit declaration of function 'is_core_idle' [-Wimplicit-function-declaration]
   9912 |         env->dst_core_idle = !sched_smt_active() || is_core_idle(env->dst_cpu);
        |                                                     ^~~~~~~~~~~~

Cheers,
Matt
-- 
Sponsored by the NGI0 Core fund.