[tip: sched/core] sched/fair: Skip misfit load accounting when the destination CPU cannot help

tip-bot2 for Ricardo Neri posted 1 patch 1 month, 3 weeks ago
kernel/sched/fair.c | 21 +++++++++++++++++----
1 file changed, 17 insertions(+), 4 deletions(-)
[tip: sched/core] sched/fair: Skip misfit load accounting when the destination CPU cannot help
Posted by tip-bot2 for Ricardo Neri 1 month, 3 weeks ago
The following commit has been merged into the sched/core branch of tip:

Commit-ID:     0fbd428d078876cefe908efdbee47a4d5adc5f24
Gitweb:        https://git.kernel.org/tip/0fbd428d078876cefe908efdbee47a4d5adc5f24
Author:        Ricardo Neri <ricardo.neri-calderon@linux.intel.com>
AuthorDate:    Mon, 20 Jul 2026 19:43:20 -07:00
Committer:     Peter Zijlstra <peterz@infradead.org>
CommitterDate: Fri, 07 Aug 2026 18:27:10 +02:00

sched/fair: Skip misfit load accounting when the destination CPU cannot help

In domains with asymmetric capacity, identifying misfit load in a
scheduling group is not useful when the destination CPU cannot help (i.e.,
its capacity exceeds the group's maximum CPU capacity by less than ~5%). In
such cases, it also prevents load balance among clusters of equal capacity
when CONFIG_SCHED_CLUSTER is enabled. This happens because
update_sd_pick_busiest() skips candidate groups of type misfit_task if the
destination CPU has similar capacity.

Skipping misfit load accounting in this situation allows the group to be
classified as has_spare or fully_busy and lets load balancing proceed. Keep
marking scheduling groups as overloaded when misfit tasks are present. The
sg_overloaded flag propagates to the root domain and allows bigger CPUs in
it to help via newly idle balance.

Signed-off-by: Ricardo Neri <ricardo.neri-calderon@linux.intel.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Reviewed-by: Christian Loehle <christian.loehle@arm.com>
Reviewed-by: Chen Yu <yu.c.chen@intel.com>
Reviewed-by: Vincent Guittot <vincent.guittot@linaro.org>
Tested-by: Christian Loehle <christian.loehle@arm.com>
Tested-by: Andrea Righi <arighi@nvidia.com>
Link: https://patch.msgid.link/20260720-rneri-fix-cas-clusters-v6-4-bb500bf4afd4@linux.intel.com
---
 kernel/sched/fair.c | 21 +++++++++++++++++----
 1 file changed, 17 insertions(+), 4 deletions(-)

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index f9d0edd..f1b4db3 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -11970,12 +11970,25 @@ static inline void update_sg_lb_stats(struct lb_env *env,
 			continue;
 
 		if (sd_flags & SD_ASYM_CPUCAPACITY) {
-			/* Check for a misfit task on the cpu */
-			if (sgs->group_misfit_task_load < rq->misfit_task_load) {
-				sgs->group_misfit_task_load = rq->misfit_task_load;
-
+			if (rq->misfit_task_load) {
+				/*
+				 * Always mark the root domain overloaded so big
+				 * CPUs can pick up misfit tasks via newly idle
+				 * balance.
+				 */
 				if (balancing_at_rd)
 					*sg_overloaded = 1;
+
+				/*
+				 * Only account misfit load if @dst_cpu can
+				 * help; otherwise, the group may be classified
+				 * as misfit_task and update_sd_pick_busiest()
+				 * will skip it.
+				 */
+				if (capacity_greater(capacity_of(env->dst_cpu),
+						     group->sgc->max_capacity) &&
+				    (sgs->group_misfit_task_load < rq->misfit_task_load))
+					sgs->group_misfit_task_load = rq->misfit_task_load;
 			}
 		} else if (env->idle && sched_reduced_capacity(rq, env->sd)) {
 			/* Check for a task running on a CPU with reduced capacity */