From nobody Mon Sep 28 23:12:43 2026 Received: from m16.mail.163.com (m16.mail.163.com [117.135.210.5]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A444D37F30D for ; Sat, 15 Aug 2026 11:04:37 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=117.135.210.5 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786791880; cv=none; b=JZhXiUI08luD9LZoyYF1ISNy63klY4Jla3bLKQQ3A51iOQcU3khwF3TUBPKqAAwAiPxsk6UjofymIkqbsohefPvsdsduexLrizeCF5rxv3q9ktM7m64eSJ0tKn5UhT1cZ+jFyYWOJZpdUbXyx8/umd/qJvBKOC6wd9+KgD8xD3o= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786791880; c=relaxed/simple; bh=YxICK9eRAEZm/b4KxMY6YgmHqrAXs3BXuBQsTXme5fo=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=eVqSD1lznYAi9EznseB4FWGj937J8H6Lgudjj+/z4EgU32+CIEYnpkTps8rhsa18ndMmgwpBugD321wnQwUBX1szwtf15lHyGIKYtvr3rKUDdin3e/q2nLpXXInVoz+76txbnItFIE3eUVKz1l2LDF+4QSFYd7GwRL114vimCXY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com; spf=pass smtp.mailfrom=163.com; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b=ZTZamVu0; arc=none smtp.client-ip=117.135.210.5 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=163.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b="ZTZamVu0" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=163.com; s=s110527; h=From:To:Subject:Date:Message-Id:MIME-Version; bh=mw Zx8AptPNgUqEDEdFiFTHEgRLhgM0jjpuxmqGdyRGw=; b=ZTZamVu0Sgqq3WNROo QUlmK2ww559L+q+KsriO10slkJDxJdN5cbFG1Tn2aZxQBZpZmq9ejuz/yHTWEYjV c/JjETir1Ek2FmvotTcjaq+R6BHwH36LPzr2EbHk3zvg/SNfEefaAOUKD5Fn0Gpx 36BKgDw7BXcv7MBBE4ZRfGvFU= Received: from zhaoxin-MS-7E12.. (unknown []) by gzsmtp3 (Coremail) with SMTP id PigvCgB3tQRjR4BqOFTVNA--.50445S3; Sat, 15 Aug 2026 19:03:02 +0800 (CST) From: Xin Zhao To: mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com Cc: linux-kernel@vger.kernel.org, Xin Zhao Subject: [PATCH 01/10] sched/fair: Do not set_rd_overloaded() if rd->online != env->cpus Date: Sat, 15 Aug 2026 19:02:48 +0800 Message-Id: <20260815110257.124354-2-jackzxcui1989@163.com> X-Mailer: git-send-email 2.34.1 In-Reply-To: <20260815110257.124354-1-jackzxcui1989@163.com> References: <20260815110257.124354-1-jackzxcui1989@163.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: PigvCgB3tQRjR4BqOFTVNA--.50445S3 X-Coremail-Antispam: 1Uf129KBjvJXoW7ZF1UKr1fArWUAF1UAry7KFg_yoW8Cr4Dpa yDKayIgr48twnxKasrAF4vg3yUJws5J3WYva1ak3yrJF1Yyw1YvrW0v34DCr4rKa4rZ3WS vrWUtFW7u3WjyFJanT9S1TB71UUUUU7qnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDUYxBIdaVFxhVjvjDU0xZFpf9x0piSAp5UUUUU= X-CM-SenderInfo: pmdfy650fxxiqzyzqiywtou0bp/xtbCvwalr2qAR2ZkzQAA3R Content-Type: text/plain; charset="utf-8" In update_sg_lb_stats(), it only traverses sched_group that belongs to env->cpus, but env->cpus may not necessarily equal rd->online. This can lead to the incorrect clearing of the overloaded flag of rd. For example, if cpuA belongs to the online CPU mask of the rd but does not belong to env->cpus, and cpuA consistently maintains nr_running >=3D 2, while other CPUs in rd->online keep rq->nr_running <=3D 1, the overloaded flag of rd will not be set until next update of update_sd_lb_stats() for that rd. During this period, sched_balance_newidle() will prematurely return due to the incorrect assumption that the rd is in a non-overloaded state. In update_sd_lb_stats(), add a check to verify whether rd->online is equal to env->cpus before calling set_rd_overloaded() to avoid such incorrect settings. Signed-off-by: Xin Zhao --- kernel/sched/fair.c | 8 ++++++-- 1 file changed, 6 insertions(+), 2 deletions(-) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index dcf860c59a14..13e873b1ef58 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -12679,8 +12679,12 @@ static inline void update_sd_lb_stats(struct lb_en= v *env, struct sd_lb_stats *sd env->fbq_type =3D fbq_classify_group(&sds->busiest_stat); =20 if (!env->sd->parent) { - /* update overload indicator if we are at root domain */ - set_rd_overloaded(env->dst_rq->rd, sg_overloaded); + /* + * Update overload indicator if we are at root domain. + * Note that env->cpus may change during sched_balance_rq(). + */ + if (cpumask_equal(env->dst_rq->rd->online, env->cpus)) + set_rd_overloaded(env->dst_rq->rd, sg_overloaded); =20 /* Update over-utilization (tipping point, U >=3D 0) indicator */ set_rd_overutilized(env->dst_rq->rd, sg_overutilized); --=20 2.34.1 From nobody Mon Sep 28 23:12:43 2026 Received: from m16.mail.163.com (m16.mail.163.com [117.135.210.2]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9C5623515CA for ; Sat, 15 Aug 2026 11:04:35 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=117.135.210.2 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786791879; cv=none; b=e/Xjk5o3fDe5xagcDo1zag3Ndr81tBydAD7oc85yHorO/I4WUKxDYjwYXdHH7U4TDjhMbHmy+Z0yMTxLjMXrg8Q4LycsgTpQFgVvo6WNONqhF4/WtALCLTSo4MAPiMXGRHZgkBX43DWVaMzwXynvRHFbSIm5g2ZG0uL7oS3pbOo= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786791879; c=relaxed/simple; bh=IYwCHGvRCE0IUY7GbipIBjw2hj/OsJ7eLxHeg+OkQyg=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=UgwQWSuDzRQeYRDbid6oZtergsWsd0poGtR80DygOFjnBxv7/rvVo1tdxzBXldPRS7Bsw6WMR0Qz/mLGeSqtNixIIZinMIJnd3ThXHo8/cP5xRTnACgINPsXotZzlKJb2r9xBiK05Y4ez1fGy1txlD4TMpoIlC6VOe8TC8t6wmU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com; spf=pass smtp.mailfrom=163.com; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b=kyae3sY4; arc=none smtp.client-ip=117.135.210.2 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=163.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b="kyae3sY4" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=163.com; s=s110527; h=From:To:Subject:Date:Message-Id:MIME-Version; bh=wn OvNfyp/TALdwlgN291rg+EW9FopoczSlJR5mBkWNc=; b=kyae3sY4fEVJEtOLLu jUfITjLl32BjStj/ozGPShwhWwtw3dlxssl6vIsH5HZcIVXig+pQzjj2vw4ATtvG 3voSPqwe/NEhs2p2iIvhvORcupXyc7USP54EOirC32d9SZzJ3hADceezurTWr15+ nzuvq90ET3CIWNvpHaJILwpSE= Received: from zhaoxin-MS-7E12.. (unknown []) by gzsmtp3 (Coremail) with SMTP id PigvCgB3tQRjR4BqOFTVNA--.50445S4; Sat, 15 Aug 2026 19:03:03 +0800 (CST) From: Xin Zhao To: mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com Cc: linux-kernel@vger.kernel.org, Xin Zhao Subject: [PATCH 02/10] scbed/fair: Remove duplicate check for busiest_cpu in active_load_balance_cpu_stop() Date: Sat, 15 Aug 2026 19:02:49 +0800 Message-Id: <20260815110257.124354-3-jackzxcui1989@163.com> X-Mailer: git-send-email 2.34.1 In-Reply-To: <20260815110257.124354-1-jackzxcui1989@163.com> References: <20260815110257.124354-1-jackzxcui1989@163.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: PigvCgB3tQRjR4BqOFTVNA--.50445S4 X-Coremail-Antispam: 1Uf129KBjvdXoWruw1UKFWxuF43KFyrZr1kZrb_yoWkZrX_ur nrurn3Kr1jvFn09ws3CrZ3Xr1F9a4YgF18G3W0gFZrCry0qrZrKrZakFn5Xr93W3yIyFZF vwn0gF1jqw1DujkaLaAFLSUrUUUUjb8apTn2vfkv8UJUUUU8Yxn0WfASr-VFAUDa7-sFnT 9fnUUvcSsGvfC2KfnxnUUI43ZEXa7xRuiiSUUUUUU== X-CM-SenderInfo: pmdfy650fxxiqzyzqiywtou0bp/xtbCwAelr2qAR2dxhQAA3z Content-Type: text/plain; charset="utf-8" The check for cpu_active(busiest_cpu) already ensures that busiest_cpu has not gone down. An additional check for busiest_cpu !=3D smp_processor_id() is redundant. After this modification, busiest_cpu will no longer be bound to smp_processor_id(), allowing the active_load_balance_cpu_stop function to accommodate more scenarios, such as preempt active balancing feature that will be addressed in later patches. Signed-off-by: Xin Zhao --- kernel/sched/fair.c | 4 +--- 1 file changed, 1 insertion(+), 3 deletions(-) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 13e873b1ef58..11c104010b2e 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -13769,9 +13769,7 @@ static int active_load_balance_cpu_stop(void *data) if (!cpu_active(busiest_cpu) || !cpu_active(target_cpu)) goto out_unlock; =20 - /* Make sure the requested CPU hasn't gone down in the meantime: */ - if (unlikely(busiest_cpu !=3D smp_processor_id() || - !busiest_rq->active_balance)) + if (unlikely(!busiest_rq->active_balance)) goto out_unlock; =20 /* Is there any task to move? */ --=20 2.34.1 From nobody Mon Sep 28 23:12:43 2026 Received: from m16.mail.163.com (m16.mail.163.com [117.135.210.4]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2FAAB3EC6A8 for ; Sat, 15 Aug 2026 11:04:40 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=117.135.210.4 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786791885; cv=none; b=MDYSY6QDpgPLsgITNv4DX2W8EXuXeogsHo7EwWzCLfcqk13JAZIwpsrfIgQWk7yospqU1TjLOTXybd/Womr6Wj/UQWOQdKu9umJbIPd3q+FFkg1YAJPUarzaMpDPJxMUldS4VojWdhxcdVqweYepQLS8YRd7BCiVm0bUfj2vX2k= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786791885; c=relaxed/simple; bh=MmE42vByQegAZN4de4r4cEy2IgaWKlOiDSWRCGZyQM8=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=Vjs6QZ2C16RBUOObr3hNImuAvW7YwCnkke/5crvMHg60+JT/Bn4l3dkXZiPmQ4nPazAQMyM3pOpDIhVbR5lSUIKzk5QdNrYHRYbGt5DRP2cIPQbuLTz0t085SqWDvH6+PPoxRsyOJGzKN8CZo6kcbdR2J3vMs0YLXZZOKSKk5ak= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com; spf=pass smtp.mailfrom=163.com; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b=JCX4eRXo; arc=none smtp.client-ip=117.135.210.4 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=163.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b="JCX4eRXo" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=163.com; s=s110527; h=From:To:Subject:Date:Message-Id:MIME-Version; bh=A6 5RVDN267qKaanvkAeYvWohUstOEpo7HrQt+WE9L84=; b=JCX4eRXoM34IpSYUeE GnYHnKtj6wT+Amume38B7Tl/3+jkWVBziOK6i6eZYKWME0MQ85ZL6r9NJn3Amv2i LERkBUjoh61+HJ3ZiDDer8wNQBzDZbkoncjk4mw3nChIBDmo5GHcdXPz5O3TquMx ZRGR4+Hw3Hch0P1LUwlVWppmo= Received: from zhaoxin-MS-7E12.. (unknown []) by gzsmtp3 (Coremail) with SMTP id PigvCgB3tQRjR4BqOFTVNA--.50445S5; Sat, 15 Aug 2026 19:03:04 +0800 (CST) From: Xin Zhao To: mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com Cc: linux-kernel@vger.kernel.org, Xin Zhao Subject: [PATCH 03/10] sched/fair: Clear active_balance at the end of active_load_balance_cpu_stop() Date: Sat, 15 Aug 2026 19:02:50 +0800 Message-Id: <20260815110257.124354-4-jackzxcui1989@163.com> X-Mailer: git-send-email 2.34.1 In-Reply-To: <20260815110257.124354-1-jackzxcui1989@163.com> References: <20260815110257.124354-1-jackzxcui1989@163.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: PigvCgB3tQRjR4BqOFTVNA--.50445S5 X-Coremail-Antispam: 1Uf129KBjvJXoW7uFy5CFW5Ar48Zr15JF1DJrb_yoW8ZF1fpr WjyayIgw4ktayYvrZ2ka18ury7uan8Jr4UGrnFqrWrZF15C3s5tw1F934xur45Zrs5CFn0 yay7tr4UCa4rGr7anT9S1TB71UUUUU7qnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDUYxBIdaVFxhVjvjDU0xZFpf9x0zRmFAZUUUUU= X-CM-SenderInfo: pmdfy650fxxiqzyzqiywtou0bp/xtbC5wmmsGqAR2mLKwAA3c Content-Type: text/plain; charset="utf-8" The rq->active_balance flag is used to prevent multiple CPUs from simultaneously dispatching active balance stop tasks. Since there can only ever be one consumer of the stop task, it is not strictly necessary to protect the setting of rq->active_balance to 0 with the rq lock in active_load_balance_cpu_stop(). Therefore, we can move the action of clearing rq->active_balance to the end of active_load_balance_cpu_stop(). The benefit of this approach is that the task load of dst_rq will change due to the execution of attach_one_task(), which helps avoid prematurely clearing rq->active_balance before attach_one_task(), thus preventing unnecessary dispatch of duplicate active balance stop tasks. Active balance stop task is triggered only when rq->active_balance flag changes from 0 to 1, and there can be at most one consumer of active balance stop task at any given time. Therefore, we should never see zero rq->active_balance in active_load_balance_cpu_stop(), use WARN_ON_ONCE instead. Signed-off-by: Xin Zhao --- kernel/sched/fair.c | 5 ++--- 1 file changed, 2 insertions(+), 3 deletions(-) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 11c104010b2e..20d03ceed9d7 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -13769,8 +13769,7 @@ static int active_load_balance_cpu_stop(void *data) if (!cpu_active(busiest_cpu) || !cpu_active(target_cpu)) goto out_unlock; =20 - if (unlikely(!busiest_rq->active_balance)) - goto out_unlock; + WARN_ON_ONCE(!busiest_rq->active_balance); =20 /* Is there any task to move? */ if (busiest_rq->nr_running <=3D 1) @@ -13815,13 +13814,13 @@ static int active_load_balance_cpu_stop(void *dat= a) } rcu_read_unlock(); out_unlock: - busiest_rq->active_balance =3D 0; rq_unlock(busiest_rq, &rf); =20 if (p) attach_one_task(target_rq, p); =20 local_irq_enable(); + busiest_rq->active_balance =3D 0; =20 return 0; } --=20 2.34.1 From nobody Mon Sep 28 23:12:43 2026 Received: from m16.mail.163.com (m16.mail.163.com [117.135.210.5]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DB7283DDDC4 for ; Sat, 15 Aug 2026 11:04:45 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=117.135.210.5 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786791892; cv=none; b=Q9VR4xuLJXVX9kqGHD1RrV96wn9iFc8EZSN3oGYXdSWhq94NHVaonVittJz/fjr3UsezZpwNajoecSfkyrL8PF9n44znGSZG786jyGG9pq5swfzxPbhzPe0JFZopNmfePb5kPSLev+k04MgGQojpiwImENl+GdfdL4tE2f882Ho= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786791892; c=relaxed/simple; bh=gQ2GOV8Dmrszsdh+wQhrIb1AXtKM9wHta2VUf59x08M=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=BSQGd4+k+Evdz1Ac9thuV3m7i/YzLd7KA4PWHbXrugskxLOG08ZlBwbWMSzDH8j0kh9SdVzbrlZ3DAb/HeE2FGUuEkQ36ocfH5VwIpu/q0OQPW78cuo1rZgLwoNa1w+KnXQsVpFuwl17SrOPgBVkn11iaQf1SQ8LcUKGvmO1A74= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com; spf=pass smtp.mailfrom=163.com; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b=RYfbiHwL; arc=none smtp.client-ip=117.135.210.5 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=163.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b="RYfbiHwL" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=163.com; s=s110527; h=From:To:Subject:Date:Message-Id:MIME-Version; bh=T7 aj3ANDLcZnLgr2CwfVQ6lQY11yYofrsfu5AMmmt8M=; b=RYfbiHwL15YEkuoq8A 6rVvlBHC8k4RJYvYQY3Hyy7p5i7xWTACuxnq2s3RLw+CCUUusidJaQO1I6SqiVN6 PHJeJrhtnR+408rR1rOjMgiHymWDLlwAYBdu+BKEKqNImpDnTOQG2A6YN5X2wEyl vRM2sqagnhI7+dvT74z7iQNQc= Received: from zhaoxin-MS-7E12.. (unknown []) by gzsmtp3 (Coremail) with SMTP id PigvCgB3tQRjR4BqOFTVNA--.50445S6; Sat, 15 Aug 2026 19:03:06 +0800 (CST) From: Xin Zhao To: mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com Cc: linux-kernel@vger.kernel.org, Xin Zhao Subject: [PATCH 04/10] sched/fair: Add LB_PROMOTE feature to enhance real-time performance of fair tasks Date: Sat, 15 Aug 2026 19:02:51 +0800 Message-Id: <20260815110257.124354-5-jackzxcui1989@163.com> X-Mailer: git-send-email 2.34.1 In-Reply-To: <20260815110257.124354-1-jackzxcui1989@163.com> References: <20260815110257.124354-1-jackzxcui1989@163.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: PigvCgB3tQRjR4BqOFTVNA--.50445S6 X-Coremail-Antispam: 1Uf129KBjvJXoWxurWUJw13JF45CFy5Xr13urg_yoWrGrykpa 93Xr4Yyr4DJryIy34fZr4xXr1ru393Cry3tr1kGw1xXas8Xr1ayw4IgF4UKFs3C397Za1j q3W2q343uF1jvaDanT9S1TB71UUUUU7qnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDUYxBIdaVFxhVjvjDU0xZFpf9x0zZg4ddUUUUU= X-CM-SenderInfo: pmdfy650fxxiqzyzqiywtou0bp/xtbCvwqmsGqAR2plAQAA3A Content-Type: text/plain; charset="utf-8" Embedded platforms commonly use CONFIG_HZ_250, and testing has revealed that there are numerous instances of unreasonable CPU idle events on such platforms. Unreasonable CPU idle refers to situations where the CPU enters an idle state for a duration of time (t), while there are tasks that can run on the idle CPU and are not limited by cgroup constraints, yet these tasks remain unscheduled for a duration greater than (t), t > 2.5 ms. Testing has shown that over 95% of these events last less than 4 ms, but there are still some instances of longer durations between 4-5ms, even occasionally between 5-10 ms. For a real-time system, scheduling delays greater than 4 ms can lead to performance spikes. Enabling this option can effectively reduce the occurrence of unreasonable CPU idle events on low HZ systems like CONFIG_HZ_250, and completely eliminate events exceeding 4 ms. Note that the feature only affects fair tasks. Note that enabling this feature will increase sys%, as it uses CPU time that would have been idle to expedite the scheduling of tasks. There will also be some CPU overhead involved in searching for suitable tasks. This feature has been split into several smaller patches, which will be elaborated on one by one later. Below are some test data: Test one compares the number of unreasonable CPU idle events and their distribution when this feature is enabled versus when it is not, under the same fillback scenario. The test duration was 60 seconds. LB_PROMOTE(on/off) 2.5-3ms 3-4ms 4ms+ index 0 on 0 0 0 index 1 off 4 13 1 index 2 on 0 0 0 index 3 off 6 3 0 index 4 on 0 0 0 index 5 off 1 1 0 Test two compares the performance of the system with and without the feature enabled, based on the same fillback scenario. Each test lasts for 25 minutes, and a total of 15 comparative tests were conducted. The results include a comparison of the maximum and median values of end-to-end latency and sys%. LB_PROMOTE(on/off) on off end-to-end latency(max) 172 180 end-to-end latency(median of avg) 166 167.68 sys%(max) 9.68 9.35 sys%(median of avg) 8.81 8.55 Signed-off-by: Xin Zhao --- kernel/sched/features.h | 24 ++++++++++++++++++++++++ 1 file changed, 24 insertions(+) diff --git a/kernel/sched/features.h b/kernel/sched/features.h index 8f0dee8fc475..4916a4b89ab3 100644 --- a/kernel/sched/features.h +++ b/kernel/sched/features.h @@ -142,3 +142,27 @@ SCHED_FEAT(LATENCY_WARN, false) */ SCHED_FEAT(NI_RANDOM, true) SCHED_FEAT(NI_RATE, true) + +/* + * Embedded platforms commonly use CONFIG_HZ_250, and testing has revealed= that + * there are numerous instances of unreasonable CPU idle events on such + * platforms. Unreasonable CPU idle refers to situations where the CPU ent= ers + * an idle state for a duration of time (t), while there are tasks that ca= n run + * on the idle CPU and are not limited by cgroup constraints, yet these ta= sks + * remain unscheduled for a duration greater than (t), t > 2.5 ms. + * + * Testing has shown that over 95% of these events last less than 4 ms, but + * there are still some instances of longer durations between 4-5ms, even + * occasionally between 5-10 ms. For a real-time system, scheduling delays + * greater than 4 ms can lead to performance spikes. + * + * Enabling this option can effectively reduce the occurrence of unreasona= ble + * CPU idle events on low HZ systems like CONFIG_HZ_250, and completely + * eliminate events exceeding 4 ms. Note that the feature only affects fair + * tasks. + * + * Note that enabling this feature will increase sys%, as it uses CPU time= that + * would have been idle to expedite the scheduling of tasks. There will al= so be + * some CPU overhead involved in searching for suitable tasks. + */ +SCHED_FEAT(LB_PROMOTE, false) --=20 2.34.1 From nobody Mon Sep 28 23:12:43 2026 Received: from m16.mail.163.com (m16.mail.163.com [220.197.31.4]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D3119364925 for ; Sat, 15 Aug 2026 11:04:36 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=220.197.31.4 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786791880; cv=none; b=ahIuLQmcORZm4OjpVPyX9Ywg8+um3KUHh5jDL5PVwdmKUNX75TGQHlI5RuuK+uw4FFJcEkVv8dIl8bAyTSnBvnDzWQLyujYVtRUlURLz4ZBSktqhhDr5F8AT+wBRof0hoJ+uIDisiHHaRW4JSH4EqX1uBmsPQqm/gpSyKt5QxXI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786791880; c=relaxed/simple; bh=7rqDzVbAOrvGHPe3+y17ORUypnLLrLSrXFuJ0eFrblU=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=Kb5J+22k24sTI/jaKHuuRhAlUtpQW8967uFMmF/Y5TcMiXSBaFHIDdQYBVi1CC1moMJ58yuHC0k3uPOChfwpdZYZIOFbhN0IA5TvwLfJ0OY8qS1Kpf71ASPz758d6gGtOl527DBD79H5ulQ+QoYuLT+t+MMbWdr8f2cbx9AOCS4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com; spf=pass smtp.mailfrom=163.com; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b=Swy55J5Q; arc=none smtp.client-ip=220.197.31.4 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=163.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b="Swy55J5Q" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=163.com; s=s110527; h=From:To:Subject:Date:Message-Id:MIME-Version; bh=4t nSggQ8oN3baWQ2VgtsX9D1878Stv467Btkc2KmqxQ=; b=Swy55J5QDALE69HNHg d7FxFnvMeEHNXb+TycnT0icqxOOqPGCUU6PW31O02UYMIpcnw9oyQ3YAC8i7Fk9e el1WB8AoyvM1OtcRKgr4Je8Q9FMecJnhP5nZ1d36QZZDlXcq7Mv9ZkH9JsBC0nHj U5xx1a62wgAi9H+PTAi4ewVeQ= Received: from zhaoxin-MS-7E12.. (unknown []) by gzsmtp3 (Coremail) with SMTP id PigvCgB3tQRjR4BqOFTVNA--.50445S7; Sat, 15 Aug 2026 19:03:07 +0800 (CST) From: Xin Zhao To: mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com Cc: linux-kernel@vger.kernel.org, Xin Zhao Subject: [PATCH 05/10] sched/fair: Introduce select_task_rq_fair_thin() to select rq when LB_PROMOTE Date: Sat, 15 Aug 2026 19:02:52 +0800 Message-Id: <20260815110257.124354-6-jackzxcui1989@163.com> X-Mailer: git-send-email 2.34.1 In-Reply-To: <20260815110257.124354-1-jackzxcui1989@163.com> References: <20260815110257.124354-1-jackzxcui1989@163.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: PigvCgB3tQRjR4BqOFTVNA--.50445S7 X-Coremail-Antispam: 1Uf129KBjvJXoWxKF4DZF1rGr4UGr1rtFy3twb_yoW7Xr15pF 4Fqw13trsrAw4Ivw1fCrs7Cr1Yv34rGw47Kr1fJF95Ca45Xryv9F1FgrnxXFyfCr1kZFyj qFW8Kr17CrWqvaUanT9S1TB71UUUUU7qnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDUYxBIdaVFxhVjvjDU0xZFpf9x0pCjqAgUUUUU= X-CM-SenderInfo: pmdfy650fxxiqzyzqiywtou0bp/xtbC6AumsGqAR2uHaQAA3d Content-Type: text/plain; charset="utf-8" The logic of select_task_rq_fair() is relatively complex, and on typical embedded systems, the number of CPUs in sd_llc domain is often limited to a maximum of 4, and sometimes only 2. This makes the complex logic of select_task_rq_fair() seem less necessary. Additionally, the update logic for the nr_idle_scan value that select_task_rq_fair() relies on is also quite time-consuming. Moreover, embedded systems require better real-time performance, and with fewer CPUs available, it becomes necessary to bind certain tasks across the sd_llc range. The default enabled feature, SD_WAKE_AFFINE, causes select_task_rq_fair() to take the fast path, often overlooking some idle CPUs across sd_llc domain, leading to increased scheduling latency. To address this, we introduce the select_task_rq_fair_thin() function, which serves as a streamlined version of select_task_rq_fair(). It can quickly perform CPU selection while also considering the real-time requirements of embedded systems. select_task_rq_fair_thin() retains the priority selection logic for prev_cpu and recent_used_cpu, and it will prioritize CPUs within the sd_llc. If there are no idle CPUs in the sd_llc, it will then look for other available CPUs to run. When the LB_PROMOTE feature is enabled, select_task_rq_fair_thin() will replace the original select_task_rq_fair(), and the update logic for nr_idle_scan that select_task_rq_fair() relies on will no longer need to be executed. Testing has shown that in our system with 18 CPUs running at 2.1GHz, where the first three sd_llc domains each contains 4 CPUs and the last sd_llc contains 2 CPUs, under same fillback scenario, select_task_rq_fair_thin() executes 25% faster than the original select_task_rq_fair(). It saves 22ms over a 10-second period, with this optimization accounting for 0.174% of total system time. Additionally, we measured the execution time of update_idle_cpu_scan, which took 0.5ms over the same 10-second period. If we use select_task_rq_fair_thin() instead, this time can be eliminated, accounting for 0.04% of total system time. Therefore, the overall optimization contributes to a reduction of 0.214% of total system time. Signed-off-by: Xin Zhao --- kernel/sched/fair.c | 52 ++++++++++++++++++++++++++++++++++++++++++++- 1 file changed, 51 insertions(+), 1 deletion(-) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 20d03ceed9d7..f10e709921fd 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -9661,6 +9661,51 @@ static int find_energy_efficient_cpu(struct task_str= uct *p, int prev_cpu) return target; } =20 +/* + * A streamlined version of select_task_rq_fair(). + * It runs faster than select_task_rq_fair, especially when there are not + * many CPUs. It will prioritize selecting an idle CPU in the following or= der: + * 1. prev_cpu + * 2. recent_used_cpu + * 3. cpu belongs to intersection of sd_llc and cpus_ptr + * 4. cpu belongs to cpus_ptr but not belongs to sd_llc + * If there is no idle CPU in cpus_ptr, it will select prev_cpu. + */ +static int select_task_rq_fair_thin(struct task_struct *p, int prev_cpu, i= nt wake_flags) +{ + int recent_used_cpu, target, cpu, start =3D nr_cpu_ids; + struct sched_domain *sd; + + if (likely(available_idle_cpu(prev_cpu))) + return prev_cpu; + + recent_used_cpu =3D p->recent_used_cpu; + p->recent_used_cpu =3D prev_cpu; + if (recent_used_cpu !=3D prev_cpu && available_idle_cpu(recent_used_cpu)) + return recent_used_cpu; + + target =3D prev_cpu; + rcu_read_lock(); + + sd =3D rcu_dereference(per_cpu(sd_llc, target)); + if (sd) + start =3D cpumask_first_and(sched_domain_span(sd), p->cpus_ptr); + if (start >=3D nr_cpu_ids) + start =3D cpumask_first(p->cpus_ptr); + + for_each_cpu_wrap(cpu, p->cpus_ptr, start) { + if (available_idle_cpu(cpu)) { + target =3D cpu; + goto unlock; + } + } + +unlock: + rcu_read_unlock(); + + return target; +} + /* * select_task_rq_fair: Select target runqueue for the waking task in doma= ins * that have the relevant SD flag set. In practice, this is SD_BALANCE_WAK= E, @@ -9682,6 +9727,9 @@ select_task_rq_fair(struct task_struct *p, int prev_c= pu, int wake_flags) /* SD_flags and WF_flags share the first nibble */ int sd_flag =3D wake_flags & 0xF; =20 + if (sched_feat(LB_PROMOTE)) + return select_task_rq_fair_thin(p, prev_cpu, wake_flags); + /* * required for stable ->cpus_allowed */ @@ -12564,8 +12612,10 @@ static void update_idle_cpu_scan(struct lb_env *en= v, * So the write of this hint only occurs during periodic load * balancing, rather than CPU_NEWLY_IDLE, because the latter * can fire way more frequently than the former. + * When LB_PROMOTE is enabled, select_task_rq_fair() is no longer + * used, and there is no need to update nr_idle_scan. */ - if (!sched_feat(SIS_UTIL) || env->idle =3D=3D CPU_NEWLY_IDLE) + if (!sched_feat(SIS_UTIL) || env->idle =3D=3D CPU_NEWLY_IDLE || sched_fea= t(LB_PROMOTE)) return; =20 sd_share =3D sd->shared; --=20 2.34.1 From nobody Mon Sep 28 23:12:43 2026 Received: from m16.mail.163.com (m16.mail.163.com [220.197.31.5]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 01C3A2EA754 for ; Sat, 15 Aug 2026 11:04:37 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=220.197.31.5 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786791880; cv=none; b=ZMbsNF/0F92T6EUJBIjgt3ubPr6qNfSJNCWxi8YLa6aidC7s8hjFfOSg/nO7DtvUIhXTz0H7jgNjHBTwsShrECcAeEe2EH7j9loOhb0OnTO0tn+C/vQ+4gfvAXEZ7X9mJD2ks8YJD7YUI1/fp64ScBOR4OlWopQY81tD/ueMHfs= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786791880; c=relaxed/simple; bh=9kEKGitU0KGtrdmgg+MeYLoaKIghma4epy46ju6KunM=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=rFTYLx9JcrDI5KHUCc24G5RN1kUkYfTuQi/ICl64SEsZVAoPCBXGNjbHnJTmq9S4di2aYVH1gSX6NL4Ai8jBVoitclTzbU0oFPpCq6QRZoCijGoruE/JHbdHcNBkFtwnimQEEiwd0oWS1HGqTSYI1DK6c7bxUo6stfTCIF3MsOw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com; spf=pass smtp.mailfrom=163.com; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b=HRqLrWx1; arc=none smtp.client-ip=220.197.31.5 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=163.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b="HRqLrWx1" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=163.com; s=s110527; h=From:To:Subject:Date:Message-Id:MIME-Version; bh=2p URnNgagNSzAGeA4erfdSykdqLoBlr6NUXAPfZFKc4=; b=HRqLrWx1CqdtRCR/8K qE1J7R4uvLEcXqwLOsdTu1eaCOymcKNSERtvR/CPg8toUEGtU7ICnI11L7olS6r6 a0FhFSe/BUDTTcYVm3YwTCiLrX+woWOmoGTt19jtKKD1YwTD31HXLRvgfX2aFVQM dZ3ssTJMm1Mjw03rE/cO8apX4= Received: from zhaoxin-MS-7E12.. (unknown []) by gzsmtp3 (Coremail) with SMTP id PigvCgB3tQRjR4BqOFTVNA--.50445S8; Sat, 15 Aug 2026 19:03:08 +0800 (CST) From: Xin Zhao To: mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com Cc: linux-kernel@vger.kernel.org, Xin Zhao Subject: [PATCH 06/10] sched/fair: Modify active_load_balance_cpu_stop() to accommodate more scenarios Date: Sat, 15 Aug 2026 19:02:53 +0800 Message-Id: <20260815110257.124354-7-jackzxcui1989@163.com> X-Mailer: git-send-email 2.34.1 In-Reply-To: <20260815110257.124354-1-jackzxcui1989@163.com> References: <20260815110257.124354-1-jackzxcui1989@163.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: PigvCgB3tQRjR4BqOFTVNA--.50445S8 X-Coremail-Antispam: 1Uf129KBjvJXoWxWr47Gw1xtr4kGrWUAr1rtFb_yoW5KF48pr WUKw4agw4DJayUZ390yF4xur1IgwsxWw47Jan2qF4rCF15G3s5tw1F9r1fuF45uFZ5CFna ya1Dtw47CFWUtFUanT9S1TB71UUUUU7qnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDUYxBIdaVFxhVjvjDU0xZFpf9x0pR5DGnUUUUU= X-CM-SenderInfo: pmdfy650fxxiqzyzqiywtou0bp/xtbCwAynsWqAR2xxvAAA3W Content-Type: text/plain; charset="utf-8" Modify active_load_balance_cpu_stop() to make it more generalized, no longer limited to the context of finding the busiest rq in the sched_balance_rq() scenario to migrate to the dst rq. Change the variables that start with 'busiest' to start with 'src'. Additionally, adjust the reference to the active_balance variable from busiest_rq->active_balance to this_rq()->active_balance. With these changes, the CPU where the stop task is located can serve as either a dst CPU or a src CPU, laying the groundwork for the upcoming preempt active balance feature. Signed-off-by: Xin Zhao --- kernel/sched/fair.c | 30 +++++++++++++++--------------- 1 file changed, 15 insertions(+), 15 deletions(-) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index f10e709921fd..9ffd01717599 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -13796,33 +13796,33 @@ update_next_balance(struct sched_domain *sd, unsi= gned long *next_balance) =20 /* * active_load_balance_cpu_stop is run by the CPU stopper. It pushes - * running tasks off the busiest CPU onto idle CPUs. It requires at + * one running task off the src CPU onto dst CPU. It requires at * least 1 task to be running on each physical CPU where possible, and * avoids physical / logical imbalances. */ static int active_load_balance_cpu_stop(void *data) { - struct rq *busiest_rq =3D data; - int busiest_cpu =3D cpu_of(busiest_rq); - int target_cpu =3D busiest_rq->push_cpu; + struct rq *src_rq =3D data; + int src_cpu =3D cpu_of(src_rq); + int target_cpu =3D this_rq()->push_cpu; struct rq *target_rq =3D cpu_rq(target_cpu); struct sched_domain *sd; struct task_struct *p =3D NULL; struct rq_flags rf; =20 - rq_lock_irq(busiest_rq, &rf); + rq_lock_irq(src_rq, &rf); /* * Between queueing the stop-work and running it is a hole in which * CPUs can become inactive. We should not move tasks from or to * inactive CPUs. */ - if (!cpu_active(busiest_cpu) || !cpu_active(target_cpu)) + if (!cpu_active(src_cpu) || !cpu_active(target_cpu)) goto out_unlock; =20 - WARN_ON_ONCE(!busiest_rq->active_balance); + WARN_ON_ONCE(!this_rq()->active_balance); =20 /* Is there any task to move? */ - if (busiest_rq->nr_running <=3D 1) + if (src_rq->nr_running <=3D 1) goto out_unlock; =20 /* @@ -13830,12 +13830,12 @@ static int active_load_balance_cpu_stop(void *dat= a) * we need to fix it. Originally reported by * Bjorn Helgaas on a 128-CPU setup. */ - WARN_ON_ONCE(busiest_rq =3D=3D target_rq); + WARN_ON_ONCE(src_rq =3D=3D target_rq); =20 /* Search for an sd spanning us and the target CPU. */ rcu_read_lock(); for_each_domain(target_cpu, sd) { - if (cpumask_test_cpu(busiest_cpu, sched_domain_span(sd))) + if (cpumask_test_cpu(src_cpu, sched_domain_span(sd))) break; } =20 @@ -13844,14 +13844,14 @@ static int active_load_balance_cpu_stop(void *dat= a) .sd =3D sd, .dst_cpu =3D target_cpu, .dst_rq =3D target_rq, - .src_cpu =3D busiest_rq->cpu, - .src_rq =3D busiest_rq, + .src_cpu =3D src_rq->cpu, + .src_rq =3D src_rq, .idle =3D CPU_IDLE, .flags =3D LBF_ACTIVE_LB, }; =20 schedstat_inc(sd->alb_count); - update_rq_clock(busiest_rq); + update_rq_clock(src_rq); =20 p =3D detach_one_task(&env); if (p) { @@ -13864,13 +13864,13 @@ static int active_load_balance_cpu_stop(void *dat= a) } rcu_read_unlock(); out_unlock: - rq_unlock(busiest_rq, &rf); + rq_unlock(src_rq, &rf); =20 if (p) attach_one_task(target_rq, p); =20 local_irq_enable(); - busiest_rq->active_balance =3D 0; + this_rq()->active_balance =3D 0; =20 return 0; } --=20 2.34.1 From nobody Mon Sep 28 23:12:43 2026 Received: from m16.mail.163.com (m16.mail.163.com [117.135.210.3]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5F6F93ECBE5 for ; Sat, 15 Aug 2026 11:04:41 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=117.135.210.3 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786791891; cv=none; b=ExkJLJhpR/yAMKnHIhtsrfKNXO8LJCaNrFDzkmljtXbOnUVlQGFKll8xnvLIPamvOuPWyoyOnhJdNyYUCxDpEJdDXL7cXI3PHtsZiw0F/ziRSBZ9jO6UKjJFgBurWnDNdW6VkZLbM31k+jpYeEx12IrVBVMxM5YtPijdtWg3y4k= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786791891; c=relaxed/simple; bh=wrEPrEhzPywTkV5zFcDwTZrPWZGDw0UPGGLR4mG1UIw=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=DY1+Z2ac7ed2opEaibzNBmvl4L81/a/fuKt2RbJ3ORtNdiyh5ZxzHmemUi3kz6nReCa+Wj0uAzr7t4PZNMyTc8v8YOvH9UTxYm+bdGw6zF+SKcK1NN7/K1BEzqy2VT07n7/AIV+hpOF9XyQhvYhrdS/LsCEoKkXewGcYKUSSulw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com; spf=pass smtp.mailfrom=163.com; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b=fZ24Uhah; arc=none smtp.client-ip=117.135.210.3 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=163.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b="fZ24Uhah" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=163.com; s=s110527; h=From:To:Subject:Date:Message-Id:MIME-Version; bh=tX BLjoyVcqoSnCe5hTKnzGr+xFkpChqXdUEH+D46IKk=; b=fZ24UhahSuTAiawy7b +JhOL9/MV41RRAg+QHVwLZf2Cv8G0igjQO+AS2+ZGeXVwY1FIDDgicSjDb5jZ6AA xAmkFvSY0f0pZ0N3+dq/n2Ls3cMA0IXspzEaPUU4IyhhHE2X1fOIa3cifjo2iimg G8jXhplpSPX2gwzDNVLpvNHO4= Received: from zhaoxin-MS-7E12.. (unknown []) by gzsmtp3 (Coremail) with SMTP id PigvCgB3tQRjR4BqOFTVNA--.50445S9; Sat, 15 Aug 2026 19:03:09 +0800 (CST) From: Xin Zhao To: mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com Cc: linux-kernel@vger.kernel.org, Xin Zhao Subject: [PATCH 07/10] sched/fair: Trigger active balance if a CFS task is preempted when LB_PROMOTE Date: Sat, 15 Aug 2026 19:02:54 +0800 Message-Id: <20260815110257.124354-8-jackzxcui1989@163.com> X-Mailer: git-send-email 2.34.1 In-Reply-To: <20260815110257.124354-1-jackzxcui1989@163.com> References: <20260815110257.124354-1-jackzxcui1989@163.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: PigvCgB3tQRjR4BqOFTVNA--.50445S9 X-Coremail-Antispam: 1Uf129KBjvJXoW3XF4fJry5KF15Zry7CFW3Jrb_yoW7trWfpr Wvya4rJa1DJ3W2q3ySkr48Zr13Wwn5Jayxtan7Jw4rAF15K34rt3Zavw13Wr45Crn5uF4a yr1jv3y2k3W8JF7anT9S1TB71UUUUU7qnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDUYxBIdaVFxhVjvjDU0xZFpf9x0zujgctUUUUU= X-CM-SenderInfo: pmdfy650fxxiqzyzqiywtou0bp/xtbCvw6nsWqAR25lJwAA3m Content-Type: text/plain; charset="utf-8" When LB_PROMOTE is on, in addition to the scenarios modified in the previous patches regarding .select_task_rq implement, preemption of CFS tasks also needs to be addressed. Once a CFS task is preempted by a real-time task, if that task is executing logic in a critical section that does not support priority inheritance, such as read-write locks or read-write semaphores, it may lead to performance issues. By adding this checkpoint for when a CFS task is preempted, we can migrate the task to an idle CPU, thus alleviating such performance problems. Additionally, the patch set for the LB_PROMOTE feature does not alter the logic of task_hot(), so this patch will not significantly increase the frequency of load balance migration. The newly introduced function preempt_active_balance() will be triggered in __schedule() when it detects that the previous task has been preempted while the LB_PROMOTE feature is enabled. preempt_active_balance() will first check whether the currently preempted task has other CPUs available to run on. If there are idle CPUs available, the preempted task will be migrated to one of those idle CPUs. The implementation of preempt_active_balance() considers the desire not to introduce additional "holes" of rq locks, so the migration triggering action is deferred to the balance_callback which is trigger_preempt_alb(). trigger_preempt_alb() will issue a stop work to allow the idle target CPU to perform the migration. After the modifications made by the previous patches, we can reuse active_load_balance_cpu_stop() to assist in this migration action. In preempt_active_balance(), we use list_move_tail to move the currently preempted task to position where detach_one_task() will first traverse, allowing active_load_balance_cpu_stop() to prioritize finding the task. Signed-off-by: Xin Zhao --- kernel/sched/core.c | 3 ++ kernel/sched/fair.c | 80 ++++++++++++++++++++++++++++++++++++++++++++ kernel/sched/sched.h | 2 ++ 3 files changed, 85 insertions(+) diff --git a/kernel/sched/core.c b/kernel/sched/core.c index 2e7cde033a31..757ea5ef303d 100644 --- a/kernel/sched/core.c +++ b/kernel/sched/core.c @@ -7227,6 +7227,9 @@ static void __sched notrace __schedule(int sched_mode) =20 trace_sched_switch(preempt, prev, next, prev_state); =20 + if (sched_feat(LB_PROMOTE) && preempt) + preempt_active_balance(prev); + /* Also unlocks the rq: */ rq =3D context_switch(rq, prev, next, &rf); } else { diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 9ffd01717599..d2b538abdf3e 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -13875,6 +13875,86 @@ static int active_load_balance_cpu_stop(void *data) return 0; } =20 +static DEFINE_PER_CPU(struct balance_callback, preempt_push_head); +static DEFINE_PER_CPU(int, push_cpu); + +static void trigger_preempt_alb(struct rq *rq) +{ + unsigned long flags; + int active_balance =3D 0; + int src_cpu =3D rq->cpu; + int idle_cpu =3D per_cpu(push_cpu, src_cpu); + struct rq *dst_rq =3D cpu_rq(idle_cpu); + + preempt_disable(); + raw_spin_rq_unlock(rq); + raw_spin_rq_lock_irqsave(dst_rq, flags); + if (dst_rq->active_balance) + goto unlock_rq; + dst_rq->active_balance =3D 1; + dst_rq->push_cpu =3D idle_cpu; + active_balance =3D 1; + +unlock_rq: + raw_spin_rq_unlock_irqrestore(dst_rq, flags); + if (active_balance) { + stop_one_cpu_nowait(idle_cpu, active_load_balance_cpu_stop, rq, + &dst_rq->active_balance_work); + } + preempt_enable(); + raw_spin_rq_lock(rq); +} + +static void queue_preempt_alb_callback(struct rq *rq, int dst_cpu) +{ + per_cpu(push_cpu, rq->cpu) =3D dst_cpu; + queue_balance_callback(rq, &per_cpu(preempt_push_head, rq->cpu), trigger_= preempt_alb); +} + +void preempt_active_balance(struct task_struct *prev) +{ + struct rq *this_rq =3D this_rq(); + int this_cpu =3D smp_processor_id(), cpu; + int start =3D nr_cpu_ids, idle_cpu =3D nr_cpu_ids; + struct sched_domain *sd; + + if (unlikely(prev->sched_class !=3D &fair_sched_class)) + return; + + if (unlikely(prev->migration_disabled)) + return; + + cpu =3D prev->recent_used_cpu; + if (cpu !=3D this_cpu && available_idle_cpu(cpu) && !cpu_rq(cpu)->active_= balance) { + idle_cpu =3D cpu; + goto active_balance; + } + + rcu_read_lock(); + + sd =3D rcu_dereference(per_cpu(sd_llc, this_cpu)); + if (sd) + start =3D cpumask_first_and(sched_domain_span(sd), prev->cpus_ptr); + if (start >=3D nr_cpu_ids) + start =3D cpumask_first(prev->cpus_ptr); + + for_each_cpu_wrap(cpu, prev->cpus_ptr, start) { + if (cpu !=3D this_cpu && available_idle_cpu(cpu) && !cpu_rq(cpu)->active= _balance) { + idle_cpu =3D cpu; + goto unlock; + } + } + +unlock: + rcu_read_unlock(); + if (idle_cpu =3D=3D nr_cpu_ids) + return; + +active_balance: + list_move_tail(&prev->se.group_node, &this_rq->cfs_tasks); + queue_preempt_alb_callback(this_rq, idle_cpu); +} + /* * Scale the max sched_balance_rq interval with the number of CPUs in the = system. * This trades load-balance latency on larger machines for less cross talk. diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h index 26ae13c86b69..11848708e5ce 100644 --- a/kernel/sched/sched.h +++ b/kernel/sched/sched.h @@ -4193,6 +4193,8 @@ extern struct balance_callback *splice_balance_callba= cks(struct rq *rq); extern void __balance_callbacks(struct rq *rq, struct rq_flags *rf); extern void balance_callbacks(struct rq *rq, struct balance_callback *head= ); =20 +extern void preempt_active_balance(struct task_struct *prev); + /* * The 'sched_change' pattern is the safe, easy and slow way of changing a * task's scheduling properties. It dequeues a task, such that the schedul= er --=20 2.34.1 From nobody Mon Sep 28 23:12:43 2026 Received: from m16.mail.163.com (m16.mail.163.com [117.135.210.2]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 66BC63A875B for ; Sat, 15 Aug 2026 11:04:38 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=117.135.210.2 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786791880; cv=none; b=Tnww5hOY6YaKcJtQ5jJw/zi27CsWo/Wc8ePtjceU7EsMfg7hCQDuKexIcTnS7JxAqVmc0OuRTFOwsc5IUByEDooeh9rsixbdMej2IUK85sIwKsccK3scBfCuzctgzIKKpCI6ELz1daJ2bvAX9WqprmebSsiBgDmSWOmPRLC1UY4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786791880; c=relaxed/simple; bh=tXOb2gM5KmdToeaqxMTKhvq6W4aDrf8KUKnU0DE+TKg=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=uZKTVOtJN0z3Pirukj+8Fp8XAIKdLCQlXq+TR3tlsmx0zKisPc6llufem5YuxLRtGVnX+lymKyuVfzdxSqH326Gv39kWxOiyrWHjicfpQGEhDhaIYnTir/5dGNoFy4wCefKtQZaYUg5afEb+NVsTCQpUGvdZ+MRJBgZ4C4pK5WI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com; spf=pass smtp.mailfrom=163.com; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b=hc+PxIXL; arc=none smtp.client-ip=117.135.210.2 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=163.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b="hc+PxIXL" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=163.com; s=s110527; h=From:To:Subject:Date:Message-Id:MIME-Version; bh=Iz Q9B9qyBge0iH1jK1v8qnI3vqWt1L1DDgzup5RG3Pw=; b=hc+PxIXLHxMxYTphUr tUlt+7Ml46D/n2WvhegSLLou2MdtQf/5FSIghb6lMJTXzUhSFQpsokXwcgEkKNPY iqan4K3zeEuv7j3Mk7VZYCtdgwt8kBxEEpT+h+EWKl//ktGNs9uaIxpn2aVgHap9 r9iXDjhDp7PzmvHPeMsIdP8WE= Received: from zhaoxin-MS-7E12.. (unknown []) by gzsmtp3 (Coremail) with SMTP id PigvCgB3tQRjR4BqOFTVNA--.50445S10; Sat, 15 Aug 2026 19:03:11 +0800 (CST) From: Xin Zhao To: mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com Cc: linux-kernel@vger.kernel.org, Xin Zhao Subject: [PATCH 08/10] sched/fair: Do not check avg_idle to prematurely exit newly idle when LB_PROMOTE Date: Sat, 15 Aug 2026 19:02:55 +0800 Message-Id: <20260815110257.124354-9-jackzxcui1989@163.com> X-Mailer: git-send-email 2.34.1 In-Reply-To: <20260815110257.124354-1-jackzxcui1989@163.com> References: <20260815110257.124354-1-jackzxcui1989@163.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: PigvCgB3tQRjR4BqOFTVNA--.50445S10 X-Coremail-Antispam: 1Uf129KBjvJXoW7WF4fCw4xJFW8Xry8Xr18Krg_yoW8Xw4fpr ZYyan5Cr1qq3Z5Ja4kAF4kWw1agwsrJa43WF1Iy343Jr15J34Fv3WSqa43WFWfCryrCF4a vr1jqw12k3W0grJanT9S1TB71UUUUU7qnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDUYxBIdaVFxhVjvjDU0xZFpf9x0pCjqAgUUUUU= X-CM-SenderInfo: pmdfy650fxxiqzyzqiywtou0bp/xtbCwA+nsWqAR29x2wAA3x Content-Type: text/plain; charset="utf-8" When LB_PROMOTE is on, there are high real-time requirements, we should avoid prematurely exiting in sched_balance_newidle() due to a short avg_idle. This is because after any newly idle state, if no tasks are found to run, there may be a long-term lack of new incoming tasks to wake up on that CPU. Meanwhile, there may be tasks waiting to run on other cores, leading to unreasonable CPU idleness. The definition of unreasonable CPU idleness can be found in the commit-log of the previous patch that introduced the LB_PROMOTE feature. Signed-off-by: Xin Zhao --- kernel/sched/fair.c | 5 +++-- 1 file changed, 3 insertions(+), 2 deletions(-) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index d2b538abdf3e..8a1d2763a923 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -14694,7 +14694,7 @@ static int sched_balance_newidle(struct rq *this_rq= , struct rq_flags *rf) goto out; =20 if (!get_rd_overloaded(this_rq->rd) || - this_rq->avg_idle < sd->max_newidle_lb_cost) { + (!sched_feat(LB_PROMOTE) && this_rq->avg_idle < sd->max_newidle_lb_co= st)) { =20 update_next_balance(sd, &next_balance); goto out; @@ -14716,7 +14716,8 @@ static int sched_balance_newidle(struct rq *this_rq= , struct rq_flags *rf) =20 update_next_balance(sd, &next_balance); =20 - if (this_rq->avg_idle < curr_cost + sd->max_newidle_lb_cost) + if (!sched_feat(LB_PROMOTE) && + this_rq->avg_idle < curr_cost + sd->max_newidle_lb_cost) break; =20 if (sd->flags & SD_BALANCE_NEWIDLE) { --=20 2.34.1 From nobody Mon Sep 28 23:12:43 2026 Received: from m16.mail.163.com (m16.mail.163.com [117.135.210.4]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8C5BE363090 for ; Sat, 15 Aug 2026 11:04:36 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=117.135.210.4 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786791879; cv=none; b=eYY0Vn5p24tiC1sLhMK8xucoMjRMQqDkoE8EI5DyFKW8Nt0evx+XmV8cPVcQRJJP/5T8a2GhGaIYBA1PIEotILsI56e7wHdQUr0TsfzO+3YBIXOPBlf7p7kj6O22o71NiJinBRAhilf4J+EIuoeppvrWYp/xzmirDwofGkiSNV4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786791879; c=relaxed/simple; bh=KsIoKJWnFpTP1QqcLS2ahGyzsxSWbGwq020JU74xWZg=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=hHx2DK51QWi42SRiPqexneftNj4QJUUwRzu4xjcRPda/pAePndEDanPEQxf4F8K1j2ylWwDLhQmn3d6R4v8iiR1xNVNYT8AGQ4JC0vJF0bTrIhuwvpPt1j5wkFILO0wK9lxFvHsZix/jS1gPjiqwFW2BebxEvKVBC/5vXgnw5w0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com; spf=pass smtp.mailfrom=163.com; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b=WUlAj4l3; arc=none smtp.client-ip=117.135.210.4 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=163.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b="WUlAj4l3" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=163.com; s=s110527; h=From:To:Subject:Date:Message-Id:MIME-Version; bh=KT QU5wwOgoyx1VpyqCpFm3lgO2ZriMIQJJ0cehW7kzQ=; b=WUlAj4l3IXc62q6mw/ eWPzr0e4lLhdTrNPYtkwdjm3lYijUduwGrFceQKw2TL1UobHDB1wkFXW66N89hs/ HTwIlbY/qza44vAe0swb3TecXs3TSr+oaNsaGhKmJlDMP72pFNPEqz7Ak36eKYTd ferhWw+rfAVx2hdTmG7ckvjo8= Received: from zhaoxin-MS-7E12.. (unknown []) by gzsmtp3 (Coremail) with SMTP id PigvCgB3tQRjR4BqOFTVNA--.50445S11; Sat, 15 Aug 2026 19:03:12 +0800 (CST) From: Xin Zhao To: mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com Cc: linux-kernel@vger.kernel.org, Xin Zhao Subject: [PATCH 09/10] sched/fair: Not goto more_balance if newly idle and has pending task when LBF_NEED_BREAK Date: Sat, 15 Aug 2026 19:02:56 +0800 Message-Id: <20260815110257.124354-10-jackzxcui1989@163.com> X-Mailer: git-send-email 2.34.1 In-Reply-To: <20260815110257.124354-1-jackzxcui1989@163.com> References: <20260815110257.124354-1-jackzxcui1989@163.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: PigvCgB3tQRjR4BqOFTVNA--.50445S11 X-Coremail-Antispam: 1Uf129KBjvJXoW7KF43KF4rWF1DWF4rXw4UXFb_yoW8Jw4xpr Wv9F45Za1qq345A39ayF48ur15uw1Sk398uFZrArWfJrn0qFWjvFZYga9xWFWjvFykA3Wr ZF1jg3429340yF7anT9S1TB71UUUUU7qnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDUYxBIdaVFxhVjvjDU0xZFpf9x0zujgctUUUUU= X-CM-SenderInfo: pmdfy650fxxiqzyzqiywtou0bp/xtbC5xCosmqAR3CLeQAA3C Content-Type: text/plain; charset="utf-8" When LBF_NEED_BREAK flag is set in env.flags, no longer unconditionally go to more_balance. Instead, we exclude the case when it is newly idle and there are pending tasks. This helps avoid unnecessary CPU wastage caused by repeatedly going to more_balance when the task load is too high during sched_balance_rq(). In another 'goto more_balance' case when LBF_DST_PINNED flag is set in env.flags, we do not need to add the check. Because LBF_DST_PINNED flag is only set within can_migrate_task(). Before setting LBF_DST_PINNED flag in can_migrate_task(), there is a check to see if it is newly idle. can_migrate_task() will exit without setting the LBF_DST_PINNED flag if it is newly idle. Signed-off-by: Xin Zhao --- kernel/sched/fair.c | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 8a1d2763a923..1ae351ea5949 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -13561,7 +13561,9 @@ static int sched_balance_rq(int this_cpu, struct rq= *this_rq, =20 if (env.flags & LBF_NEED_BREAK) { env.flags &=3D ~LBF_NEED_BREAK; - goto more_balance; + if (!(env.idle =3D=3D CPU_NEWLY_IDLE && + (env.dst_rq->nr_running > 0 || env.dst_rq->ttwu_pending))) + goto more_balance; } =20 /* --=20 2.34.1 From nobody Mon Sep 28 23:12:43 2026 Received: from m16.mail.163.com (m16.mail.163.com [220.197.31.2]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 91A9C1DE8BF for ; Sat, 15 Aug 2026 11:04:32 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=220.197.31.2 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786791877; cv=none; b=XxbhZucRoQ7eKjA4RqbdiHJmKCxOUrzGqWRGbqKH08LnhWfgnUT0z9E0LVjx/f0aPjNZcYyz+7I9Y8rtZAumS5qYF2Ps1TViB2fqFVY0u7+x0pdwe0fBIa97Mgc3Q+5AwWw/nteqpCoA5z4bVOj+S3g7XHMDExssk3HUDPabkK0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786791877; c=relaxed/simple; bh=JyuRVYUXENtFmSue4479W+r7Lt9Y/8EY2DF4juzQW6k=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=VHrXWECodBPSstdEhwhdiFvp8TEMi3oG6SLvr7qz/rOZ0do7IrARJuE8qfcHGPxzOtPyzYcpPytpDdimEj46qbyuC079NEL0M4l4A2/Y6Bln8PC//EGOtMW+lwvmcO7kywPFkpfdW3BBvo32bF3fLFmbLMxJ4C8s2MiyxwSzhOY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com; spf=pass smtp.mailfrom=163.com; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b=mTUlPXvS; arc=none smtp.client-ip=220.197.31.2 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=163.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=163.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=163.com header.i=@163.com header.b="mTUlPXvS" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=163.com; s=s110527; h=From:To:Subject:Date:Message-Id:MIME-Version; bh=L7 zCxf1hzQ1p+u7zu+dIehAq8l0mgUUnBoS0+XuNGiQ=; b=mTUlPXvSMfT/IPcxPT X/KjT/UmyjUaHu0x+hscN/yXjrbZVrFzjyFSd+Sj6iUR9bQgRGd/MwsA4HoG6XP/ tfJlaGmO8Qm+AqYqGTI/ErN4tVBXpV0HMlDjs/5kPOMhvv3XzPQGoybJgZ/Bej/o Hvx+++iTfju9YjmaFXL82Vl88= Received: from zhaoxin-MS-7E12.. (unknown []) by gzsmtp3 (Coremail) with SMTP id PigvCgB3tQRjR4BqOFTVNA--.50445S12; Sat, 15 Aug 2026 19:03:13 +0800 (CST) From: Xin Zhao To: mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com Cc: linux-kernel@vger.kernel.org, Xin Zhao Subject: [PATCH 10/10] sched/fair: Strive to find a task to migrate if newly idle when LB_PROMOTE Date: Sat, 15 Aug 2026 19:02:57 +0800 Message-Id: <20260815110257.124354-11-jackzxcui1989@163.com> X-Mailer: git-send-email 2.34.1 In-Reply-To: <20260815110257.124354-1-jackzxcui1989@163.com> References: <20260815110257.124354-1-jackzxcui1989@163.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-CM-TRANSID: PigvCgB3tQRjR4BqOFTVNA--.50445S12 X-Coremail-Antispam: 1Uf129KBjvJXoWxCr1DWw47WryDtw43GrWDXFb_yoWrAr47pr ZY9a1rKa1Dtw13t3sIkFsrZr1agws7Xr47JFZ7Jr1fCr45J34aqrnaqa43AF4rurs5Zr1a vrnrKw1j9w17trJanT9S1TB71UUUUU7qnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDUYxBIdaVFxhVjvjDU0xZFpf9x0zKjg4xUUUUU= X-CM-SenderInfo: pmdfy650fxxiqzyzqiywtou0bp/xtbC5xGosmqAR3GLigAA3x Content-Type: text/plain; charset="utf-8" This patch is a core modification within this patch set. The simple_find label in this patch implements a straightforward logic for finding a migration task. It allows any instances in sched_balance_rq() that are unable to find a migration task for various reasons to fallback to logic executing the simple_find label to find one before exiting newly idle process. This patch addresses the newly idle scenario by preventing early exit in cases when !ld_moved && !active_balance, effectively executing the logic of simple_find label. Testing has shown that the situations listed below, account for a significant proportion of early exits: 1. Failure in sched_balance_find_src_group 2. Failure in sched_balance_find_src_rq 3. ld_moved is 0 and active_balance has not been triggered Of course, even with this change, it is still possible that no migration task can be found. However, it at least ensures that all selectable CPUs within the sched_domain have been thoroughly traversed. Signed-off-by: Xin Zhao --- kernel/sched/fair.c | 38 +++++++++++++++++++++++++++++++++++++- 1 file changed, 37 insertions(+), 1 deletion(-) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 1ae351ea5949..10ec7bb9e18c 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -13470,6 +13470,8 @@ static int sched_balance_rq(int this_cpu, struct rq= *this_rq, struct rq *busiest; struct rq_flags rf; struct cpumask *cpus =3D this_cpu_cpumask_var_ptr(load_balance_mask); + int cpu; + bool sfind =3D false; struct lb_env env =3D { .sd =3D sd, .dst_cpu =3D this_cpu, @@ -13504,15 +13506,41 @@ static int sched_balance_rq(int this_cpu, struct = rq *this_rq, group =3D sched_balance_find_src_group(&env); if (!group) { schedstat_inc(sd->lb_nobusyg[idle]); + if (sched_feat(LB_PROMOTE)) + goto simple_find; goto out_balanced; } =20 busiest =3D sched_balance_find_src_rq(&env, group); if (!busiest) { schedstat_inc(sd->lb_nobusyq[idle]); + if (sched_feat(LB_PROMOTE)) + goto simple_find; goto out_balanced; } + goto begin_balance; + +simple_find: + if (env.idle !=3D CPU_NEWLY_IDLE || + (env.dst_rq->nr_running > 0 || env.dst_rq->ttwu_pending)) + goto out_balanced; =20 + sfind =3D true; + env.migration_type =3D migrate_task; + env.imbalance =3D 1; + + for_each_cpu_andnot(cpu, env.cpus, env.dst_grpmask) { + busiest =3D cpu_rq(cpu); + if (busiest->nr_running <=3D 1) { + __cpumask_clear_cpu(cpu, cpus); + continue; + } + break; + } + if (cpu >=3D nr_cpu_ids) + goto out_balanced; + +begin_balance: WARN_ON_ONCE(busiest =3D=3D env.dst_rq); =20 update_lb_imbalance_stat(&env, sd, idle); @@ -13615,6 +13643,7 @@ static int sched_balance_rq(int this_cpu, struct rq= *this_rq, =20 /* All tasks on this runqueue were pinned by CPU affinity */ if (unlikely(env.flags & LBF_ALL_PINNED)) { +check_redo: __cpumask_clear_cpu(cpu_of(busiest), cpus); /* * Attempting to continue load balancing at the current @@ -13627,6 +13656,8 @@ static int sched_balance_rq(int this_cpu, struct rq= *this_rq, if (!cpumask_subset(cpus, env.dst_grpmask)) { env.loop =3D 0; env.loop_break =3D SCHED_NR_MIGRATE_BREAK; + if (sfind) + goto simple_find; goto redo; } goto out_all_pinned; @@ -13668,8 +13699,11 @@ static int sched_balance_rq(int this_cpu, struct r= q *this_rq, * if the curr task on busiest CPU can't be * moved to this_cpu: */ - if (!cpumask_test_cpu(this_cpu, busiest->curr->cpus_ptr)) + if (!cpumask_test_cpu(this_cpu, busiest->curr->cpus_ptr)) { + if (sched_feat(LB_PROMOTE) && env.idle =3D=3D CPU_NEWLY_IDLE) + goto check_redo; goto out_one_pinned; + } =20 /* Record that we found at least one task that could run on this_cpu */ env.flags &=3D ~LBF_ALL_PINNED; @@ -13703,6 +13737,8 @@ static int sched_balance_rq(int this_cpu, struct rq= *this_rq, preempt_enable(); =20 out_unbalanced: + if (sched_feat(LB_PROMOTE) && !active_balance && env.idle =3D=3D CPU_NEWL= Y_IDLE) + goto check_redo; /* We were unbalanced, so reset the balancing interval */ sd->balance_interval =3D sd->min_interval; goto out; --=20 2.34.1