From nobody Tue Sep 29 11:19:37 2026 Received: from galois.linutronix.de (Galois.linutronix.de [193.142.43.55]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6DF383B7760; Sat, 8 Aug 2026 09:44:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=193.142.43.55 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786182299; cv=none; b=pFnijWtw3GTfZI2W87tmWJI02k1D5ghkxzTILB5EomDqNIsjC3mDQycn2+Bmwg+vBxHV3Lqm1yZ7Rgz8f5ZXintVjCTM3avc0n06cXJuR713dGiss2OZLshuJl2NGzCqkPzAbaZ2rxAI6XzIc8QFMGZ9oAG2EUGbDFuhXGQb9XQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786182299; c=relaxed/simple; bh=aOpmZqT/N38hgnzdvJEYtNnL52/PfN/FRaFUQpLnGbU=; h=Date:From:To:Subject:Cc:In-Reply-To:References:MIME-Version: Message-ID:Content-Type; b=KsqKIrikQbk3+y/d2OA0YUa093gb44vv4eK/KK1Ygi345Nkux+c8u1ZvVpIhCcVrn2MEcGLRMGwKkhK8VwC7cBkbsQwDlkJYeN/8ypZWS/Of+XX958dzce+5IitxMGB1rGKClPthwA4BywkSwqh2yiLqdPtvxyeln0GZpmt4bvs= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linutronix.de; spf=pass smtp.mailfrom=linutronix.de; dkim=pass (2048-bit key) header.d=linutronix.de header.i=@linutronix.de header.b=ALmOTaTW; dkim=permerror (0-bit key) header.d=linutronix.de header.i=@linutronix.de header.b=rTEGy56D; arc=none smtp.client-ip=193.142.43.55 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linutronix.de Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linutronix.de Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=linutronix.de header.i=@linutronix.de header.b="ALmOTaTW"; dkim=permerror (0-bit key) header.d=linutronix.de header.i=@linutronix.de header.b="rTEGy56D" Date: Sat, 08 Aug 2026 09:44:53 -0000 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linutronix.de; s=2020; t=1786182295; h=from:from:sender:sender:reply-to:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=x0TsZhITlrI0jdH7PT3cy3Uy1J51dvpxFqCyJsNCgWA=; b=ALmOTaTWxa++vd+oTMqe6W4V2BhxfthsbdZTAqLfoOkKCYfG+cYLwjduPiKq8NjYKv6X6C yhVSAX+0WbDuQF2L7JclUpeH/Fr6njGPxZmZdsrViNaF8hnbvL2f+JSatsP6SyjXrqgR1V ISRlWEQjaKRS0rC6D3Rpp45SUq4aa8Z3XCc0/sidbHT9fo4n7prHTC2Cb+ttqgNIHpkooK dOX6aeBQ2ZmgW9+yKTGFOd52t0df3Jq3oTBf+T+8X0IhiQFh8A+dLYtpmQ4h8XhFkZOAk0 Be7v4+aFPAleN5J39R1hCQK48dmRDoz90xr6T4cVgWWelQ0UrQ3BqLqUIj/yEA== DKIM-Signature: v=1; a=ed25519-sha256; c=relaxed/relaxed; d=linutronix.de; s=2020e; t=1786182295; h=from:from:sender:sender:reply-to:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=x0TsZhITlrI0jdH7PT3cy3Uy1J51dvpxFqCyJsNCgWA=; b=rTEGy56DMvzekbAEWfe117/vho/mMgZsR7tbhYKGFIL5hEgs9RFGVkpJuNEmHxMVOn6c7D ciS8/6G+rbpYU3AA== From: "tip-bot2 for Andrea Righi" Sender: tip-bot2@linutronix.de Reply-to: linux-kernel@vger.kernel.org To: linux-tip-commits@vger.kernel.org Subject: [tip: sched/core] sched/fair: Prefer fully idle cores for NOHZ balancing Cc: Andrea Righi , "Peter Zijlstra (Intel)" , Mete Durlu , Vincent Guittot , x86@kernel.org, linux-kernel@vger.kernel.org In-Reply-To: <20260804151324.918020-1-arighi@nvidia.com> References: <20260804151324.918020-1-arighi@nvidia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Message-ID: <178618229388.442315.2879525055465246887.tip-bot2@tip-bot2> Robot-ID: Robot-Unsubscribe: Contact to get blacklisted from these emails Precedence: bulk Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable The following commit has been merged into the sched/core branch of tip: Commit-ID: 293f9611ae73564febc553935830074f0f300694 Gitweb: https://git.kernel.org/tip/293f9611ae73564febc553935830074f0= f300694 Author: Andrea Righi AuthorDate: Tue, 04 Aug 2026 17:13:24 +02:00 Committer: Peter Zijlstra CommitterDate: Fri, 07 Aug 2026 18:27:09 +02:00 sched/fair: Prefer fully idle cores for NOHZ balancing find_new_ilb() selects the first idle housekeeping CPU without considering whether another thread is running on the same physical core. On an SMT system, the idle load balancer can therefore activate both siblings even when another housekeeping CPU has an entirely idle core. On most SMT systems, this is not problematic because the idle load balancer is a short-lived activity and the transient wakeup of a sibling has negligible performance impact. However, this can be particularly costly on NVIDIA Olympus cores used in Vera. Briefly activating an otherwise idle sibling can reduce the performance available to the other sibling and this effect does not necessarily end once the activated sibling becomes idle: after the ILB finishes and its CPU enters WFI, full single-thread performance is restored only after the sibling has remained idle for a qualification interval (10 Ki cycles on the tested Vera system). Repeated short sibling wakeups can therefore sustain the interference even with little actual overlap. Prevent this by preferring an idle housekeeping CPU whose entire SMT core is idle. Retain the first idle CPU as a fallback when no fully idle core is available, so NOHZ balancing continues to make forward progress. Once a partially busy core has been examined, skip its remaining SMT siblings to avoid repeating the core-idle check on wide SMT systems. Tests performed using an ad hoc GEMM benchmark running one CPU-intensive task per SMT core within its CPU affinity mask improved from approximately 6.2 TFLOP/s to 9.4 TFLOP/s. Note that this preference may wake a fully idle physical core instead of using an idle sibling of an active core, potentially increasing ILB wakeup latency or energy consumption on some architectures. It may also scan additional CPUs before selecting the one to run the ILB. The selection falls back to the first idle CPU when no fully idle SMT core is available. Non-SMT systems continue to select the first idle housekeeping CPU. Signed-off-by: Andrea Righi Signed-off-by: Peter Zijlstra (Intel) Reviewed-by: Mete Durlu Reviewed-by: Vincent Guittot Link: https://patch.msgid.link/20260804151324.918020-1-arighi@nvidia.com --- kernel/sched/fair.c | 55 +++++++++++++++++++++++++++++++++++--------- 1 file changed, 44 insertions(+), 11 deletions(-) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index df8c9c2..a24dd20 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -13962,29 +13962,62 @@ static inline int on_null_domain(struct rq *rq) */ static inline int find_new_ilb(void) { - int this_cpu =3D smp_processor_id(); - const struct cpumask *hk_mask; - int ilb_cpu; + struct cpumask *ilb_cpus; + int ilb_cpu, fallback =3D -1; + + lockdep_assert_irqs_disabled(); + + /* + * Reuse the per-CPU select_rq_mask, which is protected from concurrent + * use on this CPU by having interrupts disabled. + */ + ilb_cpus =3D this_cpu_cpumask_var_ptr(select_rq_mask); + cpumask_and(ilb_cpus, nohz.idle_cpus_mask, + housekeeping_cpumask(HK_TYPE_KERNEL_NOISE)); + + for_each_cpu(ilb_cpu, ilb_cpus) { + if (!idle_cpu(ilb_cpu)) { + /* + * Once an idle fallback exists, a busy CPU proves that + * this core cannot be fully idle. Skip its siblings. + */ + if (sched_smt_active() && fallback >=3D 0) + cpumask_andnot(ilb_cpus, ilb_cpus, cpu_smt_mask(ilb_cpu)); + continue; + } =20 - hk_mask =3D housekeeping_cpumask(HK_TYPE_KERNEL_NOISE); + /* + * Running the idle load balancer on an idle sibling of a busy + * SMT core can reduce the capacity available to its sibling. Prefer + * a CPU whose entire core is idle, but retain the first idle CPU as + * a fallback so idle balancing can still make progress when no fully + * idle core exists. + */ + if (sched_smt_active() && !is_core_idle(ilb_cpu)) { + if (fallback < 0) + fallback =3D ilb_cpu; =20 - for_each_cpu_and(ilb_cpu, nohz.idle_cpus_mask, hk_mask) { - if (ilb_cpu =3D=3D this_cpu) + /* + * The core is not idle, so there is no need to check + * any of its other SMT siblings. + */ + cpumask_andnot(ilb_cpus, ilb_cpus, + cpu_smt_mask(ilb_cpu)); continue; + } =20 - if (idle_cpu(ilb_cpu)) - return ilb_cpu; + return ilb_cpu; } =20 - return -1; + return fallback; } =20 /* * Kick a CPU to do the NOHZ balancing, if it is time for it, via a cross-= CPU * SMP function call (IPI). * - * We pick the first idle CPU in the HK_TYPE_KERNEL_NOISE housekeeping set - * (if there is one). + * Prefer a CPU on a fully idle core in the HK_TYPE_KERNEL_NOISE housekeep= ing + * set. Fall back to the first idle CPU when no fully idle core exists. */ static void kick_ilb(unsigned int flags) {