From nobody Sat Jul 25 23:05:54 2026 Received: from szxga04-in.huawei.com (szxga04-in.huawei.com [45.249.212.190]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AF6D33A4F4A for ; Sun, 12 Jul 2026 11:55:13 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=45.249.212.190 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783857318; cv=none; b=o3dOcPGhHPO7MjewDSfBVioag4dQ0/cPjyGLYIXteu8GDERFRr6QQ/Y4SWMw7aOmgJlgfUsgBIqOc6X3GgKxkkKZ1DbgIEEzSqsmATWfMKfGCiHjYikR4QJ2APJw1n6BU58VtFRFCeTMwnlXksa9SSUVPcykSL2sjJ4TOV6EkV0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783857318; c=relaxed/simple; bh=/MTZyxZOqjnjfojr8JzfEI0Nwn9m+7fTkMd5F/1aAug=; h=From:To:CC:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=QTbo05fsmN2kNmrnEn9adAE3wxpckmCUgzO07S6/5lXabQNQpXRlEfqiuDZtwWgzty80x7kSXhkOqlVlIYTy9LmnOjAOLPBZDPXPFGEziYjZuvBmFRBNFuSHHUIJePQmoW7rwS+THFgBKLKAU77LRQeVrc6/72qREI895RqZRWQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=huawei.com; spf=pass smtp.mailfrom=huawei.com; dkim=pass (1024-bit key) header.d=huawei.com header.i=@huawei.com header.b=onfGXPZJ; dkim=pass (1024-bit key) header.d=huawei.com header.i=@huawei.com header.b=onfGXPZJ; arc=none smtp.client-ip=45.249.212.190 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=huawei.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huawei.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=huawei.com header.i=@huawei.com header.b="onfGXPZJ"; dkim=pass (1024-bit key) header.d=huawei.com header.i=@huawei.com header.b="onfGXPZJ" dkim-signature: v=1; a=rsa-sha256; d=huawei.com; s=dkim; c=relaxed/relaxed; q=dns/txt; h=From; bh=ZVQuAvjfugwoB5xeaL/JmGYkINjodBsCxTEJ/awqoV4=; b=onfGXPZJuNy6QQNGInI4Cb3Y6dUt24W0r8nMuCC8SSHsn499coxHiATe1m45CDP7B3xAbq34g HIEBLX6zdUdrPlpilcyhIgIKTvpEuewvQsx1LIGNDTMS+9+HX3APvDQ9JC1QOCCtSbxZri97Yex HSxsSKNOdO6VKfGSKG2jstk= Received: from canpmsgout06.his.huawei.com (unknown [172.19.92.157]) by szxga04-in.huawei.com (SkyGuard) with ESMTPS id 4gykWb5nXXz126LrG for ; Sun, 12 Jul 2026 19:54:35 +0800 (CST) dkim-signature: v=1; a=rsa-sha256; d=huawei.com; s=dkim; c=relaxed/relaxed; q=dns/txt; h=From; bh=ZVQuAvjfugwoB5xeaL/JmGYkINjodBsCxTEJ/awqoV4=; b=onfGXPZJuNy6QQNGInI4Cb3Y6dUt24W0r8nMuCC8SSHsn499coxHiATe1m45CDP7B3xAbq34g HIEBLX6zdUdrPlpilcyhIgIKTvpEuewvQsx1LIGNDTMS+9+HX3APvDQ9JC1QOCCtSbxZri97Yex HSxsSKNOdO6VKfGSKG2jstk= Received: from mail.maildlp.com (unknown [172.19.163.104]) by canpmsgout06.his.huawei.com (SkyGuard) with ESMTPS id 4gykKJ5HBxzRhS1; Sun, 12 Jul 2026 19:45:40 +0800 (CST) Received: from kwepemr500016.china.huawei.com (unknown [7.202.195.68]) by mail.maildlp.com (Postfix) with ESMTPS id 0BC204058F; Sun, 12 Jul 2026 19:54:56 +0800 (CST) Received: from huawei.com (10.67.174.242) by kwepemr500016.china.huawei.com (7.202.195.68) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.1544.11; Sun, 12 Jul 2026 19:54:55 +0800 From: Chen Jinghuang To: CC: , , , , , , , , , , , , , , , , , , Subject: [PATCH] [Question] sched/fair: Task starvation with RUN_TO_PARITY_WAKEUP under group topologies Date: Sun, 12 Jul 2026 11:31:06 +0000 Message-ID: <20260712113106.2564384-1-chenjinghuang2@huawei.com> X-Mailer: git-send-email 2.34.1 In-Reply-To: <20240325060226.1540-2-kprateek.nayak@amd.com> References: <20240325060226.1540-2-kprateek.nayak@amd.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-ClientProxiedBy: kwepems100002.china.huawei.com (7.221.188.206) To kwepemr500016.china.huawei.com (7.202.195.68) Content-Type: text/plain; charset="utf-8" Hi all, While adapting the RUN_TO_PARITY_WAKEUP feature on top of mainline v7.2-rc2, I encountered a severe task starvation issue under specific cgroup topologi= es. Specifically, a running task completely hogs the CPU and prevents any preemption. Test Topology and Observed Behavior: - CPU 0: Bound with two tasks: A1 (non-sleeping task) from cgroup A, and B1 (frequent sleep/wake task) from cgroup B. - CPU 20: Bound with other tasks from cgroup A (A2, A3... which sleep/wake frequently) and task C1 (non-sleeping task) from cgroup C. - All cgroups maintain default shares. CPU 0 CPU 20 / Other CPUs +----------------------------+ +----------------------------+ | cgroup A cgroup B | | cgroup A cgroup C | | |- A1 (Run) |- B1 | | |- A2, A3... |- C1 | | (Non-sleeping) (50us S) | | (Freq S/W) (Non-sleeping)| | (10us W) | | | +----------------------------+ +----------------------------+ | v A2/A3 frequent sleep/wake -> Frequent reweight of gse(A) on CPU 0 -> Triggers vlag clamping -> Unidirectional avruntime drift -> gse->vruntime drops abnormally (while gse weight remains stable) -> protect_slice() returns true -> pick_eevdf always returns curr -> Task A1 constantly hogs the CPU According to function_graph traces, the frequent sleep/wake cycles of sibling tasks (A2/A3) on CPU 20 cause gse(A) on CPU 0 to undergo continuous reweighting via update_cfs_cgroup() -> reweight_entity(). This triggers the following chain reaction: 1. gse->vlag triggers the lag clamping logic in entity_lag(). 2. Due to the clamping limit, A1->vruntime drifts unidirectionally against avg_vruntime (i.e., gse->vruntime drops abnormally even though its weight remains stable). 3. protect_slice() constantly returns true, forcing pick_eevdf() to always return curr(A1) during wakeup preemption checks. 4. As a result, A1 indefinitely hogs CPU 0. Even in entity_tick(), update_curr() fails to trigger resched_curr(). Both wakeup and periodic preemptions end up picking curr. Analysis: Experimentally, either disabling RUN_TO_PARITY_WAKEUP or reverting the vlag clamping patches completely resolves this starvation issue. This indicates a problematic interaction between the extended runtime introduced by RUN_TO_PARITY_WAKEUP and the vlag clamping mechanism. While A1 runs beyond its expected share, frequent sibling-induced reweighting pushes vlag to its limit, pulling A1->vruntime artificially closer to avg_vruntime. However, A1->deadline and A1->vprot are still rescaled in rescale_entity(). This mathematical asymmetry effectively ensures that A1->vruntime < A1->vprot`always holds true. Consequently, protect_slice() permanently returns true, trapping the scheduler into believing A1 has not yet exhausted its promised slice, thereby completely blocking legitimate context switches. Discussion: I'd love to hear your thoughts on how we can fix this issue without losing the throughput gains of RUN_TO_PARITY_WAKEUP. Thanks all. Signed-off-by: Chen Jinghuang --- kernel/sched/fair.c | 18 +++++++++++++++++- 1 file changed, 17 insertions(+), 1 deletion(-) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index d78467ec6ee1..00869624d914 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -1157,7 +1157,23 @@ static struct sched_entity *pick_eevdf(struct cfs_rq= *cfs_rq, bool protect) return cfs_rq->next; } =20 - if (curr && (!curr->on_rq || !entity_eligible(cfs_rq, curr))) + if (curr && !curr->on_rq) + curr =3D NULL; + + /* + * When an entity with positive lag wakes up, it pushes the + * avg_vruntime of the runqueue backwards. This may causes the + * current entity to be ineligible soon into its run leading to + * wakeup preemption. + * + * To prevent such aggressive preemption of the current running + * entity during task wakeups, skip the eligibility check if the + * slice promised to the entity since its selection has not yet + * elapsed. + */ + if (curr && + !(sched_feat(RUN_TO_PARITY_WAKEUP) && protect && protect_slice(curr))= && + !entity_eligible(cfs_rq, curr)) curr =3D NULL; =20 if (curr && protect && protect_slice(curr)) --=20 2.34.1