From nobody Mon Sep 28 02:57:20 2026 Received: from CWXP265CU010.outbound.protection.outlook.com (mail-ukwestazon11022130.outbound.protection.outlook.com [52.101.101.130]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1CF80357CFA for ; Thu, 27 Aug 2026 22:18:18 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.101.130 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787869100; cv=fail; b=GDaknsVLKnPJMHuzI9RpEn+99exxHnVwHice70Wi5SnnVh4gFIqLGqgKWKj6IFLj1s69870J2PWsVwsdvnIJSg56G5S4bs/fQWVlT1arXAcUw0NULbOuxqAzcBWocUfl+4tUmq1MpGeaC/Mx7X4LRGpSrl76inOGeSHP1p6jPNM= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787869100; c=relaxed/simple; bh=815+FbrzLetzxB6CWo+FiLYNhY1aeYztOE4hQdqOkbs=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: Content-Type:MIME-Version; b=RXf6gar4OQrVFAdqOfzqCkORW1jscF2L1j2g2Znu/FAFy0HeEYzagOTSaqp1C8q23pgrj6fwyfUUL5iH5KAYmSZ8PSxUsehKT01wYUEfj5CSOTMgI0qT7rBgQl/dq6lF2oUm8yT7g7HtDZms6JQhy6mLtVYKeDkNO+NCU/Nz1LI= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=atomlin.com; spf=pass smtp.mailfrom=atomlin.com; arc=fail smtp.client-ip=52.101.101.130 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=atomlin.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=atomlin.com ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=KkbPKNZ5ZVQhfUU1xlv/15PwRvFYEZLKUzcKx3VPMn9W9IU8d4A+zhLQTA3n6FbKFw7Bcb3V9dBfVwudmzOF2+QM4sw4zTm/yB1JcmGRtw8vJK7RUUj8JY/YY1vk4BwMzvqXNPVpWh4ED7In7ikwrVHmB4F3QKBuADLYTpfBft0L5Yxr5oxyiIsaxsbi5k+yUghGCvoCM4YjunReb9qgwDiGYyvPb/zlpQOPXa+sX6NxqE43cL0ExLTw8flOi9o1LPRUiXDpHp04+DKfWEXJ8dxGfPMIEI7YLJBrILjs7b7E0VWwkmn7ILcBZZUkQw0H/wvV6rdD6api9i2Ew+NnxQ== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:MIME-Version; bh=wl+LjtooGTK+RDvKbW+Yyc1ghdvCxgm/m4ErHJC7Atc=; b=ioT6MYGamHGVRQHu0DKURhkUDQGBftl1zJq/XS2mGCz2ij/pRaZMff+FYdKZAuOQzvx362dQkrDkR7Y3OvYpH/xhz5gJ8hQVcoPvuCuHhstJKRVyUvxX3rFjw1/F7yT4kL7ji6pZ+t6GAhKyOUUCK0hMDYFQ1u+Lpn0GE6tMo4LTTxL/3BEd2wtDLwk1xeHiMwe4VsFnvmXv7MA0ICXKq7cdBjtQh45gtSEjSRGI0EA1uk28TmKv75Bqs+cBy8Cfhv/8O/g7L+ZiuzaDba7cZYIBCE/B4qdV1YWR6SavRSgzrP/Rb4e2MjtNIQMcqfxngJ/oJaZFr0WiHHPSoB/xEw== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass smtp.mailfrom=atomlin.com; dmarc=pass action=none header.from=atomlin.com; dkim=pass header.d=atomlin.com; arc=none Authentication-Results: dkim=none (message not signed) header.d=none;dmarc=none action=none header.from=atomlin.com; Received: from CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM (2603:10a6:400:183::5) by LO6P123MB7095.GBRP123.PROD.OUTLOOK.COM (2603:10a6:600:342::13) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.360.10; Thu, 27 Aug 2026 22:18:12 +0000 Received: from CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM ([fe80::cec4:77ab:262e:d230]) by CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM ([fe80::cec4:77ab:262e:d230%4]) with mapi id 15.21.0360.008; Thu, 27 Aug 2026 22:18:12 +0000 From: Aaron Tomlin To: mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org Cc: paulmck@kernel.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com, zhanxusheng1024@gmail.com, neelx@suse.com, atomlin@atomlin.com, chjohnst@mail.com, mproche@mail.com, sean@ashe.io, steve@abita.co, rishil1999@outlook.com, linux-kernel@vger.kernel.org Subject: [PATCH v9 1/6] sched: Annotate rq->rd with __rcu and update lockless readers Date: Thu, 27 Aug 2026 18:18:03 -0400 Message-ID: <20260827221809.988394-2-atomlin@atomlin.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260827221809.988394-1-atomlin@atomlin.com> References: <20260827221809.988394-1-atomlin@atomlin.com> Content-Transfer-Encoding: quoted-printable X-ClientProxiedBy: LO4P302CA0036.GBRP302.PROD.OUTLOOK.COM (2603:10a6:600:317::7) To CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM (2603:10a6:400:183::5) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: CWLP123MB6607:EE_|LO6P123MB7095:EE_ X-MS-Office365-Filtering-Correlation-Id: 75551f4c-e230-4e4d-f2c1-08df04891556 X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|376014|7416014|1800799024|366016|23010399003|10067099003|56012099006|5023799004|22082099003|18002099003; X-Microsoft-Antispam-Message-Info: ziYttlmd04/9zqDCyh3b3obvv68vs3+ybc8kx6apdnqWUr2gKnuMyEd6/uJOAz4QXA6bpAifEEOUJZyrZS1NZHe7TV2dTQzdMvEzo7d2hgEF7ow30+Da8M7VjYRxx7ka+DdbvMEs6MqNiPUJQFJ42vnx6vyIOWly6+OAgCwj1IDxi1mbBCAjOzsN1nSeJPPbVer2wFOusIu8iNsq7dFzdhTS5jrEPQ43MFyzj1wrEHeQK6u0gaxd4L1Obh1zZ4xgf3o7juFjBv2bNoqcIMLLbi9IcBl4OFue4s2aRm7UIiWAA/yEek9bJ1q+U6OL5UsCWXn728L/9XSKjNgu6fZPFUvh3GZaW5ZAif/JRwckuEO31J8/5P7fQvHi5XK45zeT9hSRJ/Qn/QDarbaD8aZnpdO+vhBtRvm8D5nuzG8tT43kpiHry2gZEEgW0+7GrpzvgRCZ1zBfxnvCTrnxRHt0dRLCrLCIgRgm9680rXQg+cyB7cwaFJDAAt/MswYP9pXfbzIgQHVbfS/NDsri3kuj8Atf4Ak/6l80RBwcE3xqHi2fSdzYrgCS2bjWtyVh1sOqucys5bfi2wg5LrVdp3zHJvsfwPC9FO8bytoebOTq3NtD0uBvYer9OuheJpMJdz/F/AGArcphAc5/wuGv2uwbCjZz2yB3H6KCRh5ItGggp7w= X-Forefront-Antispam-Report: CIP:255.255.255.255;CTRY:;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM;PTR:;CAT:NONE;SFS:(13230040)(376014)(7416014)(1800799024)(366016)(23010399003)(10067099003)(56012099006)(5023799004)(22082099003)(18002099003);DIR:OUT;SFP:1102; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: =?us-ascii?Q?/ps//cNuAd2s7Oli2iJmwSRrwg+vbxdUW1zN2HbAPaV8mA7ZxHoIyM8aFCfX?= =?us-ascii?Q?FzUOX2F+liyPsXrsHZaadBYt0pHdOJKkoi4QwxEy4TJd9O4fy7v5ocaawWW8?= =?us-ascii?Q?Lcabk5MAdTTbjaRVvlawDPyHeG78VUuxSRDEddKmkurrDD/Pdf5VYAy7txFN?= =?us-ascii?Q?xo5qwQ4Xk8gCAXzM6ZcXUn06mQquLayjXo/2z/XZZlOJdKi+rcXCHB+Gb2wN?= =?us-ascii?Q?7CDJqV7Vk/dQ0mgav7zX2Bww++OU5QpNJXmv+Y+WKh4T0r1eDyPnLuZpBepu?= =?us-ascii?Q?wPdLbZfQDy4ztZEnmDdVFmUuW/zTG3QSDPIaA8pvCf9cDMpGsvX0qF4A6zni?= =?us-ascii?Q?bhOcagxwQABMSSGU2aix+CkbYGvq88YEE08x48lg253X7t2xgOyYGLwA3Azm?= =?us-ascii?Q?ELvjJ9tQWFCwggc9b940f/LouNh46Esa2jzd77ZgSMHE/yHz9ixYidLBV2a8?= =?us-ascii?Q?Njoh1GDikyeZIWWwVNkBo/JEst52FUjA4zdNcJ1vsEL6bvZbEGNj3ZnbWrKQ?= =?us-ascii?Q?spToG13+acMUCKdE34Hj1om7xweIUMMoe8o2LMX/XWtcYqR/GAGKja4kHYjF?= =?us-ascii?Q?Tr80G6K5HiT2jOgDWn0ADSBAKVjupGZpH6nJDm/WLaXEreRFUVLRE2WSUdZN?= =?us-ascii?Q?ZSKsN8gchoXjEjyzRm+y3tX6K6wT7CM3y52egsIApo+/fk5TME3Zu4bWg2UK?= =?us-ascii?Q?+E8+Tx7eZyvlNGe1KowBwGtZYelRGN/D/YJ38xtkt8U+3QG0T+kfcQCQiTnw?= =?us-ascii?Q?bO3r1oIJ6Y9lW29TpRAXDhTY95P/0Tmha2CtS5vQdIIuxTdQKzC5Ow1G6lJK?= =?us-ascii?Q?B+/s767TfPj/7953nPQDzJFG5y0dsEIWv+PAUpmKsZASGnpIPdX5Y4f0hHc8?= =?us-ascii?Q?bhw7wNa8e3TN1lsUofpI4J1iau2kY1xWDeVYUmq3sMCpR8DKi8caRRgbXBwe?= =?us-ascii?Q?npJpCeLTmU0WwylIPyB9rs1smx61eTbUDYlbCI/KUjuPo+F8zHYs8Z9G/wnD?= =?us-ascii?Q?/7eYKPf32p4ZOy2NDHecS77w3Cfd28J/3Nw4usUU1hyF74JMxEPI/wTwwgLD?= =?us-ascii?Q?I4dCsjCabMQj7FJjqkxQvKOdDFmZVgONY23zvd0fNocpdKeEMSi/LbKRQdMT?= =?us-ascii?Q?laAnbxuA2c47LLBqWmZPCP8M+SpKrXh25BBe3jf9f+9d6eVnyBIxJMQHeo/0?= =?us-ascii?Q?A7dEq+0VoV33Z4HeuywtfyKjulY52r+f9ED7SiOTURuap1NhxcsK6+rpP8sp?= =?us-ascii?Q?nfynZ6DfMlt0xldvyWPhcXDTmlH/5Nab8RJesW4MRiT+JUDZ0ZcWCSkjBouw?= =?us-ascii?Q?00ASuV2IiNM0Febb/+B+9yEs3lNR/iAPCrphslGfec6HHm9PUTGSImXZpv4W?= =?us-ascii?Q?7Hm55lnS53Xg4qtDfF4M1nK5E9Qd/Hny/9i/7WWC/tISJuMO3juQ4DkZ33Z0?= =?us-ascii?Q?fwpMTXpTfbnudfCWEsEp1yzkxm4lUHx3Zz5zPprh8c6h+YactyEOITyoc1Ez?= =?us-ascii?Q?pjSViZxWf5hwqpY3b/XTECdr21KlJh1kTB12OgdEpaO7SQIa68leWkz7wWP0?= =?us-ascii?Q?DM6AlyIcsDng1G407kkmQX3R/WiNPj4V5AXfRkG472VTkXEFt1d5IFgO/31L?= =?us-ascii?Q?Ck1mrBe1zrJqubPMHz7AgZKkDsUqB5SK1uc3CxKEP/0hx7R+AAfAQpXBsBuJ?= =?us-ascii?Q?F9hGFj1lXszLl/PoAoA+i3ggavZGf8FmJ150uR7HnDpiPREM?= X-OriginatorOrg: atomlin.com X-MS-Exchange-CrossTenant-Network-Message-Id: 75551f4c-e230-4e4d-f2c1-08df04891556 X-MS-Exchange-CrossTenant-AuthSource: CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM X-MS-Exchange-CrossTenant-AuthAs: Internal X-MS-Exchange-CrossTenant-OriginalArrivalTime: 27 Aug 2026 22:18:12.6338 (UTC) X-MS-Exchange-CrossTenant-FromEntityHeader: Hosted X-MS-Exchange-CrossTenant-Id: e6a32402-7d7b-4830-9a2b-76945bbbcb57 X-MS-Exchange-CrossTenant-MailboxType: HOSTED X-MS-Exchange-CrossTenant-UserPrincipalName: oChetc7ogpVI/bEOw5OMQTTSDrTJAlXjnbUE02Od0UmSc7UZMEhOoUa30rxJy7oHFUivf00E4q1j0YBl1OFllw== X-MS-Exchange-Transport-CrossTenantHeadersStamped: LO6P123MB7095 Content-Type: text/plain; charset="utf-8" The root_domain pointer rd field in struct rq is updated dynamically using RCU, and its memory reclamation is deferred via call_rcu() in rq_attach_root(). However, struct rq's rd field was missing the __rcu compiler annotation, and several lockless readers across the scheduler subsystem accessed rq->rd directly without using RCU dereference primitives. Add the __rcu annotation to struct rq's rd field. For code clarity, introduce the rcu_dereference_root_domain() helper macro to validate access under sched_domains_mutex or active RCU-sched read-side critical sections. While mechanically identical to rcu_dereference_sched_domain(), defining rcu_dereference_root_domain() preserves the natural symmetry of rq->rd and rq->sd in struct rq, while keeping grep-ability straightforward. Update readers and accessors across kernel/sched/ to use rcu_dereference_root_domain(). This ensures proper memory ordering, enables Sparse static analysis validation, and avoids false Lockdep warnings under CONFIG_PROVE_RCU. Signed-off-by: Aaron Tomlin --- kernel/sched/core.c | 24 ++++++++----- kernel/sched/deadline.c | 77 ++++++++++++++++++++++------------------- kernel/sched/fair.c | 27 +++++++++------ kernel/sched/rt.c | 64 +++++++++++++++++++--------------- kernel/sched/sched.h | 7 ++-- kernel/sched/syscalls.c | 8 ++--- kernel/sched/topology.c | 11 +++--- 7 files changed, 126 insertions(+), 92 deletions(-) diff --git a/kernel/sched/core.c b/kernel/sched/core.c index 2e7cde033a31..9a12601d8c54 100644 --- a/kernel/sched/core.c +++ b/kernel/sched/core.c @@ -8547,8 +8547,10 @@ void set_rq_online(struct rq *rq) { if (!rq->online) { const struct sched_class *class; + struct root_domain *rd; =20 - cpumask_set_cpu(rq->cpu, rq->rd->online); + rd =3D rcu_dereference_root_domain(rq->rd); + cpumask_set_cpu(rq->cpu, rd->online); rq->online =3D 1; =20 for_each_class(class) { @@ -8562,6 +8564,7 @@ void set_rq_offline(struct rq *rq) { if (rq->online) { const struct sched_class *class; + struct root_domain *rd; =20 update_rq_clock(rq); for_each_class(class) { @@ -8569,7 +8572,8 @@ void set_rq_offline(struct rq *rq) class->rq_offline(rq); } =20 - cpumask_clear_cpu(rq->cpu, rq->rd->online); + rd =3D rcu_dereference_root_domain(rq->rd); + cpumask_clear_cpu(rq->cpu, rd->online); rq->online =3D 0; } } @@ -8577,10 +8581,12 @@ void set_rq_offline(struct rq *rq) static inline void sched_set_rq_online(struct rq *rq, int cpu) { struct rq_flags rf; + struct root_domain *rd; =20 rq_lock_irqsave(rq, &rf); - if (rq->rd) { - BUG_ON(!cpumask_test_cpu(cpu, rq->rd->span)); + rd =3D rcu_dereference_root_domain(rq->rd); + if (rd) { + BUG_ON(!cpumask_test_cpu(cpu, rd->span)); set_rq_online(rq); } rq_unlock_irqrestore(rq, &rf); @@ -8589,10 +8595,12 @@ static inline void sched_set_rq_online(struct rq *r= q, int cpu) static inline void sched_set_rq_offline(struct rq *rq, int cpu) { struct rq_flags rf; + struct root_domain *rd; =20 rq_lock_irqsave(rq, &rf); - if (rq->rd) { - BUG_ON(!cpumask_test_cpu(cpu, rq->rd->span)); + rd =3D rcu_dereference_root_domain(rq->rd); + if (rd) { + BUG_ON(!cpumask_test_cpu(cpu, rd->span)); set_rq_offline(rq); } rq_unlock_irqrestore(rq, &rf); @@ -9009,8 +9017,8 @@ void __init sched_init(void) #endif rq->next_class =3D &idle_sched_class; =20 - rq->sd =3D NULL; - rq->rd =3D NULL; + RCU_INIT_POINTER(rq->sd, NULL); + RCU_INIT_POINTER(rq->rd, NULL); rq->cpu_capacity =3D SCHED_CAPACITY_SCALE; rq->balance_callback =3D &balance_push_callback; rq->active_balance =3D 0; diff --git a/kernel/sched/deadline.c b/kernel/sched/deadline.c index 857dbe3519a8..507f056084bf 100644 --- a/kernel/sched/deadline.c +++ b/kernel/sched/deadline.c @@ -122,12 +122,12 @@ static inline struct dl_bw *dl_bw_of(int i) { RCU_LOCKDEP_WARN(!rcu_read_lock_sched_held(), "sched RCU must be held"); - return &cpu_rq(i)->rd->dl_bw; + return &rcu_dereference_root_domain(cpu_rq(i)->rd)->dl_bw; } =20 static inline int dl_bw_cpus(int i) { - struct root_domain *rd =3D cpu_rq(i)->rd; + struct root_domain *rd =3D rcu_dereference_root_domain(cpu_rq(i)->rd); =20 RCU_LOCKDEP_WARN(!rcu_read_lock_sched_held(), "sched RCU must be held"); @@ -156,16 +156,13 @@ static inline unsigned long dl_bw_capacity(int i) arch_scale_cpu_capacity(i) =3D=3D SCHED_CAPACITY_SCALE) { return dl_bw_cpus(i) << SCHED_CAPACITY_SHIFT; } else { - RCU_LOCKDEP_WARN(!rcu_read_lock_sched_held(), - "sched RCU must be held"); - - return __dl_bw_capacity(cpu_rq(i)->rd->span); + return __dl_bw_capacity(rcu_dereference_root_domain(cpu_rq(i)->rd)->span= ); } } =20 bool dl_bw_visited(int cpu, u64 cookie) { - struct root_domain *rd =3D cpu_rq(cpu)->rd; + struct root_domain *rd =3D rcu_dereference_root_domain(cpu_rq(cpu)->rd); =20 if (rd->visit_cookie =3D=3D cookie) return true; @@ -533,15 +530,18 @@ void init_dl_rq(struct dl_rq *dl_rq) =20 static inline int dl_overloaded(struct rq *rq) { - return atomic_read(&rq->rd->dlo_count); + return atomic_read(&rcu_dereference_root_domain(rq->rd)->dlo_count); } =20 static inline void dl_set_overload(struct rq *rq) { + struct root_domain *rd; + if (!rq->online) return; =20 - cpumask_set_cpu(rq->cpu, rq->rd->dlo_mask); + rd =3D rcu_dereference_root_domain(rq->rd); + cpumask_set_cpu(rq->cpu, rd->dlo_mask); /* * Must be visible before the overload count is * set (as in sched_rt.c). @@ -549,16 +549,19 @@ static inline void dl_set_overload(struct rq *rq) * Matched by the barrier in pull_dl_task(). */ smp_wmb(); - atomic_inc(&rq->rd->dlo_count); + atomic_inc(&rd->dlo_count); } =20 static inline void dl_clear_overload(struct rq *rq) { + struct root_domain *rd; + if (!rq->online) return; =20 - atomic_dec(&rq->rd->dlo_count); - cpumask_clear_cpu(rq->cpu, rq->rd->dlo_mask); + rd =3D rcu_dereference_root_domain(rq->rd); + atomic_dec(&rd->dlo_count); + cpumask_clear_cpu(rq->cpu, rd->dlo_mask); } =20 #define __node_2_pdl(node) \ @@ -699,14 +702,15 @@ static struct rq *dl_task_offline_migration(struct rq= *rq, struct task_struct *p * since p is still hanging out in the old (now moved to default) root * domain. */ - dl_b =3D &rq->rd->dl_bw; + dl_b =3D &rcu_dereference_root_domain(rq->rd)->dl_bw; raw_spin_lock(&dl_b->lock); - __dl_sub(dl_b, p->dl.dl_bw, cpumask_weight(rq->rd->span)); + __dl_sub(dl_b, p->dl.dl_bw, cpumask_weight(rcu_dereference_root_domain(rq= ->rd)->span)); raw_spin_unlock(&dl_b->lock); =20 - dl_b =3D &later_rq->rd->dl_bw; + dl_b =3D &rcu_dereference_root_domain(later_rq->rd)->dl_bw; raw_spin_lock(&dl_b->lock); - __dl_add(dl_b, p->dl.dl_bw, cpumask_weight(later_rq->rd->span)); + __dl_add(dl_b, p->dl.dl_bw, + cpumask_weight(rcu_dereference_root_domain(later_rq->rd)->span)); raw_spin_unlock(&dl_b->lock); =20 set_task_cpu(p, later_rq->cpu); @@ -2222,9 +2226,10 @@ static void inc_dl_deadline(struct dl_rq *dl_rq, u64= deadline) if (dl_rq->earliest_dl.curr =3D=3D 0 || dl_time_before(deadline, dl_rq->earliest_dl.curr)) { if (dl_rq->earliest_dl.curr =3D=3D 0) - cpupri_set(&rq->rd->cpupri, rq->cpu, CPUPRI_HIGHER); + cpupri_set(&rcu_dereference_root_domain(rq->rd)->cpupri, rq->cpu, + CPUPRI_HIGHER); dl_rq->earliest_dl.curr =3D deadline; - cpudl_set(&rq->rd->cpudl, rq->cpu, deadline); + cpudl_set(&rcu_dereference_root_domain(rq->rd)->cpudl, rq->cpu, deadline= ); } } =20 @@ -2239,14 +2244,15 @@ static void dec_dl_deadline(struct dl_rq *dl_rq, u6= 4 deadline) if (!dl_rq->dl_nr_running) { dl_rq->earliest_dl.curr =3D 0; dl_rq->earliest_dl.next =3D 0; - cpudl_clear(&rq->rd->cpudl, rq->cpu, rq->online); - cpupri_set(&rq->rd->cpupri, rq->cpu, rq->rt.highest_prio.curr); + cpudl_clear(&rcu_dereference_root_domain(rq->rd)->cpudl, rq->cpu, rq->on= line); + cpupri_set(&rcu_dereference_root_domain(rq->rd)->cpupri, rq->cpu, + rq->rt.highest_prio.curr); } else { struct rb_node *leftmost =3D rb_first_cached(&dl_rq->root); struct sched_dl_entity *entry =3D __node_2_dle(leftmost); =20 dl_rq->earliest_dl.curr =3D entry->deadline; - cpudl_set(&rq->rd->cpudl, rq->cpu, entry->deadline); + cpudl_set(&rcu_dereference_root_domain(rq->rd)->cpudl, rq->cpu, entry->d= eadline); } } =20 @@ -2686,12 +2692,14 @@ static void migrate_task_rq_dl(struct task_struct *= p, int new_cpu __maybe_unused =20 static void check_preempt_equal_dl(struct rq *rq, struct task_struct *p) { + struct root_domain *rd =3D rcu_dereference_root_domain(rq->rd); + /* * Current can't be migrated, useless to reschedule, * let's hope p can move out. */ if (rq->curr->nr_cpus_allowed =3D=3D 1 || - !cpudl_find(&rq->rd->cpudl, rq->donor, NULL)) + !cpudl_find(&rd->cpudl, rq->donor, NULL)) return; =20 /* @@ -2699,7 +2707,7 @@ static void check_preempt_equal_dl(struct rq *rq, str= uct task_struct *p) * see if it is pushed or pulled somewhere else. */ if (p->nr_cpus_allowed !=3D 1 && - cpudl_find(&rq->rd->cpudl, p, NULL)) + cpudl_find(&rd->cpudl, p, NULL)) return; =20 resched_curr(rq); @@ -2948,7 +2956,7 @@ static int find_later_rq(struct task_struct *task) * We have to consider system topology and task affinity * first, then we can look for a suitable CPU. */ - if (!cpudl_find(&task_rq(task)->rd->cpudl, task, later_mask)) + if (!cpudl_find(&rcu_dereference_root_domain(task_rq(task)->rd)->cpudl, t= ask, later_mask)) return -1; =20 /* @@ -3232,7 +3240,7 @@ static void pull_dl_task(struct rq *this_rq) */ smp_rmb(); =20 - for_each_cpu(cpu, this_rq->rd->dlo_mask) { + for_each_cpu(cpu, rcu_dereference_root_domain(this_rq->rd)->dlo_mask) { if (this_cpu =3D=3D cpu) continue; =20 @@ -3355,10 +3363,8 @@ static void set_cpus_allowed_dl(struct task_struct *= p, bool dl_task_needs_bw_move(struct task_struct *p, const struct cpumask *new_mask) { - if (!dl_task(p)) - return false; - - return !cpumask_intersects(task_rq(p)->rd->span, new_mask); + guard(rcu)(); + return !cpumask_intersects(rcu_dereference_root_domain(task_rq(p)->rd)->s= pan, new_mask); } =20 /* Assumes rq->lock is held */ @@ -3368,9 +3374,10 @@ static void rq_online_dl(struct rq *rq) dl_set_overload(rq); =20 if (rq->dl.dl_nr_running > 0) - cpudl_set(&rq->rd->cpudl, rq->cpu, rq->dl.earliest_dl.curr); + cpudl_set(&rcu_dereference_root_domain(rq->rd)->cpudl, rq->cpu, + rq->dl.earliest_dl.curr); else - cpudl_clear(&rq->rd->cpudl, rq->cpu, true); + cpudl_clear(&rcu_dereference_root_domain(rq->rd)->cpudl, rq->cpu, true); } =20 /* Assumes rq->lock is held */ @@ -3379,7 +3386,7 @@ static void rq_offline_dl(struct rq *rq) if (rq->dl.overloaded) dl_clear_overload(rq); =20 - cpudl_clear(&rq->rd->cpudl, rq->cpu, false); + cpudl_clear(&rcu_dereference_root_domain(rq->rd)->cpudl, rq->cpu, false); } =20 void __init init_sched_dl_class(void) @@ -3440,10 +3447,10 @@ void dl_add_task_root_domain(struct task_struct *p) cpu =3D cpumask_first_and(cpu_active_mask, msk); BUG_ON(cpu >=3D nr_cpu_ids); rq =3D cpu_rq(cpu); - dl_b =3D &rq->rd->dl_bw; + dl_b =3D &rcu_dereference_root_domain(rq->rd)->dl_bw; =20 raw_spin_lock(&dl_b->lock); - __dl_add(dl_b, p->dl.dl_bw, cpumask_weight(rq->rd->span)); + __dl_add(dl_b, p->dl.dl_bw, cpumask_weight(rcu_dereference_root_domain(rq= ->rd)->span)); raw_spin_unlock(&dl_b->lock); raw_spin_unlock_irqrestore(&p->pi_lock, rf.flags); } @@ -3504,7 +3511,7 @@ void dl_clear_root_domain(struct root_domain *rd) =20 void dl_clear_root_domain_cpu(int cpu) { - dl_clear_root_domain(cpu_rq(cpu)->rd); + dl_clear_root_domain(rcu_dereference_root_domain(cpu_rq(cpu)->rd)); } =20 static void switched_from_dl(struct rq *rq, struct task_struct *p) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index f79fcba4afec..687999312a7d 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -7865,13 +7865,14 @@ static inline void set_rd_overutilized(struct root_= domain *rd, bool flag) =20 static inline void check_update_overutilized_status(struct rq *rq) { + struct root_domain *rd =3D rcu_dereference_root_domain(rq->rd); + /* * overutilized field is used for load balancing decisions only * if energy aware scheduler is being used */ - - if (!is_rd_overutilized(rq->rd) && cpu_overutilized(rq->cpu)) - set_rd_overutilized(rq->rd, 1); + if (rd && !is_rd_overutilized(rd) && cpu_overutilized(rq->cpu)) + set_rd_overutilized(rd, 1); } =20 /* Runqueue only has SCHED_IDLE tasks enqueued */ @@ -9500,7 +9501,7 @@ static int find_energy_efficient_cpu(struct task_stru= ct *p, int prev_cpu) unsigned long prev_delta =3D ULONG_MAX, best_delta =3D ULONG_MAX; unsigned long p_util_min =3D uclamp_is_used() ? uclamp_eff_value(p, UCLAM= P_MIN) : 0; unsigned long p_util_max =3D uclamp_is_used() ? uclamp_eff_value(p, UCLAM= P_MAX) : 1024; - struct root_domain *rd =3D this_rq()->rd; + struct root_domain *rd =3D rcu_dereference_root_domain(this_rq()->rd); int cpu, best_energy_cpu, target =3D -1; int prev_fits =3D -1, best_fits =3D -1; unsigned long best_actual_cap =3D 0; @@ -9704,7 +9705,7 @@ select_task_rq_fair(struct task_struct *p, int prev_c= pu, int wake_flags) cpumask_test_cpu(cpu, p->cpus_ptr)) return cpu; =20 - if (!is_rd_overutilized(this_rq()->rd)) { + if (!is_rd_overutilized(rcu_dereference_root_domain(this_rq()->rd))) { new_cpu =3D find_energy_efficient_cpu(p, prev_cpu); if (new_cpu >=3D 0) return new_cpu; @@ -12690,13 +12691,15 @@ static inline void update_sd_lb_stats(struct lb_e= nv *env, struct sd_lb_stats *sd env->fbq_type =3D fbq_classify_group(&sds->busiest_stat); =20 if (!env->sd->parent) { + struct root_domain *rd =3D rcu_dereference_root_domain(env->dst_rq->rd); + /* update overload indicator if we are at root domain */ - set_rd_overloaded(env->dst_rq->rd, sg_overloaded); + set_rd_overloaded(rd, sg_overloaded); =20 /* Update over-utilization (tipping point, U >=3D 0) indicator */ - set_rd_overutilized(env->dst_rq->rd, sg_overutilized); + set_rd_overutilized(rd, sg_overutilized); } else if (sg_overutilized) { - set_rd_overutilized(env->dst_rq->rd, sg_overutilized); + set_rd_overutilized(rcu_dereference_root_domain(env->dst_rq->rd), sg_ove= rutilized); } =20 update_idle_cpu_scan(env, sum_util); @@ -12942,8 +12945,10 @@ static struct sched_group *sched_balance_find_src_= group(struct lb_env *env) if (busiest->group_type =3D=3D group_misfit_task) goto force_balance; =20 - if (!is_rd_overutilized(env->dst_rq->rd) && - rcu_dereference_all(env->dst_rq->rd->pd)) + struct root_domain *rd =3D rcu_dereference_root_domain(env->dst_rq->rd); + + if (rd && !is_rd_overutilized(rd) && + rcu_dereference_all(rd->pd)) goto out_balanced; =20 /* ASYM feature bypasses nice load balance check */ @@ -14573,7 +14578,7 @@ static int sched_balance_newidle(struct rq *this_rq= , struct rq_flags *rf) if (!sd) goto out; =20 - if (!get_rd_overloaded(this_rq->rd) || + if (!get_rd_overloaded(rcu_dereference_root_domain(this_rq->rd)) || this_rq->avg_idle < sd->max_newidle_lb_cost) { =20 update_next_balance(sd, &next_balance); diff --git a/kernel/sched/rt.c b/kernel/sched/rt.c index e6e5f8a2caaf..0bf3ea1e9957 100644 --- a/kernel/sched/rt.c +++ b/kernel/sched/rt.c @@ -338,15 +338,18 @@ static inline bool need_pull_rt_task(struct rq *rq, s= truct task_struct *prev) =20 static inline int rt_overloaded(struct rq *rq) { - return atomic_read(&rq->rd->rto_count); + return atomic_read(&rcu_dereference_root_domain(rq->rd)->rto_count); } =20 static inline void rt_set_overload(struct rq *rq) { + struct root_domain *rd; + if (!rq->online) return; =20 - cpumask_set_cpu(rq->cpu, rq->rd->rto_mask); + rd =3D rcu_dereference_root_domain(rq->rd); + cpumask_set_cpu(rq->cpu, rd->rto_mask); /* * Make sure the mask is visible before we set * the overload count. That is checked to determine @@ -357,17 +360,20 @@ static inline void rt_set_overload(struct rq *rq) * Matched by the barrier in pull_rt_task(). */ smp_wmb(); - atomic_inc(&rq->rd->rto_count); + atomic_inc(&rd->rto_count); } =20 static inline void rt_clear_overload(struct rq *rq) { + struct root_domain *rd; + if (!rq->online) return; =20 + rd =3D rcu_dereference_root_domain(rq->rd); /* the order here really doesn't matter */ - atomic_dec(&rq->rd->rto_count); - cpumask_clear_cpu(rq->cpu, rq->rd->rto_mask); + atomic_dec(&rd->rto_count); + cpumask_clear_cpu(rq->cpu, rd->rto_mask); } =20 static inline int has_pushable_tasks(struct rq *rq) @@ -580,7 +586,7 @@ static int rt_se_boosted(struct sched_rt_entity *rt_se) =20 static inline const struct cpumask *sched_rt_period_mask(void) { - return this_rq()->rd->span; + return rcu_dereference_root_domain(this_rq()->rd)->span; } =20 static inline @@ -608,7 +614,7 @@ bool sched_rt_bandwidth_account(struct rt_rq *rt_rq) static void do_balance_runtime(struct rt_rq *rt_rq) { struct rt_bandwidth *rt_b =3D sched_rt_bandwidth(rt_rq); - struct root_domain *rd =3D rq_of_rt_rq(rt_rq)->rd; + struct root_domain *rd =3D rcu_dereference_root_domain(rq_of_rt_rq(rt_rq)= ->rd); int i, weight; u64 rt_period; =20 @@ -659,7 +665,7 @@ static void do_balance_runtime(struct rt_rq *rt_rq) */ static void __disable_runtime(struct rq *rq) { - struct root_domain *rd =3D rq->rd; + struct root_domain *rd =3D rcu_dereference_root_domain(rq->rd); rt_rq_iter_t iter; struct rt_rq *rt_rq; =20 @@ -1058,7 +1064,7 @@ inc_rt_prio_smp(struct rt_rq *rt_rq, int prio, int pr= ev_prio) return; =20 if (rq->online && prio < prev_prio) - cpupri_set(&rq->rd->cpupri, rq->cpu, prio); + cpupri_set(&rcu_dereference_root_domain(rq->rd)->cpupri, rq->cpu, prio); } =20 static void @@ -1073,7 +1079,8 @@ dec_rt_prio_smp(struct rt_rq *rt_rq, int prio, int pr= ev_prio) return; =20 if (rq->online && rt_rq->highest_prio.curr !=3D prev_prio) - cpupri_set(&rq->rd->cpupri, rq->cpu, rt_rq->highest_prio.curr); + cpupri_set(&rcu_dereference_root_domain(rq->rd)->cpupri, rq->cpu, + rt_rq->highest_prio.curr); } =20 static void @@ -1575,8 +1582,10 @@ select_task_rq_rt(struct task_struct *p, int cpu, in= t flags) =20 static void check_preempt_equal_prio(struct rq *rq, struct task_struct *p) { + struct root_domain *rd =3D rcu_dereference_root_domain(rq->rd); + if (rq->curr->nr_cpus_allowed =3D=3D 1 || - !cpupri_find(&rq->rd->cpupri, rq->donor, NULL)) + !cpupri_find(&rd->cpupri, rq->donor, NULL)) return; =20 /* @@ -1584,7 +1593,7 @@ static void check_preempt_equal_prio(struct rq *rq, s= truct task_struct *p) * see if it is pushed or pulled somewhere else. */ if (p->nr_cpus_allowed !=3D 1 && - cpupri_find(&rq->rd->cpupri, p, NULL)) + cpupri_find(&rd->cpupri, p, NULL)) return; =20 /* @@ -1793,12 +1802,12 @@ static int find_lowest_rq(struct task_struct *task) */ if (sched_asym_cpucap_active()) { =20 - ret =3D cpupri_find_fitness(&task_rq(task)->rd->cpupri, + ret =3D cpupri_find_fitness(&rcu_dereference_root_domain(task_rq(task)->= rd)->cpupri, task, lowest_mask, rt_task_fits_capacity); } else { =20 - ret =3D cpupri_find(&task_rq(task)->rd->cpupri, + ret =3D cpupri_find(&rcu_dereference_root_domain(task_rq(task)->rd)->cpu= pri, task, lowest_mask); } =20 @@ -2189,16 +2198,17 @@ static inline void rto_start_unlock(atomic_t *v) =20 static void tell_cpu_to_push(struct rq *rq) { + struct root_domain *rd =3D rcu_dereference_root_domain(rq->rd); int cpu =3D -1; =20 /* Keep the loop going if the IPI is currently active */ - atomic_inc(&rq->rd->rto_loop_next); + atomic_inc(&rd->rto_loop_next); =20 /* Only one CPU can initiate a loop at a time */ - if (!rto_start_trylock(&rq->rd->rto_loop_start)) + if (!rto_start_trylock(&rd->rto_loop_start)) return; =20 - raw_spin_lock(&rq->rd->rto_lock); + raw_spin_lock(&rd->rto_lock); =20 /* * The rto_cpu is updated under the lock, if it has a valid CPU @@ -2206,17 +2216,17 @@ static void tell_cpu_to_push(struct rq *rq) * update to loop_next, and nothing needs to be done here. * Otherwise it is finishing up and an IPI needs to be sent. */ - if (rq->rd->rto_cpu < 0) - cpu =3D rto_next_cpu(rq->rd); + if (rd->rto_cpu < 0) + cpu =3D rto_next_cpu(rd); =20 - raw_spin_unlock(&rq->rd->rto_lock); + raw_spin_unlock(&rd->rto_lock); =20 - rto_start_unlock(&rq->rd->rto_loop_start); + rto_start_unlock(&rd->rto_loop_start); =20 if (cpu >=3D 0) { /* Make sure the rd does not get freed while pushing */ - sched_get_rd(rq->rd); - irq_work_queue_on(&rq->rd->rto_push_work, cpu); + sched_get_rd(rd); + irq_work_queue_on(&rd->rto_push_work, cpu); } } =20 @@ -2277,7 +2287,7 @@ static void pull_rt_task(struct rq *this_rq) =20 /* If we are the only overloaded CPU do nothing */ if (rt_overload_count =3D=3D 1 && - cpumask_test_cpu(this_rq->cpu, this_rq->rd->rto_mask)) + cpumask_test_cpu(this_rq->cpu, rcu_dereference_root_domain(this_rq->r= d)->rto_mask)) return; =20 #ifdef HAVE_RT_PUSH_IPI @@ -2287,7 +2297,7 @@ static void pull_rt_task(struct rq *this_rq) } #endif =20 - for_each_cpu(cpu, this_rq->rd->rto_mask) { + for_each_cpu(cpu, rcu_dereference_root_domain(this_rq->rd)->rto_mask) { if (this_cpu =3D=3D cpu) continue; =20 @@ -2392,7 +2402,7 @@ static void rq_online_rt(struct rq *rq) =20 __enable_runtime(rq); =20 - cpupri_set(&rq->rd->cpupri, rq->cpu, rq->rt.highest_prio.curr); + cpupri_set(&rcu_dereference_root_domain(rq->rd)->cpupri, rq->cpu, rq->rt.= highest_prio.curr); } =20 /* Assumes rq->lock is held */ @@ -2403,7 +2413,7 @@ static void rq_offline_rt(struct rq *rq) =20 __disable_runtime(rq); =20 - cpupri_set(&rq->rd->cpupri, rq->cpu, CPUPRI_INVALID); + cpupri_set(&rcu_dereference_root_domain(rq->rd)->cpupri, rq->cpu, CPUPRI_= INVALID); } =20 /* diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h index 26ae13c86b69..6d45e67bcdc3 100644 --- a/kernel/sched/sched.h +++ b/kernel/sched/sched.h @@ -1256,7 +1256,7 @@ struct rq { int membarrier_state; #endif =20 - struct root_domain *rd; + struct root_domain __rcu *rd; struct sched_domain __rcu *sd; =20 struct balance_callback *balance_callback; @@ -2143,6 +2143,9 @@ queue_balance_callback(struct rq *rq, rq->balance_callback =3D head; } =20 +#define rcu_dereference_root_domain(p) \ + rcu_dereference_all_check((p), lockdep_is_held(&sched_domains_mutex)) + #define rcu_dereference_sched_domain(p) \ rcu_dereference_all_check((p), lockdep_is_held(&sched_domains_mutex)) =20 @@ -3030,7 +3033,7 @@ static inline void add_nr_running(struct rq *rq, unsi= gned count) } =20 if (prev_nr < 2 && rq->nr_running >=3D 2) - set_rd_overloaded(rq->rd, 1); + set_rd_overloaded(rcu_dereference_root_domain(rq->rd), 1); =20 sched_update_tick_dependency(rq); } diff --git a/kernel/sched/syscalls.c b/kernel/sched/syscalls.c index b215b0ead9a6..89a2465727fd 100644 --- a/kernel/sched/syscalls.c +++ b/kernel/sched/syscalls.c @@ -621,15 +621,15 @@ int __sched_setscheduler(struct task_struct *p, #endif /* CONFIG_RT_GROUP_SCHED */ if (dl_bandwidth_enabled() && dl_policy(policy) && !(attr->sched_flags & SCHED_FLAG_SUGOV)) { - cpumask_t *span =3D rq->rd->span; + struct root_domain *rd =3D rcu_dereference_root_domain(rq->rd); =20 /* * Don't allow tasks with an affinity mask smaller than * the entire root_domain to become SCHED_DEADLINE. We * will also fail if there's no bandwidth available. */ - if (!cpumask_subset(span, p->cpus_ptr) || - rq->rd->dl_bw.bw =3D=3D 0) { + if (!cpumask_subset(rd->span, p->cpus_ptr) || + rd->dl_bw.bw =3D=3D 0) { retval =3D -EPERM; goto unlock; } @@ -1127,7 +1127,7 @@ int dl_task_check_affinity(struct task_struct *p, con= st struct cpumask *mask) * root_domain. */ guard(rcu)(); - if (!cpumask_subset(task_rq(p)->rd->span, mask)) + if (!cpumask_subset(rcu_dereference_root_domain(task_rq(p)->rd)->span, ma= sk)) return -EBUSY; =20 return 0; diff --git a/kernel/sched/topology.c b/kernel/sched/topology.c index 21e816ad23ee..bf83ceee23e9 100644 --- a/kernel/sched/topology.c +++ b/kernel/sched/topology.c @@ -413,7 +413,7 @@ static bool build_perf_domains(const struct cpumask *cp= u_map) int i; struct perf_domain *pd =3D NULL, *tmp; int cpu =3D cpumask_first(cpu_map); - struct root_domain *rd =3D cpu_rq(cpu)->rd; + struct root_domain *rd =3D rcu_dereference_root_domain(cpu_rq(cpu)->rd); =20 if (!sysctl_sched_energy_aware) goto free; @@ -478,9 +478,8 @@ void rq_attach_root(struct rq *rq, struct root_domain *= rd) =20 rq_lock_irqsave(rq, &rf); =20 - if (rq->rd) { - old_rd =3D rq->rd; - + old_rd =3D rcu_dereference_root_domain(rq->rd); + if (old_rd) { if (cpumask_test_cpu(rq->cpu, old_rd->online)) set_rq_offline(rq); =20 @@ -3461,8 +3460,10 @@ static void partition_sched_domains_locked(int ndoms= _new, cpumask_var_t doms_new /* Build perf domains: */ for (i =3D 0; i < ndoms_new; i++) { for (j =3D 0; j < n && !sched_energy_update; j++) { + int cpu =3D cpumask_first(doms_cur[j]); + if (cpumask_equal(doms_new[i], doms_cur[j]) && - cpu_rq(cpumask_first(doms_cur[j]))->rd->pd) { + rcu_dereference_root_domain(cpu_rq(cpu)->rd)->pd) { has_eas =3D true; goto match3; } --=20 2.55.0 From nobody Mon Sep 28 02:57:20 2026 Received: from CWXP265CU010.outbound.protection.outlook.com (mail-ukwestazon11022130.outbound.protection.outlook.com [52.101.101.130]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1DA2C361962 for ; Thu, 27 Aug 2026 22:18:21 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.101.130 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787869102; cv=fail; b=IQFhHQ3aJfTge4npOTluCafl6uNjGPAFzK4e0WraBU6Wwg4RvOPPzWdF5RcB3nf7M0F7Tc7WBVU6Sv5CcoJkN70P72lWt2PJVgdDZDX0U5RLHpfQ8EYfQ0t9QFpIbrED1i8JBfzwcEXRBvsH+bfnwa2zyN2fal1AtoYNhKYtiug= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787869102; c=relaxed/simple; bh=+XkfxlbpQKgMEeXYx5zk8rSq2o9H67LqImXVv43uosE=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: Content-Type:MIME-Version; b=cTZVzTcRlocbxfk5TLAJPb/vTavUy8Wlj+6Y/hQdqDSTWbIuHc81MWgooJqCH6klMVUWT1zi7w5hMeGHxeJDwrRmXFX9kJuC9pggNSQQ5lt8WnYiExI2aBhaNyOD8Z1fwCWkh2d+8N+Ipi8JNKc0LhgxVs2apsj0ueLTZtmqD9c= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=atomlin.com; spf=pass smtp.mailfrom=atomlin.com; arc=fail smtp.client-ip=52.101.101.130 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=atomlin.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=atomlin.com ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=kaU1/vrZMHTCYEO+DkjA0KN79lLlT4PGq5PkKEwcKceDUAi3+rYZpucJZJV85FowBBhEoGdNystzNqhpEDNdGf4haEGeErjHHzbbzr9j8cR+FGzZnjzxmjrW5DfaARlS/tTkKXHNntIyG55EWRn456zP2/5xHdo3OlLcrFYFlRcaZGIJQocf4SulUUjAKUmqkhHy2Q7iwEa21PHlBDz51/idLeVnP0uMMMDW8Mr52z1eM6592hE1z8J8wWCTfS6NILZF0qnw+/JwUMm34iRESVd0mm7nMPbIlZehlzWzUsHYgglOBgG98cJkPARyfNw56D+xkWzwdVxeJL60nM904A== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:MIME-Version; bh=UYwraUWxnT9P+Xj2zx8xXWpCIPcqPzZxozbnwcWDgKA=; b=wuKA55pwj7sh60etfJmM7SK5qdu1zHBIajsV/gsn/dfgSJnpvezsnUJ3YhX4DypPrQCSFWgKjROjnyW3hyzq/GBXo3+KgoNex1hlCrWc8KlcFwqAeijqWoCZveE7XJuD6+8a4IY01s69mkvY98YCuOyLycF6G0y2F+lYeiQZ9y3+tLmngTG72gnlCS/86tAhVydvQ4JI1tjpNEMd8YEY/2mYv+u9+HNWAn2Sbij6PM+KrzTv0zOtdk8OdHpqO0N53OEXiFuRPNmkvMFjGUQibisTfS4xHT0g/1o+wcfYSlC3+C27jND/JZt/kyPJ6m3v5PdgzpYSNUkJG86YzD/m7A== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass smtp.mailfrom=atomlin.com; dmarc=pass action=none header.from=atomlin.com; dkim=pass header.d=atomlin.com; arc=none Authentication-Results: dkim=none (message not signed) header.d=none;dmarc=none action=none header.from=atomlin.com; Received: from CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM (2603:10a6:400:183::5) by LO6P123MB7095.GBRP123.PROD.OUTLOOK.COM (2603:10a6:600:342::13) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.360.10; Thu, 27 Aug 2026 22:18:14 +0000 Received: from CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM ([fe80::cec4:77ab:262e:d230]) by CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM ([fe80::cec4:77ab:262e:d230%4]) with mapi id 15.21.0360.008; Thu, 27 Aug 2026 22:18:13 +0000 From: Aaron Tomlin To: mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org Cc: paulmck@kernel.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com, zhanxusheng1024@gmail.com, neelx@suse.com, atomlin@atomlin.com, chjohnst@mail.com, mproche@mail.com, sean@ashe.io, steve@abita.co, rishil1999@outlook.com, linux-kernel@vger.kernel.org Subject: [PATCH v9 2/6] sched/debug: Protect lockless rq->rd access in print_dl_rq() Date: Thu, 27 Aug 2026 18:18:04 -0400 Message-ID: <20260827221809.988394-3-atomlin@atomlin.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260827221809.988394-1-atomlin@atomlin.com> References: <20260827221809.988394-1-atomlin@atomlin.com> Content-Transfer-Encoding: quoted-printable X-ClientProxiedBy: LO4P302CA0041.GBRP302.PROD.OUTLOOK.COM (2603:10a6:600:317::19) To CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM (2603:10a6:400:183::5) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: CWLP123MB6607:EE_|LO6P123MB7095:EE_ X-MS-Office365-Filtering-Correlation-Id: d44f42f8-953d-4e31-dea0-08df0489161d X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|376014|7416014|1800799024|366016|23010399003|10067099003|56012099006|6133799003|22082099003|18002099003; X-Microsoft-Antispam-Message-Info: ZGkjyqdnhLltzwyiKCLSDRb6DZCSj8BkMEg2VQ9guqBY1ZryehavTkzUj9ilYUiMQhdsewTIUu+V+JxTXG9oqh1fqi17BZ0TqKueGg8JdWHlWvL/yHc4hFWPVRpNSg5v60Wr4/9ISTX+g6e7ufSXjwrcEgCc/8Bv3ujMbcn/PW2jcr0iB+Trn0qWVkGu4ls9lwaD58wZ3al50wljQf9IAYbyUdCe3JAvTljZQe8Ib0d4OuugMarIHTXECvN0l4zv+9Hf8fY9rb6LNhlMpxhpPnrM0p8yWvV/Jrq9P268Yz7+y+jyHDzqx4squsp9DCZ77ykfwsiMeSqMWvOcDqU3UPv0zSaxA/tiXzJr5ufFd5u+GDldRo8ZKjoSA7oLLyEMIAdZpyPpJ8fjvlFIMS57g/4rsawRFVlkhKE1Hw3BuN0Gu2vEFDougv7Bek03HZ1UVWJtBN4zAmyz0tg7TmwXfQRbflqEBTvI1qQA/g/YWy+KtkQvwavfSXocKdi7TbQgNq7ChKk9pml7LSOUGl4kFO5tvK29L8mTpCwdSUsbesD2K4BQ8pzJlp7MLlwBXG1KmSHKDF04AveVhy12hlboVEgajHMSt8IhYbhzOwWAi2NnF1qTj3zi6DZUP0jiBbtWxIPmY8pQkz9PL+nQ2rh5Cwv7vFfrcuFxM9lbbfwxo9s= X-Forefront-Antispam-Report: CIP:255.255.255.255;CTRY:;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM;PTR:;CAT:NONE;SFS:(13230040)(376014)(7416014)(1800799024)(366016)(23010399003)(10067099003)(56012099006)(6133799003)(22082099003)(18002099003);DIR:OUT;SFP:1102; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: =?us-ascii?Q?bDJ6H2iRj3uWTMlvM0og69OanrgXXtI8ehuXt9lbZjzjIUHUo7pW4Q5ykiEJ?= =?us-ascii?Q?B+tLqimgo/6UYzQQvjTrPWjV4UoOsA73AMH29Z2/D2w7C5ZJ+x3UsLwLotHX?= =?us-ascii?Q?ZnPXhqKahgf5Y5UmqElracajeGRcQVp7f931Pa/26/MaPzZYxW/P+K8r8V6y?= =?us-ascii?Q?vRnfIupuYsy50x1mHwK64Iaoj2o4hQO084JJ3x3MUyxXAkW8FpSFa2EMEF8Y?= =?us-ascii?Q?1t85vV4er6onU3eSp1wmfXE8MlH4fbBGntMA0EmMuHbVbaQL9L7TzpH31IVD?= =?us-ascii?Q?FVehKEIjhtFa8JOcV4RNN99F6KIfOfdS3KngXbUUYLEcNVozNU3cGZrWSJn3?= =?us-ascii?Q?/NLUF0I811lEhkFLM07Y8BfdgqDbYHYjSPSasekb7QFVB+rCHswSruzBy6WI?= =?us-ascii?Q?dunNjtBfrgcmXK5NBGcrZaO98bWc2JChqDoCoX98hv1aKTaTDpeMS67Xq6vX?= =?us-ascii?Q?9iMUwn3iId8ye4Ih8zeGnLwGLTnqD/zaQNZVUhQE8HVDdChE35YloteI9chA?= =?us-ascii?Q?MAccH4SEmMDY8XgatdvqK7nANtq1mzN/85sg3R83jtw48A+F7oDdaD0hV7Ob?= =?us-ascii?Q?rNuBcJm28N+e0VoH6lXWUGqs848lhrsPoY8QBQxgw5s9kha4lgyfDDOdpC7P?= =?us-ascii?Q?x4B/iVAaZ4oqQUVZmBQpGza3y3ZeuB+fCemTphGgQrt7jS6BlxxRuZm9hol6?= =?us-ascii?Q?yMM81qigJqp+g5IwNLvWDzaGOklI3sewAMAp7D3bP/JC0/ObO3dipfXL9ANu?= =?us-ascii?Q?Nb6IWWuna+ALJTk94qMdfBr4jAfcSjZyuuspzT1bsc95FecmYOAcSiexv7FB?= =?us-ascii?Q?RrTLaFPurUjRTCUgTskmJWkMoIk+SGjZGIpIONGk3Z5/Mj51hRr2w7BVtL5M?= =?us-ascii?Q?rb0ghPclXAVlbEnj8wugjoYSW7OrexHL+a3UGbLnafNW61Fzj5CcUrH8Z2+E?= =?us-ascii?Q?SvxJ2l1kWMssYy+bU8vc+3ExgWH3an28WQyTiyaXhcbvtWFbGu1hYdE0MDbn?= =?us-ascii?Q?/akLAVtQtZBTrjCY3d4FsumHIIj6V9+Ij8gvYg3bO20BygOX7TSV1iSIlIXE?= =?us-ascii?Q?GGwCvxdbrxhe9Ehb3Y9itvvG+YjJypGX7/2tkZq42xhaQG90YmE/5mhHfpGq?= =?us-ascii?Q?+pEB/9Pcsck8vTGuIegj67flhPVEE0CMov5hMjf6E8A7b30pC17Unnqueg8D?= =?us-ascii?Q?+r6ipnPiejZ4AJpJ7YTApN/R2nkb324Orwovsu/uyZH+GV3cNzfPivG2K8uZ?= =?us-ascii?Q?RM80CDcMDI345zQdSVEPtUHaTpiok+laeUI/SUjNf+zEPTjY7nCt214yDI+/?= =?us-ascii?Q?PqY6gk7BIf0+pjkBSRQ7b5zqoQvTSEJfC8ZyoiUOFMuwATL7QLuEwl88mMSP?= =?us-ascii?Q?7LgUAbXrLOPTK2M+dwRs8KioyeszBNtja2hDRdzq/pgKXd/eegLZYPRfqRv7?= =?us-ascii?Q?Tdm7A3ChraqEYvNV5O1HMFEqpA7vYNNmPJJMP+77hOWmaMFX7Y9rYm4A9iHr?= =?us-ascii?Q?ob7k3/CxMFeGWbD7yia30ef4jy76xDR5awT9GBbaR+dPFK9jCuhj51eDCMwU?= =?us-ascii?Q?RFE+nBHAoI1ssMywa5G0nSkeTqCbswF/41j5CdBwceQOYjd86yN2wA/XSBME?= =?us-ascii?Q?xjBCkFAgRImCfFra/ZAqGWYUCH13CYpXKGPwkzEo3pq3Fr9SVlgSe65QnKj5?= =?us-ascii?Q?LMSIit+3V0LmWdk+mHt+DsDczLxLWj9+onVhfm3gkqAjc3tH?= X-OriginatorOrg: atomlin.com X-MS-Exchange-CrossTenant-Network-Message-Id: d44f42f8-953d-4e31-dea0-08df0489161d X-MS-Exchange-CrossTenant-AuthSource: CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM X-MS-Exchange-CrossTenant-AuthAs: Internal X-MS-Exchange-CrossTenant-OriginalArrivalTime: 27 Aug 2026 22:18:13.8722 (UTC) X-MS-Exchange-CrossTenant-FromEntityHeader: Hosted X-MS-Exchange-CrossTenant-Id: e6a32402-7d7b-4830-9a2b-76945bbbcb57 X-MS-Exchange-CrossTenant-MailboxType: HOSTED X-MS-Exchange-CrossTenant-UserPrincipalName: N8AuIgurTdGaV3KyQjjLjrHyqyGEsY+xjph+yoHjSwTjpY0UArEKOic6vY4WQUaqDmB75d0VFonPPZooh8Of5w== X-MS-Exchange-Transport-CrossTenantHeadersStamped: LO6P123MB7095 Content-Type: text/plain; charset="utf-8" In print_dl_rq(), cpu_rq(cpu)->rd is dereferenced locklessly to display deadline bandwidth statistics. During CPU hot-unplug or cgroup cpuset repartitioning events, partition_sched_domains() calls cpu_attach_domain(), which executes rq_attach_root() to detach the CPU from its root_domain. When the reference count of the detached root_domain drops to zero, rq_attach_root() calls call_rcu(&old_rd->rcu, free_rootdomain) to schedule memory teardown after an RCU grace period. However, rq_attach_root() previously updated rq->rd using a plain C store without an RCU publication barrier (i.e., rcu_assign_pointer()). Without a release memory barrier on the writer side, CPU or compiler reordering could allow the new rq->rd pointer store to become visible to other CPUs before the initialization writes to rd->dl_bw are committed. Furthermore, because print_dl_rq() did not hold an RCU read lock while dereferencing cpu_rq(cpu)->rd, an RCU grace period could elapse concurrently while debugfs is reading the file, allowing free_rootdomain() to execute kfree(old_rd) and causing a use-after-free race condition when print_dl_rq() reads dl_bw->bw. Resolve this by using rcu_assign_pointer(rq->rd, rd) in rq_attach_root() to guarantee a release memory barrier when publishing a root_domain. Finally, fetch rq->rd using guard(rcu)() and rcu_dereference() in print_dl_= rq(). Fixes: 02968ccf7b80 ("sched: add /proc/sched_debug file") Reported-by: sashiko-bot Signed-off-by: Aaron Tomlin --- kernel/sched/debug.c | 3 ++- kernel/sched/topology.c | 2 +- 2 files changed, 3 insertions(+), 2 deletions(-) diff --git a/kernel/sched/debug.c b/kernel/sched/debug.c index 72236db67983..61932ef7ec4f 100644 --- a/kernel/sched/debug.c +++ b/kernel/sched/debug.c @@ -1172,7 +1172,8 @@ void print_dl_rq(struct seq_file *m, int cpu, struct = dl_rq *dl_rq) SEQ_printf(m, " .%-30s: %lu\n", #x, (unsigned long)(dl_rq->x)) =20 PU(dl_nr_running); - dl_bw =3D &cpu_rq(cpu)->rd->dl_bw; + guard(rcu)(); + dl_bw =3D &rcu_dereference_root_domain(cpu_rq(cpu)->rd)->dl_bw; SEQ_printf(m, " .%-30s: %lld\n", "dl_bw->bw", dl_bw->bw); SEQ_printf(m, " .%-30s: %lld\n", "dl_bw->total_bw", dl_bw->total_bw); =20 diff --git a/kernel/sched/topology.c b/kernel/sched/topology.c index bf83ceee23e9..58913dc3a8f2 100644 --- a/kernel/sched/topology.c +++ b/kernel/sched/topology.c @@ -495,7 +495,7 @@ void rq_attach_root(struct rq *rq, struct root_domain *= rd) } =20 atomic_inc(&rd->refcount); - rq->rd =3D rd; + rcu_assign_pointer(rq->rd, rd); =20 cpumask_set_cpu(rq->cpu, rd->span); if (cpumask_test_cpu(rq->cpu, cpu_active_mask)) --=20 2.55.0 From nobody Mon Sep 28 02:57:20 2026 Received: from CWXP265CU010.outbound.protection.outlook.com (mail-ukwestazon11022130.outbound.protection.outlook.com [52.101.101.130]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EAD9432D7F1 for ; Thu, 27 Aug 2026 22:18:23 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.101.130 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787869106; cv=fail; b=DPxHpPdJp9UARr7vO+q9YGC1t8L8fxliPmAXjM8AwPRW/YmgBqTH5yPp3ObnI7/5ynbdrCxmpJnDSU05pSXIc0btAGP5nZiGgyG/Db61LY4P9+/US52VIdwBI04fzgNMlK8FRjb59XD4qAG3ioDzH9la0vzs2zFji4kEe8N5IP4= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787869106; c=relaxed/simple; bh=EjWRLuMtddpGpEz0wwcKvBoRX/wKRrozMZJvimQ/Fmc=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: Content-Type:MIME-Version; b=Q3Y9rTYItFpnRG6ioW2cQzoxHsWEQncBK2w4NlYbXQP1JDj4K6hzrnw8f87W4Wtvia+T20KcZGw0ldqRFABVdSy87q0SVHhflhnGQWqZPrwCAnPs7C6qlWGuv66Uncz7QlfjeAfmSOs0ayXmIf1UKCoX+DUUc933Z79Ubx49aVA= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=atomlin.com; spf=pass smtp.mailfrom=atomlin.com; arc=fail smtp.client-ip=52.101.101.130 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=atomlin.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=atomlin.com ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=GzAFHyffu/VsYuMBAUXUJ5+sXc4/KgmkUJ36Yr6IesgePDDE2BHbkdoJicN7Sk0UubjV95NJzVOwg8BDVAAx1RzqcscS8UGzPEMskM+rj3xhlc5q06hG0gelxZdFInviXQkajA7bt678XlH6UKYbs02PCZemAUE2Azq0sALBmLcEVaxvVLEZbJeDfW9zkrFUV8cKjsme3CMpVZW8B9isrP4t7rbRruQl6vwv7urTOsoCDKyOdOpkidpQ4vvFXnL4euTlQAh5FiRhDF19rHo5T2gc2zlnFmtk0xplOHE+i3fijM1mPNMILZxtTQowU0YsyHYW/ZIF4ZFkiUkDX0lvPQ== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:MIME-Version; bh=N4L2lHhG5/Tgj/tMXo3AuwJe4R1m6EosDrnQHHDyi3c=; b=tAkuqRsKS1DiwZItq4GUB/XDwHD22zRstbxvOR/XxAhXpQzPeouLo0ub9jvsO0ueNWwrsEPN4oi6vC5fb/rXuQJkNN2OCnk6ZmglqO+6V/kx3on5Jc3LxohWBihlfhPQjNL4wlv7/FoKWMvi7ZjVP121JWjUNGqkOPOE3N94kXvDFhU/19C3W4fYJtuhuRrXfdJ2EjBemmCqd/XStxFM6VUBzOCB+eocgEshkfI/6StG2xJUDk1uKwQVccduzzNI89QDk4Lk9UuiQlCM4jch9VQocVNDp3g3OxFfJuY4pSrAYHFtnHCK7QQrk+B3qTFdIdfUDNddv8zArr78GJpBFg== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass smtp.mailfrom=atomlin.com; dmarc=pass action=none header.from=atomlin.com; dkim=pass header.d=atomlin.com; arc=none Authentication-Results: dkim=none (message not signed) header.d=none;dmarc=none action=none header.from=atomlin.com; Received: from CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM (2603:10a6:400:183::5) by LO6P123MB7095.GBRP123.PROD.OUTLOOK.COM (2603:10a6:600:342::13) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.360.10; Thu, 27 Aug 2026 22:18:15 +0000 Received: from CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM ([fe80::cec4:77ab:262e:d230]) by CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM ([fe80::cec4:77ab:262e:d230%4]) with mapi id 15.21.0360.008; Thu, 27 Aug 2026 22:18:15 +0000 From: Aaron Tomlin To: mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org Cc: paulmck@kernel.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com, zhanxusheng1024@gmail.com, neelx@suse.com, atomlin@atomlin.com, chjohnst@mail.com, mproche@mail.com, sean@ashe.io, steve@abita.co, rishil1999@outlook.com, linux-kernel@vger.kernel.org Subject: [PATCH v9 3/6] sched/debug: Protect lockless rq->curr access in print_cpu() Date: Thu, 27 Aug 2026 18:18:05 -0400 Message-ID: <20260827221809.988394-4-atomlin@atomlin.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260827221809.988394-1-atomlin@atomlin.com> References: <20260827221809.988394-1-atomlin@atomlin.com> Content-Transfer-Encoding: quoted-printable X-ClientProxiedBy: LO4P123CA0023.GBRP123.PROD.OUTLOOK.COM (2603:10a6:600:151::10) To CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM (2603:10a6:400:183::5) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: CWLP123MB6607:EE_|LO6P123MB7095:EE_ X-MS-Office365-Filtering-Correlation-Id: 1b508c76-bb19-4236-b58c-08df048916dd X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|376014|7416014|1800799024|366016|23010399003|10067099003|56012099006|22082099003|18002099003; X-Microsoft-Antispam-Message-Info: 9jRCnkNvDn0gHLV88+3I8tpTGOl/imjYKOPzrQ0GloHXy1MFnnlXKdeLVfA7QK6IupodckAbxHM+sdhwDCxvmbzXjis34mnEv3eIKJByLjpm+j+fq9G52ZAO/V06pNowtGzn9nS2JcMVai9W/yXKABSwAdhJqX15Xg0iCRiADPZp7C02bE13w577Cs/rO50aH6MrL7NKrQFyF3vcU63gJw07MojtlJko+tILuJeEqVtQU1pAZ9qz8j69CuowYOhW3/ZMEVaQnrB82kDAU8nL8TyLbenPHKhDecMPnMcbHo4KsJ25K0m+g3mDPqppSDnJtHI6CyBu/UKp/KQ78ux+wc7EfZahzp+e53oR2AsUWUxJLaeSW7GQOlpPQk9Jxxv4Xsx5Krs0rR4R6ec81OctrCoDn8LlulMEnyP2jS0gdR8tyhfrGULaPIboKe/TU42OncuqcIqFPMVsdiRSsfNXR34N7ms2FvyFPWKS4F0mGsKzps8v3dw7n2ud4iWCj2d8qkBUI0Zo/+jKe3kR+QDLaEL7dnQOJ1svAs+3Dj/gmXa9t/OC03pSDLl7nvZBSn3lHLMFBfCCAS5MZKHlxLL1R0WvilBKTos8W6WL/GH4wWd2HOux9owPWbKXO/Us46AJkXNPsZFxNXXofhACqjQUtNo21XgqPsUoZbdHzyxP8ZM= X-Forefront-Antispam-Report: CIP:255.255.255.255;CTRY:;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM;PTR:;CAT:NONE;SFS:(13230040)(376014)(7416014)(1800799024)(366016)(23010399003)(10067099003)(56012099006)(22082099003)(18002099003);DIR:OUT;SFP:1102; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: =?us-ascii?Q?X1DrbFzTDGwa9RhMJGb/exK6U9QWV1WTwExdDKjBcrHlNqjIWhuh9a+nauc9?= =?us-ascii?Q?ZMLbGj+ZVpLeWyzExjpLmo4HTvufYhdGpFdE04XuShqnWKXx/PN2G9tbMawR?= =?us-ascii?Q?g/btB8FhaGxrkd3mtzyY/xZtQY3K8E5ulZhWYWbKAMwK1OOoStMKK/gR5AWr?= =?us-ascii?Q?40LM25QJFFD0+9GnfNwGdG7J+ox/N3NRGX9NFvd6dnMUgeddxVQws90scJH7?= =?us-ascii?Q?S5ffZiyi5vTt56A/QrdEnUzg6h9/x7uVwXT3iwQCBk0/EiFa/V7ZxUogKhXN?= =?us-ascii?Q?Uj7KJ4cwZgoZ8DAL6eRWnzhYz6Tiic7hGyfEZjgqVrQxtQkikrpZUXlKJlh3?= =?us-ascii?Q?uRinFxp8qPxgrG5+sIQ0BlnHNfPZ4a1cdBjBvgnAAcKIAZws6Dzn4IzO5XDq?= =?us-ascii?Q?AFVCXAIPAmiv562Yu0hcuPcYoDWTteqW9YjW0b5LW9ybco97RlAtpfZfCjk5?= =?us-ascii?Q?lTfD9ZC7wkBkT/9lTcivBIPograxYXYehrQLGyczPOBPd5/DaQIy9/5Z2XRM?= =?us-ascii?Q?4mgooGnyz3IO1shbShe3tjVlcdzItQDqu7gjQxeXqbgpDKDc3J/oNCPWu8+b?= =?us-ascii?Q?I/e6rQpLKKYO/tXmKnGpuMf0AQfbLEZcF7YyJnDEADof0tpUkW/X3bbWQa2Z?= =?us-ascii?Q?AO1jUJw0is8RScgkxizEsGo++SOHwQ0YnAlEkAKBFiHAjaF7Q0yuC6B/M8Zg?= =?us-ascii?Q?fsJ+JuS0y2O4TkKxmSkQLz+lsuCbIKvPivUX4lFyydD3PWD5tD71DhlKG7Y9?= =?us-ascii?Q?cI6FiFh9WoVdSg1lfcJF1GxMf2NoRj6MU/Xfsn7QyAZedc6IAbVM9J3LM/em?= =?us-ascii?Q?wVg47X/FESFfrTane1ayYmYPhQOizJkWaEG7N23Xq04vEdCHxQPVQBSFGowV?= =?us-ascii?Q?v/wOTOE3pgTRtcRL+J1DWKcViTkl5WhiO+dub7Nf/1jjny80hWrNa/4/1wz6?= =?us-ascii?Q?zZQFB56w4wTP2G2uyrvFNEHezp+5oP45EuYDko7k3AqUxY+jliCVC8SjhuaJ?= =?us-ascii?Q?1QxlbJUOqUPAIe2tQjz4lJHylfHwLZC39UKhR1oHEzO95/FmINe/mAfJwGzh?= =?us-ascii?Q?jUuI5aEvAMUGla5yUiN+NFF4GUKM/DH7Rrtjow6Ms3vhYlN+u2sF56Hhk1RY?= =?us-ascii?Q?YC1VSxncwleyslmcbIaY9dBRW6r9XEDPuYjShL0XD7T1u5d7a+uY5y2wgsJy?= =?us-ascii?Q?msoYQmLV3HbOdKYO9KHQMh8w/K565Fmgh0xjBZ1EHVC2B8SSvbrXjTTcfCz9?= =?us-ascii?Q?3HkLmIGCHKIjmjhvB7uHxd0nBdkYbPPwZUjENjPw+ZoDOJwEj8gIBPTPVXkx?= =?us-ascii?Q?2B+28wO6mcirfuvak12G2F8xqF8r4Nfh0bJnosK+pjPhvwm45nf7VSPJO00o?= =?us-ascii?Q?6Z5xu75gcl77ASOm3XhARFnYGSMMi62FwPbim1BKStqmcoWE6PKy+3L2+1tQ?= =?us-ascii?Q?k9wYqsCaJiT36PH8wNfwflzDiBfYcU3duzmiO8fwUAPEEBJxr7Y9Sft1ZrXM?= =?us-ascii?Q?wjvLYkvgHHk7rCvRZLjjdtvXKDn/VYO3rEjBDXyCX1L3NWX+HHFVlePr5Vrd?= =?us-ascii?Q?gOcNUvgRWzkpFmOVWvpLNe0akR6g7ZDXAJ6vvRyEQ75m1ATSEkRA+xFAEjwl?= =?us-ascii?Q?rOXh3/RKYRA8Jy9ybF4nI2i6tVZcCvrmIaravByD71yuLZk39SgSjzZPkLzk?= =?us-ascii?Q?iTBleGPKYORJRLRF+l+ZuQMajkEbbLFt0zJTN4aq5Jq8HWbj?= X-OriginatorOrg: atomlin.com X-MS-Exchange-CrossTenant-Network-Message-Id: 1b508c76-bb19-4236-b58c-08df048916dd X-MS-Exchange-CrossTenant-AuthSource: CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM X-MS-Exchange-CrossTenant-AuthAs: Internal X-MS-Exchange-CrossTenant-OriginalArrivalTime: 27 Aug 2026 22:18:15.1592 (UTC) X-MS-Exchange-CrossTenant-FromEntityHeader: Hosted X-MS-Exchange-CrossTenant-Id: e6a32402-7d7b-4830-9a2b-76945bbbcb57 X-MS-Exchange-CrossTenant-MailboxType: HOSTED X-MS-Exchange-CrossTenant-UserPrincipalName: IEh2VCgvJA/UNXwpCUwRyUlVymUwf1wIlpAjZOQO3PECikjPbO/MWtKlj+L1D3E53Bzc0CGgt4XrJquPfKoARw== X-MS-Exchange-Transport-CrossTenantHeadersStamped: LO6P123MB7095 Content-Type: text/plain; charset="utf-8" In print_cpu(), rq->curr is dereferenced locklessly to print the current task's PID via task_pid_nr(rq->curr). While accessing /sys/kernel/debug/sched/debug is inherently best-effort only; rq->curr is indeed expected to change dynamically while print_cpu() is executing. However, if the task currently running on the CPU exits concurrently and its reference count drops to zero, put_task_struct() calls call_rcu() to schedule __put_task_struct_rcu_cb(). Because print_cpu() does not hold the RCU read lock while dereferencing rq->curr, an RCU grace period can complete concurrently and free the task structure via free_task(), creating a potential use-after-free race condition. Resolve this by reading rq->curr using rcu_dereference(rq->curr) inside an RCU read-side critical section. Holding the RCU read lock delays the invocation of __put_task_struct_rcu_cb() until after rcu_read_unlock(), ensuring that the struct task_struct memory remains valid while being accessed. Fixes: 02968ccf7b80 ("sched: add /proc/sched_debug file") Reported-by: sashiko-bot Signed-off-by: Aaron Tomlin --- kernel/sched/debug.c | 5 ++++- 1 file changed, 4 insertions(+), 1 deletion(-) diff --git a/kernel/sched/debug.c b/kernel/sched/debug.c index 61932ef7ec4f..3d6248c50dad 100644 --- a/kernel/sched/debug.c +++ b/kernel/sched/debug.c @@ -1210,7 +1210,10 @@ do { \ P(nr_switches); P(nr_uninterruptible); PN(next_balance); - SEQ_printf(m, " .%-30s: %ld\n", "curr->pid", (long)(task_pid_nr(rq->curr= ))); + rcu_read_lock(); + SEQ_printf(m, " .%-30s: %ld\n", "curr->pid", + (long)(task_pid_nr(rcu_dereference(rq->curr)))); + rcu_read_unlock(); PN(clock); PN(clock_task); #undef P --=20 2.55.0 From nobody Mon Sep 28 02:57:20 2026 Received: from CWXP265CU010.outbound.protection.outlook.com (mail-ukwestazon11022130.outbound.protection.outlook.com [52.101.101.130]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D578237F8A1 for ; Thu, 27 Aug 2026 22:18:26 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.101.130 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787869109; cv=fail; b=hm8tvhXZnAzhBaP2W86vhM/XrCeZiXxTclerSetclFBno2WoVODvMp67niiqviY+IE1s5wlcJGtCmwfiC4I0wg7Rznxp05EpMjNmRcJLDyn9x9oEvZqX1+kRbGtku4wLdWEXifkFujmmAB+euzuH5Pbx2xcuKVfm5jc61mMBGY0= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787869109; c=relaxed/simple; bh=FAMBmEKNH0E3r93mDqrGDdpD1wt5J3bOKGCk+3QO3y0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: Content-Type:MIME-Version; b=uZHr/zoer9lkVAtzD0vRl8b/et/tcFm7IzsB/AaDYr0CO/0/wWbpm4quWRS4Ph/K/G9EPhytBfVmH7k/+XYKZ0wAnMv29yezX2jOXCFB3khW8NaTHVWUg3oV+LS0K82xk1VfrnmYo++m9ZbgLoeBZ+zATxu+rJDHyL3x/3JbLH8= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=atomlin.com; spf=pass smtp.mailfrom=atomlin.com; arc=fail smtp.client-ip=52.101.101.130 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=atomlin.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=atomlin.com ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=AKua2bZa1arP673e0CESkg83jbsyZN4fwk/3gtbHUsB2IP/TgHE98zW2sZ7c1+VEjXkBG5Gs+13WwU4WY+VYRRK3jq6/es2En+i/9f3bPGUNQkxEms1rD/7dlge7XmWoraCI4J4vR/e16CyzlAiCZr0sE55vg8wdfo4CxsEd32/2ReJBZwbwXTjis19wwYWqJBl33T4jIq4xK9S9gps3rrbsQo+28bmphgEy7/DmJeGeuOvHmqnGO/QE9wYfZ4a6eLYlYGWVOVkSuo0V2ppg5sc4IL1P2o3p009wEn8HI/9kEczbNngMBj4X1b8ODS1KmtzalPYpKT7s/CnTdXUeCg== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:MIME-Version; bh=8LrLKpfCtcUw4SDnojsAv51pIUjB6Rdt3xL3H1e6+mg=; b=JjQfatwZv5Lz+3hyfHIAnuqZJBC8v4QiGyJ32wjRPywkV/3XQJes3+XBjtwl3PAtahFbsWvUahZd/uNTOL8FZVPCssjfjgoevOiQkaZ9tkuk6lSwIo245qj9rYNWS/Le67A/9A3uQjc6pSgWgYuWMs30qMU3QxyNIXVoFYB801rCZGTLrATuMHPakCrEYi4doYKGHVLA0Ljl9/0PQIKAlnlSgPUgwpemAa4P1fTt0YvSEe8teZ+ndl4vbevLDzQYfTHIVoiZgxrK7CQ2HyCxz/he7oXuVdmuPeZMR6L6PknBpqhLNKODsG1/7oIoltTemYMq1UNvbiV5/tGn48YLQg== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass smtp.mailfrom=atomlin.com; dmarc=pass action=none header.from=atomlin.com; dkim=pass header.d=atomlin.com; arc=none Authentication-Results: dkim=none (message not signed) header.d=none;dmarc=none action=none header.from=atomlin.com; Received: from CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM (2603:10a6:400:183::5) by LO6P123MB7095.GBRP123.PROD.OUTLOOK.COM (2603:10a6:600:342::13) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.360.10; Thu, 27 Aug 2026 22:18:16 +0000 Received: from CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM ([fe80::cec4:77ab:262e:d230]) by CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM ([fe80::cec4:77ab:262e:d230%4]) with mapi id 15.21.0360.008; Thu, 27 Aug 2026 22:18:16 +0000 From: Aaron Tomlin To: mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org Cc: paulmck@kernel.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com, zhanxusheng1024@gmail.com, neelx@suse.com, atomlin@atomlin.com, chjohnst@mail.com, mproche@mail.com, sean@ashe.io, steve@abita.co, rishil1999@outlook.com, linux-kernel@vger.kernel.org Subject: [PATCH v9 4/6] sched/debug: Protect p->mm access in sched_show_numa() Date: Thu, 27 Aug 2026 18:18:06 -0400 Message-ID: <20260827221809.988394-5-atomlin@atomlin.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260827221809.988394-1-atomlin@atomlin.com> References: <20260827221809.988394-1-atomlin@atomlin.com> Content-Transfer-Encoding: quoted-printable X-ClientProxiedBy: LO4P123CA0256.GBRP123.PROD.OUTLOOK.COM (2603:10a6:600:194::9) To CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM (2603:10a6:400:183::5) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: CWLP123MB6607:EE_|LO6P123MB7095:EE_ X-MS-Office365-Filtering-Correlation-Id: eeffd775-79fa-44f1-57a1-08df048917c9 X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|376014|7416014|1800799024|366016|23010399003|10067099003|56012099006|22082099003|18002099003; X-Microsoft-Antispam-Message-Info: 4VUSY4qfJalO7kPkc8JNne2c3sdFB/eFxAOffXmZCR4kV4mcgrDA0dpuT9lR1/cl0zPospxapuJAZbTpBDwrdGEXW1StKJyerEYf7pPbLXY6aFco4XvLpY2GActpRcGkQDtmNw5vXnI9KBiQsxOeKyAMnHu2mMtFQD0kjS6d2H2x+0rMc0f6dR9Kc3nuYiPSREI/TZOmBX/T8/azPo1rdbH2XeZy5i14MiIYznOeJJKuQLGmfWM/f7bdJbAyrj8iR/Yvyd02KWtVM+aXq+c3hahd2LsZxb8aakk3BnBWQn9mP1nAliRylq07Z7kDav14WI5rOQUvBBV58O80hEXiSDzxQdm6i8ffnStm3gy7KeAovWgew1gsyx4vl/xKLa0WbsllXnrdyYQvu8YJHbr/eZ0Z29FMmw+isy5/MkM0XX9aFd0HwmcJfk8cH42PsIGsO5RPhKbnVo5jWlENHwI/0033Evpfo1NQ3TSwGXPB/0PBc7uNCJLH8bBhi9TMr1nz0mh+dR6R7p3ez0zy/VQY6rZCvGKPTFnQrjNdnFccdHmxTJGmk+fCkmz1gPEZ9HQCBS9v1wMlcnGhXUMMDFEG2UfpVqSetsN2jbzB1vIUhU2qu23oshuQFACc8DMLV3nh4XjfoAXgkmYCpTiwgsB0LOWcmrCdiSzd3oj2tyl9yB0= X-Forefront-Antispam-Report: CIP:255.255.255.255;CTRY:;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM;PTR:;CAT:NONE;SFS:(13230040)(376014)(7416014)(1800799024)(366016)(23010399003)(10067099003)(56012099006)(22082099003)(18002099003);DIR:OUT;SFP:1102; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: =?us-ascii?Q?ZoWr4hFfeEMDAvhoTKmf2QnwWRMqep9k4nj6B9pSj5Q5/UnHhPxGWnARViPc?= =?us-ascii?Q?ML7BQ1l50a/o3N0nFONpWb6SC5maKI8/ZFM4qAme7Bf783agcomffvqGv20p?= =?us-ascii?Q?AEW21h99G84cd82/MJwwKNm4z940rjTNxXAf3Bsmecje7yuxF3l6n0xFiaPR?= =?us-ascii?Q?NGJBy3hwQC1FoYr+jpCsBqiF5RidMU1pAOrdQvtUjImOq/P+1y9Thv9sOyx5?= =?us-ascii?Q?4e0Ebxg9pqwAvjtBENZEFVM4gxdkqSsgc68THODM+McXSbj9LQSEjqMoj+QJ?= =?us-ascii?Q?lNJ12eQs9h7o5udwbwB+mQvCT/7NTr6okQzaLUlUW2dWRhsCElf+yZYGPD+n?= =?us-ascii?Q?i37AQGJjel64LzPUEZua1obuvpFwlypMd/SY+dAm0+UeN/wTG6UHNfhuBoAn?= =?us-ascii?Q?ytTWGetIBOdz1ECQO/1eGEb322noJzKnTEMNaRfdjT8xE//JwbHGY7jFtcof?= =?us-ascii?Q?KjqdYsEksk3yvKzZ8um3ycmoBW24LenHFK/wSY+RIipiA4LUz0ph/Jc69peG?= =?us-ascii?Q?ZRA06tMCHZk7Bo6/II3GjF4uimOQE8fOyqNFJhhnsMzpec1dzW1IcKFaSAfN?= =?us-ascii?Q?Mh1+ea8ZjUMBx/uXmM5sbLfPQQmVCrjoBzNjNSdtBNem958JuSZhpLQgd4jw?= =?us-ascii?Q?kTavsgtJDsUp5wCZC2YqtPxmZaqsIgC1ziVWX4xBXyv0hqdWG11XE6RDOFeD?= =?us-ascii?Q?bi6kOWQG95FvTdZMkZARYUa9G553RDp+dtpRJ+HQup0M6WD82Kn3XlKBWhua?= =?us-ascii?Q?XZ/WP7ou8FCvqaeIbjyQGRVgfhLh8BUW6d2XH3FNcFsYc6XK5jQnQn18fHIp?= =?us-ascii?Q?x2qYjo1Z45yUd90UzFTccKJ1NJs8EDNuMpDlFY1I6eILj27YVoEHe32yGXDM?= =?us-ascii?Q?WFgMgcvP8l3lWe4zeUMf6dGGX2dfSRBN7/h82AoyaKMpNLWMddho48emKa1t?= =?us-ascii?Q?K4+Ixg0V8kC0spyG31k70IjkHWwS3V6a0TcEoy5gOb2jI0lWdQalo93ODVM5?= =?us-ascii?Q?O+OyS7yEfBpNQdLXFYkhiADFSsXDxqw+VO8+h9YG0lEbi9W7hd8OVWrmZ3wP?= =?us-ascii?Q?kPmwDsPCF5jBNPNAAj/xdCw4yjn+0IOfx3/zLNJGdieegRqC5wPb4L0h/khD?= =?us-ascii?Q?npf6igx80sD0aNz0wOGZtRvMPimDWDWhwfEDlssHMHt2E1WyhX4+Q8a+geVX?= =?us-ascii?Q?XMP42TL+GOMt+lmf2Ri/GGIVN/TFp4bvbm1J7af+7EZlxgSz2fAk4+ayrYkU?= =?us-ascii?Q?IkfkgZ7ROdwJkhwIEoZ60vU4mwj78LUoeNYG8R/ouB1CJRFv95UNzRRISl3p?= =?us-ascii?Q?E2EHrDcDO6F0ohTs9kqBjF0+lOAzA+SrZEJXZoSLJ6aGHRAJgo3BsppUp1vE?= =?us-ascii?Q?0nY86Qlq43wTPXrIVpfEG1iNbMoyQMX/yhUDDibn3RlwAModKo/T72CE7k9M?= =?us-ascii?Q?eUIaRoaqZ274FWae1+r9lE2hBfmDe4y9SZ8SN5oSfVcwmJbx2r1IUcVJfXYo?= =?us-ascii?Q?XPsnT9TPLUg+1H0pfv+kkka7b1NblMNEDRgRM+aVZStv9lnbMjbv7Ho6gOCr?= =?us-ascii?Q?QOfhYl71YEC4dek/oYY+87x1s09Mv243qBQCS1lOeovrtLR29ib2gUtNpuGk?= =?us-ascii?Q?d2TF4I+5kjOLrye7Z9d82m9tI7St21E3WrAS06gUWuxXzBndlx/a6I7mmhjU?= =?us-ascii?Q?rHT9PZCMIWKsr9vgknE8aLPBt8/7dM47GKq3t/M2eJkc/O5d?= X-OriginatorOrg: atomlin.com X-MS-Exchange-CrossTenant-Network-Message-Id: eeffd775-79fa-44f1-57a1-08df048917c9 X-MS-Exchange-CrossTenant-AuthSource: CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM X-MS-Exchange-CrossTenant-AuthAs: Internal X-MS-Exchange-CrossTenant-OriginalArrivalTime: 27 Aug 2026 22:18:16.6918 (UTC) X-MS-Exchange-CrossTenant-FromEntityHeader: Hosted X-MS-Exchange-CrossTenant-Id: e6a32402-7d7b-4830-9a2b-76945bbbcb57 X-MS-Exchange-CrossTenant-MailboxType: HOSTED X-MS-Exchange-CrossTenant-UserPrincipalName: WZJIZ3TLMx6twZEWx7jZPI48sG3R7FNKoyr1ghb+rNnB/+N7lEaTiXl6CAcLz5S5KgTJeB7KBfHbjpQfrX9mYA== X-MS-Exchange-Transport-CrossTenantHeadersStamped: LO6P123MB7095 Content-Type: text/plain; charset="utf-8" In sched_show_numa(), p->mm is checked locklessly and then passed to the P(mm->numa_scan_seq) macro. This presents both a time-of-change to time-of-use race and a potential use-after-free vulnerability. If a task exits concurrently via exit_mm(p), another CPU can set p->mm to NULL and call mmput(mm) to free the struct mm_struct. Dereferencing mm->numa_scan_seq without holding task_lock(p) can access freed memory if mmput() runs immediately after the check. Fix this by wrapping the p->mm check and macro dereference in task_lock(p) and task_unlock(p). In exit_mm(), current->mm is set to NULL under task_lock(p) before mmput() is called, guaranteeing that p->mm cannot be set to NULL or freed while task_lock(p) is held. Fixes: b32e86b4301e ("sched/numa: Add debugging") Reported-by: sashiko-bot Signed-off-by: Aaron Tomlin --- kernel/sched/debug.c | 2 ++ 1 file changed, 2 insertions(+) diff --git a/kernel/sched/debug.c b/kernel/sched/debug.c index 3d6248c50dad..1b6d2a83d50f 100644 --- a/kernel/sched/debug.c +++ b/kernel/sched/debug.c @@ -1395,8 +1395,10 @@ void print_numa_stats(struct seq_file *m, int node, = unsigned long tsf, static void sched_show_numa(struct task_struct *p, struct seq_file *m) { #ifdef CONFIG_NUMA_BALANCING + task_lock(p); if (p->mm) P(mm->numa_scan_seq); + task_unlock(p); =20 P(numa_pages_migrated); P(numa_preferred_nid); --=20 2.55.0 From nobody Mon Sep 28 02:57:20 2026 Received: from CWXP265CU010.outbound.protection.outlook.com (mail-ukwestazon11022130.outbound.protection.outlook.com [52.101.101.130]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 095F337268C for ; Thu, 27 Aug 2026 22:18:30 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.101.130 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787869114; cv=fail; b=RFbUEz2Dypx9CdBxc49T5BKOGIHbKYfm+J8lzrSYTaSOzBEpGFiEf2ZufKGLE68FU2Gqle9nafTNujbCM7ibQWa89qWf0F3/qKFsOOOoGbbk9sd7dwLCL4y1Nc/6nDlnK8L7NxJDAHwng6xMkWIRS9ysdpU4NYTS2nReeXw+RTg= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787869114; c=relaxed/simple; bh=I6UnR6/cucTJl+cQyspHyYbEHqlSglcT4O8dqtZ8eDs=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: Content-Type:MIME-Version; b=oeyD9Dk/wpg2HnBBhtbSJIkq9fG0ls1z/IUM+kgwZPakdFySB4cIJ1GUGfCDXiidwyWtA1+r0opqVmaMIa9Y4cy8i6MdAwdIRMU767Mf/i/xHEl65OEAvlkJBAYsdYdWGqnEsyuaeX+0hSUlPmCtoM7k98nq/hc8JTyoBveTqI4= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=atomlin.com; spf=pass smtp.mailfrom=atomlin.com; arc=fail smtp.client-ip=52.101.101.130 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=atomlin.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=atomlin.com ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=pAlg4sIciTzaSKZj5Z1ntr8f+fo9xaldJ1EVyeG47TkE+/XvabTopoRjuyiU3G/Jn88v9sXzWz9bU2n2yucLg227h4beRVwwC+X7UnwPTaEfjKcPeBCywQqPmVeDx+8GHbSjZUNa1N6/25/TI26bdkVC9IBSnbd9kj5ZWOGZF/L9VhD5B3/7oau5gkiSAJEhfh/fWrJu6hN5nODROn4gNVjxP8SZgMSAnYqxIXurrBtzN23C7Lw2z5PAjpiWM7on4XIq36zis/zZWk2x+MQthcdyXOJF0OXWmuLOMeCAuS2BBTZEqYCxMWFj4wnDFSZ/tPWPPrt3bwxAMvnkqeZsOg== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:MIME-Version; bh=A7/SgoIeLjemO3i865rgZpto0uVuDj9N+2XbLdcG0so=; b=KzmgZLsbch5Exd+8J5fmFo2D4fuQskMeD7KVgQX6y6AGxEOS/ybxiTx19VK6wbHwMZVgmryqtEDcAkJBwzsEJFhjSzwCtRFFp/ryzR6CWt+UXTJtZOgDXfFRM05OZG6XnzxLy6tFoXbQUk3Mi1SQHifcStxuRl7CCCW3GxI7YCXIe5BOrz/svS0t2w+EsMkqd35qGK7gesP24SHhWYXCE5bCm6QsUwqTNCsHSKg8XkDfQ+FXxcsPMuCBwZPYbEboKyBV+f0x1gm0TdUhkEls8BT9phqEfPKiGebGrGC6mSdW7E2UgfLfT9L2NAvJwmvjLcHTvnVUa3+fsQ0+qAi5gg== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass smtp.mailfrom=atomlin.com; dmarc=pass action=none header.from=atomlin.com; dkim=pass header.d=atomlin.com; arc=none Authentication-Results: dkim=none (message not signed) header.d=none;dmarc=none action=none header.from=atomlin.com; Received: from CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM (2603:10a6:400:183::5) by LO6P123MB7095.GBRP123.PROD.OUTLOOK.COM (2603:10a6:600:342::13) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.360.10; Thu, 27 Aug 2026 22:18:18 +0000 Received: from CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM ([fe80::cec4:77ab:262e:d230]) by CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM ([fe80::cec4:77ab:262e:d230%4]) with mapi id 15.21.0360.008; Thu, 27 Aug 2026 22:18:18 +0000 From: Aaron Tomlin To: mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org Cc: paulmck@kernel.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com, zhanxusheng1024@gmail.com, neelx@suse.com, atomlin@atomlin.com, chjohnst@mail.com, mproche@mail.com, sean@ashe.io, steve@abita.co, rishil1999@outlook.com, linux-kernel@vger.kernel.org Subject: [PATCH v9 5/6] sched/fair: Use list_for_each_entry_rcu() in print_cfs_stats() Date: Thu, 27 Aug 2026 18:18:07 -0400 Message-ID: <20260827221809.988394-6-atomlin@atomlin.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260827221809.988394-1-atomlin@atomlin.com> References: <20260827221809.988394-1-atomlin@atomlin.com> Content-Transfer-Encoding: quoted-printable X-ClientProxiedBy: LO4P302CA0039.GBRP302.PROD.OUTLOOK.COM (2603:10a6:600:317::12) To CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM (2603:10a6:400:183::5) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: CWLP123MB6607:EE_|LO6P123MB7095:EE_ X-MS-Office365-Filtering-Correlation-Id: 28a78688-27cf-47c2-e6eb-08df04891887 X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|376014|7416014|1800799024|366016|23010399003|10067099003|56012099006|6133799003|3023799007|22082099003|18002099003; X-Microsoft-Antispam-Message-Info: 94kDyjob8R30W08QpWPjDKjsmL6cutp9tam13FHD3fYq6LiSprnfCq8FoWu4TW6shyPh/M5v+wk/QFP3jZkWXDW3UNtc6H1ELEA7FeT+9I8STjLNHQMHqhP0sc7YVXNQokb70twtsZWFllqLif/4WF72aRjBFzB+zWlSDPuHX2zw52IS3kv5a36HhCGIu3ob9CFGgw64XhO2nShWhQpEWl7g/n/tWg1SqoiqYBwkzUAuKZWvrDoY3a0R3DwiR10rmRPaAAxZ+dfZ1feFkM3Pe1qbjJA7X7uPIz6QSCKABPPgrLFCAn6CYKN+aNhyKtjhasJ2oPbFxwPLXT1atiHqjqcpGWiZBBNDaGLbRsBDgRjuoh2JfOtWlKBkL5JAUWUM/W0kQlIssIXA/TK/ppkQTkZlcKBvIvHSZqWyOsfFa6ex2YN74ZZfxbhkciFuY3NuMfF8N/fggoX7AZstV7j6mmNhlw+NxKfsORt8aiIQ0U2H2y7Qc0aZxbgHvM7eKA8x8WV+k2fofgtgSyb6GyCgDetr02aOfQMKDgMFu9j7lmd6Hg3iFB2NeoTykRMqJz6twYUca4DRmj02ayrpytrQF+oaimKLPvY67eJiTdSXADUjGZwLbIk7mHDAp8CZ1ySPAFJj5Aacdh+/RJFKPxm9vCIMmCC40fnhDKp3x2pKjJA= X-Forefront-Antispam-Report: CIP:255.255.255.255;CTRY:;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM;PTR:;CAT:NONE;SFS:(13230040)(376014)(7416014)(1800799024)(366016)(23010399003)(10067099003)(56012099006)(6133799003)(3023799007)(22082099003)(18002099003);DIR:OUT;SFP:1102; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: =?us-ascii?Q?RvZ38U16pCG/yb+OAIKBPVzGuZbdzBFvaoNZ9tHEiiK4f2uAtg44QZqzqsMN?= =?us-ascii?Q?u5BCC6Bjzp6k27XV4q5dSzrnrjgqMvMEviSKBnTfStCqO/JRaSmU8hUiOjIl?= =?us-ascii?Q?hh5uVhnbeW4JxuQIdkMik5KuIBaEQGC4vTXjhmbrUVnq8IgfTmmHU1KC4H3O?= =?us-ascii?Q?s8yhdcgmOtL9/eLqHYflTUMFcd5DyjA0xXeJIlYu1nnry0MLhXbw+tnCC0cc?= =?us-ascii?Q?0w/zhEoC6wWbiD0bNbrfZw4L6GRqIhj/ohzvXapXU96HNnYIPD1+mMpIiuyK?= =?us-ascii?Q?Qe9+KaANiyIGIjCvMsJb2DyWj84REfcwi2oko+OwuZAMKq+2WqbLi/oBuf+4?= =?us-ascii?Q?Yskvrqj9j3V/77grNBW6j7S03yuXd4H4pdDeYR9n7tbDBPUAqzfzkdxPVbOI?= =?us-ascii?Q?4szSV2EK+T8VrSXk3v/mfhqqZB33e12PWApVXR0paxrXNNSoox7fX/t0Qj5g?= =?us-ascii?Q?M0Z81oQiuElfXN8ehzPmkiKnBleYnOCT1JJnBi+u7I5tPxO4L56kztrrUT2o?= =?us-ascii?Q?++04OCKrqCx8WC+7b0uq1tr+hrC3+HfGJMS7MFcem62KgwgAbctVm8D0+cN3?= =?us-ascii?Q?TjfTvYLasdtxFb1qFh22mAUInbPSHSk9dHCyy29x6fsoUAwJaHBWSN3jik05?= =?us-ascii?Q?uczE9OYnlfQZtzvYBCa/ibHVqYKlzcd6298sWgYKMKmmln4troBHqYen3/mo?= =?us-ascii?Q?d7nkIKjMesfZbb2aEpJl68J6M60+W2NmAYIm1gv5SiXsgeEXy+fQ82ffvgNg?= =?us-ascii?Q?b9CJr4zrRNX/d157CAZ1pNHa3vPBDoPft4Ubv/1hr9IzGkDciQQn1wq1R9b8?= =?us-ascii?Q?CxaS8rl3SOs+udUXda/Cl9QZDGBh4QM3bKA4ei06CaVMlFsEuPlOdj1oOwYG?= =?us-ascii?Q?qf5VKZQS7VnrpyENrIcmQdssX5L/+be9FFFxUfgjtlkqGfMfehICY5p72k3i?= =?us-ascii?Q?nS1Xn2LvJa8DDLIuGq13rsNfvuO3UClaxjLP0hJC6/7Wghaj3WWImum2s9bK?= =?us-ascii?Q?t2QKT4BWNtOR/ZsB93zaDfTBfcwPVlEQW4YTiRnznUianII6Esw+703ZT16e?= =?us-ascii?Q?MA6Q/0mX7w+modALQFP5KZ/UbowRLaBoqeQLfjw1cVoe8gjQwIw9hOFMBMf6?= =?us-ascii?Q?IVtFhw+yWkHbg8FWnIr1Y3Gh+noqVreKwHoWWa0//9kn0ZpsJBQ+96veOHFf?= =?us-ascii?Q?Ghx61TEsjnL+ysclQDK2F1Ty1ZaVcK+r/3uehZYJCrpW7Tr1QAcL+i2wORBX?= =?us-ascii?Q?34t0sU6CgvlnO4Of4LOvTmWATX/LqECFPL/001Ycgob1nLCLhhLca3sJXNiv?= =?us-ascii?Q?sTtJ8YNFWE1AuZ0G/9efYRNH5Yq8QUe4ekgddAtv2I1XMHvfBl7JUmK2azHH?= =?us-ascii?Q?P3ApgO0d767Y00ZY77ecEJPl4Dq0g796cil1WZVTVm4JpLCye4yRwqHie7if?= =?us-ascii?Q?rQjvXgB7EhphRb9JEdx2AHwjfCtG0f6PtkPjrzn6Se6eydHlFRa5j7JAXikF?= =?us-ascii?Q?3ppoP0VPzza2oj3iJ+00K0iz3vkwbrNiMUugi72JNNEp8BIgfq4kBQySwh3Q?= =?us-ascii?Q?iztA76f+Qfa9hrOSeNXh5et0uwuKWtazItiL8KJrI+mivAQHU14XjyPCEuPf?= =?us-ascii?Q?lG9rZ9c8dWywt0e8AWhdPk6+lQaok2SXC1Vpq0MXvYhcHrtbNiHaWypVoLFk?= =?us-ascii?Q?p+RbEAT7GwoFQ9V1VaVGwDqaSpP2ZNfdxjuTHgZG2LgTPOrT?= X-OriginatorOrg: atomlin.com X-MS-Exchange-CrossTenant-Network-Message-Id: 28a78688-27cf-47c2-e6eb-08df04891887 X-MS-Exchange-CrossTenant-AuthSource: CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM X-MS-Exchange-CrossTenant-AuthAs: Internal X-MS-Exchange-CrossTenant-OriginalArrivalTime: 27 Aug 2026 22:18:17.9439 (UTC) X-MS-Exchange-CrossTenant-FromEntityHeader: Hosted X-MS-Exchange-CrossTenant-Id: e6a32402-7d7b-4830-9a2b-76945bbbcb57 X-MS-Exchange-CrossTenant-MailboxType: HOSTED X-MS-Exchange-CrossTenant-UserPrincipalName: FXrU4Qd6IEGrjFBDQa914uZLeapYwaGgy+1HBxLLCqSuOIiY8gZ/pHo1IzWNAFL7/Jamg+zF4VUIh0hP5B5nVQ== X-MS-Exchange-Transport-CrossTenantHeadersStamped: LO6P123MB7095 Content-Type: text/plain; charset="utf-8" In print_cfs_stats(), rq->leaf_cfs_rq_list is traversed locklessly under RCU using for_each_leaf_cfs_rq_safe(), which expands to list_for_each_entry_safe(). Although rq->leaf_cfs_rq_list is modified using list_add_rcu(), list_for_each_entry_safe() is a non-RCU iteration macro that dereferences pointer links without READ_ONCE() and pre-fetches the next pointer without memory ordering guarantees. Without READ_ONCE(), the compiler is free to re-fetch pointers or reorder instructions. As a result, a lockless reader can observe a newly inserted cfs_rq's pointer before its internal fields are fully visible, leading to reading uninitialised data or dereferencing invalid pointers. Furthermore, in the core scheduler, cfs_rq structures are embedded in struct task_group and are enqueued/dequeued dynamically on each CPU's leaf_cfs_rq_list during task wakeups and throttling. Because this occurs in atomic fast paths under rq->lock, deferring list deletion with an RCU grace period or allocating dynamic proxy nodes is not feasible. While the underlying task_group/cfs_rq memory backing each node is safely reclaimed via call_rcu() (i.e., sched_free_group_rcu()), immediate node re-insertion can modify cfs_rq->next while print_cfs_stats() is executing locklessly. Under continuous list churn, lockless readers could experience backward jumps, resulting in unbounded list traversal inside the RCU read-side critical section and triggering an RCU CPU stall warning. Address this by: 1. Introducing for_each_leaf_cfs_rq_rcu(), which expands to list_for_each_entry_rcu(). This enforces READ_ONCE() and proper data-dependency ordering on all architectures during list traversal 2. Using guard(rcu)() to ensure the backing task_group/cfs_rq memory remains valid throughout traversal 3. Capping the lockless list traversal with a per-CPU circuit-breaker ceiling. This represents a generous upper bound for active leaf cfs_rqs on an individual core while guaranteeing loop termination and preventing RCU stalls under list churn 4. Emitting an explicit truncation notice if the ceiling is ever reached, ensuring transparency in debugfs Fixes: 039ae8bcf7a5 ("sched/fair: Fix O(nr_cgroups) in the load balancing p= ath") Reported-by: sashiko-bot Signed-off-by: Aaron Tomlin --- kernel/sched/debug.c | 39 +++------------------------------ kernel/sched/fair.c | 33 ++++++++++++++++++++++++---- kernel/sched/sched.h | 51 ++++++++++++++++++++++++++++++++++++++++++++ 3 files changed, 83 insertions(+), 40 deletions(-) diff --git a/kernel/sched/debug.c b/kernel/sched/debug.c index 1b6d2a83d50f..ffc55b89801a 100644 --- a/kernel/sched/debug.c +++ b/kernel/sched/debug.c @@ -11,17 +11,6 @@ #include #include "sched.h" =20 -/* - * This allows printing both to /sys/kernel/debug/sched/debug and - * to the console - */ -#define SEQ_printf(m, x...) \ - do { \ - if (m) \ - seq_printf(m, x); \ - else \ - pr_cont(x); \ - } while (0) =20 /* * Ease the printing of nsec fields: @@ -933,38 +922,16 @@ static void print_cfs_group_stats(struct seq_file *m,= int cpu, struct task_group #endif /* CONFIG_FAIR_GROUP_SCHED */ =20 #ifdef CONFIG_CGROUP_SCHED -static DEFINE_SPINLOCK(sched_debug_lock); -static char group_path[PATH_MAX]; +DEFINE_SPINLOCK(sched_debug_lock); +char sched_debug_group_path[PATH_MAX]; =20 -static void task_group_path(struct task_group *tg, char *path, int plen) +void task_group_path(struct task_group *tg, char *path, int plen) { if (autogroup_path(tg, path, plen)) return; =20 cgroup_path(tg->css.cgroup, path, plen); } - -/* - * Only 1 SEQ_printf_task_group_path() caller can use the full length - * group_path[] for cgroup path. Other simultaneous callers will have - * to use a shorter stack buffer. A "..." suffix is appended at the end - * of the stack buffer so that it will show up in case the output length - * matches the given buffer size to indicate possible path name truncation. - */ -#define SEQ_printf_task_group_path(m, tg, fmt...) \ -{ \ - if (spin_trylock(&sched_debug_lock)) { \ - task_group_path(tg, group_path, sizeof(group_path)); \ - SEQ_printf(m, fmt, group_path); \ - spin_unlock(&sched_debug_lock); \ - } else { \ - char buf[128]; \ - char *bufend =3D buf + sizeof(buf) - 3; \ - task_group_path(tg, buf, bufend - buf); \ - strcpy(bufend - 1, "..."); \ - SEQ_printf(m, fmt, buf); \ - } \ -} #endif =20 static void diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 687999312a7d..c6b25ac2b844 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -416,6 +416,10 @@ static inline void assert_list_leaf_cfs_rq(struct rq *= rq) list_for_each_entry_safe(cfs_rq, pos, &rq->leaf_cfs_rq_list, \ leaf_cfs_rq_list) =20 +#define for_each_leaf_cfs_rq_rcu(rq, cfs_rq) \ + list_for_each_entry_rcu(cfs_rq, &(rq)->leaf_cfs_rq_list, \ + leaf_cfs_rq_list) + /* Do the two (enqueued) entities belong to the same group ? */ static inline struct cfs_rq * is_same_group(struct sched_entity *se, struct sched_entity *pse) @@ -469,6 +473,9 @@ static inline void assert_list_leaf_cfs_rq(struct rq *r= q) #define for_each_leaf_cfs_rq_safe(rq, cfs_rq, pos) \ for (cfs_rq =3D &rq->cfs, pos =3D NULL; cfs_rq; cfs_rq =3D pos) =20 +#define for_each_leaf_cfs_rq_rcu(rq, cfs_rq) \ + for (cfs_rq =3D &rq->cfs; cfs_rq; cfs_rq =3D NULL) + static inline struct sched_entity *parent_entity(struct sched_entity *se) { return NULL; @@ -15580,14 +15587,32 @@ DEFINE_SCHED_CLASS(fair) =3D { #endif }; =20 +#define SCHED_DEBUG_MAX_ITER 4096 +#define SCHED_DEBUG_TRUNCATED_MSG \ + "stats truncated at " __stringify(SCHED_DEBUG_MAX_ITER) " iterations\n" + void print_cfs_stats(struct seq_file *m, int cpu) { - struct cfs_rq *cfs_rq, *pos; + struct cfs_rq *cfs_rq; + int max_iter =3D SCHED_DEBUG_MAX_ITER; =20 - rcu_read_lock(); - for_each_leaf_cfs_rq_safe(cpu_rq(cpu), cfs_rq, pos) + guard(rcu)(); + for_each_leaf_cfs_rq_rcu(cpu_rq(cpu), cfs_rq) { + if (--max_iter < 0) { + SEQ_printf(m, "\n"); + if (IS_ENABLED(CONFIG_FAIR_GROUP_SCHED)) { + SEQ_printf_task_group_path(m, cfs_rq_tg(cfs_rq), + "cfs_rq[%d]:%s ... " + SCHED_DEBUG_TRUNCATED_MSG, + cpu); + } else { + SEQ_printf(m, "cfs_rq[%d]: " + SCHED_DEBUG_TRUNCATED_MSG, cpu); + } + break; + } print_cfs_rq(m, cpu, cfs_rq); - rcu_read_unlock(); + } } =20 #ifdef CONFIG_NUMA_BALANCING diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h index 6d45e67bcdc3..c70807f17104 100644 --- a/kernel/sched/sched.h +++ b/kernel/sched/sched.h @@ -777,6 +777,18 @@ struct cfs_rq { #endif /* CONFIG_FAIR_GROUP_SCHED */ }; =20 +#ifdef CONFIG_FAIR_GROUP_SCHED +static inline struct task_group *cfs_rq_tg(struct cfs_rq *cfs_rq) +{ + return cfs_rq->tg; +} +#else +static inline struct task_group *cfs_rq_tg(struct cfs_rq *cfs_rq) +{ + return NULL; +} +#endif + #ifdef CONFIG_SCHED_CLASS_EXT /* scx_rq->flags, protected by the rq lock */ enum scx_rq_flags { @@ -3413,6 +3425,45 @@ extern struct sched_entity *__pick_root_entity(struc= t cfs_rq *cfs_rq); extern struct sched_entity *__pick_first_entity(struct cfs_rq *cfs_rq); extern struct sched_entity *__pick_last_entity(struct cfs_rq *cfs_rq); =20 +/* + * This allows printing both to /sys/kernel/debug/sched/debug and + * to the console + */ +#define SEQ_printf(m, x...) \ +do { \ + if (m) \ + seq_printf(m, x); \ + else \ + pr_cont(x); \ +} while (0) + +#ifdef CONFIG_CGROUP_SCHED +extern spinlock_t sched_debug_lock; +extern char sched_debug_group_path[PATH_MAX]; +extern void task_group_path(struct task_group *tg, char *path, int plen); + +#define SEQ_printf_task_group_path(m, tg, fmt...) \ +{ \ + if (spin_trylock(&sched_debug_lock)) { \ + task_group_path(tg, sched_debug_group_path, sizeof(sched_debug_group_pat= h)); \ + SEQ_printf(m, fmt, sched_debug_group_path); \ + spin_unlock(&sched_debug_lock); \ + } else { \ + char buf[128]; \ + char *bufend =3D buf + sizeof(buf) - 3; \ + task_group_path(tg, buf, bufend - buf); \ + strscpy(bufend - 1, "...", sizeof("...")); \ + SEQ_printf(m, fmt, buf); \ + } \ +} +#else +static inline void +SEQ_printf_task_group_path(struct seq_file *m, struct task_group *tg, + const char *fmt, ...) +{ +} +#endif + extern bool sched_debug_verbose; =20 extern void print_cfs_stats(struct seq_file *m, int cpu); --=20 2.55.0 From nobody Mon Sep 28 02:57:20 2026 Received: from CWXP265CU010.outbound.protection.outlook.com (mail-ukwestazon11022130.outbound.protection.outlook.com [52.101.101.130]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 191723921C6 for ; Thu, 27 Aug 2026 22:18:35 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.101.130 ARC-Seal: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787869123; cv=fail; b=ND0mm1PwVL1+HVJyrJ6NcFn7PVcEwIMkxvOcssTfohM5WxVeszyxUdC9WRT0eKrRMHv1AeXdv6K1bTXAWYeIIvX7qDJT0ZMTuh71QI0HvpRg/K73qw6OeOqd3UAOhyH5itKihqx3Cis82woqFPy8iZayegpztzRG7p0oXgEikYU= ARC-Message-Signature: i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787869123; c=relaxed/simple; bh=J7pwT1edrOVUk/npmmI7+krU2MCAyo/qiiNK5NBCXNM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: Content-Type:MIME-Version; b=SdScVyx5QFTxuqVtvViOn4YZhh4IAmCsHn+g6Js2pjU125kkckoDRyDeKcURkEqDP3aIUVmo9KYEcNutjma4p8IQyEKBMdRVDz5wYuR0awrP7MBzpWGX7ccqYtdfRXVW4z6sSA0trQmSt7dWavEboktXIuoZQx+QusPBW06qPzQ= ARC-Authentication-Results: i=2; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=atomlin.com; spf=pass smtp.mailfrom=atomlin.com; arc=fail smtp.client-ip=52.101.101.130 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=atomlin.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=atomlin.com ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=MM8ZOJCwyZymeZK11AiL7vCTqFy5QspiZKK9+F7+8m9uovafshbf5LMmKgp+K155s0wkv96GSYVHEh45/jV7Slhpw7cd2p3qS5JlEklFjTrITyMNRPdepEgE8TXe2lTwwPucuLdZHSrRQN+1/fonga/g7OTNiFmkG0JiLXhnszBM3ugG9Exrz2jALLwltZUfBbfOYSiX7PYBOKO9HIUWNdZDcke2p7zDioonQ6d+Yf3R+DMpUWAZuc03wGmHMCP9jqEvfAnLKJudK5x7BrQvK26wo+09ErH0LXiQdbtQaC7cvXfVtfaj/qyoppkswTGHVtjtGhqVJFNy+2LMS8ik3w== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:MIME-Version; bh=rZ1aBVvW2tdu5aY4cnFprHRqoBAJnLIy8xFYAwiA1Vk=; b=hYpNDl5otNo1Iz+byP+GoDc/7gDaMe20JraiwQOFHU1Qeqwqpv63w+Dei68Ghhsw1t7Xrn1fQkmUzs+qkKkzA9IWaObYHslYbi/TOvwMsLpyokMpoNHspAJSC5gfRC5XGVsdqN8R5Bmd4CBEz3SmgEBHBHFoo9GLCEIvDYY0Zt+zErR66DPSFlL8JzweerRoE8m1i8G6m6N2jc2ruFqViedMSp/ViGSVeW1ANLGydbiz5Jhj5BGSvNdgedEzx2/m3OBG900ZGCeLwxYBjacf1Mm/gen9gNKfq1QGpsBGGPFT0Yl5f7BiyDWaRf/B5ZIDKU/TeWXmb9xKBHTarqRZQg== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass smtp.mailfrom=atomlin.com; dmarc=pass action=none header.from=atomlin.com; dkim=pass header.d=atomlin.com; arc=none Authentication-Results: dkim=none (message not signed) header.d=none;dmarc=none action=none header.from=atomlin.com; Received: from CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM (2603:10a6:400:183::5) by LO6P123MB7095.GBRP123.PROD.OUTLOOK.COM (2603:10a6:600:342::13) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.360.10; Thu, 27 Aug 2026 22:18:19 +0000 Received: from CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM ([fe80::cec4:77ab:262e:d230]) by CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM ([fe80::cec4:77ab:262e:d230%4]) with mapi id 15.21.0360.008; Thu, 27 Aug 2026 22:18:19 +0000 From: Aaron Tomlin To: mingo@redhat.com, peterz@infradead.org, juri.lelli@redhat.com, vincent.guittot@linaro.org Cc: paulmck@kernel.org, dietmar.eggemann@arm.com, rostedt@goodmis.org, bsegall@google.com, mgorman@suse.de, vschneid@redhat.com, kprateek.nayak@amd.com, zhanxusheng1024@gmail.com, neelx@suse.com, atomlin@atomlin.com, chjohnst@mail.com, mproche@mail.com, sean@ashe.io, steve@abita.co, rishil1999@outlook.com, linux-kernel@vger.kernel.org Subject: [PATCH v9 6/6] sched/debug: Introduce per-CPU debugfs files Date: Thu, 27 Aug 2026 18:18:08 -0400 Message-ID: <20260827221809.988394-7-atomlin@atomlin.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260827221809.988394-1-atomlin@atomlin.com> References: <20260827221809.988394-1-atomlin@atomlin.com> Content-Transfer-Encoding: quoted-printable X-ClientProxiedBy: LO4P123CA0092.GBRP123.PROD.OUTLOOK.COM (2603:10a6:600:191::7) To CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM (2603:10a6:400:183::5) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: CWLP123MB6607:EE_|LO6P123MB7095:EE_ X-MS-Office365-Filtering-Correlation-Id: 7c7d27e4-b8d9-49ef-a785-08df0489194e X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|376014|7416014|1800799024|366016|23010399003|10067099003|56012099006|5023799004|3023799007|22082099003|18002099003; X-Microsoft-Antispam-Message-Info: oX3QoN6gAPTxTSi1DV/lsbNIlcesV0ttNTeRJccVks7hBvcRvmTucmJLosG+buRFMEPz5GrK8uCnShwIRXvghwgYTDmYzcruxI0b8z70qv4EVM8xDdU+u7cJ5Bnh5ZxR329asI/2AEoFphoYVGDNDRQDXIAlLm+SPg+l11/VEWYxKT+HaviuHBfk4lFrDsdfVGclAzpXNxOQfe7ZqlnFIcfo9SlYxxH/+hI05ktWIoyDtDkbRA+ZP4aV83ffAXE7Golawn4HCBJmV2IoTnRj0urXfDmaU3VZi3JVfb6RHKuFuiN8HhjNCMHUtK3H+MXp+uUno8VDWiDCXrSkZnHBBHTAZZbWJmDdVFXUXBgAdUVV3cJ82qF8Ls79T7nQvAl63RtaxgYejgLjDXKwI1VQJoaWfIDS3wiUzydx6w98XNhYuqiRC1Hq3TPtObGAySKVqiAUv3a4pqmeHqVSJlGRqwkVkn28mAnW5UX5C1FRX3G56KAX03w1/ZeGAOZikvJIdjuu+RRC/pXTF0J/a0fhJa0itWle4lo8kJ8kVIGvrLdDAujKiSGi9qjZ/xLz7sKqvDN/5ARr1h9ASqnvszOa6NG1sYzJLSoY3ZuxMIkfPnqtNLyysujW0s2mwTyaNCljrf0IEJkNMKsJkcIAi8ATP59pJgu8fayZ6SRRdcSoLXs= X-Forefront-Antispam-Report: CIP:255.255.255.255;CTRY:;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM;PTR:;CAT:NONE;SFS:(13230040)(376014)(7416014)(1800799024)(366016)(23010399003)(10067099003)(56012099006)(5023799004)(3023799007)(22082099003)(18002099003);DIR:OUT;SFP:1102; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: =?us-ascii?Q?HQJbKzXlQHb+si2Zuw6FVQ8ZsYtFplWXwMIuayLXpGFxQbOR5rmj95MjBWE8?= =?us-ascii?Q?vFdKzSKdSsJzfnYS31TJYxOkmedYa1ZqD2jB1XynOGrljrKv5zU9+KAL88Z/?= =?us-ascii?Q?IBZWPNm9YrdIcspvSYTMAuCQIxURjfpLwrkvYEEm+Qoxz3FD+PqSGLua/Q6w?= =?us-ascii?Q?Aae/SheI4Ab3QvDDHJxuLNwd5YyyjWFDrQEhfZabu9AXn4YcOwMU1ME94FUW?= =?us-ascii?Q?tNyjaS8zjptFqEbc/Ac4gc4LREqUFghkEji09l9M79+2lfsuQ+q8OWIfVYbJ?= =?us-ascii?Q?eQSe5bOq3i0KC/vTijal4OXqU/SHx65BgHv7c/5urzG4vI6i34jMJA2aQvkG?= =?us-ascii?Q?KDvmTsnh7o7JwJf5FL3pqdqLeCR1qRWncHUeR8OyNj0tzecaE/NTYCRCMHtL?= =?us-ascii?Q?6fBbtWpTI4vdOjArco05B5PF7W9+f1a1C+aB7wkQYbHt77DAefCbeZP6il/9?= =?us-ascii?Q?CFLv2cbf/7TlJy9aHgQbFNqvmL7CYWp804zubWCEunpejXXzRBiEUTUFW9At?= =?us-ascii?Q?IUAJIg/9kcNSSlZxKmbURE4pVgdLHlQOKVeqb8XpQbVVhfQCYnxv/Ad+6Q+E?= =?us-ascii?Q?xU9NcjcNrHDrwMAq2LCarST3Qiy1RGKH46H+N/e5OFrk0K8P7bzACb7erzCs?= =?us-ascii?Q?/LzbqvtlxYa2oxFoeV1GhxT/0oqqkeRfYi5QYgvlPsbFCImRCTNbfCA71Niv?= =?us-ascii?Q?bqJFM9kbWJaPCyFsQ+Jzx63wFPzr2bArqE2XpAE170cTzQX/i0crz/L6CDG3?= =?us-ascii?Q?UhG7YIo+ivLpKARcl5/GSzTpMXK19Bdk10xMB7f9El6upi2Wj1dN0ar6aY7C?= =?us-ascii?Q?+DTDhLvKURI8MDROs2RibM6uhMQ/Z+OUKy9LzypDziSdgsN6ketT+ESWEJK1?= =?us-ascii?Q?/wCLZdLq5vJUWm4nWQk0UhuI7m3XnicKQv3hcUWMNV1u3w1JLzruC6VL4rkp?= =?us-ascii?Q?uYRt4ZpicA4Iw9t4OsD+5G+Oa2nTWjx0aiYiRsSf/UKGDcujGodzveB+EfqA?= =?us-ascii?Q?DQGMR02VaVTSWv/LNnRZs7fO2Xlj4BTYcc43/QBR8ifUgUIu77MXjCgtnUs/?= =?us-ascii?Q?bTArkFZkiq4+0AnXFhJze0uE82ow6ncWjAOpUi5pv13RGI4BKPfdZMt9TRTZ?= =?us-ascii?Q?biLnXX+jLrbAmzeKXeVKwyu586maWYzg9iDa9LgVJz7e7o+wauhWcsSMlu4Y?= =?us-ascii?Q?Puw0kUryjp1Wr/8qp4PkgTDAzhcA6EStiYbzp+qF3fX7o8/LaX6/EN0g4hBf?= =?us-ascii?Q?m+XU0AkmlQMhfk6e0BXUIMZEtoBY1vD70C+h5TGz2VKCyHskmZM5y2RtEYgM?= =?us-ascii?Q?iYFA2qiyE6ijivHVxcoZYG9/6g0lk2GBumKj+qasxGHovbR61zqEnfmU9xLY?= =?us-ascii?Q?TX04bBhGfYZIQrawsFj/sxihH0TrM0q0ffORdiDUPCfh8+ygM22yLrlpLuqQ?= =?us-ascii?Q?YbtZE+x1r39qK0/NkArp1jv83QRWPlpyc1neqxY0WbBP3z0pEI9yK1oqNIcE?= =?us-ascii?Q?5bLYl4CjuNkeTZcuLShAx5A4Ixfxh3rgn0NxSqgxq/cF6/XPLuW4e/P3RT3W?= =?us-ascii?Q?3wnAhW3l2WS5NiDOTudfhkt5GssEobWYlaFAmgX5Q9CfkVmcezWi/jiwUS9V?= =?us-ascii?Q?Bc3mkxtxtg5f50mKoNf9/AVXOK22aGp1+GuDWQFjjce7x+NppqDWijGY25Fz?= =?us-ascii?Q?CYXD29tmkT9irt9L3Dpg/VgKElfLQmjbEnZp2OFHN7R/aPKa?= X-OriginatorOrg: atomlin.com X-MS-Exchange-CrossTenant-Network-Message-Id: 7c7d27e4-b8d9-49ef-a785-08df0489194e X-MS-Exchange-CrossTenant-AuthSource: CWLP123MB6607.GBRP123.PROD.OUTLOOK.COM X-MS-Exchange-CrossTenant-AuthAs: Internal X-MS-Exchange-CrossTenant-OriginalArrivalTime: 27 Aug 2026 22:18:19.2873 (UTC) X-MS-Exchange-CrossTenant-FromEntityHeader: Hosted X-MS-Exchange-CrossTenant-Id: e6a32402-7d7b-4830-9a2b-76945bbbcb57 X-MS-Exchange-CrossTenant-MailboxType: HOSTED X-MS-Exchange-CrossTenant-UserPrincipalName: YWNS3iFglDZZYgyscxR5sv/hJU7mW7VXTC0TUdGvTlT4bnczATTsCVF1yt7paNT7DXGytK3Q1T6l0aH3V+N3jg== X-MS-Exchange-Transport-CrossTenantHeadersStamped: LO6P123MB7095 Content-Type: text/plain; charset="utf-8" Currently, accessing scheduler debugging details for a specific CPU requires reading /sys/kernel/debug/sched/debug, which outputs information for all online CPUs. When investigating a latency anomaly or scheduling issue isolated to a specific CPU, accessing /sys/kernel/debug/sched/cpu/cpu/debug provides an immediate, targeted view of that runqueue. Add support for per-CPU debug files under: /sys/kernel/debug/sched/cpu/cpu/debug. Reading /sys/kernel/debug/sched/cpu/cpu/debug calls print_cpu() specifically for CPU , exposing CPU-specific runqueue details on demand. If the target CPU is currently offline, reading its file returns -ENODEV. Signed-off-by: Aaron Tomlin --- kernel/sched/debug.c | 43 +++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 43 insertions(+) diff --git a/kernel/sched/debug.c b/kernel/sched/debug.c index ffc55b89801a..0c1122007543 100644 --- a/kernel/sched/debug.c +++ b/kernel/sched/debug.c @@ -345,6 +345,7 @@ static const struct file_operations sched_verbose_fops = =3D { }; =20 static const struct seq_operations sched_debug_sops; +static void print_cpu(struct seq_file *m, int cpu); =20 static int sched_debug_open(struct inode *inode, struct file *filp) { @@ -698,6 +699,47 @@ static const struct file_operations sched_cgroup_fops = =3D { }; #endif =20 +static int sched_debug_cpu_show(struct seq_file *m, void *v) +{ + unsigned long cpu =3D (unsigned long) m->private; + + if (!cpu_online(cpu)) + return -ENODEV; + + print_cpu(m, cpu); + return 0; +} + +static int sched_debug_cpu_open(struct inode *inode, struct file *filp) +{ + return single_open(filp, sched_debug_cpu_show, inode->i_private); +} + +static const struct file_operations sched_debug_cpu_fops =3D { + .open =3D sched_debug_cpu_open, + .read =3D seq_read, + .llseek =3D seq_lseek, + .release =3D single_release, +}; + +static __init void debugfs_cpu_init(void) +{ + struct dentry *d_cpu_dir; + unsigned long cpu; + char buf[16]; + + d_cpu_dir =3D debugfs_create_dir("cpu", debugfs_sched); + + for_each_possible_cpu(cpu) { + struct dentry *d_cpu; + + snprintf(buf, sizeof(buf), "cpu%lu", cpu); + d_cpu =3D debugfs_create_dir(buf, d_cpu_dir); + + debugfs_create_file("debug", 0444, d_cpu, (void *) cpu, &sched_debug_cpu= _fops); + } +} + static __init int sched_init_debug(void) { struct dentry __maybe_unused *numa, *llc; @@ -759,6 +801,7 @@ static __init int sched_init_debug(void) #ifdef CONFIG_SCHED_CLASS_EXT debugfs_ext_server_init(); #endif + debugfs_cpu_init(); =20 return 0; } --=20 2.55.0