From nobody Sun Jul 26 01:47:31 2026 Received: from mail-pl1-f175.google.com (mail-pl1-f175.google.com [209.85.214.175]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 025B6366DC1 for ; Fri, 10 Jul 2026 03:28:33 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.175 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783654115; cv=none; b=LHTUXSGqJrArT5FycXusZWi/jAxlzsBOnn/aJfa8Mezt41KcyNqapUj2YmLGrJ9obfifysiZa/4kM4bvBEhA2G0Mu9fJSLEiVzBbCJcjv0pT+K88orfU34ckrDHIqr0wYMwthpDQ6W7DLCMpEAL4sEzhxZ4sQ/UPshO3lYrSE+M= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783654115; c=relaxed/simple; bh=OWzkBZj94kwxwBi57rZUXfo+HI1CfaJQHHKsPhN/SsE=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=WpPJ/bCFYePOZrHcGLyWdaeLwTzqwGU8LdC121So3ifKoAx/2W6Rty/vsuHNxyQrCMyQ2yupKRcqk2eFM4Qb252X6ok7K1UoL7gNScPUDrnGjk8wQGGnJMtafQMu8hoYjzWfFTKbC28c3F3tKXaUbOL1zBKtmWJ3ityhMaaRX7U= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=bmpmGQCT; arc=none smtp.client-ip=209.85.214.175 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="bmpmGQCT" Received: by mail-pl1-f175.google.com with SMTP id d9443c01a7336-2c7c61b5292so7019055ad.0 for ; Thu, 09 Jul 2026 20:28:33 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1783654113; x=1784258913; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=w+HqsKCBa8kVrZBI4dYL0PYb+R87jVrp5SVVk7foZ+g=; b=bmpmGQCT3Rvh1t88mGtFGv1b+QNcN2ItJcPpVGzoeLFqRc3MXc3TZIsYV2FvuQl2XD aHbRq2AmxZ+CNrPgfGCOv1qker64KXHBtoKD6GI7TYV+vIIbYUDCs1rQrmZcpFL7RDms D00BpVRBDBGdVYW1hNooulbEQU+nkyegX/TKNtK/v41BhqDcJpvpkmnFtxe5AqWPaqDv e2SHFn1Qn1xN+r/m97sqC4oh2adTCT577O8jD0xACHlltBou0bGEwEA7DNK+LnFj+P1j s+aP4j0wh7tv4bnSr3bZQkS0/xUiESFPTAirLHSBemdqs7A2Z+18rcnrV6sm2Tj/uvhD DaIg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783654113; x=1784258913; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=w+HqsKCBa8kVrZBI4dYL0PYb+R87jVrp5SVVk7foZ+g=; b=XbGzj6f2Z45Ycip0squjmpnGc+O4pzXCNEOVqBcCZ/5iVZKgAwVedenO4JBCdpF0ED HsKdKnj7LdahQ/VLEA9ZaP7wBuxt6EG74bQwvmqMPMC/o6SKhaFmSgG4GmFoP7x5tbYi RZfpXlKFFru0q4niN6rx0+86fmdcHBCYakWuGSaKj2GwuF+U8bilPaRuluoPkuLPXqr7 3R0mZ/u+wabmKtee+08ctB27s9IVbYh2kAzslupOj0eYAFDIyp3Qw06O9v2T5e+h8ZEt mJr8K09QeRfWcFAoOP/2ckT8vnBE//sZSXuXPHaY8rRahexziMMBZPhC5xj8GGxN/gUA JO8A== X-Forwarded-Encrypted: i=1; AHgh+RrbU23yDJ/d08HP1yei+ngSKtF2+0+AqqDcesUPk9EuB76YnZMjSAz07FjYLKDxxOUJ2y51s1h4mLfjvtk=@vger.kernel.org X-Gm-Message-State: AOJu0YzJ/P0fghQwsOFHQsYs3WnFawLkxQTgAVz2IRFwYB9Q93wn3YMz nu2gvAmM/mhEPkUp1Uc2b5cKBkL0djjuyaBpxS+ery06ePRcEF6cesye X-Gm-Gg: AfdE7cma1rOjnMAkgfj2QTxR6pjHh0E/MNwv9ypbfxHcMU32IQb+ikfZCHPiaU0bl37 fKyTH2vC7StXoTgmf3fRPGanPPD1KQwa/dCpkyR/9EwiurK7ZN4H7lNyul9X98H0ca4p57tyspM ZzJc9LI0CziRvPoAIvGUHb9NjMDdCEejicbts7tXImrHzslhTtYXPH3htixvgHPmhsRrl2RWOOL QSHCsuugwQNvLovzqPNgx8jrOL654S0V6pj9wWOsmcZeaqfeYO58mD/sGlJ07N0qdvWT+FffPlN 9NLsNb83EC9n3KTZ6fCVIeeEuK39KttEQVzGf9j46r2Iv7nx2mXImVxBhGFfqO5cS07DMy3AKOg PyMPMje2L6xA3+520iLLfQ3dECjXZzxItzxA9mAqys7A4+VVSj11AVkoi3iiEzq3Gdw28XKSJ/j qjPQ4PADWX0UM= X-Received: by 2002:a17:902:ce03:b0:2cc:ee78:3236 with SMTP id d9443c01a7336-2ccee7832d2mr96667315ad.33.1783654112959; Thu, 09 Jul 2026 20:28:32 -0700 (PDT) Received: from [127.0.1.1] ([138.199.21.246]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2ccc9bdb56fsm53436465ad.15.2026.07.09.20.28.24 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 09 Jul 2026 20:28:32 -0700 (PDT) From: Jing Wu Date: Fri, 10 Jul 2026 11:28:12 +0800 Subject: [PATCH v4 01/11] sched/isolation: Add runtime housekeeping mask updates with boot snapshots Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260710-wujing-dhm-v4-1-2e912e5d9645@gmail.com> References: <20260710-wujing-dhm-v4-0-2e912e5d9645@gmail.com> In-Reply-To: <20260710-wujing-dhm-v4-0-2e912e5d9645@gmail.com> To: Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Josh Triplett , Boqun Feng , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Anna-Maria Behnsen , Tejun Heo , Jonathan Corbet , Shuah Khan , Shuah Khan , Thomas Gleixner Cc: Waiman Long , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, cgroups@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, Jing Wu , Qiliang Yuan X-Mailer: b4 0.13.0 The housekeeping cpumasks for kernel noise (HK_TYPE_KERNEL_NOISE) and managed interrupts (HK_TYPE_MANAGED_IRQ) are computed once at boot from the nohz_full=3D and isolcpus=3Dmanaged_irq arguments. Dynamic CPU isolation driven by cpuset isolated partitions needs to update these masks after boot. Add housekeeping_update_types() to recompute one or more housekeeping masks as (boot snapshot & ~isolated), publish them via RCU and free the old masks after a grace period. Introduce HK_TYPE_KERNEL_NOISE_BOOT and HK_TYPE_MANAGED_IRQ_BOOT to record the immutable boot configuration. Anchoring every update on the boot snapshot keeps the runtime mask a subset of the boot set and lets de-isolation restore the boot configuration exactly. When neither nohz_full=3D nor isolcpus=3Dnohz was given at boot no boot snapshot exists, so cpu_possible_mask is the implicit boot set and HK_TYPE_KERNEL_NOISE is enabled on the first runtime call. In that case housekeeping_update_types() also calls sched_tick_offload_init() to allocate the tick offload percpu data that would otherwise only be allocated by housekeeping_init() when nohz_full=3D is present at boot. Remove __init from sched_tick_offload_init() and guard it with a tick_work_cpu check so it is safe to call at runtime. Co-developed-by: Qiliang Yuan Signed-off-by: Qiliang Yuan Signed-off-by: Jing Wu --- include/linux/sched/isolation.h | 32 ++++++- kernel/sched/core.c | 4 +- kernel/sched/isolation.c | 197 +++++++++++++++++++++++++++++++++++-= ---- kernel/sched/sched.h | 2 +- 4 files changed, 209 insertions(+), 26 deletions(-) diff --git a/include/linux/sched/isolation.h b/include/linux/sched/isolatio= n.h index cf0fd03dd7a24..70602a74c1410 100644 --- a/include/linux/sched/isolation.h +++ b/include/linux/sched/isolation.h @@ -14,10 +14,26 @@ enum hk_type { * is always a subset of HK_TYPE_DOMAIN_BOOT. */ HK_TYPE_DOMAIN, - /* Inverse of boot-time isolcpus=3Dmanaged_irq argument */ - HK_TYPE_MANAGED_IRQ, - /* Inverse of boot-time nohz_full=3D or isolcpus=3Dnohz arguments */ + + /* + * Inverse of the boot-time nohz_full=3D or isolcpus=3Dnohz arguments. + * When neither is given, DHM still records cpu_possible_mask here so + * that kernel-noise isolation can be enabled purely at runtime. + */ + HK_TYPE_KERNEL_NOISE_BOOT, + /* + * A subset of HK_TYPE_KERNEL_NOISE_BOOT: it may exclude additional + * CPUs isolated at runtime via cpuset isolated partitions. + */ HK_TYPE_KERNEL_NOISE, + + /* Inverse of the boot-time isolcpus=3Dmanaged_irq argument */ + HK_TYPE_MANAGED_IRQ_BOOT, + /* + * A subset of HK_TYPE_MANAGED_IRQ_BOOT: it may exclude additional + * CPUs isolated at runtime via cpuset isolated partitions. + */ + HK_TYPE_MANAGED_IRQ, HK_TYPE_MAX, =20 /* @@ -40,10 +56,13 @@ enum hk_type { DECLARE_STATIC_KEY_FALSE(housekeeping_overridden); extern int housekeeping_any_cpu(enum hk_type type); extern const struct cpumask *housekeeping_cpumask(enum hk_type type); +extern const struct cpumask *housekeeping_cpumask_rcu(enum hk_type type); extern bool housekeeping_enabled(enum hk_type type); extern void housekeeping_affine(struct task_struct *t, enum hk_type type); extern bool housekeeping_test_cpu(int cpu, enum hk_type type); extern int housekeeping_update(struct cpumask *isol_mask); +extern int housekeeping_update_types(unsigned long type_mask, + struct cpumask *isol_mask); extern void __init housekeeping_init(void); =20 #else @@ -58,6 +77,11 @@ static inline const struct cpumask *housekeeping_cpumask= (enum hk_type type) return cpu_possible_mask; } =20 +static inline const struct cpumask *housekeeping_cpumask_rcu(enum hk_type = type) +{ + return cpu_possible_mask; +} + static inline bool housekeeping_enabled(enum hk_type type) { return false; @@ -72,6 +96,8 @@ static inline bool housekeeping_test_cpu(int cpu, enum hk= _type type) } =20 static inline int housekeeping_update(struct cpumask *isol_mask) { return = 0; } +static inline int housekeeping_update_types(unsigned long type_mask, + struct cpumask *isol_mask) { return 0; } static inline void housekeeping_init(void) { } #endif /* CONFIG_CPU_ISOLATION */ =20 diff --git a/kernel/sched/core.c b/kernel/sched/core.c index b8871449d3c69..79c3349f65bb4 100644 --- a/kernel/sched/core.c +++ b/kernel/sched/core.c @@ -5813,8 +5813,10 @@ static void sched_tick_stop(int cpu) } #endif /* CONFIG_HOTPLUG_CPU */ =20 -int __init sched_tick_offload_init(void) +int sched_tick_offload_init(void) { + if (tick_work_cpu) + return 0; tick_work_cpu =3D alloc_percpu(struct tick_work); BUG_ON(!tick_work_cpu); return 0; diff --git a/kernel/sched/isolation.c b/kernel/sched/isolation.c index ef152d401fe20..4602c8d0108e4 100644 --- a/kernel/sched/isolation.c +++ b/kernel/sched/isolation.c @@ -12,10 +12,12 @@ #include "sched.h" =20 enum hk_flags { - HK_FLAG_DOMAIN_BOOT =3D BIT(HK_TYPE_DOMAIN_BOOT), - HK_FLAG_DOMAIN =3D BIT(HK_TYPE_DOMAIN), - HK_FLAG_MANAGED_IRQ =3D BIT(HK_TYPE_MANAGED_IRQ), - HK_FLAG_KERNEL_NOISE =3D BIT(HK_TYPE_KERNEL_NOISE), + HK_FLAG_DOMAIN_BOOT =3D BIT(HK_TYPE_DOMAIN_BOOT), + HK_FLAG_DOMAIN =3D BIT(HK_TYPE_DOMAIN), + HK_FLAG_KERNEL_NOISE_BOOT =3D BIT(HK_TYPE_KERNEL_NOISE_BOOT), + HK_FLAG_KERNEL_NOISE =3D BIT(HK_TYPE_KERNEL_NOISE), + HK_FLAG_MANAGED_IRQ_BOOT =3D BIT(HK_TYPE_MANAGED_IRQ_BOOT), + HK_FLAG_MANAGED_IRQ =3D BIT(HK_TYPE_MANAGED_IRQ), }; =20 DEFINE_STATIC_KEY_FALSE(housekeeping_overridden); @@ -34,25 +36,40 @@ bool housekeeping_enabled(enum hk_type type) } EXPORT_SYMBOL_GPL(housekeeping_enabled); =20 +/* + * Types that can change at runtime via cpuset isolated partitions. + * Boot-only types (DOMAIN_BOOT) are always safe to read without lockdep. + */ +static bool housekeeping_type_can_change(enum hk_type type) +{ + switch (type) { + case HK_TYPE_DOMAIN: + case HK_TYPE_KERNEL_NOISE: + case HK_TYPE_MANAGED_IRQ: + return true; + default: + return false; + } +} + static bool housekeeping_dereference_check(enum hk_type type) { - if (IS_ENABLED(CONFIG_LOCKDEP) && type =3D=3D HK_TYPE_DOMAIN) { - /* Cpuset isn't even writable yet? */ - if (system_state <=3D SYSTEM_SCHEDULING) - return true; + if (!IS_ENABLED(CONFIG_LOCKDEP) || !housekeeping_type_can_change(type)) + return true; =20 - /* CPU hotplug write locked, so cpuset partition can't be overwritten */ - if (IS_ENABLED(CONFIG_HOTPLUG_CPU) && lockdep_is_cpus_write_held()) - return true; + /* Cpuset isn't even writable yet? */ + if (system_state <=3D SYSTEM_SCHEDULING) + return true; =20 - /* Cpuset lock held, partitions not writable */ - if (IS_ENABLED(CONFIG_CPUSETS) && lockdep_is_cpuset_held()) - return true; + /* CPU hotplug write locked, so cpuset partition can't be overwritten */ + if (IS_ENABLED(CONFIG_HOTPLUG_CPU) && lockdep_is_cpus_write_held()) + return true; =20 - return false; - } + /* Cpuset lock held, partitions not writable */ + if (IS_ENABLED(CONFIG_CPUSETS) && lockdep_is_cpuset_held()) + return true; =20 - return true; + return false; } =20 static inline struct cpumask *housekeeping_cpumask_dereference(enum hk_typ= e type) @@ -75,12 +92,26 @@ const struct cpumask *housekeeping_cpumask(enum hk_type= type) } EXPORT_SYMBOL_GPL(housekeeping_cpumask); =20 +const struct cpumask *housekeeping_cpumask_rcu(enum hk_type type) +{ + const struct cpumask *mask =3D NULL; + + if (static_branch_unlikely(&housekeeping_overridden)) { + if (READ_ONCE(housekeeping.flags) & BIT(type)) + mask =3D rcu_dereference(housekeeping.cpumasks[type]); + } + if (!mask) + mask =3D cpu_possible_mask; + return mask; +} +EXPORT_SYMBOL_GPL(housekeeping_cpumask_rcu); + int housekeeping_any_cpu(enum hk_type type) { int cpu; =20 if (static_branch_unlikely(&housekeeping_overridden)) { - if (housekeeping.flags & BIT(type)) { + if (READ_ONCE(housekeeping.flags) & BIT(type)) { cpu =3D sched_numa_find_closest(housekeeping_cpumask(type), smp_process= or_id()); if (cpu < nr_cpu_ids) return cpu; @@ -162,6 +193,130 @@ int housekeeping_update(struct cpumask *isol_mask) return 0; } =20 +/** + * housekeeping_update_types - Update housekeeping masks for specified typ= es + * @type_mask: Bitmask of housekeeping types to update + * @isol_mask: CPUs being added to the isolation set + * + * For each type in @type_mask, compute the trial mask as + * (boot snapshot & ~@isol_mask), validate it against @cpu_online_mask, + * then swap the RCU mask pointer and free the old mask after + * synchronize_rcu(). Anchoring on the immutable boot snapshot + * (HK_TYPE_*_BOOT) keeps the runtime mask a subset of the boot set and + * lets de-isolation restore the boot configuration exactly. + * + * The updated mask only takes effect for subsystems as CPUs cycle + * through hotplug; callers isolate CPUs via the CPU hotplug machinery so + * that tick, RCU and interrupt state is reconfigured by the existing + * online/offline callbacks rather than reconfigured in place. + * + * HK_TYPE_KERNEL_NOISE also supports runtime first-enable: when neither + * nohz_full=3D nor isolcpus=3Dnohz was given at boot, no boot snapshot + * exists, so cpu_possible_mask is the implicit boot set and the type + * flag is set in housekeeping.flags on the first call. + * + * Return: 0 on success, -ENOMEM on allocation failure, -EINVAL if + * a trial mask has no online CPUs. + */ +int housekeeping_update_types(unsigned long type_mask, + struct cpumask *isol_mask) +{ + struct cpumask *trials[HK_TYPE_MAX] =3D {}; + struct cpumask *old_masks[HK_TYPE_MAX] =3D {}; + enum hk_type type; + int ret =3D 0; + + for_each_set_bit(type, &type_mask, HK_TYPE_MAX) { + const struct cpumask *base; + + if (type =3D=3D HK_TYPE_DOMAIN_BOOT) + continue; + if (!housekeeping_enabled(type)) { + /* + * HK_TYPE_KERNEL_NOISE supports runtime first-enable + * for DHM isolated partitions created without nohz_full=3D + * at boot. All other types must be boot-enabled. + */ + if (type !=3D HK_TYPE_KERNEL_NOISE) + continue; + } + + /* + * Compute the trial mask relative to the immutable boot + * snapshot, never relative to the current (already shrunk) + * mask. Using the current mask would let it shrink + * monotonically across isolation/de-isolation cycles and would + * never recover CPUs once de-isolated. Anchoring on the boot + * snapshot keeps the runtime mask a subset of the boot set and + * lets de-isolation restore exactly the boot configuration. + * + * HK_TYPE_KERNEL_NOISE additionally supports runtime + * first-enable: when no nohz_full=3D/isolcpus=3Dnohz was given at + * boot, no boot snapshot exists, so cpu_possible_mask is the + * implicit boot set. + */ + if (type =3D=3D HK_TYPE_KERNEL_NOISE && + !(housekeeping.flags & HK_FLAG_KERNEL_NOISE_BOOT)) + base =3D cpu_possible_mask; + else if (type =3D=3D HK_TYPE_KERNEL_NOISE) + base =3D housekeeping_cpumask(HK_TYPE_KERNEL_NOISE_BOOT); + else if (type =3D=3D HK_TYPE_MANAGED_IRQ) + base =3D housekeeping_cpumask(HK_TYPE_MANAGED_IRQ_BOOT); + else + base =3D housekeeping_cpumask(type); + trials[type] =3D kmalloc(cpumask_size(), GFP_KERNEL); + if (!trials[type]) { + ret =3D -ENOMEM; + goto err_free; + } + cpumask_andnot(trials[type], base, isol_mask); + if (!cpumask_intersects(trials[type], cpu_online_mask)) { + ret =3D -EINVAL; + goto err_free; + } + } + + if (!housekeeping.flags) { + ret =3D -EINVAL; + goto err_free; + } + + for_each_set_bit(type, &type_mask, HK_TYPE_MAX) { + if (!trials[type]) + continue; + old_masks[type] =3D housekeeping_cpumask_dereference(type); + /* First-time runtime enable: register the type now. */ + if (!housekeeping_enabled(type)) { + WRITE_ONCE(housekeeping.flags, + housekeeping.flags | BIT(type)); + /* + * HK_TYPE_KERNEL_NOISE first-enable at runtime + * (zero-boot-param path): tick offload percpu data + * was never allocated at boot since nohz_full=3D was + * absent. Allocate it now before CPUs cycle through + * hotplug and sched_tick_stop() dereferences + * tick_work_cpu. + */ + if (type =3D=3D HK_TYPE_KERNEL_NOISE) + WARN_ON_ONCE(sched_tick_offload_init()); + } + rcu_assign_pointer(housekeeping.cpumasks[type], trials[type]); + trials[type] =3D NULL; + } + + synchronize_rcu(); + + for_each_set_bit(type, &type_mask, HK_TYPE_MAX) + kfree(old_masks[type]); + + return 0; + +err_free: + for_each_set_bit(type, &type_mask, HK_TYPE_MAX) + kfree(trials[type]); + return ret; +} + void __init housekeeping_init(void) { enum hk_type type; @@ -305,7 +460,7 @@ static int __init housekeeping_nohz_full_setup(char *st= r) { unsigned long flags; =20 - flags =3D HK_FLAG_KERNEL_NOISE; + flags =3D HK_FLAG_KERNEL_NOISE | HK_FLAG_KERNEL_NOISE_BOOT; =20 return housekeeping_setup(str, flags); } @@ -324,7 +479,7 @@ static int __init housekeeping_isolcpus_setup(char *str) */ if (!strncmp(str, "nohz,", 5)) { str +=3D 5; - flags |=3D HK_FLAG_KERNEL_NOISE; + flags |=3D HK_FLAG_KERNEL_NOISE | HK_FLAG_KERNEL_NOISE_BOOT; continue; } =20 @@ -336,7 +491,7 @@ static int __init housekeeping_isolcpus_setup(char *str) =20 if (!strncmp(str, "managed_irq,", 12)) { str +=3D 12; - flags |=3D HK_FLAG_MANAGED_IRQ; + flags |=3D HK_FLAG_MANAGED_IRQ | HK_FLAG_MANAGED_IRQ_BOOT; continue; } =20 diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h index 9f63b15d309d1..9d54ac3f2251d 100644 --- a/kernel/sched/sched.h +++ b/kernel/sched/sched.h @@ -2925,7 +2925,7 @@ extern void post_init_entity_util_avg(struct task_str= uct *p); =20 #ifdef CONFIG_NO_HZ_FULL extern bool sched_can_stop_tick(struct rq *rq); -extern int __init sched_tick_offload_init(void); +extern int sched_tick_offload_init(void); =20 /* * Tick may be needed by tasks in the runqueue depending on their policy a= nd --=20 2.43.0 From nobody Sun Jul 26 01:47:31 2026 Received: from mail-pl1-f169.google.com (mail-pl1-f169.google.com [209.85.214.169]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D6DFE3672B3 for ; Fri, 10 Jul 2026 03:28:42 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.169 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783654124; cv=none; b=deQkgla+gb7p0Bxiya3gAXcTHsFCK+/E8wWOD7B04xyw0kWne1s8wBJ4pthQ2D2Wv0VpKkoQrhbSB7mWaiDvMJM7rNRnsdlPD2uSXO5DK9+72NqeISOig9Ggo3Ieneji47pu1X+2X0eYwbQnBMTSfP05PIQgLnxJg0/kg/CVsSY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783654124; c=relaxed/simple; bh=tA9wD3TNXeEArI8ZoxLwOtsaod6AseL+D5l+XBb7lvs=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=UyWBEjlEm8kEsyEEEDvN9M+EUQM3rznAYxwdSLp+7vG7RCX4tb5VGYA0tuFUSjeevxx2Ct5Xcux2s3SjdZJAD2KmtIU4+s6vwv2oT3h2/xOSmJCIyfOKqgd7G8Utkrm52hDpsizQUaot4qIh6izfcvSuBVv6gKEqentvrCteYJQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=cnDmH4dh; arc=none smtp.client-ip=209.85.214.169 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="cnDmH4dh" Received: by mail-pl1-f169.google.com with SMTP id d9443c01a7336-2cc891373e0so4002005ad.2 for ; Thu, 09 Jul 2026 20:28:42 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1783654122; x=1784258922; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=WRlCpxoiJUvKc2NSwy5g73yIeI0cTYhVnCLSI2ilvTQ=; b=cnDmH4dhfFzZIKi56f3Z4JsrCm7444rLwqcUq0UDgzOHlxqIKVw3lY96f7ZeP+Z3P/ +A3SVdvam2tgi3wXs1lJh9R3j5MKtfImQBFyrZ2mGE3Vcp+rXByWnHTZSJKq6ol5ijDp 7dNu3QsPyQgy9btje65H/Fg5HUDfRIVw2RqW44iWePsysbFS7hv6n+4DoR7C5+NuP0XM A3BshLDLUxcllOyYU61e25E1tBwEaEDCZ7kcy284Gb39/t08SSi3L5sLn01vs83mHda/ qhYvTsMPINHJhrxpRFwXpSTbHPcyqWHTu4SAWd6JvXUeVbNfovE51xiod3AdLqY9UulA S2vg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783654122; x=1784258922; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=WRlCpxoiJUvKc2NSwy5g73yIeI0cTYhVnCLSI2ilvTQ=; b=rRlLndZNNccdFFosrRnEHohAwQf/nM5dRrfotQYBIrRHbGxjKUEbqUA8ObEL8ObPkl 1RnPa+5S5CuD0zT1IoE/h7/WqLmuNyfZOa2CsxqdI//TfEydFw+5S9t41EzQEgpU7CuQ yM/8kkGOxm3pLcu6Cdt4bKQOVXulMaR4j5E2Ll8/ls34eLI7UZT4+w0WYmjDypY1LKer OiqdA0sFsKv0DoESMt8MF6g+NFdaIOweOupT8K98DrhOb8eoKF/ty9STYny3QcpramN/ yS9PtTmSv1t/zL1O4LeOZF80/JNSJ/OYJaMSi3xuZQ8FQUGLR9g9pUpDzpdtU98DWMUt qbqQ== X-Forwarded-Encrypted: i=1; AHgh+RqGTUVDrizhLBBAantWTPyov0V22lQTteXGRB9xKnxKqrZFXNAtALov9EcNikn3hxrRvMK+tDHZYMQoXIs=@vger.kernel.org X-Gm-Message-State: AOJu0YxUlLRoSqJskhGqNTeWMcwm+XTInwMMPOGzjmUQYfj6AbL3Uacn elRJ98kfNFEIirjZqZ3bciRoiQljWg2XJqU3Q93/nCQAjPNZR+3+FqTv3swvZHsqdb4= X-Gm-Gg: AfdE7clTGaaC1JJ5EFzE+9KG0MokT4Dc5q7dKccna0x2RZDfpAv2kNYwJQnaddY38dx vwfoOGJm22JiV0JedJQL0du3EA8gJqO8NColDOiPa7ELXUPZAUg5o02JilpaYhV8GuL/KvczWfu HltoSWT+PsZ4RaUj5r4+sLi4Ag9Iva95ilKmyunYVjwnb+uOoLwe+TzxwJrdMkgQskWl3Epfz2H EnUIDoHfI0Utk7WfpJlDdJrjQX7Ku09WmG4QtdTmNwycrLkralendU7nETzVfptEYB5lcHDoNF5 y/MRWAHVB3knFp9LsiLR+0WBcyuxrtsZTE7y/Rm9ejkmi/zIHKn/Iwrb0GP6+mvpD/g/l1NuKcP ZOxRrNcTreMeoThbkfGO2FKq5PVOI8s28z2NIB/MAq2mhBrPISJzmtSm8/G5pYu33P5tNbCRnZK 3xZj/BXCbnBbY= X-Received: by 2002:a17:902:db09:b0:2c9:b48c:fdec with SMTP id d9443c01a7336-2ccea3b1266mr100601535ad.12.1783654121956; Thu, 09 Jul 2026 20:28:41 -0700 (PDT) Received: from [127.0.1.1] ([138.199.21.246]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2ccc9bdb56fsm53436465ad.15.2026.07.09.20.28.33 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 09 Jul 2026 20:28:41 -0700 (PDT) From: Jing Wu Date: Fri, 10 Jul 2026 11:28:13 +0800 Subject: [PATCH v4 02/11] sched/isolation: RCU-protect runtime-mutable housekeeping cpumask readers Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260710-wujing-dhm-v4-2-2e912e5d9645@gmail.com> References: <20260710-wujing-dhm-v4-0-2e912e5d9645@gmail.com> In-Reply-To: <20260710-wujing-dhm-v4-0-2e912e5d9645@gmail.com> To: Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Josh Triplett , Boqun Feng , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Anna-Maria Behnsen , Tejun Heo , Jonathan Corbet , Shuah Khan , Shuah Khan , Thomas Gleixner Cc: Waiman Long , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, cgroups@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, Jing Wu , Qiliang Yuan X-Mailer: b4 0.13.0 Now that HK_TYPE_KERNEL_NOISE and HK_TYPE_MANAGED_IRQ can be updated at runtime, their cpumask pointers are swapped and the old masks freed after an RCU grace period. Readers that dereference these masks must do so inside an RCU read-side critical section, otherwise the mask can be freed while it is still in use. Convert the runtime-mutable readers to housekeeping_cpumask_rcu() under rcu_read_lock(): - get_nohz_timer_target() (HK_TYPE_KERNEL_NOISE) - hrtimer target selection (HK_TYPE_TIMER) - arm64 topology (HK_TYPE_TICK) - Hyper-V channel management (HK_TYPE_MANAGED_IRQ) - the housekeeping sysfs attribute (HK_TYPE_KERNEL_NOISE) The watchdog boot-time cpumask initialisation is switched from the HK_TYPE_TIMER alias to HK_TYPE_KERNEL_NOISE for consistency; both alias the same value. Co-developed-by: Qiliang Yuan Signed-off-by: Qiliang Yuan Signed-off-by: Jing Wu --- arch/arm64/kernel/topology.c | 9 ++++++-- drivers/base/cpu.c | 20 +++++++++++++----- drivers/hv/channel_mgmt.c | 50 ++++++++++++++++++++++++++++++----------= ---- kernel/sched/core.c | 3 +-- kernel/time/hrtimer.c | 5 ++++- kernel/watchdog.c | 2 +- 6 files changed, 62 insertions(+), 27 deletions(-) diff --git a/arch/arm64/kernel/topology.c b/arch/arm64/kernel/topology.c index b32f13358fbb1..8f4329b57cea7 100644 --- a/arch/arm64/kernel/topology.c +++ b/arch/arm64/kernel/topology.c @@ -212,8 +212,13 @@ int arch_freq_get_on_cpu(int cpu) if (!policy) return -EINVAL; =20 - if (!cpumask_intersects(policy->related_cpus, - housekeeping_cpumask(HK_TYPE_TICK))) { + bool no_hk_in_policy; + + rcu_read_lock(); + no_hk_in_policy =3D !cpumask_intersects(policy->related_cpus, + housekeeping_cpumask_rcu(HK_TYPE_TICK)); + rcu_read_unlock(); + if (no_hk_in_policy) { cpufreq_cpu_put(policy); return -EOPNOTSUPP; } diff --git a/drivers/base/cpu.c b/drivers/base/cpu.c index 875abdc9942e1..e0aa3d34bc41a 100644 --- a/drivers/base/cpu.c +++ b/drivers/base/cpu.c @@ -303,13 +303,23 @@ static DEVICE_ATTR(isolated, 0444, print_cpus_isolate= d, NULL); static ssize_t housekeeping_show(struct device *dev, struct device_attribute *attr, char *buf) { - const struct cpumask *hk_mask; + ssize_t len; =20 - hk_mask =3D housekeeping_cpumask(HK_TYPE_KERNEL_NOISE); + if (!housekeeping_enabled(HK_TYPE_KERNEL_NOISE)) + return sysfs_emit(buf, "\n"); =20 - if (housekeeping_enabled(HK_TYPE_KERNEL_NOISE)) - return sysfs_emit(buf, "%*pbl\n", cpumask_pr_args(hk_mask)); - return sysfs_emit(buf, "\n"); + /* + * HK_TYPE_KERNEL_NOISE is runtime-mutable: the mask pointer can be + * swapped and the old mask freed after an RCU grace period. Hold the + * RCU read lock across the dereference and the format so the mask + * cannot be freed while it is being printed. + */ + rcu_read_lock(); + len =3D sysfs_emit(buf, "%*pbl\n", + cpumask_pr_args(housekeeping_cpumask_rcu(HK_TYPE_KERNEL_NOISE))); + rcu_read_unlock(); + + return len; } static DEVICE_ATTR_RO(housekeeping); =20 diff --git a/drivers/hv/channel_mgmt.c b/drivers/hv/channel_mgmt.c index 84eb0a6a0b546..fc5247e92e1b3 100644 --- a/drivers/hv/channel_mgmt.c +++ b/drivers/hv/channel_mgmt.c @@ -750,26 +750,43 @@ static void init_vp_index(struct vmbus_channel *chann= el) { bool perf_chn =3D hv_is_perf_channel(channel); u32 i, ncpu =3D num_online_cpus(); - cpumask_var_t available_mask; + cpumask_var_t available_mask, hk_snap; struct cpumask *allocated_mask; - const struct cpumask *hk_mask =3D housekeeping_cpumask(HK_TYPE_MANAGED_IR= Q); u32 target_cpu; int numa_node; =20 - if (!perf_chn || - !alloc_cpumask_var(&available_mask, GFP_KERNEL) || - cpumask_empty(hk_mask)) { - /* - * If the channel is not a performance critical - * channel, bind it to VMBUS_CONNECT_CPU. - * In case alloc_cpumask_var() fails, bind it to - * VMBUS_CONNECT_CPU. - * If all the cpus are isolated, bind it to - * VMBUS_CONNECT_CPU. - */ + if (!perf_chn) { + channel->target_cpu =3D VMBUS_CONNECT_CPU; + return; + } + + if (!alloc_cpumask_var(&available_mask, GFP_KERNEL)) { + channel->target_cpu =3D VMBUS_CONNECT_CPU; + hv_set_allocated_cpu(VMBUS_CONNECT_CPU); + return; + } + + /* + * Snapshot HK_TYPE_MANAGED_IRQ cpumask under RCU read lock. + * housekeeping_update_types() frees the old cpumask after + * synchronize_rcu(), so we must not hold the pointer beyond an + * RCU read-side critical section. + */ + if (!alloc_cpumask_var(&hk_snap, GFP_KERNEL)) { + free_cpumask_var(available_mask); + channel->target_cpu =3D VMBUS_CONNECT_CPU; + hv_set_allocated_cpu(VMBUS_CONNECT_CPU); + return; + } + rcu_read_lock(); + cpumask_copy(hk_snap, housekeeping_cpumask_rcu(HK_TYPE_MANAGED_IRQ)); + rcu_read_unlock(); + + if (cpumask_empty(hk_snap)) { + free_cpumask_var(hk_snap); + free_cpumask_var(available_mask); channel->target_cpu =3D VMBUS_CONNECT_CPU; - if (perf_chn) - hv_set_allocated_cpu(VMBUS_CONNECT_CPU); + hv_set_allocated_cpu(VMBUS_CONNECT_CPU); return; } =20 @@ -788,7 +805,7 @@ static void init_vp_index(struct vmbus_channel *channel) =20 retry: cpumask_xor(available_mask, allocated_mask, cpumask_of_node(numa_node)); - cpumask_and(available_mask, available_mask, hk_mask); + cpumask_and(available_mask, available_mask, hk_snap); =20 if (cpumask_empty(available_mask)) { /* @@ -809,6 +826,7 @@ static void init_vp_index(struct vmbus_channel *channel) =20 channel->target_cpu =3D target_cpu; =20 + free_cpumask_var(hk_snap); free_cpumask_var(available_mask); } =20 diff --git a/kernel/sched/core.c b/kernel/sched/core.c index 79c3349f65bb4..3eb39103f5469 100644 --- a/kernel/sched/core.c +++ b/kernel/sched/core.c @@ -1272,9 +1272,8 @@ int get_nohz_timer_target(void) default_cpu =3D cpu; } =20 - hk_mask =3D housekeeping_cpumask(HK_TYPE_KERNEL_NOISE); - guard(rcu)(); + hk_mask =3D housekeeping_cpumask_rcu(HK_TYPE_KERNEL_NOISE); =20 for_each_domain(cpu, sd) { for_each_cpu_and(i, sched_domain_span(sd), hk_mask) { diff --git a/kernel/time/hrtimer.c b/kernel/time/hrtimer.c index 5bd6efe598f0f..18e17a9dad67b 100644 --- a/kernel/time/hrtimer.c +++ b/kernel/time/hrtimer.c @@ -242,8 +242,11 @@ static bool hrtimer_suitable_target(struct hrtimer *ti= mer, struct hrtimer_clock_ static inline struct hrtimer_cpu_base *get_target_base(struct hrtimer_cpu_= base *base, bool pinned) { if (!hrtimer_base_is_online(base)) { - int cpu =3D cpumask_any_and(cpu_online_mask, housekeeping_cpumask(HK_TYP= E_TIMER)); + int cpu; =20 + rcu_read_lock(); + cpu =3D cpumask_any_and(cpu_online_mask, housekeeping_cpumask_rcu(HK_TYP= E_TIMER)); + rcu_read_unlock(); return &per_cpu(hrtimer_bases, cpu); } =20 diff --git a/kernel/watchdog.c b/kernel/watchdog.c index 87dd5e0f6968d..c18c3e9781d7b 100644 --- a/kernel/watchdog.c +++ b/kernel/watchdog.c @@ -1389,7 +1389,7 @@ void __init lockup_detector_init(void) pr_info("Disabling watchdog on nohz_full cores by default\n"); =20 cpumask_copy(&watchdog_cpumask, - housekeeping_cpumask(HK_TYPE_TIMER)); + housekeeping_cpumask(HK_TYPE_KERNEL_NOISE)); =20 if (!watchdog_hardlockup_probe()) watchdog_hardlockup_available =3D true; --=20 2.43.0 From nobody Sun Jul 26 01:47:31 2026 Received: from mail-pl1-f173.google.com (mail-pl1-f173.google.com [209.85.214.173]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4A5933672A0 for ; Fri, 10 Jul 2026 03:28:50 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.173 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783654132; cv=none; b=kjeIZLSIcNLSx7cRFKigWS0vmALHQFqzXSRCrJ3/waY5vlcA/YC9PZu6zJnICY7sVD4TZ7uQUNgRcSNKrkvCWYUpTFhC/kY4B59ZAni0ZD2nMERkocXVrFUCaDYyERuSqcOUxV0BwYBDxpIflsl275YZ/GLmC4rJ9EJG/HqhpkU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783654132; c=relaxed/simple; bh=DSH1dsD7ufyJUSitPhujJbKPkynswivD0Ay115xuL8k=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=ovWF/IuKyN3kD+2nwrEBqcADGHbkfvHXObBSLsj3du5DdHgAuEokTSIeffYnJF1iSsltB0X6nEURXH8kSwc0xlofkQWOpkNB5hGJMgTPK8s+Cs39cDOAJ34hwmV16M8aahdVbHpJSVw9cBPrw2EcNxpY4Hh0v6uJqduDsxj7/Xk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=TOqZeRep; arc=none smtp.client-ip=209.85.214.173 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="TOqZeRep" Received: by mail-pl1-f173.google.com with SMTP id d9443c01a7336-2caced6038eso5709865ad.0 for ; Thu, 09 Jul 2026 20:28:50 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1783654130; x=1784258930; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=8OHQR36SH5DHkwv4ahYjdM9a+G5jG41FdFMZ2BcZsNM=; b=TOqZeRep6oegsbqogMXnoxA1Rl9xdeE5wTH4F3P1nrrQZoy03wVa7K5pG04W8Je6N6 iYxhrF4ccDJjIVym2aNTT6Pf4ElV2wq6jDKLnoVG5X49+Y5/+e9WeRxIG5KjHM7UM+S7 pNa9ryubNJodOkwUmNDEbOYJl+JOnZznJAPznk9MmvXxk+1E/YUfxSJ1+teWOatFBujw U4DyC85XyO0CgHkS1GO8DQHlxXOkmAHNIpHbqEu4ZnxFBgE2eF3cipGp4QbWKlA/EW7k KIwSgYtL2cP9/itQ/izW2ZTrJ0hZph4+HRXDuXMwc6EmBhj6V4wX6QzHzyGLvoX2gEdf kpqg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783654130; x=1784258930; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=8OHQR36SH5DHkwv4ahYjdM9a+G5jG41FdFMZ2BcZsNM=; b=BIULMBtyOW9S1w+Xv6RpnX/Ya2Ji7DJnHb/1e9d1XKAVuiyKrUC0zsNpA4RMbwQUje eGMUYzmY4lZ4iP1gAeactVD2pP1cJbdKYRhv2UQUgdKkf1syyIVEbNGmTxEw7gIyZHrr ElWxBergBYN3UbPpm6ODL6wMi7gV+2D8qKxk7Jn7BMhOh7D2kkSAMkh36eicEUzVX0ID Bhya8eSqbv9bz9O1jQI2kKorAzbpsz677kNqgIcBPd5cs6veneQKgh1fLZZ8rkLHlNIc 4qRkGecim4lnP5wK8jkgVF9b2RRmeqJ9LNeKnnawCDQyeE7ZfOGODviYCKS7OWDVfuYJ K0sg== X-Forwarded-Encrypted: i=1; AHgh+Rq0SVfH6yUEIUMcIZ9BggW8vJqjlJcPoae2FnBY2FS1bYS/NrPFXqxc2THn6BpOghhZtHnHhItEGogeI14=@vger.kernel.org X-Gm-Message-State: AOJu0Ywg/8sTcRAz+yj4EUGnw9NIdp5tlAMYLlmDLTARNmTgTtBsCwpp 5GixF7gzhsomo3nScl6TD1cu5t8wNSkVVeAyQ2HKpeiBlvhmu1zOSRBw X-Gm-Gg: AfdE7cm2HtyZUyqKrgMNtPsRzlmgk4kriyUUWmiKHfpuo857DoX0ybVYGuKxnLhbGL8 MIxG2Rx273zguoUZ8luz+M17rzEArtuGxIPz4qo3I6J6J9g8R11oB8rHdAvJr1CoxMtfNkOuhw4 fHEOtHTunhyKG5/53hQYTLd+K0L0wrj6sI5hoO5TBywFfRSGZElnKYwW/hrtBKlip4Es+bGMtWx Uo0+bqFRB6fKSbi3DatzirDcAo2UI7u0wDthRQq8Ufo+h6qFB1vOxjRD5wpB7gq46Jr/a1osjza oGW/PROUiB3953lggr8FxeEi8IK85W5erIRL+oBS1CTaWf8fu7OgO0kam9UTR8q3r/uc0fsfBJ6 7Er2yaRnc6TDrNzpGWvvkKvxmfwlh4xpQAOXDBT3gBVfQcBn8kzgtONTb8rBLaDK7XIXgqlWu0k OJ5elAutJi0Zs= X-Received: by 2002:a17:902:d2c7:b0:2c9:994c:9a5 with SMTP id d9443c01a7336-2ce8298a501mr18775855ad.30.1783654129682; Thu, 09 Jul 2026 20:28:49 -0700 (PDT) Received: from [127.0.1.1] ([138.199.21.246]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2ccc9bdb56fsm53436465ad.15.2026.07.09.20.28.42 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 09 Jul 2026 20:28:49 -0700 (PDT) From: Jing Wu Date: Fri, 10 Jul 2026 11:28:14 +0800 Subject: [PATCH v4 03/11] cgroup/cpuset: Drive kernel-noise housekeeping updates from isolated partitions Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260710-wujing-dhm-v4-3-2e912e5d9645@gmail.com> References: <20260710-wujing-dhm-v4-0-2e912e5d9645@gmail.com> In-Reply-To: <20260710-wujing-dhm-v4-0-2e912e5d9645@gmail.com> To: Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Josh Triplett , Boqun Feng , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Anna-Maria Behnsen , Tejun Heo , Jonathan Corbet , Shuah Khan , Shuah Khan , Thomas Gleixner Cc: Waiman Long , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, cgroups@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, Jing Wu , Qiliang Yuan X-Mailer: b4 0.13.0 An isolated cpuset partition already updates the HK_TYPE_DOMAIN housekeeping mask. Extend it to also update the kernel-noise masks (HK_TYPE_KERNEL_NOISE and HK_TYPE_MANAGED_IRQ) so that creating or destroying an isolated partition reconfigures the full set of housekeeping cpumasks. The sched domain mask is updated first because the workqueue flush and timer migration paths depend on it; the kernel-noise masks are updated afterwards via housekeeping_update_types(). housekeeping_update() and housekeeping_update_types() are called after dropping cpus_read_lock and cpuset_mutex, with only cpuset_top_mutex held for mutual exclusion. Co-developed-by: Qiliang Yuan Signed-off-by: Qiliang Yuan Signed-off-by: Jing Wu --- kernel/cgroup/cpuset.c | 23 +++++++++++++++++++++-- 1 file changed, 21 insertions(+), 2 deletions(-) diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c index 5c33ab20cc208..80f43a24d3c8a 100644 --- a/kernel/cgroup/cpuset.c +++ b/kernel/cgroup/cpuset.c @@ -1347,17 +1347,36 @@ static void cpuset_update_sd_hk_unlock(void) rebuild_sched_domains_locked(); =20 if (update_housekeeping) { + static const unsigned long noise_types =3D + BIT(HK_TYPE_KERNEL_NOISE) | BIT(HK_TYPE_MANAGED_IRQ); + update_housekeeping =3D false; cpumask_copy(isolated_hk_cpus, isolated_cpus); =20 + mutex_unlock(&cpuset_mutex); + cpus_read_unlock(); + /* * housekeeping_update() is now called without holding * cpus_read_lock and cpuset_mutex. Only cpuset_top_mutex * is still being held for mutual exclusion. */ - mutex_unlock(&cpuset_mutex); - cpus_read_unlock(); + + /* + * Update the sched domain mask first; it must succeed + * before the kernel-noise types because workqueue flush + * and timer migration depend on the sched domain mask. + */ WARN_ON_ONCE(housekeeping_update(isolated_hk_cpus)); + + /* + * Update the kernel-noise housekeeping masks + * (HK_TYPE_KERNEL_NOISE and HK_TYPE_MANAGED_IRQ). The tick, + * RCU and managed-interrupt state is reconfigured as the + * affected CPUs are cycled through the CPU hotplug machinery. + */ + WARN_ON_ONCE(housekeeping_update_types(noise_types, + isolated_hk_cpus)); mutex_unlock(&cpuset_top_mutex); } else { cpuset_full_unlock(); --=20 2.43.0 From nobody Sun Jul 26 01:47:31 2026 Received: from mail-pl1-f181.google.com (mail-pl1-f181.google.com [209.85.214.181]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 334F63672A0 for ; Fri, 10 Jul 2026 03:29:00 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.181 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783654141; cv=none; b=evXOltQyUBgKh1dh2yuf3NKaESK7TWJ7YZ/Nf93XyHNAZw/S4uvcYBrS00gz5SpThLgr+9+hfETlIj3skkvRfe68zyYE/1ziRUj8acqH865puelFLTFCkf667t045vkFYVGhA2Gmc7loid2GmxYl4R6nsVT0dpMrVQKDyOzXSCM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783654141; c=relaxed/simple; bh=/7g5rPjTXZCVa2/uRZYg9xonGTIXf23dBk7iXIO7Ns0=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=Wym8fHKf8iCNOp8wCNatq+OmBJ+mj4QgX5P366eA8x82GJXoMe2VzJ3wFYjCiob1R8MtJ2zlq2C8S0BJ9RhY4U03GQEZJcviKKW3kehf/X5mGmTatDQwrs1+COxzmCYnWJpqC2Y7mfqfX5ownjZ225U0ql+O7SOhpfwWUyzi/50= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=ctctuR/1; arc=none smtp.client-ip=209.85.214.181 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="ctctuR/1" Received: by mail-pl1-f181.google.com with SMTP id d9443c01a7336-2ce87c7e3bbso2750665ad.1 for ; Thu, 09 Jul 2026 20:29:00 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1783654140; x=1784258940; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=CdWcLivIGaN7u5tA3Cryx3dVQvByPYpqsiikxpuxuiQ=; b=ctctuR/1r3DuO1xZ+F+26ZuVuh9k62emAj6nfw8tg4tNihn5UI3STRkKRb7pJEHgKO gTg8x9kqXW+KRyQ7hLsgTCUsuxyHRHJ1qz2/ibVxZFUgRqxU98LpMP7OVP9mtfG8WCJu Kwf8nz/PZ7JkWKLwoDFxFWzgTG6mq/V3ofvtzsQ9TUT61VYhBXSaDN55rZSWLYdfxsE/ 9hG6mKhIQuie5XEYE7PcILmUD/t7gytXx8qzqrxLF4li2f8zcCSdJxFJX6sBDNdSqAVU z6IEDwN9h8XvV1ZXOPqPtTprZDy/g+52goUjSwKSI4dP24HyBdIxgU/DqWx4/MKdsFi9 tVLQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783654140; x=1784258940; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=CdWcLivIGaN7u5tA3Cryx3dVQvByPYpqsiikxpuxuiQ=; b=QsL5mAIcqimFXUGNHdQM3VMwoR1wku4p3fi8qmcRkgzNMECJa4HDQ4bdVpiLXxUERg 94q68LCFBMd32pjD3UfxUvdI7xM0SJejBhcYOzK4u6u7G2rxyWwxWBnB7YI+EGAkriLq swYXpOClTZ8o4XLXIFYWv00oIaVUAp2DE6LU00pRalaU89VFzgaA8ztZqvNHg+vwhCni LkOVuxajqs1YiluP9YY9FzWXtPPuUPh8e0ncBPWOxh6ROfmB+zjKQQPb2aOYmsJ6ffqu S+ijd1TXksnYw8MuDhhTNwbNFO2CtItB3AESwnA2bAGAmvlyVrMcYaiEP0fxgSPyKOma BR3w== X-Forwarded-Encrypted: i=1; AHgh+RrNr2TytwgDaUHeP9+UaVBCcKmp/+Du15NVaQUH/iE7mFSgQAx3NbjyQejy1i2+B9S+/Zod0+yN6Se5YEs=@vger.kernel.org X-Gm-Message-State: AOJu0YyvVK328rQ9ypX+nZ4BhRIleHtMBHQYUU8Ey+kru2pFYMhVUIvR MZ4oIWaHWP4D+aJvBcup0wjKWVbSYVLwyRRnfyi5gKloN7Fy3Q18sDBU X-Gm-Gg: AfdE7cnYzmGvBS2RKtvWE++qnPScdjd3n8XUbHVFCWSwOkHXKOc5ST33zLhWfTwC/J7 S7VV4gvTmjEHsVk5WN+UjZEvEn7ecnGnrH+GAuqrXrDCbxfflyBemyx0lrZHOVMNOY5J94soJZR PuoYEuczN7lB8JZHB4tTbwDyZlAIuaZx52ccPrvQKh8TxK+U23/p2R0eZw2EDnuqCCPmYrjjz6w pvVU5G89tB8OGDpoF1cJ5WfMBvwKzZne8XCPvxAwVTWCw4/NAAO5gSw0JenJWYsWTfIizQibmIE LTKZQo1A4f3nJZDFgbqb0n47U3qzuTPrCygJYkZNHIWKWeYFgHNj6HF0MHrcaVlUb49bqIUgR16 OFvVPj75T2hEc0I2MsbFUIMog9xAJ7fdTjcQzlhJovIQJ1LvO7e01i12wvBxkS36MQr+ccWzRJA gVpuKwnFSqzIs= X-Received: by 2002:a17:903:22c1:b0:2bf:9760:b94d with SMTP id d9443c01a7336-2ccea37ace4mr104298875ad.15.1783654139660; Thu, 09 Jul 2026 20:28:59 -0700 (PDT) Received: from [127.0.1.1] ([138.199.21.246]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2ccc9bdb56fsm53436465ad.15.2026.07.09.20.28.49 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 09 Jul 2026 20:28:59 -0700 (PDT) From: Jing Wu Date: Fri, 10 Jul 2026 11:28:15 +0800 Subject: [PATCH v4 04/11] context_tracking: allow runtime per-CPU user tracking enable/disable Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260710-wujing-dhm-v4-4-2e912e5d9645@gmail.com> References: <20260710-wujing-dhm-v4-0-2e912e5d9645@gmail.com> In-Reply-To: <20260710-wujing-dhm-v4-0-2e912e5d9645@gmail.com> To: Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Josh Triplett , Boqun Feng , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Anna-Maria Behnsen , Tejun Heo , Jonathan Corbet , Shuah Khan , Shuah Khan , Thomas Gleixner Cc: Waiman Long , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, cgroups@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, Jing Wu , Qiliang Yuan X-Mailer: b4 0.13.0 ct_cpu_track_user() and the context_tracking_key static key are currently restricted to boot-time use: the key is __ro_after_init and the function is __init with __initdata state. This prevents enabling nohz_full context tracking for CPUs isolated at runtime via cpuset partitions. Split ct_cpu_track_user() into three functions: ct_cpu_track_user(cpu) - sets per_cpu(context_tracking.active) and increments context_tracking_key; callable at runtime with the CPU offline. ct_cpu_untrack_user(cpu) - reverses the above; for de-isolation. ct_cpu_track_user_init(cpu) - __init wrapper; calls ct_cpu_track_user() and handles TIF_NOHZ / tasklist setup. Change context_tracking_key from DEFINE_STATIC_KEY_FALSE_RO to DEFINE_STATIC_KEY_FALSE so that static_branch_inc/dec() can be called after the __ro_after_init window closes. Update tick_nohz_init() to call ct_cpu_track_user_init() so boot behaviour is unchanged. This is a prerequisite for DHM (Dynamic Housekeeping Management) runtime CPU noise isolation without boot parameters. Co-developed-by: Qiliang Yuan Signed-off-by: Qiliang Yuan Signed-off-by: Jing Wu --- include/linux/context_tracking.h | 2 ++ kernel/context_tracking.c | 38 ++++++++++++++++++++++++++++++++++--= -- kernel/time/tick-sched.c | 2 +- 3 files changed, 37 insertions(+), 5 deletions(-) diff --git a/include/linux/context_tracking.h b/include/linux/context_track= ing.h index af9fe87a09225..735d353d87560 100644 --- a/include/linux/context_tracking.h +++ b/include/linux/context_tracking.h @@ -12,6 +12,8 @@ =20 #ifdef CONFIG_CONTEXT_TRACKING_USER extern void ct_cpu_track_user(int cpu); +extern void ct_cpu_untrack_user(int cpu); +extern void __init ct_cpu_track_user_init(int cpu); =20 /* Called with interrupts disabled. */ extern void __ct_user_enter(enum ctx_state state); diff --git a/kernel/context_tracking.c b/kernel/context_tracking.c index a743e7ffa6c00..a81d9f8b85eed 100644 --- a/kernel/context_tracking.c +++ b/kernel/context_tracking.c @@ -411,7 +411,7 @@ static __always_inline void ct_kernel_enter(bool user, = int offset) { } #define CREATE_TRACE_POINTS #include =20 -DEFINE_STATIC_KEY_FALSE_RO(context_tracking_key); +DEFINE_STATIC_KEY_FALSE(context_tracking_key); EXPORT_SYMBOL_GPL(context_tracking_key); =20 static noinstr bool context_tracking_recursion_enter(void) @@ -674,14 +674,44 @@ void user_exit_callable(void) } NOKPROBE_SYMBOL(user_exit_callable); =20 -void __init ct_cpu_track_user(int cpu) +/** + * ct_cpu_track_user - enable context tracking for a CPU + * @cpu: target CPU (must be offline when called at runtime) + * + * Marks @cpu as actively tracking user/kernel transitions and increments + * the context_tracking_key refcount. Safe to call at runtime provided + * the CPU is offline so no context-tracking readers are active on it. + */ +void ct_cpu_track_user(int cpu) { - static __initdata bool initialized =3D false; - if (!per_cpu(context_tracking.active, cpu)) { per_cpu(context_tracking.active, cpu) =3D true; static_branch_inc(&context_tracking_key); } +} +EXPORT_SYMBOL_GPL(ct_cpu_track_user); + +/** + * ct_cpu_untrack_user - disable context tracking for a CPU + * @cpu: target CPU (must be offline when called) + * + * Reverses ct_cpu_track_user(). The CPU must be offline so that no + * context-tracking readers are active on it. + */ +void ct_cpu_untrack_user(int cpu) +{ + if (per_cpu(context_tracking.active, cpu)) { + per_cpu(context_tracking.active, cpu) =3D false; + static_branch_dec(&context_tracking_key); + } +} +EXPORT_SYMBOL_GPL(ct_cpu_untrack_user); + +void __init ct_cpu_track_user_init(int cpu) +{ + static __initdata bool initialized =3D false; + + ct_cpu_track_user(cpu); =20 if (initialized) return; diff --git a/kernel/time/tick-sched.c b/kernel/time/tick-sched.c index cbbb87a0c6e7c..ba7adc671c580 100644 --- a/kernel/time/tick-sched.c +++ b/kernel/time/tick-sched.c @@ -677,7 +677,7 @@ void __init tick_nohz_init(void) } =20 for_each_cpu(cpu, tick_nohz_full_mask) - ct_cpu_track_user(cpu); + ct_cpu_track_user_init(cpu); =20 ret =3D cpuhp_setup_state_nocalls(CPUHP_AP_ONLINE_DYN, "kernel/nohz:predown", NULL, --=20 2.43.0 From nobody Sun Jul 26 01:47:31 2026 Received: from mail-pl1-f177.google.com (mail-pl1-f177.google.com [209.85.214.177]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 788693672A5 for ; Fri, 10 Jul 2026 03:29:10 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.177 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783654153; cv=none; b=GPaZ2PlEeIxGALvc9cMwWt8sK2yUVJuFMFMvCeyU0ZiMlsRzzfFXimKs3xdYcMInknLCaGcDHNE3nR7HA7yf2SGypvrVJRD+7yg/Im7Gg70ZyTLPqlbRUASFk5675Qerq2gOxL9pu3Cdec27S53/9rUA6tGjcfqwsQAqOCgBoRw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783654153; c=relaxed/simple; bh=cZmplDdace8ka2XP9vD8p3ABtAIzwqaOpdlJ1+3LPG8=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=PgvK7Ups8Mfqu/dyxVlyKX8Wsx+64/3oQqjet2eA5nc2M6xxhdACRcsPh50M4I1+8RnT7BOwBO15y1uiwC81OCGTWhObTB2g7WaJ5rhSJ2Q+ZI+PaZt+Ox3lDiLCHXINU1K70NiJQJ9S1BD0ymLBFhFQMUuDwouLn2upZhamdXE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=BJWYlpQe; arc=none smtp.client-ip=209.85.214.177 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="BJWYlpQe" Received: by mail-pl1-f177.google.com with SMTP id d9443c01a7336-2ca64c3ce5fso5853815ad.3 for ; Thu, 09 Jul 2026 20:29:10 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1783654150; x=1784258950; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=2v7qKSZXWIDgV3PiMinRsvjWO8t8BKfsRKjtUJUHEIk=; b=BJWYlpQePs5LgU89Ap/ztAaLrVUh9Ys/ozCiulN8O8u5mNxv1MOK3uzC1IYFUGz7wn eePjfSa4axkfO4knzEvIEYSVURceFD79DmS5bKamCHTBZcKVAfLPQootUhF/t4M3cSif I2FfqBksYwJagrKjOC6RdJkAJBu3hJK7HPdaQ3v1q8G1gp8ZN2Sf0nHGE86OkE25nw12 4rwYkll3DLBPL96HWUv3lRes024hE5y8M4LjZP4Upsjp5wKXN00tRQ35qkRseaBbSQdn JEiYudHwzGv9o/BFNWuxuDE6R9bkC938pHlCfYFaH763xN2jocs6lsjMRLX3UE1qE0iJ 88Kg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783654150; x=1784258950; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=2v7qKSZXWIDgV3PiMinRsvjWO8t8BKfsRKjtUJUHEIk=; b=oQzj6BieJpaOAq7HiUb5IMLFbT9msW9QBlYYiGxl9Ai27Msvv2TutciiMawFp8nMEb +yJv+2eC0s69IU60NuCudygGzfXv2TgGfHcruAbf4rD7q1ebvXHROv7darBl2w9uYc0S 8OCsxfB72VoRIDwc093MdG1yR+ZlgDGnfVQbrZ6T7OWWMcumH7UUiqDG83h22LGUb+dB e1neCucHYX3lwsN1E2MRI5ie1+4n1h57lRTyfsKfiBz0Cp1RGy9DHzFjgffzmJXkLf0X id/mfhGZvNHt5oQoXB1qYA5jrM0Qh+fGPJRK6zCbEjY6tXjX0Yb0818sYpM5ht0LmC5g UyNQ== X-Forwarded-Encrypted: i=1; AHgh+Rre05uLO6Luz2ewtL/k52lBAkFs0zhQD5MwaHgWzhCIWga9uN25Qu1TdxPOSR3I2SlMkosfeNL1YB2L3lk=@vger.kernel.org X-Gm-Message-State: AOJu0Yzo5Ud+R0CZEWgtA2yOPzAILifmV+RaqHAygkZPuO2ZxEy1l0Zh N3bhgxCEgJ7zN+0qKVjs4MWB/taaLj1iCyGDOTgFAIz2C/uD2LwDlclO X-Gm-Gg: AfdE7ckLQ8BIrDFa0O5GqpG7lCCcnj0OBu+D19wRsHpDAwrhMK7CnE13vCiPNgUx9wk MtGmZU22vx+JeX/OC7jhU2XT1yxp5CqybeZTmwAzDY2pGs3O4GhBH61f3vDnsaerBgZoLHU8nQH /GxKjeAGcjtG+L8a+UopjZQ41PxohVerCTSTjJcrgXr3LfAzw4FryiM0VkBTuLzT+Rh+rCUprpG aL8HZKDCod2orzxfYqkTU25vldiAjFaB9VQXsVj8Bl1+iELIyT7iVb020E+meeRp5N0EcePoWxy nbVISW+JZNQzHBEz9/3nUePLAQGaroiQA1Woh0Cv71ESBKFnIFDXdb4PXFYqLJecwj6pGMd5JU2 qAKqhGIxql2rzgzJFIZ1H26fB39VXDlM9z6NXrH8mNcEgaTEC0NDekuNxhnsFYZ0kB0v25EThKp dCsnf0+abnjyg= X-Received: by 2002:a17:903:3803:b0:2c9:d298:6c06 with SMTP id d9443c01a7336-2ccea45f203mr97612355ad.25.1783654149819; Thu, 09 Jul 2026 20:29:09 -0700 (PDT) Received: from [127.0.1.1] ([138.199.21.246]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2ccc9bdb56fsm53436465ad.15.2026.07.09.20.28.59 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 09 Jul 2026 20:29:08 -0700 (PDT) From: Jing Wu Date: Fri, 10 Jul 2026 11:28:16 +0800 Subject: [PATCH v4 05/11] rcu/nocb: support lazy init for runtime CPU isolation Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260710-wujing-dhm-v4-5-2e912e5d9645@gmail.com> References: <20260710-wujing-dhm-v4-0-2e912e5d9645@gmail.com> In-Reply-To: <20260710-wujing-dhm-v4-0-2e912e5d9645@gmail.com> To: Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Josh Triplett , Boqun Feng , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Anna-Maria Behnsen , Tejun Heo , Jonathan Corbet , Shuah Khan , Shuah Khan , Thomas Gleixner Cc: Waiman Long , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, cgroups@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, Jing Wu , Qiliang Yuan X-Mailer: b4 0.13.0 When a cpuset isolated partition first requests kernel-noise isolation, it needs to offload RCU callbacks from the affected CPUs. The existing rcu_nocb_cpu_offload() requires rcu_nocbs=3D or nohz_full=3D at boot so that rcu_organize_nocb_kthreads() has already set rdp->nocb_gp_rdp for each CPU. Without a boot parameter, nocb_gp_rdp is NULL and offload fails immediately. Introduce lazy nocb initialization so that the first call into the isolation path triggers the one-time setup automatically: rcu_nocb_lazy_init() - allocates rcu_nocb_mask and calls rcu_organize_nocb_kthreads() (now without __init) to set nocb_gp_rdp for every possible CPU. Uses a dedicated mutex for serialization with a fast-path read of nocb_is_setup. rcu_nocb_cpu_isolate() - exported entry point called per-CPU while the CPU is offline. Calls rcu_nocb_lazy_init() for the one-time setup, spawns the GP and CB kthreads via the existing rcu_spawn_cpu_nocb_kthread(), then finalizes offload through rcu_nocb_cpu_offload(). Adding the CPU to the GP kthread's nocb_head_rdp list is handled by nocb_gp_toggle_rdp() in the GP kthread, so rcu_organize_nocb_kthreads() can run with an empty mask. Remove __init from rcu_organize_nocb_kthreads() to allow this runtime call path; the function itself has no __initdata dependencies. Co-developed-by: Qiliang Yuan Signed-off-by: Qiliang Yuan Signed-off-by: Jing Wu --- include/linux/rcupdate.h | 2 ++ kernel/rcu/tree.h | 2 +- kernel/rcu/tree_nocb.h | 43 ++++++++++++++++++++++++++++++++++++++++++- 3 files changed, 45 insertions(+), 2 deletions(-) diff --git a/include/linux/rcupdate.h b/include/linux/rcupdate.h index bfa765132de85..5937485d59118 100644 --- a/include/linux/rcupdate.h +++ b/include/linux/rcupdate.h @@ -149,6 +149,7 @@ static __always_inline void rcu_irq_work_resched(void) = { } void rcu_init_nohz(void); int rcu_nocb_cpu_offload(int cpu); int rcu_nocb_cpu_deoffload(int cpu); +int rcu_nocb_cpu_isolate(int cpu); void rcu_nocb_flush_deferred_wakeup(void); =20 #define RCU_NOCB_LOCKDEP_WARN(c, s) RCU_LOCKDEP_WARN(c, s) @@ -158,6 +159,7 @@ void rcu_nocb_flush_deferred_wakeup(void); static inline void rcu_init_nohz(void) { } static inline int rcu_nocb_cpu_offload(int cpu) { return -EINVAL; } static inline int rcu_nocb_cpu_deoffload(int cpu) { return 0; } +static inline int rcu_nocb_cpu_isolate(int cpu) { return -EINVAL; } static inline void rcu_nocb_flush_deferred_wakeup(void) { } =20 #define RCU_NOCB_LOCKDEP_WARN(c, s) diff --git a/kernel/rcu/tree.h b/kernel/rcu/tree.h index 7dfc57e9adb18..f3d31918ea322 100644 --- a/kernel/rcu/tree.h +++ b/kernel/rcu/tree.h @@ -517,7 +517,7 @@ static void rcu_nocb_unlock_irqrestore(struct rcu_data = *rdp, unsigned long flags); static void rcu_lockdep_assert_cblist_protected(struct rcu_data *rdp); #ifdef CONFIG_RCU_NOCB_CPU -static void __init rcu_organize_nocb_kthreads(void); +static void rcu_organize_nocb_kthreads(void); =20 /* * Disable IRQs before checking offloaded state so that local diff --git a/kernel/rcu/tree_nocb.h b/kernel/rcu/tree_nocb.h index 1047b30cd46b7..694fd615f1809 100644 --- a/kernel/rcu/tree_nocb.h +++ b/kernel/rcu/tree_nocb.h @@ -1347,6 +1347,47 @@ void __init rcu_init_nohz(void) rcu_organize_nocb_kthreads(); } =20 +static DEFINE_MUTEX(rcu_nocb_lazy_mutex); + +/* + * Lazily initialize nocb infrastructure on the first call. Allocates + * rcu_nocb_mask and sets nocb_gp_rdp for every possible CPU so that + * rcu_nocb_cpu_isolate() can offload callbacks without rcu_nocbs=3D at bo= ot. + */ +static noinline int rcu_nocb_lazy_init(void) +{ + if (rcu_state.nocb_is_setup) + return 0; + + mutex_lock(&rcu_nocb_lazy_mutex); + if (!rcu_state.nocb_is_setup) { + if (!zalloc_cpumask_var(&rcu_nocb_mask, GFP_KERNEL)) { + mutex_unlock(&rcu_nocb_lazy_mutex); + return -ENOMEM; + } + rcu_organize_nocb_kthreads(); + rcu_state.nocb_is_setup =3D true; + } + mutex_unlock(&rcu_nocb_lazy_mutex); + return 0; +} + +/* + * Offload RCU callbacks for a CPU entering a kernel-noise isolated partit= ion. + * @cpu must be offline. Lazily initializes nocb infrastructure on first u= se. + */ +int rcu_nocb_cpu_isolate(int cpu) +{ + int ret; + + ret =3D rcu_nocb_lazy_init(); + if (ret) + return ret; + rcu_spawn_cpu_nocb_kthread(cpu); + return rcu_nocb_cpu_offload(cpu); +} +EXPORT_SYMBOL_GPL(rcu_nocb_cpu_isolate); + /* Initialize per-rcu_data variables for no-CBs CPUs. */ static void __init rcu_boot_init_nocb_percpu_data(struct rcu_data *rdp) { @@ -1439,7 +1480,7 @@ module_param(rcu_nocb_gp_stride, int, 0444); /* * Initialize GP-CB relationships for all no-CBs CPU. */ -static void __init rcu_organize_nocb_kthreads(void) +static void rcu_organize_nocb_kthreads(void) { int cpu; bool firsttime =3D true; --=20 2.43.0 From nobody Sun Jul 26 01:47:31 2026 Received: from mail-pl1-f174.google.com (mail-pl1-f174.google.com [209.85.214.174]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CD7A035F5E4 for ; Fri, 10 Jul 2026 03:29:18 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.174 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783654160; cv=none; b=fsegBTITJrmIQ2TSJsCgbeSUjb4bMFf3XGfpt7ofYMCvx2Bzks+Y5d6+XMKLnO2imzl5o7HJePi2+IitQRFsvkYb+y1leOWOZxf63eYluu689aGeH6FBBp2aGaVeoAYj/GB17r/IBnO/56/Cs83vFWJyh4PJ9YyAxzu08px/CKA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783654160; c=relaxed/simple; bh=XLXzIgc0ems38hTXhytSCLWNSTJ6/lnczDaOsWqzk80=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=kEbRTJA2ONfoIJ3gXDTwUgyhkmiSRjavBS1A2ftbXKeolQQtxVXjyttyTDpO8LZukmCBIrMTzyBjSErWLXgApPsSf7JrsOkpDx2OLHErFIMLiP8JOki7TxGopv1TXXiFYr8zcKVIRO00DPMGKmleDco33MLCj5G/5dM2QGPmkwg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=FAwlXME/; arc=none smtp.client-ip=209.85.214.174 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="FAwlXME/" Received: by mail-pl1-f174.google.com with SMTP id d9443c01a7336-2ca64c3ce5fso5855375ad.3 for ; Thu, 09 Jul 2026 20:29:18 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1783654158; x=1784258958; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=foKf2OzW/IJk9qSRNXtj3QHT8NO8zJPjssJk6y4CFnQ=; b=FAwlXME/ZU8CIggt766MxOWFrtD1DEAq+1QJV89HtJ/4qNAdyqlrf/mhagoGSGEkRn simFefKR9u1kk0tTchJvpGxeImcfcMQ+B3g9eTnnkDEH9XXJLiKeNBiBoFJhpiNoTROV JPfNXiOvjRdsXCNgbxO6KYKDP20enDacVWXEzfyJzG3eTFqBg1o+IgDB+qgBK4SpH1bV yjUrAOOQSD0c67wB5qVWQYidOMm3XPFPFvlrZxewjKtLcSQg5kZIti0sNZqwA4jlgV1T 5fy1b02DQ9S97U8n8/oJN3AhnVlmiLR1Dx9eL1uSHP/O9VPhJb8mU8FiWzuxGLj7+ByE oXkg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783654158; x=1784258958; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=foKf2OzW/IJk9qSRNXtj3QHT8NO8zJPjssJk6y4CFnQ=; b=i+u5h6yDiPTbf9sZhLNzprYqCCWLbCGE7+ZDw+tbWi6rJ9zY6NAsTgwfMGnViQNbFf vz5H8CJ5lRjtWUZFg5Ncc5TVrlV5ghkyNwV7D6fp/zsoxSteMVPOCTWBbk8WuV1IIU2c 2F0/KwpFgKLZCrIjGo8Ddi8gSjEYmD1R3L7rfDbEIobFPcrpd+RZqmnur/vyLolbTu63 6QcXzO2DQlfpSafhsWXUw5P0ffnwpvPVqVbE5KG7e0px3FoKv0LU4BX+nzljZy5jtaWA HWz23nh9jZkzI4Vr/kvrDGtv4I9/S3a5motjMZb5bDJG1Zupikg/lP/T8UbcLkQqa4qb 5VWQ== X-Forwarded-Encrypted: i=1; AHgh+Rr7QEpLZ6kYjpaRayjoWwNLbA/SffzoUI+TYQB4/krZubZHlfKTpSyotQ14SXXkqlV1H+BK7wk8xqnhReQ=@vger.kernel.org X-Gm-Message-State: AOJu0YylymPrv+qaMMwG8lVfSe5LbSnxb6clvWTK594SY0hH73+w3gd4 e+XX6SLap+iwivEI9EW1YuUYk6acAHoXBLzLnNRwYlF/WFcJf0P4EsSf X-Gm-Gg: AfdE7clmpvZ2+QgjsA0OntVcKakix9epeW9e86oYdM8hNWcJgzCDnEhv/ee2rXpJo2B oS5BkhBOssi+O7um71meOGi8fo8LUfXA8TL/W8pvYNn2VMyW1XnqtImT9sctN05hkSz+hIU4Qlo glPmBIB4gnHy2x1ioV8tqGE0Uh+AiT5JqDLq2A0fjYtkSUSyms1tSEiStcLGXslobPT9hGxqJtT eUEfTWVUqf6ve3MqY4EeggYoonHnwv3H/4Z82d5pu1alx6NSOl8uRpHQwyGio6sQlnGWj9PpffL Da41WFvbL7z6Y5SohdJ0FPslGdjEg6+/vbxrrUdi6bPbtWeGLzD+6VkoptFeK/MuP19vm2APCro xanIaw1CNqspKcX8uKmHhUL9oRDkoO1k8TmRnL20A6/LwjVheRN7inuLMRynjOHuJoAgMLB3N71 qptdfGTIrLhsA= X-Received: by 2002:a17:902:f78b:b0:2ce:6d4a:5b8a with SMTP id d9443c01a7336-2ce6d4a5c71mr56853935ad.38.1783654158257; Thu, 09 Jul 2026 20:29:18 -0700 (PDT) Received: from [127.0.1.1] ([138.199.21.246]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2ccc9bdb56fsm53436465ad.15.2026.07.09.20.29.10 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 09 Jul 2026 20:29:17 -0700 (PDT) From: Jing Wu Date: Fri, 10 Jul 2026 11:28:17 +0800 Subject: [PATCH v4 06/11] watchdog: sync watchdog_cpumask with HK_TYPE_KERNEL_NOISE on isolation Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260710-wujing-dhm-v4-6-2e912e5d9645@gmail.com> References: <20260710-wujing-dhm-v4-0-2e912e5d9645@gmail.com> In-Reply-To: <20260710-wujing-dhm-v4-0-2e912e5d9645@gmail.com> To: Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Josh Triplett , Boqun Feng , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Anna-Maria Behnsen , Tejun Heo , Jonathan Corbet , Shuah Khan , Shuah Khan , Thomas Gleixner Cc: Waiman Long , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, cgroups@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, Jing Wu , Qiliang Yuan X-Mailer: b4 0.13.0 The watchdog is initialized at boot to run on all housekeeping CPUs (HK_TYPE_KERNEL_NOISE). When a cpuset isolated partition removes CPUs from that mask at runtime, watchdog continues running on those CPUs because nothing updates watchdog_cpumask. Save the boot-time watchdog_cpumask as watchdog_cpumask_boot, which captures the user's intended coverage (possibly narrowed via kernel parameter or sysctl) before any runtime isolation. Introduce lockup_detector_hk_update() which intersects this boot snapshot with the current HK_TYPE_KERNEL_NOISE mask and reconfigures the detector. This ensures that isolated CPUs are excluded while honoring any manual narrowing the admin applied at or after boot. lockup_detector_hk_update() snapshots the RCU-protected housekeeping mask under rcu_read_lock(), then updates watchdog_cpumask and calls __lockup_detector_reconfigure() under watchdog_mutex, matching the same locking discipline used by proc_watchdog_cpumask(). Co-developed-by: Qiliang Yuan Signed-off-by: Qiliang Yuan Signed-off-by: Jing Wu --- include/linux/nmi.h | 2 ++ kernel/watchdog.c | 24 ++++++++++++++++++++++++ 2 files changed, 26 insertions(+) diff --git a/include/linux/nmi.h b/include/linux/nmi.h index bc1162895f355..0bbe562de67b7 100644 --- a/include/linux/nmi.h +++ b/include/linux/nmi.h @@ -37,6 +37,7 @@ extern int sysctl_hardlockup_all_cpu_backtrace; static inline void lockup_detector_init(void) { } static inline void lockup_detector_retry_init(void) { } static inline void lockup_detector_soft_poweroff(void) { } +static inline void lockup_detector_hk_update(void) { } #endif /* !CONFIG_LOCKUP_DETECTOR */ =20 #ifdef CONFIG_SOFTLOCKUP_DETECTOR @@ -120,6 +121,7 @@ void watchdog_hardlockup_enable(unsigned int cpu); void watchdog_hardlockup_disable(unsigned int cpu); =20 void lockup_detector_reconfigure(void); +void lockup_detector_hk_update(void); =20 #ifdef CONFIG_HARDLOCKUP_DETECTOR_BUDDY void watchdog_buddy_check_hardlockup(int hrtimer_interrupts); diff --git a/kernel/watchdog.c b/kernel/watchdog.c index c18c3e9781d7b..26463f6d3a39d 100644 --- a/kernel/watchdog.c +++ b/kernel/watchdog.c @@ -53,6 +53,8 @@ static int __read_mostly watchdog_hardlockup_available; =20 struct cpumask watchdog_cpumask __read_mostly; unsigned long *watchdog_cpumask_bits =3D cpumask_bits(&watchdog_cpumask); +/* Boot snapshot: user's intended watchdog mask before any runtime isolati= on. */ +static struct cpumask watchdog_cpumask_boot __ro_after_init; =20 #ifdef CONFIG_HARDLOCKUP_DETECTOR =20 @@ -1348,6 +1350,27 @@ static void __init lockup_detector_delay_init(struct= work_struct *work) lockup_detector_setup(); } =20 +void lockup_detector_hk_update(void) +{ + cpumask_var_t new_mask; + + if (!alloc_cpumask_var(&new_mask, GFP_KERNEL)) + return; + + rcu_read_lock(); + cpumask_and(new_mask, &watchdog_cpumask_boot, + housekeeping_cpumask_rcu(HK_TYPE_KERNEL_NOISE)); + rcu_read_unlock(); + + mutex_lock(&watchdog_mutex); + cpumask_copy(&watchdog_cpumask, new_mask); + __lockup_detector_reconfigure(false); + mutex_unlock(&watchdog_mutex); + + free_cpumask_var(new_mask); +} +EXPORT_SYMBOL_GPL(lockup_detector_hk_update); + /* * lockup_detector_retry_init - retry init lockup detector if possible. * @@ -1390,6 +1413,7 @@ void __init lockup_detector_init(void) =20 cpumask_copy(&watchdog_cpumask, housekeeping_cpumask(HK_TYPE_KERNEL_NOISE)); + cpumask_copy(&watchdog_cpumask_boot, &watchdog_cpumask); =20 if (!watchdog_hardlockup_probe()) watchdog_hardlockup_available =3D true; --=20 2.43.0 From nobody Sun Jul 26 01:47:31 2026 Received: from mail-pl1-f171.google.com (mail-pl1-f171.google.com [209.85.214.171]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 67A201F3D56 for ; Fri, 10 Jul 2026 03:29:28 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.171 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783654169; cv=none; b=GjF4qObviIOi5GY2WsrBZ4d0QDAirgB8sVU6VVMSGhy21NbSebAls6q4b2ldnkZd6CqjvGCBGYNAk4mud1YFQ3Kd1sQ9v/hEyiP+gNshDUzHClsmPZ1AxKvr7vMbUdyKiPoHJnH/Zdy0GJJ9lmTYA+yjx2dJX1Q2HNi3rU8HXEc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783654169; c=relaxed/simple; bh=Gwt0vIVRdUfUurXNz0U9um8/c6hFVw+sQjNLaLm168U=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=qQ7xR6ip1SJMJa0B+G1pZiEQQyENoPFv/4yws5tGyzRdmv6Pa3fqsix3gVAr3lACg3p/uL3oZI1qTEdXdLlR/n6wa0Q7OjUES5aFH7K4KtoEQvdFYtdpym2HapTEpXI6A6+tICZ218lH5RRzkgTas9OOKugBmlPbJhkHpBt2B68= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=jfRS0gus; arc=none smtp.client-ip=209.85.214.171 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="jfRS0gus" Received: by mail-pl1-f171.google.com with SMTP id d9443c01a7336-2c9bd2f8bf7so21464435ad.1 for ; Thu, 09 Jul 2026 20:29:28 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1783654168; x=1784258968; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=6HyVQ93m8I4DnhXDv5rSatBwFHcjRnR8kBcgkmMTv+c=; b=jfRS0gusrHnEQMRd1dhgGrBgdfHybLeQ9ECOlV9GUjmEr3cT0oYfIiX3gea9Wmf2Uy U7pHaSzygXfxwMA6JX/+Mp2DF3Yat95RQu4rTarDlE0EK1twr6zceJfVlxYP0uTJ3BkS SiHdPXVaYYEiNN4bMvZVAie3epJWEdtHrUeQyOQBhS1VAn2dYrKlniXvb92s62idar3y NaUCEvTYX4Exl0JQ0NMe8FyrHge9eGJ97kfDoHmbQf4Vxe7wzgs0O69fEqwSVDXtzwYx c0/2doFPj7ukUyNyPBYqzLiamn4Qi37xwr1FRlVN+SE6QEfeei6IP2wcvIqmVYFSIti8 hXZw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783654168; x=1784258968; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=6HyVQ93m8I4DnhXDv5rSatBwFHcjRnR8kBcgkmMTv+c=; b=PtZgEV+kKwmUanNlkru8SzHGHme1bHFmqbk4p7vVe9H73/AtMnW1uebEaBL7DUZMYF r7Gko9JnZGoE6oEFSSBbsXIw05v/K+p+O4R8VfCQYWIIEBtCTp1B2kN+A/gBolUu9gEY YRSSZCFUQhXa/Ud1Z9zpuKIvk3eszmnQa5HGf3oBacNaJ0nz1HHe6iKRRosGpWGgdFvu +70LnGEX2RO8g9XuGxXuKcFbzz8gazBRB+lkyDFDjbrbEJyLvf3f5WjBjOrWgqNqC5D4 wsuYEQ+Z2qSmyZRGx8/nCssGCIGXRMPsA5ofr0LSTYLzvd8b3lw3AMNLuKydeb6uC522 1uBQ== X-Forwarded-Encrypted: i=1; AHgh+RpavvL8B2rXXuuZgJfc+u8ztDnCgeI/PHTGzrMZ4GYgX4DSzIxNJao1L7TnBcmO93wU8s1aL2ItvjHi4aE=@vger.kernel.org X-Gm-Message-State: AOJu0Yy0xGfzZsWMJHqPZ8/KUCn0kDhOjCxREn5Qi28UJcfIsneWkHyu xCrSg4CPajF8MMZSkulu+qM2gCpznP+V5p0hUrq2IMghZnP2t5WrR3ZN X-Gm-Gg: AfdE7clhxXaUngROAQVGqGTdOPe4LrDTgQ9ns8NXHF98+ludkM0GEgFo5E3Yp1mYii/ 608EJF8zT7vSjuO/bIEFANkOajo89W3W1L3jl8jBVoqZQJvPY+kjpBSvxmbG+zcVItlUV6x8N5r BcCY8tuvd4sRoBZKIRLKDsz7PztB35MDvLdW1vQAwkpa3+IExuvsGILUCjc0RFz8igSNXgNE3DN tdT1w5COEH4EA7XrYTT9plYj8K10HVVrZBQuIQYhymi0ZkttbjUBPs6AexQ2PwwRB6u5z81Yzzn mmr0vV3Aa4wBrzvSWnrNWbiaQdbiy3TA4moJMfN8ilQCJuVm1Zm6ulQbt5kPfta3KBBEF7sZple r57FBR78+nO+ro8hor0xP6suWKGuzkU615N0QzUnQZeKZ9+9ng8wvneGpNnNjD8IOUqNa7M3EV2 gR8qSbwKv1bo0= X-Received: by 2002:a17:902:d98b:b0:2c6:8d95:fd7e with SMTP id d9443c01a7336-2ce8285abefmr20003265ad.6.1783654167651; Thu, 09 Jul 2026 20:29:27 -0700 (PDT) Received: from [127.0.1.1] ([138.199.21.246]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2ccc9bdb56fsm53436465ad.15.2026.07.09.20.29.18 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 09 Jul 2026 20:29:27 -0700 (PDT) From: Jing Wu Date: Fri, 10 Jul 2026 11:28:18 +0800 Subject: [PATCH v4 07/11] tick/nohz: add runtime tick_nohz_full_mask update for CPU isolation Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260710-wujing-dhm-v4-7-2e912e5d9645@gmail.com> References: <20260710-wujing-dhm-v4-0-2e912e5d9645@gmail.com> In-Reply-To: <20260710-wujing-dhm-v4-0-2e912e5d9645@gmail.com> To: Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Josh Triplett , Boqun Feng , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Anna-Maria Behnsen , Tejun Heo , Jonathan Corbet , Shuah Khan , Shuah Khan , Thomas Gleixner Cc: Waiman Long , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, cgroups@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, Jing Wu , Qiliang Yuan X-Mailer: b4 0.13.0 When kernel-noise isolation is requested for a CPU at runtime via a cpuset isolated partition, tick suppression must be activated on the affected CPU without requiring nohz_full=3D at boot. tick_nohz_full_mask and tick_nohz_full_running are currently set only during boot by tick_nohz_full_setup(). There is no runtime path to add or remove CPUs from the full-dynticks set. Add tick_nohz_cpu_isolate() and tick_nohz_cpu_deisolate(), both called with the target CPU offline (between remove_cpu and add_cpu in the hotplug cycling path): tick_nohz_cpu_isolate(cpu) - lazily allocates tick_nohz_full_mask if it was never set up at boot, sets the CPU's bit in the mask, sets tick_nohz_full_running, and calls ct_cpu_track_user() to activate per-CPU context tracking so kernel/user transitions suppress the scheduler tick. tick_nohz_cpu_deisolate(cpu) - reverses the above: deactivates context tracking via ct_cpu_untrack_user(), clears the CPU's bit, and clears tick_nohz_full_running when the mask becomes empty. A per-function mutex guards the lazy allocation and running flag updates against concurrent isolation requests. Co-developed-by: Qiliang Yuan Signed-off-by: Qiliang Yuan Signed-off-by: Jing Wu --- include/linux/tick.h | 4 ++++ kernel/time/tick-sched.c | 45 +++++++++++++++++++++++++++++++++++++++++++++ 2 files changed, 49 insertions(+) diff --git a/include/linux/tick.h b/include/linux/tick.h index 738007d6f577f..c4038e62533ba 100644 --- a/include/linux/tick.h +++ b/include/linux/tick.h @@ -210,6 +210,8 @@ extern void tick_nohz_dep_set_signal(struct task_struct= *tsk, extern void tick_nohz_dep_clear_signal(struct signal_struct *signal, enum tick_dep_bits bit); extern bool tick_nohz_cpu_hotpluggable(unsigned int cpu); +extern int tick_nohz_cpu_isolate(int cpu); +extern void tick_nohz_cpu_deisolate(int cpu); =20 /* * The below are tick_nohz_[set,clear]_dep() wrappers that optimize off-ca= ses @@ -281,6 +283,8 @@ static inline bool tick_nohz_full_cpu(int cpu) { return= false; } static inline void tick_nohz_dep_set_cpu(int cpu, enum tick_dep_bits bit) = { } static inline void tick_nohz_dep_clear_cpu(int cpu, enum tick_dep_bits bit= ) { } static inline bool tick_nohz_cpu_hotpluggable(unsigned int cpu) { return t= rue; } +static inline int tick_nohz_cpu_isolate(int cpu) { return -EINVAL; } +static inline void tick_nohz_cpu_deisolate(int cpu) { } =20 static inline void tick_dep_set(enum tick_dep_bits bit) { } static inline void tick_dep_clear(enum tick_dep_bits bit) { } diff --git a/kernel/time/tick-sched.c b/kernel/time/tick-sched.c index ba7adc671c580..1bc01e2ab525b 100644 --- a/kernel/time/tick-sched.c +++ b/kernel/time/tick-sched.c @@ -621,6 +621,51 @@ void __tick_nohz_task_switch(void) } } =20 +static DEFINE_MUTEX(tick_nohz_cpu_isolate_mutex); + +/* + * tick_nohz_cpu_isolate - Add a CPU to the full-dynticks set at runtime. + * @cpu: the CPU to isolate; must be offline. + * + * Lazily allocates tick_nohz_full_mask on the first call so that no + * nohz_full=3D boot parameter is required. Activates per-CPU context + * tracking so that kernel/user transitions suppress the scheduler tick. + */ +int tick_nohz_cpu_isolate(int cpu) +{ + mutex_lock(&tick_nohz_cpu_isolate_mutex); + if (!cpumask_available(tick_nohz_full_mask)) { + if (!zalloc_cpumask_var(&tick_nohz_full_mask, GFP_KERNEL)) { + mutex_unlock(&tick_nohz_cpu_isolate_mutex); + return -ENOMEM; + } + } + cpumask_set_cpu(cpu, tick_nohz_full_mask); + if (!tick_nohz_full_running) + tick_nohz_full_running =3D true; + mutex_unlock(&tick_nohz_cpu_isolate_mutex); + + ct_cpu_track_user(cpu); + return 0; +} +EXPORT_SYMBOL_GPL(tick_nohz_cpu_isolate); + +/* + * tick_nohz_cpu_deisolate - Remove a CPU from the full-dynticks set. + * @cpu: the CPU to de-isolate; must be offline. + */ +void tick_nohz_cpu_deisolate(int cpu) +{ + ct_cpu_untrack_user(cpu); + + mutex_lock(&tick_nohz_cpu_isolate_mutex); + cpumask_clear_cpu(cpu, tick_nohz_full_mask); + if (cpumask_empty(tick_nohz_full_mask)) + tick_nohz_full_running =3D false; + mutex_unlock(&tick_nohz_cpu_isolate_mutex); +} +EXPORT_SYMBOL_GPL(tick_nohz_cpu_deisolate); + /* Get the boot-time nohz CPU list from the kernel parameters. */ void __init tick_nohz_full_setup(cpumask_var_t cpumask) { --=20 2.43.0 From nobody Sun Jul 26 01:47:31 2026 Received: from mail-pl1-f175.google.com (mail-pl1-f175.google.com [209.85.214.175]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9764E367B7B for ; Fri, 10 Jul 2026 03:29:39 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.175 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783654181; cv=none; b=EztltVScVYOHjoaJDJakixj/yk/fjt8Z67Ox+FOSRr14bST0O1zJ+xItAO/Gza9RLGJMTYUEWgUHwg0PDxDpN2DAHsH5TatyBgmi1vcUWvkALySF75yvuo6xzu105i0d4UClwzOmpHkzbavhUeGCj3NCDbh0xNVqCpETi52nO/E= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783654181; c=relaxed/simple; bh=0kVUx60+5alzmI6kgDLzOujtVEhari+4UwkUd1VMwS0=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=XfaeOnNwUi3AowO940CHZnX7YrI1P+0Jku5cU7i+KhOtNjpP0fClT5SqyO5t0Hh8OP4PYAl91KV+kGKTXeUsH7Y9MJwwUa+6JaqwuhhyzcmKM7qARF3WwyolZ/tDzO6VNsKb7FHNXs5GNzBIrt2vgI8xkajfwbolLIML4J2GuQE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=Bt8EBiUI; arc=none smtp.client-ip=209.85.214.175 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="Bt8EBiUI" Received: by mail-pl1-f175.google.com with SMTP id d9443c01a7336-2ca64c3ce5fso5858515ad.3 for ; Thu, 09 Jul 2026 20:29:39 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1783654179; x=1784258979; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=mrKXut03HXmQzwP+VjNkq+G+JdZqPXgKjkWKKkzd7+Y=; b=Bt8EBiUI2V1wO/K1l4Lhczu1qbvAKxWWLcDVIaRqro2iD8040Zp0B29lShkQSOZqIF RaOhkFfd1jw/2JevWdAHN8MDl5/Rz0LVgkys/NG3qmkDACsvZMr340yo6cR/p02sVzfa Wk9sXELAFSfRZ5D4FDm71dyKQu+mri5cQXtazs71jcl8Mm9Yjp/z6zpCtpJ9GVDo1TVn RpxfAuT8075tF4tdEsYCPsgxzgZGLVEixqLXb1RCniZ7rBI2z7gnaoRFyKFDhjnNO0j5 L8Y1cQ1BlE36/M4ZUDZCg3C5jESYopsnCGY1gF8nyvHYKsqEr6NamSU+iOItO9Pg8Wa+ 4WJw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783654179; x=1784258979; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=mrKXut03HXmQzwP+VjNkq+G+JdZqPXgKjkWKKkzd7+Y=; b=YZVcMuk4ONZzx1xb4sdodukXeg5g1cmpMbaRnhSJFZZULn5vtbuOvZRDNBNBH8NuUK zEspVNpUOMUEqfn74GHylyCcPcyXouyUAYLtq+AXxNXmNK8jh7O2V+C33GHWZoTOuAX4 h/8r41KHG36hXCOy0u+EHqvV47bApiwxtmY9Or1a/UzjW28mYlf6Ckg2P0O4+Bb6t937 HxkhYqdiIFW2yZX5Srm7vyWUochuFodRPo7I3reSnhdtCcHlX9NE+2oxNHTRMhykhAC7 cFPEoXtJ+VIm/gl1WcgPJc7K3nBEKeqRPILtP8VNbQCkoTrfur83NvUlz+hVMnDif+bA crgg== X-Forwarded-Encrypted: i=1; AHgh+RqurkwqNRL5NX5upeEDdLBQhzzeEgDB027p6VLbgIAgoITv9o1QQmbiXXsiucApVGxKYWIc97dUUP0SbbA=@vger.kernel.org X-Gm-Message-State: AOJu0YxPbP9Fl1BxPORSCu8ev/GolHMxfZGymkcs83mcALGRryShJSM+ id2/jvg+D+FIafdmS9mQVcDtBwiSuV3QB9hRGzuq+FVIMPvM2lB4FxwM X-Gm-Gg: AfdE7cmscmMQtVTEe7/br7d5npVQySzm8EMx/bd1qlzvEo47aHyPKKNzS6gvrLfj0u4 YjnqLvxf/Ua6aX3I1mJl7NNH83vibqaeiVKyWLUEnluw1xKxTiA2jO1lPhk31R/Kn+Zrc78jJW4 uHxGMuv4FPQHtiCXyWnGQw+wTEb69PHBqMVOSebumheQf+LHPdWECtyJ3r3x3bUyLmEpBbxK66I ccMUjj4JeRTIUFgVo1v62YzaLeOeUbAusZPk0j/5xPm4PDb8xOYZ9hwZ+FJn3icwkaJJWnddgGm RcboaEGRTGQE/sNgjz80GuRbaf8wzmQ4Bpq+iZxWL7N2Tpj+4wZRbJw98cke21jMmO3xGgF9P/n 4BK36cGBTkqhC3Y5kwB7Ss5qbKBQ7hOBBk6MvtxH+lgZv8+Lhily7a45F4rlupun53DQDUrR7qS dW5R8OTjInSngmcPs/H627Pg== X-Received: by 2002:a17:903:2411:b0:2c9:b396:1a55 with SMTP id d9443c01a7336-2ccea394d5dmr102447495ad.12.1783654178893; Thu, 09 Jul 2026 20:29:38 -0700 (PDT) Received: from [127.0.1.1] ([138.199.21.246]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2ccc9bdb56fsm53436465ad.15.2026.07.09.20.29.27 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 09 Jul 2026 20:29:38 -0700 (PDT) From: Jing Wu Date: Fri, 10 Jul 2026 11:28:19 +0800 Subject: [PATCH v4 08/11] cpuset: add dhm_cycling_cpus mask to suppress transient invalidation Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260710-wujing-dhm-v4-8-2e912e5d9645@gmail.com> References: <20260710-wujing-dhm-v4-0-2e912e5d9645@gmail.com> In-Reply-To: <20260710-wujing-dhm-v4-0-2e912e5d9645@gmail.com> To: Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Josh Triplett , Boqun Feng , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Anna-Maria Behnsen , Tejun Heo , Jonathan Corbet , Shuah Khan , Shuah Khan , Thomas Gleixner Cc: Waiman Long , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, cgroups@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, Jing Wu , Qiliang Yuan X-Mailer: b4 0.13.0 When the cpuset code cycles a CPU through hotplug to apply kernel-noise isolation (remove_cpu followed by add_cpu), the CPU disappears from cpu_active_mask temporarily. cpuset_hotplug_update_tasks() sees an empty effective CPU set on the isolated partition and issues partcmd_invalidate, tearing down the partition. The subsequent add_cpu brings the CPU back online, but the partition has already been marked invalid and requires manual user intervention to restore. Add a global dhm_cycling_cpus cpumask protected by dhm_cycling_lock. The isolation cycling path sets the bits for CPUs being cycled before calling remove_cpu(), clears them after add_cpu() completes. cpuset_hotplug_update_tasks() checks whether any of the cpuset's effective exclusive CPUs are in dhm_cycling_cpus and skips the invalidation command when they are, treating the transient empty-CPU state as expected rather than an error. A global cpumask avoids the need to walk the cpuset tree to find the owning cpuset during the cycling loop which runs without cpuset locks. Co-developed-by: Qiliang Yuan Signed-off-by: Qiliang Yuan Signed-off-by: Jing Wu --- kernel/cgroup/cpuset.c | 40 ++++++++++++++++++++++++++++++++++++++-- 1 file changed, 38 insertions(+), 2 deletions(-) diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c index 80f43a24d3c8a..62eb6798a0c3e 100644 --- a/kernel/cgroup/cpuset.c +++ b/kernel/cgroup/cpuset.c @@ -156,6 +156,14 @@ static bool update_housekeeping; /* RWCS */ */ static cpumask_var_t isolated_hk_cpus; /* T */ =20 +/* + * CPUs currently being cycled through hotplug for kernel-noise isolation. + * Protected by dhm_cycling_lock; read in cpuset_hotplug_update_tasks() to + * suppress transient partition invalidation during the offline step. + */ +static DEFINE_SPINLOCK(dhm_cycling_lock); +static cpumask_var_t dhm_cycling_cpus; + /* * A flag to force sched domain rebuild at the end of an operation. * It can be set in @@ -3708,6 +3716,7 @@ int __init cpuset_init(void) BUG_ON(!zalloc_cpumask_var(&subpartitions_cpus, GFP_KERNEL)); BUG_ON(!zalloc_cpumask_var(&isolated_cpus, GFP_KERNEL)); BUG_ON(!zalloc_cpumask_var(&isolated_hk_cpus, GFP_KERNEL)); + BUG_ON(!zalloc_cpumask_var(&dhm_cycling_cpus, GFP_KERNEL)); =20 cpumask_setall(top_cpuset.cpus_allowed); nodes_setall(top_cpuset.mems_allowed); @@ -3804,6 +3813,20 @@ static void cpuset_hotplug_update_tasks(struct cpuse= t *cs, struct tmpmasks *tmp) if (remote && (cpumask_empty(subpartitions_cpus) || (cpumask_empty(&new_cpus) && partition_is_populated(cs, NULL)))) { + bool cycling; + + /* + * Suppress transient invalidation when the offline is part + * of a hotplug cycling step for kernel-noise isolation. + */ + spin_lock(&dhm_cycling_lock); + cycling =3D cpumask_available(dhm_cycling_cpus) && + cpumask_intersects(cs->effective_xcpus, + dhm_cycling_cpus); + spin_unlock(&dhm_cycling_lock); + if (cycling) + goto unlock; + cs->prs_err =3D PERR_HOTPLUG; remote_partition_disable(cs, tmp); compute_effective_cpumask(&new_cpus, cs, parent); @@ -3821,8 +3844,21 @@ static void cpuset_hotplug_update_tasks(struct cpuse= t *cs, struct tmpmasks *tmp) if (is_local_partition(cs) && (!is_partition_valid(parent) || tasks_nocpu_error(parent, cs, &new_cpus) || - cpumask_empty(subpartitions_cpus))) - partcmd =3D partcmd_invalidate; + cpumask_empty(subpartitions_cpus))) { + bool cycling; + + /* + * Suppress transient invalidation when the offline is part + * of a hotplug cycling step for kernel-noise isolation. + */ + spin_lock(&dhm_cycling_lock); + cycling =3D cpumask_available(dhm_cycling_cpus) && + cpumask_intersects(cs->effective_xcpus, + dhm_cycling_cpus); + spin_unlock(&dhm_cycling_lock); + if (!cycling) + partcmd =3D partcmd_invalidate; + } /* * On the other hand, an invalid partition root may be transitioned * back to a regular one with a non-empty effective xcpus. --=20 2.43.0 From nobody Sun Jul 26 01:47:31 2026 Received: from mail-pl1-f173.google.com (mail-pl1-f173.google.com [209.85.214.173]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 711801F3D56 for ; Fri, 10 Jul 2026 03:29:48 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.173 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783654189; cv=none; b=Mj2uB1OPRgpSaAU1y9Zxn/x310laaG5eQg43x89e1jlCXTzY/x6ZKzIoCjQ94iKgdABRfPSPYA6H852k2T+T2oCTki/lzPIy6biwSYhjVC4mvhvlkluJ8bxy7C6dRpPU6OHT67kZELu23CFc/59O9WdZsUocGKFdcKHOIC2OCU8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783654189; c=relaxed/simple; bh=Fu7ulrEF9vMTmeYkdDPOqk6yNZOJFGTqTgSUSx/HiFE=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=TIjGs9K5WDfI3zN9oc1xEqCWLVpTWGAmolK4K51FGYubivdd63sl9LofWGPGZV3VE5lmx1UJB4LENZssvyUCT953rWmRZpNyzn8+7XsxP6pf0bCHMDaf0zXnq9TETkloi7dlfm2+Ss4UMGWe3Y0zB6LIdBpJn+aOdWGu8LWnx+A= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=CPjSwQ1+; arc=none smtp.client-ip=209.85.214.173 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="CPjSwQ1+" Received: by mail-pl1-f173.google.com with SMTP id d9443c01a7336-2cacd69a9c0so3805895ad.1 for ; Thu, 09 Jul 2026 20:29:48 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1783654188; x=1784258988; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=csN4cJukbkvLe5YqSWgDpZ0W4w2lD/0lFTHmWQ8kViM=; b=CPjSwQ1+Uh+vPpiUN05+zQHyfSzuyXF4e3gROBnIIoc2rU4Ll0devqaPDVA9FCir8o 5LVYRJITVLiku3cPATljGHQ1eB1yHBR20dcFXTuo+talptIz++27a0LM6nhRWs0K7Hwa wxACCtUNOziunFnY7mbMCo627vNHOktUsBl5DIJ0G+xZQ8SEnWua8BPiGENm0X3fT7lv uhg9/zkRaDLCxUU4fiaSUPbryuaJ4bQUwalWRnKnkP6dnv3kNXATI2umjMHm4zsmx8Ju ybXTCKBjPEVV0n3jNhEOJbweOsBBCq69kmyrSwG0/rnJZUu/9eX01RydiiYGnLkwtJsn MvDw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783654188; x=1784258988; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=csN4cJukbkvLe5YqSWgDpZ0W4w2lD/0lFTHmWQ8kViM=; b=h13HlXuaPmZp0yVizwenPJBylje+HLGZ6NgbnZY74YQ3rorH8OewQVHxcqjOkHP8yN /jJoe8+sV31XiM5QfBdYLh33+lmduVzf8JOrJSFLtwxKy75LJLDyk/n/g6cvQKEi9az9 1fpEOu2+NrTrLJnyzQQ6CdKJqT+Iy6l0bFiwyYFHgf/pmZEEKOBsrQnsdaDil6x7qZD2 +XMBmi2lbg4pkvU6eeswHGcucAQWK9DmWA6riXTOzTXAsswp+khEwF0JToycoVEr0IW9 EwH23+QzbnrXKCSBIQpQYZfihp4VWDmnZgedRczuNrRd4+U/hLB7WHgK6jgT36ovh/6q gDlg== X-Forwarded-Encrypted: i=1; AHgh+RqzdlHJmegzKo+BX3KP9+sozr0/tUXKzUC3bze5PDddkLH12YKbvLms2zQiPtMN/inNnysBuLRTOZqPzLI=@vger.kernel.org X-Gm-Message-State: AOJu0Yw+b6DfaxUWuwtjNEBLmIWLLcWHb5m7xNqxkvpNn32bExZOMAmB FvtK6rF8thCIG9ZApkKJDFCPT+h81wRlRXa9Vl2nBmNilpB8xmTlC1MF X-Gm-Gg: AfdE7ckd0HgqPJtsK+fbIn1BOTj2ADJ8uyv/mkrfObNMxbqMlDvnZQbtye5hCQulq+u jwnrp2HOPTIqkCD96MMxNZqSJsSPjrurxvnmn29RmRYtl6gWk2Lga7V0c/4036rm+4f31teiull 4vpYqMn8rT4RrWmnJjA9wEB1sfS1t0amofLL23s8HqBkhQPcCGZs3nSHfGsqcGvosBQhuDcxfQk i4gDii5RCWBEAcI5r8RQTxbH2yarsWxqe/eIjhqWmCMGRha/M2oIt5ESJXbxaP0qs4BldNisb02 NLH3wJ2d6509bourCO+LQuD7O0e+NZX9HWQAcM2UpUlxanvU9Qdx3uXvTk/1PWFn6tMDRTF66i/ 83Bxj6nELdyoaKeu3i/XBHZCj90PZbmd7mgDxYcrJixKNfznpLse0sFY2ekWeJQOO+SpGQrhUAm u4V/o0Vmn0oKs= X-Received: by 2002:a17:903:2990:b0:2cc:f4d4:29a7 with SMTP id d9443c01a7336-2ccf4d46bd1mr83112295ad.24.1783654187634; Thu, 09 Jul 2026 20:29:47 -0700 (PDT) Received: from [127.0.1.1] ([138.199.21.246]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2ccc9bdb56fsm53436465ad.15.2026.07.09.20.29.39 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 09 Jul 2026 20:29:47 -0700 (PDT) From: Jing Wu Date: Fri, 10 Jul 2026 11:28:20 +0800 Subject: [PATCH v4 09/11] cpuset: drive kernel-noise isolation via per-CPU hotplug cycling Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260710-wujing-dhm-v4-9-2e912e5d9645@gmail.com> References: <20260710-wujing-dhm-v4-0-2e912e5d9645@gmail.com> In-Reply-To: <20260710-wujing-dhm-v4-0-2e912e5d9645@gmail.com> To: Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Josh Triplett , Boqun Feng , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Anna-Maria Behnsen , Tejun Heo , Jonathan Corbet , Shuah Khan , Shuah Khan , Thomas Gleixner Cc: Waiman Long , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, cgroups@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, Jing Wu , Qiliang Yuan X-Mailer: b4 0.13.0 Track A (cpuset_update_sd_hk_unlock) updates the HK_TYPE_KERNEL_NOISE and HK_TYPE_MANAGED_IRQ cpumasks but performs no per-CPU reconfiguration. Tick suppression, RCU callback offloading and managed-IRQ remapping only take effect when the affected CPUs pass through the CPU hotplug machinery. Implement dhm_cycle_isolated_cpus() and call it from cpuset_update_sd_hk_unlock() after all cpuset and hotplug locks are released so that remove_cpu()/add_cpu() may acquire cpus_write_lock without violating the cpu_hotplug_lock > cpuset_top_mutex order. On isolation, for each newly-isolated CPU: 1. remove_cpu() - offline; dying callbacks migrate IRQs 2. tick_nohz_cpu_isolate() - add to tick_nohz_full_mask, enable context tracking (B0/B3) 3. rcu_nocb_cpu_isolate() - lazy nocb init, spawn kthreads, offload callbacks (B1) 4. add_cpu() - online; tick and IRQ online callbacks reconfigure against updated HK masks On de-isolation, the reverse order is applied. The managed-IRQ remapping requires no explicit call: irq_migrate_all_off_this_cpu() (dying callback) and irq_affinity_online_cpu() (online callback) already consult the updated HK_TYPE_MANAGED_IRQ mask. dhm_prev_isolated tracks the previous isolation set so that only CPUs whose state changed are cycled rather than the full isolation set. lockup_detector_hk_update() (B2) is called once after all CPUs are cycled to update the watchdog mask. CPUs with hotplug disabled (e.g. x86-64 boot CPU) cannot be taken offline and are skipped. On any remove_cpu() failure the corresponding CPU is cleared from dhm_prev_isolated so the next isolation attempt will retry rather than silently treating it as already isolated. Symmetrically, a de-isolation remove_cpu() failure re-sets the bit in dhm_prev_isolated so the CPU remains tracked as still isolated. Co-developed-by: Qiliang Yuan Signed-off-by: Qiliang Yuan Signed-off-by: Jing Wu --- kernel/cgroup/cpuset.c | 98 ++++++++++++++++++++++++++++++++++++++++++++++= ++++ 1 file changed, 98 insertions(+) diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c index 62eb6798a0c3e..d9e121bf14292 100644 --- a/kernel/cgroup/cpuset.c +++ b/kernel/cgroup/cpuset.c @@ -20,6 +20,8 @@ */ #include "cpuset-internal.h" =20 +#include +#include #include #include #include @@ -33,7 +35,9 @@ #include #include #include +#include #include +#include #include #include #include @@ -164,6 +168,14 @@ static cpumask_var_t isolated_hk_cpus; /* T */ static DEFINE_SPINLOCK(dhm_cycling_lock); static cpumask_var_t dhm_cycling_cpus; =20 +/* + * Snapshot of the isolated CPUs from the previous housekeeping update. + * Used to compute the delta (newly isolated / newly de-isolated) so that + * only the changed CPUs are cycled rather than the full isolation set. + * Protected by cpuset_top_mutex. + */ +static cpumask_var_t dhm_prev_isolated; + /* * A flag to force sched domain rebuild at the end of an operation. * It can be set in @@ -1339,6 +1351,84 @@ static bool prstate_housekeeping_conflict(int prstat= e, struct cpumask *new_cpus) return false; } =20 +/* + * dhm_cycle_isolated_cpus - Apply kernel-noise isolation via hotplug cycl= ing + * + * For each CPU newly entering isolation: cycle it offline, configure tick + * suppression and RCU callback offloading while it is offline, then bring + * it back online. The managed-IRQ state is handled automatically by the + * existing irq_migrate_all_off_this_cpu() dying callback and the + * irq_affinity_online_cpu() online callback which both consult the + * already-updated HK_TYPE_MANAGED_IRQ mask. + * + * For each CPU leaving isolation: cycle it offline, de-offload RCU and + * restore the tick, then bring it back online. + * + * Must be called without any cpuset or hotplug locks held. + */ +static void dhm_cycle_isolated_cpus(const struct cpumask *new_isolated) +{ + cpumask_var_t newly_isolated, newly_deisolated; + int cpu; + + if (!alloc_cpumask_var(&newly_isolated, GFP_KERNEL) || + !alloc_cpumask_var(&newly_deisolated, GFP_KERNEL)) { + free_cpumask_var(newly_isolated); + return; + } + + cpumask_andnot(newly_isolated, new_isolated, dhm_prev_isolated); + cpumask_andnot(newly_deisolated, dhm_prev_isolated, new_isolated); + cpumask_copy(dhm_prev_isolated, new_isolated); + + if (cpumask_empty(newly_isolated) && cpumask_empty(newly_deisolated)) + return; + + /* Mark cycling CPUs so cpuset_hotplug_update_tasks skips invalidation */ + spin_lock(&dhm_cycling_lock); + cpumask_or(dhm_cycling_cpus, newly_isolated, newly_deisolated); + spin_unlock(&dhm_cycling_lock); + + for_each_cpu(cpu, newly_isolated) { + if (!cpu_is_hotpluggable(cpu)) { + pr_warn_once("cpuset: CPU%d cannot be isolated (hotplug disabled)\n", + cpu); + cpumask_clear_cpu(cpu, dhm_prev_isolated); + continue; + } + if (remove_cpu(cpu)) { + pr_warn_once("cpuset: failed to offline CPU%d for isolation\n", + cpu); + cpumask_clear_cpu(cpu, dhm_prev_isolated); + continue; + } + WARN_ON_ONCE(tick_nohz_cpu_isolate(cpu)); + WARN_ON_ONCE(rcu_nocb_cpu_isolate(cpu)); + WARN_ON_ONCE(add_cpu(cpu)); + } + + for_each_cpu(cpu, newly_deisolated) { + if (remove_cpu(cpu)) { + pr_warn_once("cpuset: failed to offline CPU%d for de-isolation\n", + cpu); + cpumask_set_cpu(cpu, dhm_prev_isolated); + continue; + } + WARN_ON_ONCE(rcu_nocb_cpu_deoffload(cpu)); + tick_nohz_cpu_deisolate(cpu); + WARN_ON_ONCE(add_cpu(cpu)); + } + + spin_lock(&dhm_cycling_lock); + cpumask_clear(dhm_cycling_cpus); + spin_unlock(&dhm_cycling_lock); + + lockup_detector_hk_update(); + + free_cpumask_var(newly_isolated); + free_cpumask_var(newly_deisolated); +} + /* * cpuset_update_sd_hk_unlock - Rebuild sched domains, update HK & unlock * @@ -1386,6 +1476,13 @@ static void cpuset_update_sd_hk_unlock(void) WARN_ON_ONCE(housekeeping_update_types(noise_types, isolated_hk_cpus)); mutex_unlock(&cpuset_top_mutex); + + /* + * All cpuset and hotplug locks are released. Cycle each + * affected CPU through hotplug to activate tick suppression, + * RCU callback offloading and managed-IRQ remapping. + */ + dhm_cycle_isolated_cpus(isolated_hk_cpus); } else { cpuset_full_unlock(); } @@ -3717,6 +3814,7 @@ int __init cpuset_init(void) BUG_ON(!zalloc_cpumask_var(&isolated_cpus, GFP_KERNEL)); BUG_ON(!zalloc_cpumask_var(&isolated_hk_cpus, GFP_KERNEL)); BUG_ON(!zalloc_cpumask_var(&dhm_cycling_cpus, GFP_KERNEL)); + BUG_ON(!zalloc_cpumask_var(&dhm_prev_isolated, GFP_KERNEL)); =20 cpumask_setall(top_cpuset.cpus_allowed); nodes_setall(top_cpuset.mems_allowed); --=20 2.43.0 From nobody Sun Jul 26 01:47:31 2026 Received: from mail-pl1-f170.google.com (mail-pl1-f170.google.com [209.85.214.170]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6FDA732B131 for ; Fri, 10 Jul 2026 03:29:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.170 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783654198; cv=none; b=iBt0wb9ExjgJtlCsfvwKfcqucRBDFTITm0B59cxuyf5ucutov6bqOE0LQIfz0bBB3qXDKTeR+FYGvq5b4S66r5ufbCbbpIRaxyAyRvHNMtUCk85KSUxHD+gk00Tbwz0dDKK1Z+4DkJ01HCZFA4J2UvNJQPx/ic9KPJLQpUW4zOI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783654198; c=relaxed/simple; bh=MNtb2Qrr3nNt+ub4rjVucp13CHEPWughXQYJLAqNpy4=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=a2HAQ44ba/ynskx+08HJY2HfFVOqM0/Y8EyA1Xn8ItGWlxiqJGML0sOPtbOp5g+BhgbtP4WrDqlKTCYblyxzCr/SdxwkwOqKAW3Z9YVTVoKUkh2HNgaEq64fIN1ymjC7QYU/3xljgg2rj86loAdMBhmKZ0juhqfO7Y8LcLAnfhk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=FOw3hfY3; arc=none smtp.client-ip=209.85.214.170 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="FOw3hfY3" Received: by mail-pl1-f170.google.com with SMTP id d9443c01a7336-2cacf197759so5795375ad.2 for ; Thu, 09 Jul 2026 20:29:57 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1783654197; x=1784258997; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=Cx1US55ZcQE2IXAHU74DmTXGa2F7H+sEKzQcFe01wfY=; b=FOw3hfY3Oj5i9E3RdC0nhn5AdEWFb5cLVgRgArGb4bb6kp/M6RdpixAkzx0OUPJDS/ OCzD2dxGCeiV1VB4CXInwLgPPekz+xhtXDl4y2+YmSa5pu0nmL8y3IIK72XOV5kZDmln Utgz1Mii77S2ODs5veYjNP1vVJpupanprFN5EE7ZUg4o85bOuNpcpiMghscP5Ijsz71E rXtG6SZTWDTNwYG77d3fQGou5JIQZ8OqqhKiRxQ/0WaXggRpIzvnSYDPbftCBkkPlogi 9Utiz4nYmMvhnv8wEbcHsy3/cBDSIf/YTUzVFSx/EMMjTv//mq/relr7ibH0sRUlLyJa bZmw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783654197; x=1784258997; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=Cx1US55ZcQE2IXAHU74DmTXGa2F7H+sEKzQcFe01wfY=; b=SyA+W20bS/Y7+JPmexsQ8P1AtB1zdHgqZTTKjKK3Y9efz9/V02QNut2hxcE2f3pAXa 9tmhVbBYVsVuwZkcgqGXupbst4PKohksyZMzjNlsXfpdADPnTmdex+bMizvsameQaelN saLDyuXR86/yOZ3RzkN4JIlN7pFbE0zah/YruH5iHcjkH8cjhkeb/dMeVCtN09udzVH2 48hOKkrL9ORy8J1DcO3qeZybaIwxPfuv10dAx7gRRlAHVIrsujcgru/nkoY8gP3wism+ ZlNovw5DkiFRZwiU5zlCYzm/4FfhfKKTXtVCE1/XOr5/gCJQQQWZO9E/iyaWFwkUWIow YfOA== X-Forwarded-Encrypted: i=1; AHgh+Rrku7yR63Luk3KDT8bT9R4bP9353wGbMenUAljCNCUPgCJNIWe0+8SCkKv5UDiIHxMkj1DdcbVDcHhn1Z0=@vger.kernel.org X-Gm-Message-State: AOJu0YwWQr7NAfSfqlPmGj+RvZwVQgbhj1L9fQu9WVq/Dodb4NoG/xqa snWhPbDsX0q6OSEINx91apniLxrsmstqB09bNOVTWXWjbgRadGiLDmcR X-Gm-Gg: AfdE7ckOne3y+Nu/3ZC26L7TKMY1wv1WyM3RyTlJlyzTrUASTnnaV+oyCxAtiuNZN0O ArrAeHFAe/XcRNr7XDxS/r1b0PRA0ff26ZBn6SkL06WaHR28Ivmk9hnr+vl3OR1kWhwOPiDliak py8ycaKt19h9EMbM+Q/5SX+hpsjKn5yC+CO+79ZMbKIVcfh3CzoIqrjtEJFxqyU+0k2v/9EACAb Pct/K8pAggy3ZkyhAK5eI03A62xSEo2HqFW8nJE70b27RU4M+DuP9g2Ox+woQ6xDTFATx3hNHKG egSEgR29qf+A27RWTCuUvQM0vuI1vO3K62xj4Marl4qb+ap6stjs//Idclhfzh3xOM2mAhK6/Zh HR9xEBNbolAvQefd/4Sv/0eUcJR5RTVGjdCqv6wi9L2EmuXZRWx4MKMxzmIVeB1V4wu+n0b6YaN 06JSpJwNfg5h8= X-Received: by 2002:a17:903:244e:b0:2ca:d91d:d3a7 with SMTP id d9443c01a7336-2ccea2a5de9mr109420575ad.10.1783654196716; Thu, 09 Jul 2026 20:29:56 -0700 (PDT) Received: from [127.0.1.1] ([138.199.21.246]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2ccc9bdb56fsm53436465ad.15.2026.07.09.20.29.48 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 09 Jul 2026 20:29:56 -0700 (PDT) From: Jing Wu Date: Fri, 10 Jul 2026 11:28:21 +0800 Subject: [PATCH v4 10/11] docs: cgroup-v2: document kernel-noise isolation via isolated partitions Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260710-wujing-dhm-v4-10-2e912e5d9645@gmail.com> References: <20260710-wujing-dhm-v4-0-2e912e5d9645@gmail.com> In-Reply-To: <20260710-wujing-dhm-v4-0-2e912e5d9645@gmail.com> To: Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Josh Triplett , Boqun Feng , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Anna-Maria Behnsen , Tejun Heo , Jonathan Corbet , Shuah Khan , Shuah Khan , Thomas Gleixner Cc: Waiman Long , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, cgroups@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, Jing Wu , Qiliang Yuan X-Mailer: b4 0.13.0 Document that creating a cpuset isolated partition updates the kernel-noise housekeeping masks (HK_TYPE_KERNEL_NOISE and HK_TYPE_MANAGED_IRQ) in addition to the sched-domain mask, and that destroying it restores the boot configuration. No boot-time kernel parameters such as nohz_full=3D or rcu_nocbs=3D are required; writing "isolated" to cpuset.cpus.partition is the only mechanism needed. Co-developed-by: Qiliang Yuan Signed-off-by: Qiliang Yuan Signed-off-by: Jing Wu --- Documentation/admin-guide/cgroup-v2.rst | 17 +++++++++++++++++ 1 file changed, 17 insertions(+) diff --git a/Documentation/admin-guide/cgroup-v2.rst b/Documentation/admin-= guide/cgroup-v2.rst index 6efd0095ed995..eaafe6d88c0e5 100644 --- a/Documentation/admin-guide/cgroup-v2.rst +++ b/Documentation/admin-guide/cgroup-v2.rst @@ -2721,6 +2721,23 @@ Cpuset Interface Files kernel boot command line option. If those CPUs are to be put into a partition, they have to be used in an isolated partition. =20 + When an isolated partition is created or destroyed, the kernel + automatically drives runtime updates of the housekeeping masks + for kernel-noise types (nohz_full, RCU NOCB, managed IRQ + interrupts). This extends isolation beyond scheduler domains: + the tick is stopped on isolated CPUs, RCU callbacks are + offloaded to housekeeping cores, and managed interrupts are + migrated away. No boot-time kernel parameters such as + ``nohz_full=3D`` or ``rcu_nocbs=3D`` are required; writing + ``isolated`` to ``cpuset.cpus.partition`` is the only mechanism + needed. No additional cgroupfs files are required. + + CPUs with hotplug disabled (typically the boot CPU, CPU 0, on + x86-64) cannot be cycled offline for kernel-noise isolation. + The kernel emits a one-time warning and keeps those CPUs in + the tick and RCU-NOCB housekeeping set, even when they appear + in an isolated partition. + =20 Device controller ----------------- --=20 2.43.0 From nobody Sun Jul 26 01:47:31 2026 Received: from mail-pl1-f174.google.com (mail-pl1-f174.google.com [209.85.214.174]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3EEC2205E02 for ; Fri, 10 Jul 2026 03:30:06 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.174 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783654208; cv=none; b=Sni/6maNDzH7cnq0atPd2lSBAYfoJqmsi1vlzBlrYGlNJkJUbd6ZQPqSVJ4kPIeGIE2+3N2pUZ3vkVIQSNx7VhW/yhjKsYbWrI1ml8ZrNxaKZdIk4Yl55RJ3TWoIgqWk17eCVhbedySqI/dg+GJ3sPUb6Mfh66pMQjBOs/xRFnQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783654208; c=relaxed/simple; bh=nWNwzN91YYBr3Xi6bx2bqBw0sf4Bqsx5uZpg5NO3D9g=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=BocSCkwLaDu9XuGX8M3GE9QBjGLZlA5HWHSB5G7eVZDJao717vdYmiNQkWIItEOynFtgDE/HLqB0y4eVfp5zZt4LAgFWSuvU/HdUfl/T/c5YYeUfPL4g8ixXIlgWlvTnRqB1uyKMGz8hXyfWe/TqjLEFxFs/28tYKALt/mdyJZo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=glN+hGKn; arc=none smtp.client-ip=209.85.214.174 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="glN+hGKn" Received: by mail-pl1-f174.google.com with SMTP id d9443c01a7336-2cc8e87f29bso3189245ad.2 for ; Thu, 09 Jul 2026 20:30:06 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1783654205; x=1784259005; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=QvfnVYRZ3tKllx+hig17bTx3GgJpR8NyPboIOl5+PGA=; b=glN+hGKnY9/f3LddK8IkYraGH43ZoanQxvmOpQjElLGCxS9D4pYmRkPR9945RhTgXk ACGY2Etp9TRFx3yEJBWlSPOBO8rSYoluG/sUVfDUF04E3snhXDoarhraKdrb1ozvPJgl mmU4KSHad21hPwWFw4FXQ3swSSVzM1UGn7w9Qsp1BqhPcLWRlnRMZNZPEqiDJnuC4k9z R2BE6iZMqwRtYXl9aKt97wHskFEZ7CErWDbK7SFR5JckXPGEtT/Zw9l+VXQjEQvD5swI AlGW+eWPoENq2+ljFpf3Bsw1t7NNZ9KjisbmwFN6fPt/g4LhcL3kRr8Rf44D9WrE7dJB hRDQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1783654205; x=1784259005; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=QvfnVYRZ3tKllx+hig17bTx3GgJpR8NyPboIOl5+PGA=; b=R8zFEvCU/MQZXxCRQsnj12B5VIuh5ujpdQZ3y6lwiPFQcpB5r9rgYlCzwDFsCK85U+ Okt2uiki6AmNkjZeQQnH1lHNHCZLj+PhEgnQMDN31S7My02ijb+NOZ+g9ULHsmIY84/H eVmta6CmBuVdkbnrvXlPzgQjaIW7wjY8WJU6fUNWMCsyim0B1LVtW079vxdfsYuRe8+t cLAiOQWpxk47zGVHAQyJcjsZoajU8FomE8SbNfOHQBYjO2jbgZ6svul72U0ZN51uPYBd noI3Izz2UYOjEPLNIUOfSRmov/41TLD0wRjwgQRPWOGMFRgE936mS7m/b6hYSlPTHD14 oJDA== X-Forwarded-Encrypted: i=1; AHgh+RpzosuAYmy4pH8wL6hvT75d8fagHIJCO+puw3Nep1o1Nbdv/5aMjfqKeaMGrpG4aMAJ5pqY4g2GZKbIcI0=@vger.kernel.org X-Gm-Message-State: AOJu0YzgqB1OzO+Yix+/vEyFrO0jNuHCkXJtWyps+lupAR1ni/EDwA+n /ENHvq8xQpsDQZV9dhn6uf5souqXqw1NLP7DgcYJz2ukAM6LVeOiAo44 X-Gm-Gg: AfdE7ckhp7j+ZDzCY4enV2b7f6WAK8p4Gvv3V3eCxGmuDz145FLiTW7UG8wU8fL5K25 ME/QiO1KnuaQnY7Qxfy4GMPNHWqsHp/MAOiFUvoUf7ybuqCcW0UcQ7VuoYrm+j7vKw6GSIt7a3J sz4oQqlhMRAfzzdLV99ahYev8kOfuI1AyFcCF7VYze/cY5hbQVwYAlyuehWY1U5vdIpDHW6zDLP ZgoURC09lYzlEpf269NdomhG/XBFjKmYP9h8Xpe/jfGrpVdQQ7lDYHUbPmyKLgiuCD8lZGGqkn5 jZlJ0zI6MvuPd+CzZXtr96yrQWXDSo6d0r6+dl+RGdyJ6UEobUX4d/MAVCHaMgC7L937RwUqhHB 1iyFxS1v1hpkYi4tccFbFK2P6ANF+NyclUHQp/pRU0b4KDTyXvXuJ3tzbA5pXOlANhFiScTTbRS 4bwZW2va3fYSQ= X-Received: by 2002:a17:902:e950:b0:2ca:3df2:919a with SMTP id d9443c01a7336-2ccea3a28c4mr91454235ad.33.1783654205335; Thu, 09 Jul 2026 20:30:05 -0700 (PDT) Received: from [127.0.1.1] ([138.199.21.246]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2ccc9bdb56fsm53436465ad.15.2026.07.09.20.29.56 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 09 Jul 2026 20:30:04 -0700 (PDT) From: Jing Wu Date: Fri, 10 Jul 2026 11:28:22 +0800 Subject: [PATCH v4 11/11] selftests/cgroup: add kernel-noise isolation test to cpuset selftest Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260710-wujing-dhm-v4-11-2e912e5d9645@gmail.com> References: <20260710-wujing-dhm-v4-0-2e912e5d9645@gmail.com> In-Reply-To: <20260710-wujing-dhm-v4-0-2e912e5d9645@gmail.com> To: Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , "Paul E. McKenney" , Frederic Weisbecker , Neeraj Upadhyay , Joel Fernandes , Josh Triplett , Boqun Feng , Uladzislau Rezki , Mathieu Desnoyers , Lai Jiangshan , Zqiang , Anna-Maria Behnsen , Tejun Heo , Jonathan Corbet , Shuah Khan , Shuah Khan , Thomas Gleixner Cc: Waiman Long , linux-kernel@vger.kernel.org, rcu@vger.kernel.org, cgroups@vger.kernel.org, linux-doc@vger.kernel.org, linux-kselftest@vger.kernel.org, Jing Wu , Qiliang Yuan X-Mailer: b4 0.13.0 Add test_hk_noise_isolated() to test_cpuset_prs.sh to verify that creating and destroying an isolated partition updates the kernel-noise housekeeping state, including the /sys/devices/system/cpu/nohz_full attribute. Add the cpu_in_cpulist() helper to correctly test membership against a cpulist that may contain ranges. Also detect and report whether the test is running in zero-boot-param mode (no nohz_full=3D in /proc/cmdline). When in zero-boot-param mode the test confirms that nohz_full is activated by the DHM runtime path, verifying that dhm_cycle_isolated_cpus() correctly enables tick isolation without any boot-time setup. Co-developed-by: Qiliang Yuan Signed-off-by: Qiliang Yuan Signed-off-by: Jing Wu --- tools/testing/selftests/cgroup/test_cpuset_prs.sh | 580 ++++++++++++++++++= +++- 1 file changed, 579 insertions(+), 1 deletion(-) diff --git a/tools/testing/selftests/cgroup/test_cpuset_prs.sh b/tools/test= ing/selftests/cgroup/test_cpuset_prs.sh index a56f4153c64df..ce16734545505 100755 --- a/tools/testing/selftests/cgroup/test_cpuset_prs.sh +++ b/tools/testing/selftests/cgroup/test_cpuset_prs.sh @@ -20,7 +20,7 @@ skip_test() { WAIT_INOTIFY=3D$(cd $(dirname $0); pwd)/wait_inotify =20 # Find cgroup v2 mount point -CGROUP2=3D$(mount -t cgroup2 | head -1 | awk -e '{print $3}') +CGROUP2=3D$(mount -t cgroup2 | head -1 | awk '{print $3}') [[ -n "$CGROUP2" ]] || skip_test "Cgroup v2 mount point not found!" SUBPARTS_CPUS=3D$CGROUP2/.__DEBUG__.cpuset.cpus.subpartitions CPULIST=3D$(cat $CGROUP2/cpuset.cpus.effective) @@ -1204,9 +1204,587 @@ test_inotify() echo "" > cpuset.cpus } =20 +# +# cpu_in_cpulist +# +# Return 0 if appears in (a kernel cpumask list such as +# "0-3,8-31"), non-zero otherwise. The kernel cpulist format uses ranges +# ("lo-hi") and comma-separated items; a simple grep cannot detect that a +# number falls in the middle of a range, so walk each element explicitly. +# +cpu_in_cpulist() +{ + local cpu=3D$1 list=3D$2 range lo hi + for range in $(echo "$list" | tr ',' ' '); do + if [[ "$range" =3D=3D *-* ]]; then + lo=3D${range%-*} + hi=3D${range#*-} + [[ $cpu -ge $lo && $cpu -le $hi ]] && return 0 + else + [[ $cpu -eq $range ]] && return 0 + fi + done + return 1 +} + +# +# verify_nohz_exact +# +# Verify that nohz_full equals =E2=88=AA . +# may be a cpulist range ("4-7") or empty string (""). +# Catches both missing CPUs and unexpected extra CPUs. +# +verify_nohz_exact() +{ + local baseline=3D$1 active=3D$2 current=3D$3 cpu exp got + for cpu in $(seq 0 $((NR_CPUS - 1))); do + exp=3D0; got=3D0 + cpu_in_cpulist $cpu "$baseline" && exp=3D1 + [[ -n "$active" ]] && cpu_in_cpulist $cpu "$active" && exp=3D1 + cpu_in_cpulist $cpu "$current" && got=3D1 + [[ $exp -eq $got ]] || { + if [[ $got -eq 0 ]]; then + echo "FAIL: cpu${cpu} expected in nohz_full but absent" \ + "(baseline=3D'$baseline' active=3D'$active'" \ + "current=3D'$current')" + else + echo "FAIL: cpu${cpu} unexpectedly in nohz_full" \ + "(baseline=3D'$baseline' active=3D'$active'" \ + "current=3D'$current')" + fi + return 1 + } + done + return 0 +} + +# +# Test that isolated partition creation/destruction drives kernel-noise +# housekeeping mask updates and remains correct under pressure. +# +# Requires: >=3D8 CPUs, no isolcpus=3D boot conflict, root +# + +# +# hk_noise_check_nocb_affinity +# +# When expect_isolated=3D1: verify rcuop/N kthreads for CPUs in cpulist do= NOT +# include those CPUs in their scheduler affinity (RCU NOCB active =E2=80= =94 callbacks +# for the isolated CPU are offloaded to a different CPU). +# When expect_isolated=3D0: verify CPUs are back in affinity (NOCB restore= d). +# Silently skips CPUs whose rcuop/N thread is absent (no NOCB support). +# +hk_noise_check_nocb_affinity() +{ + local cpulist=3D$1 expect_isolated=3D$2 + local cpu lo hi range pid aff_hex rev_hex nibble_pos nibble_char + local nibble_val bit_in_nibble bit failed=3D0 + + command -v taskset > /dev/null 2>&1 || return 0 + + for range in $(echo "$cpulist" | tr ',' ' '); do + if [[ "$range" =3D=3D *-* ]]; then + lo=3D${range%-*}; hi=3D${range#*-} + else + lo=3D$range; hi=3D$range + fi + for cpu in $(seq "$lo" "$hi"); do + pid=3D$(ps -eo pid,comm | awk -v c=3D"rcuop/$cpu" '$2=3D=3Dc{print $1}') + [[ -n "$pid" ]] || continue + aff_hex=3D$(taskset -p "$pid" 2>/dev/null | awk '{print $NF}') + [[ -n "$aff_hex" ]] || continue + + # Extract bit from the hex affinity mask. + # Each hex digit covers 4 CPUs; reverse the string to + # work from the LSB side. + rev_hex=3D$(echo "$aff_hex" | rev) + nibble_pos=3D$((cpu / 4)) + nibble_char=3D${rev_hex:$nibble_pos:1} + if [[ -z "$nibble_char" ]]; then + nibble_val=3D0 + else + nibble_val=3D$((16#$nibble_char)) + fi + bit_in_nibble=3D$((cpu % 4)) + bit=3D$(( (nibble_val >> bit_in_nibble) & 1 )) + + if [[ $expect_isolated -eq 1 && $bit -eq 1 ]]; then + echo "FAIL: rcuop/$cpu affinity still includes" \ + "CPU$cpu after isolation (mask=3D0x$aff_hex)" + failed=3D1 + elif [[ $expect_isolated -eq 0 && $bit -eq 0 ]]; then + echo "FAIL: rcuop/$cpu affinity still excludes" \ + "CPU$cpu after de-isolation (mask=3D0x$aff_hex)" + failed=3D1 + fi + done + done + return $failed +} + +test_hk_noise_isolated() +{ + local ISOL_BEFORE TEST_CPUS i PART ISOL_AFTER ISOL_RESTORE + local NOHZ_FILE NOHZ_BEFORE NOHZ_AFTER NOHZ_RESTORE + local HK_NOHZ_CHECK=3D0 + local LOOPS=3D100 + local CMDLINE HAS_BOOT_NOHZ=3D0 + local DMESG_LINES_START + DMESG_LINES_START=3D$(dmesg | wc -l) + + [[ $NR_CPUS -ge 8 ]] || { + echo "HK-noise test skipped: need >=3D8 CPUs, have $NR_CPUS" + return 0 + } + + # Detect whether CONFIG_NO_HZ_FULL is active: the sysfs attribute + # /sys/devices/system/cpu/nohz_full exposes the current nohz_full + # cpumask and is only present when NO_HZ_FULL is enabled. + NOHZ_FILE=3D/sys/devices/system/cpu/nohz_full + [[ -r "$NOHZ_FILE" ]] && HK_NOHZ_CHECK=3D1 + + # Determine if running in zero-boot-param mode. DHM activates tick + # and RCU-NOCB isolation at runtime; no nohz_full=3D or rcu_nocbs=3D + # kernel boot parameters are required. + { read -r CMDLINE < /proc/cmdline; } 2>/dev/null || CMDLINE=3D"" + [[ $CMDLINE =3D *nohz_full=3D* ]] && HAS_BOOT_NOHZ=3D1 + if [[ $HAS_BOOT_NOHZ -eq 0 ]]; then + console_msg "HK-noise: zero-boot-param mode" \ + "(no nohz_full=3D in /proc/cmdline -- testing DHM runtime pa= th)" + else + console_msg "HK-noise: boot-param mode (nohz_full=3D present at boot)" + fi + + cd $CGROUP2/test + echo member > cpuset.cpus.partition 2>/dev/null + echo "" > cpuset.cpus 2>/dev/null + + ISOL_BEFORE=3D$(cat $CGROUP2/cpuset.cpus.isolated) + [[ $HK_NOHZ_CHECK -eq 1 ]] && NOHZ_BEFORE=3D$(cat $NOHZ_FILE) + TEST_CPUS=3D"4-7" + echo $TEST_CPUS > cpuset.cpus + + # + # Basic create/destroy cycle =E2=80=94 verify domain isolation and + # kernel-noise (nohz_full) changes together. + # + console_msg "HK-noise: basic create/destroy cycle" + echo isolated > cpuset.cpus.partition + + ISOL_AFTER=3D$(cat $CGROUP2/cpuset.cpus.isolated) + [[ $ISOL_AFTER !=3D "$ISOL_BEFORE" ]] || { + echo "FAIL: isolated set unchanged after partition create" + exit 1 + } + + if [[ $HK_NOHZ_CHECK -eq 1 ]]; then + NOHZ_AFTER=3D$(cat $NOHZ_FILE) + verify_nohz_exact "$NOHZ_BEFORE" "$TEST_CPUS" "$NOHZ_AFTER" || exit 1 + console_msg "HK-noise: nohz_full after isolation: $NOHZ_AFTER" + fi + + # Verify RCU NOCB: rcuop/N kthreads for isolated CPUs must have those + # CPUs removed from their scheduler affinity mask. + hk_noise_check_nocb_affinity "$TEST_CPUS" 1 || exit 1 + + echo member > cpuset.cpus.partition + + ISOL_RESTORE=3D$(cat $CGROUP2/cpuset.cpus.isolated) + [[ $ISOL_RESTORE =3D "$ISOL_BEFORE" ]] || { + echo "FAIL: expected '$ISOL_BEFORE' after destroy, got '$ISOL_RESTORE'" + exit 1 + } + + if [[ $HK_NOHZ_CHECK -eq 1 ]]; then + NOHZ_RESTORE=3D$(cat $NOHZ_FILE) + verify_nohz_exact "$NOHZ_BEFORE" "" "$NOHZ_RESTORE" || exit 1 + fi + + # Verify RCU NOCB restored: isolated CPUs must reappear in rcuop/N affini= ty. + hk_noise_check_nocb_affinity "$TEST_CPUS" 0 || exit 1 + + # + # Reject all-CPU isolation (must leave at least one housekeeping CPU) + # + console_msg "HK-noise: reject all-CPU isolation" + echo 0-$((NR_CPUS - 1)) > cpuset.cpus + echo isolated > cpuset.cpus.partition + PART=3D$(cat cpuset.cpus.partition) + [[ $PART =3D *invalid* || $PART =3D member ]] || { + echo "FAIL: all-CPU isolation was not rejected, got '$PART'" + exit 1 + } + + # + # SMT safety: partial sibling isolation + # + console_msg "HK-noise: SMT sibling constraint" + echo $TEST_CPUS > cpuset.cpus + echo isolated > cpuset.cpus.partition + PART=3D$(cat cpuset.cpus.partition) + [[ $PART =3D isolated ]] || { + echo "FAIL: could not create isolated partition, got '$PART'" + exit 1 + } + echo member > cpuset.cpus.partition + + # + # Non-hotpluggable CPU: must be skipped with a kernel warning without + # rejecting the partition; hotpluggable peers must still be isolated. + # + # A CPU whose online file is absent (e.g. CPU 0 on x86-64) has hotplug + # disabled. DHM emits pr_warn_once and keeps it in the tick/RCU-NOCB + # housekeeping set; it must not appear in nohz_full after isolation. + # The remaining hotpluggable CPUs in the partition must still be isolated. + # + local FIXED_CPU=3D"" c NOHZ_NOW + for c in $(seq 0 $((NR_CPUS - 1))); do + [[ -f /sys/devices/system/cpu/cpu${c}/online ]] || { + FIXED_CPU=3D$c + break + } + done + if [[ -n "$FIXED_CPU" ]]; then + console_msg "HK-noise: non-hotpluggable CPU${FIXED_CPU} skip" + echo "${FIXED_CPU},${TEST_CPUS}" > cpuset.cpus + echo isolated > cpuset.cpus.partition + PART=3D$(cat cpuset.cpus.partition) + [[ $PART =3D isolated ]] || { + echo "FAIL: partition rejected when including non-hotpluggable" \ + "CPU${FIXED_CPU}: got '$PART'" + exit 1 + } + if [[ $HK_NOHZ_CHECK -eq 1 ]]; then + NOHZ_NOW=3D$(cat $NOHZ_FILE) + if cpu_in_cpulist $FIXED_CPU "$NOHZ_NOW"; then + echo "FAIL: non-hotpluggable CPU${FIXED_CPU} appeared" \ + "in nohz_full (should be skipped)" + exit 1 + fi + local lo hi + lo=3D${TEST_CPUS%%-*} + hi=3D${TEST_CPUS##*-} + for cpu in $(seq "$lo" "$hi"); do + if ! cpu_in_cpulist $cpu "$NOHZ_NOW"; then + echo "FAIL: hotpluggable cpu${cpu} missing from" \ + "nohz_full in mixed partition (got: '$NOHZ_NOW')" + exit 1 + fi + done + console_msg "HK-noise: CPU${FIXED_CPU} absent, ${TEST_CPUS} present" \ + "in nohz_full: $NOHZ_NOW" + fi + echo member > cpuset.cpus.partition + echo $TEST_CPUS > cpuset.cpus + else + console_msg "HK-noise: all CPUs hotpluggable; skip non-hotpluggable subt= est" + fi + + # + # Delta isolation: modify cpuset.cpus while the partition is isolated. + # dhm_prev_isolated must track the delta and update nohz_full in step. + # + local lo hi mid lower_cpus upper_cpu + lo=3D${TEST_CPUS%%-*} + hi=3D${TEST_CPUS##*-} + mid=3D$(( lo + (hi - lo) / 2 )) + lower_cpus=3D"${lo}-${mid}" + upper_cpu=3D$(( mid + 1 )) + console_msg "HK-noise: delta isolation (shrink ${TEST_CPUS} =E2=86=92 ${l= ower_cpus})" + echo $TEST_CPUS > cpuset.cpus + echo isolated > cpuset.cpus.partition + PART=3D$(cat cpuset.cpus.partition) + [[ $PART =3D isolated ]] || { + echo "FAIL: delta test: initial isolation failed, got '$PART'" + exit 1 + } + echo $lower_cpus > cpuset.cpus + PART=3D$(cat cpuset.cpus.partition) + if [[ $PART =3D isolated ]]; then + if [[ $HK_NOHZ_CHECK -eq 1 ]]; then + NOHZ_NOW=3D$(cat $NOHZ_FILE) + verify_nohz_exact "$NOHZ_BEFORE" "$lower_cpus" "$NOHZ_NOW" || exit 1 + fi + # Expand back to full TEST_CPUS and re-verify + echo $TEST_CPUS > cpuset.cpus + if [[ $HK_NOHZ_CHECK -eq 1 ]]; then + NOHZ_NOW=3D$(cat $NOHZ_FILE) + verify_nohz_exact "$NOHZ_BEFORE" "$TEST_CPUS" "$NOHZ_NOW" || exit 1 + fi + echo member > cpuset.cpus.partition + if [[ $HK_NOHZ_CHECK -eq 1 ]]; then + NOHZ_NOW=3D$(cat $NOHZ_FILE) + verify_nohz_exact "$NOHZ_BEFORE" "" "$NOHZ_NOW" || exit 1 + fi + else + console_msg "HK-noise: delta test: partition invalidated on shrink" \ + "('$PART') -- cpuset constraint, not a DHM bug; skipping" + echo member > cpuset.cpus.partition 2>/dev/null || true + fi + echo $TEST_CPUS > cpuset.cpus + + # + # Nested partition: parent root =E2=86=92 child isolated + # + console_msg "HK-noise: nested partition inheritance" + echo $TEST_CPUS > cpuset.cpus + test_partition root + mkdir -p HK_SUB + cd HK_SUB + echo "${lo}-$((lo + 1))" > cpuset.cpus + echo isolated > cpuset.cpus.partition + ISOL_AFTER=3D$(cat $CGROUP2/cpuset.cpus.isolated) + [[ -n $ISOL_AFTER ]] || { + echo "FAIL: nested isolated partition not reflected in cpuset.cpus.isola= ted" + exit 1 + } + echo member > cpuset.cpus.partition + cd $CGROUP2/test + echo member > cpuset.cpus.partition + rmdir HK_SUB 2>/dev/null + + # + # Pressure test: 100 create/destroy cycles with nohz_full verified + # on every cycle to catch mid-run state corruption. + # + console_msg "HK-noise: pressure test ($LOOPS cycles)" + echo $TEST_CPUS > cpuset.cpus + for i in $(seq 1 $LOOPS); do + echo isolated > cpuset.cpus.partition + PART=3D$(cat cpuset.cpus.partition) + [[ $PART =3D isolated ]] || { + echo "FAIL: cycle $i create failed, got '$PART'" + exit 1 + } + if [[ $HK_NOHZ_CHECK -eq 1 ]]; then + NOHZ_NOW=3D$(cat $NOHZ_FILE) + verify_nohz_exact "$NOHZ_BEFORE" "$TEST_CPUS" "$NOHZ_NOW" || { + echo "FAIL: nohz_full wrong at cycle $i (isolated)" + exit 1 + } + fi + echo member > cpuset.cpus.partition + PART=3D$(cat cpuset.cpus.partition) + [[ $PART =3D member ]] || { + echo "FAIL: cycle $i destroy failed, got '$PART'" + exit 1 + } + if [[ $HK_NOHZ_CHECK -eq 1 ]]; then + NOHZ_NOW=3D$(cat $NOHZ_FILE) + verify_nohz_exact "$NOHZ_BEFORE" "" "$NOHZ_NOW" || { + echo "FAIL: nohz_full wrong at cycle $i (member)" + exit 1 + } + fi + done + + # + # Stability: after pressure test, verify final state + # + console_msg "HK-noise: post-pressure cleanup" + echo isolated > cpuset.cpus.partition + ISOL_AFTER=3D$(cat $CGROUP2/cpuset.cpus.isolated) + [[ -n $ISOL_AFTER ]] || { + echo "FAIL: isolated set empty after pressure test" + exit 1 + } + echo member > cpuset.cpus.partition + echo "" > cpuset.cpus + ISOL_RESTORE=3D$(cat $CGROUP2/cpuset.cpus.isolated) + [[ $ISOL_RESTORE =3D "$ISOL_BEFORE" ]] || { + echo "FAIL: final isolated '$ISOL_RESTORE' !=3D '$ISOL_BEFORE'" + exit 1 + } + + if [[ $HK_NOHZ_CHECK -eq 1 ]]; then + NOHZ_RESTORE=3D$(cat $NOHZ_FILE) + [[ "$NOHZ_RESTORE" =3D "$NOHZ_BEFORE" ]] || { + echo "FAIL: nohz_full not restored after pressure test:" \ + "expected '$NOHZ_BEFORE', got '$NOHZ_RESTORE'" + exit 1 + } + fi + + # + # Pressure with resident task: create/destroy cycles while a sleeping + # task occupies the isolated partition. Exercises the dhm_cycling_cpus + # suppression path that prevents false partition invalidation when a + # task is present during hotplug cycling steps. + # + console_msg "HK-noise: pressure with resident task ($LOOPS cycles)" + echo $TEST_CPUS > cpuset.cpus + sleep 600 & + local TASK_PID=3D$! + echo $TASK_PID > cgroup.procs + for i in $(seq 1 $LOOPS); do + echo isolated > cpuset.cpus.partition + PART=3D$(cat cpuset.cpus.partition) + [[ $PART =3D isolated ]] || { + echo "FAIL: task-occupied cycle $i create failed, got '$PART'" + kill $TASK_PID 2>/dev/null; wait $TASK_PID 2>/dev/null + exit 1 + } + if [[ $HK_NOHZ_CHECK -eq 1 ]]; then + NOHZ_NOW=3D$(cat $NOHZ_FILE) + verify_nohz_exact "$NOHZ_BEFORE" "$TEST_CPUS" "$NOHZ_NOW" || { + kill $TASK_PID 2>/dev/null; wait $TASK_PID 2>/dev/null + exit 1 + } + fi + echo member > cpuset.cpus.partition + PART=3D$(cat cpuset.cpus.partition) + [[ $PART =3D member ]] || { + echo "FAIL: task-occupied cycle $i destroy failed, got '$PART'" + kill $TASK_PID 2>/dev/null; wait $TASK_PID 2>/dev/null + exit 1 + } + done + kill $TASK_PID 2>/dev/null + wait $TASK_PID 2>/dev/null + echo "" > cpuset.cpus + + # + # Concurrent partitions: two sibling cgroups each holding half of + # TEST_CPUS simultaneously isolated. Verifies independent isolation + # and correct union in cpuset.cpus.isolated / nohz_full. + # + local lo hi mid lower_half upper_half NOHZ_BOTH + lo=3D${TEST_CPUS%%-*}; hi=3D${TEST_CPUS##*-} + mid=3D$(( lo + (hi - lo) / 2 )) + lower_half=3D"${lo}-${mid}" + upper_half=3D"$((mid + 1))-${hi}" + console_msg "HK-noise: concurrent partitions ($lower_half and $upper_half= )" + echo $TEST_CPUS > cpuset.cpus + echo root > cpuset.cpus.partition + mkdir -p HK_A HK_B + echo "$lower_half" > HK_A/cpuset.cpus + echo isolated > HK_A/cpuset.cpus.partition + echo "$upper_half" > HK_B/cpuset.cpus + echo isolated > HK_B/cpuset.cpus.partition + local PA PB + PA=3D$(cat HK_A/cpuset.cpus.partition) + PB=3D$(cat HK_B/cpuset.cpus.partition) + [[ $PA =3D isolated ]] || { + echo "FAIL: concurrent partition HK_A not isolated, got '$PA'" + echo member > HK_A/cpuset.cpus.partition 2>/dev/null + echo member > HK_B/cpuset.cpus.partition 2>/dev/null + rmdir HK_A HK_B 2>/dev/null + echo member > cpuset.cpus.partition 2>/dev/null + exit 1 + } + [[ $PB =3D isolated ]] || { + echo "FAIL: concurrent partition HK_B not isolated, got '$PB'" + echo member > HK_A/cpuset.cpus.partition 2>/dev/null + echo member > HK_B/cpuset.cpus.partition 2>/dev/null + rmdir HK_A HK_B 2>/dev/null + echo member > cpuset.cpus.partition 2>/dev/null + exit 1 + } + if [[ $HK_NOHZ_CHECK -eq 1 ]]; then + NOHZ_BOTH=3D$(cat $NOHZ_FILE) + verify_nohz_exact "$NOHZ_BEFORE" "$TEST_CPUS" "$NOHZ_BOTH" || { + echo member > HK_A/cpuset.cpus.partition 2>/dev/null + echo member > HK_B/cpuset.cpus.partition 2>/dev/null + rmdir HK_A HK_B 2>/dev/null + echo member > cpuset.cpus.partition 2>/dev/null + exit 1 + } + fi + hk_noise_check_nocb_affinity "$TEST_CPUS" 1 || { + echo member > HK_A/cpuset.cpus.partition 2>/dev/null + echo member > HK_B/cpuset.cpus.partition 2>/dev/null + rmdir HK_A HK_B 2>/dev/null + echo member > cpuset.cpus.partition 2>/dev/null + exit 1 + } + echo member > HK_A/cpuset.cpus.partition + echo member > HK_B/cpuset.cpus.partition + rmdir HK_A HK_B + echo member > cpuset.cpus.partition + if [[ $HK_NOHZ_CHECK -eq 1 ]]; then + NOHZ_NOW=3D$(cat $NOHZ_FILE) + verify_nohz_exact "$NOHZ_BEFORE" "" "$NOHZ_NOW" || exit 1 + fi + hk_noise_check_nocb_affinity "$TEST_CPUS" 0 || exit 1 + echo "" > cpuset.cpus + + # + # Concurrent create/destroy: two background subshells race on the same + # partition simultaneously. Verifies that concurrent cpuset writes + # do not corrupt kernel state or trigger warnings. + # + console_msg "HK-noise: concurrent create/destroy race" + echo $TEST_CPUS > cpuset.cpus + local RACE_LOOPS=3D30 RACE_PID1 RACE_PID2 + (for i in $(seq 1 $RACE_LOOPS); do + echo isolated > cpuset.cpus.partition 2>/dev/null + echo member > cpuset.cpus.partition 2>/dev/null + done) & + RACE_PID1=3D$! + (for i in $(seq 1 $RACE_LOOPS); do + echo member > cpuset.cpus.partition 2>/dev/null + echo isolated > cpuset.cpus.partition 2>/dev/null + done) & + RACE_PID2=3D$! + wait $RACE_PID1 $RACE_PID2 + # Drive to a known-good state regardless of who won the last write. + echo member > cpuset.cpus.partition 2>/dev/null || true + PART=3D$(cat cpuset.cpus.partition) + [[ $PART =3D member ]] || { + echo "FAIL: concurrent race left partition in bad state: '$PART'" + exit 1 + } + if [[ $HK_NOHZ_CHECK -eq 1 ]]; then + NOHZ_NOW=3D$(cat $NOHZ_FILE) + verify_nohz_exact "$NOHZ_BEFORE" "" "$NOHZ_NOW" || exit 1 + fi + echo "" > cpuset.cpus + + # + # Kernel hard-error check: none of the above scenarios must have + # triggered a BUG, OOPS, panic, or RCU stall in the kernel log. + # Also catch WARNINGs that implicate our subsystems (cpuset / rcu / + # nohz / housekeeping / irq_affinity); ignore unrelated WARNINGs from + # other kernel subsystems or user-space processes. + # + console_msg "HK-noise: checking kernel log for errors" + local new_errors + new_errors=3D$(dmesg | tail -n "+$((DMESG_LINES_START + 1))" | \ + grep -c -E \ + 'kernel BUG at|OOPS|Kernel panic|RCU Stall|scheduling while atomic' \ + || true) + local new_subsys_warns + new_subsys_warns=3D$(dmesg | tail -n "+$((DMESG_LINES_START + 1))" | \ + grep 'WARNING:' | \ + grep -c -E 'cpuset|rcu|nohz|housekeeping|irq_affinity|dhm' \ + || true) + local total_errors=3D$(( new_errors + new_subsys_warns )) + [[ $total_errors -eq 0 ]] || { + echo "FAIL: $total_errors kernel error(s)/warning(s) during HK-noise tes= t:" + dmesg | tail -n "+$((DMESG_LINES_START + 1))" | \ + grep -E 'kernel BUG at|OOPS|Kernel panic|RCU Stall|scheduling while ato= mic' | head -10 + dmesg | tail -n "+$((DMESG_LINES_START + 1))" | \ + grep 'WARNING:' | grep -E 'cpuset|rcu|nohz|housekeeping|irq_affinity|dh= m' | head -10 + exit 1 + } + + cd $CGROUP2 + if [[ $HK_NOHZ_CHECK -eq 1 ]]; then + if [[ $HAS_BOOT_NOHZ -eq 0 ]]; then + console_msg "HK-noise: PASSED" \ + "(zero-boot-param: nohz_full verified via DHM runtime path)" + else + console_msg "HK-noise: PASSED (with nohz_full verification)" + fi + else + console_msg "HK-noise: PASSED (nohz_full skipped: CONFIG_NO_HZ_FULL not = active)" + fi +} + trap cleanup 0 2 3 6 run_state_test TEST_MATRIX run_remote_state_test REMOTE_TEST_MATRIX test_isolated test_inotify +test_hk_noise_isolated echo "All tests PASSED." --=20 2.43.0