From nobody Thu Sep 24 17:53:07 2026 Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.15]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 097A8316905; Tue, 22 Sep 2026 00:32:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=192.198.163.15 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790037134; cv=none; b=ThFGQ8kfjY6Q9g6v1/SNrei4hRwjM6kbwHe0Q26ZykmG/9fx3HLBiENVFZJizd7pK3g4POu+27AhkKM6efm/Vsq55Pf/JqsIpEuUinAiD6F9HmeigdNmyaKL0mGILEuj1FgrFoCe03R3sdraRAzlrVCMLLy29dl2u5qYJrFPl0M= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790037134; c=relaxed/simple; bh=TLr+muP9IJPeDswlAwrJ9DVaMebbbJPEc9g+IiyTn3Q=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=oPv0i8F1rOCveWF3GvgDtx+7CWxkwFatd2YyIyaHAk8v46wDB0j7QiNoXipJNsPkWRYSW50dkZC4kZjtFXLmSG7AcDfD2B4XJYw2u5yPDRXUv59Sc5GwRLm6uBeBhtIe/7Fwma8VNTWp4bKFF2fZhfPg83zI6ChSajXSaVi1XIo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com; spf=pass smtp.mailfrom=linux.intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=K7aVy/LA; arc=none smtp.client-ip=192.198.163.15 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="K7aVy/LA" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1790037133; x=1821573133; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=TLr+muP9IJPeDswlAwrJ9DVaMebbbJPEc9g+IiyTn3Q=; b=K7aVy/LAJAY+dtrFY78rakivrFWttR/CDM02IsMCD/lZhT0p4G+LChp/ rU5/ISy4/LzJLPk6t0nWgveznpHAK/q88XY4P/4TzJI6gh+Kzjh+nNKwu xmnfVtWYSUMiRoa9xyR+Dm/7CnR3F/W0hC0xiMTUtFtgh4tT0CxaQmG4M qf+eVQBpbOpLo1gGjim7IdZ1ga1JYs9khznrkNTpUStEiTLrg6woT7sEq DMSYZRLLHfycYhebm0pbDGdNuD9SCc0YXDmNJR2j78xyObJ34Ri5jYhp/ INhQHapuGH9v2/XbQvjYtQjha76A+Hf7acIUF+T49X4v+l/2kZoQSwa7w A==; X-CSE-ConnectionGUID: PMPX99m5TEaalj0Y3EO+gQ== X-CSE-MsgGUID: RYUIYs9jQkGzBm2xVTjYpg== X-IronPort-AV: E=McAfee;i="6800,10657,11912"; a="90716634" X-IronPort-AV: E=Sophos;i="6.27,115,1787036400"; d="scan'208";a="90716634" Received: from fmviesa004.fm.intel.com ([10.60.135.144]) by fmvoesa109.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 21 Sep 2026 17:32:12 -0700 X-CSE-ConnectionGUID: twlC2qkvS7CeVw4fkj5Ksw== X-CSE-MsgGUID: bQa3l3LNTVy7qIDuPDoQrA== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.27,115,1787036400"; d="scan'208";a="277671820" Received: from b04f130c83f2.jf.intel.com ([10.165.154.98]) by fmviesa004.fm.intel.com with ESMTP; 21 Sep 2026 17:32:10 -0700 From: Tim Chen To: Peter Zijlstra , Ingo Molnar Cc: Tim Chen , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , K Prateek Nayak , Kees Cook , Christian Brauner , Alexander Viro , Jan Kara , Shrikanth Hegde , Qais Yousef , Aaron Lu , Srikar Dronamraju , Vineeth Remanan Pillai , Ricardo Neri-Calderon , Chen Yu , Lu Wang , Hyunwoo Kim , Zhan Xusheng , Zhan Xusheng , Yi Lai , "Rafael J . Wysocki" , Greg Kroah-Hartman , Danilo Krummrich , Zenghui Yu , linux-kernel@vger.kernel.org, linux-mm@kvack.org, linux-fsdevel@vger.kernel.org, Kayra Cizmeci , stable@kernel.org Subject: [PATCH v2 1/6] sched/cache: Keep nr_pref_llc_running in the runnable domain Date: Mon, 21 Sep 2026 17:37:22 -0700 Message-Id: <06af61afedac32e6477f57feb4d658f6c411c3af.1790035273.git.tim.c.chen@linux.intel.com> X-Mailer: git-send-email 2.32.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" alb_break_llc() decides whether to break LLC preference during active load balance. It does so by testing that every runnable fair task on the source rq prefers its LLC: env->src_rq->nr_pref_llc_running =3D=3D env->src_rq->cfs.h_nr_runnable But the two counters cover different sets. nr_pref_llc_running is updated in account_llc_enqueue()/account_llc_dequeue(), next to cfs_rq->nr_queued, so it follows queued tasks. h_nr_runnable is updated in set_delayed()/ clear_delayed() and drops delay-dequeued tasks. So under DELAY_DEQUEUE, a preferring task that goes to sleep stays counted in nr_pref_llc_running while h_nr_runnable falls. The equality then breaks, alb_break_llc() returns false, and active balance is free to pull a task off its preferred LLC. Active balance only moves runnable tasks, and this is the only LLC check it consults: once the stopper runs, LBF_ACTIVE_LB skips the per-task test in can_migrate_task(). The runnable set is the one we want. Fix it on the counter side. A task should be counted in nr_pref_llc_running exactly while it is both queued on its preferred LLC (pref_llc_queued) and runnable (!sched_delayed). Define that membership once in task_pref_llc_runnable(), and adjust the counter only through pref_llc_running_inc()/pref_llc_running_dec() from the four sites that change either input: account_llc_enqueue(), account_llc_dequeue(), set_delayed() and clear_delayed(). Gating every update on the same predicate keeps the delay, wake and dequeue paths from double-counting or underflowing; see the comments at those sites for the ordering. nr_llc_running and sd->llc_counts are not touched and stay on queued semantics. Fixes: 714059f79ff0 ("sched/cache: Handle moving single tasks to/from their= preferred LLC") Reported-by: Zhan Xusheng Closes: https://lore.kernel.org/lkml/20260827135000.735138-1-zhanxusheng@xi= aomi.com/ Suggested-by: Chen Yu Reviewed-by: Kayra Cizmeci Cc: stable@kernel.org #7.2.x Signed-off-by: Tim Chen --- kernel/sched/fair.c | 63 +++++++++++++++++++++++++++++++++++++++++---- 1 file changed, 58 insertions(+), 5 deletions(-) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index d5989b53adef..19765ee1af83 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -1551,6 +1551,28 @@ static bool invalid_llc_nr(struct mm_struct *mm, str= uct task_struct *p, (scale * per_cpu(sd_llc_size, cpu))); } =20 +/* + * A task counts in nr_pref_llc_running while it is queued on its preferred + * LLC (pref_llc_queued) and runnable (!sched_delayed), keeping the counte= r in + * the runnable domain so alb_break_llc() can compare it with h_nr_runnabl= e. + */ +static bool task_pref_llc_runnable(struct task_struct *p) +{ + return p->pref_llc_queued && !p->se.sched_delayed; +} + +static void pref_llc_running_inc(struct rq *rq, struct task_struct *p) +{ + if (task_pref_llc_runnable(p)) + rq->nr_pref_llc_running++; +} + +static void pref_llc_running_dec(struct rq *rq, struct task_struct *p) +{ + if (task_pref_llc_runnable(p)) + rq->nr_pref_llc_running--; +} + static void account_llc_enqueue(struct rq *rq, struct task_struct *p) { int pref_llc, pref_llc_queued; @@ -1562,7 +1584,6 @@ static void account_llc_enqueue(struct rq *rq, struct= task_struct *p) =20 pref_llc_queued =3D (pref_llc =3D=3D task_llc(p)); rq->nr_llc_running++; - rq->nr_pref_llc_running +=3D pref_llc_queued; =20 /* * Record whether p is enqueued on its preferred @@ -1580,6 +1601,9 @@ static void account_llc_enqueue(struct rq *rq, struct= task_struct *p) */ p->pref_llc_queued =3D pref_llc_queued; =20 + /* Skipped while delayed; clear_delayed() adds it back on wake. */ + pref_llc_running_inc(rq, p); + sd =3D rcu_dereference_all(rq->sd); if (sd && (unsigned int)pref_llc < sd->llc_max) sd->llc_counts[pref_llc]++; @@ -1596,7 +1620,12 @@ static void account_llc_dequeue(struct rq *rq, struc= t task_struct *p) =20 rq->nr_llc_running--; if (p->pref_llc_queued) { - rq->nr_pref_llc_running--; + /* + * Skipped if still delayed (set_delayed() already removed it); + * clearing pref_llc_queued below also stops clear_delayed() + * from re-adding it. + */ + pref_llc_running_dec(rq, p); /* * Update the status in case * other logic might query @@ -2000,6 +2029,7 @@ void init_sched_mm(struct task_struct *p) * polluting account_llc_enqueue(). */ p->preferred_llc =3D -1; + p->pref_llc_queued =3D 0; } =20 #else /* CONFIG_SCHED_CACHE */ @@ -2021,6 +2051,10 @@ static void account_llc_enqueue(struct rq *rq, struc= t task_struct *p) {} =20 static void account_llc_dequeue(struct rq *rq, struct task_struct *p) {} =20 +static void pref_llc_running_inc(struct rq *rq, struct task_struct *p) {} + +static void pref_llc_running_dec(struct rq *rq, struct task_struct *p) {} + #endif /* CONFIG_SCHED_CACHE */ =20 /* @@ -6395,15 +6429,27 @@ static __always_inline void return_cfs_rq_runtime(s= truct cfs_rq *cfs_rq); =20 static void set_delayed(struct sched_entity *se) { - se->sched_delayed =3D 1; - /* * Delayed se of cfs_rq have no tasks queued on them. * Do not adjust h_nr_runnable since __dequeue_task() * will account it for blocked tasks. + * + * This check can be removed because when flat pick + * patches get merged as only task can get delayed, + * same for clear_delayed(). */ - if (!entity_is_task(se)) + if (!entity_is_task(se)) { + se->sched_delayed =3D 1; return; + } + + /* + * Drop a task leaving the runnable set. + * Needs to be called before sched_delayed is set. + * clear_delayed() mirrors this after clearing the flag. + */ + pref_llc_running_dec(rq_of(cfs_rq_of(se)), task_of(se)); + se->sched_delayed =3D 1; =20 for_each_sched_entity(se) { struct cfs_rq *cfs_rq =3D cfs_rq_of(se); @@ -6425,6 +6471,13 @@ static void clear_delayed(struct sched_entity *se) if (!entity_is_task(se)) return; =20 + /* + * Re-add on wake, after sched_delayed is cleared. On a final delayed + * dequeue account_llc_dequeue() already cleared pref_llc_queued, so + * this does nothing. + */ + pref_llc_running_inc(rq_of(cfs_rq_of(se)), task_of(se)); + for_each_sched_entity(se) { struct cfs_rq *cfs_rq =3D cfs_rq_of(se); =20 --=20 2.32.0 From nobody Thu Sep 24 17:53:07 2026 Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.15]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A193931E824; Tue, 22 Sep 2026 00:32:13 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=192.198.163.15 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790037135; cv=none; b=uD2x7PTda8UbC2DEc2iVTtz8mLF4pH0P6Tv7MdrgGsJe45VpTAF2MekTwxwWJM5F6YlbVVd8dskAV4Xlj33m2ti/++PTaM+7tdhK62BgU2/82MuoPNXkM5f/nX/0Smc4m2KUu8pSw4UOd6BH/tBxkT8L7lIahMzCS3P5PL5RPXo= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790037135; c=relaxed/simple; bh=3SNad8cSGX/yLWI8E3JE9cy9dkKCeA1Qu3xf/C4B2+Q=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=MvySleftWDs8uPNxhl+iAR3lukcKRVk2KuTe6vB7BM67/tTJPcpXgU2xiY1eK827oGuu3t06s2hcIpYYTNjnXeaw2u+TqgmxPfGZBg2wilYUcNykHBiMl1ZZxMtWvFdvheBXThd6tI57ll6EWkv0SK2IoYOIBj0qJZSzoOWM80E= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com; spf=pass smtp.mailfrom=linux.intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=mp8cMHpM; arc=none smtp.client-ip=192.198.163.15 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="mp8cMHpM" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1790037134; x=1821573134; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=3SNad8cSGX/yLWI8E3JE9cy9dkKCeA1Qu3xf/C4B2+Q=; b=mp8cMHpMPkuFeW8D2S78kLhkxTVDMdV7YGOM0IUbizQmBh8yp17gsqxG xQ4bCO7fCs7Cbhp8taQxTVM78s+vLgPP+lDrzID9sRvqQR2FXehZs49D6 vmtyxTduULowF4YMOrHBIrViKtW3PUW2SmpCvUpjMaMOhrW5jCPuvnYl/ 8G6gzqWpJrWD2O2KJ3O06UugNx8DTX56ZaaDRhZPcz3jkgdBfH1s3K1Bh KPJh5ep2lmWUwS9oQ7PNlhHA6m+WiIqp9819w1AVt6+H11Q/cw3FxQF00 ViYFlKFU2yWYYnvFLukk9Gp8uk13jXycE5Oi79AMbkkvtCsA09vXat0F5 w==; X-CSE-ConnectionGUID: QBPprYNYS8ichcCwp4eANA== X-CSE-MsgGUID: Vb/KKmTWSLOqJhsE6zmTdQ== X-IronPort-AV: E=McAfee;i="6800,10657,11912"; a="90716655" X-IronPort-AV: E=Sophos;i="6.27,115,1787036400"; d="scan'208";a="90716655" Received: from fmviesa004.fm.intel.com ([10.60.135.144]) by fmvoesa109.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 21 Sep 2026 17:32:13 -0700 X-CSE-ConnectionGUID: 4p7YC+7bQ++tsxep43jVwA== X-CSE-MsgGUID: xD69XWy0QB2tetf9x/f/Wg== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.27,115,1787036400"; d="scan'208";a="277671840" Received: from b04f130c83f2.jf.intel.com ([10.165.154.98]) by fmviesa004.fm.intel.com with ESMTP; 21 Sep 2026 17:32:12 -0700 From: Tim Chen To: Peter Zijlstra , Ingo Molnar Cc: Lu Wang , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , K Prateek Nayak , Kees Cook , Christian Brauner , Alexander Viro , Jan Kara , Shrikanth Hegde , Qais Yousef , Aaron Lu , Srikar Dronamraju , Vineeth Remanan Pillai , Ricardo Neri-Calderon , Chen Yu , Hyunwoo Kim , Zhan Xusheng , Zhan Xusheng , Yi Lai , Tim Chen , "Rafael J . Wysocki" , Greg Kroah-Hartman , Danilo Krummrich , Zenghui Yu , linux-kernel@vger.kernel.org, linux-mm@kvack.org, linux-fsdevel@vger.kernel.org, stable@kernel.org Subject: [PATCH v2 2/6] sched/cache: Honor migrate_llc_task semantics in active load balance Date: Mon, 21 Sep 2026 17:37:23 -0700 Message-Id: X-Mailer: git-send-email 2.32.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Lu Wang CAS introduced the migrate_llc_task migration type to direct tasks toward their preferred LLC, but its semantics can be lost when passive load balance falls back to active load balance. This may allow ALB to select a candidate whose preferred LLC does not match the destination, moving it away from its preferred LLC. Example scenario: src_rq has two runnable tasks, p1 and p2. p1 prefers dst_rq (dst_llc), while p2 prefers src_rq (src_llc). In this case, migrate_llc_task is set because src_rq has at least one task, p1, that wants to migrate to dst_rq. In ALB, can_migrate_task() finds p2 and returns true for it, thus moving p2 out of its preferred LLC. Solution: The CPU stopper in ALB constructs a fresh lb_env that does not inherit migration_type from the passive load-balance pass. Two approaches are possible: (a) Add a new member to struct rq so ALB can inherit migrate_llc_task from the passive LB that triggered it. (b) Define a new flag LBF_ACTIVE_LB_LLC and select the stopper callback at kick time to preserve the migration semantics across the asynchronous boundary. We choose (b) because it avoids passing migration_type through the stopper, which would affect the meaning of migration_type for delayed-dequeue tasks. Fixes: e4c9a4cb244a ("sched/cache: Add migrate_llc_task migration type for = cache-aware balancing") Suggested-by: Chen Yu Signed-off-by: Lu Wang Reviewed-by: Tim Chen Reviewed-by: Chen Yu Cc: stable@kernel.org #7.2.x Signed-off-by: Tim Chen --- kernel/sched/fair.c | 57 ++++++++++++++++++++++++++++++++++++++++----- 1 file changed, 51 insertions(+), 6 deletions(-) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 19765ee1af83..6f1939d17e9e 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -10453,6 +10453,7 @@ enum migration_type { #define LBF_SOME_PINNED 0x08 #define LBF_ACTIVE_LB 0x10 #define LBF_LLC_PINNED 0x20 +#define LBF_ACTIVE_LB_LLC 0x40 =20 struct lb_env { struct sched_domain *sd; @@ -10871,6 +10872,21 @@ alb_break_llc(struct lb_env *env) return false; } =20 +/* + * Returns true if p's preferred LLC does not match the destination CPU + * under migrate_llc_task semantics. Passive LB passes migrate_llc_task + * in env->migration_type, while active LB carries LBF_ACTIVE_LB_LLC in + * env->flags to avoid overwriting env->migration_type. + */ +static inline bool +migrate_llc_task_wrong_dst(struct task_struct *p, struct lb_env *env) +{ + return sched_cache_enabled() && + (env->migration_type =3D=3D migrate_llc_task || + env->flags & LBF_ACTIVE_LB_LLC) && + READ_ONCE(p->preferred_llc) !=3D llc_id(env->dst_cpu); +} + /* * Check if migrating task p from env->src_cpu to * env->dst_cpu breaks LLC localiy. @@ -10899,8 +10915,7 @@ static bool migrate_degrades_llc(struct task_struct= *p, struct lb_env *env) * run on env->dst_cpu, skip the tasks do not prefer * env->dst_cpu, and find the one that prefers. */ - if (env->migration_type =3D=3D migrate_llc_task && - READ_ONCE(p->preferred_llc) !=3D llc_id(env->dst_cpu)) + if (migrate_llc_task_wrong_dst(p, env)) return true; =20 if (can_migrate_llc_task(env, p) !=3D mig_forbid) @@ -10922,6 +10937,12 @@ alb_break_llc(struct lb_env *env) return false; } =20 +static inline bool +migrate_llc_task_wrong_dst(struct task_struct *p, struct lb_env *env) +{ + return false; +} + static inline bool migrate_degrades_llc(struct task_struct *p, struct lb_env *env) { @@ -11021,7 +11042,7 @@ int can_migrate_task(struct task_struct *p, struct = lb_env *env) * 4) too many balance attempts have failed. */ if (env->flags & LBF_ACTIVE_LB) - return 1; + return !migrate_llc_task_wrong_dst(p, env); =20 degrades =3D migrate_degrades_locality(p, env); if (!degrades) { @@ -13420,6 +13441,20 @@ static int need_active_balance(struct lb_env *env) } =20 static int active_load_balance_cpu_stop(void *data); +static int active_load_balance_llc_cpu_stop(void *data); + +/* + * migration_type is checked elsewhere to decide migration policy, so + * it shouldn't be repurposed just to flag an LLC-directed active + * balance across the stopper. Pick the callback here instead. + */ +static inline cpu_stop_fn_t alb_stop_fn(struct lb_env *env) +{ + if (env->migration_type =3D=3D migrate_llc_task) + return active_load_balance_llc_cpu_stop; + + return active_load_balance_cpu_stop; +} =20 static int should_we_balance(struct lb_env *env) { @@ -13765,7 +13800,7 @@ static int sched_balance_rq(int this_cpu, struct rq= *this_rq, } if (active_balance) { stop_one_cpu_nowait(cpu_of(busiest), - active_load_balance_cpu_stop, busiest, + alb_stop_fn(&env), busiest, &busiest->active_balance_work); } preempt_enable(); @@ -13870,7 +13905,7 @@ update_next_balance(struct sched_domain *sd, unsign= ed long *next_balance) * least 1 task to be running on each physical CPU where possible, and * avoids physical / logical imbalances. */ -static int active_load_balance_cpu_stop(void *data) +static int __active_load_balance_cpu_stop(void *data, unsigned int lb_flag= s) { struct rq *busiest_rq =3D data; int busiest_cpu =3D cpu_of(busiest_rq); @@ -13920,7 +13955,7 @@ static int active_load_balance_cpu_stop(void *data) .src_cpu =3D busiest_rq->cpu, .src_rq =3D busiest_rq, .idle =3D CPU_IDLE, - .flags =3D LBF_ACTIVE_LB, + .flags =3D LBF_ACTIVE_LB | lb_flags, }; =20 schedstat_inc(sd->alb_count); @@ -13948,6 +13983,16 @@ static int active_load_balance_cpu_stop(void *data) return 0; } =20 +static int active_load_balance_cpu_stop(void *data) +{ + return __active_load_balance_cpu_stop(data, 0); +} + +static int active_load_balance_llc_cpu_stop(void *data) +{ + return __active_load_balance_cpu_stop(data, LBF_ACTIVE_LB_LLC); +} + /* * Scale the max sched_balance_rq interval with the number of CPUs in the = system. * This trades load-balance latency on larger machines for less cross talk. --=20 2.32.0 From nobody Thu Sep 24 17:53:07 2026 Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.15]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0B03A334C3B; Tue, 22 Sep 2026 00:32:15 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=192.198.163.15 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790037138; cv=none; b=VfWJv4uRYLu8EKTzmbX3CdP8F4Xd6zxFPWz2zJmca+5zdnFsHabyuMliXTy96cJFnYVpRauzLLhwhLIyBpSzacKf6qazPIJiA2fYTXE3D68FAu0GPjzGyHLTWnJeo7StYeYXtcLS9F/EBMppt//Bi+liOlaNH+ppN6ZJ5OMn4ZA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790037138; c=relaxed/simple; bh=NiligWz0gyX4ezHC8gHM6+vA0T6/3jJREB3L+QxjAf8=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=Xe2lcarmq2UN6rQPYJg9QjiYF66Cz4XYUyPsQUOjl5u+orSRdiBOR/aewl8fBumFnbgB+ZR//RTYFAOW1G7+p12EV8+CrCkTUawN4tHlO8UFbI04jLH74rY6th0y6IWZo0SLqsfjefVRZcOgKPib5HDsdtfAMusOBqjDrCbns0o= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com; spf=pass smtp.mailfrom=linux.intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=dpWcsFa1; arc=none smtp.client-ip=192.198.163.15 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="dpWcsFa1" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1790037136; x=1821573136; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=NiligWz0gyX4ezHC8gHM6+vA0T6/3jJREB3L+QxjAf8=; b=dpWcsFa1Kmiw88a7SiWxSWJVEKLvpspZaxKyAHEPSFEkMLbFkAF/QtAC apFDx50vmiTO7j36aeYjCtv6UDEm+28aquIHBgDBfxPYZUm5KxpZeKMmF gXUOmo8JYtREpQaa9/B/xF7GZuBk6Sy6DczXs8IoLTwPUiSSicMjUlMAA SCkb71cIaVzi0QUYxG6BkkX/ymybHY1vbZtU5hmk5n2G6eh0w8Dc3GxYh gyx702TeWfViYW4RviM0oKR+RAjxMUUVGn4yosiQqcTmDbvz6ZcY6wKnJ 8jBrYaQaAn136dV76UiswMJB0As8NygjwYx9pEsVmQes514zwFYmDlvSE A==; X-CSE-ConnectionGUID: pZOPAGR+RXueqVGQy5HyHg== X-CSE-MsgGUID: KoyYb0klQvqrTZG0dkFckg== X-IronPort-AV: E=McAfee;i="6800,10657,11912"; a="90716676" X-IronPort-AV: E=Sophos;i="6.27,115,1787036400"; d="scan'208";a="90716676" Received: from fmviesa004.fm.intel.com ([10.60.135.144]) by fmvoesa109.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 21 Sep 2026 17:32:15 -0700 X-CSE-ConnectionGUID: eZ2TKfSZTymQKBVejoNhvw== X-CSE-MsgGUID: h98emjb1Rre/qLRRl6AFIQ== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.27,115,1787036400"; d="scan'208";a="277671856" Received: from b04f130c83f2.jf.intel.com ([10.165.154.98]) by fmviesa004.fm.intel.com with ESMTP; 21 Sep 2026 17:32:14 -0700 From: Tim Chen To: Peter Zijlstra , Ingo Molnar Cc: Tim Chen , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , K Prateek Nayak , Kees Cook , Christian Brauner , Alexander Viro , Jan Kara , Shrikanth Hegde , Qais Yousef , Aaron Lu , Srikar Dronamraju , Vineeth Remanan Pillai , Ricardo Neri-Calderon , Chen Yu , Lu Wang , Hyunwoo Kim , Zhan Xusheng , Zhan Xusheng , Yi Lai , "Rafael J . Wysocki" , Greg Kroah-Hartman , Danilo Krummrich , Zenghui Yu , linux-kernel@vger.kernel.org, linux-mm@kvack.org, linux-fsdevel@vger.kernel.org, stable@kernel.org Subject: [PATCH v2 3/6] sched/cache: Decouple sched_cache_group from mm Date: Mon, 21 Sep 2026 17:37:24 -0700 Message-Id: <91fd1e3266707c865bc9abecfb3e17bc676712df.1790035273.git.tim.c.chen@linux.intel.com> X-Mailer: git-send-email 2.32.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Currently the sched cache grouping is by mm and the scheduling statistics sched_cache_stat lives in the mm structure. This ties the life cycle of scheduling stats with mm. In account_mm_sched(), the scheduling stats are accessed by task->mm->sc_stat. However, a task may be switching mm on one CPU when another CPU is running account_mm_sched(), and possibly accessing the old mm that was freed. This problem was found when running tests with KASAN by Hyunwoo. https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/ Instead of serializing the mm access by introducing extra acquisition of rq lock in the mm free path, extract sched_cache_stat from mm_struct, rename it as sched_cache_group and manage its life cycle apart from mm_struct with its own ref counting. This allows us in the next patch access sched_cache_group directly from task, and add a refcount on sched_cache_group when a task links to it. This prevents the use after free issue when accessing stale and released old mm and its sched cache stat a task switches to a new mm while account_mm_sched() is done elsewhere. The other benefit of this restructure is in the future, the grouping of tasks to a LLC would have the flexibility to be associated with a user defined grouping, or cgroup, cookie group, numa_group or others instead of just with a single mm address space. Rename sched_cache_stat to sched_cache_group and turn it into a refcounted object allocated from mm_struct. The mm_struct now holds a pointer (sched_cache_grp) to this object instead of embedding it. Introduce kernel/sched/cache_sched.c to host the cache aware scheduling helpers and define sched_cache_group_put() there. Fixes: df0d98475954 ("sched/cache: Introduce infrastructure for cache-aware= load balancing") Reported-by: Hyunwoo Kim Closes: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/ Reported-by: Zenghui Yu (Huawei) Closes: https://lore.kernel.org/all/343a7e07-7fad-4979-9c9b-82ec038c293c@li= nux.dev/ Cc: stable@kernel.org #7.2.x Co-developed-by: Chen Yu Signed-off-by: Chen Yu Signed-off-by: Tim Chen --- include/linux/mm_types.h | 15 ++-- include/linux/sched.h | 8 +- kernel/exit.c | 11 ++- kernel/sched/build_utility.c | 4 + kernel/sched/cache_sched.c | 19 +++++ kernel/sched/fair.c | 156 ++++++++++++++++++++++++----------- 6 files changed, 152 insertions(+), 61 deletions(-) create mode 100644 kernel/sched/cache_sched.c diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h index 6d815f6440c9..f3e5a2fadbe5 100644 --- a/include/linux/mm_types.h +++ b/include/linux/mm_types.h @@ -1226,7 +1226,7 @@ struct mm_struct { struct mm_mm_cid mm_cid; =20 /* sched_cache related statistics */ - struct sched_cache_stat sc_stat; + struct sched_cache_group *sched_cache_grp; #ifdef CONFIG_MMU atomic_long_t pgtables_bytes; /* size of all page tables */ #endif @@ -1624,8 +1624,9 @@ static inline unsigned int mm_cid_size(void) #endif /* CONFIG_SCHED_MM_CID */ =20 #ifdef CONFIG_SCHED_CACHE -void mm_init_sched(struct mm_struct *mm, - struct sched_cache_time __percpu *pcpu_sched); +int mm_init_sched(struct mm_struct *mm, + struct sched_cache_time __percpu *pcpu_sched); +void mm_destroy_sched(struct mm_struct *mm); =20 static inline int mm_alloc_sched_noprof(struct mm_struct *mm) { @@ -1635,17 +1636,11 @@ static inline int mm_alloc_sched_noprof(struct mm_s= truct *mm) if (!pcpu_sched) return -ENOMEM; =20 - mm_init_sched(mm, pcpu_sched); - return 0; + return mm_init_sched(mm, pcpu_sched); } =20 #define mm_alloc_sched(...) alloc_hooks(mm_alloc_sched_noprof(__VA_ARGS__)) =20 -static inline void mm_destroy_sched(struct mm_struct *mm) -{ - free_percpu(mm->sc_stat.pcpu_sched); - mm->sc_stat.pcpu_sched =3D NULL; -} #else /* !CONFIG_SCHED_CACHE */ =20 static inline int mm_alloc_sched(struct mm_struct *mm) { return 0; } diff --git a/include/linux/sched.h b/include/linux/sched.h index 8b3d47a325cc..1f254364f216 100644 --- a/include/linux/sched.h +++ b/include/linux/sched.h @@ -2405,7 +2405,7 @@ struct sched_cache_time { unsigned long epoch; }; =20 -struct sched_cache_stat { +struct sched_cache_group { struct sched_cache_time __percpu *pcpu_sched; raw_spinlock_t lock; unsigned long epoch; @@ -2413,11 +2413,15 @@ struct sched_cache_stat { unsigned long next_scan; unsigned long footprint; int cpu; + refcount_t refcnt; + struct rcu_head rcu; } ____cacheline_aligned_in_smp; =20 +void sched_cache_group_put(struct sched_cache_group *grp); + #else =20 -struct sched_cache_stat { }; +struct sched_cache_group { }; =20 #endif =20 diff --git a/kernel/exit.c b/kernel/exit.c index 97686af89501..16abbe1cf682 100644 --- a/kernel/exit.c +++ b/kernel/exit.c @@ -554,18 +554,23 @@ void mm_update_next_owner(struct mm_struct *mm) */ static void exit_mm_sched_cache(struct mm_struct *mm) { + struct sched_cache_group *grp; unsigned long fp, sub; =20 if (!current->total_numa_faults) return; /* * No lock protection due to performance considerations. - * Make sure mm->sc_stat.footprint does not become + * Make sure the group footprint does not become * negative. */ - fp =3D READ_ONCE(mm->sc_stat.footprint); + grp =3D READ_ONCE(mm->sched_cache_grp); + if (!grp) + return; + + fp =3D READ_ONCE(grp->footprint); sub =3D min(fp, current->total_numa_faults); - WRITE_ONCE(mm->sc_stat.footprint, fp - sub); + WRITE_ONCE(grp->footprint, fp - sub); } #else static inline void exit_mm_sched_cache(struct mm_struct *mm) diff --git a/kernel/sched/build_utility.c b/kernel/sched/build_utility.c index e2cf3b08d4e9..24202893b262 100644 --- a/kernel/sched/build_utility.c +++ b/kernel/sched/build_utility.c @@ -89,6 +89,10 @@ # include "core_sched.c" #endif =20 +#ifdef CONFIG_SCHED_CACHE +# include "cache_sched.c" +#endif + #ifdef CONFIG_PSI # include "psi.c" #endif diff --git a/kernel/sched/cache_sched.c b/kernel/sched/cache_sched.c new file mode 100644 index 000000000000..ff3d9538e7a4 --- /dev/null +++ b/kernel/sched/cache_sched.c @@ -0,0 +1,19 @@ +// SPDX-License-Identifier: GPL-2.0-only +#include "sched.h" + +static void sched_cache_group_free_rcu(struct rcu_head *rcu) +{ + struct sched_cache_group *grp =3D + container_of(rcu, struct sched_cache_group, rcu); + + free_percpu(grp->pcpu_sched); + kfree(grp); +} + +void sched_cache_group_put(struct sched_cache_group *grp) +{ + if (!grp || !refcount_dec_and_test(&grp->refcnt)) + return; + + call_rcu(&grp->rcu, sched_cache_group_free_rcu); +} diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 6f1939d17e9e..6e939807dff2 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -1497,12 +1497,17 @@ static bool exceed_llc_capacity(struct mm_struct *m= m, int cpu) return true; =20 if (static_branch_likely(&sched_numa_balancing)) { + struct sched_cache_group *grp =3D READ_ONCE(mm->sched_cache_grp); + + if (!grp) + return true; + /* * TBD: RDT exclusive LLC ways reserved should be * excluded. */ llc =3D sd->llc_bytes; - footprint =3D READ_ONCE(mm->sc_stat.footprint); + footprint =3D READ_ONCE(grp->footprint); =20 /* * Scale the LLC size by 256*llc_aggr_tolerance @@ -1534,6 +1539,7 @@ static bool exceed_llc_capacity(struct mm_struct *mm,= int cpu) static bool invalid_llc_nr(struct mm_struct *mm, struct task_struct *p, int cpu) { + struct sched_cache_group *grp; int scale; =20 if (get_nr_threads(p) <=3D 1) @@ -1547,7 +1553,11 @@ static bool invalid_llc_nr(struct mm_struct *mm, str= uct task_struct *p, if (scale =3D=3D INT_MAX) return false; =20 - return !fits_capacity((mm->sc_stat.nr_running_avg * cpu_smt_num_threads), + grp =3D READ_ONCE(mm->sched_cache_grp); + if (!grp) + return true; + + return !fits_capacity((READ_ONCE(grp->nr_running_avg) * cpu_smt_num_threa= ds), (scale * per_cpu(sd_llc_size, cpu))); } =20 @@ -1653,12 +1663,20 @@ static void account_llc_dequeue(struct rq *rq, stru= ct task_struct *p) } } =20 -void mm_init_sched(struct mm_struct *mm, - struct sched_cache_time __percpu *_pcpu_sched) +int mm_init_sched(struct mm_struct *mm, + struct sched_cache_time __percpu *_pcpu_sched) { + struct sched_cache_group *grp; unsigned long epoch =3D 0; int i; =20 + grp =3D kzalloc_obj(*grp); + if (!grp) { + free_percpu(_pcpu_sched); + mm->sched_cache_grp =3D NULL; + return -ENOMEM; + } + for_each_possible_cpu(i) { struct sched_cache_time *pcpu_sched =3D per_cpu_ptr(_pcpu_sched, i); struct rq *rq =3D cpu_rq(i); @@ -1669,18 +1687,34 @@ void mm_init_sched(struct mm_struct *mm, epoch =3D rq->cpu_epoch; } =20 - raw_spin_lock_init(&mm->sc_stat.lock); - mm->sc_stat.epoch =3D epoch; - mm->sc_stat.cpu =3D -1; - mm->sc_stat.next_scan =3D jiffies; - mm->sc_stat.nr_running_avg =3D 0; - mm->sc_stat.footprint =3D 0; + raw_spin_lock_init(&grp->lock); + grp->epoch =3D epoch; + grp->cpu =3D -1; + grp->next_scan =3D jiffies; + grp->nr_running_avg =3D 0; + grp->footprint =3D 0; + refcount_set(&grp->refcnt, 1); /* - * The update to mm->sc_stat should not be reordered - * before initialization to mm's other fields, in case + * The update to grp->pcpu_sched should not be reordered + * before initialization to grp's other fields, in case * the readers may get invalid mm_sched_epoch, etc. */ - smp_store_release(&mm->sc_stat.pcpu_sched, _pcpu_sched); + smp_store_release(&grp->pcpu_sched, _pcpu_sched); + /* + * Publish the group last. Not every reader qualifies it by + * grp->pcpu_sched - can_migrate_llc_task() only checks that the + * pointer is non-NULL before reading grp->footprint and + * grp->nr_running_avg - so a reachable group must already be + * fully initialized. + */ + smp_store_release(&mm->sched_cache_grp, grp); + return 0; +} + +void mm_destroy_sched(struct mm_struct *mm) +{ + sched_cache_group_put(mm->sched_cache_grp); + mm->sched_cache_grp =3D NULL; } =20 /* because why would C be fully specified */ @@ -1734,11 +1768,16 @@ static unsigned long fraction_mm_sched(struct rq *r= q, static int get_pref_llc(struct task_struct *p, struct mm_struct *mm) { int mm_sched_llc =3D -1, mm_sched_cpu; + struct sched_cache_group *grp; =20 if (!mm) return -1; =20 - mm_sched_cpu =3D READ_ONCE(mm->sc_stat.cpu); + grp =3D READ_ONCE(mm->sched_cache_grp); + if (!grp) + return -1; + + mm_sched_cpu =3D READ_ONCE(grp->cpu); if (mm_sched_cpu !=3D -1) { mm_sched_llc =3D llc_id(mm_sched_cpu); =20 @@ -1769,6 +1808,7 @@ static inline void account_mm_sched(struct rq *rq, struct task_struct *p, s64 delta_exec) { struct sched_cache_time *pcpu_sched; + struct sched_cache_group *grp; struct mm_struct *mm =3D p->mm; int mm_sched_llc =3D -1; unsigned long epoch; @@ -1782,10 +1822,14 @@ void account_mm_sched(struct rq *rq, struct task_st= ruct *p, s64 delta_exec) * init_task, kthreads and user thread created * by user_mode_thread() don't have mm. */ - if (!mm || !mm->sc_stat.pcpu_sched) + if (!mm) + return; + + grp =3D READ_ONCE(mm->sched_cache_grp); + if (!grp || !grp->pcpu_sched) return; =20 - pcpu_sched =3D per_cpu_ptr(mm->sc_stat.pcpu_sched, cpu_of(rq)); + pcpu_sched =3D per_cpu_ptr(grp->pcpu_sched, cpu_of(rq)); =20 scoped_guard (raw_spinlock, &rq->cpu_epoch_lock) { __update_mm_sched(rq, pcpu_sched); @@ -1798,11 +1842,11 @@ void account_mm_sched(struct rq *rq, struct task_st= ruct *p, s64 delta_exec) * If this process hasn't hit task_cache_work() for a while invalidate * its preferred state. */ - if ((long)(epoch - READ_ONCE(mm->sc_stat.epoch)) > llc_epoch_affinity_tim= eout || + if ((long)(epoch - READ_ONCE(grp->epoch)) > llc_epoch_affinity_timeout || invalid_llc_nr(mm, p, cpu_of(rq)) || exceed_llc_capacity(mm, cpu_of(rq))) { - if (READ_ONCE(mm->sc_stat.cpu) !=3D -1) - WRITE_ONCE(mm->sc_stat.cpu, -1); + if (READ_ONCE(grp->cpu) !=3D -1) + WRITE_ONCE(grp->cpu, -1); } =20 mm_sched_llc =3D get_pref_llc(p, mm); @@ -1819,30 +1863,35 @@ void account_mm_sched(struct rq *rq, struct task_st= ruct *p, s64 delta_exec) static void task_tick_cache(struct rq *rq, struct task_struct *p) { struct callback_head *work =3D &p->cache_work; + struct sched_cache_group *grp; struct mm_struct *mm =3D p->mm; unsigned long epoch; =20 if (!sched_cache_enabled()) return; =20 - if (!mm || p->flags & PF_KTHREAD || - !mm->sc_stat.pcpu_sched) + if (!mm || p->flags & PF_KTHREAD) + return; + + grp =3D READ_ONCE(mm->sched_cache_grp); + if (!grp || !grp->pcpu_sched) return; =20 epoch =3D rq->cpu_epoch; /* avoid moving backwards */ - if (time_after_eq(mm->sc_stat.epoch, epoch)) + if (time_after_eq(grp->epoch, epoch)) return; =20 - guard(raw_spinlock)(&mm->sc_stat.lock); + guard(raw_spinlock)(&grp->lock); =20 if (work->next =3D=3D work) { task_work_add(p, work, TWA_RESUME); - WRITE_ONCE(mm->sc_stat.epoch, epoch); + WRITE_ONCE(grp->epoch, epoch); } } =20 -static void get_scan_cpumasks(cpumask_var_t cpus, struct task_struct *p) +static void get_scan_cpumasks(cpumask_var_t cpus, struct task_struct *p, + struct sched_cache_group *grp) { #ifdef CONFIG_NUMA_BALANCING int cpu, curr_cpu, nid, pref_nid; @@ -1850,7 +1899,7 @@ static void get_scan_cpumasks(cpumask_var_t cpus, str= uct task_struct *p) if (!static_branch_likely(&sched_numa_balancing)) goto out; =20 - cpu =3D READ_ONCE(p->mm->sc_stat.cpu); + cpu =3D READ_ONCE(grp->cpu); if (cpu !=3D -1) nid =3D cpu_to_node(cpu); curr_cpu =3D task_cpu(p); @@ -1911,6 +1960,7 @@ static void task_cache_work(struct callback_head *wor= k) unsigned long next_scan, now =3D jiffies; struct task_struct *p =3D current, *cur; unsigned long curr_m_a_occ =3D 0; + struct sched_cache_group *grp; struct mm_struct *mm =3D p->mm; unsigned long m_a_occ =3D 0; cpumask_var_t cpus; @@ -1922,12 +1972,16 @@ static void task_cache_work(struct callback_head *w= ork) if (p->flags & PF_EXITING) return; =20 - next_scan =3D READ_ONCE(mm->sc_stat.next_scan); + grp =3D READ_ONCE(mm->sched_cache_grp); + if (!grp) + return; + + next_scan =3D READ_ONCE(grp->next_scan); if (time_before(now, next_scan)) return; =20 /* only 1 thread is allowed to scan */ - if (!try_cmpxchg(&mm->sc_stat.next_scan, &next_scan, + if (!try_cmpxchg(&grp->next_scan, &next_scan, now + max_t(unsigned long, READ_ONCE(llc_epoch_period), 1))) return; @@ -1935,8 +1989,8 @@ static void task_cache_work(struct callback_head *wor= k) curr_cpu =3D task_cpu(p); if (invalid_llc_nr(mm, p, curr_cpu) || exceed_llc_capacity(mm, curr_cpu)) { - if (READ_ONCE(mm->sc_stat.cpu) !=3D -1) - WRITE_ONCE(mm->sc_stat.cpu, -1); + if (READ_ONCE(grp->cpu) !=3D -1) + WRITE_ONCE(grp->cpu, -1); =20 return; } @@ -1947,7 +2001,7 @@ static void task_cache_work(struct callback_head *wor= k) scoped_guard (cpus_read_lock) { guard(rcu)(); =20 - get_scan_cpumasks(cpus, p); + get_scan_cpumasks(cpus, p, grp); =20 for_each_cpu(cpu, cpus) { /* XXX sched_cluster_active */ @@ -1960,7 +2014,7 @@ static void task_cache_work(struct callback_head *wor= k) =20 for_each_cpu(i, sched_domain_span(sd)) { occ =3D fraction_mm_sched(cpu_rq(i), - per_cpu_ptr(mm->sc_stat.pcpu_sched, i)); + per_cpu_ptr(grp->pcpu_sched, i)); a_occ +=3D occ; if (occ > m_occ) { m_occ =3D occ; @@ -1993,7 +2047,7 @@ static void task_cache_work(struct callback_head *wor= k) m_a_cpu =3D m_cpu; } =20 - if (llc_id(cpu) =3D=3D llc_id(READ_ONCE(mm->sc_stat.cpu))) + if (llc_id(cpu) =3D=3D llc_id(READ_ONCE(grp->cpu))) curr_m_a_occ =3D a_occ; =20 cpumask_andnot(cpus, cpus, sched_domain_span(sd)); @@ -2002,7 +2056,7 @@ static void task_cache_work(struct callback_head *wor= k) =20 if (m_a_occ > (2 * curr_m_a_occ)) { /* - * Avoid switching sc_stat.cpu too fast. + * Avoid switching sched_cache_grp->cpu too fast. * The reason to choose 2X is because: * 1. It is better to keep the preferred LLC stable, * rather than changing it frequently and cause migrations @@ -2011,10 +2065,10 @@ static void task_cache_work(struct callback_head *w= ork) * 3. 2X is chosen based on test results, as it delivers * the optimal performance gain so far. */ - WRITE_ONCE(mm->sc_stat.cpu, m_a_cpu); + WRITE_ONCE(grp->cpu, m_a_cpu); } =20 - update_avg_scale(&mm->sc_stat.nr_running_avg, nr_running); + update_avg_scale(&grp->nr_running_avg, nr_running); free_cpumask_var(cpus); } =20 @@ -3731,6 +3785,7 @@ static int preferred_group_nid(struct task_struct *p,= int nid) static void task_numa_placement(struct task_struct *p) __context_unsafe(/* conditional locking */) { + struct sched_cache_group __maybe_unused *grp; int seq, nid, max_nid =3D NUMA_NO_NODE; unsigned long max_faults =3D 0; unsigned long fault_types[2] =3D { 0, 0 }; @@ -3823,19 +3878,23 @@ static void task_numa_placement(struct task_struct = *p) * heuristic and occasional lost updates are tolerable. * * If a task exits, its corresponding footprint must - * be subtracted from the mm->sc_stat.footprint, otherwise - * the mm->sc_stat.footprint will not converge: - * the exiting thread's footprint remains unchanged/undecayed - * in mm->sc_stat.footprint. See exit_mm(). + * be subtracted from the mm->sched_cache_grp->footprint, + * otherwise the mm->sched_cache_grp->footprint will not + * converge: the exiting thread's footprint remains + * unchanged/undecayed in mm->sched_cache_grp->footprint. + * See exit_mm(). * * Lost updates and unsynchronized subtraction * in exit_mm() can cause footprint + diff to * go negative. Clamp to zero to prevent the * unsigned footprint from wrapping. */ - new_fp =3D (long)READ_ONCE(p->mm->sc_stat.footprint) + diff; - WRITE_ONCE(p->mm->sc_stat.footprint, - max(new_fp, 0L)); + grp =3D READ_ONCE(p->mm->sched_cache_grp); + if (!grp) + continue; + + new_fp =3D (long)READ_ONCE(grp->footprint) + diff; + WRITE_ONCE(grp->footprint, max(new_fp, 0L)); #endif } =20 @@ -10783,6 +10842,7 @@ static inline bool task_misfits_asym_cpu(struct lb_= env *env, struct task_struct static enum llc_mig can_migrate_llc_task(struct lb_env *env, struct task_struct *p) { + struct sched_cache_group *grp; struct mm_struct *mm; bool to_pref; int cpu, src_cpu, dst_cpu; @@ -10796,15 +10856,19 @@ static enum llc_mig can_migrate_llc_task(struct l= b_env *env, if (!mm) return mig_unrestricted; =20 - cpu =3D READ_ONCE(mm->sc_stat.cpu); + grp =3D READ_ONCE(mm->sched_cache_grp); + if (!grp) + return mig_unrestricted; + + cpu =3D READ_ONCE(grp->cpu); if (cpu < 0 || cpus_share_cache(src_cpu, dst_cpu)) return mig_unrestricted; =20 /* skip cache aware load balance for too many threads */ if (invalid_llc_nr(mm, p, dst_cpu) || exceed_llc_capacity(mm, dst_cpu)) { - if (READ_ONCE(mm->sc_stat.cpu) !=3D -1) - WRITE_ONCE(mm->sc_stat.cpu, -1); + if (READ_ONCE(grp->cpu) !=3D -1) + WRITE_ONCE(grp->cpu, -1); return mig_unrestricted; } =20 --=20 2.32.0 From nobody Thu Sep 24 17:53:07 2026 Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.15]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8733D23D7F4; Tue, 22 Sep 2026 00:32:17 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=192.198.163.15 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790037140; cv=none; b=Lc/cyg4XXBdX61wRdYhfAQxFxOpYYcrmhJ5ckzXqWaY8Ea4D6LwHReXxGicXmgBxl7sDizrsis5nvy769zdeWKpiEMaZYG2Nkf08EHVwbB87cSRoIHNpXO36IUMAWGamIap9AZeh5WM4TmwLWrQbRC7ap/AYADi+79fUQFIYbEw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790037140; c=relaxed/simple; bh=+1tSKOiwkomk8D3mtv6M2i5ogdNO51o7OpMzTMwjWRQ=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=UEAJyz3IG6BuX97wDl0/tUURHhJy86eORaKZVlWKuC5WAz22GanXBaxEa5hh9MHxcCZH6tlAvzblCZG9yXcR6w8dqbcGLMOqBqDLgO6RJmZK4wFkA016H3MyNqpEWVBnr6ExJGg55PC1inTTRcJ9abnS5j9nUf04JGUXsfCimhM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com; spf=pass smtp.mailfrom=linux.intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=NTijbP2Z; arc=none smtp.client-ip=192.198.163.15 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="NTijbP2Z" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1790037138; x=1821573138; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=+1tSKOiwkomk8D3mtv6M2i5ogdNO51o7OpMzTMwjWRQ=; b=NTijbP2Zm8yD7YFX3REb2JR71kyFlB2U6BGKOXAP8CHLY6CXonaLDeZQ 5DGazQ07C/zOycaHidrC03778HcAff7RWQyX9vnDR+bsel1hjf+slpNa3 JXPVW0fJ0NnAr6dEUjK3HF/71QHc+9vGhwTBlgzbuoaW8oqJyIrDYiIPd 7UqgUHFhdrfSmA35D9/iwXrbw5ynId9POs7zNGcl0bgEQLpJGxso9usYk vQHpsKqjao1GzXkmIuOjoNPvBsF6oRqKYvhEXMJ0y/YS2iYDXjiGHH2RR rw0owsmK8p7Yvjq2qubB9z59QqNqNVf8N8VElAE/Jb19NUHqgun2f9s4s A==; X-CSE-ConnectionGUID: UXbdTV4ySHCWtmIV8vWv7Q== X-CSE-MsgGUID: wrhrJ96YRm2T7J+JUFFQ4A== X-IronPort-AV: E=McAfee;i="6800,10657,11912"; a="90716698" X-IronPort-AV: E=Sophos;i="6.27,115,1787036400"; d="scan'208";a="90716698" Received: from fmviesa004.fm.intel.com ([10.60.135.144]) by fmvoesa109.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 21 Sep 2026 17:32:17 -0700 X-CSE-ConnectionGUID: suUgIRFJSGqtNw+SpAWtNg== X-CSE-MsgGUID: VHXHKVwgRJqRI/9S0w5HsQ== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.27,115,1787036400"; d="scan'208";a="277671883" Received: from b04f130c83f2.jf.intel.com ([10.165.154.98]) by fmviesa004.fm.intel.com with ESMTP; 21 Sep 2026 17:32:15 -0700 From: Tim Chen To: Peter Zijlstra , Ingo Molnar Cc: Tim Chen , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , K Prateek Nayak , Kees Cook , Christian Brauner , Alexander Viro , Jan Kara , Shrikanth Hegde , Qais Yousef , Aaron Lu , Srikar Dronamraju , Vineeth Remanan Pillai , Ricardo Neri-Calderon , Chen Yu , Lu Wang , Hyunwoo Kim , Zhan Xusheng , Zhan Xusheng , Yi Lai , "Rafael J . Wysocki" , Greg Kroah-Hartman , Danilo Krummrich , Zenghui Yu , linux-kernel@vger.kernel.org, linux-mm@kvack.org, linux-fsdevel@vger.kernel.org, stable@kernel.org Subject: [PATCH v2 4/6] sched/cache: Introduce task_struct->sched_cache_grp Date: Mon, 21 Sep 2026 17:37:25 -0700 Message-Id: X-Mailer: git-send-email 2.32.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add a sched_cache_grp pointer to task_struct so that scheduler code can access the cache group directly via the task, without going through mm->sched_cache_grp. This decouples the scheduler's hot-path accesses from the mm_struct. Each task holds its own refcount on the sched_cache_group, separate from the reference held by its mm_struct. The reference is acquired in copy_mm() (fork) and exec_mmap() (exec), and released in exit_mm(). This fixes the use-after-free when account_mm_sched() reaches the group through a task whose mm is being switched, as reported by Hyunwoo: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/ Convert all scheduler code in fair.c and exit.c to use p->sched_cache_grp instead of p->mm->sched_cache_grp. Keep the fork/exec/exit reference management out of the generic mm paths: add sched_cache_fork(), sched_cache_fork_cleanup(), sched_cache_exec_mmap() and sched_cache_exit_mm() in kernel/sched/cache_sched.c (with empty stubs for !CONFIG_SCHED_CACHE), so fs/exec.c, kernel/fork.c and kernel/exit.c each call one helper instead of open-coding the refcounting under #ifdef. Also add sched_cache_group_get() and task_cache_group_get(). Fixes: df0d98475954 ("sched/cache: Introduce infrastructure for cache-aware= load balancing") Reported-by: Hyunwoo Kim Closes: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/ Reported-by: Zenghui Yu (Huawei) Closes: https://lore.kernel.org/all/343a7e07-7fad-4979-9c9b-82ec038c293c@li= nux.dev/ Cc: stable@kernel.org #7.2.x Co-developed-by: Chen Yu Signed-off-by: Chen Yu Signed-off-by: Tim Chen --- fs/exec.c | 1 + include/linux/sched.h | 13 +++++ kernel/exit.c | 33 +----------- kernel/fork.c | 2 + kernel/sched/cache_sched.c | 87 ++++++++++++++++++++++++++++++ kernel/sched/fair.c | 106 ++++++++++++++++--------------------- kernel/sched/sched.h | 3 ++ 7 files changed, 153 insertions(+), 92 deletions(-) diff --git a/fs/exec.c b/fs/exec.c index 745f6eb5279e..6400bae97a32 100644 --- a/fs/exec.c +++ b/fs/exec.c @@ -882,6 +882,7 @@ static int exec_mmap(struct linux_binprm *bprm) active_mm =3D tsk->active_mm; tsk->active_mm =3D mm; tsk->mm =3D mm; + sched_cache_exec_mmap(tsk, mm); mm_init_cid(mm, tsk); exec_state =3D task_exec_state_replace(tsk, exec_state); /* diff --git a/include/linux/sched.h b/include/linux/sched.h index 1f254364f216..1aa81cb6637b 100644 --- a/include/linux/sched.h +++ b/include/linux/sched.h @@ -1434,6 +1434,7 @@ struct task_struct { #ifdef CONFIG_SCHED_CACHE struct callback_head cache_work; int preferred_llc; + struct sched_cache_group __rcu *sched_cache_grp; /* 1: task was enqueued to its preferred LLC, 0 otherwise */ int pref_llc_queued; #endif @@ -2418,11 +2419,23 @@ struct sched_cache_group { } ____cacheline_aligned_in_smp; =20 void sched_cache_group_put(struct sched_cache_group *grp); +struct sched_cache_group *sched_cache_group_get(struct sched_cache_group *= grp); +struct sched_cache_group *task_cache_group_get(struct task_struct *p); + +void sched_cache_fork(struct task_struct *p); +void sched_cache_fork_cleanup(struct task_struct *p); +void sched_cache_exec_mmap(struct task_struct *p, struct mm_struct *mm); +void sched_cache_exit_mm(struct task_struct *p); =20 #else =20 struct sched_cache_group { }; =20 +static inline void sched_cache_fork(struct task_struct *p) { } +static inline void sched_cache_fork_cleanup(struct task_struct *p) { } +static inline void sched_cache_exec_mmap(struct task_struct *p, struct mm_= struct *mm) { } +static inline void sched_cache_exit_mm(struct task_struct *p) { } + #endif =20 #ifndef MODULE diff --git a/kernel/exit.c b/kernel/exit.c index 16abbe1cf682..35afa1d2d251 100644 --- a/kernel/exit.c +++ b/kernel/exit.c @@ -547,37 +547,6 @@ void mm_update_next_owner(struct mm_struct *mm) } #endif /* CONFIG_MEMCG */ =20 -#if defined(CONFIG_SCHED_CACHE) && defined(CONFIG_NUMA_BALANCING) -/* - * Subtract the memory footprint of the current task from - * mm. - */ -static void exit_mm_sched_cache(struct mm_struct *mm) -{ - struct sched_cache_group *grp; - unsigned long fp, sub; - - if (!current->total_numa_faults) - return; - /* - * No lock protection due to performance considerations. - * Make sure the group footprint does not become - * negative. - */ - grp =3D READ_ONCE(mm->sched_cache_grp); - if (!grp) - return; - - fp =3D READ_ONCE(grp->footprint); - sub =3D min(fp, current->total_numa_faults); - WRITE_ONCE(grp->footprint, fp - sub); -} -#else -static inline void exit_mm_sched_cache(struct mm_struct *mm) -{ -} -#endif /* CONFIG_SCHED_CACHE CONFIG_NUMA_BALANCING */ - /* * Turn us into a lazy TLB process if we * aren't already.. @@ -590,7 +559,7 @@ static void exit_mm(void) if (!mm) return; =20 - exit_mm_sched_cache(mm); + sched_cache_exit_mm(current); =20 mmap_read_lock(mm); mmgrab_lazy_tlb(mm); diff --git a/kernel/fork.c b/kernel/fork.c index 416758c8a3d4..d9b263a32471 100644 --- a/kernel/fork.c +++ b/kernel/fork.c @@ -1599,6 +1599,7 @@ static int copy_mm(u64 clone_flags, struct task_struc= t *tsk) =20 tsk->mm =3D mm; tsk->active_mm =3D mm; + sched_cache_fork(tsk); return 0; } =20 @@ -2599,6 +2600,7 @@ __latent_entropy struct task_struct *copy_process( bad_fork_cleanup_namespaces: exit_nsproxy_namespaces(p); bad_fork_cleanup_mm: + sched_cache_fork_cleanup(p); if (p->mm) { mm_clear_owner(p->mm, p); mmput(p->mm); diff --git a/kernel/sched/cache_sched.c b/kernel/sched/cache_sched.c index ff3d9538e7a4..7c5d23a1e09c 100644 --- a/kernel/sched/cache_sched.c +++ b/kernel/sched/cache_sched.c @@ -1,6 +1,93 @@ // SPDX-License-Identifier: GPL-2.0-only #include "sched.h" =20 +#define rcu_deref_sched_cache_grp(tsk) \ + rcu_dereference_check((tsk)->sched_cache_grp, (tsk) =3D=3D current) + +static struct sched_cache_group *sched_cache_replace_grp(struct task_struc= t *p, + struct sched_cache_group *new) +{ + struct sched_cache_group *old; + + old =3D rcu_deref_sched_cache_grp(p); + rcu_assign_pointer(p->sched_cache_grp, new); + + return old; +} + +struct sched_cache_group *sched_cache_group_get(struct sched_cache_group *= grp) +{ + /* + * refcount_inc_not_zero() is the acquire primitive for lockless + * (RCU) lookups; plain refcount_inc() would scribble the count if + * it already reached zero. Return NULL in that case. + */ + if (grp && !refcount_inc_not_zero(&grp->refcnt)) + grp =3D NULL; + + return grp; +} + +struct sched_cache_group *task_cache_group_get(struct task_struct *p) +{ + guard(rcu)(); + return sched_cache_group_get(rcu_dereference(p->sched_cache_grp)); +} + +void sched_cache_fork(struct task_struct *p) +{ + /* + * The child takes its own reference on the mm's cache group, separate + * from the reference held by the mm. @p is not yet visible to readers, + * so a plain initializing store is enough. + */ + RCU_INIT_POINTER(p->sched_cache_grp, + sched_cache_group_get(p->mm->sched_cache_grp)); +} + +void sched_cache_fork_cleanup(struct task_struct *p) +{ + /* + * A fork that fails after sched_cache_fork() never reaches exit_mm(), + * so drop the reference here. @p never became visible, so there are no + * concurrent readers and the reference we hold keeps the group alive. + */ + sched_cache_group_put(rcu_access_pointer(p->sched_cache_grp)); + RCU_INIT_POINTER(p->sched_cache_grp, NULL); +} + +void sched_cache_exec_mmap(struct task_struct *p, struct mm_struct *mm) +{ + struct sched_cache_group *old; + + /* + * Acquire the new reference before publishing the pointer, then drop + * the old one. @p is current and the only writer of its own pointer. + */ + old =3D sched_cache_replace_grp(p, sched_cache_group_get(mm->sched_cache_= grp)); + sched_cache_group_put(old); +} + +void sched_cache_exit_mm(struct task_struct *p) +{ + struct sched_cache_group *grp =3D sched_cache_replace_grp(p, NULL); + +#ifdef CONFIG_NUMA_BALANCING + /* + * Subtract this task's footprint from the group before dropping the + * reference, so the group footprint converges as its threads exit. + * Unlocked for performance; clamp to avoid underflow. + */ + if (grp && p->total_numa_faults) { + unsigned long fp =3D READ_ONCE(grp->footprint); + unsigned long sub =3D min(fp, p->total_numa_faults); + + WRITE_ONCE(grp->footprint, fp - sub); + } +#endif + sched_cache_group_put(grp); +} + static void sched_cache_group_free_rcu(struct rcu_head *rcu) { struct sched_cache_group *grp =3D diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 6e939807dff2..3e2d236be2a3 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -1483,7 +1483,7 @@ static inline int get_sched_cache_scale(int mul) return (1 + (tol - 1) * mul); } =20 -static bool exceed_llc_capacity(struct mm_struct *mm, int cpu) +static bool exceed_llc_capacity(struct sched_cache_group *grp, int cpu) { #ifdef CONFIG_NUMA_BALANCING unsigned long llc, footprint; @@ -1497,11 +1497,6 @@ static bool exceed_llc_capacity(struct mm_struct *mm= , int cpu) return true; =20 if (static_branch_likely(&sched_numa_balancing)) { - struct sched_cache_group *grp =3D READ_ONCE(mm->sched_cache_grp); - - if (!grp) - return true; - /* * TBD: RDT exclusive LLC ways reserved should be * excluded. @@ -1536,10 +1531,9 @@ static bool exceed_llc_capacity(struct mm_struct *mm= , int cpu) return false; } =20 -static bool invalid_llc_nr(struct mm_struct *mm, struct task_struct *p, +static bool invalid_llc_nr(struct sched_cache_group *grp, struct task_stru= ct *p, int cpu) { - struct sched_cache_group *grp; int scale; =20 if (get_nr_threads(p) <=3D 1) @@ -1553,10 +1547,6 @@ static bool invalid_llc_nr(struct mm_struct *mm, str= uct task_struct *p, if (scale =3D=3D INT_MAX) return false; =20 - grp =3D READ_ONCE(mm->sched_cache_grp); - if (!grp) - return true; - return !fits_capacity((READ_ONCE(grp->nr_running_avg) * cpu_smt_num_threa= ds), (scale * per_cpu(sd_llc_size, cpu))); } @@ -1765,15 +1755,10 @@ static unsigned long fraction_mm_sched(struct rq *r= q, return div64_u64(NICE_0_LOAD * pcpu_sched->runtime, rq->cpu_runtime + 1); } =20 -static int get_pref_llc(struct task_struct *p, struct mm_struct *mm) +static int get_pref_llc(struct task_struct *p, struct sched_cache_group *g= rp) { int mm_sched_llc =3D -1, mm_sched_cpu; - struct sched_cache_group *grp; =20 - if (!mm) - return -1; - - grp =3D READ_ONCE(mm->sched_cache_grp); if (!grp) return -1; =20 @@ -1807,9 +1792,8 @@ static unsigned int task_running_on_cpu(int cpu, stru= ct task_struct *p); static inline void account_mm_sched(struct rq *rq, struct task_struct *p, s64 delta_exec) { + struct sched_cache_group *grp =3D rcu_dereference_all(p->sched_cache_grp); struct sched_cache_time *pcpu_sched; - struct sched_cache_group *grp; - struct mm_struct *mm =3D p->mm; int mm_sched_llc =3D -1; unsigned long epoch; =20 @@ -1820,12 +1804,8 @@ void account_mm_sched(struct rq *rq, struct task_str= uct *p, s64 delta_exec) return; /* * init_task, kthreads and user thread created - * by user_mode_thread() don't have mm. + * by user_mode_thread() don't have a cache group. */ - if (!mm) - return; - - grp =3D READ_ONCE(mm->sched_cache_grp); if (!grp || !grp->pcpu_sched) return; =20 @@ -1843,13 +1823,13 @@ void account_mm_sched(struct rq *rq, struct task_st= ruct *p, s64 delta_exec) * its preferred state. */ if ((long)(epoch - READ_ONCE(grp->epoch)) > llc_epoch_affinity_timeout || - invalid_llc_nr(mm, p, cpu_of(rq)) || - exceed_llc_capacity(mm, cpu_of(rq))) { + invalid_llc_nr(grp, p, cpu_of(rq)) || + exceed_llc_capacity(grp, cpu_of(rq))) { if (READ_ONCE(grp->cpu) !=3D -1) WRITE_ONCE(grp->cpu, -1); } =20 - mm_sched_llc =3D get_pref_llc(p, mm); + mm_sched_llc =3D get_pref_llc(p, grp); =20 /* task not on rq accounted later in account_entity_enqueue() */ if (task_running_on_cpu(rq->cpu, p) && @@ -1862,19 +1842,15 @@ void account_mm_sched(struct rq *rq, struct task_st= ruct *p, s64 delta_exec) =20 static void task_tick_cache(struct rq *rq, struct task_struct *p) { + struct sched_cache_group *grp =3D rcu_dereference_all(p->sched_cache_grp); struct callback_head *work =3D &p->cache_work; - struct sched_cache_group *grp; - struct mm_struct *mm =3D p->mm; unsigned long epoch; =20 if (!sched_cache_enabled()) return; =20 - if (!mm || p->flags & PF_KTHREAD) - return; - - grp =3D READ_ONCE(mm->sched_cache_grp); - if (!grp || !grp->pcpu_sched) + if (!grp || p->flags & PF_KTHREAD || + !grp->pcpu_sched) return; =20 epoch =3D rq->cpu_epoch; @@ -1956,14 +1932,13 @@ static inline void update_avg_scale(u64 *avg, u64 s= ample) =20 static void task_cache_work(struct callback_head *work) { + struct sched_cache_group *grp __free(sched_cache_group_put) =3D NULL; + cpumask_var_t cpus __free(free_cpumask_var) =3D CPUMASK_VAR_NULL; int cpu, m_a_cpu =3D -1, nr_running =3D 0, curr_cpu; unsigned long next_scan, now =3D jiffies; struct task_struct *p =3D current, *cur; unsigned long curr_m_a_occ =3D 0; - struct sched_cache_group *grp; - struct mm_struct *mm =3D p->mm; unsigned long m_a_occ =3D 0; - cpumask_var_t cpus; =20 WARN_ON_ONCE(work !=3D &p->cache_work); =20 @@ -1972,7 +1947,12 @@ static void task_cache_work(struct callback_head *wo= rk) if (p->flags & PF_EXITING) return; =20 - grp =3D READ_ONCE(mm->sched_cache_grp); + /* + * A reference makes sure grp is not released by others. The rcu + * lock can not be held till after zalloc_cpumask_var() below, + * because the latter might sleep. + */ + grp =3D task_cache_group_get(p); if (!grp) return; =20 @@ -1987,8 +1967,8 @@ static void task_cache_work(struct callback_head *wor= k) return; =20 curr_cpu =3D task_cpu(p); - if (invalid_llc_nr(mm, p, curr_cpu) || - exceed_llc_capacity(mm, curr_cpu)) { + if (invalid_llc_nr(grp, p, curr_cpu) || + exceed_llc_capacity(grp, curr_cpu)) { if (READ_ONCE(grp->cpu) !=3D -1) WRITE_ONCE(grp->cpu, -1); =20 @@ -2021,9 +2001,13 @@ static void task_cache_work(struct callback_head *wo= rk) m_cpu =3D i; } =20 + /* + * rcu_access_pointer() is used because the + * pointer is only compared, never dereferenced. + */ cur =3D rcu_dereference_all(cpu_rq(i)->curr); if (cur && !(cur->flags & (PF_EXITING | PF_KTHREAD)) && - cur->mm =3D=3D mm) + rcu_access_pointer(cur->sched_cache_grp) =3D=3D grp) nr_running++; } =20 @@ -2069,7 +2053,6 @@ static void task_cache_work(struct callback_head *wor= k) } =20 update_avg_scale(&grp->nr_running_avg, nr_running); - free_cpumask_var(cpus); } =20 void init_sched_mm(struct task_struct *p) @@ -2078,6 +2061,13 @@ void init_sched_mm(struct task_struct *p) =20 init_task_work(work, task_cache_work); work->next =3D work; + /* + * dup_task_struct() copies the parent's task_struct, including its + * sched_cache_grp, for which the child holds no reference. Clear it + * here - before copy_mm() runs - so the child never carries a + * borrowed pointer that the fork error path would put. + */ + RCU_INIT_POINTER(p->sched_cache_grp, NULL); /* * Reset new task's preference to avoid * polluting account_llc_enqueue(). @@ -3878,10 +3868,9 @@ static void task_numa_placement(struct task_struct *= p) * heuristic and occasional lost updates are tolerable. * * If a task exits, its corresponding footprint must - * be subtracted from the mm->sched_cache_grp->footprint, - * otherwise the mm->sched_cache_grp->footprint will not - * converge: the exiting thread's footprint remains - * unchanged/undecayed in mm->sched_cache_grp->footprint. + * be subtracted from p->sched_cache_grp->footprint, + * otherwise the footprint will not converge: the + * exiting thread's footprint remains unchanged/undecayed. * See exit_mm(). * * Lost updates and unsynchronized subtraction @@ -3889,12 +3878,14 @@ static void task_numa_placement(struct task_struct = *p) * go negative. Clamp to zero to prevent the * unsigned footprint from wrapping. */ - grp =3D READ_ONCE(p->mm->sched_cache_grp); - if (!grp) - continue; + scoped_guard(rcu) { + grp =3D rcu_dereference(p->sched_cache_grp); =20 - new_fp =3D (long)READ_ONCE(grp->footprint) + diff; - WRITE_ONCE(grp->footprint, max(new_fp, 0L)); + if (grp) { + new_fp =3D (long)READ_ONCE(grp->footprint) + diff; + WRITE_ONCE(grp->footprint, max(new_fp, 0L)); + } + } #endif } =20 @@ -10843,7 +10834,6 @@ static enum llc_mig can_migrate_llc_task(struct lb_= env *env, struct task_struct *p) { struct sched_cache_group *grp; - struct mm_struct *mm; bool to_pref; int cpu, src_cpu, dst_cpu; =20 @@ -10852,11 +10842,7 @@ static enum llc_mig can_migrate_llc_task(struct lb= _env *env, =20 src_cpu =3D env->src_cpu; dst_cpu =3D env->dst_cpu; - mm =3D p->mm; - if (!mm) - return mig_unrestricted; - - grp =3D READ_ONCE(mm->sched_cache_grp); + grp =3D rcu_dereference_all(p->sched_cache_grp); if (!grp) return mig_unrestricted; =20 @@ -10865,8 +10851,8 @@ static enum llc_mig can_migrate_llc_task(struct lb_= env *env, return mig_unrestricted; =20 /* skip cache aware load balance for too many threads */ - if (invalid_llc_nr(mm, p, dst_cpu) || - exceed_llc_capacity(mm, dst_cpu)) { + if (invalid_llc_nr(grp, p, dst_cpu) || + exceed_llc_capacity(grp, dst_cpu)) { if (READ_ONCE(grp->cpu) !=3D -1) WRITE_ONCE(grp->cpu, -1); return mig_unrestricted; diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h index e656c7059bf8..8b67af28a471 100644 --- a/kernel/sched/sched.h +++ b/kernel/sched/sched.h @@ -4145,6 +4145,9 @@ static inline bool sched_cache_enabled(void) return static_branch_unlikely(&sched_cache_active); } =20 +DEFINE_FREE(sched_cache_group_put, struct sched_cache_group *, + sched_cache_group_put(_T)); + extern void sched_cache_active_set(void); =20 #endif --=20 2.32.0 From nobody Thu Sep 24 17:53:07 2026 Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.15]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B2D4333F5B6; Tue, 22 Sep 2026 00:32:18 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=192.198.163.15 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790037140; cv=none; b=DNsRqQt8gIzDFQY/yIehkWtj66PjI1z1Gz9KeD3cPIkKPw1CS2rKaunQ29KjsD4BBE/4l8W1ZYHelzpBU/zMRrT2Blc/7x42btKzvNAJ0Gw9GNUKqhNwxp3MW7Xe3Sf75i8l4zU+pA72XbhY+fakjNCJsCDY7ajbt91XtAa/yM0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790037140; c=relaxed/simple; bh=pOlHJ5o3h0eTKEGfK5htB/gt/IweJdH5Z9ZDtcA9SHU=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=LKXNDIITufTSXpem7p4rYBeZ9zrxriWeUbmP9jUFCnoQlekU45jxHbWeBi1eGGjrqEkp/T+aTI6/hViIXOYeoLZnSwLvo77okQmVMGNA0CNs4JM/6vWM1LAfwuXWKzrYy1tsFbmJ8BOrY1D25/F7Or3nXwEdlnK01bfKnlgL1RY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com; spf=pass smtp.mailfrom=linux.intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=FhudifD4; arc=none smtp.client-ip=192.198.163.15 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="FhudifD4" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1790037139; x=1821573139; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=pOlHJ5o3h0eTKEGfK5htB/gt/IweJdH5Z9ZDtcA9SHU=; b=FhudifD4556rpGjT5Cp4KrW82omp9ZPr1vgjyNzZPOuv5G8pQKfah8Gq 8sFZTFg2ICiuXAbayqyagGQvwKtzns6NufKwl6qq3MHVW4oofmxSdj2NU 8I9RtsAUx1lIVonKfxFqYaEzXAJI8Vg6B5W3hmLibe0GhbY5Yxie1AD2A aOvXTU7PDpNHvSl66bc/JtaFmKCvMvrwvC9yG8bhfNThVy2Nr520133Zl kt0XuAnqNWFHjrmcPY1vE8eykoIROUQNVMJj0ITUiSVbO1/dnnOGqECEv P/4a5rJv8y6ato2oVnDCms2OSeZvFY1PWnBnPsuqg2b6UzvLvt7phnTrr w==; X-CSE-ConnectionGUID: 2HVDypgkSkqi9/Ni2KNgmw== X-CSE-MsgGUID: 6/Z09I8yR7+3cv9RoSXsaw== X-IronPort-AV: E=McAfee;i="6800,10657,11912"; a="90716718" X-IronPort-AV: E=Sophos;i="6.27,115,1787036400"; d="scan'208";a="90716718" Received: from fmviesa004.fm.intel.com ([10.60.135.144]) by fmvoesa109.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 21 Sep 2026 17:32:18 -0700 X-CSE-ConnectionGUID: FK5iBX5YTM+jt9J2U/V1QQ== X-CSE-MsgGUID: dAaD2tWDQOCsMtP+fuKQ9Q== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.27,115,1787036400"; d="scan'208";a="277671907" Received: from b04f130c83f2.jf.intel.com ([10.165.154.98]) by fmviesa004.fm.intel.com with ESMTP; 21 Sep 2026 17:32:17 -0700 From: Tim Chen To: Peter Zijlstra , Ingo Molnar Cc: Chen Yu , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , K Prateek Nayak , Kees Cook , Christian Brauner , Alexander Viro , Jan Kara , Shrikanth Hegde , Qais Yousef , Aaron Lu , Srikar Dronamraju , Vineeth Remanan Pillai , Ricardo Neri-Calderon , Lu Wang , Hyunwoo Kim , Zhan Xusheng , Zhan Xusheng , Yi Lai , Tim Chen , "Rafael J . Wysocki" , Greg Kroah-Hartman , Danilo Krummrich , Zenghui Yu , linux-kernel@vger.kernel.org, linux-mm@kvack.org, linux-fsdevel@vger.kernel.org, stable@kernel.org Subject: [PATCH v2 5/6] sched/cache: Skip kernel thread for cache aware scheduling Date: Mon, 21 Sep 2026 17:37:26 -0700 Message-Id: <058f0c6ea7b991c177a17de347fa3157f25489a7.1790035273.git.tim.c.chen@linux.intel.com> X-Mailer: git-send-email 2.32.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Chen Yu Kernel thread should not be covered by cache aware scheduling as it borrows the statistics from the user space thread. Filter the kernel thread in account_mm_sched(). In theory a kernel thread does not have any valid cache group, so !grp should gate the kernel thread. Add the PF_KTHREAD check explicitly here for safety reasons, to guard against future modifications and to pair with task_tick_cache(). Fixes: df0d98475954 ("sched/cache: Introduce infrastructure for cache-aware= load balancing") Cc: stable@kernel.org #7.2.x Signed-off-by: Chen Yu Signed-off-by: Tim Chen --- kernel/sched/fair.c | 10 ++++++++-- 1 file changed, 8 insertions(+), 2 deletions(-) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 3e2d236be2a3..341d2f9ed72b 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -1805,8 +1805,14 @@ void account_mm_sched(struct rq *rq, struct task_str= uct *p, s64 delta_exec) /* * init_task, kthreads and user thread created * by user_mode_thread() don't have a cache group. - */ - if (!grp || !grp->pcpu_sched) + * In theory a kernel thread does not have any valid + * cache group, because sched_cache_fork() is not + * invoked for a kernel thread - !grp should gate the + * kernel thread. Add the PF_KTHREAD check explicitly + * here for safety reasons, to guard against future + * modifications and to pair with task_tick_cache(). + */ + if (!grp || p->flags & PF_KTHREAD || !grp->pcpu_sched) return; =20 pcpu_sched =3D per_cpu_ptr(grp->pcpu_sched, cpu_of(rq)); --=20 2.32.0 From nobody Thu Sep 24 17:53:07 2026 Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.15]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D5AAB2EC54A; Tue, 22 Sep 2026 00:32:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=192.198.163.15 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790037142; cv=none; b=hMCTlZvGp4Z5p6pf2AGikQabdVmUPrUBPdowXM4806oeDVmtifTCzcYBQz1b2ymLIygMMDRKKloyTujsSPU7vldcwQSqPmjd3SNPT0HALBs8nWFxhvFXTZYqOdDAqMLMbCUkIbrcJwGu/41LUPZW359tRdISBQZlXQbsRbdFyI8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790037142; c=relaxed/simple; bh=eXYuTenHkmviiP89DV79U/M3G01VwUN3IcB9EPfEUW4=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=B/Cl0cJA7P+cV92IVzAt56+/DDVsMXinWJ5dGbomR00gt5M+pET+BufUK/fsrZE0itwIFD/XAkEKSjBpR6W6D6SVmg/GEqBnSOrcDJsyvdVvtTICB6kLOIFrgrz0MOr7em5JjUk1BWxAY2EdX6NJySd4xnmZGFo8I/IUOgn63So= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com; spf=pass smtp.mailfrom=linux.intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=U9XRyQBI; arc=none smtp.client-ip=192.198.163.15 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="U9XRyQBI" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1790037141; x=1821573141; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=eXYuTenHkmviiP89DV79U/M3G01VwUN3IcB9EPfEUW4=; b=U9XRyQBIhTTNdK4VVoWiUZ89RPpp+PS1pe0RGsJRAK3ijKxClGCvELgW UlwF/oAWb0hN+VeJ7cyQjwAw8Xs7uRxkbN1nzVX8c/vQvYrtCHM0r/EmQ 5XFenk12usnV3pkK2DgVq74McDz6kRPmsVNB+Q9/Mxb3KXuoQEOUNET7p rRmZynx+sVCqPyghi7KbCQpTl976Ti7AvGmj8+r147tm11KfP1CyhFcP0 zGe23Xcv5FEw4k+pUegEyYgnWeqPIzpVeUxmRkNs1LVr+fGHYAAbXGLIS X4s0wLTjj2ErzVEZYjF8ayu0P8gFD+feFkVraXPT9oF+WW0RfIi96Y8U7 A==; X-CSE-ConnectionGUID: wnmUxw7ORdaYBM9mx5Io/w== X-CSE-MsgGUID: 8nca5NqhQcGqJ8QgLNsFEg== X-IronPort-AV: E=McAfee;i="6800,10657,11912"; a="90716741" X-IronPort-AV: E=Sophos;i="6.27,115,1787036400"; d="scan'208";a="90716741" Received: from fmviesa004.fm.intel.com ([10.60.135.144]) by fmvoesa109.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 21 Sep 2026 17:32:20 -0700 X-CSE-ConnectionGUID: gRU2xwr9RXGJBzfMN92iYw== X-CSE-MsgGUID: NTHk847jRTG3cUQ0MrdRsw== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.27,115,1787036400"; d="scan'208";a="277671922" Received: from b04f130c83f2.jf.intel.com ([10.165.154.98]) by fmviesa004.fm.intel.com with ESMTP; 21 Sep 2026 17:32:19 -0700 From: Tim Chen To: Peter Zijlstra , Ingo Molnar Cc: Davi Chaves Azevedo , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , K Prateek Nayak , Kees Cook , Christian Brauner , Alexander Viro , Jan Kara , Shrikanth Hegde , Qais Yousef , Aaron Lu , Srikar Dronamraju , Vineeth Remanan Pillai , Ricardo Neri-Calderon , Chen Yu , Lu Wang , Hyunwoo Kim , Zhan Xusheng , Zhan Xusheng , Yi Lai , Tim Chen , "Rafael J . Wysocki" , Greg Kroah-Hartman , Danilo Krummrich , Zenghui Yu , linux-kernel@vger.kernel.org, linux-mm@kvack.org, linux-fsdevel@vger.kernel.org, stable@kernel.org Subject: [PATCH v2 6/6] sched/cache: Refresh LLC capacity across CPU hotplug Date: Mon, 21 Sep 2026 17:37:27 -0700 Message-Id: <6751d93e15889e624796c74db0bfe66603d60b1b.1790035273.git.tim.c.chen@linux.intel.com> X-Mailer: git-send-email 2.32.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Davi Chaves Azevedo The scheduler scales LLC capacity by the fraction of cache-sharing CPUs covered by a domain: llc_bytes =3D cache_size * span_weight / shared_weight During CPU teardown, sched_cpu_deactivate() rebuilds scheduler domains before cacheinfo_cpu_pre_down() removes the CPU from shared_cpu_map. The new domains therefore use the old sharing weight. The later call to sched_update_llc_bytes() looks up the departing CPU's sd_llc, which has already been detached, and returns without correcting the surviving CPUs. On a Ryzen 5 7535U with twelve logical CPUs sharing a 16 MiB LLC, offlining one SMT sibling left the remaining CPUs with: llc_bytes =3D floor(16777216 * 11 / 12) =3D 15379114 bytes The correct capacity is still 16777216 bytes. On systems with active cache-aware scheduling, an underestimated capacity can cause exceed_llc_capacity() to reject aggregation for a process whose footprint would fit. Unchanged cpuset partitions sharing the physical cache can also retain stale capacity when a CPU comes online in another partition. Pass the cache-sharing mask already retained by cacheinfo to the scheduler update. Refresh every surviving CPU using its own LLC domain so that each partition receives the correct share. This also preserves the correction needed as cache-sharing maps grow during boot. Keep the existing CPU-hotplug and scheduler-domain synchronization. The update remains on the hotplug path; no steady-state scheduling operation or persistent allocation is added. Fixes: 7030513a0877 ("sched/cache: Calculate the LLC size and store it in s= ched_domain") Signed-off-by: Davi Chaves Azevedo Reviewed-by: Chen Yu Tested-by: Chen Yu Reviewed-by: Tim Chen Reviewed-by: K Prateek Nayak Tested-by: K Prateek Nayak Cc: stable@kernel.org #7.2.x Signed-off-by: Tim Chen --- drivers/base/cacheinfo.c | 11 ++++++----- include/linux/sched/topology.h | 4 ++-- kernel/sched/topology.c | 22 +++++++++++++--------- 3 files changed, 21 insertions(+), 16 deletions(-) diff --git a/drivers/base/cacheinfo.c b/drivers/base/cacheinfo.c index 9f9c72727a05..7a47a392568a 100644 --- a/drivers/base/cacheinfo.c +++ b/drivers/base/cacheinfo.c @@ -1040,9 +1040,10 @@ static int cacheinfo_cpu_online(unsigned int cpu) rc =3D cache_add_dev(cpu); if (rc) goto err; - if (cpu_map_shared_cache(true, cpu, &cpu_map)) + if (cpu_map_shared_cache(true, cpu, &cpu_map)) { update_per_cpu_data_slice_size(true, cpu, cpu_map); - sched_update_llc_bytes(cpu); + sched_update_llc_bytes(cpu_map); + } return 0; err: free_cache_attributes(cpu); @@ -1059,10 +1060,10 @@ static int cacheinfo_cpu_pre_down(unsigned int cpu) cpu_cache_sysfs_exit(cpu); =20 free_cache_attributes(cpu); - if (nr_shared > 1) + if (nr_shared > 1) { update_per_cpu_data_slice_size(false, cpu, cpu_map); - - sched_update_llc_bytes(cpu); + sched_update_llc_bytes(cpu_map); + } =20 return 0; } diff --git a/include/linux/sched/topology.h b/include/linux/sched/topology.h index b5d9d7c2b8ad..f96812d71c51 100644 --- a/include/linux/sched/topology.h +++ b/include/linux/sched/topology.h @@ -281,9 +281,9 @@ static inline int task_node(const struct task_struct *p) } =20 #ifdef CONFIG_SCHED_CACHE -extern void sched_update_llc_bytes(unsigned int cpu); +extern void sched_update_llc_bytes(const struct cpumask *cpus); #else -static inline void sched_update_llc_bytes(unsigned int cpu) { } +static inline void sched_update_llc_bytes(const struct cpumask *cpus) { } #endif =20 #endif /* _LINUX_SCHED_TOPOLOGY_H */ diff --git a/kernel/sched/topology.c b/kernel/sched/topology.c index 0248227d983a..3dab0253976f 100644 --- a/kernel/sched/topology.c +++ b/kernel/sched/topology.c @@ -985,8 +985,8 @@ void sched_cache_active_set(void) } =20 /* - * Update the bottom sched_domain's llc_bytes for @cpu and all its - * LLC siblings. Called from cacheinfo_cpu_online() or + * Update the bottom sched_domain's llc_bytes for @cpus sharing a physical + * LLC. Called from cacheinfo_cpu_online() or * cacheinfo_cpu_pre_down() with cpu hotplug lock held. * * Note: get_effective_llc_bytes() returns 0 on PowerPC. @@ -996,17 +996,13 @@ void sched_cache_active_set(void) * and does not populates the per-CPU struct cpu_cacheinfo array * that get_cpu_cacheinfo_llc() reads. */ -void sched_update_llc_bytes(unsigned int cpu) +void sched_update_llc_bytes(const struct cpumask *cpus) { struct sched_domain *sd, *sdp; unsigned int i; =20 sched_domains_mutex_lock(); =20 - sdp =3D rcu_dereference_sched_domain(per_cpu(sd_llc, cpu)); - if (!sdp) - goto unlock; - /* * ci->shared_cpu_map is built incrementally as CPUs come * online, so the first CPU in an LLC initially sees @@ -1014,14 +1010,22 @@ void sched_update_llc_bytes(unsigned int cpu) * get_effective_llc_bytes(). Re-evaluating every LLC * sibling on each online event corrects this once the full * shared_cpu_map is known. + * + * The departing CPU's domains have already been detached when + * cacheinfo removes it. Use the surviving cache siblings instead. + * They may belong to different cpuset partitions, so use each CPU's + * own LLC domain to scale its share of the physical cache. */ - for_each_cpu(i, sched_domain_span(sdp)) { + for_each_cpu(i, cpus) { + sdp =3D rcu_dereference_sched_domain(per_cpu(sd_llc, i)); + if (!sdp) + continue; + sd =3D rcu_dereference_sched_domain(cpu_rq(i)->sd); if (sd) sd->llc_bytes =3D get_effective_llc_bytes(i, sdp); } =20 -unlock: sched_domains_mutex_unlock(); } =20 --=20 2.32.0