From nobody Sat Sep 26 11:46:28 2026 Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.8]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2B55D2836F for ; Wed, 2 Sep 2026 00:04:01 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=192.198.163.8 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788307444; cv=none; b=XCGyeFO2JA+ThrQjrvcCZ20dQWJ3Tij1neFjBkSZJ3owMOvEiNACxLHBlOCvyqEClC75Y7hTvdqcg/ejAJLrua8VV+x1VyVkeOKTqOge8PFU9MvVPTpEd5lBUoK6vZgGLTPOwcgJi4B/ga7X7/0rkkYsxbXf+72MymjxOJwRv5Y= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788307444; c=relaxed/simple; bh=4HA4NOSLIzsRJHIxI6uRGIPpQMeaoKh+cm0awgGMBMQ=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=EQBVc04IG2etOf6YjHjSj3Xo9CumCnV5nnPD5yLniNGheQl/wqg6Ar5OAzz4i/qgwGqmcqyDCGYZ+ApM7Yirsinw5kKJDgCaNF7/Qq3+ESIML8ZzVpwKsHFWetDsRL33we4X/46ePFJuP5hOkOh1Fj/9man/6YgcM0YoFttWOF0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com; spf=pass smtp.mailfrom=linux.intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=BWd8Eosu; arc=none smtp.client-ip=192.198.163.8 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="BWd8Eosu" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1788307442; x=1819843442; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=4HA4NOSLIzsRJHIxI6uRGIPpQMeaoKh+cm0awgGMBMQ=; b=BWd8EosuxRjjOe07cahsc2Ae/iBhFRCamV3EsPgH+xAvVe1/TF99sJK3 eElpdmnrrqkwP0dAxsz8pewWOX6fA2G+FDuSN6qvx/4pqaFCz3Z8MRH5h YFD8uTSr3CgkvWC4BG7B8cveu9iJjj1duQALrtSyZjthly02MXXpKBS31 0Rt1z7UoDY/7wuPJVcTPJPctMDpdOVuncPlPaijoij/3obkCCzlnPTt6+ w5D3V6r19wyistjYnqUUmkOMPV7zIPro8wLkWocpDMoDwrDwxikR72nSP U4BPUvCfhXHFKf2Gg6PhN24IIfIYbay4bLbZPCrixfCfSDjz8ozW2ORgv Q==; X-CSE-ConnectionGUID: 9DmEIDxfRG2JN6OgjXcYEA== X-CSE-MsgGUID: v7zPn3pESN2EUTv1WCRwQg== X-IronPort-AV: E=McAfee;i="6800,10657,11893"; a="106272982" X-IronPort-AV: E=Sophos;i="6.25,256,1779174000"; d="scan'208";a="106272982" Received: from fmviesa010.fm.intel.com ([10.60.135.150]) by fmvoesa102.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 01 Sep 2026 17:03:57 -0700 X-CSE-ConnectionGUID: /z/QCYmlRButVJ3u2YlN8A== X-CSE-MsgGUID: W5+9GGHUTz2HpPZKxYFwFw== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.25,256,1779174000"; d="scan'208";a="265531792" Received: from b04f130c83f2.jf.intel.com ([10.165.154.98]) by fmviesa010.fm.intel.com with ESMTP; 01 Sep 2026 17:03:42 -0700 From: Tim Chen To: Peter Zijlstra , Ingo Molnar Cc: Tim Chen , Chen Yu , Hyunwoo Kim , Kees Cook , Christian Brauner , Alexander Viro , Jan Kara , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , K Prateek Nayak , Shrikanth Hegde , Qais Yousef , Aaron Lu , Srikar Dronamraju , Vineeth Remanan Pillai , linux-kernel@vger.kernel.org, linux-mm@kvack.org, "chen . yu @ linux . dev" Subject: [PATCH 1/2] sched/cache: Decouple sched_cache_group from mm Date: Tue, 1 Sep 2026 17:08:55 -0700 Message-Id: X-Mailer: git-send-email 2.32.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Currently the sched cache grouping is by mm and the scheduling statistics sched_cache_stat lives in the mm structure. This ties the life cycle of scheduling stats with mm. In account_mm_sched(), the scheduling stats are accessed by task->mm->sc_stat. However, a task may be switching mm on one CPU when another CPU is running account_mm_sched(), and possibly accessing the old mm that was freed. This problem was found when running tests with KASAN by Hyunwoo. https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/ Instead of serializing the mm access by introducing extra acquisition of rq lock in the mm free path, extract sched_cache_stat from mm_struct, rename it as sched_cache_group and manage its life cycle apart from mm_struct with its own ref counting. Access sched_cache_group directly from a task instead of having to go through a task's mm. This allows us to later add a refcount on sched_cache_group when a task links to it, preventing the use after free issue when accessing stale and released old mm and its sched cache stat a task switches to a new mm while account_mm_sched() is done elsewhere. The other benefit of this restructure is in the future, the grouping of tasks to a LLC would have the flexibility to be associated with cgroup, cookie group, numa_group or others instead of just with a single mm address space. Rename sched_cache_stat to sched_cache_group and turn it into a refcounted object allocated from mm_struct. The mm_struct now holds a pointer (sched_cache_grp) to this object instead of embedding it. Introduce kernel/sched/cache_sched.c to host the cache aware scheduling helpers and define sched_cache_group_put() there. Meanwhile skip kthreads in account_mm_sched(), consistent with task_tick_cache(). Reported-by: Hyunwoo Kim Closes: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/ Tested-by: Hyunwoo Kim Fixes: df0d98475954 ("sched/cache: Introduce infrastructure for cache-aware= load balancing") Co-developed-by: Chen Yu Signed-off-by: Chen Yu Signed-off-by: Tim Chen --- include/linux/mm_types.h | 15 ++--- include/linux/sched.h | 8 ++- kernel/exit.c | 6 +- kernel/sched/build_utility.c | 4 ++ kernel/sched/cache_sched.c | 20 ++++++ kernel/sched/fair.c | 118 ++++++++++++++++++++++------------- 6 files changed, 113 insertions(+), 58 deletions(-) create mode 100644 kernel/sched/cache_sched.c diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h index 6d815f6440c9..f3e5a2fadbe5 100644 --- a/include/linux/mm_types.h +++ b/include/linux/mm_types.h @@ -1226,7 +1226,7 @@ struct mm_struct { struct mm_mm_cid mm_cid; =20 /* sched_cache related statistics */ - struct sched_cache_stat sc_stat; + struct sched_cache_group *sched_cache_grp; #ifdef CONFIG_MMU atomic_long_t pgtables_bytes; /* size of all page tables */ #endif @@ -1624,8 +1624,9 @@ static inline unsigned int mm_cid_size(void) #endif /* CONFIG_SCHED_MM_CID */ =20 #ifdef CONFIG_SCHED_CACHE -void mm_init_sched(struct mm_struct *mm, - struct sched_cache_time __percpu *pcpu_sched); +int mm_init_sched(struct mm_struct *mm, + struct sched_cache_time __percpu *pcpu_sched); +void mm_destroy_sched(struct mm_struct *mm); =20 static inline int mm_alloc_sched_noprof(struct mm_struct *mm) { @@ -1635,17 +1636,11 @@ static inline int mm_alloc_sched_noprof(struct mm_s= truct *mm) if (!pcpu_sched) return -ENOMEM; =20 - mm_init_sched(mm, pcpu_sched); - return 0; + return mm_init_sched(mm, pcpu_sched); } =20 #define mm_alloc_sched(...) alloc_hooks(mm_alloc_sched_noprof(__VA_ARGS__)) =20 -static inline void mm_destroy_sched(struct mm_struct *mm) -{ - free_percpu(mm->sc_stat.pcpu_sched); - mm->sc_stat.pcpu_sched =3D NULL; -} #else /* !CONFIG_SCHED_CACHE */ =20 static inline int mm_alloc_sched(struct mm_struct *mm) { return 0; } diff --git a/include/linux/sched.h b/include/linux/sched.h index 8b3d47a325cc..1f254364f216 100644 --- a/include/linux/sched.h +++ b/include/linux/sched.h @@ -2405,7 +2405,7 @@ struct sched_cache_time { unsigned long epoch; }; =20 -struct sched_cache_stat { +struct sched_cache_group { struct sched_cache_time __percpu *pcpu_sched; raw_spinlock_t lock; unsigned long epoch; @@ -2413,11 +2413,15 @@ struct sched_cache_stat { unsigned long next_scan; unsigned long footprint; int cpu; + refcount_t refcnt; + struct rcu_head rcu; } ____cacheline_aligned_in_smp; =20 +void sched_cache_group_put(struct sched_cache_group *grp); + #else =20 -struct sched_cache_stat { }; +struct sched_cache_group { }; =20 #endif =20 diff --git a/kernel/exit.c b/kernel/exit.c index 97686af89501..006edcc0c2c5 100644 --- a/kernel/exit.c +++ b/kernel/exit.c @@ -560,12 +560,12 @@ static void exit_mm_sched_cache(struct mm_struct *mm) return; /* * No lock protection due to performance considerations. - * Make sure mm->sc_stat.footprint does not become + * Make sure the group footprint does not become * negative. */ - fp =3D READ_ONCE(mm->sc_stat.footprint); + fp =3D READ_ONCE(mm->sched_cache_grp->footprint); sub =3D min(fp, current->total_numa_faults); - WRITE_ONCE(mm->sc_stat.footprint, fp - sub); + WRITE_ONCE(mm->sched_cache_grp->footprint, fp - sub); } #else static inline void exit_mm_sched_cache(struct mm_struct *mm) diff --git a/kernel/sched/build_utility.c b/kernel/sched/build_utility.c index e2cf3b08d4e9..24202893b262 100644 --- a/kernel/sched/build_utility.c +++ b/kernel/sched/build_utility.c @@ -89,6 +89,10 @@ # include "core_sched.c" #endif =20 +#ifdef CONFIG_SCHED_CACHE +# include "cache_sched.c" +#endif + #ifdef CONFIG_PSI # include "psi.c" #endif diff --git a/kernel/sched/cache_sched.c b/kernel/sched/cache_sched.c new file mode 100644 index 000000000000..d492df55f9d5 --- /dev/null +++ b/kernel/sched/cache_sched.c @@ -0,0 +1,20 @@ +// SPDX-License-Identifier: GPL-2.0-only +#include "sched.h" + +static void sched_cache_group_free_rcu(struct rcu_head *rcu) +{ + struct sched_cache_group *grp =3D + container_of(rcu, struct sched_cache_group, rcu); + + /* free_percpu() may be called from atomic context. */ + free_percpu(grp->pcpu_sched); + kfree(grp); +} + +void sched_cache_group_put(struct sched_cache_group *grp) +{ + if (!grp || !refcount_dec_and_test(&grp->refcnt)) + return; + + call_rcu(&grp->rcu, sched_cache_group_free_rcu); +} diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 8dff37059faf..8587dcbaa1cf 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -1489,7 +1489,7 @@ static bool exceed_llc_capacity(struct mm_struct *mm,= int cpu) * excluded. */ llc =3D sd->llc_bytes; - footprint =3D READ_ONCE(mm->sc_stat.footprint); + footprint =3D READ_ONCE(mm->sched_cache_grp->footprint); =20 /* * Scale the LLC size by 256*llc_aggr_tolerance @@ -1534,7 +1534,7 @@ static bool invalid_llc_nr(struct mm_struct *mm, stru= ct task_struct *p, if (scale =3D=3D INT_MAX) return false; =20 - return !fits_capacity((mm->sc_stat.nr_running_avg * cpu_smt_num_threads), + return !fits_capacity((mm->sched_cache_grp->nr_running_avg * cpu_smt_num_= threads), (scale * per_cpu(sd_llc_size, cpu))); } =20 @@ -1611,12 +1611,20 @@ static void account_llc_dequeue(struct rq *rq, stru= ct task_struct *p) } } =20 -void mm_init_sched(struct mm_struct *mm, - struct sched_cache_time __percpu *_pcpu_sched) +int mm_init_sched(struct mm_struct *mm, + struct sched_cache_time __percpu *_pcpu_sched) { + struct sched_cache_group *grp; unsigned long epoch =3D 0; int i; =20 + grp =3D kzalloc_obj(*grp); + if (!grp) { + free_percpu(_pcpu_sched); + mm->sched_cache_grp =3D NULL; + return -ENOMEM; + } + for_each_possible_cpu(i) { struct sched_cache_time *pcpu_sched =3D per_cpu_ptr(_pcpu_sched, i); struct rq *rq =3D cpu_rq(i); @@ -1627,18 +1635,35 @@ void mm_init_sched(struct mm_struct *mm, epoch =3D rq->cpu_epoch; } =20 - raw_spin_lock_init(&mm->sc_stat.lock); - mm->sc_stat.epoch =3D epoch; - mm->sc_stat.cpu =3D -1; - mm->sc_stat.next_scan =3D jiffies; - mm->sc_stat.nr_running_avg =3D 0; - mm->sc_stat.footprint =3D 0; + raw_spin_lock_init(&grp->lock); + grp->epoch =3D epoch; + grp->cpu =3D -1; + grp->next_scan =3D jiffies; + grp->nr_running_avg =3D 0; + grp->footprint =3D 0; + refcount_set(&grp->refcnt, 1); /* - * The update to mm->sc_stat should not be reordered - * before initialization to mm's other fields, in case + * The update to grp->pcpu_sched should not be reordered + * before initialization to grp's other fields, in case * the readers may get invalid mm_sched_epoch, etc. */ - smp_store_release(&mm->sc_stat.pcpu_sched, _pcpu_sched); + smp_store_release(&grp->pcpu_sched, _pcpu_sched); + /* + * Publish the group last. Not every reader qualifies it by + * grp->pcpu_sched - can_migrate_llc_task() only checks that the + * pointer is non-NULL before reading grp->footprint and + * grp->nr_running_avg - so a reachable group must already be + * fully initialized. + */ + mm->sched_cache_grp =3D grp; + return 0; +} + +void mm_destroy_sched(struct mm_struct *mm) +{ + if (mm->sched_cache_grp) + sched_cache_group_put(mm->sched_cache_grp); + mm->sched_cache_grp =3D NULL; } =20 /* because why would C be fully specified */ @@ -1696,7 +1721,7 @@ static int get_pref_llc(struct task_struct *p, struct= mm_struct *mm) if (!mm) return -1; =20 - mm_sched_cpu =3D READ_ONCE(mm->sc_stat.cpu); + mm_sched_cpu =3D READ_ONCE(mm->sched_cache_grp->cpu); if (mm_sched_cpu !=3D -1) { mm_sched_llc =3D llc_id(mm_sched_cpu); =20 @@ -1739,11 +1764,15 @@ void account_mm_sched(struct rq *rq, struct task_st= ruct *p, s64 delta_exec) /* * init_task, kthreads and user thread created * by user_mode_thread() don't have mm. + * + * A kthread can temporarily adopt an mm via kthread_use_mm(), + * so p->mm alone does not imply a user task. */ - if (!mm || !mm->sc_stat.pcpu_sched) + if (!mm || p->flags & PF_KTHREAD || !mm->sched_cache_grp || + !mm->sched_cache_grp->pcpu_sched) return; =20 - pcpu_sched =3D per_cpu_ptr(mm->sc_stat.pcpu_sched, cpu_of(rq)); + pcpu_sched =3D per_cpu_ptr(mm->sched_cache_grp->pcpu_sched, cpu_of(rq)); =20 scoped_guard (raw_spinlock, &rq->cpu_epoch_lock) { __update_mm_sched(rq, pcpu_sched); @@ -1756,11 +1785,11 @@ void account_mm_sched(struct rq *rq, struct task_st= ruct *p, s64 delta_exec) * If this process hasn't hit task_cache_work() for a while invalidate * its preferred state. */ - if ((long)(epoch - READ_ONCE(mm->sc_stat.epoch)) > llc_epoch_affinity_tim= eout || + if ((long)(epoch - READ_ONCE(mm->sched_cache_grp->epoch)) > llc_epoch_aff= inity_timeout || invalid_llc_nr(mm, p, cpu_of(rq)) || exceed_llc_capacity(mm, cpu_of(rq))) { - if (READ_ONCE(mm->sc_stat.cpu) !=3D -1) - WRITE_ONCE(mm->sc_stat.cpu, -1); + if (READ_ONCE(mm->sched_cache_grp->cpu) !=3D -1) + WRITE_ONCE(mm->sched_cache_grp->cpu, -1); } =20 mm_sched_llc =3D get_pref_llc(p, mm); @@ -1784,19 +1813,19 @@ static void task_tick_cache(struct rq *rq, struct t= ask_struct *p) return; =20 if (!mm || p->flags & PF_KTHREAD || - !mm->sc_stat.pcpu_sched) + !mm->sched_cache_grp->pcpu_sched) return; =20 epoch =3D rq->cpu_epoch; /* avoid moving backwards */ - if (time_after_eq(mm->sc_stat.epoch, epoch)) + if (time_after_eq(mm->sched_cache_grp->epoch, epoch)) return; =20 - guard(raw_spinlock)(&mm->sc_stat.lock); + guard(raw_spinlock)(&mm->sched_cache_grp->lock); =20 if (work->next =3D=3D work) { task_work_add(p, work, TWA_RESUME); - WRITE_ONCE(mm->sc_stat.epoch, epoch); + WRITE_ONCE(mm->sched_cache_grp->epoch, epoch); } } =20 @@ -1808,7 +1837,7 @@ static void get_scan_cpumasks(cpumask_var_t cpus, str= uct task_struct *p) if (!static_branch_likely(&sched_numa_balancing)) goto out; =20 - cpu =3D READ_ONCE(p->mm->sc_stat.cpu); + cpu =3D READ_ONCE(p->mm->sched_cache_grp->cpu); if (cpu !=3D -1) nid =3D cpu_to_node(cpu); curr_cpu =3D task_cpu(p); @@ -1880,12 +1909,12 @@ static void task_cache_work(struct callback_head *w= ork) if (p->flags & PF_EXITING) return; =20 - next_scan =3D READ_ONCE(mm->sc_stat.next_scan); + next_scan =3D READ_ONCE(mm->sched_cache_grp->next_scan); if (time_before(now, next_scan)) return; =20 /* only 1 thread is allowed to scan */ - if (!try_cmpxchg(&mm->sc_stat.next_scan, &next_scan, + if (!try_cmpxchg(&mm->sched_cache_grp->next_scan, &next_scan, now + max_t(unsigned long, READ_ONCE(llc_epoch_period), 1))) return; @@ -1893,8 +1922,8 @@ static void task_cache_work(struct callback_head *wor= k) curr_cpu =3D task_cpu(p); if (invalid_llc_nr(mm, p, curr_cpu) || exceed_llc_capacity(mm, curr_cpu)) { - if (READ_ONCE(mm->sc_stat.cpu) !=3D -1) - WRITE_ONCE(mm->sc_stat.cpu, -1); + if (READ_ONCE(mm->sched_cache_grp->cpu) !=3D -1) + WRITE_ONCE(mm->sched_cache_grp->cpu, -1); =20 return; } @@ -1917,8 +1946,10 @@ static void task_cache_work(struct callback_head *wo= rk) continue; =20 for_each_cpu(i, sched_domain_span(sd)) { + struct sched_cache_group *grp =3D mm->sched_cache_grp; + occ =3D fraction_mm_sched(cpu_rq(i), - per_cpu_ptr(mm->sc_stat.pcpu_sched, i)); + per_cpu_ptr(grp->pcpu_sched, i)); a_occ +=3D occ; if (occ > m_occ) { m_occ =3D occ; @@ -1951,7 +1982,7 @@ static void task_cache_work(struct callback_head *wor= k) m_a_cpu =3D m_cpu; } =20 - if (llc_id(cpu) =3D=3D llc_id(READ_ONCE(mm->sc_stat.cpu))) + if (llc_id(cpu) =3D=3D llc_id(READ_ONCE(mm->sched_cache_grp->cpu))) curr_m_a_occ =3D a_occ; =20 cpumask_andnot(cpus, cpus, sched_domain_span(sd)); @@ -1960,7 +1991,7 @@ static void task_cache_work(struct callback_head *wor= k) =20 if (m_a_occ > (2 * curr_m_a_occ)) { /* - * Avoid switching sc_stat.cpu too fast. + * Avoid switching sched_cache_grp->cpu too fast. * The reason to choose 2X is because: * 1. It is better to keep the preferred LLC stable, * rather than changing it frequently and cause migrations @@ -1969,10 +2000,10 @@ static void task_cache_work(struct callback_head *w= ork) * 3. 2X is chosen based on test results, as it delivers * the optimal performance gain so far. */ - WRITE_ONCE(mm->sc_stat.cpu, m_a_cpu); + WRITE_ONCE(mm->sched_cache_grp->cpu, m_a_cpu); } =20 - update_avg_scale(&mm->sc_stat.nr_running_avg, nr_running); + update_avg_scale(&mm->sched_cache_grp->nr_running_avg, nr_running); free_cpumask_var(cpus); } =20 @@ -3776,18 +3807,19 @@ static void task_numa_placement(struct task_struct = *p) * heuristic and occasional lost updates are tolerable. * * If a task exits, its corresponding footprint must - * be subtracted from the mm->sc_stat.footprint, otherwise - * the mm->sc_stat.footprint will not converge: - * the exiting thread's footprint remains unchanged/undecayed - * in mm->sc_stat.footprint. See exit_mm(). + * be subtracted from the mm->sched_cache_grp->footprint, + * otherwise the mm->sched_cache_grp->footprint will not + * converge: the exiting thread's footprint remains + * unchanged/undecayed in mm->sched_cache_grp->footprint. + * See exit_mm(). * * Lost updates and unsynchronized subtraction * in exit_mm() can cause footprint + diff to * go negative. Clamp to zero to prevent the * unsigned footprint from wrapping. */ - new_fp =3D (long)READ_ONCE(p->mm->sc_stat.footprint) + diff; - WRITE_ONCE(p->mm->sc_stat.footprint, + new_fp =3D (long)READ_ONCE(p->mm->sched_cache_grp->footprint) + diff; + WRITE_ONCE(p->mm->sched_cache_grp->footprint, max(new_fp, 0L)); #endif } @@ -10703,18 +10735,18 @@ static enum llc_mig can_migrate_llc_task(int src_= cpu, int dst_cpu, int cpu; =20 mm =3D p->mm; - if (!mm) + if (!mm || !mm->sched_cache_grp) return mig_unrestricted; =20 - cpu =3D READ_ONCE(mm->sc_stat.cpu); + cpu =3D READ_ONCE(mm->sched_cache_grp->cpu); if (cpu < 0 || cpus_share_cache(src_cpu, dst_cpu)) return mig_unrestricted; =20 /* skip cache aware load balance for too many threads */ if (invalid_llc_nr(mm, p, dst_cpu) || exceed_llc_capacity(mm, dst_cpu)) { - if (READ_ONCE(mm->sc_stat.cpu) !=3D -1) - WRITE_ONCE(mm->sc_stat.cpu, -1); + if (READ_ONCE(mm->sched_cache_grp->cpu) !=3D -1) + WRITE_ONCE(mm->sched_cache_grp->cpu, -1); return mig_unrestricted; } =20 --=20 2.32.0 From nobody Sat Sep 26 11:46:28 2026 Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.8]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AEA351DA23 for ; Wed, 2 Sep 2026 00:04:01 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=192.198.163.8 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788307444; cv=none; b=Q1vHrY+zmULXpEz9wv/CresLXDES2HYaM2Q0DmZBFJamwj/ZMzmeynN+TE6MUjWEdNwuxO4ieEmFVWELPV7SRe6ohlH5G3+5jhOW+Zdr37eSBurJAWa84S88ol83S6MknZNAtUXlmDu4JfvDdSXQY+Pr8ySZkWx2wweOMcmOqXE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788307444; c=relaxed/simple; bh=hYvTU1//GIKznZB6qx0LqpWLR5jjVcFMotNPjEt1Klg=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=T7mnXkx82wYJkE3P4o773Rq5Zx+tPheM16KWreUVx8t3k1amaX4T0qGWT47XVSF+8PNBgx/Cvn1ZbmaD1luTdCPDinnWC+SwfW1ZWCdr6n8+nnGDNuedtgJL3tr565hvh0zTU8ErUI6UvEsF/P2zrLHGk9Rhnv2XPrrox6SkwUU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com; spf=pass smtp.mailfrom=linux.intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=UGzVQwo3; arc=none smtp.client-ip=192.198.163.8 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="UGzVQwo3" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1788307442; x=1819843442; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=hYvTU1//GIKznZB6qx0LqpWLR5jjVcFMotNPjEt1Klg=; b=UGzVQwo35QP1MlW888O6yDwwrt7awXl++XhSMAGoFzTO/cX6CBM/x7V3 l9K0K+YXAtSbBq+r7ueFrvBSNZ7c8+GnSYzvjtn8Ho4DLcE5rTdobtuBx 9n/QOPLHWudhxAIK/HDwtK5ypEj04dfNI8/izlzYuV8FJCkQFt54o5Ixq 7RzH9QQPxgXstYNR8jMz87Jdk+UfaTIH6kFeTlsXkxjfPngZvCn/5/lYj yU/QbmGgK/r5ujfzk1EJ5Pi10cyZ9Q4KeXzllP6oVnuL6GCfxthJvl6b4 /l+qQS/OVfFRVEmniiYsR0vZ3gTXz21t3hESsrgcqP6WnZXC8z4rmfRQi Q==; X-CSE-ConnectionGUID: z0Jy2U5aSWazVsn/5S6tOQ== X-CSE-MsgGUID: xvgeGZs4T+uAa8iqXk5xwQ== X-IronPort-AV: E=McAfee;i="6800,10657,11893"; a="106272963" X-IronPort-AV: E=Sophos;i="6.25,256,1779174000"; d="scan'208";a="106272963" Received: from fmviesa010.fm.intel.com ([10.60.135.150]) by fmvoesa102.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 01 Sep 2026 17:03:57 -0700 X-CSE-ConnectionGUID: Abi1DXElS7aY064lABRYYg== X-CSE-MsgGUID: SMcmwTVlToS56ClLA5mEcQ== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.25,256,1779174000"; d="scan'208";a="265531794" Received: from b04f130c83f2.jf.intel.com ([10.165.154.98]) by fmviesa010.fm.intel.com with ESMTP; 01 Sep 2026 17:03:45 -0700 From: Tim Chen To: Peter Zijlstra , Ingo Molnar Cc: Tim Chen , Chen Yu , Hyunwoo Kim , Kees Cook , Christian Brauner , Alexander Viro , Jan Kara , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , K Prateek Nayak , Shrikanth Hegde , Qais Yousef , Aaron Lu , Srikar Dronamraju , Vineeth Remanan Pillai , linux-kernel@vger.kernel.org, linux-mm@kvack.org, "chen . yu @ linux . dev" Subject: [PATCH 2/2] sched/cache: Introduce task_struct->sched_cache_grp Date: Tue, 1 Sep 2026 17:08:56 -0700 Message-Id: X-Mailer: git-send-email 2.32.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add a sched_cache_grp pointer to task_struct so that scheduler code can access the cache group directly via the task, without going through mm->sched_cache_grp. This decouples the scheduler's hot-path accesses from the mm_struct. Each task holds its own refcount on the sched_cache_group, separate from the reference held by its mm_struct. The reference is acquired in copy_mm() (fork) and exec_mmap() (exec), and released in exit_mm(). This fixes use after free problem when accessing sched_cache_grp in account_mm_sched() via mm as reported in https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/ Convert all scheduler code in fair.c and exit.c to use p->sched_cache_grp instead of p->mm->sched_cache_grp. Add sched_cache_group_get() to kernel/sched/cache_sched.c. Reported-by: Hyunwoo Kim Closes: https://lore.kernel.org/lkml/apPb-Dr4nPYuHQOK@v4bel/ Tested-by: Hyunwoo Kim Fixes: df0d98475954 ("sched/cache: Introduce infrastructure for cache-aware= load balancing") Co-developed-by: Chen Yu Signed-off-by: Chen Yu Signed-off-by: Tim Chen --- fs/exec.c | 14 ++++ include/linux/sched.h | 3 + kernel/exit.c | 26 +++++-- kernel/fork.c | 23 ++++++ kernel/sched/cache_sched.c | 19 +++++ kernel/sched/fair.c | 142 +++++++++++++++++++++---------------- kernel/sched/sched.h | 3 + 7 files changed, 164 insertions(+), 66 deletions(-) diff --git a/fs/exec.c b/fs/exec.c index 745f6eb5279e..7a8a9954343e 100644 --- a/fs/exec.c +++ b/fs/exec.c @@ -882,6 +882,20 @@ static int exec_mmap(struct linux_binprm *bprm) active_mm =3D tsk->active_mm; tsk->active_mm =3D mm; tsk->mm =3D mm; +#ifdef CONFIG_SCHED_CACHE + { + struct sched_cache_group *old_grp, *new_grp; + + old_grp =3D rcu_dereference_protected(tsk->sched_cache_grp, true); + + /* Acquire the reference before publishing the pointer. */ + new_grp =3D sched_cache_group_get(mm->sched_cache_grp); + + rcu_assign_pointer(tsk->sched_cache_grp, new_grp); + if (old_grp) + sched_cache_group_put(old_grp); + } +#endif mm_init_cid(mm, tsk); exec_state =3D task_exec_state_replace(tsk, exec_state); /* diff --git a/include/linux/sched.h b/include/linux/sched.h index 1f254364f216..cab8e89b1462 100644 --- a/include/linux/sched.h +++ b/include/linux/sched.h @@ -1434,6 +1434,7 @@ struct task_struct { #ifdef CONFIG_SCHED_CACHE struct callback_head cache_work; int preferred_llc; + struct sched_cache_group __rcu *sched_cache_grp; /* 1: task was enqueued to its preferred LLC, 0 otherwise */ int pref_llc_queued; #endif @@ -2418,6 +2419,8 @@ struct sched_cache_group { } ____cacheline_aligned_in_smp; =20 void sched_cache_group_put(struct sched_cache_group *grp); +struct sched_cache_group *sched_cache_group_get(struct sched_cache_group *= grp); +struct sched_cache_group *task_cache_group_get(struct task_struct *p); =20 #else =20 diff --git a/kernel/exit.c b/kernel/exit.c index 006edcc0c2c5..442535778ce1 100644 --- a/kernel/exit.c +++ b/kernel/exit.c @@ -552,23 +552,25 @@ void mm_update_next_owner(struct mm_struct *mm) * Subtract the memory footprint of the current task from * mm. */ -static void exit_mm_sched_cache(struct mm_struct *mm) +static void exit_mm_sched_cache(void) { + struct sched_cache_group *grp =3D + rcu_dereference_protected(current->sched_cache_grp, true); unsigned long fp, sub; =20 - if (!current->total_numa_faults) + if (!grp || !current->total_numa_faults) return; /* * No lock protection due to performance considerations. * Make sure the group footprint does not become * negative. */ - fp =3D READ_ONCE(mm->sched_cache_grp->footprint); + fp =3D READ_ONCE(grp->footprint); sub =3D min(fp, current->total_numa_faults); - WRITE_ONCE(mm->sched_cache_grp->footprint, fp - sub); + WRITE_ONCE(grp->footprint, fp - sub); } #else -static inline void exit_mm_sched_cache(struct mm_struct *mm) +static inline void exit_mm_sched_cache(void) { } #endif /* CONFIG_SCHED_CACHE CONFIG_NUMA_BALANCING */ @@ -585,7 +587,19 @@ static void exit_mm(void) if (!mm) return; =20 - exit_mm_sched_cache(mm); + exit_mm_sched_cache(); + +#ifdef CONFIG_SCHED_CACHE + { + struct sched_cache_group *grp =3D + rcu_dereference_protected(current->sched_cache_grp, true); + + rcu_assign_pointer(current->sched_cache_grp, NULL); + + if (grp) + sched_cache_group_put(grp); + } +#endif =20 mmap_read_lock(mm); mmgrab_lazy_tlb(mm); diff --git a/kernel/fork.c b/kernel/fork.c index 416758c8a3d4..2e79548cb7c1 100644 --- a/kernel/fork.c +++ b/kernel/fork.c @@ -1599,6 +1599,19 @@ static int copy_mm(u64 clone_flags, struct task_stru= ct *tsk) =20 tsk->mm =3D mm; tsk->active_mm =3D mm; +#ifdef CONFIG_SCHED_CACHE + { + /* + * A task holds its own reference on the group, separate from + * the reference held by its mm_struct. Acquire it before + * publishing the pointer. + */ + struct sched_cache_group *grp =3D + sched_cache_group_get(mm->sched_cache_grp); + + rcu_assign_pointer(tsk->sched_cache_grp, grp); + } +#endif return 0; } =20 @@ -2599,6 +2612,16 @@ __latent_entropy struct task_struct *copy_process( bad_fork_cleanup_namespaces: exit_nsproxy_namespaces(p); bad_fork_cleanup_mm: +#ifdef CONFIG_SCHED_CACHE + /* + * copy_mm() took a task reference on the cache group; a failed fork + * never reaches exit_mm(), so release it here to avoid leaking the + * group and its per-CPU buffer. + */ + sched_cache_group_put(rcu_dereference_protected(p->sched_cache_grp, true)= ); + RCU_INIT_POINTER(p->sched_cache_grp, NULL); +#endif + if (p->mm) { mm_clear_owner(p->mm, p); mmput(p->mm); diff --git a/kernel/sched/cache_sched.c b/kernel/sched/cache_sched.c index d492df55f9d5..99d07e1e067c 100644 --- a/kernel/sched/cache_sched.c +++ b/kernel/sched/cache_sched.c @@ -1,6 +1,25 @@ // SPDX-License-Identifier: GPL-2.0-only #include "sched.h" =20 +struct sched_cache_group *sched_cache_group_get(struct sched_cache_group *= grp) +{ + /* + * refcount_inc_not_zero() is the acquire primitive for lockless + * (RCU) lookups; plain refcount_inc() would scribble the count if + * it already reached zero. Return NULL in that case. + */ + if (grp && !refcount_inc_not_zero(&grp->refcnt)) + grp =3D NULL; + + return grp; +} + +struct sched_cache_group *task_cache_group_get(struct task_struct *p) +{ + guard(rcu)(); + return sched_cache_group_get(rcu_dereference(p->sched_cache_grp)); +} + static void sched_cache_group_free_rcu(struct rcu_head *rcu) { struct sched_cache_group *grp =3D diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 8587dcbaa1cf..4226af32728a 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -1470,7 +1470,7 @@ static inline int get_sched_cache_scale(int mul) return (1 + (tol - 1) * mul); } =20 -static bool exceed_llc_capacity(struct mm_struct *mm, int cpu) +static bool exceed_llc_capacity(struct sched_cache_group *grp, int cpu) { #ifdef CONFIG_NUMA_BALANCING unsigned long llc, footprint; @@ -1489,7 +1489,7 @@ static bool exceed_llc_capacity(struct mm_struct *mm,= int cpu) * excluded. */ llc =3D sd->llc_bytes; - footprint =3D READ_ONCE(mm->sched_cache_grp->footprint); + footprint =3D READ_ONCE(grp->footprint); =20 /* * Scale the LLC size by 256*llc_aggr_tolerance @@ -1518,7 +1518,7 @@ static bool exceed_llc_capacity(struct mm_struct *mm,= int cpu) return false; } =20 -static bool invalid_llc_nr(struct mm_struct *mm, struct task_struct *p, +static bool invalid_llc_nr(struct sched_cache_group *grp, struct task_stru= ct *p, int cpu) { int scale; @@ -1534,7 +1534,7 @@ static bool invalid_llc_nr(struct mm_struct *mm, stru= ct task_struct *p, if (scale =3D=3D INT_MAX) return false; =20 - return !fits_capacity((mm->sched_cache_grp->nr_running_avg * cpu_smt_num_= threads), + return !fits_capacity((grp->nr_running_avg * cpu_smt_num_threads), (scale * per_cpu(sd_llc_size, cpu))); } =20 @@ -1714,14 +1714,14 @@ static unsigned long fraction_mm_sched(struct rq *r= q, return div64_u64(NICE_0_LOAD * pcpu_sched->runtime, rq->cpu_runtime + 1); } =20 -static int get_pref_llc(struct task_struct *p, struct mm_struct *mm) +static int get_pref_llc(struct task_struct *p, struct sched_cache_group *g= rp) { int mm_sched_llc =3D -1, mm_sched_cpu; =20 - if (!mm) + if (!grp) return -1; =20 - mm_sched_cpu =3D READ_ONCE(mm->sched_cache_grp->cpu); + mm_sched_cpu =3D READ_ONCE(grp->cpu); if (mm_sched_cpu !=3D -1) { mm_sched_llc =3D llc_id(mm_sched_cpu); =20 @@ -1751,8 +1751,8 @@ static unsigned int task_running_on_cpu(int cpu, stru= ct task_struct *p); static inline void account_mm_sched(struct rq *rq, struct task_struct *p, s64 delta_exec) { + struct sched_cache_group *grp =3D rcu_dereference_all(p->sched_cache_grp); struct sched_cache_time *pcpu_sched; - struct mm_struct *mm =3D p->mm; int mm_sched_llc =3D -1; unsigned long epoch; =20 @@ -1763,16 +1763,12 @@ void account_mm_sched(struct rq *rq, struct task_st= ruct *p, s64 delta_exec) return; /* * init_task, kthreads and user thread created - * by user_mode_thread() don't have mm. - * - * A kthread can temporarily adopt an mm via kthread_use_mm(), - * so p->mm alone does not imply a user task. + * by user_mode_thread() don't have a cache group. */ - if (!mm || p->flags & PF_KTHREAD || !mm->sched_cache_grp || - !mm->sched_cache_grp->pcpu_sched) + if (!grp || p->flags & PF_KTHREAD || !grp->pcpu_sched) return; =20 - pcpu_sched =3D per_cpu_ptr(mm->sched_cache_grp->pcpu_sched, cpu_of(rq)); + pcpu_sched =3D per_cpu_ptr(grp->pcpu_sched, cpu_of(rq)); =20 scoped_guard (raw_spinlock, &rq->cpu_epoch_lock) { __update_mm_sched(rq, pcpu_sched); @@ -1785,14 +1781,14 @@ void account_mm_sched(struct rq *rq, struct task_st= ruct *p, s64 delta_exec) * If this process hasn't hit task_cache_work() for a while invalidate * its preferred state. */ - if ((long)(epoch - READ_ONCE(mm->sched_cache_grp->epoch)) > llc_epoch_aff= inity_timeout || - invalid_llc_nr(mm, p, cpu_of(rq)) || - exceed_llc_capacity(mm, cpu_of(rq))) { - if (READ_ONCE(mm->sched_cache_grp->cpu) !=3D -1) - WRITE_ONCE(mm->sched_cache_grp->cpu, -1); + if ((long)(epoch - READ_ONCE(grp->epoch)) > llc_epoch_affinity_timeout || + invalid_llc_nr(grp, p, cpu_of(rq)) || + exceed_llc_capacity(grp, cpu_of(rq))) { + if (READ_ONCE(grp->cpu) !=3D -1) + WRITE_ONCE(grp->cpu, -1); } =20 - mm_sched_llc =3D get_pref_llc(p, mm); + mm_sched_llc =3D get_pref_llc(p, grp); =20 /* task not on rq accounted later in account_entity_enqueue() */ if (task_running_on_cpu(rq->cpu, p) && @@ -1805,31 +1801,32 @@ void account_mm_sched(struct rq *rq, struct task_st= ruct *p, s64 delta_exec) =20 static void task_tick_cache(struct rq *rq, struct task_struct *p) { + struct sched_cache_group *grp =3D rcu_dereference_all(p->sched_cache_grp); struct callback_head *work =3D &p->cache_work; - struct mm_struct *mm =3D p->mm; unsigned long epoch; =20 if (!sched_cache_enabled()) return; =20 - if (!mm || p->flags & PF_KTHREAD || - !mm->sched_cache_grp->pcpu_sched) + if (!grp || p->flags & PF_KTHREAD || + !grp->pcpu_sched) return; =20 epoch =3D rq->cpu_epoch; /* avoid moving backwards */ - if (time_after_eq(mm->sched_cache_grp->epoch, epoch)) + if (time_after_eq(grp->epoch, epoch)) return; =20 - guard(raw_spinlock)(&mm->sched_cache_grp->lock); + guard(raw_spinlock)(&grp->lock); =20 if (work->next =3D=3D work) { task_work_add(p, work, TWA_RESUME); - WRITE_ONCE(mm->sched_cache_grp->epoch, epoch); + WRITE_ONCE(grp->epoch, epoch); } } =20 -static void get_scan_cpumasks(cpumask_var_t cpus, struct task_struct *p) +static void get_scan_cpumasks(cpumask_var_t cpus, struct task_struct *p, + struct sched_cache_group *grp) { #ifdef CONFIG_NUMA_BALANCING int cpu, curr_cpu, nid, pref_nid; @@ -1837,7 +1834,7 @@ static void get_scan_cpumasks(cpumask_var_t cpus, str= uct task_struct *p) if (!static_branch_likely(&sched_numa_balancing)) goto out; =20 - cpu =3D READ_ONCE(p->mm->sched_cache_grp->cpu); + cpu =3D READ_ONCE(grp->cpu); if (cpu !=3D -1) nid =3D cpu_to_node(cpu); curr_cpu =3D task_cpu(p); @@ -1898,9 +1895,7 @@ static void task_cache_work(struct callback_head *wor= k) unsigned long next_scan, now =3D jiffies; struct task_struct *p =3D current, *cur; unsigned long curr_m_a_occ =3D 0; - struct mm_struct *mm =3D p->mm; unsigned long m_a_occ =3D 0; - cpumask_var_t cpus; =20 WARN_ON_ONCE(work !=3D &p->cache_work); =20 @@ -1909,32 +1904,44 @@ static void task_cache_work(struct callback_head *w= ork) if (p->flags & PF_EXITING) return; =20 - next_scan =3D READ_ONCE(mm->sched_cache_grp->next_scan); + /* + * A reference makes sure grp is not released by others. The rcu + * lock can not be held till after zalloc_cpumask_var() below, + * because the latter might sleep. + */ + struct sched_cache_group *grp __free(sched_cache_group_put) =3D + task_cache_group_get(p); + if (!grp) + return; + + next_scan =3D READ_ONCE(grp->next_scan); if (time_before(now, next_scan)) return; =20 /* only 1 thread is allowed to scan */ - if (!try_cmpxchg(&mm->sched_cache_grp->next_scan, &next_scan, + if (!try_cmpxchg(&grp->next_scan, &next_scan, now + max_t(unsigned long, READ_ONCE(llc_epoch_period), 1))) return; =20 curr_cpu =3D task_cpu(p); - if (invalid_llc_nr(mm, p, curr_cpu) || - exceed_llc_capacity(mm, curr_cpu)) { - if (READ_ONCE(mm->sched_cache_grp->cpu) !=3D -1) - WRITE_ONCE(mm->sched_cache_grp->cpu, -1); + if (invalid_llc_nr(grp, p, curr_cpu) || + exceed_llc_capacity(grp, curr_cpu)) { + if (READ_ONCE(grp->cpu) !=3D -1) + WRITE_ONCE(grp->cpu, -1); =20 return; } =20 + cpumask_var_t cpus __free(free_cpumask_var) =3D CPUMASK_VAR_NULL; + if (!zalloc_cpumask_var(&cpus, GFP_KERNEL)) return; =20 scoped_guard (cpus_read_lock) { guard(rcu)(); =20 - get_scan_cpumasks(cpus, p); + get_scan_cpumasks(cpus, p, grp); =20 for_each_cpu(cpu, cpus) { /* XXX sched_cluster_active */ @@ -1946,8 +1953,6 @@ static void task_cache_work(struct callback_head *wor= k) continue; =20 for_each_cpu(i, sched_domain_span(sd)) { - struct sched_cache_group *grp =3D mm->sched_cache_grp; - occ =3D fraction_mm_sched(cpu_rq(i), per_cpu_ptr(grp->pcpu_sched, i)); a_occ +=3D occ; @@ -1956,9 +1961,13 @@ static void task_cache_work(struct callback_head *wo= rk) m_cpu =3D i; } =20 + /* + * rcu_access_pointer() is used because the + * pointer is only compared, never dereferenced. + */ cur =3D rcu_dereference_all(cpu_rq(i)->curr); if (cur && !(cur->flags & (PF_EXITING | PF_KTHREAD)) && - cur->mm =3D=3D mm) + rcu_access_pointer(cur->sched_cache_grp) =3D=3D grp) nr_running++; } =20 @@ -1982,7 +1991,7 @@ static void task_cache_work(struct callback_head *wor= k) m_a_cpu =3D m_cpu; } =20 - if (llc_id(cpu) =3D=3D llc_id(READ_ONCE(mm->sched_cache_grp->cpu))) + if (llc_id(cpu) =3D=3D llc_id(READ_ONCE(grp->cpu))) curr_m_a_occ =3D a_occ; =20 cpumask_andnot(cpus, cpus, sched_domain_span(sd)); @@ -2000,11 +2009,10 @@ static void task_cache_work(struct callback_head *w= ork) * 3. 2X is chosen based on test results, as it delivers * the optimal performance gain so far. */ - WRITE_ONCE(mm->sched_cache_grp->cpu, m_a_cpu); + WRITE_ONCE(grp->cpu, m_a_cpu); } =20 - update_avg_scale(&mm->sched_cache_grp->nr_running_avg, nr_running); - free_cpumask_var(cpus); + update_avg_scale(&grp->nr_running_avg, nr_running); } =20 void init_sched_mm(struct task_struct *p) @@ -2013,6 +2021,13 @@ void init_sched_mm(struct task_struct *p) =20 init_task_work(work, task_cache_work); work->next =3D work; + /* + * dup_task_struct() copies the parent's task_struct, including its + * sched_cache_grp, for which the child holds no reference. Clear it + * here - before copy_mm() runs - so the child never carries a + * borrowed pointer that the fork error path would put. + */ + RCU_INIT_POINTER(p->sched_cache_grp, NULL); /* * Reset new task's preference to avoid * polluting account_llc_enqueue(). @@ -3807,10 +3822,9 @@ static void task_numa_placement(struct task_struct *= p) * heuristic and occasional lost updates are tolerable. * * If a task exits, its corresponding footprint must - * be subtracted from the mm->sched_cache_grp->footprint, - * otherwise the mm->sched_cache_grp->footprint will not - * converge: the exiting thread's footprint remains - * unchanged/undecayed in mm->sched_cache_grp->footprint. + * be subtracted from p->sched_cache_grp->footprint, + * otherwise the footprint will not converge: the + * exiting thread's footprint remains unchanged/undecayed. * See exit_mm(). * * Lost updates and unsynchronized subtraction @@ -3818,9 +3832,17 @@ static void task_numa_placement(struct task_struct *= p) * go negative. Clamp to zero to prevent the * unsigned footprint from wrapping. */ - new_fp =3D (long)READ_ONCE(p->mm->sched_cache_grp->footprint) + diff; - WRITE_ONCE(p->mm->sched_cache_grp->footprint, - max(new_fp, 0L)); + { + struct sched_cache_group *grp; + + guard(rcu)(); + grp =3D rcu_dereference(p->sched_cache_grp); + + if (grp) { + new_fp =3D (long)READ_ONCE(grp->footprint) + diff; + WRITE_ONCE(grp->footprint, max(new_fp, 0L)); + } + } #endif } =20 @@ -10730,23 +10752,23 @@ static enum llc_mig can_migrate_llc(int src_cpu, = int dst_cpu, static enum llc_mig can_migrate_llc_task(int src_cpu, int dst_cpu, struct task_struct *p) { - struct mm_struct *mm; + struct sched_cache_group *grp; bool to_pref; int cpu; =20 - mm =3D p->mm; - if (!mm || !mm->sched_cache_grp) + grp =3D rcu_dereference_all(p->sched_cache_grp); + if (!grp) return mig_unrestricted; =20 - cpu =3D READ_ONCE(mm->sched_cache_grp->cpu); + cpu =3D READ_ONCE(grp->cpu); if (cpu < 0 || cpus_share_cache(src_cpu, dst_cpu)) return mig_unrestricted; =20 /* skip cache aware load balance for too many threads */ - if (invalid_llc_nr(mm, p, dst_cpu) || - exceed_llc_capacity(mm, dst_cpu)) { - if (READ_ONCE(mm->sched_cache_grp->cpu) !=3D -1) - WRITE_ONCE(mm->sched_cache_grp->cpu, -1); + if (invalid_llc_nr(grp, p, dst_cpu) || + exceed_llc_capacity(grp, dst_cpu)) { + if (READ_ONCE(grp->cpu) !=3D -1) + WRITE_ONCE(grp->cpu, -1); return mig_unrestricted; } =20 diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h index e656c7059bf8..8b67af28a471 100644 --- a/kernel/sched/sched.h +++ b/kernel/sched/sched.h @@ -4145,6 +4145,9 @@ static inline bool sched_cache_enabled(void) return static_branch_unlikely(&sched_cache_active); } =20 +DEFINE_FREE(sched_cache_group_put, struct sched_cache_group *, + sched_cache_group_put(_T)); + extern void sched_cache_active_set(void); =20 #endif --=20 2.32.0