From nobody Fri Jul 24 23:35:55 2026 Received: from out203-205-221-164.mail.qq.com (out203-205-221-164.mail.qq.com [203.205.221.164]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CC9713E49ED; Wed, 22 Jul 2026 09:09:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=203.205.221.164 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784711403; cv=none; b=YOoh3vbmLTpD2wIE1Wjk/4UwnBO42xK6uZH2Vs1/2jJcved8clPfzKXcku/E27Oaa6gC/J+3nyCzcyMhNfu8ndThaobWKIUYurqdV6ja8bcEe4e5IsIyNGP1InAUOJjOZ21Z3vTF+6kqEzCe/N+wdowALA2/qo8anw+ale4wEkY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784711403; c=relaxed/simple; bh=0r0C9D0mfxL2AiUlPsM0pyeLcwtNb3baZZNoy/w5ubw=; h=Message-ID:From:To:Cc:Subject:Date:In-Reply-To:References: MIME-Version; b=pFKWqPANnTJHTdvmixL31s30PGHVBJm/VAVI0SYstPBx5L0+qzwKSXcv2dIpxvI2x5AseWHKEc1rDzvImDGZ/O2OEHuyIkNhwp98CMuREcZZd9RChtH3csNLlrCSqVslymZSUXYvDa6QMVwyzUE2fSffeLdX4CKoFCWsk80Mk5k= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=cyyself.name; spf=pass smtp.mailfrom=cyyself.name; dkim=pass (1024-bit key) header.d=qq.com header.i=@qq.com header.b=y1q1Q1sk; arc=none smtp.client-ip=203.205.221.164 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=cyyself.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=cyyself.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=qq.com header.i=@qq.com header.b="y1q1Q1sk" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=qq.com; s=s201512; t=1784711389; bh=h6dWYlj5CbfoQlGCMPOx9K1HpD8J9nXPEjfkkiJDMMY=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=y1q1Q1skeNPCmLWXT8lIz4/mm+aYwJ+0dXV6/mavbOX7AJ2AEpz6eJVnr0h+R2YQz DKgEUq3swAdqT5CXaB1ylO3a8zYGzZYe3BqP5+VupN8+pwFNRsDn639i/LBweQ7XAn UKh2qIOTPJ6laVDkQ0v71mEEiEyexUzDKclnge9Q= Received: from localhost.localdomain ([240e:37c:2242:cf00:5054:a3ff:fe89:345b]) by newxmesmtplogicsvrszc56-0.qq.com (NewEsmtp) with SMTP id 26DB9E5A; Wed, 22 Jul 2026 17:09:45 +0800 X-QQ-mid: xmsmtpt1784711386t4eyhdya9 Message-ID: X-QQ-XMAILINFO: MOnz+xTS1+9inM1CBaz9W4ULK2eHsYaTpl+qMM4rB9HaHyTaXyiYxWcrHx3P4W w8ibug0boTt1+JAsa3q+Ay3V8jFSDqouCOD3odhB2Jf4GdxnK0gTFuSKINRNG0RE+a4yjGb64LaL 7TVV8NlWS8UHavVR7fVAyH0BdMkTS0XwHHl/xdRIskyUyXwxxB8/Ngkp2PFKKb8v5yEvk1RWm7fN qf70ySncbILTqO8lsrdvxz8CENi9vJ+vCIKNIwFoe4sDHn4XPU/+RMM9E1+s/7jCsuAAWdJqP9FP Z4u8dp46/1O9mBaIqkT7aP28TMrkgUr76B+oSq8AUtTcV8WNf5codfj+RhueuI/DjO/8MpEWC2KU eCh8JA7a3JhzYW382ZqB9wtmDQMwHuQYR8NKCP0aUdTCbDdSJU5VjdaTNUypSVmknG4ZwsXSuK6X L51Xak618YWK90TQsHjgZxiw31H+5kRNMPh0MP8++MpR2EFMqe/F0A1sXUgSWKOzTuzDA54dBAKl YJFFsoT9oWuSSR/8t7mE6M/GSvgU4tJFEN3pJUCra8XA6rR64Niq4f0my+K7IRBVSH827oWILtI/ 068sI0ToLBF/zQ7FP8Hksd3slzQcR3/caKsBlWDvHuZxeWaZMHQxA333CYzHfocoXUp059MxG8Nm H465u5/la36AznSQ4wNyET7DWuecV8yGPrtw6KYR5K+1B+bqUZ4bGQXcCkH3v/lG1OjgbbIQxid0 6HmZuuj5g+figdz1jR16dJRPN7jBswCrnzO3dNCBjICMbv+KtfbnffhjPMaKEr5torUF8U1og6VO y3rpEBwJVUjwrfWNldWmzpRwA//3ND78IBY6k/YKdAoArCIUUVuEvLEv/Td7fuEN7sIhPqIezDjr JyOJg83HuMG+NYADUfOBJSaGAogogLieqN0ihVOfLitSVRNHVOPOLocgLaD+kM39Ar+MvPYZ+mb4 lSYNz+Tt0qWD3XF45U6Ju6zlQioZcuRWxCrr1jW6OncO62d/q7INRdmhN8psvI0eXc+uhKE3gbMW 1RzSydaAma6OhR2JP7h40mghxf8F4eljOFyucqaH2sW+XncNcB2Iyn7OeylufuFzTMeHFKodJRUW 5skOgs X-QQ-XMRINFO: NS+P29fieYNwqS3WCnRCOn9D1NpZuCnCRA== From: Yangyu Chen To: Peter Zijlstra , Ingo Molnar , Juri Lelli , Vincent Guittot Cc: Chen Yu , Tim Chen , K Prateek Nayak , Dietmar Eggemann , Valentin Schneider , Jonathan Corbet , Shuah Khan , Yangyu Chen , linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-doc@vger.kernel.org, Yangyu Chen Subject: [PATCH 1/4] sched/cache: Split llc_aggr_tolerance into nr and size tolerances Date: Wed, 22 Jul 2026 17:09:44 +0800 X-OQ-MSGID: <20260722090944.1591886-1-cyy@cyyself.name> X-Mailer: git-send-email 2.47.3 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" llc_aggr_tolerance scales two independent gates with a single value: the number of cores in an LLC compared against the process's active thread count in invalid_llc_nr(), and the LLC size compared against the process's memory footprint in exceed_llc_capacity(). The right settings are workload and platform specific. A Verilator run is one example: its RSS is large while only a small part of it is hot, so an RSS-based footprint estimate overstates its cache usage and should not decide whether it is aggregated. And running its threads on the SMT siblings of a single LLC can beat spreading them across LLCs on some platforms (e.g. AMD EPYC Turin) but not on others (e.g. EPYC Milan). Tuning for such a workload means a larger thread-count tolerance while the footprint tolerance keeps its own setting - which a single combined knob cannot express. Split the knob into llc_aggr_tolerance_nr for the thread-count gate and llc_aggr_tolerance_size for the footprint gate, and let get_sched_cache_scale() take the tolerance value from the caller. Both new debugfs files keep the semantics of the old knob (0 disables aggregation for that gate, values >=3D 100 mean unlimited). While making the tunables independently settable, also make the percentage knobs safe for any value: llc_overaggr_pct and llc_imb_pct are exposed via debugfs as unbounded u32 values and feed percentage multiplications: util * 100 < max * aggr_pct (fits_llc_capacity()) util1 * 100 > util2 * (100 + llc_imb_pct) (util_greater()) CONFIG_SCHED_CACHE only depends on SMP, so this also builds on 32-bit where unsigned long is 32 bits: a large enough percentage makes max * aggr_pct (or util2 * (100 + pct)) wrap and produce a garbage comparison, and the SMT-1 bump aggr_pct * 3 / 2 can wrap the u32 itself. Rather than capping the tunables, detect the overflow at the point of use with check_mul_overflow()/check_add_overflow() and fall back to the saturated meaning of the comparison: an overflowing threshold is effectively unlimited (the LLC always has room), and an overflowing bias is effectively infinite (the destination utilization is never considered noticeably greater). Assisted-by: Claude:claude-fable-5 Signed-off-by: Yangyu Chen --- kernel/sched/debug.c | 6 +++-- kernel/sched/fair.c | 62 +++++++++++++++++++++++++++++++------------- kernel/sched/sched.h | 3 ++- 3 files changed, 50 insertions(+), 21 deletions(-) diff --git a/kernel/sched/debug.c b/kernel/sched/debug.c index 40584b27ea0c..b7be245b9fc2 100644 --- a/kernel/sched/debug.c +++ b/kernel/sched/debug.c @@ -672,8 +672,10 @@ static __init int sched_init_debug(void) llc =3D debugfs_create_dir("llc_balancing", debugfs_sched); debugfs_create_file("enabled", 0644, llc, NULL, &sched_cache_enable_fops); - debugfs_create_u32("aggr_tolerance", 0644, llc, - &llc_aggr_tolerance); + debugfs_create_u32("aggr_tolerance_nr", 0644, llc, + &llc_aggr_tolerance_nr); + debugfs_create_u32("aggr_tolerance_size", 0644, llc, + &llc_aggr_tolerance_size); debugfs_create_u32("epoch_period", 0644, llc, &llc_epoch_period); debugfs_create_u32("epoch_affinity_timeout", 0644, llc, diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index d78467ec6ee1..c31acbfa3247 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -1404,7 +1404,8 @@ static void set_next_buddy(struct sched_entity *se); */ #define EPOCH_PERIOD (HZ / 100) /* 10 ms */ #define EPOCH_LLC_AFFINITY_TIMEOUT 5 /* 50 ms */ -__read_mostly unsigned int llc_aggr_tolerance =3D 1; +__read_mostly unsigned int llc_aggr_tolerance_nr =3D 1; +__read_mostly unsigned int llc_aggr_tolerance_size =3D 1; __read_mostly unsigned int llc_epoch_period =3D EPOCH_PERIOD; __read_mostly unsigned int llc_epoch_affinity_timeout =3D EPOCH_LLC_AFFINI= TY_TIMEOUT; __read_mostly unsigned int llc_imb_pct =3D 20; @@ -1418,10 +1419,8 @@ static int llc_id(int cpu) return per_cpu(sd_llc_id, cpu); } =20 -static inline int get_sched_cache_scale(int mul) +static inline int get_sched_cache_scale(unsigned int tol, int mul) { - unsigned int tol =3D READ_ONCE(llc_aggr_tolerance); - if (!tol) return 0; =20 @@ -1453,23 +1452,23 @@ static bool exceed_llc_capacity(struct mm_struct *m= m, int cpu) footprint =3D READ_ONCE(mm->sc_stat.footprint); =20 /* - * Scale the LLC size by 256*llc_aggr_tolerance + * Scale the LLC size by 256*llc_aggr_tolerance_size * and compare it to the task's footprint. * * Suppose the L3 size is 32MB. If the - * llc_aggr_tolerance is 1: + * llc_aggr_tolerance_size is 1: * When the footprint is larger than 32MB, the * process is regarded as exceeding the LLC - * capacity. If the llc_aggr_tolerance is 99: + * capacity. If the llc_aggr_tolerance_size is 99: * When the footprint is larger than 784GB, the * process is regarded as exceeding the LLC * capacity: * 784GB =3D (1 + (99 - 1) * 256) * 32MB - * If the llc_aggr_tolerance is 100: + * If the llc_aggr_tolerance_size is 100: * ignore the footprint and do the aggregation * anyway. */ - scale =3D get_sched_cache_scale(256); + scale =3D get_sched_cache_scale(READ_ONCE(llc_aggr_tolerance_size), 256); if (scale =3D=3D INT_MAX) return false; =20 @@ -1488,10 +1487,10 @@ static bool invalid_llc_nr(struct mm_struct *mm, st= ruct task_struct *p, return true; =20 /* - * Scale the number of 'cores' in a LLC by llc_aggr_tolerance + * Scale the number of 'cores' in a LLC by llc_aggr_tolerance_nr * and compare it to the task's active threads. */ - scale =3D get_sched_cache_scale(1); + scale =3D get_sched_cache_scale(READ_ONCE(llc_aggr_tolerance_nr), 1); if (scale =3D=3D INT_MAX) return false; =20 @@ -10418,20 +10417,34 @@ static inline int task_is_ineligible_on_dst_cpu(s= truct task_struct *p, int dest_ * done. * Derived from fits_capacity(). * + * llc_overaggr_pct is an unbounded debugfs u32 and CONFIG_SCHED_CACHE + * only depends on SMP, so max * aggr_pct can wrap on 32-bit. A + * threshold large enough to overflow is effectively unlimited: the + * LLC always has room, so treat that as "fits" rather than letting + * the multiplication wrap. + * * (default: ~50%, tunable via debugfs) */ static bool fits_llc_capacity(unsigned long util, unsigned long max) { - u32 aggr_pct =3D llc_overaggr_pct; + u32 aggr_pct =3D READ_ONCE(llc_overaggr_pct); + unsigned long thresh; + u32 bumped; =20 /* * For single core systems, raise the aggregation * threshold to accommodate more tasks. */ - if (cpu_smt_num_threads =3D=3D 1) - aggr_pct =3D (aggr_pct * 3 / 2); + if (cpu_smt_num_threads =3D=3D 1) { + if (check_mul_overflow(aggr_pct, 3U, &bumped)) + return true; + aggr_pct =3D bumped / 2; + } + + if (check_mul_overflow(max, (unsigned long)aggr_pct, &thresh)) + return true; =20 - return util * 100 < max * aggr_pct; + return util * 100 < thresh; } =20 /* @@ -10439,10 +10452,23 @@ static bool fits_llc_capacity(unsigned long util,= unsigned long max) * is 'util1' noticeably greater than 'util2' * Derived from capacity_greater(). * Bias is in perentage. + * + * Allows dst util to be bigger than src util by up to bias percent. + * llc_imb_pct is an unbounded debugfs u32; a bias large enough to + * overflow util2 * (100 + llc_imb_pct) is effectively infinite, so + * util1 is never noticeably greater - treat that as "not greater" + * rather than letting the multiplication wrap. */ -/* Allows dst util to be bigger than src util by up to bias percent */ -#define util_greater(util1, util2) \ - ((util1) * 100 > (util2) * (100 + llc_imb_pct)) +static bool util_greater(unsigned long util1, unsigned long util2) +{ + unsigned long bias, rhs; + + if (check_add_overflow(100UL, (unsigned long)READ_ONCE(llc_imb_pct), &bia= s) || + check_mul_overflow(util2, bias, &rhs)) + return false; + + return util1 * 100 > rhs; +} =20 static __maybe_unused bool get_llc_stats(int cpu, unsigned long *util, unsigned long *cap) diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h index 56acf502ba26..0ebafdcf2c7f 100644 --- a/kernel/sched/sched.h +++ b/kernel/sched/sched.h @@ -4095,7 +4095,8 @@ static inline void mm_cid_switch_to(struct task_struc= t *prev, struct task_struct DECLARE_STATIC_KEY_FALSE(sched_cache_present); DECLARE_STATIC_KEY_FALSE(sched_cache_active); extern int sysctl_sched_cache_user; -extern unsigned int llc_aggr_tolerance; +extern unsigned int llc_aggr_tolerance_nr; +extern unsigned int llc_aggr_tolerance_size; extern unsigned int llc_epoch_period; extern unsigned int llc_epoch_affinity_timeout; extern unsigned int llc_imb_pct; --=20 2.47.3 From nobody Fri Jul 24 23:35:55 2026 Received: from xmbghk7.mail.qq.com (xmbghk7.mail.qq.com [43.163.128.54]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 079AC492507; Wed, 22 Jul 2026 09:11:29 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=43.163.128.54 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784711498; cv=none; b=TUMN9qIEZKR8wT4FJ8z92Ryyb7T+0wyKbHXOj41Sx7cL5Q9+xH9m8FiTwSg91J9MS1ez3pXz56kN3SYNLpxiDSODTAY42YHELrQMA2qIyf9PO5KgnEyojHV60K026nAXxHvx/wzZPpv3n8HuGbAdbDuqR4BXTK54pQjqpO8EqyA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784711498; c=relaxed/simple; bh=FleOJXrlCtujhY/pztO3ScnodOopKfJNGLR3/6VreNs=; h=Message-ID:From:To:Cc:Subject:Date:In-Reply-To:References: MIME-Version; b=V+FrEkg65KR74qTE/1d3PBQZl16VBr+fBetwHS6Hli+P3pMDD1buffG/XIqL1dOq8+bFUQr/GcdZZYBHdo8v8LEH2keyCuxm//57x+vp5gI/qkevtmpnKjoBQQv6VbXh5icrp/hEkBRETAEVaBWeuUB1USc8C6oMzrXyOkDJ+fU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=cyyself.name; spf=pass smtp.mailfrom=cyyself.name; dkim=pass (1024-bit key) header.d=qq.com header.i=@qq.com header.b=gqD/UTr7; arc=none smtp.client-ip=43.163.128.54 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=cyyself.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=cyyself.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=qq.com header.i=@qq.com header.b="gqD/UTr7" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=qq.com; s=s201512; t=1784711479; bh=pBceSZ87OtYFt6jo6r3CVmVMXRv/Zcf/pKvpkXkFoso=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=gqD/UTr7ZOqPuyNVwJeuNqc/Mu1cnzSx6D6qz9QMdErKQJHwpuKcNJP9AzMjFvibL qgnCgY9d8tKP6V2Y0q5sAP12FAviyZ37D6fdLKfAIo4n2hM5OUGKaVinJ9v1J61LSu 6PFZf3I67G1UAayQwhOvqILc4Zx6VI4b/GYXz05A= Received: from localhost.localdomain ([240e:37c:2242:cf00:5054:a3ff:fe89:345b]) by newxmesmtplogicsvrsza63-0.qq.com (NewEsmtp) with SMTP id 2878B841; Wed, 22 Jul 2026 17:10:07 +0800 X-QQ-mid: xmsmtpt1784711407tplekwvb1 Message-ID: X-QQ-XMAILINFO: Obw6anrw5fL41fkMqrg7zMZByN+QVzXMQcfldQjGYtr+rwTyu5o92iwqvHIFh6 riTlsN5e9Vs1QQXXEIQXhqcSOx+Np04ZJI1fdsYavcyVDExztZgTjDM91sL/WtmAJF4GdYfFfMWn e38IoW2ExOak+HR5GndXJXHNi3AEBn1q4xGBDZUOEtA8br5caaXJfjTCwVrJTK0tZ6BrZl/FQtO1 pCVekQkezDQDhdizIM0oNoc9+hAPcSsNFZ9toNA9UrjDbUQ9Z6gnP4b06LfKFCqzPqYbcNPo+8Zp sk0kZghQnQIq7rQcvCpxvWtPbfHpWurtJ+ugqdiFKstCUzDibmtB201LkcgsJlXzuqhp5UNSgxmw HietZ/AgzH0n4vfqettqkZIhzKXNOy2Ee/obAfbxOL8OfjpJcmsGuUkPRLUlTata+icbtQ91EoRk cbDBW7OG70sHMmHv7zzVB6IqOyzECOs4xzVyD3MfSvK08SUEcXFg47EKtd86ja+nPqV6C+iyOdgt mYM34CFQMMYFBG5xngkTrPwx6/mrCKn78TMHvxTeKZjBjlRaYzW9WUWe07E0hXfNvTK/wM45Imyg z1D0DIccX6DAtyKtVGo9tmNHjAsRzBzxZOz5YrwzMa3asdxEdPEfYf6G+I0HsPm05hitW2XiW8oO aP4XnLBFofw026Nxb0r8oiN2gUT1L9quMpqkmQWmXCCaf7kpP+13lqq/ZOhMIipAVzaaoFVfoczz tZwTBxzmgWPgATTRMnA3LsYlmjh5CUKpwWZdn2S93vHKkbZyGTrxu4lFXAfuk4AwrwUs2JDLXEU4 nPw4hQuFH4SBVWXqSOlEbr5PR5k0jgb2mhCprGzV67Y4CbbAjVfCQenP9vGoCVLNWYJRcjO5KAzz jgkQiseviVJWM842vEHroZxnXlDHo2rmlBm6txXLwDR4/Svpg+wdQmXqXPn5YxWCfDk/ybZIRL1h zGXZhM/emVzd04d6Boigg/7HxEd2Koqtg9vM3Bx7TsG/4BzASq7SsH5WG4gZRTaeu44U541qxDRe j+EZ5qpH7UGJOaCqV/MoMb/pS9zMFPNqqKgITBFVpEigerHDzKivaGgT6Kw83H5/LyQJ0v4GkK9M yS5w4dumGPqQi2SA4VTSm9AksR4/1tzkE4q7gnXpBSAhJUj4lweWx2F+N8bA== X-QQ-XMRINFO: NI4Ajvh11aEjEMj13RCX7UuhPEoou2bs1g== From: Yangyu Chen To: Peter Zijlstra , Ingo Molnar , Juri Lelli , Vincent Guittot Cc: Chen Yu , Tim Chen , K Prateek Nayak , Dietmar Eggemann , Valentin Schneider , Jonathan Corbet , Shuah Khan , Yangyu Chen , linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-doc@vger.kernel.org, Yangyu Chen Subject: [PATCH 2/4] sched/cache: Add PR_SCHED_CACHE prctl for per-mm control Date: Wed, 22 Jul 2026 17:10:03 +0800 X-OQ-MSGID: <20260722091003.1593022-1-cyy@cyyself.name> X-Mailer: git-send-email 2.47.3 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Cache aware scheduling is currently controlled only through global debugfs knobs, but the right aggressiveness is workload and platform specific. A multi-threaded Verilator run is one example: its RSS is large while only a small part of it is hot, so an RSS-based footprint estimate should not decide whether it is aggregated; and packing its threads onto the SMT siblings of one LLC beats spreading them across LLCs on some platforms (e.g. AMD EPYC Turin) but not on others (e.g. EPYC Milan). Such choices cannot be made globally for the whole machine. Add a prctl interface to override the knobs per process (per mm_struct): prctl(PR_SCHED_CACHE, PR_SCHED_CACHE_SET, attr, value, 0); prctl(PR_SCHED_CACHE, PR_SCHED_CACHE_GET, attr, &value, 0); A single prctl command implements both directions, like PR_RSEQ_SLICE_EXTENSION. The attributes are a per-process enable (effective only while the feature is globally active), the two aggregation tolerances, the overaggr percentage (applied where a task's own migration is admitted; group level statistics span many processes and keep using the global value), and an inherit mask selecting which attributes an mm created by execve() keeps. fork() always inherits everything, and the mask itself lives on the task_struct so it survives both, which lets a numactl-like launcher configure a workload and exec it. The overrides live in mm->sc_stat with -1 meaning "follow the global default"; GET stores the raw value through an int pointer so this sentinel round-trips without being mistaken for an errno. mm_init_sched() gains the creating task to tell fork (p !=3D current) from exec (p =3D=3D current) apart. A disabled mm has its preferred LLC invalidated at the existing invalidation points, so all group-level statistics self-neutralize. Also sync the tools/perf/trace/beauty copy of prctl.h. Assisted-by: Claude:claude-fable-5 Signed-off-by: Yangyu Chen --- include/linux/mm_types.h | 10 +- include/linux/sched.h | 16 ++ include/uapi/linux/prctl.h | 38 +++++ kernel/fork.c | 2 +- kernel/sched/fair.c | 149 +++++++++++++++--- kernel/sched/syscalls.c | 120 ++++++++++++++ kernel/sys.c | 5 + .../trace/beauty/include/uapi/linux/prctl.h | 37 +++++ 8 files changed, 352 insertions(+), 25 deletions(-) diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h index b18c2b2e7d2c..eb8e77d6e476 100644 --- a/include/linux/mm_types.h +++ b/include/linux/mm_types.h @@ -1609,10 +1609,11 @@ static inline unsigned int mm_cid_size(void) #endif /* CONFIG_SCHED_MM_CID */ =20 #ifdef CONFIG_SCHED_CACHE -void mm_init_sched(struct mm_struct *mm, +void mm_init_sched(struct mm_struct *mm, struct task_struct *p, struct sched_cache_time __percpu *pcpu_sched); =20 -static inline int mm_alloc_sched_noprof(struct mm_struct *mm) +static inline int mm_alloc_sched_noprof(struct mm_struct *mm, + struct task_struct *p) { struct sched_cache_time __percpu *pcpu_sched =3D alloc_percpu_noprof(struct sched_cache_time); @@ -1620,7 +1621,7 @@ static inline int mm_alloc_sched_noprof(struct mm_str= uct *mm) if (!pcpu_sched) return -ENOMEM; =20 - mm_init_sched(mm, pcpu_sched); + mm_init_sched(mm, p, pcpu_sched); return 0; } =20 @@ -1633,7 +1634,8 @@ static inline void mm_destroy_sched(struct mm_struct = *mm) } #else /* !CONFIG_SCHED_CACHE */ =20 -static inline int mm_alloc_sched(struct mm_struct *mm) { return 0; } +static inline int mm_alloc_sched(struct mm_struct *mm, + struct task_struct *p) { return 0; } static inline void mm_destroy_sched(struct mm_struct *mm) { } =20 #endif /* CONFIG_SCHED_CACHE */ diff --git a/include/linux/sched.h b/include/linux/sched.h index 373bcc0598d1..5a3fd080676a 100644 --- a/include/linux/sched.h +++ b/include/linux/sched.h @@ -1424,6 +1424,8 @@ struct task_struct { int preferred_llc; /* 1: task was enqueued to its preferred LLC, 0 otherwise */ int pref_llc_queued; + /* PR_SCHED_CACHE_INHERIT flags kept across execve() */ + unsigned int sched_cache_inherit; #endif =20 struct rseq_data rseq; @@ -2338,6 +2340,11 @@ static inline void sched_core_fork(struct task_struc= t *p) { } static inline int sched_core_idle_cpu(int cpu) { return idle_cpu(cpu); } #endif =20 +#ifdef CONFIG_SCHED_CACHE +extern int sched_cache_prctl(unsigned long opt, unsigned long attr, + unsigned long val, unsigned long arg5); +#endif + extern void sched_set_stop_task(int cpu, struct task_struct *stop); =20 #ifdef CONFIG_MEM_ALLOC_PROFILING @@ -2398,6 +2405,15 @@ struct sched_cache_stat { unsigned long next_scan; unsigned long footprint; int cpu; + /* + * Per-process overrides of the cache aware scheduling + * knobs, set via prctl(PR_SCHED_CACHE). -1 makes an + * attribute follow the system-wide default. + */ + int user_enabled; + int aggr_tolerance_nr; + int aggr_tolerance_size; + int overaggr_pct; } ____cacheline_aligned_in_smp; =20 #else diff --git a/include/uapi/linux/prctl.h b/include/uapi/linux/prctl.h index b6ec6f693719..dbc1c511b284 100644 --- a/include/uapi/linux/prctl.h +++ b/include/uapi/linux/prctl.h @@ -416,4 +416,42 @@ struct prctl_mm_map { # define PR_CFI_DISABLE _BITUL(1) # define PR_CFI_LOCK _BITUL(2) =20 +/* + * Get or set the per-process (per address space) cache aware + * scheduling attributes. + * + * PR_SCHED_CACHE_GET stores the attribute selected by arg3 into the + * int pointed to by arg4. PR_SCHED_CACHE_SET sets the attribute + * selected by arg3 to the value in arg4. + */ +#define PR_SCHED_CACHE 82 +# define PR_SCHED_CACHE_GET 1 +# define PR_SCHED_CACHE_SET 2 +/* Attributes for PR_SCHED_CACHE_GET/PR_SCHED_CACHE_SET */ +# define PR_SCHED_CACHE_ENABLE 1 +# define PR_SCHED_CACHE_AGGR_TOLERANCE_NR 2 +# define PR_SCHED_CACHE_AGGR_TOLERANCE_SIZE 3 +# define PR_SCHED_CACHE_OVERAGGR_PCT 4 +# define PR_SCHED_CACHE_INHERIT 5 +/* + * Attribute value that resets an attribute to the system default; + * an unset value attribute also reads back as this via + * PR_SCHED_CACHE_GET. + */ +# define PR_SCHED_CACHE_DEFAULT (-1) +/* + * Flags for PR_SCHED_CACHE_INHERIT: which attributes the new address + * space keeps across execve(). New address spaces created by fork() + * always inherit all attributes. + */ +# define PR_SCHED_CACHE_INHERIT_ENABLE (1UL << 0) +# define PR_SCHED_CACHE_INHERIT_AGGR_TOLERANCE_NR (1UL << 1) +# define PR_SCHED_CACHE_INHERIT_AGGR_TOLERANCE_SIZE (1UL << 2) +# define PR_SCHED_CACHE_INHERIT_OVERAGGR_PCT (1UL << 3) +# define PR_SCHED_CACHE_INHERIT_MASK \ + (PR_SCHED_CACHE_INHERIT_ENABLE | \ + PR_SCHED_CACHE_INHERIT_AGGR_TOLERANCE_NR | \ + PR_SCHED_CACHE_INHERIT_AGGR_TOLERANCE_SIZE | \ + PR_SCHED_CACHE_INHERIT_OVERAGGR_PCT) + #endif /* _LINUX_PRCTL_H */ diff --git a/kernel/fork.c b/kernel/fork.c index f0e2e131a9a5..3c7c979e1d3d 100644 --- a/kernel/fork.c +++ b/kernel/fork.c @@ -1135,7 +1135,7 @@ static struct mm_struct *mm_init(struct mm_struct *mm= , struct task_struct *p) if (mm_alloc_cid(mm, p)) goto fail_cid; =20 - if (mm_alloc_sched(mm)) + if (mm_alloc_sched(mm, p)) goto fail_sched; =20 if (percpu_counter_init_many(mm->rss_stat, 0, GFP_KERNEL_ACCOUNT, diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index c31acbfa3247..d085a8438d3d 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -44,6 +44,7 @@ #include #include #include +#include #include #include #include @@ -1430,6 +1431,61 @@ static inline int get_sched_cache_scale(unsigned int= tol, int mul) return (1 + (tol - 1) * mul); } =20 +/* + * Effective cache aware scheduling state of @mm. The per-process + * prctl(PR_SCHED_CACHE_ENABLE) attribute overrides the global + * default. Only meaningful when sched_cache_enabled(). + */ +static bool sched_cache_mm_enabled(struct mm_struct *mm) +{ + int enabled; + + /* + * Statically allocated mms (init_mm, efi_mm) never go through + * mm_init_sched(): their sc_stat is zero-initialized rather + * than set up, recognizable by the NULL pcpu_sched. Treat them + * as disabled instead of interpreting the zeroes (e.g. + * sc_stat.cpu =3D=3D 0 would read as a valid preferred CPU). + */ + if (!mm || !mm->sc_stat.pcpu_sched) + return false; + + /* + * -1 means no per-process override: follow the global enable, + * which is on in every path that reaches this - they are all + * behind sched_cache_enabled(). + */ + enabled =3D READ_ONCE(mm->sc_stat.user_enabled); + + return enabled !=3D 0; +} + +/* + * The following helpers return the effective value of a cache aware + * scheduling knob for @mm: the per-process attribute if one was set + * via prctl(PR_SCHED_CACHE), the global tunable otherwise. + */ +static inline unsigned int mm_aggr_tolerance_nr(struct mm_struct *mm) +{ + int tol =3D mm ? READ_ONCE(mm->sc_stat.aggr_tolerance_nr) : -1; + + return tol >=3D 0 ? tol : READ_ONCE(llc_aggr_tolerance_nr); +} + +static inline unsigned int mm_aggr_tolerance_size(struct mm_struct *mm) +{ + int tol =3D mm ? READ_ONCE(mm->sc_stat.aggr_tolerance_size) : -1; + + return tol >=3D 0 ? tol : READ_ONCE(llc_aggr_tolerance_size); +} + +static inline unsigned int mm_overaggr_pct(struct mm_struct *mm) +{ + int pct =3D mm ? READ_ONCE(mm->sc_stat.overaggr_pct) : -1; + + return pct >=3D 0 ? pct : READ_ONCE(llc_overaggr_pct); +} + static bool exceed_llc_capacity(struct mm_struct *mm, int cpu) { #ifdef CONFIG_NUMA_BALANCING @@ -1468,7 +1524,7 @@ static bool exceed_llc_capacity(struct mm_struct *mm,= int cpu) * ignore the footprint and do the aggregation * anyway. */ - scale =3D get_sched_cache_scale(READ_ONCE(llc_aggr_tolerance_size), 256); + scale =3D get_sched_cache_scale(mm_aggr_tolerance_size(mm), 256); if (scale =3D=3D INT_MAX) return false; =20 @@ -1490,7 +1546,7 @@ static bool invalid_llc_nr(struct mm_struct *mm, stru= ct task_struct *p, * Scale the number of 'cores' in a LLC by llc_aggr_tolerance_nr * and compare it to the task's active threads. */ - scale =3D get_sched_cache_scale(READ_ONCE(llc_aggr_tolerance_nr), 1); + scale =3D get_sched_cache_scale(mm_aggr_tolerance_nr(mm), 1); if (scale =3D=3D INT_MAX) return false; =20 @@ -1571,7 +1627,7 @@ static void account_llc_dequeue(struct rq *rq, struct= task_struct *p) } } =20 -void mm_init_sched(struct mm_struct *mm, +void mm_init_sched(struct mm_struct *mm, struct task_struct *p, struct sched_cache_time __percpu *_pcpu_sched) { unsigned long epoch =3D 0; @@ -1593,6 +1649,34 @@ void mm_init_sched(struct mm_struct *mm, mm->sc_stat.next_scan =3D jiffies; mm->sc_stat.nr_running_avg =3D 0; mm->sc_stat.footprint =3D 0; + mm->sc_stat.user_enabled =3D -1; + mm->sc_stat.aggr_tolerance_nr =3D -1; + mm->sc_stat.aggr_tolerance_size =3D -1; + mm->sc_stat.overaggr_pct =3D -1; + + /* + * A new mm created by fork() (@p is the new child) inherits all + * of the parent's prctl(PR_SCHED_CACHE) attributes. Across + * execve() (@p is current) only the attributes marked in the + * calling thread's PR_SCHED_CACHE_INHERIT mask survive. + */ + if (current->mm) { + struct sched_cache_stat *src =3D ¤t->mm->sc_stat; + unsigned int inherit =3D ~0U; + + if (p =3D=3D current) + inherit =3D READ_ONCE(current->sched_cache_inherit); + + if (inherit & PR_SCHED_CACHE_INHERIT_ENABLE) + mm->sc_stat.user_enabled =3D READ_ONCE(src->user_enabled); + if (inherit & PR_SCHED_CACHE_INHERIT_AGGR_TOLERANCE_NR) + mm->sc_stat.aggr_tolerance_nr =3D READ_ONCE(src->aggr_tolerance_nr); + if (inherit & PR_SCHED_CACHE_INHERIT_AGGR_TOLERANCE_SIZE) + mm->sc_stat.aggr_tolerance_size =3D READ_ONCE(src->aggr_tolerance_size); + if (inherit & PR_SCHED_CACHE_INHERIT_OVERAGGR_PCT) + mm->sc_stat.overaggr_pct =3D READ_ONCE(src->overaggr_pct); + } + /* * The update to mm->sc_stat should not be reordered * before initialization to mm's other fields, in case @@ -1713,10 +1797,12 @@ void account_mm_sched(struct rq *rq, struct task_st= ruct *p, s64 delta_exec) } =20 /* - * If this process hasn't hit task_cache_work() for a while invalidate - * its preferred state. + * If this process hasn't hit task_cache_work() for a while, or + * has cache aware scheduling disabled via prctl(PR_SCHED_CACHE), + * invalidate its preferred state. */ if ((long)(epoch - READ_ONCE(mm->sc_stat.epoch)) > llc_epoch_affinity_tim= eout || + !sched_cache_mm_enabled(mm) || invalid_llc_nr(mm, p, cpu_of(rq)) || exceed_llc_capacity(mm, cpu_of(rq))) { if (READ_ONCE(mm->sc_stat.cpu) !=3D -1) @@ -1747,6 +1833,9 @@ static void task_tick_cache(struct rq *rq, struct tas= k_struct *p) !mm->sc_stat.pcpu_sched) return; =20 + if (!sched_cache_mm_enabled(mm)) + return; + epoch =3D rq->cpu_epoch; /* avoid moving backwards */ if (time_after_eq(mm->sc_stat.epoch, epoch)) @@ -1851,7 +1940,8 @@ static void task_cache_work(struct callback_head *wor= k) return; =20 curr_cpu =3D task_cpu(p); - if (invalid_llc_nr(mm, p, curr_cpu) || + if (!sched_cache_mm_enabled(mm) || + invalid_llc_nr(mm, p, curr_cpu) || exceed_llc_capacity(mm, curr_cpu)) { if (READ_ONCE(mm->sc_stat.cpu) !=3D -1) WRITE_ONCE(mm->sc_stat.cpu, -1); @@ -1918,7 +2008,14 @@ static void task_cache_work(struct callback_head *wo= rk) } } =20 - if (m_a_occ > (2 * curr_m_a_occ)) { + /* + * Re-check the per-mm enable after the scan: a concurrent + * prctl() may have disabled cache aware scheduling for this + * mm and reset sc_stat.cpu while we were scanning - do not + * undo that reset. The check is best effort; a lost race is + * corrected at the next tick. + */ + if (sched_cache_mm_enabled(mm) && m_a_occ > (2 * curr_m_a_occ)) { /* * Avoid switching sc_stat.cpu too fast. * The reason to choose 2X is because: @@ -10417,17 +10514,19 @@ static inline int task_is_ineligible_on_dst_cpu(s= truct task_struct *p, int dest_ * done. * Derived from fits_capacity(). * + * The per-mm percentage is bounded by the prctl, but the global * llc_overaggr_pct is an unbounded debugfs u32 and CONFIG_SCHED_CACHE * only depends on SMP, so max * aggr_pct can wrap on 32-bit. A * threshold large enough to overflow is effectively unlimited: the * LLC always has room, so treat that as "fits" rather than letting * the multiplication wrap. * - * (default: ~50%, tunable via debugfs) + * (default: ~50%, tunable via debugfs and prctl(PR_SCHED_CACHE)) */ -static bool fits_llc_capacity(unsigned long util, unsigned long max) +static bool fits_llc_capacity(unsigned long util, unsigned long max, + struct mm_struct *mm) { - u32 aggr_pct =3D READ_ONCE(llc_overaggr_pct); + u32 aggr_pct =3D mm_overaggr_pct(mm); unsigned long thresh; u32 bumped; =20 @@ -10549,7 +10648,8 @@ enum llc_mig { */ static enum llc_mig can_migrate_llc(int src_cpu, int dst_cpu, unsigned long tsk_util, - bool to_pref) + bool to_pref, + struct mm_struct *mm) { unsigned long src_util, dst_util, src_cap, dst_cap; =20 @@ -10560,8 +10660,8 @@ static enum llc_mig can_migrate_llc(int src_cpu, in= t dst_cpu, src_util =3D src_util < tsk_util ? 0 : src_util - tsk_util; dst_util =3D dst_util + tsk_util; =20 - if (!fits_llc_capacity(dst_util, dst_cap) && - !fits_llc_capacity(src_util, src_cap)) + if (!fits_llc_capacity(dst_util, dst_cap, mm) && + !fits_llc_capacity(src_util, src_cap, mm)) return mig_unrestricted; =20 if (to_pref) { @@ -10571,7 +10671,7 @@ static enum llc_mig can_migrate_llc(int src_cpu, in= t dst_cpu, * than the src, in which case migration will * increase the imbalance too much. */ - if (!fits_llc_capacity(dst_util, dst_cap) && + if (!fits_llc_capacity(dst_util, dst_cap, mm) && util_greater(dst_util, src_util)) return mig_forbid; } else { @@ -10582,7 +10682,7 @@ static enum llc_mig can_migrate_llc(int src_cpu, in= t dst_cpu, * of preferred LLC, leading to migration again * back to preferred LLC. */ - if (fits_llc_capacity(src_util, src_cap) || + if (fits_llc_capacity(src_util, src_cap, mm) || !util_greater(src_util, dst_util)) return mig_forbid; } @@ -10608,8 +10708,9 @@ static enum llc_mig can_migrate_llc_task(int src_cp= u, int dst_cpu, if (cpu < 0 || cpus_share_cache(src_cpu, dst_cpu)) return mig_unrestricted; =20 - /* skip cache aware load balance for too many threads */ - if (invalid_llc_nr(mm, p, dst_cpu) || + /* skip cache aware load balance for disabled mm or too many threads */ + if (!sched_cache_mm_enabled(mm) || + invalid_llc_nr(mm, p, dst_cpu) || exceed_llc_capacity(mm, dst_cpu)) { if (READ_ONCE(mm->sc_stat.cpu) !=3D -1) WRITE_ONCE(mm->sc_stat.cpu, -1); @@ -10624,7 +10725,7 @@ static enum llc_mig can_migrate_llc_task(int src_cp= u, int dst_cpu, return mig_unrestricted; =20 return can_migrate_llc(src_cpu, dst_cpu, - task_util(p), to_pref); + task_util(p), to_pref, mm); } =20 /* @@ -10663,8 +10764,12 @@ alb_break_llc(struct lb_env *env) if (cur && cur->sched_class =3D=3D &fair_sched_class) util =3D task_util(cur); =20 + /* + * No stable mm context here: rq->curr's mm may be + * dropped at any time, use the global threshold. + */ if (can_migrate_llc(env->src_cpu, env->dst_cpu, - util, false) =3D=3D mig_forbid) + util, false, NULL) =3D=3D mig_forbid) return true; } =20 @@ -11775,9 +11880,13 @@ static inline bool llc_balance(struct lb_env *env,= struct sg_lb_stats *sgs, if (env->sd->nr_balance_failed >=3D env->sd->cache_nice_tries + 1) return false; =20 + /* + * Group level statistics aggregate tasks of many processes, + * there is no single owning mm: use the global threshold. + */ if (sgs->nr_pref_dst_llc && can_migrate_llc(cpumask_first(sched_group_span(group)), - env->dst_cpu, 0, true) =3D=3D mig_llc) + env->dst_cpu, 0, true, NULL) =3D=3D mig_llc) return true; =20 return false; diff --git a/kernel/sched/syscalls.c b/kernel/sched/syscalls.c index b215b0ead9a6..dec114eff269 100644 --- a/kernel/sched/syscalls.c +++ b/kernel/sched/syscalls.c @@ -7,6 +7,7 @@ * Copyright (C) 1991-2002 Linus Torvalds * Copyright (C) 1998-2024 Ingo Molnar, Red Hat */ +#include #include #include #include @@ -1576,3 +1577,122 @@ SYSCALL_DEFINE2(sched_rr_get_interval_time32, pid_t= , pid, return retval; } #endif + +#ifdef CONFIG_SCHED_CACHE +/* + * PR_SCHED_CACHE_DEFAULT is the int -1. Depending on how userspace + * passed it (int through prctl()'s varargs, long, or from a 32-bit + * task) it arrives either sign-extended (the first comparison, -1 + * converts to ULONG_MAX) or zero-extended to 0xffffffff (the second); + * accept both. + */ +static bool sched_cache_val_default(unsigned long val) +{ + return val =3D=3D (unsigned long)PR_SCHED_CACHE_DEFAULT || + val =3D=3D (unsigned int)PR_SCHED_CACHE_DEFAULT; +} + +static int sched_cache_set_attr(unsigned long attr, unsigned long val) +{ + struct mm_struct *mm =3D current->mm; + bool def =3D sched_cache_val_default(val); + int ival =3D def ? -1 : (int)val; + + switch (attr) { + case PR_SCHED_CACHE_ENABLE: + if (!def && val > 1) + return -EINVAL; + WRITE_ONCE(mm->sc_stat.user_enabled, ival); + /* + * Drop the preferred LLC hint on any change: a process + * that became disabled must stop being honored right + * away, and one that became enabled re-establishes the + * hint within an epoch anyway. This is best effort: an + * in-flight task_cache_work() scan re-checks the enable + * before publishing a new preference, and a lost race + * is corrected at the next tick. + */ + WRITE_ONCE(mm->sc_stat.cpu, -1); + break; + case PR_SCHED_CACHE_AGGR_TOLERANCE_NR: + if (!def && val > 100) + return -EINVAL; + WRITE_ONCE(mm->sc_stat.aggr_tolerance_nr, ival); + break; + case PR_SCHED_CACHE_AGGR_TOLERANCE_SIZE: + if (!def && val > 100) + return -EINVAL; + WRITE_ONCE(mm->sc_stat.aggr_tolerance_size, ival); + break; + case PR_SCHED_CACHE_OVERAGGR_PCT: + /* + * Bound the percentage so that scaling an LLC capacity + * by it cannot overflow, even on 32-bit. Anything in the + * hundreds already means "never treat the LLC as busy". + */ + if (!def && val > 1000) + return -EINVAL; + WRITE_ONCE(mm->sc_stat.overaggr_pct, ival); + break; + case PR_SCHED_CACHE_INHERIT: + /* the default inherit mask is empty */ + if (def) + val =3D 0; + if (val & ~PR_SCHED_CACHE_INHERIT_MASK) + return -EINVAL; + WRITE_ONCE(current->sched_cache_inherit, val); + break; + default: + return -EINVAL; + } + + return 0; +} + +static int sched_cache_get_attr(unsigned long attr, unsigned long uptr) +{ + struct mm_struct *mm =3D current->mm; + int val; + + switch (attr) { + case PR_SCHED_CACHE_ENABLE: + val =3D READ_ONCE(mm->sc_stat.user_enabled); + break; + case PR_SCHED_CACHE_AGGR_TOLERANCE_NR: + val =3D READ_ONCE(mm->sc_stat.aggr_tolerance_nr); + break; + case PR_SCHED_CACHE_AGGR_TOLERANCE_SIZE: + val =3D READ_ONCE(mm->sc_stat.aggr_tolerance_size); + break; + case PR_SCHED_CACHE_OVERAGGR_PCT: + val =3D READ_ONCE(mm->sc_stat.overaggr_pct); + break; + case PR_SCHED_CACHE_INHERIT: + val =3D READ_ONCE(current->sched_cache_inherit); + break; + default: + return -EINVAL; + } + + return put_user(val, (int __user *)uptr); +} + +int sched_cache_prctl(unsigned long opt, unsigned long attr, + unsigned long val, unsigned long arg5) +{ + if (arg5) + return -EINVAL; + + if (!current->mm) + return -EINVAL; + + switch (opt) { + case PR_SCHED_CACHE_GET: + return sched_cache_get_attr(attr, val); + case PR_SCHED_CACHE_SET: + return sched_cache_set_attr(attr, val); + default: + return -EINVAL; + } +} +#endif /* CONFIG_SCHED_CACHE */ diff --git a/kernel/sys.c b/kernel/sys.c index df69bd71de03..e2b105d72c7e 100644 --- a/kernel/sys.c +++ b/kernel/sys.c @@ -2807,6 +2807,11 @@ SYSCALL_DEFINE5(prctl, int, option, unsigned long, a= rg2, unsigned long, arg3, case PR_SCHED_CORE: error =3D sched_core_share_pid(arg2, arg3, arg4, arg5); break; +#endif +#ifdef CONFIG_SCHED_CACHE + case PR_SCHED_CACHE: + error =3D sched_cache_prctl(arg2, arg3, arg4, arg5); + break; #endif case PR_SET_MDWE: error =3D prctl_set_mdwe(arg2, arg3, arg4, arg5); diff --git a/tools/perf/trace/beauty/include/uapi/linux/prctl.h b/tools/per= f/trace/beauty/include/uapi/linux/prctl.h index 560f99bc4782..dbc1c511b284 100644 --- a/tools/perf/trace/beauty/include/uapi/linux/prctl.h +++ b/tools/perf/trace/beauty/include/uapi/linux/prctl.h @@ -416,5 +416,42 @@ struct prctl_mm_map { # define PR_CFI_DISABLE _BITUL(1) # define PR_CFI_LOCK _BITUL(2) =20 +/* + * Get or set the per-process (per address space) cache aware + * scheduling attributes. + * + * PR_SCHED_CACHE_GET stores the attribute selected by arg3 into the + * int pointed to by arg4. PR_SCHED_CACHE_SET sets the attribute + * selected by arg3 to the value in arg4. + */ +#define PR_SCHED_CACHE 82 +# define PR_SCHED_CACHE_GET 1 +# define PR_SCHED_CACHE_SET 2 +/* Attributes for PR_SCHED_CACHE_GET/PR_SCHED_CACHE_SET */ +# define PR_SCHED_CACHE_ENABLE 1 +# define PR_SCHED_CACHE_AGGR_TOLERANCE_NR 2 +# define PR_SCHED_CACHE_AGGR_TOLERANCE_SIZE 3 +# define PR_SCHED_CACHE_OVERAGGR_PCT 4 +# define PR_SCHED_CACHE_INHERIT 5 +/* + * Attribute value that resets an attribute to the system default; + * an unset value attribute also reads back as this via + * PR_SCHED_CACHE_GET. + */ +# define PR_SCHED_CACHE_DEFAULT (-1) +/* + * Flags for PR_SCHED_CACHE_INHERIT: which attributes the new address + * space keeps across execve(). New address spaces created by fork() + * always inherit all attributes. + */ +# define PR_SCHED_CACHE_INHERIT_ENABLE (1UL << 0) +# define PR_SCHED_CACHE_INHERIT_AGGR_TOLERANCE_NR (1UL << 1) +# define PR_SCHED_CACHE_INHERIT_AGGR_TOLERANCE_SIZE (1UL << 2) +# define PR_SCHED_CACHE_INHERIT_OVERAGGR_PCT (1UL << 3) +# define PR_SCHED_CACHE_INHERIT_MASK \ + (PR_SCHED_CACHE_INHERIT_ENABLE | \ + PR_SCHED_CACHE_INHERIT_AGGR_TOLERANCE_NR | \ + PR_SCHED_CACHE_INHERIT_AGGR_TOLERANCE_SIZE | \ + PR_SCHED_CACHE_INHERIT_OVERAGGR_PCT) =20 #endif /* _LINUX_PRCTL_H */ --=20 2.47.3 From nobody Fri Jul 24 23:35:55 2026 Received: from out203-205-221-210.mail.qq.com (out203-205-221-210.mail.qq.com [203.205.221.210]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DF5ED46D2DB; Wed, 22 Jul 2026 09:10:36 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=203.205.221.210 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784711442; cv=none; b=N9t0dK5CrqTVXpPeMZoaa0nhnjeyhwzIiFCY87a98xe1zDaFUTnAuFvuCdg/LKP8rC0cWL+LJgDTn8IYUQGdXIVaGJh5BZnrRcNE8kfFFc/DwDW2E0jWev1EBmYFdOuuUcT6i+Cz/730FpDa7Elpw84+SSR5A2UB07yMjDRBbBQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784711442; c=relaxed/simple; bh=Kq34AL8sTs1lk9mVauctyl4w6kvH5J+t3A+tMnnwzow=; h=Message-ID:From:To:Cc:Subject:Date:In-Reply-To:References: MIME-Version; b=XnvBgvRupE/zDQ1qR4h1cLhmgEnDtdmqcLTJ59hPnFFiLaMUCaK8G+9qWxDidjJJiF+IAuOqZN5DZ8av5ebipD1GCZJzOwVUKqQ4K4xnUNF87jE2Q6E3pGNxdx5jktEytX1fUt4RkQoLiOBGQCEVAprB0WmUg6V1xj3mncg+T3s= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=cyyself.name; spf=pass smtp.mailfrom=cyyself.name; dkim=pass (1024-bit key) header.d=qq.com header.i=@qq.com header.b=ZzMF9oZB; arc=none smtp.client-ip=203.205.221.210 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=cyyself.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=cyyself.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=qq.com header.i=@qq.com header.b="ZzMF9oZB" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=qq.com; s=s201512; t=1784711424; bh=ywO2kDac9NDoTIOZ6pj/hyQ52rfOMrUFIVHCDa9+1W0=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=ZzMF9oZB5wpw397pR8PX+02Ia67YlWmSPQvD/Mc/SpOC0GNIFb12QlzhM7dC4rEoT JAjUkTMxWmmHMoUQ7j4NkIIciwzYdyBYlj7fbdB8vFE3Pdv0sugFyDWJVcSLpzMFdZ xb7/6rqiw6ppPGCDPZWkdh3kABJ0QptMyGHmKrwE= Received: from localhost.localdomain ([240e:37c:2242:cf00:5054:a3ff:fe89:345b]) by newxmesmtplogicsvrszc43-0.qq.com (NewEsmtp) with SMTP id 2950BA47; Wed, 22 Jul 2026 17:10:21 +0800 X-QQ-mid: xmsmtpt1784711421t6wjxam2q Message-ID: X-QQ-XMAILINFO: N/WmRbclY25GpmgsbjapwPVKCTfLcbqp0pCLKbaxab7tvRE4lCIiWmgiNv7H24 8H945UHmbxl4KzHsvWO3FFN9K2cVydShUjmFdr9Hda8b8TGCPKVgYDRkBMwAu57F++v5k8H8zD18 rYwoOMjbBmyPDaNU9ZEbBZSzasD3Mpt3RR0QVWtvHQKRalaIfJYLnrZmSEIeii+df5H6S2nzkKT7 osn1AbUP4QNAsM6HI0xWlgj26CqqsGEJUgbudjg0MLILTm1KvzfhVU70gWNS5O0KuyZ7oBlR9TL2 XXaz0WScFHWDeLOVqtbqVu6V1ntfi2EtVB4hz6cwaDha18C0OX/XVioj727mJhhRlyzEQeHmm7oh CqYBJJT3dflWITei4Kwgk1FsH357BWU9ozN8/NLUgI1gmsv9ZZh5A8U7le0k6rVTFKsvU7E7G8fn XHlxv2i47/dtzgtj7jj0HjV/R0i92Hv8Ph3pL7kqAiyOoi0/ZKQfg+3MvVRInVs7jBddlwEzzoir m//y4etMaQLr4r4yZnej+yiZutr2balD034tv87gDTKTfGvNkkbPbkXblGC4BrT6GCbumvGVn7rD Aa/f2j9kmpsAdqqCc8v5LPzql7hCtTvJX8cURcwoYOpmZ6WUMMxU5fmaax4ncm/Fe3YjyerX5Y5P fvldSv+/7PnHKd1/SL9iX0WPdYYQbHAjP3Lcn7QFBaxvU4HN6PRJDsNI06evk4SQ3x1wDHXQWrJh gmhtmYCpdk68yR6iTbrhJzw+5o9+1LONoj5sLPUju+4B/NejokVvHZOX+sFcsJtaNS5ygbRNyt+A mXwGr+1cQE3CmKFF32GXY79UcjtvYi/wPv/3SY6HlQD6GCAJVKVwqTE7seMM5uPfFmkhxFh0toPx f5XSvNZFROlvTKsBkB2UunmNh+jTWu9BXWQ4cLTm8lrnCfHQs3qCuJNV2f5NS2aqeKD2aZSY+U/K nFWSPajMVjDwSO00039izQF3jLwtpCcFlSnn7cr6rVEpoaBJYG9TvBlIZiZhc1xoqTSkwBgSZLWa eUpUvIgItgNm9mFcMq4QX8XnBSjMsL3TvBJyRw7EVJT14Ssq+toHD9peThzGaHWZIRQCuGOlsvPW tw0TizHG8976RGu6u3fTzwPyEIAtR23eG558NKnxCINWCzOA4= X-QQ-XMRINFO: NS+P29fieYNwqS3WCnRCOn9D1NpZuCnCRA== From: Yangyu Chen To: Peter Zijlstra , Ingo Molnar , Juri Lelli , Vincent Guittot Cc: Chen Yu , Tim Chen , K Prateek Nayak , Dietmar Eggemann , Valentin Schneider , Jonathan Corbet , Shuah Khan , Yangyu Chen , linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-doc@vger.kernel.org, Yangyu Chen Subject: [PATCH 3/4] selftests/prctl: Add PR_SCHED_CACHE tests Date: Wed, 22 Jul 2026 17:10:20 +0800 X-OQ-MSGID: <20260722091020.1594137-1-cyy@cyyself.name> X-Mailer: git-send-email 2.47.3 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Test the per-mm cache aware scheduling prctl: get/set round trips for every attribute including resets via PR_SCHED_CACHE_DEFAULT in both its sign-extended and 32-bit truncated form, argument validation (out of range values, unknown attributes and sub-commands, nonzero unused arguments, a faulting GET pointer, and that rejected values are not stored), that attribute values live on the shared mm while the inherit mask stays per thread, that fork() always inherits without changes leaking back to the parent, and that execve() resets attributes unless selected in the inherit mask, verified by re-executing the test binary in a checker mode. All tests SKIP on kernels without PR_SCHED_CACHE. Assisted-by: Claude:claude-fable-5 Signed-off-by: Yangyu Chen --- tools/testing/selftests/prctl/.gitignore | 1 + tools/testing/selftests/prctl/Makefile | 15 +- tools/testing/selftests/prctl/config | 1 + tools/testing/selftests/prctl/sched-cache.c | 394 ++++++++++++++++++++ 4 files changed, 406 insertions(+), 5 deletions(-) create mode 100644 tools/testing/selftests/prctl/sched-cache.c diff --git a/tools/testing/selftests/prctl/.gitignore b/tools/testing/selft= ests/prctl/.gitignore index 05d5e31661df..1b22c9f45c4b 100644 --- a/tools/testing/selftests/prctl/.gitignore +++ b/tools/testing/selftests/prctl/.gitignore @@ -2,5 +2,6 @@ disable-tsc-ctxt-sw-stress-test disable-tsc-on-off-stress-test disable-tsc-test +sched-cache set-anon-vma-name-test set-process-name diff --git a/tools/testing/selftests/prctl/Makefile b/tools/testing/selftes= ts/prctl/Makefile index e770e86fad9a..f889721a0749 100644 --- a/tools/testing/selftests/prctl/Makefile +++ b/tools/testing/selftests/prctl/Makefile @@ -1,14 +1,19 @@ # SPDX-License-Identifier: GPL-2.0 -ifndef CROSS_COMPILE ARCH ?=3D $(shell uname -m 2>/dev/null || echo not) override ARCH :=3D $(shell echo $(ARCH) | sed -e s/i.86/x86/ -e s/x86_64/x= 86/) =20 +# sched-cache tests an arch-independent prctl +TEST_PROGS :=3D sched-cache + +ifndef CROSS_COMPILE ifeq ($(ARCH),x86) -TEST_PROGS :=3D disable-tsc-ctxt-sw-stress-test disable-tsc-on-off-stress-= test \ +TEST_PROGS +=3D disable-tsc-ctxt-sw-stress-test disable-tsc-on-off-stress-= test \ disable-tsc-test set-anon-vma-name-test set-process-name +endif +endif + +LDLIBS +=3D -pthread + all: $(TEST_PROGS) =20 include ../lib.mk - -endif -endif diff --git a/tools/testing/selftests/prctl/config b/tools/testing/selftests= /prctl/config index c6ed03c544e5..f44bf7e888e6 100644 --- a/tools/testing/selftests/prctl/config +++ b/tools/testing/selftests/prctl/config @@ -1 +1,2 @@ CONFIG_ANON_VMA_NAME=3Dy +CONFIG_SCHED_CACHE=3Dy diff --git a/tools/testing/selftests/prctl/sched-cache.c b/tools/testing/se= lftests/prctl/sched-cache.c new file mode 100644 index 000000000000..6babbf205a0f --- /dev/null +++ b/tools/testing/selftests/prctl/sched-cache.c @@ -0,0 +1,394 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Tests for prctl(PR_SCHED_CACHE): per-process (per-mm) control of + * cache aware scheduling. + * + * The prctl stores attributes on the calling process's mm, so values + * are shared by all threads, always inherited over fork(), and + * inherited over execve() according to the per-thread + * PR_SCHED_CACHE_INHERIT mask. + */ +#define _GNU_SOURCE + +#include +#include +#include +#include +#include +#include +#include +#include + +#include "kselftest_harness.h" + +#ifndef PR_SCHED_CACHE +#define PR_SCHED_CACHE 82 +# define PR_SCHED_CACHE_GET 1 +# define PR_SCHED_CACHE_SET 2 +# define PR_SCHED_CACHE_ENABLE 1 +# define PR_SCHED_CACHE_AGGR_TOLERANCE_NR 2 +# define PR_SCHED_CACHE_AGGR_TOLERANCE_SIZE 3 +# define PR_SCHED_CACHE_OVERAGGR_PCT 4 +# define PR_SCHED_CACHE_INHERIT 5 +# define PR_SCHED_CACHE_DEFAULT (-1) +# define PR_SCHED_CACHE_INHERIT_ENABLE (1UL << 0) +# define PR_SCHED_CACHE_INHERIT_AGGR_TOLERANCE_NR (1UL << 1) +# define PR_SCHED_CACHE_INHERIT_AGGR_TOLERANCE_SIZE (1UL << 2) +# define PR_SCHED_CACHE_INHERIT_OVERAGGR_PCT (1UL << 3) +# define PR_SCHED_CACHE_INHERIT_MASK \ + (PR_SCHED_CACHE_INHERIT_ENABLE | \ + PR_SCHED_CACHE_INHERIT_AGGR_TOLERANCE_NR | \ + PR_SCHED_CACHE_INHERIT_AGGR_TOLERANCE_SIZE | \ + PR_SCHED_CACHE_INHERIT_OVERAGGR_PCT) +#endif + +static int sc_get(unsigned long attr, int *val) +{ + if (prctl(PR_SCHED_CACHE, PR_SCHED_CACHE_GET, attr, val, 0)) + return -errno; + return 0; +} + +static int sc_set(unsigned long attr, unsigned long val) +{ + if (prctl(PR_SCHED_CACHE, PR_SCHED_CACHE_SET, attr, val, 0)) + return -errno; + return 0; +} + +static int sc_supported(void) +{ + int val; + + return sc_get(PR_SCHED_CACHE_ENABLE, &val) =3D=3D 0; +} + +static void sc_reset_all(void) +{ + sc_set(PR_SCHED_CACHE_ENABLE, PR_SCHED_CACHE_DEFAULT); + sc_set(PR_SCHED_CACHE_AGGR_TOLERANCE_NR, PR_SCHED_CACHE_DEFAULT); + sc_set(PR_SCHED_CACHE_AGGR_TOLERANCE_SIZE, PR_SCHED_CACHE_DEFAULT); + sc_set(PR_SCHED_CACHE_OVERAGGR_PCT, PR_SCHED_CACHE_DEFAULT); + sc_set(PR_SCHED_CACHE_INHERIT, 0); +} + +TEST(get_set_roundtrip) +{ + int val; + + if (!sc_supported()) + SKIP(return, "PR_SCHED_CACHE not supported"); + + sc_reset_all(); + + ASSERT_EQ(0, sc_get(PR_SCHED_CACHE_ENABLE, &val)); + ASSERT_EQ(PR_SCHED_CACHE_DEFAULT, val); + ASSERT_EQ(0, sc_get(PR_SCHED_CACHE_AGGR_TOLERANCE_NR, &val)); + ASSERT_EQ(PR_SCHED_CACHE_DEFAULT, val); + ASSERT_EQ(0, sc_get(PR_SCHED_CACHE_AGGR_TOLERANCE_SIZE, &val)); + ASSERT_EQ(PR_SCHED_CACHE_DEFAULT, val); + ASSERT_EQ(0, sc_get(PR_SCHED_CACHE_OVERAGGR_PCT, &val)); + ASSERT_EQ(PR_SCHED_CACHE_DEFAULT, val); + ASSERT_EQ(0, sc_get(PR_SCHED_CACHE_INHERIT, &val)); + ASSERT_EQ(0, val); + + ASSERT_EQ(0, sc_set(PR_SCHED_CACHE_ENABLE, 0)); + ASSERT_EQ(0, sc_get(PR_SCHED_CACHE_ENABLE, &val)); + ASSERT_EQ(0, val); + ASSERT_EQ(0, sc_set(PR_SCHED_CACHE_ENABLE, 1)); + ASSERT_EQ(0, sc_get(PR_SCHED_CACHE_ENABLE, &val)); + ASSERT_EQ(1, val); + + ASSERT_EQ(0, sc_set(PR_SCHED_CACHE_AGGR_TOLERANCE_NR, 7)); + ASSERT_EQ(0, sc_get(PR_SCHED_CACHE_AGGR_TOLERANCE_NR, &val)); + ASSERT_EQ(7, val); + + ASSERT_EQ(0, sc_set(PR_SCHED_CACHE_AGGR_TOLERANCE_SIZE, 13)); + ASSERT_EQ(0, sc_get(PR_SCHED_CACHE_AGGR_TOLERANCE_SIZE, &val)); + ASSERT_EQ(13, val); + + ASSERT_EQ(0, sc_set(PR_SCHED_CACHE_OVERAGGR_PCT, 155)); + ASSERT_EQ(0, sc_get(PR_SCHED_CACHE_OVERAGGR_PCT, &val)); + ASSERT_EQ(155, val); + + ASSERT_EQ(0, sc_set(PR_SCHED_CACHE_INHERIT, + PR_SCHED_CACHE_INHERIT_AGGR_TOLERANCE_NR)); + ASSERT_EQ(0, sc_get(PR_SCHED_CACHE_INHERIT, &val)); + ASSERT_EQ(PR_SCHED_CACHE_INHERIT_AGGR_TOLERANCE_NR, (unsigned long)val); + + /* DEFAULT resets the inherit mask to empty */ + ASSERT_EQ(0, sc_set(PR_SCHED_CACHE_INHERIT, PR_SCHED_CACHE_DEFAULT)); + ASSERT_EQ(0, sc_get(PR_SCHED_CACHE_INHERIT, &val)); + ASSERT_EQ(0, val); + + /* both the sign-extended and the 32-bit truncated DEFAULT reset */ + ASSERT_EQ(0, sc_set(PR_SCHED_CACHE_AGGR_TOLERANCE_NR, + PR_SCHED_CACHE_DEFAULT)); + ASSERT_EQ(0, sc_get(PR_SCHED_CACHE_AGGR_TOLERANCE_NR, &val)); + ASSERT_EQ(PR_SCHED_CACHE_DEFAULT, val); + ASSERT_EQ(0, sc_set(PR_SCHED_CACHE_AGGR_TOLERANCE_SIZE, 0xffffffffUL)); + ASSERT_EQ(0, sc_get(PR_SCHED_CACHE_AGGR_TOLERANCE_SIZE, &val)); + ASSERT_EQ(PR_SCHED_CACHE_DEFAULT, val); + + sc_reset_all(); +} + +TEST(invalid_arguments) +{ + int val; + + if (!sc_supported()) + SKIP(return, "PR_SCHED_CACHE not supported"); + + sc_reset_all(); + + /* out of range values */ + ASSERT_EQ(-EINVAL, sc_set(PR_SCHED_CACHE_ENABLE, 2)); + ASSERT_EQ(-EINVAL, sc_set(PR_SCHED_CACHE_AGGR_TOLERANCE_NR, 101)); + ASSERT_EQ(-EINVAL, sc_set(PR_SCHED_CACHE_AGGR_TOLERANCE_SIZE, 101)); + ASSERT_EQ(-EINVAL, sc_set(PR_SCHED_CACHE_OVERAGGR_PCT, 1001)); + ASSERT_EQ(-EINVAL, sc_set(PR_SCHED_CACHE_INHERIT, + PR_SCHED_CACHE_INHERIT_MASK + 1)); + + /* unknown attribute */ + ASSERT_EQ(-EINVAL, sc_set(0, 1)); + ASSERT_EQ(-EINVAL, sc_set(PR_SCHED_CACHE_INHERIT + 1, 1)); + ASSERT_EQ(-EINVAL, sc_get(0, &val)); + ASSERT_EQ(-EINVAL, sc_get(PR_SCHED_CACHE_INHERIT + 1, &val)); + + /* unknown op */ + errno =3D 0; + ASSERT_EQ(-1, prctl(PR_SCHED_CACHE, 0, PR_SCHED_CACHE_ENABLE, 0, 0)); + ASSERT_EQ(EINVAL, errno); + errno =3D 0; + ASSERT_EQ(-1, prctl(PR_SCHED_CACHE, 3, PR_SCHED_CACHE_ENABLE, 0, 0)); + ASSERT_EQ(EINVAL, errno); + + /* nonzero unused argument */ + errno =3D 0; + ASSERT_EQ(-1, prctl(PR_SCHED_CACHE, PR_SCHED_CACHE_SET, + PR_SCHED_CACHE_ENABLE, 1, 1)); + ASSERT_EQ(EINVAL, errno); + + /* bad GET pointer */ + errno =3D 0; + ASSERT_EQ(-1, prctl(PR_SCHED_CACHE, PR_SCHED_CACHE_GET, + PR_SCHED_CACHE_ENABLE, NULL, 0)); + ASSERT_EQ(EFAULT, errno); + + /* values rejected with -EINVAL must not be stored */ + ASSERT_EQ(0, sc_get(PR_SCHED_CACHE_ENABLE, &val)); + ASSERT_EQ(PR_SCHED_CACHE_DEFAULT, val); +} + +struct thread_ctx { + int nr_seen; + int inherit_seen; + int ret; +}; + +static void *thread_fn(void *arg) +{ + struct thread_ctx *ctx =3D arg; + + ctx->ret =3D sc_get(PR_SCHED_CACHE_AGGR_TOLERANCE_NR, &ctx->nr_seen); + if (!ctx->ret) + ctx->ret =3D sc_get(PR_SCHED_CACHE_INHERIT, &ctx->inherit_seen); + if (!ctx->ret) + ctx->ret =3D sc_set(PR_SCHED_CACHE_AGGR_TOLERANCE_NR, 42); + /* clearing this thread's mask must not affect the creator's */ + if (!ctx->ret) + ctx->ret =3D sc_set(PR_SCHED_CACHE_INHERIT, 0); + + return NULL; +} + +TEST(values_shared_by_threads_inherit_mask_is_not) +{ + struct thread_ctx ctx =3D { .nr_seen =3D -2, .inherit_seen =3D -2 }; + pthread_t thread; + int val; + + if (!sc_supported()) + SKIP(return, "PR_SCHED_CACHE not supported"); + + sc_reset_all(); + + ASSERT_EQ(0, sc_set(PR_SCHED_CACHE_AGGR_TOLERANCE_NR, 31)); + ASSERT_EQ(0, sc_set(PR_SCHED_CACHE_INHERIT, + PR_SCHED_CACHE_INHERIT_ENABLE)); + + ASSERT_EQ(0, pthread_create(&thread, NULL, thread_fn, &ctx)); + ASSERT_EQ(0, pthread_join(thread, NULL)); + ASSERT_EQ(0, ctx.ret); + + /* attribute values live on the shared mm */ + ASSERT_EQ(31, ctx.nr_seen); + ASSERT_EQ(0, sc_get(PR_SCHED_CACHE_AGGR_TOLERANCE_NR, &val)); + ASSERT_EQ(42, val); + + /* + * The inherit mask is per thread: the new thread starts with a + * copy of its creator's mask, and clearing it in the thread + * does not touch the creator's. + */ + ASSERT_EQ(PR_SCHED_CACHE_INHERIT_ENABLE, + (unsigned long)ctx.inherit_seen); + ASSERT_EQ(0, sc_get(PR_SCHED_CACHE_INHERIT, &val)); + ASSERT_EQ(PR_SCHED_CACHE_INHERIT_ENABLE, (unsigned long)val); + + sc_reset_all(); +} + +TEST(fork_always_inherits) +{ + pid_t pid; + int status; + int val; + + if (!sc_supported()) + SKIP(return, "PR_SCHED_CACHE not supported"); + + sc_reset_all(); + + ASSERT_EQ(0, sc_set(PR_SCHED_CACHE_ENABLE, 1)); + ASSERT_EQ(0, sc_set(PR_SCHED_CACHE_AGGR_TOLERANCE_NR, 9)); + ASSERT_EQ(0, sc_set(PR_SCHED_CACHE_AGGR_TOLERANCE_SIZE, 11)); + ASSERT_EQ(0, sc_set(PR_SCHED_CACHE_OVERAGGR_PCT, 77)); + + pid =3D fork(); + ASSERT_LE(0, pid); + if (pid =3D=3D 0) { + int val; + + if (sc_get(PR_SCHED_CACHE_ENABLE, &val) || val !=3D 1) + _exit(1); + if (sc_get(PR_SCHED_CACHE_AGGR_TOLERANCE_NR, &val) || val !=3D 9) + _exit(2); + if (sc_get(PR_SCHED_CACHE_AGGR_TOLERANCE_SIZE, &val) || val !=3D 11) + _exit(3); + if (sc_get(PR_SCHED_CACHE_OVERAGGR_PCT, &val) || val !=3D 77) + _exit(4); + + /* changes in the child must not leak back to the parent */ + if (sc_set(PR_SCHED_CACHE_AGGR_TOLERANCE_NR, 50)) + _exit(5); + + _exit(0); + } + + ASSERT_EQ(pid, waitpid(pid, &status, 0)); + ASSERT_TRUE(WIFEXITED(status)); + ASSERT_EQ(0, WEXITSTATUS(status)); + + ASSERT_EQ(0, sc_get(PR_SCHED_CACHE_AGGR_TOLERANCE_NR, &val)); + ASSERT_EQ(9, val); + + sc_reset_all(); +} + +static int exec_check(int expect_enable, int expect_nr, + int expect_size, int expect_pct) +{ + char enable[16], nr[16], size[16], pct[16]; + pid_t pid; + int status; + + snprintf(enable, sizeof(enable), "%d", expect_enable); + snprintf(nr, sizeof(nr), "%d", expect_nr); + snprintf(size, sizeof(size), "%d", expect_size); + snprintf(pct, sizeof(pct), "%d", expect_pct); + + pid =3D fork(); + if (pid < 0) + return -1; + + if (pid =3D=3D 0) { + setenv("SCHED_CACHE_EXEC_MODE", "1", 1); + setenv("SCHED_CACHE_EXPECT_ENABLE", enable, 1); + setenv("SCHED_CACHE_EXPECT_NR", nr, 1); + setenv("SCHED_CACHE_EXPECT_SIZE", size, 1); + setenv("SCHED_CACHE_EXPECT_PCT", pct, 1); + execl("/proc/self/exe", "sched-cache", NULL); + _exit(126); + } + + if (waitpid(pid, &status, 0) !=3D pid) + return -1; + if (!WIFEXITED(status)) + return -1; + + return WEXITSTATUS(status); +} + +TEST(execve_inheritance) +{ + if (!sc_supported()) + SKIP(return, "PR_SCHED_CACHE not supported"); + + sc_reset_all(); + + ASSERT_EQ(0, sc_set(PR_SCHED_CACHE_ENABLE, 1)); + ASSERT_EQ(0, sc_set(PR_SCHED_CACHE_AGGR_TOLERANCE_NR, 21)); + ASSERT_EQ(0, sc_set(PR_SCHED_CACHE_AGGR_TOLERANCE_SIZE, 22)); + ASSERT_EQ(0, sc_set(PR_SCHED_CACHE_OVERAGGR_PCT, 23)); + + /* without inherit flags execve() resets everything */ + ASSERT_EQ(0, sc_set(PR_SCHED_CACHE_INHERIT, 0)); + ASSERT_EQ(0, exec_check(PR_SCHED_CACHE_DEFAULT, + PR_SCHED_CACHE_DEFAULT, + PR_SCHED_CACHE_DEFAULT, + PR_SCHED_CACHE_DEFAULT)); + + /* selected attributes survive execve() */ + ASSERT_EQ(0, sc_set(PR_SCHED_CACHE_INHERIT, + PR_SCHED_CACHE_INHERIT_AGGR_TOLERANCE_NR | + PR_SCHED_CACHE_INHERIT_OVERAGGR_PCT)); + ASSERT_EQ(0, exec_check(PR_SCHED_CACHE_DEFAULT, 21, + PR_SCHED_CACHE_DEFAULT, 23)); + + /* all of them */ + ASSERT_EQ(0, sc_set(PR_SCHED_CACHE_INHERIT, + PR_SCHED_CACHE_INHERIT_MASK)); + ASSERT_EQ(0, exec_check(1, 21, 22, 23)); + + sc_reset_all(); +} + +static int exec_expect(const char *env, unsigned long attr) +{ + const char *str =3D getenv(env); + int val; + + if (!str) + return 1; + if (sc_get(attr, &val)) + return 1; + if (val !=3D atoi(str)) + return 1; + + return 0; +} + +static int exec_check_main(void) +{ + int bad =3D 0; + + bad |=3D exec_expect("SCHED_CACHE_EXPECT_ENABLE", PR_SCHED_CACHE_ENABLE); + bad |=3D exec_expect("SCHED_CACHE_EXPECT_NR", + PR_SCHED_CACHE_AGGR_TOLERANCE_NR); + bad |=3D exec_expect("SCHED_CACHE_EXPECT_SIZE", + PR_SCHED_CACHE_AGGR_TOLERANCE_SIZE); + bad |=3D exec_expect("SCHED_CACHE_EXPECT_PCT", + PR_SCHED_CACHE_OVERAGGR_PCT); + + return bad; +} + +int main(int argc, char **argv) +{ + if (getenv("SCHED_CACHE_EXEC_MODE")) + return exec_check_main(); + + return test_harness_run(argc, argv); +} --=20 2.47.3 From nobody Fri Jul 24 23:35:55 2026 Received: from out203-205-221-231.mail.qq.com (out203-205-221-231.mail.qq.com [203.205.221.231]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 275D847ECD7; Wed, 22 Jul 2026 09:10:39 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=203.205.221.231 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784711443; cv=none; b=f4LyowF931XDaQ67zwZYclVpSh8yTp6UV5flNzYCFj1id4pXhoIjkPmhwpB5EYst5b4t6r6XNnapRiV3qMgnHk/v24oRnFEfRp0hRb1jSUIVs3p1x2ybBOsETbpyCvAVj0zWQzR3GJe2+pywFzbYGgeb7Hm/pkPMh7XJktFbV6c= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784711443; c=relaxed/simple; bh=HtRRQ7FvUvFUqrNqKU6qaN0koHzhkGJVahBmFnXPI0E=; h=Message-ID:From:To:Cc:Subject:Date:In-Reply-To:References: MIME-Version; b=jZb8r8vdns1hCLB0UgFJUNIB/3vDpHFXfOC8kfi++XSSLY5no2z/OaVJD0E4/RS9tasIdnBOdpe4emYui1in1ExF+R2es16j85Tt+fq2QSE+SyZgTsMd9rbVOPEoLYGm8XHxt9MSmgGKQxqXXWyfHepoxKKBGvQ5O2Zrb+yx+sk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=cyyself.name; spf=pass smtp.mailfrom=cyyself.name; dkim=pass (1024-bit key) header.d=qq.com header.i=@qq.com header.b=DxxkqMg+; arc=none smtp.client-ip=203.205.221.231 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=cyyself.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=cyyself.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=qq.com header.i=@qq.com header.b="DxxkqMg+" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=qq.com; s=s201512; t=1784711436; bh=15c3lDh4VOZgnc1cxDbEs540nVC466TqGoV/PidAeLE=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=DxxkqMg+FCTsD0K3UGdNsE/tgJtj1TqsEX5KrTJEDr1YNDf2jFnKnKzRlGABKUfGL EgNmPyrn4a3F1VoznRUqwX5kOtZkQMmuDJ4/ohv+ADkTX8SWJtyu9crxm38MHSH5Ku JXjHLi4BC3WtlAq27iQNtTIha5r8IX759cgNd0JY= Received: from localhost.localdomain ([240e:37c:2242:cf00:5054:a3ff:fe89:345b]) by newxmesmtplogicsvrszc43-0.qq.com (NewEsmtp) with SMTP id 2A18FC86; Wed, 22 Jul 2026 17:10:33 +0800 X-QQ-mid: xmsmtpt1784711433t0k849hic Message-ID: X-QQ-XMAILINFO: OATpkVjS499u0vvGgzUsiACipxHsREnOxfr5ZTGpVur0HETK5BlTd3aSVcQyXd w6jhJaSjgl8iiQ1g/qO7qjasL43Por9ofY9tOr8o2HBMUvt5uK2iTAVc7hTA4XzJBzXx9NffxhA2 dr9EtoEOLZ/GLQjWTNaQ/eBeMO4fXpVA9viU8YBzkWMHCSn6EyacSPONq6vBXYUL/o+aSKSZT20a ZAyvcKAkIyG59ph+Yf0MPHzSLz0QDqkkgd6vc/3rnhi9rviYhKa+d9L8lGDmyCS289gJwZsAXN63 UQ8IW3opAsYMJgBBFS5it0JHDnpQJC0CJzYjatyzbN06esD+Xpwt2HfbqLppl7XjKf+yWuDajh6c vegWFh665rjlxq+ZPRre+R4W5XEUItrh6Hn4ERYsFQYHe3Mr9EPyoc98OrspFRoktiGTvj/P67kB jRLNUD3t/mc6MkqKErDAdxUDbom4gUtnWFDlynCC7Suqox3Dh+d1uHci57iHk7lVrJEEWXzgTV53 DqLb1LguiffhYPe6XjWwdpmBBZR/31iIFo3Pnwf00RLYOaBW7+APiVLeklDw2pq5DqYc5+yGIkd2 6OMQy2CM3HGUvuvnEp0o2BUGcckCP2Ltr7bBXjlzXlQkm43N9kAg3HSOYHNGgVDGZ+8sEDg/f+a8 UEca/yfz+AYVuWPblAtZzyTwy2RrQE0mNanyOjmc4ogi/k93iXBDSC0jQui0oO9aOVUSYCZqAaaZ kECEtEG4TPaOxmemjHNiHxQnRUc+I7+nKwtVGEcK996m6ZwGmFKzCOKp9H8KuLlaU766Kw91yLEn W1gALMB95U/4i1VKx5ENAy6xmmSbYETRlZtru7PM9dnUqJJJjvQfNBiEnkdtbbAb8aFCtix2z9X4 MSdQaTjkKizJIwGjsuENPuOoLIi55LC82saXV2tV5RB0p71ZdXiRgE8V7SLVmvHriJNbBul/a/Kr m/3Fyx0JsMtD9oUyzwKqVAOfJ0vmFnDG4397Z/PRjt8ZeoBshprNNlO4coizouIttrhIGhuG4WeF veUr6PUEENrjaoHPeyDpQHoBKx1B04zvHcTN9vy59nXqXLWSD5iZ83OicmaRJbgKxeGgcMYWNpMz Fz9pOHaH3TYloCCl6iebDdEhm08pyPf9tKp8esKTj4rlSOU5I= X-QQ-XMRINFO: Nq+8W0+stu50tPAe92KXseR0ZZmBTk3gLg== From: Yangyu Chen To: Peter Zijlstra , Ingo Molnar , Juri Lelli , Vincent Guittot Cc: Chen Yu , Tim Chen , K Prateek Nayak , Dietmar Eggemann , Valentin Schneider , Jonathan Corbet , Shuah Khan , Yangyu Chen , linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-doc@vger.kernel.org, Yangyu Chen Subject: [PATCH 4/4] docs/scheduler: Document cache aware scheduling controls Date: Wed, 22 Jul 2026 17:10:32 +0800 X-OQ-MSGID: <20260722091032.1595192-1-cyy@cyyself.name> X-Mailer: git-send-email 2.47.3 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Cache aware scheduling gained a debugfs control surface when it was merged and now also a per-process prctl (PR_SCHED_CACHE), but neither is documented. Add Documentation/scheduler/sched-cache.rst describing what the feature does and when aggregation is skipped, the /sys/kernel/debug/sched/llc_balancing/ knobs including the split aggr_tolerance_nr/aggr_tolerance_size tolerances, the prctl attributes with their ranges and the GET/SET calling convention, and the inheritance semantics across fork() and execve() with a numactl-like launcher example. Assisted-by: Claude:claude-fable-5 Signed-off-by: Yangyu Chen --- Documentation/scheduler/index.rst | 1 + Documentation/scheduler/sched-cache.rst | 125 ++++++++++++++++++++++++ 2 files changed, 126 insertions(+) create mode 100644 Documentation/scheduler/sched-cache.rst diff --git a/Documentation/scheduler/index.rst b/Documentation/scheduler/in= dex.rst index 17ce8d76befc..8c6e633d4d79 100644 --- a/Documentation/scheduler/index.rst +++ b/Documentation/scheduler/index.rst @@ -10,6 +10,7 @@ Scheduler membarrier sched-arch sched-bwc + sched-cache sched-deadline sched-design-CFS sched-eevdf diff --git a/Documentation/scheduler/sched-cache.rst b/Documentation/schedu= ler/sched-cache.rst new file mode 100644 index 000000000000..8bd92c7c7d3d --- /dev/null +++ b/Documentation/scheduler/sched-cache.rst @@ -0,0 +1,125 @@ +.. SPDX-License-Identifier: GPL-2.0 + +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D +Cache Aware Scheduling +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +Cache aware scheduling (``CONFIG_SCHED_CACHE``) aggregates the threads +of a process on a preferred last level cache (LLC) domain when that is +expected to improve cache locality: the scheduler samples the per-LLC +CPU occupancy of every multi-threaded process and biases load balance +towards the LLC the process already uses the most. + +Aggregation is skipped for processes that would not benefit from it: +single-threaded processes, processes with more active threads than the +LLC has cores, and processes whose memory footprint exceeds the LLC +size (footprint tracking requires ``CONFIG_NUMA_BALANCING``). + +Global control (debugfs) +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +The feature is controlled globally through +``/sys/kernel/debug/sched/llc_balancing/``: + +``enabled`` + Boolean. When on (the default), every process participates + unless it opted out via ``prctl(PR_SCHED_CACHE)``; when off, + the feature is inactive for every process. + +``aggr_tolerance_nr`` + Scales the number of cores of an LLC before it is compared to + the process's number of active threads. 0 disables aggregation + through this gate, 1 is the strict default, values of 100 or + more disable the thread count check entirely. + +``aggr_tolerance_size`` + Scales the LLC size (by 256 per step) before it is compared to + the process's memory footprint. Same 0/1/100 semantics as + ``aggr_tolerance_nr``. Only effective with + ``CONFIG_NUMA_BALANCING``. + +``overaggr_pct`` + The LLC utilization threshold, in percent of the LLC capacity, + below which the LLC is considered idle enough to aggregate more + tasks into it (default 50). + +``imb_pct`` + Utilization imbalance hysteresis, in percent, used when + comparing two LLCs (default 20). + +``epoch_period``, ``epoch_affinity_timeout`` + Occupancy sampling period (jiffies) and the number of epochs + after which a process's LLC preference expires. + +Per-process control (prctl) +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D= =3D=3D=3D + +``prctl(PR_SCHED_CACHE)`` overrides the global knobs for one process. +The attributes live on the process's address space (``mm_struct``): +they are shared by all threads of the process, and a single prctl +command implements both directions:: + + prctl(PR_SCHED_CACHE, PR_SCHED_CACHE_SET, attr, value, 0); + prctl(PR_SCHED_CACHE, PR_SCHED_CACHE_GET, attr, &value, 0); + +``PR_SCHED_CACHE_GET`` stores the raw attribute value into the ``int`` +pointed to by the fourth argument and returns 0. A value attribute +that was never set reads back as -1 (=3D=3D ``PR_SCHED_CACHE_DEFAULT``), +meaning "follow the global default"; the ``PR_SCHED_CACHE_INHERIT`` +mask reads back as the mask itself (default 0). Setting an attribute +to ``PR_SCHED_CACHE_DEFAULT`` resets it (for the inherit mask this +clears all bits). + +Attributes: + +``PR_SCHED_CACHE_ENABLE`` + 0 or 1. A process can opt out of the globally enabled feature + with 0; 1 (like the default) participates. It cannot activate + the feature when the global knob is off or the hardware has a + single LLC. + +``PR_SCHED_CACHE_AGGR_TOLERANCE_NR`` + 0..100. Per-process version of ``aggr_tolerance_nr``. + +``PR_SCHED_CACHE_AGGR_TOLERANCE_SIZE`` + 0..100. Per-process version of ``aggr_tolerance_size``. + +``PR_SCHED_CACHE_OVERAGGR_PCT`` + 0..1000. Per-process version of ``overaggr_pct``. It applies + where a task's own migration is admitted; group-level load + balance statistics keep using the global value, because they + aggregate tasks of many processes. + +``PR_SCHED_CACHE_INHERIT`` + A bitmask selecting which attributes a new address space + created by execve() keeps: ``PR_SCHED_CACHE_INHERIT_ENABLE``, + ``..._AGGR_TOLERANCE_NR``, ``..._AGGR_TOLERANCE_SIZE`` and + ``..._OVERAGGR_PCT``. The mask is per thread and is itself kept + across fork() and execve(). + +Inheritance +=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + +fork() always copies all attributes to the child, like other process +properties. execve() resets attributes to the default unless the +corresponding ``PR_SCHED_CACHE_INHERIT`` bit is set in the calling +thread. + +This allows a numactl-like launcher to configure a workload:: + + prctl(PR_SCHED_CACHE, PR_SCHED_CACHE_SET, + PR_SCHED_CACHE_AGGR_TOLERANCE_NR, 100, 0); + prctl(PR_SCHED_CACHE, PR_SCHED_CACHE_SET, + PR_SCHED_CACHE_INHERIT, + PR_SCHED_CACHE_INHERIT_AGGR_TOLERANCE_NR, 0); + execve(workload, ...); + +Errors +=3D=3D=3D=3D=3D=3D + +``EINVAL`` + Unknown sub-command or attribute, value out of range, nonzero + unused argument, calling task has no mm, or the kernel was + built without ``CONFIG_SCHED_CACHE``. +``EFAULT`` + Invalid pointer passed to ``PR_SCHED_CACHE_GET``. --=20 2.47.3