From nobody Tue Sep 29 01:18:19 2026 Received: from mail-pl1-f170.google.com (mail-pl1-f170.google.com [209.85.214.170]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D87DD4968EC for ; Thu, 13 Aug 2026 18:58:51 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.170 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786647534; cv=none; b=GWYSb2KUEMyg5aUnBjlkHqiSnkBvX6ajxzZZ/+k/j00VmvQyqKBT14nMCejCxij2xpV4M7uMRIm+ycvQmJa1oT+VrGcPdOfy7VUc9KJWh8BTPfYkC5BU7wW5s21HisAaEsZCihDkS1Cp7D3erZqV4eetJ49FMNCUrZNFHFwdTag= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786647534; c=relaxed/simple; bh=qeyDf1zf1toovwRK9sFVxz1irGFidwtP1E1vr6ae6BY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=U4S78CmgIV/Cb+ue1hQhytEn8s/aORZ9gSQe74Y7PwTq1U97mDvHaVS1EvKVM3FRB38jwSxQrpuzCMU+N7opD1bJKJvIIs+jol/Mq5XtTCp0myWTL1Gtczpkkr7B/pnS5cHMIzEbSrv8Eyr85+6uO8idFdy84zsI/HGw5i7E2Mk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=hqPlzD4T; arc=none smtp.client-ip=209.85.214.170 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="hqPlzD4T" Received: by mail-pl1-f170.google.com with SMTP id d9443c01a7336-2cc891373e0so4440655ad.2 for ; Thu, 13 Aug 2026 11:58:51 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786647531; x=1787252331; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=j7MPifKM0g+V499r8MTdHDoJMUDv6wW/Ghb2RIQSLiw=; b=hqPlzD4TwzSpNtLlKBi0A7aiZru5FAZVYFnxRG0ruRWbe1iE4qu55CfW58jmHgs3H+ aNr8CsjSLMS4oU0wi3iHzzF96aBztl6mYR9UPTCqTfqh9veqqT4nuzdpy72qKmWIMwVS b+Wk7XjD/y+kSOCGUrjrgaQiv4z1RAvT4QP0dUlv5uScWMhPP1gw1erOLtGwLTXs/V+I mQgrRjW0YuagvOhQfX+KGBCtrHvq56fNbh44x1x3dx5htpfLblkb/pdvuNyrte4Q3t64 5iRVBO4DMmPijo/I8mWhqW9qhcKdhXOBeBID4P/tLcg4al8W2EDaX4AXrSWU4xuUA9rj 5yOQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786647531; x=1787252331; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=j7MPifKM0g+V499r8MTdHDoJMUDv6wW/Ghb2RIQSLiw=; b=Vx2uG4HyM9s2OSTeX0KMjHunwKuyzcpUaD0eu9lZ3gtuIqDS4WTKuOdANwhQo+ajHt 3K7og2xPWIEldaz/rxySVCDnV2ZWdNWIK/5eXoCo1NX5cv8C1nc/kttd1qLiGmo90jHQ 56irG15md6pTPStIqG+17EBHq706en3oOXMoJnym9W0SQpTGQFSTtBzeJiZ6CvOmjU9L P5EyUmAxF32XB1/OcUjhjIN/UUenMsATgvyiJ21Xyeg4pKQX/B3gAbM/4vdRMWeC3PyP E/qoZWROPeJ3M7guSA+yLMKU0nXFZuJ8Kx37eyQiCplX3STwU5I4cGDamrfNMKoO4pJk n2Cw== X-Forwarded-Encrypted: i=1; AHgh+RqDvK0xa+LexRx/VVeqgFmRqhU9uKrFRdFLkj3Tl54o0JiCzTrI8CGHsRGfPJpV3ox2I+H+bmvCiiJLxkk=@vger.kernel.org X-Gm-Message-State: AOJu0Yw1UE2p08RpuWdKlM7PK1J/D/m5EcEYZJxIn3uTnNC6iNzcV55N Cu5AeK6exq3sQ6i2XWeIrRdewwADC44FhK6XiGCj0TTRkik3Sp7FpfwD X-Gm-Gg: AR+sD11sPxQvEr8OTdMYaYiKnwbeEir+jr+Tj299LP2IJt1DP9ZqjQeV0mBAQS9GJED t5Tik/avOl5szvgzZI7iBOi0cwSJXerXzhiJA7O103lK7yCqR8dmwMhgxeBy6fh6hJxYVJAHqVd cAL14o1GivGtiX+UxX2DomzcpxlDYLckdH1pXhmjnyE2fhW1msJOTm03HImXT4Db/ft3wPkzwpS IILNjcV/vYKKLstQfkTxcP29uF7RCCmLQu2g3vTAtyk9J1rYAYfdxyL2QkYOyTmb2w/IqFnSObS 6bbrTnD4AFC9VLUtaJNZ9kYcgh4TDEo+VSwj7IFEfby5jlP7fxH48opK7V4Y/UXPHnnUl5QEhyf khZRBmP/CRDcUyDjvbaFqi6YYG7EZOlTZlQ2Z79CfduJo8PanQswGL8cQ69BACSpFNqpXE8sT0E ZxmwlIzQ1aENdAPtSJcKg1V2nSgRdY8bRW9noElhPLbspKHUcb543xJ8Q= X-Received: by 2002:a05:6a20:3d1c:b0:3c6:3c5b:f2e3 with SMTP id adf61e73a8af0-3cc5542340bmr10139073637.33.1786647531022; Thu, 13 Aug 2026 11:58:51 -0700 (PDT) Received: from localhost ([2a03:2880:9ff:66::]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-31ebc75d8d6sm11208926eec.4.2026.08.13.11.58.50 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 13 Aug 2026 11:58:50 -0700 (PDT) From: Ziyang Men To: Tejun Heo , Johannes Weiner , =?UTF-8?q?Michal=20Koutn=C3=BD?= , Ingo Molnar , Peter Zijlstra , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , Eduard Zingerman , Kumar Kartikeya Dwivedi Cc: Ben Segall , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Mel Gorman , Valentin Schneider , K Prateek Nayak , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Shuah Khan , Roman Gushchin , Shakeel Butt , JP Kobryn , Mykola Lysenko , Ziyang Men , kernel-team@meta.com, bpf@vger.kernel.org, cgroups@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org Subject: [PATCH 1/2] cgroup, sched: add BPF kfuncs to read a cpu cgroup's stats Date: Thu, 13 Aug 2026 11:58:45 -0700 Message-ID: <20260813185846.1216892-2-ziyang.meme@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260813185846.1216892-1-ziyang.meme@gmail.com> References: <20260813185846.1216892-1-ziyang.meme@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" This series adds bpf kfuncs for the cgroup CPU controller, following the memory controller kfuncs in mm/bpf_memcontrol.c. Collecting cgroup statistics is expensive: the existing method is to open and parse a cgroup file. memcg already has an efficient alternative through BPF; this series extends that idea to cpu. Design: - Leave reading the CFS bandwidth counters to the BPF program. They are plain fields of tg->cfs_bandwidth, so they need no kernel code. - Add one kfunc to compute the throttled time. This is necessary because it is a sum over every possible cpu, which a user cannot do itself. - The bpf_cpu_cgroup_cputime() returns all five base CPU-time values in one call with one cputime_adjust(). The only part it touches the scheduler part is to discard the static for throttled_time_self() in order to use externally. The two kfuncs that take the rstat lock are KF_SLEEPABLE following idea in the mm/bpf_memcontrol.c Suggested-by: Shakeel Butt Assisted-by: Claude:claude-opus-5 Signed-off-by: Ziyang Men --- include/linux/cgroup.h | 15 +++++++ kernel/cgroup/Makefile | 2 + kernel/cgroup/bpf_cpu.c | 80 +++++++++++++++++++++++++++++++++ kernel/cgroup/cgroup-internal.h | 3 ++ kernel/cgroup/rstat.c | 42 +++++++++++++++++ kernel/sched/core.c | 2 +- 6 files changed, 143 insertions(+), 1 deletion(-) create mode 100644 kernel/cgroup/bpf_cpu.c diff --git a/include/linux/cgroup.h b/include/linux/cgroup.h index f2aa46a4f871..d2a6b5efad51 100644 --- a/include/linux/cgroup.h +++ b/include/linux/cgroup.h @@ -923,4 +923,19 @@ struct cgroup *task_get_cgroup1(struct task_struct *ts= k, int hierarchy_id); =20 struct cgroup_of_peak *of_peak(struct kernfs_open_file *of); =20 +/* A cgroup's base CPU-time counters in microseconds, as cpu.stat prints t= hem */ +struct cpu_cgroup_cputime { + u64 usage_usec; + u64 user_usec; + u64 system_usec; + u64 nice_usec; + u64 forceidle_usec; /* 0 without CONFIG_SCHED_CORE */ +}; + +/* A task_group's own throttled time in nanoseconds; see cpu.stat.local */ +struct task_group; +#ifdef CONFIG_CFS_BANDWIDTH +u64 throttled_time_self(struct task_group *tg); +#endif + #endif /* _LINUX_CGROUP_H */ diff --git a/kernel/cgroup/Makefile b/kernel/cgroup/Makefile index ede31601a363..0ba59b7eef48 100644 --- a/kernel/cgroup/Makefile +++ b/kernel/cgroup/Makefile @@ -1,6 +1,8 @@ # SPDX-License-Identifier: GPL-2.0 obj-y :=3D cgroup.o rstat.o namespace.o cgroup-v1.o freezer.o =20 +obj-$(CONFIG_BPF_SYSCALL) +=3D bpf_cpu.o + obj-$(CONFIG_CGROUP_FREEZER) +=3D legacy_freezer.o obj-$(CONFIG_CGROUP_PIDS) +=3D pids.o obj-$(CONFIG_CGROUP_RDMA) +=3D rdma.o diff --git a/kernel/cgroup/bpf_cpu.c b/kernel/cgroup/bpf_cpu.c new file mode 100644 index 000000000000..6eb89c8e84fd --- /dev/null +++ b/kernel/cgroup/bpf_cpu.c @@ -0,0 +1,80 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * CPU Controller-related BPF kfuncs + * + * bpf_cpu_cgroup_cputime() is defined in rstat.c, which owns the locking = it + * needs, and only registered here. + * + * Author: Ziyang Men + */ + +#include +#include +#include + +#include "cgroup-internal.h" + +__bpf_kfunc_start_defs(); + +/** + * bpf_cpu_cgroup_flush_stats - Flush a cgroup's base CPU-time statistics + * @cgrp: cgroup to flush + * + * Propagate the cgroup's base CPU-time statistics up the cgroup tree. + */ +__bpf_kfunc void bpf_cpu_cgroup_flush_stats(struct cgroup *cgrp) +{ + css_rstat_flush(&cgrp->self); +} + +/** + * bpf_cpu_cgroup_throttled_self - Read a cgroup's own throttled time + * @cgrp: cgroup to read from + * + * Return: The throttled time in microseconds, or 0 if config is off. + */ +__bpf_kfunc u64 bpf_cpu_cgroup_throttled_self(struct cgroup *cgrp) +{ +/* cpu_cgrp_id needs the cpu controller, which CFS bandwidth depends on */ +#ifdef CONFIG_CFS_BANDWIDTH + struct cgroup_subsys_state *css; + + guard(rcu)(); + + css =3D rcu_dereference(cgrp->subsys[cpu_cgrp_id]); + if (!css) + return 0; + + return div_u64(throttled_time_self((struct task_group *)css), + NSEC_PER_USEC); +#else + return 0; +#endif +} + +__bpf_kfunc_end_defs(); + +/* KF_SLEEPABLE keeps the rstat spinlock out of NMI */ +BTF_KFUNCS_START(bpf_cpu_cgroup_kfunc_ids) +BTF_ID_FLAGS(func, bpf_cpu_cgroup_flush_stats, KF_SLEEPABLE) +BTF_ID_FLAGS(func, bpf_cpu_cgroup_cputime, KF_SLEEPABLE) +BTF_ID_FLAGS(func, bpf_cpu_cgroup_throttled_self) +BTF_KFUNCS_END(bpf_cpu_cgroup_kfunc_ids) + +static const struct btf_kfunc_id_set bpf_cpu_cgroup_kfunc_set =3D { + .owner =3D THIS_MODULE, + .set =3D &bpf_cpu_cgroup_kfunc_ids, +}; + +static int __init bpf_cpu_cgroup_kfunc_init(void) +{ + int err; + + err =3D register_btf_kfunc_id_set(BPF_PROG_TYPE_UNSPEC, + &bpf_cpu_cgroup_kfunc_set); + if (err) + pr_warn("error while registering cpu cgroup kfuncs: %d\n", err); + + return err; +} +late_initcall(bpf_cpu_cgroup_kfunc_init); diff --git a/kernel/cgroup/cgroup-internal.h b/kernel/cgroup/cgroup-interna= l.h index 58797123b752..65f5b6318289 100644 --- a/kernel/cgroup/cgroup-internal.h +++ b/kernel/cgroup/cgroup-internal.h @@ -271,6 +271,9 @@ int css_rstat_init(struct cgroup_subsys_state *css); void css_rstat_exit(struct cgroup_subsys_state *css); int ss_rstat_init(struct cgroup_subsys *ss); void cgroup_base_stat_cputime_show(struct seq_file *seq); +#ifdef CONFIG_BPF_SYSCALL +void bpf_cpu_cgroup_cputime(struct cgroup *cgrp, struct cpu_cgroup_cputime= *out); +#endif =20 /* * namespace.c diff --git a/kernel/cgroup/rstat.c b/kernel/cgroup/rstat.c index de816a43db9f..f9e30719068e 100644 --- a/kernel/cgroup/rstat.c +++ b/kernel/cgroup/rstat.c @@ -752,6 +752,48 @@ void cgroup_base_stat_cputime_show(struct seq_file *se= q) cgroup_force_idle_show(seq, &bstat); } =20 +#ifdef CONFIG_BPF_SYSCALL + +__bpf_kfunc_start_defs(); + +/** + * bpf_cpu_cgroup_cputime - Read a cgroup's base CPU-time data + * @cgrp: cgroup to read from + * @out: the data in microseconds. Zero it first: the verifier reads the + * whole struct. + * + * Adjust once and fill all values. + */ +__bpf_kfunc void bpf_cpu_cgroup_cputime(struct cgroup *cgrp, + struct cpu_cgroup_cputime *out) +{ + struct cgroup_base_stat bstat; + + if (cgroup_parent(cgrp)) { + __css_rstat_lock(&cgrp->self, -1); + bstat =3D cgrp->bstat; + cputime_adjust(&cgrp->bstat.cputime, &cgrp->prev_cputime, + &bstat.cputime.utime, &bstat.cputime.stime); + __css_rstat_unlock(&cgrp->self, -1); + } else { + root_cgroup_cputime(&bstat); + } + + out->usage_usec =3D div_u64(bstat.cputime.sum_exec_runtime, NSEC_PER_USEC= ); + out->user_usec =3D div_u64(bstat.cputime.utime, NSEC_PER_USEC); + out->system_usec =3D div_u64(bstat.cputime.stime, NSEC_PER_USEC); + out->nice_usec =3D div_u64(bstat.ntime, NSEC_PER_USEC); +#ifdef CONFIG_SCHED_CORE + out->forceidle_usec =3D div_u64(bstat.forceidle_sum, NSEC_PER_USEC); +#else + out->forceidle_usec =3D 0; +#endif +} + +__bpf_kfunc_end_defs(); + +#endif /* CONFIG_BPF_SYSCALL */ + /* Add bpf kfuncs for css_rstat_updated() and css_rstat_flush() */ BTF_KFUNCS_START(bpf_rstat_kfunc_ids) BTF_ID_FLAGS(func, css_rstat_updated) diff --git a/kernel/sched/core.c b/kernel/sched/core.c index 96226707c2f6..75735e0e81ef 100644 --- a/kernel/sched/core.c +++ b/kernel/sched/core.c @@ -10027,7 +10027,7 @@ static int cpu_cfs_stat_show(struct seq_file *sf, v= oid *v) return 0; } =20 -static u64 throttled_time_self(struct task_group *tg) +u64 throttled_time_self(struct task_group *tg) { int i; u64 total =3D 0; --=20 2.53.0-Meta From nobody Tue Sep 29 01:18:19 2026 Received: from mail-pl1-f172.google.com (mail-pl1-f172.google.com [209.85.214.172]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A5B614968F0 for ; Thu, 13 Aug 2026 18:58:53 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.172 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786647538; cv=none; b=VJqmbahV2qRNCZ2Agd1Mf6kRJo5vGZiiT48PFkRODhRR0t8Sf6HW0wBWFupqe9bhd0l9lHfCkz6nQTaRze1kmnWWWnlNXEdo5IIYdlrm503DM1ZWHRAfBzs8VVYcJFMtoas9M/1+PZap+sjyPPiVfUG6SDMTWhmNoBgwiRQdUnU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786647538; c=relaxed/simple; bh=GOCeIpJRgFyYHMOfCqGUNzxzujcQWvp3QaqDw0/DsqY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=VQsJO42slOb3EqzpfZozS717+dsk99pf/ru02uORLWZkuFMOGvIrAQOUo5PSFUAc++hahndzmggynbifer7LXWHZq2Ght6puzTTgF0bWFhycHpOK3v9dSbzNo1hjC2WINZ+DKHCYTRwiYtP2reYt9yCt3gN1Lmyqwy5cs7dUP6s= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=f0mz3B09; arc=none smtp.client-ip=209.85.214.172 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="f0mz3B09" Received: by mail-pl1-f172.google.com with SMTP id d9443c01a7336-2cc61541f8cso18103945ad.0 for ; Thu, 13 Aug 2026 11:58:53 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786647533; x=1787252333; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=r8vCfOJqlyeUUrhQ0TEVig3puBrFV2IWvi7C6KR3I1g=; b=f0mz3B09o42+sgDyicQmHskQV9bYAolwzXKpc9qBHM9Xti973wqzKUBSNZdRdDmMMr U8POGVoOCWOTq2sDp8KA1pnfaOebuz+Xrm+ca1Zn46vz9bSeO3KA939NXpOLbfkVGIdE xJQ5KNUNj0+egcwiihBd8dZcQmE9dNVRoFkgM+Rs3K4Vxou1TS4oXmIqp3jjM0VIbCl2 giZuvLtzFMYlFnUKOI84SOBE/RikV71OUt00NdjgLxreJaEr9NPGqCNZrrvnPcrrJxyI qK+BDyHLhb3bl0tDto9f1QZ7NjDQtGBMa8jaISbuDKi5+FTxI0oJSIzgPhvZH9FvDJHv xUhA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786647533; x=1787252333; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=r8vCfOJqlyeUUrhQ0TEVig3puBrFV2IWvi7C6KR3I1g=; b=lcSP3ULiw0lS/W2kZR8JOqexRk3Zqxssps2x86Rq/id1bGqZ4o2Ao39HH3D4Jz7UQr zCrDmxhRwp4Xo5rhQdAc8M90oBvyUVDiO0vSN6NCssfN8A8nDfDQfYHK7pyxJ1yA+Ace trCM2dSrjguCDPJg7EQcNAYpxY+3n1+Lf4LZ0cyOfLQEojX6xh9I4WJ17biQUj+Haw5p dY58W6Bn15fyN50iV3TTK3CfBowsqn5gzz3WeLJt3dJVi9dMOKdumUGRpcjZ0tuwwqIH /IdIegCzvwAO/Ht1XQA2P4f1J48iB5ECwuE2NKYiGQHY0cQLpGcCgGXGgqb3+0iFyh+H /sJQ== X-Forwarded-Encrypted: i=1; AHgh+RoXTXEr0S13TuEdHDZRhYkPLpBJdh1HQz7h5IjDCL5cxrVLeEIaQnB9Nj5zXrtCQQNZDBwKKt1X1ddx+JQ=@vger.kernel.org X-Gm-Message-State: AOJu0YyCanLoqyTjR/kBL0FwpVVAhVNsZr0On2oOJ5NTdWLlcZnzjqPZ e5vzHOnEuG/zzOESU4JYID0KlqYpZ9NbXVGa8COCBT3WPKbwnDxN6XNB X-Gm-Gg: AR+sD10IMzQZpyvc4CpWqdR67TDow4bqVWnQchiyVmSBEcBh+VsC4dR1VJlAEtTYs2j Gd/4S0ikLZsj2OC3eMpYKffDcxsRdaa85A+4CbT6ln7kwD+GUojx+mbtWF2cogVNuv8RPSjMECT L6zFpMvBCageRLhVHytrBgl60Lwn7O4Bi0piUaY+tMUyu3vv9VuAzki8AN31+4XiEjFo/IiJk+J tJrX8tLnKqFs09Nd/ruGQ7JJf+S+j1y+uMzDNVRaQAc6rHHzk8w8AKhFnSi0FaCV3ErGFwvK4rm XS+sjN6Mgw4PzlxRKSeu2NyxhXs3oCpfGBWu3UtB2A22f+tH9lAISePiiN79NGqyioIxg29G5MK Af556q9YdZPhdAwdWyWjRjDRZo6X0+wevIf/Q+eUn/C1sLxLFMm5Gfx0LZBgNULS6qq4CkDPq10 U8GU6DCvcUGBK9fl19qEJBbRyZWctVdIEl1EKbTsCt0ZijO57UHE+LRmE= X-Received: by 2002:a17:902:f64e:b0:2cf:70d2:da7c with SMTP id d9443c01a7336-2d37f77fdfemr99897105ad.12.1786647532683; Thu, 13 Aug 2026 11:58:52 -0700 (PDT) Received: from localhost ([2a03:2880:9ff:63::]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-31eb99e2556sm10969714eec.0.2026.08.13.11.58.51 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 13 Aug 2026 11:58:52 -0700 (PDT) From: Ziyang Men To: Tejun Heo , Johannes Weiner , =?UTF-8?q?Michal=20Koutn=C3=BD?= , Ingo Molnar , Peter Zijlstra , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , Eduard Zingerman , Kumar Kartikeya Dwivedi Cc: Ben Segall , Juri Lelli , Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Mel Gorman , Valentin Schneider , K Prateek Nayak , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Shuah Khan , Roman Gushchin , Shakeel Butt , JP Kobryn , Mykola Lysenko , Ziyang Men , kernel-team@meta.com, bpf@vger.kernel.org, cgroups@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org Subject: [PATCH 2/2] selftests/bpf: add cgroup_iter_cpu test for cpu cgroup kfuncs Date: Thu, 13 Aug 2026 11:58:46 -0700 Message-ID: <20260813185846.1216892-3-ziyang.meme@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260813185846.1216892-1-ziyang.meme@gmail.com> References: <20260813185846.1216892-1-ziyang.meme@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add cgroup_iter_cpu, a selftest for the added CPU controller BPF kfuncs. The userspace side runs a CPU hog in a test cgroup with cpu.max settled then: - checks the CPU-time and throttling counters are nonzero, - compares whether all values the program read are same as those reading from cgroup file. CONFIG_CGROUP_SCHED, CONFIG_FAIR_GROUP_SCHED and CONFIG_CFS_BANDWIDTH are added to the test config. Tests passed on v7.2-rc5. Suggested-by: Shakeel Butt Assisted-by: Claude:claude-opus-5 Signed-off-by: Ziyang Men --- tools/testing/selftests/bpf/cgroup_iter_cpu.h | 22 ++ tools/testing/selftests/bpf/config | 3 + .../bpf/prog_tests/cgroup_iter_cpu.c | 259 ++++++++++++++++++ .../selftests/bpf/progs/cgroup_iter_cpu.c | 53 ++++ 4 files changed, 337 insertions(+) create mode 100644 tools/testing/selftests/bpf/cgroup_iter_cpu.h create mode 100644 tools/testing/selftests/bpf/prog_tests/cgroup_iter_cpu.c create mode 100644 tools/testing/selftests/bpf/progs/cgroup_iter_cpu.c diff --git a/tools/testing/selftests/bpf/cgroup_iter_cpu.h b/tools/testing/= selftests/bpf/cgroup_iter_cpu.h new file mode 100644 index 000000000000..74599a5c0e4d --- /dev/null +++ b/tools/testing/selftests/bpf/cgroup_iter_cpu.h @@ -0,0 +1,22 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* Copyright (c) 2025 Meta Platforms, Inc. and affiliates. */ +#ifndef __CGROUP_ITER_CPU_H +#define __CGROUP_ITER_CPU_H + +struct cpu_query { + /* base cpu time, from cpu.stat */ + __u64 usage_usec; + __u64 user_usec; + __u64 system_usec; + __u64 nice_usec; + __u64 forceidle_usec; + /* CFS bandwidth throttling, from cpu.stat and cpu.stat.local */ + __u64 nr_periods; + __u64 nr_throttled; + __u64 throttled_usec; + __u64 nr_bursts; + __u64 burst_usec; + __u64 throttled_self_usec; +}; + +#endif /* __CGROUP_ITER_CPU_H */ diff --git a/tools/testing/selftests/bpf/config b/tools/testing/selftests/b= pf/config index ea7044f30adc..482b40dde2f9 100644 --- a/tools/testing/selftests/bpf/config +++ b/tools/testing/selftests/bpf/config @@ -11,6 +11,9 @@ CONFIG_BPF_STREAM_PARSER=3Dy CONFIG_BPF_SYSCALL=3Dy # CONFIG_BPF_UNPRIV_DEFAULT_OFF is not set CONFIG_CGROUP_BPF=3Dy +CONFIG_CGROUP_SCHED=3Dy +CONFIG_FAIR_GROUP_SCHED=3Dy +CONFIG_CFS_BANDWIDTH=3Dy CONFIG_CRYPTO_HMAC=3Dy CONFIG_CRYPTO_SHA256=3Dy CONFIG_CRYPTO_USER_API=3Dy diff --git a/tools/testing/selftests/bpf/prog_tests/cgroup_iter_cpu.c b/too= ls/testing/selftests/bpf/prog_tests/cgroup_iter_cpu.c new file mode 100644 index 000000000000..cd7e92ababfb --- /dev/null +++ b/tools/testing/selftests/bpf/prog_tests/cgroup_iter_cpu.c @@ -0,0 +1,259 @@ +// SPDX-License-Identifier: GPL-2.0 +/* Copyright (c) 2025 Meta Platforms, Inc. and affiliates. */ +#include +#include +#include +#include +#include +#include +#include +#include "cgroup_helpers.h" +#include "cgroup_iter_cpu.h" +#include "cgroup_iter_cpu.skel.h" + +static int read_stats(struct bpf_link *link) +{ + int fd, ret =3D 0; + ssize_t bytes; + + fd =3D bpf_iter_create(bpf_link__fd(link)); + if (!ASSERT_OK_FD(fd, "bpf_iter_create")) + return 1; + + bytes =3D read(fd, NULL, 0); + if (!ASSERT_EQ(bytes, 0, "read fd")) + ret =3D 1; + + close(fd); + return ret; +} + +/* Read cgroup file @name into @buf. */ +static int read_cgroup_file(int cgroup_fd, const char *name, char *buf, + size_t size) +{ + ssize_t n; + int fd; + + fd =3D openat(cgroup_fd, name, O_RDONLY); + if (fd < 0) + return -1; + n =3D read(fd, buf, size - 1); + close(fd); + if (n <=3D 0) + return -1; + buf[n] =3D '\0'; + return 0; +} + +/* Parse the "cpu.stat" file into @out. */ +static int parse_cpu_stat(int cgroup_fd, struct cpu_query *out) +{ + char buf[4096], *line, *sp; + unsigned long long v; + + if (read_cgroup_file(cgroup_fd, "cpu.stat", buf, sizeof(buf))) + return -1; + + for (line =3D strtok_r(buf, "\n", &sp); line; + line =3D strtok_r(NULL, "\n", &sp)) { + if (sscanf(line, "usage_usec %llu", &v) =3D=3D 1) + out->usage_usec =3D v; + else if (sscanf(line, "user_usec %llu", &v) =3D=3D 1) + out->user_usec =3D v; + else if (sscanf(line, "system_usec %llu", &v) =3D=3D 1) + out->system_usec =3D v; + else if (sscanf(line, "nice_usec %llu", &v) =3D=3D 1) + out->nice_usec =3D v; + else if (sscanf(line, "core_sched.force_idle_usec %llu", &v) =3D=3D 1) + out->forceidle_usec =3D v; + else if (sscanf(line, "nr_periods %llu", &v) =3D=3D 1) + out->nr_periods =3D v; + else if (sscanf(line, "nr_throttled %llu", &v) =3D=3D 1) + out->nr_throttled =3D v; + else if (sscanf(line, "throttled_usec %llu", &v) =3D=3D 1) + out->throttled_usec =3D v; + else if (sscanf(line, "nr_bursts %llu", &v) =3D=3D 1) + out->nr_bursts =3D v; + else if (sscanf(line, "burst_usec %llu", &v) =3D=3D 1) + out->burst_usec =3D v; + } + return 0; +} + +/* + * Parse the "cpu.stat.local" file into @out. + */ +static int parse_cpu_stat_local(int cgroup_fd, struct cpu_query *out) +{ + unsigned long long v; + char buf[256]; + + if (read_cgroup_file(cgroup_fd, "cpu.stat.local", buf, sizeof(buf))) + return -1; + if (sscanf(buf, "throttled_usec %llu", &v) !=3D 1) + return -1; + out->throttled_self_usec =3D v; + return 0; +} + +/* Read file value the bpf program reads. */ +static int parse_stats(int cgroup_fd, struct cpu_query *out, bool have_bw) +{ + if (parse_cpu_stat(cgroup_fd, out)) + return -1; + if (have_bw && parse_cpu_stat_local(cgroup_fd, out)) + return -1; + return 0; +} + +/* + * Check whether this kernel accounts CFS bandwidth. + */ +static bool cgroup_has_bw_stat(int cgroup_fd) +{ + char buf[4096]; + + if (read_cgroup_file(cgroup_fd, "cpu.stat", buf, sizeof(buf))) + return false; + return strstr(buf, "nr_periods "); +} + +/* Fork a child that spins in the current cgroup, kill it if the test exit= s. */ +static pid_t spawn_cpu_hog(void) +{ + pid_t pid =3D fork(); + + if (pid =3D=3D 0) { + prctl(PR_SET_PDEATHSIG, SIGKILL); + while (1) + ; + } + return pid; +} + +void test_cgroup_iter_cpu(void) +{ + char *cgroup_rel_path =3D "/cgroup_iter_cpu_test"; + struct cgroup_iter_cpu *skel; + struct cpu_query *q; + struct bpf_link *link; + bool wrote_max, have_bw; + int cgroup_fd; + pid_t hog; + + cgroup_fd =3D cgroup_setup_and_join(cgroup_rel_path); + if (!ASSERT_OK_FD(cgroup_fd, "cgroup_setup_and_join")) + return; + + wrote_max =3D !write_cgroup_file(cgroup_rel_path, "cpu.max", "10000 10000= 0"); + + skel =3D cgroup_iter_cpu__open_and_load(); + if (!ASSERT_OK_PTR(skel, "cgroup_iter_cpu__open_and_load")) + goto cleanup_cgroup_fd; + + DECLARE_LIBBPF_OPTS(bpf_iter_attach_opts, opts); + union bpf_iter_link_info linfo =3D { + .cgroup.cgroup_fd =3D cgroup_fd, + .cgroup.order =3D BPF_CGROUP_ITER_SELF_ONLY, + }; + opts.link_info =3D &linfo; + opts.link_info_len =3D sizeof(linfo); + + link =3D bpf_program__attach_iter(skel->progs.cgroup_cpu_query, &opts); + if (!ASSERT_OK_PTR(link, "bpf_program__attach_iter")) + goto cleanup_skel; + + q =3D &skel->data_query->cpu_query; + + hog =3D spawn_cpu_hog(); + if (!ASSERT_GT(hog, 0, "spawn_cpu_hog")) + goto cleanup_link; + + sleep(1); + + /* Run the bpf program before anything here reads cpu.stat. */ + if (!ASSERT_OK(read_stats(link), "read stats")) + goto cleanup_hog; + + have_bw =3D wrote_max && cgroup_has_bw_stat(cgroup_fd); + + if (test__start_subtest("cgroup_iter_cpu__cputime")) { + ASSERT_GT(q->usage_usec, 0, "usage_usec"); + ASSERT_GT(q->user_usec + q->system_usec, 0, "user+system_usec"); + } + if (test__start_subtest("cgroup_iter_cpu__throttling")) { + if (!have_bw) { + test__skip(); + } else { + ASSERT_GT(q->nr_periods, 0, "nr_periods"); + ASSERT_GT(q->nr_throttled, 0, "nr_throttled"); + ASSERT_GT(q->throttled_usec, 0, "throttled_usec"); + ASSERT_GT(q->throttled_self_usec, 0, "throttled_self_usec"); + } + } + + /* + * cpu.stat cputime grows on every tick a task in the cgroup runs, so + * stop them all before comparing + */ + if (test__start_subtest("cgroup_iter_cpu__match")) { + struct cpu_query filev =3D {}; + int i, stable =3D 0; + + kill(hog, SIGSTOP); + waitpid(hog, NULL, WUNTRACED); + if (!ASSERT_OK(join_root_cgroup(), "join_root_cgroup")) + goto cleanup_hog; + + /* + * The period timer keeps adding to nr_periods for a while + * after the hog stops + */ + for (i =3D 0; i < 20; i++) { + struct cpu_query before =3D {}, after =3D {}; + + if (!ASSERT_OK(parse_stats(cgroup_fd, &before, have_bw), "cpu.stat") || + !ASSERT_OK(read_stats(link), "read stats") || + !ASSERT_OK(parse_stats(cgroup_fd, &after, have_bw), "cpu.stat")) + goto cleanup_hog; + + if (!memcmp(&before, &after, sizeof(before))) { + filev =3D before; + stable =3D 1; + break; + } + usleep(100000); + } + + if (!ASSERT_TRUE(stable, "cpu.stat stable")) + goto cleanup_hog; + + ASSERT_EQ(q->usage_usec, filev.usage_usec, "usage_usec"); + ASSERT_EQ(q->user_usec, filev.user_usec, "user_usec"); + ASSERT_EQ(q->system_usec, filev.system_usec, "system_usec"); + ASSERT_EQ(q->nice_usec, filev.nice_usec, "nice_usec"); + ASSERT_EQ(q->forceidle_usec, filev.forceidle_usec, "forceidle_usec"); + + if (have_bw) { + ASSERT_EQ(q->nr_periods, filev.nr_periods, "nr_periods"); + ASSERT_EQ(q->nr_throttled, filev.nr_throttled, "nr_throttled"); + ASSERT_EQ(q->throttled_usec, filev.throttled_usec, "throttled_usec"); + ASSERT_EQ(q->nr_bursts, filev.nr_bursts, "nr_bursts"); + ASSERT_EQ(q->burst_usec, filev.burst_usec, "burst_usec"); + ASSERT_EQ(q->throttled_self_usec, filev.throttled_self_usec, + "throttled_self_usec"); + } + } + +cleanup_hog: + kill(hog, SIGKILL); + waitpid(hog, NULL, 0); +cleanup_link: + bpf_link__destroy(link); +cleanup_skel: + cgroup_iter_cpu__destroy(skel); +cleanup_cgroup_fd: + close(cgroup_fd); + cleanup_cgroup_environment(); +} diff --git a/tools/testing/selftests/bpf/progs/cgroup_iter_cpu.c b/tools/te= sting/selftests/bpf/progs/cgroup_iter_cpu.c new file mode 100644 index 000000000000..6a288a00c25a --- /dev/null +++ b/tools/testing/selftests/bpf/progs/cgroup_iter_cpu.c @@ -0,0 +1,53 @@ +// SPDX-License-Identifier: GPL-2.0 +/* Copyright (c) 2025 Meta Platforms, Inc. and affiliates. */ +#include +#include +#include +#include "cgroup_iter_cpu.h" + +char _license[] SEC("license") =3D "GPL"; + +struct cpu_query cpu_query SEC(".data.query"); + +SEC("iter.s/cgroup") +int cgroup_cpu_query(struct bpf_iter__cgroup *ctx) +{ + struct cpu_cgroup_cputime ct =3D {}; + struct cgroup *cgrp =3D ctx->cgroup; + struct cgroup_subsys_state *css; + struct task_group *tg; + + if (!cgrp) + return 1; + + bpf_cpu_cgroup_flush_stats(cgrp); + bpf_cpu_cgroup_cputime(cgrp, &ct); + + cpu_query.usage_usec =3D ct.usage_usec; + cpu_query.user_usec =3D ct.user_usec; + cpu_query.system_usec =3D ct.system_usec; + cpu_query.nice_usec =3D ct.nice_usec; + cpu_query.forceidle_usec =3D ct.forceidle_usec; + + bpf_rcu_read_lock(); + css =3D cgrp->subsys[cpu_cgrp_id]; + tg =3D (struct task_group *)css; + if (tg && bpf_core_field_exists(tg->cfs_bandwidth.nr_periods)) { + cpu_query.nr_periods =3D + (__u32)BPF_CORE_READ(tg, cfs_bandwidth.nr_periods); + cpu_query.nr_throttled =3D + (__u32)BPF_CORE_READ(tg, cfs_bandwidth.nr_throttled); + cpu_query.throttled_usec =3D + BPF_CORE_READ(tg, cfs_bandwidth.throttled_time) / 1000; + cpu_query.nr_bursts =3D + (__u32)BPF_CORE_READ(tg, cfs_bandwidth.nr_burst); + cpu_query.burst_usec =3D + BPF_CORE_READ(tg, cfs_bandwidth.burst_time) / 1000; + } + bpf_rcu_read_unlock(); + + /* a sum over every possible cpu, so the test uses a kfunc */ + cpu_query.throttled_self_usec =3D bpf_cpu_cgroup_throttled_self(cgrp); + + return 0; +} --=20 2.53.0-Meta