From nobody Tue Sep 29 13:19:45 2026 Received: from mail-pg1-f175.google.com (mail-pg1-f175.google.com [209.85.215.175]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 29A6A3A544A for ; Fri, 7 Aug 2026 19:37:37 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.175 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786131459; cv=none; b=ojy2PF6x/lUtrsjKOGjwsKk/23rHfRQrJZT+ne6tq9g5iBJaFEz+UUUGtqJJ64MB9cmkbXnQD9FUogs6jUpD8a4uFlW15y05tzj2/G7n7L1zDzWP3pnh/D/z8si56oZPIQUBMPVQMVciIdwY5qH51Hvuj3IZQ9MNX2HH18n9HbE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786131459; c=relaxed/simple; bh=LY1zudeQGpKe93Wm1RjYLV4w4I6hRdLzm8Tu6l8oMRA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=rHDBEx8VsfQfZgS/UjtG7gBwIh1WuRb2wiHz0aQnvhBbLafCkcNOZlziEbpapGqjf7U3ZNm11YhURcprf4L3fkt4M2zP0+1WA0e6z19rT+2Z0qt6MNgpuWtwkaHtBQu/rZ+opaGVLkGMiIjdJ0YBqSeZK+M+rbADCLgj2YGmtxY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=bOoEnoKB; arc=none smtp.client-ip=209.85.215.175 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="bOoEnoKB" Received: by mail-pg1-f175.google.com with SMTP id 41be03b00d2f7-cbb8b54fcf8so3382504a12.0 for ; Fri, 07 Aug 2026 12:37:37 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786131456; x=1786736256; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=oyaDPzZxAYZ6dkJ1RtbEqn7iDEULJIHV8/+9+Tra4y8=; b=bOoEnoKB+sUbVvsTr/PHFYuU2ucycPs6Xse04sWmn42rdI6T6ORcpgPO1XQTCw0v/I eyEnfdiVsvf2gcQ2NR5r2OAm2rS6Je34MlzNFGYXNsmEKtBN5skgQsd5JKh2bmOEHMSO fwdUhggBadPEIIyEAjFbTkwHzbFRvrQJre7IwmeZwyGK7/vc0zDM2PszMoxrFtD4Rtft OA3YlddsEb+nDOIHV1i7rJUo8m2S3L0cV/5Oa2qLX6i4x2BU51OIS64y0vSWHJEHumoE EURFTVy9UXc13eS6pWE8nj1WYHFWGjYje/88JROJbocrUcyL+uQHCL7UAhdyTmbbFFAq Q7zA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786131456; x=1786736256; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=oyaDPzZxAYZ6dkJ1RtbEqn7iDEULJIHV8/+9+Tra4y8=; b=XNYZlMTbTDXesxi+MoVTfsS2FNc8PJ3nAAez1Ob89o9KyW10SIxm3QnQXFQUM3Yn6i lGcftJqVZgrgYj3PnnYHLFz0qgDzb/2qP3n5xmL5ShEMY7htlCxEXs4zAwOLtcnZ043Z pBTvf9tih9npy0j5/BwbBIlw4h2ZwnfzYivuT7il8Hb2UAWJT94douehdDkw7fISBJf6 IeAdwKXA3/Nt6P+a1chC4DdS/AYuEXux50+Fc3e/cVPwJTpUtJb8G3jTYPFuYbbZBDG5 2vBtwyJNIBv+KzCgEOdbx7pM0i9qezPGtAhKi2Fojvat4zrL4TulWMjRJaNLU5LdHXxv EKuA== X-Forwarded-Encrypted: i=1; AHgh+RpsMzKvlojCNJTkLqf/kb2z4NnxYZjMza6vyQAmTEMLFJUmNmy0AbvRvm43Paeld0aZCYTqWQet/ZwBg/k=@vger.kernel.org X-Gm-Message-State: AOJu0YwP/WtMBC7taLj23Kzx0CPo+mleKlLmNZIHlFT+aPthctUSL28l HA789krrkUGtnJalx43X4ikpd9vbAp37II/zecC/KPpgtTydMbZcPhm4 X-Gm-Gg: AR+sD13Yk0EjiI1ZkuDKRP8uOuX0+vG2/1hyiRRLzok+wsCwNwxEDeLv0vMz5KCC288 4Igc1t5HcYdyk+O2WbJJNowFlFIeLRSOxkqTqvv49E4QgvcYk2+624AnxfF4zJXvoUKvqCNpqqu QWL32VYB3ORydTBslRbBKqmjmCdUx1DK4BagI6b6y82RuWnvZgb/7N9x/qohkQD9s2E9rwKRX4d WCbFsBZyJfxKWVH1kUuJrP8TkXJnzbL2rIuQna0TSMd/KJzpc+0PFD7G4xDGu+Rd6zwTugVQ0PJ nNRXC6sy+GG0rbBYaWn1etXNgARPaD/FxcuOZncTlpLMp9unhvaVkWB9CIWKzkesRU5aG884aQr MhoHFgXljl6EBWtW7tBBtkBZKiPMyMfLBdoJtmZwUps4PgbiJLp/KTvxXTvEIOLwjJT0yE+Ph00 LThWHcYJ1CqgEFVv0/Axh5zq4UGeY8nxkIDPP5znCijV4uHBb73sOXRnI= X-Received: by 2002:a05:6a20:6f02:b0:3c4:511e:26ab with SMTP id adf61e73a8af0-3cb85ded984mr34425035637.3.1786131456465; Fri, 07 Aug 2026 12:37:36 -0700 (PDT) Received: from localhost ([2a03:2880:9ff:43::]) by smtp.gmail.com with ESMTPSA id a92af1059eb24-1410199bfd8sm9547703c88.2.2026.08.07.12.37.35 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 07 Aug 2026 12:37:36 -0700 (PDT) From: Ziyang Men To: Jens Axboe , Tejun Heo , Josef Bacik , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , Eduard Zingerman , Kumar Kartikeya Dwivedi Cc: Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Shuah Khan , Johannes Weiner , =?UTF-8?q?Michal=20Koutn=C3=BD?= , Roman Gushchin , Shakeel Butt , JP Kobryn , Mykola Lysenko , kernel-team@meta.com, linux-block@vger.kernel.org, bpf@vger.kernel.org, cgroups@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, Ziyang Men Subject: [PATCH 1/2] block: add BPF kfuncs to read blkcg io.stat Date: Fri, 7 Aug 2026 12:37:31 -0700 Message-ID: <20260807193732.4073299-2-ziyang.meme@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260807193732.4073299-1-ziyang.meme@gmail.com> References: <20260807193732.4073299-1-ziyang.meme@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Expose the block I/O controller's per-device statistics to BPF, mirroring the memory controller kfuncs in mm/bpf_memcontrol.c. A BPF program gets a blkcg from a cgroup's css with bpf_get_blkcg() (or bpf_get_root_blkcg() for the root), flushes the stats with bpf_blkcg_flush_stats(), then walks the cgroup's per-device blkgs with the bpf_iter_blkg open-coded iterator and reads each device's counters with bpf_blkg_iostat_bytes() and bpf_blkg_iostat_ios(). bpf_blkg_dev() returns the device id for labelling. The reference is released with bpf_put_blkcg(). Unlike the memory controller, blkcg keeps one blkg (and one io.stat line) per block device, so the reader kfuncs take a blkg and the iterator yields them under RCU. The counters are read under the same u64_stats seqlock the io.stat file uses, so the kfuncs add no fast-path cost: accounting stays in the per-cpu blkg iostat and is only folded on flush. bpf_blkcg_flush_stats() branches the way blkcg_print_stat() does. A non-root cgroup is flushed through rstat. The root cgroup is not accounted through rstat at all - blkcg_rstat_flush() returns early for it and __blkcg_rstat_flush() stops propagating one level short - so its per-device aggregates are refilled from the disks' own statistics instead, by blkcg_fill_root_iostats(), which is no longer static for that reason. Without this a program reading the root cgroup would see zeroes. Those numbers cover every cgroup's I/O, exactly as the root io.stat file reports them. Two details are worth calling out: bpf_iter_blkg_next() forgets the list head once the walk ends, not just the position. process_iter_next_call() in the verifier requires an iterator to keep returning NULL once it has returned it, and stops checking the loop for termination at that point; restarting the walk would let such a loop spin forever. The counter readers give up instead of retrying when called from NMI on 32-bit. There the u64_stats read is a real seqcount loop, every writer of blkg->iostat keeps interrupts off, and a perf event program can call these kfuncs from NMI, where the loop would never end. On 64-bit the loop compiles away. Signed-off-by: Ziyang Men --- MAINTAINERS | 1 + block/Makefile | 3 + block/blk-cgroup.c | 2 +- block/blk-cgroup.h | 1 + block/bpf_blkcg.c | 315 +++++++++++++++++++++++++++++++++++++++++++++ 5 files changed, 321 insertions(+), 1 deletion(-) create mode 100644 block/bpf_blkcg.c diff --git a/MAINTAINERS b/MAINTAINERS index 2f9472c1a090..87c56e955577 100644 --- a/MAINTAINERS +++ b/MAINTAINERS @@ -6617,6 +6617,7 @@ F: block/blk-cgroup.c F: block/blk-iocost.c F: block/blk-iolatency.c F: block/blk-throttle.c +F: block/bpf_blkcg.c F: include/linux/blk-cgroup.h =20 CONTROL GROUP - CPUSET diff --git a/block/Makefile b/block/Makefile index e7bd320e3d69..572e49988c8e 100644 --- a/block/Makefile +++ b/block/Makefile @@ -17,6 +17,9 @@ obj-$(CONFIG_BLK_ERROR_INJECTION) +=3D error-injection.o obj-$(CONFIG_BLK_DEV_BSG_COMMON) +=3D bsg.o obj-$(CONFIG_BLK_DEV_BSGLIB) +=3D bsg-lib.o obj-$(CONFIG_BLK_CGROUP) +=3D blk-cgroup.o +ifdef CONFIG_BPF_SYSCALL +obj-$(CONFIG_BLK_CGROUP) +=3D bpf_blkcg.o +endif obj-$(CONFIG_BLK_CGROUP_RWSTAT) +=3D blk-cgroup-rwstat.o obj-$(CONFIG_BLK_CGROUP_FC_APPID) +=3D blk-cgroup-fc-appid.o obj-$(CONFIG_BLK_DEV_THROTTLING) +=3D blk-throttle.o diff --git a/block/blk-cgroup.c b/block/blk-cgroup.c index d9676126c5b5..8d538ad4e861 100644 --- a/block/blk-cgroup.c +++ b/block/blk-cgroup.c @@ -1086,7 +1086,7 @@ static void blkcg_rstat_flush(struct cgroup_subsys_st= ate *css, int cpu) * flushing the root cgroup's stats by explicitly filling in the iostat * with disk level statistics. */ -static void blkcg_fill_root_iostats(void) +void blkcg_fill_root_iostats(void) { struct class_dev_iter iter; struct device *dev; diff --git a/block/blk-cgroup.h b/block/blk-cgroup.h index 615390f751aa..8c9c2a1adfaa 100644 --- a/block/blk-cgroup.h +++ b/block/blk-cgroup.h @@ -205,6 +205,7 @@ void blkcg_deactivate_policy(struct gendisk *disk, const struct blkcg_policy *pol); =20 const char *blkg_dev_name(struct blkcg_gq *blkg); +void blkcg_fill_root_iostats(void); void blkcg_print_blkgs(struct seq_file *sf, struct blkcg *blkcg, u64 (*prfill)(struct seq_file *, struct blkg_policy_data *, int), diff --git a/block/bpf_blkcg.c b/block/bpf_blkcg.c new file mode 100644 index 000000000000..48a86f07e198 --- /dev/null +++ b/block/bpf_blkcg.c @@ -0,0 +1,315 @@ +// SPDX-License-Identifier: GPL-2.0-or-later +/* + * Block I/O Controller-related BPF kfuncs and auxiliary code. + * + * These let a BPF program read a cgroup's io.stat counters. A program tur= ns a + * cgroup's css into a struct blkcg with bpf_get_blkcg(), flushes the stat= s with + * bpf_blkcg_flush_stats(), then walks the cgroup's per-device blkgs with = the + * bpf_iter_blkg open-coded iterator, reading each device's counters with + * bpf_blkg_iostat_bytes()/bpf_blkg_iostat_ios(). It mirrors the memory + * controller kfuncs in mm/bpf_memcontrol.c, but adds a per-device dimensi= on: + * unlike memcg, blkcg keeps one blkg (and one io.stat line) per block dev= ice. + * + * This file lives in block/ because the blkcg/blkg struct layouts are pri= vate + * to block/blk-cgroup.h. + */ + +#include "blk-cgroup.h" + +#include +#include +#include +#include + +__bpf_kfunc_start_defs(); + +/** + * bpf_get_root_blkcg - Returns a pointer to the root block cgroup + * + * The function has KF_ACQUIRE semantics, even though the root block cgrou= p is + * never destroyed and doesn't require reference counting. It's safe to pa= ss it + * to bpf_put_blkcg(). + * + * Note that the root cgroup is special: its counters are the disks' own + * statistics, so they cover every cgroup's I/O rather than only the root'= s. + * This matches what the root io.stat file prints. + * + * Return: A pointer to the root block cgroup. + */ +__bpf_kfunc struct blkcg *bpf_get_root_blkcg(void) +{ + /* css_get() is not needed */ + return &blkcg_root; +} + +/** + * bpf_get_blkcg - Get a reference to a block cgroup + * @css: pointer to the css structure + * + * It's fine to pass a css which belongs to any cgroup controller, + * e.g. unified hierarchy's main css. + * + * Implements KF_ACQUIRE semantics. + * + * Return: A pointer to a blkcg structure after bumping the corresponding = css's + * reference counter, or NULL if the io controller is not enabled on the c= group. + */ +__bpf_kfunc struct blkcg *bpf_get_blkcg(struct cgroup_subsys_state *css) +{ + struct blkcg *blkcg =3D NULL; + + if (css->ss =3D=3D &io_cgrp_subsys) + return css_tryget(css) ? css_to_blkcg(css) : NULL; + + /* + * Some other controller's css, or the cgroup's own one. Look up the io + * controller's css; rcu keeps it alive between the load and the tryget. + * Acquire and release rcu on one straight path, so that block/'s lock + * context analysis can follow it. + */ + rcu_read_lock(); + css =3D rcu_dereference_raw(css->cgroup->subsys[io_cgrp_id]); + if (css && css_tryget(css)) + blkcg =3D css_to_blkcg(css); + rcu_read_unlock(); + + return blkcg; +} + +/** + * bpf_put_blkcg - Put a reference to a block cgroup + * @blkcg: block cgroup to release + * + * Releases a previously acquired blkcg reference. + * Implements KF_RELEASE semantics. + */ +__bpf_kfunc void bpf_put_blkcg(struct blkcg *blkcg) +{ + css_put(&blkcg->css); +} + +/** + * bpf_blkcg_flush_stats - Flush a block cgroup's io statistics + * @blkcg: block cgroup + * + * Call this before reading counters for up-to-date values. Sleepable. + * + * It does what reading the io.stat file does, which differs by cgroup. Fo= r a + * non-root cgroup it folds the per-cpu deltas into the per-device aggrega= tes + * and up the cgroup tree. The root cgroup is not accounted through rstat = at + * all, so for it the per-device aggregates are refilled from the disks' + * own statistics, which count every cgroup's I/O. + * + * The root branch is not self-limiting the way the rstat one is: it rerea= ds + * every disk on every call, while a second rstat flush finds nothing left= to + * fold. The numbers it produces are the same for every cgroup, so read th= em + * once rather than once per cgroup of a walk. + */ +__bpf_kfunc void bpf_blkcg_flush_stats(struct blkcg *blkcg) +{ + if (!blkcg->css.parent) + blkcg_fill_root_iostats(); + else + css_rstat_flush(&blkcg->css); +} + +struct bpf_iter_blkg { + __u64 __opaque[2]; +} __attribute__((aligned(8))); + +struct bpf_iter_blkg_kern { + struct blkcg *blkcg; + struct blkcg_gq *pos; +} __attribute__((aligned(8))); + +/** + * bpf_iter_blkg_new - Start iterating a block cgroup's per-device blkgs + * @it: iterator to initialize + * @blkcg: block cgroup whose devices to walk + * + * Each yielded blkg holds one block device's counters, the same ones behi= nd a + * per-device line of the io.stat file. Offline blkgs are skipped, as the = file + * skips them. One case differs: the file prints no line for a blkg whose = disk + * is gone, while the walk still yields it, and bpf_blkg_dev() returns 0 f= or + * it. Must be used inside an RCU read section. + * + * Return: 0 on success. + */ +__bpf_kfunc int bpf_iter_blkg_new(struct bpf_iter_blkg *it, struct blkcg *= blkcg) +{ + struct bpf_iter_blkg_kern *kit =3D (void *)it; + + BUILD_BUG_ON(sizeof(struct bpf_iter_blkg_kern) > sizeof(struct bpf_iter_b= lkg)); + BUILD_BUG_ON(__alignof__(struct bpf_iter_blkg_kern) !=3D + __alignof__(struct bpf_iter_blkg)); + + kit->blkcg =3D blkcg; + kit->pos =3D NULL; + return 0; +} + +/** + * bpf_iter_blkg_next - Return the next online blkg of the iterated block = cgroup + * @it: iterator + * + * Return: the next online blkg, or NULL when the walk is done. + */ +__bpf_kfunc struct blkcg_gq *bpf_iter_blkg_next(struct bpf_iter_blkg *it) +{ + struct bpf_iter_blkg_kern *kit =3D (void *)it; + struct blkcg_gq *blkg =3D kit->pos; + struct hlist_node *node; + + /* Cleared once the walk is done, see below. */ + if (!kit->blkcg) + return NULL; + + if (!blkg) + node =3D rcu_dereference(hlist_first_rcu(&kit->blkcg->blkg_list)); + else + node =3D rcu_dereference(hlist_next_rcu(&blkg->blkcg_node)); + + /* Skip offline blkgs, matching the io.stat file. */ + while (node) { + blkg =3D hlist_entry(node, struct blkcg_gq, blkcg_node); + if (blkg->online) { + kit->pos =3D blkg; + return blkg; + } + node =3D rcu_dereference(hlist_next_rcu(&blkg->blkcg_node)); + } + + /* + * Forget the list head as well. The verifier assumes that an iterator + * which returned NULL keeps returning NULL, and stops checking the + * loop for termination once it has; starting the walk over would let + * such a loop spin forever. + */ + kit->pos =3D NULL; + kit->blkcg =3D NULL; + return NULL; +} + +/** + * bpf_iter_blkg_destroy - Tear down a blkg iterator + * @it: iterator + */ +__bpf_kfunc void bpf_iter_blkg_destroy(struct bpf_iter_blkg *it) +{ +} + +/* + * Read one counter out of @blkg's flushed io.stat aggregate. @counters is= one + * of the two arrays in blkg->iostat.cur; both are guarded by that struct's + * seqlock, the one the io.stat file uses. Returns (u64)-1 if the counter + * cannot be read. + */ +static u64 blkg_iostat_read(struct blkcg_gq *blkg, const u64 *counters, + enum blkg_iostat_type rw) +{ + struct blkg_iostat_set *bis =3D &blkg->iostat; + unsigned int seq; + u64 val; + + if ((unsigned int)rw >=3D BLKG_IOSTAT_NR) + return (u64)-1; + + /* + * On 32-bit the loop below really is a seqcount retry loop. Every + * writer of blkg->iostat keeps interrupts off, so only an NMI can land + * inside an update, and then the loop would never end. These kfuncs + * are reachable from a perf event program, which does run in NMI, so + * give up rather than spin. On 64-bit the loop compiles away. + */ + if (BITS_PER_LONG =3D=3D 32 && in_nmi()) + return (u64)-1; + + do { + seq =3D u64_stats_fetch_begin(&bis->sync); + val =3D counters[rw]; + } while (u64_stats_fetch_retry(&bis->sync, seq)); + + return val; +} + +/** + * bpf_blkg_iostat_bytes - Read a device's io.stat byte counter + * @blkg: block group (one device of a block cgroup) + * @rw: which counter (BLKG_IOSTAT_READ / _WRITE / _DISCARD) + * + * Reads the flushed aggregate, so call bpf_blkcg_flush_stats() first for + * up-to-date values. The read uses the u64_stats seqlock, like the io.stat + * file. + * + * Return: the number of bytes, or (u64)-1 if @rw is out of range or the + * counter cannot be read. + */ +__bpf_kfunc u64 bpf_blkg_iostat_bytes(struct blkcg_gq *blkg, + enum blkg_iostat_type rw) +{ + return blkg_iostat_read(blkg, blkg->iostat.cur.bytes, rw); +} + +/** + * bpf_blkg_iostat_ios - Read a device's io.stat I/O count + * @blkg: block group (one device of a block cgroup) + * @rw: which counter (BLKG_IOSTAT_READ / _WRITE / _DISCARD) + * + * Return: the number of I/Os, or (u64)-1 if @rw is out of range or the + * counter cannot be read. + */ +__bpf_kfunc u64 bpf_blkg_iostat_ios(struct blkcg_gq *blkg, + enum blkg_iostat_type rw) +{ + return blkg_iostat_read(blkg, blkg->iostat.cur.ios, rw); +} + +/** + * bpf_blkg_dev - Return a blkg's device id + * @blkg: block group + * + * Return: the device's dev_t (use MAJOR()/MINOR() to split), or 0 if the = blkg + * has no disk. + */ +__bpf_kfunc u64 bpf_blkg_dev(struct blkcg_gq *blkg) +{ + if (!blkg->q || !blkg->q->disk) + return 0; + + return blkg->q->disk->part0->bd_dev; +} + +__bpf_kfunc_end_defs(); + +BTF_KFUNCS_START(bpf_blkcg_kfuncs) +BTF_ID_FLAGS(func, bpf_get_root_blkcg, KF_ACQUIRE | KF_RET_NULL) +BTF_ID_FLAGS(func, bpf_get_blkcg, KF_ACQUIRE | KF_RET_NULL | KF_RCU) +BTF_ID_FLAGS(func, bpf_put_blkcg, KF_RELEASE) +BTF_ID_FLAGS(func, bpf_blkcg_flush_stats, KF_SLEEPABLE) + +BTF_ID_FLAGS(func, bpf_iter_blkg_new, KF_ITER_NEW | KF_RCU_PROTECTED) +BTF_ID_FLAGS(func, bpf_iter_blkg_next, KF_ITER_NEXT | KF_RET_NULL) +BTF_ID_FLAGS(func, bpf_iter_blkg_destroy, KF_ITER_DESTROY) + +BTF_ID_FLAGS(func, bpf_blkg_iostat_bytes, KF_RCU) +BTF_ID_FLAGS(func, bpf_blkg_iostat_ios, KF_RCU) +BTF_ID_FLAGS(func, bpf_blkg_dev, KF_RCU) +BTF_KFUNCS_END(bpf_blkcg_kfuncs) + +static const struct btf_kfunc_id_set bpf_blkcg_kfunc_set =3D { + .owner =3D THIS_MODULE, + .set =3D &bpf_blkcg_kfuncs, +}; + +static int __init bpf_blkcg_init(void) +{ + int err; + + err =3D register_btf_kfunc_id_set(BPF_PROG_TYPE_UNSPEC, + &bpf_blkcg_kfunc_set); + if (err) + pr_warn("error while registering bpf blkcg kfuncs: %d\n", err); + + return err; +} +late_initcall(bpf_blkcg_init); --=20 2.53.0-Meta From nobody Tue Sep 29 13:19:45 2026 Received: from mail-pg1-f174.google.com (mail-pg1-f174.google.com [209.85.215.174]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 74E42419305 for ; Fri, 7 Aug 2026 19:37:38 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.174 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786131462; cv=none; b=Hgl/0+BF800gKiXkSqfccuIn19E4fL79aZYewL5iG60tF6cqn45PiCAfmIGsPP2EOoniturdG9lSgTtZD2EnN3EY25BWVNM1Dsddi/R7NER5+teHLYynLYZFUxYEo1nd3PqYIZAnhYRLcVYRI9l+hGvIVDhAuQv+3XE/fr+MMuU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786131462; c=relaxed/simple; bh=8L2jNqqneTrWjIhxBAyTlYEDrqUxG5JMpaCEtZLJpN8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=GIOP0VwlEIwOK1NSF0/xRj2jXiiQMBArKPnJwtk41X+8wOucphbSTiSB/xl6mGmSk4/qHmA+70AgC8DpWO3plmbU0QSge7dco3At054sAl4kKjXuB2xcAV4R+js6GMLlUJ6hJkp+xNieV4TBmvWGKvohgdCd1Ta3b3S80jJY6BI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=Lyh7K97d; arc=none smtp.client-ip=209.85.215.174 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="Lyh7K97d" Received: by mail-pg1-f174.google.com with SMTP id 41be03b00d2f7-c9cf07d2df6so2943155a12.2 for ; Fri, 07 Aug 2026 12:37:38 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1786131458; x=1786736258; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=516fNSEVFZidFTbnujDBUpNRDc+xNXskAxbQE+m29cc=; b=Lyh7K97dEOnrAUVmkZ3bm2ItCwlZiM8HfVOkM92N9L+eJnTQ7DmiJzm2v05gFaRC4Y Rje5IB/6lNYNAtxqkvMdEZ0faE/cjUzfRTmgvqPzxb6Q/GI/5/qirKs5Bmxygd6e8ly+ Y6TzYCHsYO110blHkHcuKG5orOYxcgmVEHou7G1DRA294YJXJEKIZgl8d3IIw1BA89m2 vqIvX/VvOoL6KnxdFppfm/7xNSlTuBKsiBYmI+Pe3B0N8MYxJRzTYAlBtFa+xOonr8wQ /2N6vSgsAzBVL5gDDU/3B1PftkOaMfATvGNWa9jtoubi+oalaM9krrT21+zubbIqkWgN 4Mog== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786131458; x=1786736258; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=516fNSEVFZidFTbnujDBUpNRDc+xNXskAxbQE+m29cc=; b=Dp9obrSzdUtMI4C8isUllPjh17Zkv3XPAQnITu0PNx8gxEOc5c591r4ThZIx4XQG34 j/0eqGdfrmy8w14C7ji47h2dTl9RMhrwq+zzUpBjhOf61iDerxM+UeQAi/nJWK6mAMxW afo8TmHH/8u/f+HIaUgVKocXx85vzuo/niGeRNJCHhY3XoYfJjvBHD1KBALbO4GfF1/L 4VyFwGGequyvcRngE5vUqJox8CzzButCWa8V/G47ciiAZER5vIVKSascZUkVoIxDmguO IP8GIG8POuorOIFLoof9NKhnxHPJ0xb4mydUcwI5jmxiMJ8QXDxp3JvpH0Ryw7TBQzR+ STOQ== X-Forwarded-Encrypted: i=1; AHgh+RojfxM1kBq3MpMYNuMSuTQy6Bgx9hOEiuGTzWY7bS1L1/urKO7tL/DDqI5brtKxNyexHGjmN4m3AKHE2jA=@vger.kernel.org X-Gm-Message-State: AOJu0YwDaZZdJ0pkDVpAHkWe3eRJCeyeQgR3ARhCmdbW1VNeGG9M1zz5 3msABK26USLNEMGkstTIAAqJgqfq4mF2L8rN3DX8ofnQ8IHfAvzxB39O X-Gm-Gg: AR+sD1009d9YSUP1fsadj8N3+eRMx6TQpfc7RnQ6C330EGwaB1WJfpZ3ItfIUo5D3kw wYixCunCZJ7q/zxFiHmEGI6W/e/uRIlGAbevimNyk+qLpxTRIwqLZA0wvpS3DvQ4NQo4qAsaRI+ pTMHYYM9phQNlw3N/O3C/i/qSfIyKO9PLrylgPdg8RieP3828W5OX5/vQqiDvVgempr8/fMftK2 oPGIa1uKE2S+z8SZ8RvWhPpTdznLMPyEbhNaRKSsEz8lrQP9HFZY2J627P2w2fUUw0baRezQMwz APyh4VceA4/QVpK0/2VI4yuNlKOroWyKX8a50IoqeXJdBgiLk4ZKwMZe5xq+EcIq2uT2mNpwbio XgxViVzDhN5acBP4xbeoQ6d/N5h0BxDteztjEUB/JS34JG3hCGz9ilDDjd3eRtqw47u54ClM6sy bTzhah5Emfc42h+NyNOxaDorEa94RUKPsuIufE79UpOs5Ny3H9aAXMftH9AgExSo2qDw== X-Received: by 2002:a05:6a20:9183:b0:3bf:63af:855 with SMTP id adf61e73a8af0-3cb85dee60cmr28788437637.1.1786131458255; Fri, 07 Aug 2026 12:37:38 -0700 (PDT) Received: from localhost ([2a03:2880:9ff:72::]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-315bebde308sm10594920eec.20.2026.08.07.12.37.37 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 07 Aug 2026 12:37:37 -0700 (PDT) From: Ziyang Men To: Jens Axboe , Tejun Heo , Josef Bacik , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , Eduard Zingerman , Kumar Kartikeya Dwivedi Cc: Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Shuah Khan , Johannes Weiner , =?UTF-8?q?Michal=20Koutn=C3=BD?= , Roman Gushchin , Shakeel Butt , JP Kobryn , Mykola Lysenko , kernel-team@meta.com, linux-block@vger.kernel.org, bpf@vger.kernel.org, cgroups@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, Ziyang Men Subject: [PATCH 2/2] selftests/bpf: add test for blkcg io.stat BPF kfuncs Date: Fri, 7 Aug 2026 12:37:32 -0700 Message-ID: <20260807193732.4073299-3-ziyang.meme@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260807193732.4073299-1-ziyang.meme@gmail.com> References: <20260807193732.4073299-1-ziyang.meme@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add cgroup_iter_io, a test_progs test for the block I/O controller BPF kfuncs. A SEC("iter.s/cgroup") program acquires the cgroup's blkcg, flushes stats, iterates its blkgs and reads the io.stat counters for a target device. The userspace side attaches a loop device, generates O_DIRECT read and write I/O charged to a test cgroup, and then: - checks the write and read byte/io counters are nonzero, - checks the reported device id, - compares every kfunc-read value against the cgroup's io.stat file for the same device and requires an exact match, - reads the same device through bpf_get_root_blkcg() and checks the root counters are at or above the test cgroup's. The measured device is pinned to the loop device, which has no asynchronous writeback, so the kfunc snapshot and the io.stat file snapshot are identical rather than merely close. The root cgroup's numbers for a device come from the disk itself and so cover every cgroup's I/O to it, which is why the root check is "at or above" rather than an exact match. CONFIG_BLK_CGROUP is added to the test config; CONFIG_BLK_DEV_LOOP is already present. Signed-off-by: Ziyang Men --- tools/testing/selftests/bpf/cgroup_iter_io.h | 17 + tools/testing/selftests/bpf/config | 1 + .../selftests/bpf/prog_tests/cgroup_iter_io.c | 310 ++++++++++++++++++ .../selftests/bpf/progs/cgroup_iter_io.c | 107 ++++++ 4 files changed, 435 insertions(+) create mode 100644 tools/testing/selftests/bpf/cgroup_iter_io.h create mode 100644 tools/testing/selftests/bpf/prog_tests/cgroup_iter_io.c create mode 100644 tools/testing/selftests/bpf/progs/cgroup_iter_io.c diff --git a/tools/testing/selftests/bpf/cgroup_iter_io.h b/tools/testing/s= elftests/bpf/cgroup_iter_io.h new file mode 100644 index 000000000000..f4bbaaccdf71 --- /dev/null +++ b/tools/testing/selftests/bpf/cgroup_iter_io.h @@ -0,0 +1,17 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* Copyright (c) 2025 Meta Platforms, Inc. and affiliates. */ +#ifndef __CGROUP_ITER_IO_H +#define __CGROUP_ITER_IO_H + +struct io_query { + /* one device's io.stat counters */ + __u64 rbytes; + __u64 wbytes; + __u64 rios; + __u64 wios; + __u64 dbytes; + __u64 dios; + __u64 dev; /* dev_t of the device the counters belong to */ +}; + +#endif /* __CGROUP_ITER_IO_H */ diff --git a/tools/testing/selftests/bpf/config b/tools/testing/selftests/b= pf/config index ea7044f30adc..270e6bf9194d 100644 --- a/tools/testing/selftests/bpf/config +++ b/tools/testing/selftests/bpf/config @@ -1,3 +1,4 @@ +CONFIG_BLK_CGROUP=3Dy CONFIG_BLK_DEV_LOOP=3Dy CONFIG_BOOTPARAM_HARDLOCKUP_PANIC=3Dy CONFIG_BOOTPARAM_SOFTLOCKUP_PANIC=3D1 diff --git a/tools/testing/selftests/bpf/prog_tests/cgroup_iter_io.c b/tool= s/testing/selftests/bpf/prog_tests/cgroup_iter_io.c new file mode 100644 index 000000000000..32cda8243318 --- /dev/null +++ b/tools/testing/selftests/bpf/prog_tests/cgroup_iter_io.c @@ -0,0 +1,310 @@ +// SPDX-License-Identifier: GPL-2.0 +/* Copyright (c) 2025 Meta Platforms, Inc. and affiliates. */ +#define _GNU_SOURCE +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include "cgroup_helpers.h" +#include "cgroup_iter_io.h" +#include "cgroup_iter_io.skel.h" + +#define IO_SIZE (4 * 1024 * 1024) + +static int read_stats(struct bpf_link *link) +{ + int fd, ret =3D 0; + ssize_t bytes; + + fd =3D bpf_iter_create(bpf_link__fd(link)); + if (!ASSERT_OK_FD(fd, "bpf_iter_create")) + return 1; + + /* Results land in skel->data_query; the read itself returns no data. */ + bytes =3D read(fd, NULL, 0); + if (!ASSERT_EQ(bytes, 0, "read fd")) + ret =3D 1; + + close(fd); + return ret; +} + +/* + * Attach a loop device to an anonymous temp file so we have a real block + * device to generate cgroup-charged I/O against. Returns 0 on success, or= -1 + * if loop devices are unavailable (non-root / no CONFIG_BLK_DEV_LOOP) so = the + * caller can skip. + */ +static int loop_setup(char *loop_path, size_t sz, int *ctl_fd, int *loop_f= d, + int *back_fd) +{ + char back_path[] =3D "/tmp/cgroup_iter_io.XXXXXX"; + int nr; + + *ctl_fd =3D *loop_fd =3D *back_fd =3D -1; + + *ctl_fd =3D open("/dev/loop-control", O_RDWR | O_CLOEXEC); + if (*ctl_fd < 0) + return -1; + + nr =3D ioctl(*ctl_fd, LOOP_CTL_GET_FREE); + if (nr < 0) + goto err; + snprintf(loop_path, sz, "/dev/loop%d", nr); + + *back_fd =3D mkstemp(back_path); + if (*back_fd < 0) + goto err; + unlink(back_path); + if (ftruncate(*back_fd, (off_t)IO_SIZE * 4)) + goto err; + + *loop_fd =3D open(loop_path, O_RDWR | O_CLOEXEC); + if (*loop_fd < 0) + goto err; + if (ioctl(*loop_fd, LOOP_SET_FD, *back_fd)) + goto err; + + return 0; +err: + if (*loop_fd >=3D 0) + close(*loop_fd); + if (*back_fd >=3D 0) + close(*back_fd); + close(*ctl_fd); + *ctl_fd =3D *loop_fd =3D *back_fd =3D -1; + return -1; +} + +static void loop_teardown(const char *loop_path, int ctl_fd, int loop_fd, + int back_fd) +{ + int nr =3D -1; + + if (loop_fd >=3D 0) { + ioctl(loop_fd, LOOP_CLR_FD, 0); + close(loop_fd); + } + if (back_fd >=3D 0) + close(back_fd); + if (ctl_fd >=3D 0) { + if (sscanf(loop_path, "/dev/loop%d", &nr) =3D=3D 1 && nr >=3D 0) + ioctl(ctl_fd, LOOP_CTL_REMOVE, nr); + close(ctl_fd); + } +} + +/* O_DIRECT I/O to the loop device, charged to the current cgroup. */ +static int do_direct_io(const char *loop_path) +{ + void *buf; + int fd, ret =3D -1; + + fd =3D open(loop_path, O_RDWR | O_DIRECT | O_CLOEXEC); + if (fd < 0) + return -1; + if (posix_memalign(&buf, 4096, IO_SIZE)) + goto out_fd; + memset(buf, 0xab, IO_SIZE); + + if (pwrite(fd, buf, IO_SIZE, 0) !=3D IO_SIZE) + goto out_buf; + fsync(fd); + if (pread(fd, buf, IO_SIZE, 0) !=3D IO_SIZE) + goto out_buf; + ret =3D 0; +out_buf: + free(buf); +out_fd: + close(fd); + return ret; +} + +/* + * Parse the io.stat line for device @dev out of the cgroup's io.stat file= and + * fill @out. @dev is a kernel dev_t (as returned by bpf_blkg_dev), whose + * major:minor split matches how io.stat prints the device. Returns 0 if t= he + * device's line was found. + */ +static int parse_io_stat(int cgroup_fd, __u64 dev, struct io_query *out) +{ + unsigned int want_maj =3D dev >> 20, want_min =3D dev & ((1U << 20) - 1); + char buf[4096], *line, *saveptr; + int fd, n, ret =3D -1; + + fd =3D openat(cgroup_fd, "io.stat", O_RDONLY); + if (fd < 0) + return -1; + n =3D read(fd, buf, sizeof(buf) - 1); + close(fd); + if (n <=3D 0) + return -1; + buf[n] =3D '\0'; + + for (line =3D strtok_r(buf, "\n", &saveptr); line; + line =3D strtok_r(NULL, "\n", &saveptr)) { + unsigned long long rb =3D 0, wb =3D 0, ri =3D 0, wi =3D 0, db =3D 0, di = =3D 0; + unsigned int maj, min; + + /* + * The "maj:min" token is always present; the field block is + * optional (the kernel omits it for a device with no read/write + * I/O), so a match of >=3D 2 is enough and absent fields stay 0. + */ + if (sscanf(line, + "%u:%u rbytes=3D%llu wbytes=3D%llu rios=3D%llu wios=3D%llu dbytes=3D= %llu dios=3D%llu", + &maj, &min, &rb, &wb, &ri, &wi, &db, &di) < 2) + continue; + if (maj !=3D want_maj || min !=3D want_min) + continue; + + out->rbytes =3D rb; + out->wbytes =3D wb; + out->rios =3D ri; + out->wios =3D wi; + out->dbytes =3D db; + out->dios =3D di; + ret =3D 0; + break; + } + return ret; +} + +void test_cgroup_iter_io(void) +{ + char *cgroup_rel_path =3D "/cgroup_iter_io_test"; + int ctl_fd =3D -1, loop_fd =3D -1, back_fd =3D -1; + struct cgroup_iter_io *skel =3D NULL; + struct bpf_link *link =3D NULL; + char loop_path[64]; + struct io_query *q; + int cgroup_fd; + + cgroup_fd =3D cgroup_setup_and_join(cgroup_rel_path); + if (!ASSERT_OK_FD(cgroup_fd, "cgroup_setup_and_join")) + return; + + if (loop_setup(loop_path, sizeof(loop_path), &ctl_fd, &loop_fd, &back_fd)= ) { + test__skip(); /* needs root + CONFIG_BLK_DEV_LOOP */ + goto cleanup_cgroup_fd; + } + + skel =3D cgroup_iter_io__open_and_load(); + if (!ASSERT_OK_PTR(skel, "cgroup_iter_io__open_and_load")) + goto cleanup_loop; + + /* + * Pin the read to the loop device so the measured device is stable and + * quiesced. Convert the glibc-encoded st_rdev to the kernel dev_t + * encoding (major << 20 | minor) that bpf_blkg_dev returns. + */ + { + struct stat lst; + + if (!ASSERT_OK(fstat(loop_fd, &lst), "fstat loop")) + goto cleanup_skel; + skel->data_query->target_dev =3D + ((__u64)major(lst.st_rdev) << 20) | minor(lst.st_rdev); + } + + DECLARE_LIBBPF_OPTS(bpf_iter_attach_opts, opts); + union bpf_iter_link_info linfo =3D { + .cgroup.cgroup_fd =3D cgroup_fd, + .cgroup.order =3D BPF_CGROUP_ITER_SELF_ONLY, + }; + opts.link_info =3D &linfo; + opts.link_info_len =3D sizeof(linfo); + + link =3D bpf_program__attach_iter(skel->progs.cgroup_io_query, &opts); + if (!ASSERT_OK_PTR(link, "bpf_program__attach_iter")) + goto cleanup_skel; + + /* This process is in the test cgroup, so the loop I/O is charged here. */ + if (!ASSERT_OK(do_direct_io(loop_path), "do_direct_io")) + goto cleanup_link; + + if (!ASSERT_OK(read_stats(link), "read stats")) + goto cleanup_link; + + /* + * Weak check: we did I/O, so the numbers must be non-zero. Follows the + * pattern in cgroup_iter_memcg. + */ + q =3D &skel->data_query->io_query; + if (test__start_subtest("cgroup_iter_io__write")) { + ASSERT_GT(q->wbytes, 0, "wbytes"); + ASSERT_GT(q->wios, 0, "wios"); + } + if (test__start_subtest("cgroup_iter_io__read")) { + ASSERT_GT(q->rbytes, 0, "rbytes"); + ASSERT_GT(q->rios, 0, "rios"); + } + if (test__start_subtest("cgroup_iter_io__dev")) + ASSERT_GT(q->dev, 0, "dev"); + + /* + * Stronger check: the kfunc-read values must equal what the io.stat + * file reports for the same device. Refresh via the prog, then read + * the file with no I/O in between, so both flushed snapshots match + * exactly. + */ + if (test__start_subtest("cgroup_iter_io__match")) { + struct io_query filev =3D {}; + + if (ASSERT_OK(read_stats(link), "read stats") && + ASSERT_OK(parse_io_stat(cgroup_fd, q->dev, &filev), + "parse io.stat")) { + ASSERT_EQ(q->rbytes, filev.rbytes, "rbytes"); + ASSERT_EQ(q->wbytes, filev.wbytes, "wbytes"); + ASSERT_EQ(q->rios, filev.rios, "rios"); + ASSERT_EQ(q->wios, filev.wios, "wios"); + ASSERT_EQ(q->dbytes, filev.dbytes, "dbytes"); + ASSERT_EQ(q->dios, filev.dios, "dios"); + } + } + + /* + * Separate program for the root block cgroup. Its counters do not come + * from rstat, they are refilled from the disks themselves, so this + * covers the other half of bpf_blkcg_flush_stats(). They cover every + * cgroup's I/O to the loop device, and only this test touches it, so + * they must be at or above what the test cgroup was charged. + */ + if (test__start_subtest("cgroup_iter_io__root")) { + struct bpf_link *root_link; + struct io_query *r; + + skel->data_query->got_root_blkcg =3D 0; + root_link =3D bpf_program__attach_iter(skel->progs.cgroup_root_blkcg_que= ry, + &opts); + if (ASSERT_OK_PTR(root_link, "attach root iter")) { + if (ASSERT_OK(read_stats(root_link), "read root stats")) { + r =3D &skel->data_query->root_query; + ASSERT_EQ(skel->data_query->got_root_blkcg, 1, + "got_root_blkcg"); + ASSERT_EQ(r->dev, q->dev, "root dev"); + ASSERT_GE(r->wbytes, q->wbytes, "root wbytes"); + ASSERT_GE(r->wios, q->wios, "root wios"); + ASSERT_GE(r->rbytes, q->rbytes, "root rbytes"); + ASSERT_GE(r->rios, q->rios, "root rios"); + } + bpf_link__destroy(root_link); + } + } + +cleanup_link: + bpf_link__destroy(link); +cleanup_skel: + cgroup_iter_io__destroy(skel); +cleanup_loop: + loop_teardown(loop_path, ctl_fd, loop_fd, back_fd); +cleanup_cgroup_fd: + close(cgroup_fd); + cleanup_cgroup_environment(); +} diff --git a/tools/testing/selftests/bpf/progs/cgroup_iter_io.c b/tools/tes= ting/selftests/bpf/progs/cgroup_iter_io.c new file mode 100644 index 000000000000..b839def94508 --- /dev/null +++ b/tools/testing/selftests/bpf/progs/cgroup_iter_io.c @@ -0,0 +1,107 @@ +// SPDX-License-Identifier: GPL-2.0 +/* Copyright (c) 2025 Meta Platforms, Inc. and affiliates. */ +#include +#include +#include +#include "bpf_experimental.h" +#include "cgroup_iter_io.h" + +char _license[] SEC("license") =3D "GPL"; + +/* The counters of the device named by target_dev are stored here. */ +struct io_query io_query SEC(".data.query"); + +/* The same device's counters read through the root block cgroup. */ +struct io_query root_query SEC(".data.query"); + +/* Set to 1 by cgroup_root_blkcg_query when bpf_get_root_blkcg() succeeds.= */ +__u64 got_root_blkcg SEC(".data.query"); + +/* Device to read, set by userspace (kernel dev_t). Pinning the device kee= ps + * the read deterministic and lets the value be compared to io.stat exactl= y. + */ +__u64 target_dev SEC(".data.query"); + +/* + * Flush @blkcg and copy the target device's counters into @out. Reading o= nly + * the one pinned device keeps the result deterministic: that device is + * quiesced, so its counters match io.stat exactly, while picking "any dev= ice + * with I/O" would race with backing-store writeback. + */ +static __always_inline void read_target_dev(struct blkcg *blkcg, + struct io_query *out) +{ + struct blkcg_gq *pos; + + /* io.stat needs a flush before it can be read (sleepable). */ + bpf_blkcg_flush_stats(blkcg); + + /* The per-device blkg walk needs an RCU section. */ + bpf_rcu_read_lock(); + bpf_for_each(blkg, pos, blkcg) { + if (bpf_blkg_dev(pos) !=3D target_dev) + continue; + + out->dev =3D bpf_blkg_dev(pos); + out->rbytes =3D bpf_blkg_iostat_bytes(pos, BLKG_IOSTAT_READ); + out->wbytes =3D bpf_blkg_iostat_bytes(pos, BLKG_IOSTAT_WRITE); + out->rios =3D bpf_blkg_iostat_ios(pos, BLKG_IOSTAT_READ); + out->wios =3D bpf_blkg_iostat_ios(pos, BLKG_IOSTAT_WRITE); + out->dbytes =3D bpf_blkg_iostat_bytes(pos, BLKG_IOSTAT_DISCARD); + out->dios =3D bpf_blkg_iostat_ios(pos, BLKG_IOSTAT_DISCARD); + break; + } + bpf_rcu_read_unlock(); +} + +SEC("iter.s/cgroup") +int cgroup_io_query(struct bpf_iter__cgroup *ctx) +{ + struct cgroup *cgrp =3D ctx->cgroup; + struct blkcg *blkcg; + + /* The last iteration has a NULL cgroup, skip it. */ + if (!cgrp) + return 1; + + /* Start fresh so a device that is not found stays all-zero. */ + __builtin_memset(&io_query, 0, sizeof(io_query)); + + blkcg =3D bpf_get_blkcg(&cgrp->self); + if (!blkcg) + return 0; + + read_target_dev(blkcg, &io_query); + + bpf_put_blkcg(blkcg); + return 0; +} + +SEC("iter.s/cgroup") +int cgroup_root_blkcg_query(struct bpf_iter__cgroup *ctx) +{ + struct cgroup *cgrp =3D ctx->cgroup; + struct blkcg *blkcg; + + /* The last iteration has a NULL cgroup, skip it. */ + if (!cgrp) + return 1; + + __builtin_memset(&root_query, 0, sizeof(root_query)); + + blkcg =3D bpf_get_root_blkcg(); + if (!blkcg) + return 0; + + /* + * The root cgroup takes its numbers from the disks themselves rather + * than from rstat, so this also covers the root side of + * bpf_blkcg_flush_stats(). The counters cover every cgroup's I/O, so + * they can only be at or above what this test's own cgroup did. + */ + read_target_dev(blkcg, &root_query); + + got_root_blkcg =3D 1; + bpf_put_blkcg(blkcg); + return 0; +} --=20 2.53.0-Meta