From nobody Sat Jul 25 00:11:12 2026 Received: from mail-pj2-f2.google.com (mail-pj2-f2.google.com [74.125.227.130]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1CDE64DBD8B for ; Tue, 21 Jul 2026 17:48:40 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.227.130 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784656122; cv=none; b=rXMnsFBGlzH09/j7OTk+xjGcFlABHrl1XodcXK1y0AQcO24NQhvlTUmpQqueqHmWPD+pgeh4VGS/HnuJ9cILVHLHidPiqjMkPRwfIo65gWNPnELgWCN3rup3EqJkVS0KJzikuyLv0pduj4Dq/3iT0sQSwSDlCfep5QGxFliTtME= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784656122; c=relaxed/simple; bh=/cw6DBp4kTloBb6x9j/7235qDRGnsiYevOiWZeXJ+CM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=pd5afD2/x9hj2U4OG6R/maLpK2ztuHfptl0SYXuqX/76v2kOEnDvUoKuDOCSmge2xOuOiPNqQw63CoaDojjiv4Nvwg3pJK+qsvYmpmGbMUV7kdR7JRnpjF/K7cXzv5jdjm3Gfzez7bVOnAuHMhA0J7Eqj9ZXdDJ/0Tu+9D6oqHU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=EcMCnNgR; arc=none smtp.client-ip=74.125.227.130 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="EcMCnNgR" Received: by mail-pj2-f2.google.com with SMTP id d9443c01a7336-2ccc2e84048so67477825ad.1 for ; Tue, 21 Jul 2026 10:48:40 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784656119; x=1785260919; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=wb99tNhyMGyQHq+Tc4HfWLLbXPl8M5pWqDy6JsoUgIg=; b=EcMCnNgRmqeNHKtf1WnrJqG7hl9bGsk898ScN1WqjRPUXhsa5N+6Pi4y+0qpZNarQt 7XVj9d86ESVss6dizDU2ZEXxwzhCATHyBvmxrQle0lW95A4SMOFsD2NQmebkYQe3YJ2o 0N9yechaI+rBm1UygX//pb7deUtfnZreCVud6CvWzipgKI9alWtRa8UDZSuID1qxkGcK i/RHp2IxbXcS+1Qe+ZUYQ++V5lYMEzOmtDC15Nk8Jybe7DPLEYrJC8RwSFCiyIp959Rg sp/DotI26uB9uTGaKvVZNZgqctIOvBh70/cTgmnhb6LQOnB/i0QrWXd5p0otZANM2POK N/lw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784656119; x=1785260919; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=wb99tNhyMGyQHq+Tc4HfWLLbXPl8M5pWqDy6JsoUgIg=; b=oSAA81lkhCzrFlHJhNuSoeNd8yutIT0Cp0RFtHI3QCrr2aPmlTEComovCblHfGxj0M aI1fNyTRf58NZIL9KpgBKD+gxYwmmfVJiCvDzOE03iwdQdnibgZbZWNLobBkDQPx4169 Ev3nNY/FIjZBorvVEsRh5yCMFnW7QGrz7u9zJsZDUzcKpnbIiAtWPCPpuA5Czg3BMSps Zb60Bf5HPHALN9gfEi84J2zdujfzgNQ6f0rImwPfeX6R482TnNfqcZZzBFEouDmodnIK UC/4m7wBE19RvLFfywjbHiNJ93Pg/wPVs0Q3xG1mK2FnNlK1mV22yjitXGBlXOSHS2dV Gobg== X-Forwarded-Encrypted: i=1; AHgh+RqAhbnSdfgmcs4y37/cLyZConeTEEYbj+vi/SZAjxPMhRP8/6aa29BRGAYJICDwGl3P00oKGR5BNHcbL9g=@vger.kernel.org X-Gm-Message-State: AOJu0YzuGXOdV9VmmASDJ4CoNEEA0Azc0hQBoTb4zEdjEPZ0hvKqx1nc Fh4AgNnUorWnRvIDtvOdcoxoQBs2UxV3PrjE0Uue6rJ7MIMfaBeEewLN X-Gm-Gg: AR+sD11roxZ/Ter2Kw36+kvADkDcvBTDHCfh//Ei2NIpgVkRrlPjoLnuJIssqTbXtZe N9vw+q5u53ptBMayHQb/J9Qlwe+dBKnZR4rYajXeV34WcIHlUcvOTlz8RS/XiSqG2YU6eshsnoP F3LRJUaoNSIHYHFX8UY2KBwTDHc62wHBkYuzn3F0Ie4Z6kfncrtsZyceqlZSvv5DsVvSUBus4YJ xuUH/FpfTb/DHP2hZle8sXnhOtt1af/MmRkVhS3HtlX9xh09btEVa1gsEYmJzzd5T/TWikSJ8X1 PC243hFbWLmp4gyHC2AFQt0cezx1fZWwoJe3LTvUDUAMZgcGd9Fks4RNeFYEI4f8iVMo7JFuQJY KSvVoS9MW1eLfYtqdveK/LjZtl22dcDYXqrHlyvXa4FEp39dRawnG8exJ6kBVtsrz368f7/Bf X-Received: by 2002:a17:903:1a43:b0:2c9:de53:f84f with SMTP id d9443c01a7336-2cf3487e1e8mr224944495ad.19.1784656119107; Tue, 21 Jul 2026 10:48:39 -0700 (PDT) Received: from localhost ([2a03:2880:9ff:65::]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-3147e070538sm814563eec.20.2026.07.21.10.48.37 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 21 Jul 2026 10:48:37 -0700 (PDT) From: Ziyang Men To: Shuah Khan , Tejun Heo , Johannes Weiner , =?UTF-8?q?Michal=20Koutn=C3=BD?= , Jiri Kosina , Benjamin Tissoires , David Vernet , Eduard Zingerman Cc: Andrea Righi , Changwoo Min , Michal Hocko , Roman Gushchin , Shakeel Butt , Muchun Song , Andrew Morton , JP Kobryn , Mykola Lysenko , Nathan Chancellor , linux-kselftest@vger.kernel.org, cgroups@vger.kernel.org, linux-input@vger.kernel.org, sched-ext@lists.linux.dev, linux-mm@kvack.org, kernel-team@meta.com, bpf@vger.kernel.org, llvm@lists.linux.dev, linux-kernel@vger.kernel.org, Ziyang Men Subject: [PATCH v2 1/4] selftests: add shared lib.bpf.mk to build BPF progs and skeletons Date: Tue, 21 Jul 2026 10:48:30 -0700 Message-ID: <20260721174833.1232771-2-ziyang.meme@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260721174833.1232771-1-ziyang.meme@gmail.com> References: <20260721174833.1232771-1-ziyang.meme@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" The libbpf + bpftool + vmlinux.h + BPF-object + skeleton build tool-chain is currently duplicated across tools/testing/selftests/{bpf,sched_ext, hid}/, each carrying ~100-140 lines of near-identical Makefile. As more subsystems grow BPF-based selftests, the duplication scales poorly. Add tools/testing/selftests/lib.bpf.mk, a single includable fragment that provides the whole chain end-to-end. It builds the in-tree libbpf.a and a host bpftool, generates vmlinux.h from the kernel's BTF, compiles *.bpf.c into BPF objects (clang --target=3Dbpf) and generates their skeletons. It also provides what the user-space test binary needs to build: the header search paths (so #includes resolve), a bpf_link macro that wraps the link command, and BPF_LDLIBS, so the test can be statically linked against libbpf.a. To use: set BPF_SRCS and OVERRIDE_TARGETS :=3D 1 before including ../lib.mk (so lib.mk's default link rule is suppressed), then include ../lib.bpf.mk and list $(BPF_SKELS) as prerequisites of the test binary, e.g.,: BPF_SRCS :=3D progs/foo.bpf.c OVERRIDE_TARGETS :=3D 1 include ../lib.mk include ../lib.bpf.mk $(OUTPUT)/foo_test: foo_test.c $(BPF_SKELS) $(call bpf_link,$@,$<) This saves much work for configuring selftests in other folder, such the cg= roup. net/bpf.mk (which only builds *.bpf.o, without skeleton or vmlinux.h generation) is left unchanged; replacing the existing duplication mentioned above is the next step. Suggested-by: Shakeel Butt Suggested-by: Eduard Zingerman Suggested-by: Mykola Lysenko Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Ziyang Men --- tools/testing/selftests/lib.bpf.mk | 247 +++++++++++++++++++++++++++++ 1 file changed, 247 insertions(+) create mode 100644 tools/testing/selftests/lib.bpf.mk diff --git a/tools/testing/selftests/lib.bpf.mk b/tools/testing/selftests/l= ib.bpf.mk new file mode 100644 index 000000000000..6f175c6568e9 --- /dev/null +++ b/tools/testing/selftests/lib.bpf.mk @@ -0,0 +1,247 @@ +# SPDX-License-Identifier: GPL-2.0 +# +# Shared fragment for selftests that compile *.bpf.c into BPF objects + +# skeletons and link them into userspace test binaries, without each +# subsystem's Makefile re-implementing the libbpf/bpftool/vmlinux.h machin= ery. +# +# Caller contract (per-test Makefile): +# +# BPF_SRCS :=3D foo.bpf.c bar.bpf.c +# TEST_GEN_PROGS :=3D foo_test +# OVERRIDE_TARGETS :=3D 1 # MUST be set before lib.mk +# include ../lib.mk # defines OUTPUT, CC, Q, msg, +# include ../lib.bpf.mk # selfdir, top_srcdir; honours = OVERRIDE +# +# $(OUTPUT)/foo_test: foo_test.c $(BPF_SKELS) +# $(call bpf_link,$@,$<) +# +# Optional knobs (set before including lib.bpf.mk): +# BPF_PROG_EXT - source suffix, default .bpf.c; set to .c for the le= gacy +# "progs/foo.c -> foo.bpf.o" layout. +# BPF_EXTRA_HDRS - extra prerequisites (headers) for the BPF objects. +# BPF_EXTRA_CFLAGS - appended to BPF_CFLAGS for the BPF compile. +# BPF_SKEL_EXT - skeleton suffix, default .skel.h (e.g. .bpf.skel.h). +# BPF_GEN_SUBSKEL - if set, also emit a subskeleton next to each skelet= on. +# BPF_OBJ_DIR - dir for generated *.bpf.o (default $(OUTPUT)). +# BPF_SKEL_DIR - dir for generated skeletons (default $(OUTPUT)). +# A caller with unusual compile needs may override BPF_CFLAGS wholesale af= ter +# the include (the object recipe expands it lazily). +# BPF_SRCS entries may live in a subdirectory (e.g. progs/foo.bpf.c); the +# objects and skeletons are always emitted flat under $(OUTPUT), keyed by = the +# source basename (foo.bpf.o / foo.skel.h). +# +# lib.mk MUST be included first: this fragment consumes the vars it defines +# (OUTPUT, top_srcdir, CC, CLANG, Q, msg) and needs OVERRIDE_TARGETS to ha= ve +# already suppressed lib.mk's default link rule. + +include $(top_srcdir)/tools/scripts/Makefile.arch # ARCH / SRCARCH +# Pull in the shared toolchain definitions (HOSTCC/HOSTLD/CLANG) so this +# fragment picks the *same* host compiler the libbpf and bpftool sub-makes +# will: in particular HOSTCC becomes clang under LLVM=3D1 (gcc otherwise). +# Without this the host bpftool and its bootstrap libbpf can end up built = with +# a mix of gcc and clang, which trips clang on gcc-only flags (-Wstrict-al= iasing=3D3). +# +# Makefile.include's allow-override resets CC to a bare "clang", which wou= ld +# drop the --target=3D flag lib.mk set for LLVM=3D1 cross builds (ma= ke LLVM=3D1 +# ARCH=3Darm64). lib.mk ran first and configured CC for the target, so sa= ve it +# across the include and restore it afterwards. +lib_bpf_mk_saved_cc :=3D $(CC) +include $(top_srcdir)/tools/scripts/Makefile.include +CC :=3D $(lib_bpf_mk_saved_cc) + +CLANG ?=3D clang +HOSTCC ?=3D gcc +HOSTLD ?=3D ld +ifneq ($(V),1) +submake_extras :=3D feature_display=3D0 +endif + +# ---- paths & tools -----------------------------------------------------= --- +TOOLSDIR :=3D $(top_srcdir)/tools +LIBDIR :=3D $(TOOLSDIR)/lib +BPFDIR :=3D $(LIBDIR)/bpf +TOOLSINCDIR :=3D $(TOOLSDIR)/include +BPFTOOLDIR :=3D $(TOOLSDIR)/bpf/bpftool +APIDIR :=3D $(TOOLSINCDIR)/uapi + +# Everything generated lives under $(OUTPUT) so O=3D and in-tree both work= and +# per-test builds (distinct $(OUTPUT)) never collide. +SCRATCH_DIR :=3D $(OUTPUT)/tools +BUILD_DIR :=3D $(SCRATCH_DIR)/build +INCLUDE_DIR :=3D $(SCRATCH_DIR)/include +BPFOBJ :=3D $(BUILD_DIR)/libbpf/libbpf.a + +# bpftool must run on the *host*; split the host toolchain out when cross-= building. +ifneq ($(CROSS_COMPILE),) +HOST_BUILD_DIR :=3D $(BUILD_DIR)/host +HOST_SCRATCH_DIR :=3D $(OUTPUT)/host-tools +else +HOST_BUILD_DIR :=3D $(BUILD_DIR) +HOST_SCRATCH_DIR :=3D $(SCRATCH_DIR) +endif +HOST_BPFOBJ :=3D $(HOST_BUILD_DIR)/libbpf/libbpf.a +DEFAULT_BPFTOOL :=3D $(HOST_SCRATCH_DIR)/sbin/bpftool +BPFTOOL ?=3D $(DEFAULT_BPFTOOL) + +# ---- vmlinux BTF discovery ---------------------------------------------= -- +VMLINUX_BTF_PATHS ?=3D $(if $(O),$(O)/vmlinux) \ + $(if $(KBUILD_OUTPUT),$(KBUILD_OUTPUT)/vmlinux) \ + $(top_srcdir)/vmlinux \ + /sys/kernel/btf/vmlinux \ + /boot/vmlinux-$(shell uname -r) +VMLINUX_BTF ?=3D $(abspath $(firstword $(wildcard $(VMLINUX_BTF_PATHS)))) +ifeq ($(VMLINUX_BTF),) +$(error Cannot find a vmlinux for VMLINUX_BTF at any of "$(VMLINUX_BTF_PAT= HS)") +endif + +# ---- clang flags -------------------------------------------------------= --- +# Clang's default system includes (not the ones seen under --target=3Dbpf)= ; fixes +# "missing" asm/byteorder.h etc. '-idirafter' so we never shadow real incl= udes. +define get_sys_includes +$(shell $(1) $(2) -v -E - &1 \ + | sed -n '/<...> search starts here:/,/End of search list./{ s| \(/.*\)|-= idirafter \1|p }') \ +$(shell $(1) $(2) -dM -E - &1 | gre= p -q 'v3' \ + && echo v3 || echo v2) + +# -fms-extensions + -Wno-microsoft-anon-tag: required so clang accepts the +# anonymous nested struct/union members bpftool emits into vmlinux.h. +BPF_CFLAGS =3D -g -Wall -Werror -D__TARGET_ARCH_$(SRCARCH) $(MENDIAN) \ + -I$(INCLUDE_DIR) -I$(APIDIR) -I$(TOOLSINCDIR) \ + -std=3Dgnu11 \ + -fno-strict-aliasing \ + -fms-extensions -Wno-microsoft-anon-tag \ + -Wno-compare-distinct-pointer-types \ + $(CLANG_SYS_INCLUDES) $(BPF_EXTRA_CFLAGS) + +# $1 =3D src .bpf.c, $2 =3D dst .bpf.o +define BPF_BUILD_RULE + $(call msg,CLNG-BPF,,$2) + $(Q)$(CLANG) $(BPF_CFLAGS) -O2 --target=3Dbpf -mcpu=3D$(CLANG_BPF_CPU) -c= $1 -o $2 +endef + +# ---- output dirs for generated objects/skeletons -----------------------= --- +# Default: flat under $(OUTPUT) (what bpf/, hid/, cgroup/ do). A caller m= ay +# segregate the generated files into subdirs, e.g. BPF_SKEL_DIR :=3D $(OUT= PUT)/... +BPF_OBJ_DIR ?=3D $(OUTPUT) +BPF_SKEL_DIR ?=3D $(OUTPUT) + +# ---- scratch dirs ------------------------------------------------------= --- +MAKE_DIRS :=3D $(sort $(BUILD_DIR)/libbpf $(HOST_BUILD_DIR)/libbpf \ + $(HOST_BUILD_DIR)/bpftool $(INCLUDE_DIR) \ + $(filter-out $(OUTPUT),$(BPF_OBJ_DIR) $(BPF_SKEL_DIR))) +$(MAKE_DIRS): + $(call msg,MKDIR,,$@) + $(Q)mkdir -p $@ + +# ---- libbpf (target) ---------------------------------------------------= --- +# Pass ARCH/CROSS_COMPILE/CC through: lib.mk's CC is file-origin and is not +# exported, so without this the libbpf sub-make would rebuild for the host= under +# a pure-LLVM cross build (make LLVM=3D1 ARCH=3D). -fPIC keeps the = static +# libbpf linkable into position-independent (PIE) test binaries. +$(BPFOBJ): $(wildcard $(BPFDIR)/*.[ch] $(BPFDIR)/Makefile) \ + $(APIDIR)/linux/bpf.h | $(BUILD_DIR)/libbpf + $(Q)$(MAKE) $(submake_extras) -C $(BPFDIR) OUTPUT=3D$(BUILD_DIR)/libbpf/ \ + ARCH=3D$(ARCH) CROSS_COMPILE=3D$(CROSS_COMPILE) CC=3D"$(CC)" \ + EXTRA_CFLAGS=3D'-g -O0 -fPIC' \ + DESTDIR=3D$(SCRATCH_DIR) prefix=3D all install_headers + +# ---- libbpf (host) -- a distinct rule only when cross-compiling --------= --- +ifneq ($(BPFOBJ),$(HOST_BPFOBJ)) +$(HOST_BPFOBJ): $(wildcard $(BPFDIR)/*.[ch] $(BPFDIR)/Makefile) \ + | $(HOST_BUILD_DIR)/libbpf + $(Q)$(MAKE) $(submake_extras) -C $(BPFDIR) ARCH=3D CROSS_COMPILE=3D = \ + OUTPUT=3D$(HOST_BUILD_DIR)/libbpf/ CC=3D$(HOSTCC) LD=3D$(HOSTLD) \ + EXTRA_CFLAGS=3D'-g -O0' \ + DESTDIR=3D$(HOST_SCRATCH_DIR) prefix=3D all install_headers +endif + +# ---- bpftool (host) ----------------------------------------------------= --- +$(DEFAULT_BPFTOOL): $(wildcard $(BPFTOOLDIR)/*.[ch] $(BPFTOOLDIR)/Makefile= ) \ + $(HOST_BPFOBJ) | $(HOST_BUILD_DIR)/bpftool + $(Q)$(MAKE) $(submake_extras) -C $(BPFTOOLDIR) \ + ARCH=3D CROSS_COMPILE=3D CC=3D$(HOSTCC) LD=3D$(HOSTLD) \ + EXTRA_CFLAGS=3D'-g -O0' \ + OUTPUT=3D$(HOST_BUILD_DIR)/bpftool/ \ + LIBBPF_OUTPUT=3D$(HOST_BUILD_DIR)/libbpf/ \ + LIBBPF_DESTDIR=3D$(HOST_SCRATCH_DIR)/ \ + prefix=3D DESTDIR=3D$(HOST_SCRATCH_DIR)/ install-bin + +# ---- vmlinux.h ---------------------------------------------------------= --- +$(INCLUDE_DIR)/vmlinux.h: $(VMLINUX_BTF) $(BPFTOOL) | $(INCLUDE_DIR) +ifeq ($(VMLINUX_H),) + $(call msg,GEN,,$@) + $(Q)$(BPFTOOL) btf dump file $(VMLINUX_BTF) format c > $@ +else + $(call msg,CP,,$@) + $(Q)cp "$(VMLINUX_H)" $@ +endif + +# ---- BPF objects + skeletons -------------------------------------------= -- +# Sources may sit in a subdir and use the modern *.bpf.c or the legacy *.c +# suffix (BPF_PROG_EXT); objects go in $(BPF_OBJ_DIR) and skeletons in +# $(BPF_SKEL_DIR) (both default to $(OUTPUT)), keyed by the source basenam= e. +BPF_PROG_EXT ?=3D .bpf.c +bpf_stems :=3D $(patsubst %$(BPF_PROG_EXT),%,$(notdir $(BPF_SRCS))) +# The output namespace is flat, so two sources with the same basename would +# collapse into one object/skeleton; fail loudly instead of silently dropp= ing. +ifneq ($(words $(bpf_stems)),$(words $(sort $(bpf_stems)))) +$(error lib.bpf.mk: BPF_SRCS has colliding basenames: $(BPF_SRCS)) +endif +# BPF_SKEL_EXT lets a caller pick the skeleton suffix (default .skel.h; e.= g. +# .bpf.skel.h). With BPF_GEN_SUBSKEL set, a matching subskeleton is emitt= ed +# alongside (foo.subskel.h / foo.bpf.subskel.h). +BPF_SKEL_EXT ?=3D .skel.h +BPF_SUBSKEL_EXT :=3D $(patsubst %skel.h,%subskel.h,$(BPF_SKEL_EXT)) +BPF_OBJS :=3D $(addprefix $(BPF_OBJ_DIR)/,$(addsuffix .bpf.o,$(bpf_stems)= )) +BPF_SKELS :=3D $(addprefix $(BPF_SKEL_DIR)/,$(addsuffix $(BPF_SKEL_EXT),$(= bpf_stems))) + +# Locate the sources wherever the caller keeps them (e.g. progs/). +vpath %$(BPF_PROG_EXT) $(sort $(dir $(BPF_SRCS))) + +$(BPF_OBJS): $(BPF_OBJ_DIR)/%.bpf.o: %$(BPF_PROG_EXT) $(BPF_EXTRA_HDRS) \ + $(wildcard *.bpf.h) $(INCLUDE_DIR)/vmlinux.h | $(BPF_OBJ_DIR) $(BPFO= BJ) + $(call BPF_BUILD_RULE,$<,$@) + +$(BPF_SKELS): $(BPF_SKEL_DIR)/%$(BPF_SKEL_EXT): $(BPF_OBJ_DIR)/%.bpf.o $(B= PFTOOL) | $(BPF_SKEL_DIR) + $(call msg,GEN-SKEL,,$@) + $(Q)$(BPFTOOL) gen object $(<:.o=3D.linked.o) $< + $(Q)$(BPFTOOL) gen skeleton $(<:.o=3D.linked.o) name $(notdir $(<:.bpf.o= =3D)) > $@ +ifneq ($(BPF_GEN_SUBSKEL),) + $(Q)$(BPFTOOL) gen subskeleton $(<:.o=3D.linked.o) name $(notdir $(<:.bpf= .o=3D)) > $(@:$(BPF_SKEL_EXT)=3D$(BPF_SUBSKEL_EXT)) +endif + +# ---- exports consumed by the caller ------------------------------------= -- +# -I$(OUTPUT)/-I$(BPF_SKEL_DIR): so the test .c can #include "foo.skel.h". +# -I$(INCLUDE_DIR): so userspace can pull in the generated vmlinux.h if ne= eded. +CFLAGS +=3D -I$(OUTPUT) -I$(BPF_SKEL_DIR) -I$(INCLUDE_DIR) + +# Static libbpf.a first, then its deps. libbpf may pull in zstd (BTF decom= press) +# only when built against it; link -lzstd only if libzstd is present. +BPF_LDLIBS :=3D $(BPFOBJ) -lelf -lz +ifneq ($(shell pkg-config --exists libzstd 2>/dev/null && echo y),) +BPF_LDLIBS +=3D -lzstd +endif + +TEST_GEN_FILES +=3D $(BPF_OBJS) + +# Link helper: $1 =3D output binary, $2 =3D test .c (skels are the target'= s deps). +define bpf_link + $(call msg,BINARY,,$1) + $(Q)$(CC) $(CFLAGS) $2 $(BPF_LDLIBS) $(LDLIBS) -o $1 +endef + +EXTRA_CLEAN +=3D $(SCRATCH_DIR) $(HOST_SCRATCH_DIR) \ + $(addprefix $(BPF_OBJ_DIR)/,*.bpf.o *.linked.o) \ + $(addprefix $(BPF_SKEL_DIR)/,*$(BPF_SKEL_EXT) *$(BPF_SUBSKEL_EXT)) --=20 2.53.0-Meta From nobody Sat Jul 25 00:11:12 2026 Received: from mail-pj2-f11.google.com (mail-pj2-f11.google.com [74.125.227.139]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 36B5A4F7994 for ; Tue, 21 Jul 2026 17:48:42 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.227.139 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784656127; cv=none; b=JvaawKT7DMbJbZRbkMyQlG+Ityda7wE/1f5eEBSdutd2BhzCS1YJeagR3uTcbmxgH6ZgzjPm9TAv/cFRaLjQHHxZALbsYgkHMPn5zG/qPqgfHSpfKHV+cOD7YCem3zVm77JSVz99R6unRggXw8Mo7NjyrN4PbrffIBJWfmFwovk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784656127; c=relaxed/simple; bh=vp2J4n0CbSrGTpuHrSDYdICgPaD/f9fPdyLRV/2i+Ic=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Q/lPMgYIL4LmnCiejsQLa+nZkGIyLF8DK0MFKr9fRTyWXTAaICgYZDLSmqv1DOWpgfKIiyKcJ+v243MRhhlV5GR1TTr97OSX2XPTEQYAKuin9WuCKW8ng/kNoejm9dhmIEIM1JGvOLpWFFAGcBuI4+MoEDWy2LkEq12+CcXE8Go= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=QZxwMPJb; arc=none smtp.client-ip=74.125.227.139 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="QZxwMPJb" Received: by mail-pj2-f11.google.com with SMTP id 98e67ed59e1d1-38e47da25f8so2618610a91.1 for ; Tue, 21 Jul 2026 10:48:42 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784656121; x=1785260921; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=AKxksonGf+ONkWkO5QIWgd6TDCo2B078e/Xr4EoxJss=; b=QZxwMPJbOmtByA0lssLtKg10RxRAPef/5LKxEAfDlsu39+yplZiFqhECr8iz9JmWtc sHQ4H8LxNTecSajJmfOOlQOQMVUoT5RAeyfLWQDj54o9ZjWgyFu/gPgVDXYefScpGv3W sNPJwDQqAgNVmMm/BUcGepqSf4EzUrdIpXnJ1iyQWweZSPn/xjHyYkIHWL4Q3L9LYmux GNmafpUm0LGSR0Vzqs/aL01VPrT16goZdKb4KhJN0HrqlfixozpaymlqTpMhH3kZb4Sh csREo0wjjLgqILoDuVSpq+S8rK1g8FFM56ZZe/sn+KA3CfCBIkL932s6rzRYXu5eRWyB zxXw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784656121; x=1785260921; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=AKxksonGf+ONkWkO5QIWgd6TDCo2B078e/Xr4EoxJss=; b=BWSQY0HcC6/Cdnf9sOMLxMkW5RxtZfOZHtIL7/qEE64+U2H60uGTsRcAjucNVNNNkA 9TpmGRnyyCPjNMD/ZkBo4kUrP6mo4ugkqbRq7n1WHdAyivVK7ZBoCZsFq8it3NaySaR7 VTgoLBA7rsmp6g+hVi6gqQqyM4bO7a0rl46JJg8VfbJXpU0IqBkfPCsmQBuiMqQDUSpk q8ZIkQyGCEupi5ZdB6+V+WNqtfJBKxSScbQHVULmpW9NJdSK74Exj8fwploCLsl1RcDz aRU5otU1sCwMCk6qLLY6W7K/Q2+2HAYBjkk0q0W30TqCm68nt5aFh/5zwOWaFXfEWu1q MkSg== X-Forwarded-Encrypted: i=1; AHgh+RodiP59sX4z5b4XLjHGO/5KatQrJ0NVFaXD9mpIH2KLMMxnr2Gg5iudRX/vD8LizKnSGwXvqIsSYfogMPI=@vger.kernel.org X-Gm-Message-State: AOJu0YxlQ4T0mzknAYUjoUnS0lLujlpdRiuyFtt8gZAHdYj2ERVCoQI+ ssD9eaL2jBlsLIMuDAk7W3mD5wk0WRTMi2v4HVDiNk5cYcyitP4kLYR9 X-Gm-Gg: AR+sD11atSNfEz3oDgyEKKhkxr25l5y8NCzNSu49sdkL8sLsLYuQXzuICKmdS9MCFsq ygQVPbwA4eLK3haasZjvMYON6vc1YMJmoWKLjqglG/ruETVJb8Hc+nFWKiDRqP7I8NC+MGIcteG hmisTfppIPyWbPbxdRc1oCOyVGO7xNQHfDsafdXE+iEzhXZ9nMsTHdRN+GXYUuBtm6Nz6y7VopE ODMK2kJH3gmOmoJDwgNKQwuQLdnxpgLbC9SGKrSJ4qkfcwO0rBQ6OOHYrkWpnkvX2xQrfHMh1vN xUedQYJfkkOSas4yyKFRv1YPsV7BzWYOWAgxo7j7EKe2Mxq0kOk2nJ8XFbRpu/l/Fr3+szm2T0G x5CRu2lWuFqRv9iKS0De1XmpyPI31hF2JF3BgjoAunb2mOBLu86iFov5LS3sBJkagKJQq0U1c X-Received: by 2002:a05:6a21:1404:b0:3c3:856e:dccd with SMTP id adf61e73a8af0-3c3ad7ec9camr21040818637.23.1784656121245; Tue, 21 Jul 2026 10:48:41 -0700 (PDT) Received: from localhost ([2a03:2880:9ff:68::]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-3147e070538sm814778eec.20.2026.07.21.10.48.39 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 21 Jul 2026 10:48:40 -0700 (PDT) From: Ziyang Men To: Shuah Khan , Tejun Heo , Johannes Weiner , =?UTF-8?q?Michal=20Koutn=C3=BD?= , Jiri Kosina , Benjamin Tissoires , David Vernet , Eduard Zingerman Cc: Andrea Righi , Changwoo Min , Michal Hocko , Roman Gushchin , Shakeel Butt , Muchun Song , Andrew Morton , JP Kobryn , Mykola Lysenko , Nathan Chancellor , linux-kselftest@vger.kernel.org, cgroups@vger.kernel.org, linux-input@vger.kernel.org, sched-ext@lists.linux.dev, linux-mm@kvack.org, kernel-team@meta.com, bpf@vger.kernel.org, llvm@lists.linux.dev, linux-kernel@vger.kernel.org, Ziyang Men Subject: [PATCH v2 2/4] selftests/cgroup: add memcg_stat_cross_cpu correctness test for flush Date: Tue, 21 Jul 2026 10:48:31 -0700 Message-ID: <20260721174833.1232771-3-ziyang.meme@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260721174833.1232771-1-ziyang.meme@gmail.com> References: <20260721174833.1232771-1-ziyang.meme@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add a test_progs selftest that verifies the memory-cgroup BPF kfuncs return values that agree with what userspace reads from cgroupfs, across a whole cgroup subtree that has been charged on many CPUs. It complements the existing cgroup_iter_memcg test. cgroup_iter_memcg calls the memcg kfuncs on a single cgroup (BPF_CGROUP_ITER_SELF_ONLY) and only asserts each value is greater than zero, so it never checks that a value is actually correct. This test compares the kfunc values against what userspace reads from memory.stat, and checks every node in the cgroup tree. Moreover, cgroup_iter_memcg does not exercise the rstat flush (mem_cgroup_flush_stats()). This test does: it launches a process on each leaf that charges and holds memory across multiple CPUs (so the leaf's rstat is dirty on K per-cpu trees), then confirms the flush works for both the BPF and the file path by checking two properties: 1. after the flush, each leaf's anon is at least what the process charged there; 2. after the flush, the sum of the leaves' charged memory equals the amount at the root. The BPF reader and the file reader run in two separate rounds. Each round builds the same cgroup tree from scratch and charges the same amount of memory, so both rounds start from the same state and the numbers are comparable. In detail: - round 1 (BPF) builds the subtree, forks one child per leaf that charges the leaf across K CPUs and then blocks holding the charge, walks the tree with a SEC("iter.s/cgroup") program that flushes the subtree at the root and reads each cgroup via the memcg kfuncs (bpf_get_mem_cgroup, bpf_mem_cgroup_flush_stats, bpf_mem_cgroup_page_state, bpf_mem_cgroup_vm_events, bpf_put_mem_cgroup) into a hash map; - round 2 (cgroupfs) builds and charges an identical tree the same way, then reads every cgroup's memory.stat / memory.current from userspace. The subtests differ in how many CPUs each leaf is charged on (a single CPU or across K CPUs), and run on two cgroup tree sizes. The charging children pin CPUs and there is one per leaf, so the test is registered serial. The traditional path reads memory.stat / memory.current through a new read_cgroup_file() helper added to cgroup_helpers (the read counterpart of write_cgroup_file). When the memcg kfuncs are unavailable (CONFIG_MEMCG=3Dn) the test skips cleanly; the base selftest config now selects CONFIG_MEMCG=3Dy. Suggested-by: Shakeel Butt Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Ziyang Men --- tools/testing/selftests/cgroup/.gitignore | 6 + tools/testing/selftests/cgroup/Makefile | 21 + tools/testing/selftests/cgroup/config | 4 + .../selftests/cgroup/lib/cgroup_util.c | 48 ++ .../cgroup/lib/include/cgroup_util.h | 1 + .../cgroup/memcg_stat_cross_cpu.bpf.c | 86 ++ .../selftests/cgroup/memcg_stat_cross_cpu.h | 27 + .../cgroup/test_memcg_stat_cross_cpu.c | 780 ++++++++++++++++++ 8 files changed, 973 insertions(+) create mode 100644 tools/testing/selftests/cgroup/memcg_stat_cross_cpu.bpf= .c create mode 100644 tools/testing/selftests/cgroup/memcg_stat_cross_cpu.h create mode 100644 tools/testing/selftests/cgroup/test_memcg_stat_cross_cp= u.c diff --git a/tools/testing/selftests/cgroup/.gitignore b/tools/testing/self= tests/cgroup/.gitignore index 952e4448bf07..9e6616af832e 100644 --- a/tools/testing/selftests/cgroup/.gitignore +++ b/tools/testing/selftests/cgroup/.gitignore @@ -6,7 +6,13 @@ test_freezer test_hugetlb_memcg test_kill test_kmem +test_memcg_stat_cross_cpu test_memcontrol test_pids test_zswap wait_inotify +# Artifacts generated by lib.bpf.mk +/tools +*.bpf.o +*.linked.o +*.skel.h diff --git a/tools/testing/selftests/cgroup/Makefile b/tools/testing/selfte= sts/cgroup/Makefile index e01584c2189a..daebd35aec79 100644 --- a/tools/testing/selftests/cgroup/Makefile +++ b/tools/testing/selftests/cgroup/Makefile @@ -14,14 +14,29 @@ TEST_GEN_PROGS +=3D test_freezer TEST_GEN_PROGS +=3D test_hugetlb_memcg TEST_GEN_PROGS +=3D test_kill TEST_GEN_PROGS +=3D test_kmem +TEST_GEN_PROGS +=3D test_memcg_stat_cross_cpu TEST_GEN_PROGS +=3D test_memcontrol TEST_GEN_PROGS +=3D test_pids TEST_GEN_PROGS +=3D test_zswap =20 LOCAL_HDRS +=3D $(selfdir)/clone3/clone3_selftests.h $(selfdir)/pidfd/pidf= d.h =20 +# test_memcg_stat_cross_cpu builds a BPF program + skeleton through lib.bp= f.mk. +# OVERRIDE_TARGETS suppresses lib.mk's default C link rule (re-supplied be= low); +# it must be set before ../lib.mk is included. +BPF_SRCS :=3D memcg_stat_cross_cpu.bpf.c +OVERRIDE_TARGETS :=3D 1 + include ../lib.mk include lib/libcgroup.mk +include ../lib.bpf.mk + +# Re-supply the default C link rule that OVERRIDE_TARGETS removed, for the= plain +# cgroup tests. test_memcg_stat_cross_cpu has its own recipe further belo= w. +LOCAL_HDRS +=3D $(selfdir)/kselftest_harness.h $(selfdir)/kselftest.h +$(OUTPUT)/%: %.c $(LOCAL_HDRS) + $(call msg,CC,,$@) + $(Q)$(LINK.c) $(filter-out $(LOCAL_HDRS),$^) $(LDLIBS) -o $@ =20 $(OUTPUT)/test_core: $(LIBCGROUP_O) $(OUTPUT)/test_cpu: $(LIBCGROUP_O) @@ -33,3 +48,9 @@ $(OUTPUT)/test_kmem: $(LIBCGROUP_O) $(OUTPUT)/test_memcontrol: $(LIBCGROUP_O) $(OUTPUT)/test_pids: $(LIBCGROUP_O) $(OUTPUT)/test_zswap: $(LIBCGROUP_O) + +# test_memcg_stat_cross_cpu links cgroup_util and the generated BPF skelet= on +# against the in-tree static libbpf that lib.bpf.mk built. +$(OUTPUT)/test_memcg_stat_cross_cpu: test_memcg_stat_cross_cpu.c \ + $(BPF_SKELS) $(LIBCGROUP_O) + $(call bpf_link,$@,$< $(LIBCGROUP_O)) diff --git a/tools/testing/selftests/cgroup/config b/tools/testing/selftest= s/cgroup/config index 39f979690dd3..9457bf604b23 100644 --- a/tools/testing/selftests/cgroup/config +++ b/tools/testing/selftests/cgroup/config @@ -4,3 +4,7 @@ CONFIG_CGROUP_FREEZER=3Dy CONFIG_CGROUP_SCHED=3Dy CONFIG_MEMCG=3Dy CONFIG_PAGE_COUNTER=3Dy +CONFIG_BPF=3Dy +CONFIG_BPF_SYSCALL=3Dy +CONFIG_CGROUP_BPF=3Dy +CONFIG_DEBUG_INFO_BTF=3Dy diff --git a/tools/testing/selftests/cgroup/lib/cgroup_util.c b/tools/testi= ng/selftests/cgroup/lib/cgroup_util.c index 2596c12cd864..3a30557855d3 100644 --- a/tools/testing/selftests/cgroup/lib/cgroup_util.c +++ b/tools/testing/selftests/cgroup/lib/cgroup_util.c @@ -54,6 +54,54 @@ ssize_t write_text(const char *path, char *buf, ssize_t = len) return len < 0 ? -errno : len; } =20 +/* + * cg_get_id - return the kernfs id of a cgroup directory + * @cgroup: absolute path to the cgroup directory + * + * Returns the cgroup's kernfs node id (cgrp->kn->id) -- the same value the + * kernel exposes to BPF as cgrp->kn->id and via bpf_get_current_cgroup_id= (). + * This is obtained from the cgroupfs file handle and is NOT the directory= 's + * st_ino. Returns 0 (an invalid id) on failure. + */ +unsigned long long cg_get_id(const char *cgroup) +{ + union { + unsigned long long id; + unsigned char raw[8]; + } handle; + struct file_handle *fhp, *fhp2; + int mount_id, fhsize, err; + unsigned long long ret =3D 0; + + fhsize =3D sizeof(*fhp); + fhp =3D calloc(1, fhsize); + if (!fhp) + return 0; + + /* + * The probe call is expected to fail (EOVERFLOW) and report the real + * handle size in fhp->handle_bytes; a cgroupfs handle is always 8 bytes. + */ + err =3D name_to_handle_at(AT_FDCWD, cgroup, fhp, &mount_id, 0); + if (err >=3D 0 || fhp->handle_bytes !=3D 8) + goto out; + + fhsize =3D sizeof(*fhp) + fhp->handle_bytes; + fhp2 =3D realloc(fhp, fhsize); + if (!fhp2) + goto out; + fhp =3D fhp2; + + if (name_to_handle_at(AT_FDCWD, cgroup, fhp, &mount_id, 0) < 0) + goto out; + + memcpy(handle.raw, fhp->f_handle, 8); + ret =3D handle.id; +out: + free(fhp); + return ret; +} + char *cg_name(const char *root, const char *name) { size_t len =3D strlen(root) + strlen(name) + 2; diff --git a/tools/testing/selftests/cgroup/lib/include/cgroup_util.h b/too= ls/testing/selftests/cgroup/lib/include/cgroup_util.h index 8ebb2b4d4ec0..832fcefe60a4 100644 --- a/tools/testing/selftests/cgroup/lib/include/cgroup_util.h +++ b/tools/testing/selftests/cgroup/lib/include/cgroup_util.h @@ -54,6 +54,7 @@ extern ssize_t write_text(const char *path, char *buf, ss= ize_t len); extern int cg_find_controller_root(char *root, size_t len, const char *con= troller); extern int cg_find_unified_root(char *root, size_t len, bool *nsdelegate); extern char *cg_name(const char *root, const char *name); +extern unsigned long long cg_get_id(const char *cgroup); extern char *cg_name_indexed(const char *root, const char *name, int index= ); extern char *cg_control(const char *cgroup, const char *control); extern int cg_create(const char *cgroup); diff --git a/tools/testing/selftests/cgroup/memcg_stat_cross_cpu.bpf.c b/to= ols/testing/selftests/cgroup/memcg_stat_cross_cpu.bpf.c new file mode 100644 index 000000000000..3b8c716e8d01 --- /dev/null +++ b/tools/testing/selftests/cgroup/memcg_stat_cross_cpu.bpf.c @@ -0,0 +1,86 @@ +// SPDX-License-Identifier: GPL-2.0 +/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */ +#include +#include +#include +#include "memcg_stat_cross_cpu.h" + +char _license[] SEC("license") =3D "GPL"; + +/* + * Per-cgroup results, keyed by cgroup id. The BPF-side id (cgrp->kn->id) + * equals the userspace get_cgroup_id() value, so the test can correlate m= ap + * entries back to the cgroups it created. max_entries is resized by user= space + * (bpf_map__set_max_entries) to the size of the subtree before load. + */ +struct { + __uint(type, BPF_MAP_TYPE_HASH); + __uint(max_entries, 1); + __type(key, __u64); + __type(value, struct memcg_stat_snapshot); +} results SEC(".maps"); + +/* + * Sleepable cgroup iterator: flush the subtree's rstat once at the root, = then + * for every cgroup in the walked subtree read a fixed set of memcg statis= tics + * through the memcg kfuncs and stash them in the hash map for userspace to + * compare against memory.stat. + * + * The flush kfunc may sleep, hence SEC("iter.s/cgroup"). + */ +SEC("iter.s/cgroup") +int cgroup_memcg_stat_cross_cpu(struct bpf_iter__cgroup *ctx) +{ + struct cgroup *cgrp =3D ctx->cgroup; + struct memcg_stat_snapshot snap =3D {}; + struct cgroup_subsys_state *css; + struct mem_cgroup *memcg; + int idx_anon, idx_file, idx_shmem, idx_fmapped, idx_pgfault; + __u64 cg_id; + + /* + * DESCENDANTS_PRE ends with a terminal element where cgroup =3D=3D NULL. + * Return 0 (not 1) so the walk runs to completion. + */ + if (!cgrp) + return 0; + + css =3D &cgrp->self; + memcg =3D bpf_get_mem_cgroup(css); + if (!memcg) + return 0; + + /* + * Flush once, at the subtree root -- the first element visited in + * DESCENDANTS_PRE order (seq_num =3D=3D 0). css_rstat_flush() is + * subtree-wide, so this one flush brings the whole walked subtree + * up to date and every descendant read afterwards is accurate; + * flushing again per-cgroup would only hit the no-op threshold gate. + */ + if (ctx->meta->seq_num =3D=3D 0) + bpf_mem_cgroup_flush_stats(memcg); + + cg_id =3D BPF_CORE_READ(cgrp, kn, id); + snap.cgroup_id =3D cg_id; + + idx_anon =3D bpf_core_enum_value(enum node_stat_item, NR_ANON_MAPPED); + idx_file =3D bpf_core_enum_value(enum node_stat_item, NR_FILE_PAGES); + idx_shmem =3D bpf_core_enum_value(enum node_stat_item, NR_SHMEM); + idx_fmapped =3D bpf_core_enum_value(enum node_stat_item, NR_FILE_MAPPED); + idx_pgfault =3D bpf_core_enum_value(enum vm_event_item, PGFAULT); + + snap.anon =3D bpf_mem_cgroup_page_state(memcg, idx_anon); + snap.file =3D bpf_mem_cgroup_page_state(memcg, idx_file); + snap.shmem =3D bpf_mem_cgroup_page_state(memcg, idx_shmem); + snap.file_mapped =3D bpf_mem_cgroup_page_state(memcg, idx_fmapped); + snap.pgfault =3D bpf_mem_cgroup_vm_events(memcg, idx_pgfault); + + /* page_counter fields need no kfunc; read them off the trusted ptr. */ + snap.usage_pages =3D BPF_CORE_READ(memcg, memory.usage.counter); + snap.max_pages =3D BPF_CORE_READ(memcg, memory.max); + + bpf_map_update_elem(&results, &cg_id, &snap, BPF_ANY); + + bpf_put_mem_cgroup(memcg); + return 0; +} diff --git a/tools/testing/selftests/cgroup/memcg_stat_cross_cpu.h b/tools/= testing/selftests/cgroup/memcg_stat_cross_cpu.h new file mode 100644 index 000000000000..1583cafdab7e --- /dev/null +++ b/tools/testing/selftests/cgroup/memcg_stat_cross_cpu.h @@ -0,0 +1,27 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */ +#ifndef __MEMCG_STAT_CROSS_CPU_H +#define __MEMCG_STAT_CROSS_CPU_H + +/* + * One per-cgroup snapshot, produced by the BPF cgroup iterator and read b= ack + * from a BPF hash map keyed by cgroup id. After the charge has quiesced = the + * test compares every field against what userspace parses from + * memory.stat / memory.current / memory.max, so the two must agree. + * + * Page-state values are in bytes (already unit-scaled by the kernel), so = they + * compare directly against memory.stat. usage_pages / max_pages come str= aight + * off the page_counter and are in PAGES. + */ +struct memcg_stat_snapshot { + __u64 cgroup_id; + __u64 anon; /* NR_ANON_MAPPED, bytes */ + __u64 file; /* NR_FILE_PAGES, bytes */ + __u64 shmem; /* NR_SHMEM, bytes */ + __u64 file_mapped; /* NR_FILE_MAPPED, bytes */ + __u64 pgfault; /* PGFAULT, count */ + __u64 usage_pages; /* page_counter memory.usage, in PAGES */ + __u64 max_pages; /* page_counter memory.max, in PAGES */ +}; + +#endif /* __MEMCG_STAT_CROSS_CPU_H */ diff --git a/tools/testing/selftests/cgroup/test_memcg_stat_cross_cpu.c b/t= ools/testing/selftests/cgroup/test_memcg_stat_cross_cpu.c new file mode 100644 index 000000000000..673b3da467ba --- /dev/null +++ b/tools/testing/selftests/cgroup/test_memcg_stat_cross_cpu.c @@ -0,0 +1,780 @@ +// SPDX-License-Identifier: GPL-2.0 +/* Copyright (c) 2026 Meta Platforms, Inc. and affiliates. */ + +/* + * memcg_stat_cross_cpu + * =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D + * A memory-cgroup statistics correctness test. It compares the memcg + * statistics read through the BPF memcg kfuncs against what userspace rea= ds + * from memory.stat, over a whole cgroup tree that has been charged across= many + * CPUs. Where a plain kfunc smoke test only checks that a single cgroup's + * values are non-zero, this test checks the values are actually correct a= nd + * that a cross-CPU rstat flush aggregates every per-CPU slice. + * + * The BPF reader and the file reader are run in two separate rounds, each= on + * its own freshly built and charged tree: + * + * round 1: build the subtree; fork one process per leaf that charges th= e leaf + * across K CPUs and then blocks holding the charge; walk the t= ree + * with a SEC("iter.s/cgroup") program that flushes and reads e= ach + * cgroup via the memcg kfuncs into a hash map; record one snap= shot + * per node; clean the tree. + * round 2: build and charge an identical tree the same way, then read e= very + * cgroup's memory.stat / memory.current from userspace. + * + * The correctness is ensured as follows: each leaf was charged a known am= ount + * of anon spread over K CPUs, so the flushed anon must be at least that a= mount, + * and the subtree root's recursive anon must equal the sum of the leaves'= anon. + * + * The charging children are CPU-pinning and there is one per leaf. + */ +#define _GNU_SOURCE + +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include +#include +#include +#include + +#include "kselftest.h" +#include "cgroup_util.h" +#include "memcg_stat_cross_cpu.h" +#include "memcg_stat_cross_cpu.skel.h" + +#define SUBTREE_NAME "mcg_xcpu" + +static char root[PATH_MAX]; /* cgroup2 mount root */ +static char *subtree_root; /* /mcg_xcpu, from cg_name() */ + +struct cg_node { + char path[PATH_MAX]; + __u64 id; + bool is_leaf; +}; + +/* Field subset parsed from memory.stat / memory.current. */ +struct file_snap { + __u64 anon, file, shmem, file_mapped, pgfault; + __u64 current; /* memory.current, bytes */ + __u64 max; /* memory.max, bytes (valid unless max_is_max) */ + bool max_is_max; +}; + +static long page_size; + +/* ---- allowed CPU set --------------------------------------------------= - */ + +static int *cpu_list; /* ids of the CPUs this process may run on */ +static int n_cpu; /* number of such CPUs */ + +/* + * Collect the CPUs the test is allowed to run on. + * Bounded by CPU_SETSIZE (1024). + */ +static int collect_cpus(void) +{ + cpu_set_t set; + int i, want, n =3D 0; + + CPU_ZERO(&set); + if (sched_getaffinity(0, sizeof(set), &set)) + return -1; + want =3D CPU_COUNT(&set); + if (want <=3D 0) + return -1; + cpu_list =3D calloc(want, sizeof(*cpu_list)); + if (!cpu_list) + return -1; + for (i =3D 0; i < CPU_SETSIZE && n < want; i++) + if (CPU_ISSET(i, &set)) + cpu_list[n++] =3D i; + n_cpu =3D n; + return 0; +} + +/* Pin the calling task to a single CPU. */ +static int pin_cpu(int cpu) +{ + cpu_set_t set; + + CPU_ZERO(&set); + CPU_SET(cpu, &set); + return sched_setaffinity(0, sizeof(set), &set); +} + +/* ---- tree construction ------------------------------------------------= - */ + +static struct cg_node *nodes; +static int n_nodes; +static int n_leaves; + +static int add_node(const char *path, bool is_leaf, int *keep_fd) +{ + if (cg_create(path)) + return -1; + if (keep_fd) { + *keep_fd =3D open(path, O_RDONLY); + if (*keep_fd < 0) + return -1; + } + + strncpy(nodes[n_nodes].path, path, sizeof(nodes[n_nodes].path) - 1); + nodes[n_nodes].path[sizeof(nodes[n_nodes].path) - 1] =3D '\0'; + nodes[n_nodes].id =3D cg_get_id(path); + nodes[n_nodes].is_leaf =3D is_leaf; + if (is_leaf) + n_leaves++; + n_nodes++; + return 0; +} + +/* Recursively create children of @path. @path must already exist and be r= ecorded. */ +static int build_children(const char *path, int fanout, int depth) +{ + char child[PATH_MAX]; + int i; + + if (depth =3D=3D 0) + return 0; + + /* Enable memory on this interior node so its children get a memcg. */ + if (cg_write(path, "cgroup.subtree_control", "+memory")) + return -1; + + for (i =3D 0; i < fanout; i++) { + snprintf(child, sizeof(child), "%s/c%d", path, i); + if (add_node(child, depth =3D=3D 1, NULL)) + return -1; + if (build_children(child, fanout, depth - 1)) + return -1; + } + return 0; +} + +static size_t tree_capacity(int fanout, int depth) +{ + size_t total =3D 1, level =3D 1; + int d; + + for (d =3D 0; d < depth; d++) { + level *=3D fanout; + total +=3D level; + } + return total; +} + +/* The tree is arranged in the DFS order within an array */ +static int build_tree(int fanout, int depth, int *root_fd) +{ + n_nodes =3D 0; + n_leaves =3D 0; + nodes =3D calloc(tree_capacity(fanout, depth), sizeof(*nodes)); + if (!nodes) + return -1; + + if (add_node(subtree_root, depth =3D=3D 0, root_fd)) + return -1; + return build_children(subtree_root, fanout, depth); +} + +/* ---- cross-CPU charge (one pinning child per leaf) --------------------= - */ + +static pid_t *charger_pids; +static int n_chargers; +/* parent pid, for the children's PR_SET_PDEATHSIG race check */ +static pid_t test_pid; +static int charge_ready[2] =3D { -1, -1 }; /* child -> parent "ready" barr= ier */ +static int charge_ctrl[2] =3D { -1, -1 }; /* parent -> child "exit" (close= to signal) */ + +/* + * One charging child, dedicated to a single leaf and spread over K CPUs. = It + * joins its leaf, maps a resident anon region, then faults the region in K + * slices, each on a different CPU, so this leaf's rstat ends up dirty on K + * per-cpu trees. The region stays mapped, so the charge persists while t= he + * parent reads. After signalling readiness the child blocks (holding the + * charge) until the parent closes the control pipe. Never returns. + * + * @base is this child's starting index into cpu_list; its K CPUs are + * (base + 0..K-1) mod n_cpu. + */ +static void charger_child(const struct cg_node *leaf, int base, int k, + size_t resident_bytes) +{ + size_t per, off; + char *region; + char c; + int j; + + /* + * If the parent dies without running the cleanup that closes the control + * pipe -- e.g. a CI timeout SIGKILLs the whole test -- ask the kernel to + * SIGKILL this child too, so it can never hang as an orphan holding a + * charge. + */ + prctl(PR_SET_PDEATHSIG, SIGKILL); + if (getppid() !=3D test_pid) + _exit(0); + + close(charge_ready[0]); + close(charge_ctrl[1]); + + if (cg_enter_current(leaf->path)) + _exit(1); + + region =3D mmap(NULL, resident_bytes, PROT_READ | PROT_WRITE, + MAP_ANONYMOUS | MAP_PRIVATE, -1, 0); + if (region =3D=3D MAP_FAILED) + _exit(2); + + /* + * Fault the region in K slices, each on a different CPU, so the charge + * for this leaf is scattered across K per-cpu rstat trees. A correct + * flush must gather all K slices. + */ + per =3D resident_bytes / k; + for (j =3D 0; j < k; j++) { + off =3D (size_t)j * per; + if (pin_cpu(cpu_list[(base + j) % n_cpu])) + _exit(3); + memset(region + off, 1, + (j =3D=3D k - 1) ? resident_bytes - off : per); + } + + /* Ready: the charge is in place and spread across K CPUs. */ + if (write(charge_ready[1], "x", 1) !=3D 1) + _exit(4); + close(charge_ready[1]); + + /* Hold the charge (region stays mapped) until the parent tells + * us to exit by closing the control pipe. + */ + while (read(charge_ctrl[0], &c, 1) > 0) + ; + + munmap(region, resident_bytes); + _exit(0); +} + +/* + * Fork one charging child per leaf, each spread over K CPUs. Returns 0 o= nce + * every child has charged its leaf and is holding the charge, so the tree= is + * under a steady, quiesced load. On failure the caller's cleanup path ca= lls + * stop_chargers(). + */ +static int start_chargers(int cpus_per_leaf, size_t resident_bytes) +{ + int k_eff, i, h =3D 0; + + k_eff =3D (cpus_per_leaf > 0 && cpus_per_leaf <=3D n_cpu) ? cpus_per_leaf= : + n_cpu; + + if (pipe(charge_ready) || pipe(charge_ctrl)) { + ksft_print_msg("pipe: %s\n", strerror(errno)); + return -1; + } + + charger_pids =3D calloc(n_leaves, sizeof(*charger_pids)); + if (!charger_pids) { + ksft_print_msg("calloc charger_pids failed\n"); + return -1; + } + + /* recorded before the fork so each child can PR_SET_PDEATHSIG against us= */ + test_pid =3D getpid(); + + for (i =3D 0; i < n_nodes; i++) { + pid_t pid; + + if (!nodes[i].is_leaf) + continue; + + pid =3D fork(); + if (pid < 0) { + ksft_print_msg("fork charger: %s\n", strerror(errno)); + return -1; + } + if (pid =3D=3D 0) + charger_child(&nodes[i], h * k_eff, k_eff, + resident_bytes); + + charger_pids[n_chargers++] =3D pid; + h++; + } + + /* parent: keeps only the ready-read end and the ctrl-write end */ + close(charge_ready[1]); + charge_ready[1] =3D -1; + close(charge_ctrl[0]); + charge_ctrl[0] =3D -1; + + /* wait until every child has charged its leaf and is holding it */ + for (i =3D 0; i < n_chargers; i++) { + char c; + ssize_t r =3D read(charge_ready[0], &c, 1); + + if (r !=3D 1) { + ksft_print_msg("charger exited before ready (setup failed?)\n"); + return -1; + } + } + return 0; +} + +static void stop_chargers(void) +{ + int i, status; + + /* closing the ctrl write end unblocks every child -> they munmap + exit = */ + if (charge_ctrl[1] >=3D 0) { + close(charge_ctrl[1]); + charge_ctrl[1] =3D -1; + } + if (charge_ctrl[0] >=3D 0) { + close(charge_ctrl[0]); + charge_ctrl[0] =3D -1; + } + if (charge_ready[0] >=3D 0) { + close(charge_ready[0]); + charge_ready[0] =3D -1; + } + if (charge_ready[1] >=3D 0) { + close(charge_ready[1]); + charge_ready[1] =3D -1; + } + + for (i =3D 0; i < n_chargers; i++) { + if (!charger_pids || charger_pids[i] <=3D 0) + continue; + if (waitpid(charger_pids[i], &status, 0) =3D=3D charger_pids[i] && + (!WIFEXITED(status) || WEXITSTATUS(status) !=3D 0)) + ksft_print_msg("charger %d exited abnormally (status=3D0x%x)\n", + charger_pids[i], status); + } + + free(charger_pids); + charger_pids =3D NULL; + n_chargers =3D 0; +} + +/* ---- file (traditional) reader ----------------------------------------= - */ + +static void parse_stat(char *buf, struct file_snap *o) +{ + char *save, *line; + + for (line =3D strtok_r(buf, "\n", &save); line; + line =3D strtok_r(NULL, "\n", &save)) { + unsigned long long val; + char name[64]; + + if (sscanf(line, "%63s %llu", name, &val) !=3D 2) + continue; + if (!strcmp(name, "anon")) + o->anon =3D val; + else if (!strcmp(name, "file")) + o->file =3D val; + else if (!strcmp(name, "shmem")) + o->shmem =3D val; + else if (!strcmp(name, "file_mapped")) + o->file_mapped =3D val; + else if (!strcmp(name, "pgfault")) + o->pgfault =3D val; + } +} + +static int file_read_node(const char *path, struct file_snap *o) +{ + char buf[8192]; + + memset(o, 0, sizeof(*o)); + + if (cg_read(path, "memory.stat", buf, sizeof(buf))) + return -1; + parse_stat(buf, o); + + if (!cg_read(path, "memory.current", buf, sizeof(buf))) + o->current =3D strtoull(buf, NULL, 10); + if (!cg_read(path, "memory.max", buf, sizeof(buf))) { + if (!strncmp(buf, "max", 3)) + o->max_is_max =3D true; + else + o->max =3D strtoull(buf, NULL, 10); + } + return 0; +} + +/* ---- BPF reader -------------------------------------------------------= - */ + +static int bpf_walk_once(struct bpf_link *link) +{ + char buf[4096]; + ssize_t r; + int fd; + + fd =3D bpf_iter_create(bpf_link__fd(link)); + if (fd < 0) + return -1; + while ((r =3D read(fd, buf, sizeof(buf))) > 0) + ; + close(fd); + return r =3D=3D 0 ? 0 : -1; +} + +/* ---- correctness comparison -------------------------------------------= - */ + +static bool close_enough(__u64 a, __u64 b, __u64 tol) +{ + return (a > b ? a - b : b - a) <=3D tol; +} + +/* Dump one node's bpf-vs-file stats; called when a mismatch is detected. = */ +static void dump_node(int i, const struct memcg_stat_snapshot *b, + const struct file_snap *f) +{ + ksft_print_msg("node %d bpf : anon=3D%llu file=3D%llu shmem=3D%llu fmappe= d=3D%llu pgfault=3D%llu\n", + i, b->anon, b->file, b->shmem, b->file_mapped, b->pgfault); + ksft_print_msg("node %d file: anon=3D%llu file=3D%llu shmem=3D%llu fmappe= d=3D%llu pgfault=3D%llu\n", + i, f->anon, f->file, f->shmem, f->file_mapped, f->pgfault); +} + +/* + * Compare the BPF kfunc snapshots (round 1) against the memory.stat values + * (round 2), node by node. The two rounds are independent, equivalently + * charged trees, so the flushed stats are compared within a small toleran= ce + * that absorbs per-round overhead (i.e., a charging child's own stack pag= es). + * A wrong unit, enum or field in the kfunc path would miss by far more. + * + * Two per-round checks verify the flush itself: each leaf was charged + * resident_bytes of anon spread over K CPUs, so its flushed anon must be = at + * least that; and the root's recursive anon must equal the sum of the lea= ves' + * anon (rstat propagated the charge up the tree). + * + * Returns 0 if every check passes, -1 otherwise. + */ +static int check_correctness(const struct memcg_stat_snapshot *bpf, + const struct file_snap *file, const bool *is_leaf, + int n, size_t resident_bytes) +{ + __u64 stat_tol =3D 64 * page_size; + __u64 pgf_tol =3D 1024; + __u64 broot =3D 0, bsum =3D 0, froot =3D 0, fsum =3D 0; + int i, mism =3D 0, flush_bad =3D 0; + + for (i =3D 0; i < n; i++) { + const struct memcg_stat_snapshot *b =3D &bpf[i]; + const struct file_snap *f =3D &file[i]; + __u64 bcur =3D b->usage_pages * page_size; + + /* kfunc path (round 1) compared with memory.stat path (round 2) */ + if (!close_enough(b->anon, f->anon, stat_tol) || + !close_enough(b->file, f->file, stat_tol) || + !close_enough(b->shmem, f->shmem, stat_tol) || + !close_enough(b->file_mapped, f->file_mapped, stat_tol) || + !close_enough(b->pgfault, f->pgfault, pgf_tol)) { + mism++; + dump_node(i, b, f); + } + + /* memory.current is live (no flush) and must bound the flushed anon */ + if (b->anon =3D=3D 0 || b->anon > bcur) { + flush_bad++; + ksft_print_msg("node %d: anon=3D%llu exceeds current=3D%llu\n", + i, b->anon, bcur); + } + + /* each leaf's flush must have gathered the full cross-CPU charge */ + if (is_leaf[i]) { + if (b->anon < resident_bytes || f->anon < resident_bytes) { + flush_bad++; + ksft_print_msg("node %d: short flush bpf=3D%llu file=3D%llu\n", + i, b->anon, f->anon); + } + bsum +=3D b->anon; + fsum +=3D f->anon; + } + if (i =3D=3D 0) { /* nodes[0] =3D=3D subtree_root */ + broot =3D b->anon; + froot =3D f->anon; + } + } + + if (mism) { + ksft_print_msg("bpf (round 1) disagrees with memory.stat (round 2)\n"); + return -1; + } + if (flush_bad) { + ksft_print_msg("flush did not aggregate the cross-cpu charge\n"); + return -1; + } + if (broot !=3D bsum || froot !=3D fsum) { + ksft_print_msg("root anon !=3D sum of leaf anon: bpf %llu/%llu file %llu= /%llu\n", + broot, bsum, froot, fsum); + return -1; + } + if (bsum =3D=3D 0) { + ksft_print_msg("tree carries no anon\n"); + return -1; + } + return 0; +} + +/* ---- one case ---------------------------------------------------------= - */ + +struct testcase { + const char *name; + int fanout; + int depth; + int cpus_per_leaf; /* K: CPUs each leaf is charged on; 0 =3D all CPUs */ + size_t resident_bytes; /* anon charged per leaf */ +}; + +/* + * Remove the subtree in reverse creation order. Nodes are recorded in DFS + * order (a parent precedes all its descendants), so iterating backwards + * removes every child before its parent. + */ +static void destroy_tree(void) +{ + int i; + + if (!nodes) + return; + for (i =3D n_nodes - 1; i >=3D 0; i--) + cg_destroy(nodes[i].path); + free(nodes); + nodes =3D NULL; +} + +/* + * Round 1: build and charge a fresh tree, walk it with the BPF iterator (= which + * flushes and reads each cgroup via the memcg kfuncs), and capture one sn= apshot + * per node into @snap. @is_leaf records the tree shape so the later comp= arison + * can run after the tree is gone. Returns the node count, or -1 on failu= re. + * The tree is always torn down before returning. + */ +static int capture_bpf_round(const struct testcase *tc, + struct memcg_stat_snapshot *snap, bool *is_leaf) +{ + struct memcg_stat_cross_cpu *skel =3D NULL; + struct bpf_link *link =3D NULL; + int root_fd =3D -1, ret =3D -1, i, mfd; + + if (build_tree(tc->fanout, tc->depth, &root_fd)) { + ksft_print_msg("build tree (bpf) failed\n"); + goto out; + } + if (start_chargers(tc->cpus_per_leaf, tc->resident_bytes)) + goto out; + + skel =3D memcg_stat_cross_cpu__open(); + if (!skel) { + ksft_print_msg("skel open failed\n"); + goto out; + } + if (bpf_map__set_max_entries(skel->maps.results, n_nodes + 8)) { + ksft_print_msg("set max_entries failed\n"); + goto out; + } + if (memcg_stat_cross_cpu__load(skel)) { + ksft_print_msg("skel load failed\n"); + goto out; + } + + DECLARE_LIBBPF_OPTS(bpf_iter_attach_opts, opts); + union bpf_iter_link_info linfo =3D {}; + + linfo.cgroup.cgroup_fd =3D root_fd; + linfo.cgroup.order =3D BPF_CGROUP_ITER_DESCENDANTS_PRE; + opts.link_info =3D &linfo; + opts.link_info_len =3D sizeof(linfo); + + link =3D bpf_program__attach_iter(skel->progs.cgroup_memcg_stat_cross_cpu, + &opts); + if (!link) { + ksft_print_msg("attach iter failed\n"); + goto out; + } + + /* bpf walk through the cgroup tree and fetch result */ + if (bpf_walk_once(link)) { + ksft_print_msg("bpf walk failed\n"); + goto out; + } + + mfd =3D bpf_map__fd(skel->maps.results); + for (i =3D 0; i < n_nodes; i++) { + /* Save the position for the leaf, used for later correctness check */ + is_leaf[i] =3D nodes[i].is_leaf; + if (bpf_map_lookup_elem(mfd, &nodes[i].id, &snap[i])) { + ksft_print_msg("map lookup failed for node %d\n", i); + goto out; + } + } + ret =3D n_nodes; +out: + bpf_link__destroy(link); + memcg_stat_cross_cpu__destroy(skel); + if (root_fd >=3D 0) + close(root_fd); + stop_chargers(); + destroy_tree(); + return ret; +} + +/* + * Round 2: build and charge an identical fresh tree, then read every cgro= up via + * memory.stat / memory.current. The tree is reset after bpf read, so each + * read does a real rstat flush. Returns the node count, or -1 on failure. + * The tree is always clear before returning. + */ +static int capture_file_round(const struct testcase *tc, struct file_snap = *snap) +{ + int root_fd =3D -1, ret =3D -1, i; + + if (build_tree(tc->fanout, tc->depth, &root_fd)) { + ksft_print_msg("build tree (file) failed\n"); + goto out; + } + if (start_chargers(tc->cpus_per_leaf, tc->resident_bytes)) + goto out; + + for (i =3D 0; i < n_nodes; i++) + if (file_read_node(nodes[i].path, &snap[i])) { + ksft_print_msg("file read failed for node %d\n", i); + goto out; + } + ret =3D n_nodes; +out: + if (root_fd >=3D 0) + close(root_fd); + stop_chargers(); + destroy_tree(); + return ret; +} + +static int run_case(const struct testcase *tc) +{ + struct memcg_stat_snapshot *bpf =3D NULL; + struct file_snap *file =3D NULL; + bool *is_leaf =3D NULL; + size_t cap =3D tree_capacity(tc->fanout, tc->depth); + int nb, nf, ret =3D KSFT_FAIL; + + bpf =3D calloc(cap, sizeof(*bpf)); + file =3D calloc(cap, sizeof(*file)); + is_leaf =3D calloc(cap, sizeof(*is_leaf)); + if (!bpf || !file || !is_leaf) { + ksft_print_msg("calloc failed\n"); + goto out; + } + + ksft_print_msg("%s: fanout=3D%d depth=3D%d cpus=3D%d/%d resident=3D%zuKB/= leaf\n", + tc->name, tc->fanout, tc->depth, n_cpu, tc->cpus_per_leaf, + tc->resident_bytes >> 10); + + /* round 1: BPF reader flushes and reads its own charged tree */ + nb =3D capture_bpf_round(tc, bpf, is_leaf); + if (nb < 0) + goto out; + + /* round 2: file reader flushes and reads a fresh, equivalent tree */ + nf =3D capture_file_round(tc, file); + if (nf < 0) + goto out; + + if (nb !=3D nf) { + ksft_print_msg("node count differs between rounds: %d vs %d\n", + nb, nf); + goto out; + } + + if (!check_correctness(bpf, file, is_leaf, nb, tc->resident_bytes)) + ret =3D KSFT_PASS; +out: + free(bpf); + free(file); + free(is_leaf); + return ret; +} + +static const struct testcase cases[] =3D { + /* name, fan, depth, K, resident anon */ + { "single_cpu_small_tree", 4, 2, 1, 2 << 20 }, + { "cross_cpu_small_tree", 4, 2, 0, 2 << 20 }, + { "single_cpu_large_tree", 10, 3, 1, 256 << 10 }, + { "cross_cpu_large_tree", 10, 3, 0, 256 << 10 }, +}; + +static bool memcg_kfuncs_available(void) +{ + struct btf *btf; + bool ok; + + btf =3D btf__load_vmlinux_btf(); + if (!btf) + return false; + ok =3D btf__find_by_name_kind(btf, "bpf_get_mem_cgroup", BTF_KIND_FUNC) >= 0; + btf__free(btf); + return ok; +} + +int main(int argc, char **argv) +{ + int i; + + ksft_print_header(); + ksft_set_plan(ARRAY_SIZE(cases)); + + /* Feature gate first: a read-only BTF probe, no privilege needed. */ + if (!memcg_kfuncs_available()) + ksft_exit_skip("memcg BPF kfuncs are not available\n"); + + if (cg_find_unified_root(root, sizeof(root), NULL)) + ksft_exit_skip("cgroup v2 isn't mounted\n"); + + if (cg_read_strstr(root, "cgroup.controllers", "memory")) + ksft_exit_skip("memory controller isn't available\n"); + + if (cg_read_strstr(root, "cgroup.subtree_control", "memory")) + if (cg_write(root, "cgroup.subtree_control", "+memory")) + ksft_exit_skip("Failed to set memory controller\n"); + + if (collect_cpus()) + ksft_exit_skip("cannot read CPU affinity\n"); + + page_size =3D sysconf(_SC_PAGESIZE); + subtree_root =3D cg_name(root, SUBTREE_NAME); + if (!subtree_root) + ksft_exit_skip("cannot build subtree root path\n"); + + for (i =3D 0; i < ARRAY_SIZE(cases); i++) { + switch (run_case(&cases[i])) { + case KSFT_PASS: + ksft_test_result_pass("%s\n", cases[i].name); + break; + case KSFT_SKIP: + ksft_test_result_skip("%s\n", cases[i].name); + break; + default: + ksft_test_result_fail("%s\n", cases[i].name); + break; + } + } + + free(cpu_list); + cpu_list =3D NULL; + n_cpu =3D 0; + + ksft_finished(); +} --=20 2.53.0-Meta From nobody Sat Jul 25 00:11:12 2026 Received: from mail-pj2-f4.google.com (mail-pj2-f4.google.com [74.125.227.132]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3E2D0415F21 for ; Tue, 21 Jul 2026 17:48:43 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.227.132 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784656127; cv=none; b=sSIBfD4if+YWvkphRCytD+hMAJhEn9cV4X9xHZlYKXOtTBjAIQuON7U5FBd1F+IiCCclMiyDYpB2S+2ym14suyWSfR48XO+TfU+UQCzsFWpb7qU6tfJVMm7RsuZTmtbvDAruULLQW2XS9L6KR4lnXhs0ZuZXzB+J6iLJDvtZDvY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784656127; c=relaxed/simple; bh=sRoAOqve++B63kDPpInmDmu+xh4/uy82qPpeQLmoTZw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=dU995DfgtGLzCuUW0k5q5sONvyGwnHoyDC4k3c6V2IIHQWk4BnDU2eFeGIT+9EepvPGAd//PeH4nWofsSEy8ufbT/5cJg6BoDpcOF6P5tOehrbFz2gjr0uTNiMuoup1HVakYu0QrmEKPmGeoQ22YaCXKfJQViDTxNu0CVGSZpf8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=UtUM5tXN; arc=none smtp.client-ip=74.125.227.132 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="UtUM5tXN" Received: by mail-pj2-f4.google.com with SMTP id 98e67ed59e1d1-38111ea8a88so7789775a91.1 for ; Tue, 21 Jul 2026 10:48:43 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784656123; x=1785260923; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=jApOwXpBhXqEU1tu6BspXbjxNN8qabGRryStwscMjoo=; b=UtUM5tXNxah0MeaB2u4x9Z/b6ueSU8GdalGAbwyRqFFnrNMo43HQAGLjmb2DRAs9UV FdYMo75b95znPf4jVK2MxidJyVilzE0xALTBe1xQmCfqHEAxJZb//O0FBK5O/vBlmxd8 qziNSexSv8AJ4EMEg3D151EarB/I6EWyxvy6RYaZpE3WG9OV8gO2fY767ofBf28YfF3z bSnlgbLJYkHCOOR6+15ue9Ds3py+N6IpPOZnY9dUTLpBxyT9+pd4wJrl+tq4wsQkrmHU lTXl9iSrZi1qAreQtr+Xf6hPUmes9oA593ITV2ziqt+9e/VegmlXr6ummtq3P5P7C1aW L22w== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784656123; x=1785260923; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=jApOwXpBhXqEU1tu6BspXbjxNN8qabGRryStwscMjoo=; b=JshotMVF1FGGjdoZFL9v1KxJcl9CCRPkA02YKUZ2G4z30j6n0C9boKwXWtL5cAle33 0lANOyHEgxl1tCIV+RKRuYZWeApkYYdQJwUuz6hbVrz5SzAL/9L9MbVE9Jt2bfQNXE/s rAVOCZ/EnGPyHIU+CizgT9mdUYYDYzF1Xsioyg2i2xgPxUqNlnX3s8E9vev/8eF1u2RX 8jQbtd6GIjjqPmMzmqA6lRZcN5328esIqAw44LExHxvZr5/o1DPOQ2wrPg4n6xy26qp6 N/QNYevmftPYSBkv1UCxifivOULj4/7d65Fb533CjjsMp+rIygeCVQC0489IuqRREiQN xBKg== X-Forwarded-Encrypted: i=1; AHgh+Row5vuMKVPMG05oZ6Mg9emiBNI/d4j1YZkCyuByzAWwBC7Sl0yj3WXT+C8GDhuO2UbPlr8m5sGpHHfCOxE=@vger.kernel.org X-Gm-Message-State: AOJu0YxphTJ0bfP8YbjL6BCP6fFzesMzRj1wuwJpR+8EDkLisU+Lt0ld VVOtUQP6tKKwBjS0Pr5tDX2zXukWsGiuSJkKsiZWQKsiaI7x2f5E6dlc X-Gm-Gg: AR+sD13Ibk9Koh4ozOiqOpnfD4tYwjw4/hEPy+oageY+I9fUa8ILu8IfwJJy5v5D31R qmzH7fOJ6lyj4U6wZkyPjnMfVS+cBb1g8fzfHyruVcCXWRUzUgNKoRoSHyDH0vn2CKjLJsW3vQR fJeKK8Ntr6M+fDyOx0hx2pzrwr0+jpxKULr46RG6XcN/3Br60DwCvbGT04MGlhwJMR5RhbzuUbL hFXMepOzKcCt6wSRpv+wZp4GPPwj/TWag4cU7I3JF+eS0Tfo6UzDnDcfplOwNOFFtGgoOorUuXy iihmpVzjCXW7U+37H3cVnpdA575kbf7hrHU6ZbwDdZeiY/wQn6dAR3YqOOF+Pli2DDwK4ZgsSNf heTI4ozK23CCHFGm2MsCPk/BGUCV8JyjTvReWzn3aLl7wCVcAtwBsJVnIDH4W/58/3f0CmEwj X-Received: by 2002:a05:6a21:a393:b0:3c3:9251:fe49 with SMTP id adf61e73a8af0-3c3ad972d5bmr20981964637.50.1784656123174; Tue, 21 Jul 2026 10:48:43 -0700 (PDT) Received: from localhost ([2a03:2880:9ff:66::]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-3147dc193f8sm2331163eec.1.2026.07.21.10.48.42 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 21 Jul 2026 10:48:42 -0700 (PDT) From: Ziyang Men To: Shuah Khan , Tejun Heo , Johannes Weiner , =?UTF-8?q?Michal=20Koutn=C3=BD?= , Jiri Kosina , Benjamin Tissoires , David Vernet , Eduard Zingerman Cc: Andrea Righi , Changwoo Min , Michal Hocko , Roman Gushchin , Shakeel Butt , Muchun Song , Andrew Morton , JP Kobryn , Mykola Lysenko , Nathan Chancellor , linux-kselftest@vger.kernel.org, cgroups@vger.kernel.org, linux-input@vger.kernel.org, sched-ext@lists.linux.dev, linux-mm@kvack.org, kernel-team@meta.com, bpf@vger.kernel.org, llvm@lists.linux.dev, linux-kernel@vger.kernel.org, Ziyang Men Subject: [PATCH v2 3/4] selftests/hid: build the BPF program via the shared lib.bpf.mk Date: Tue, 21 Jul 2026 10:48:32 -0700 Message-ID: <20260721174833.1232771-4-ziyang.meme@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260721174833.1232771-1-ziyang.meme@gmail.com> References: <20260721174833.1232771-1-ziyang.meme@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" hid carries its own ~150 lines of libbpf + bpftool + vmlinux.h + BPF-object + skeleton build machinery, copied from selftests/bpf. Replace it with an include of the shared tools/testing/selftests/lib.bpf.mk, so that the previous ~150 lines of BPF build configuration can now be achieved in only ~10 lines. hid keeps the legacy progs/.c layout, so it sets BPF_PROG_EXT :=3D .c and passes its shared BPF headers through BPF_EXTRA_HDRS. It also gains the fragment's -Wall on the BPF compile, which its old Makefile did not set; hid.c has a few pre-existing unused locals, so BPF_EXTRA_CFLAGS :=3D -Wno-unused-variable keeps the previous behaviour without touching the program source. The conversion drops two pieces of dead machinery hid had copied from selftests/bpf: a resolve_btfids build that was never a prerequisite of any target, and a CLANG_CFLAGS variable that was never passed to any compile. No functional change: the generated hid.skel.h public API is byte-identical before and after, hid_bpf and hidraw are built and linked the same way, and the folder builds cleanly under both plain make and LLVM=3D1. Suggested-by: Shakeel Butt Suggested-by: Eduard Zingerman Suggested-by: Mykola Lysenko Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Ziyang Men --- tools/testing/selftests/hid/Makefile | 182 +++------------------------ 1 file changed, 20 insertions(+), 162 deletions(-) diff --git a/tools/testing/selftests/hid/Makefile b/tools/testing/selftests= /hid/Makefile index 2f423de83147..48def484703c 100644 --- a/tools/testing/selftests/hid/Makefile +++ b/tools/testing/selftests/hid/Makefile @@ -58,172 +58,31 @@ override define CLEAN $(Q)$(RM) -r $(EXTRA_CLEAN) endef =20 -include ../lib.mk +# The BPF program (progs/hid.c) is compiled into a skeleton by the shared +# tools/testing/selftests/lib.bpf.mk fragment. hid keeps the legacy +# progs/.c layout, so point BPF_PROG_EXT at .c and hand the fragment +# hid's shared BPF headers as extra prerequisites. BPF_EXTRA_HDRS is defe= rred +# because it references $(BPFDIR), which lib.bpf.mk defines. +BPF_SRCS :=3D progs/hid.c +BPF_PROG_EXT :=3D .c +BPF_EXTRA_HDRS =3D $(wildcard progs/*.h) $(wildcard $(BPFDIR)/hid_bpf_*.h= ) \ + $(wildcard $(BPFDIR)/*.bpf.h) +# hid.c predates the shared -Wall (hid's old Makefile didn't set it) and h= as a +# few unused locals; tolerate that pre-existing warning without touching t= he +# program source. +BPF_EXTRA_CFLAGS :=3D -Wno-unused-variable =20 -TOOLSDIR :=3D $(top_srcdir)/tools -LIBDIR :=3D $(TOOLSDIR)/lib -BPFDIR :=3D $(LIBDIR)/bpf -TOOLSINCDIR :=3D $(TOOLSDIR)/include -BPFTOOLDIR :=3D $(TOOLSDIR)/bpf/bpftool -SCRATCH_DIR :=3D $(OUTPUT)/tools -BUILD_DIR :=3D $(SCRATCH_DIR)/build -INCLUDE_DIR :=3D $(SCRATCH_DIR)/include -BPFOBJ :=3D $(BUILD_DIR)/libbpf/libbpf.a -ifneq ($(CROSS_COMPILE),) -HOST_BUILD_DIR :=3D $(BUILD_DIR)/host -HOST_SCRATCH_DIR :=3D $(OUTPUT)/host-tools -HOST_INCLUDE_DIR :=3D $(HOST_SCRATCH_DIR)/include -else -HOST_BUILD_DIR :=3D $(BUILD_DIR) -HOST_SCRATCH_DIR :=3D $(SCRATCH_DIR) -HOST_INCLUDE_DIR :=3D $(INCLUDE_DIR) -endif -HOST_BPFOBJ :=3D $(HOST_BUILD_DIR)/libbpf/libbpf.a -RESOLVE_BTFIDS :=3D $(HOST_BUILD_DIR)/resolve_btfids/resolve_btfids - -VMLINUX_BTF_PATHS ?=3D $(if $(O),$(O)/vmlinux) \ - $(if $(KBUILD_OUTPUT),$(KBUILD_OUTPUT)/vmlinux) \ - ../../../../vmlinux \ - /sys/kernel/btf/vmlinux \ - /boot/vmlinux-$(shell uname -r) -VMLINUX_BTF ?=3D $(abspath $(firstword $(wildcard $(VMLINUX_BTF_PATHS)))) -ifeq ($(VMLINUX_BTF),) -$(error Cannot find a vmlinux for VMLINUX_BTF at any of "$(VMLINUX_BTF_PAT= HS)") -endif +include ../lib.mk +include ../lib.bpf.mk =20 -# Define simple and short `make test_progs`, `make test_sysctl`, etc targe= ts -# to build individual tests. +# Define simple and short `make hid_bpf`, `make hidraw` targets. # NOTE: Semicolon at the end is critical to override lib.mk's default stat= ic # rule for binaries. $(notdir $(TEST_GEN_PROGS)): %: $(OUTPUT)/% ; =20 -# sort removes libbpf duplicates when not cross-building -MAKE_DIRS :=3D $(sort $(BUILD_DIR)/libbpf $(HOST_BUILD_DIR)/libbpf \ - $(HOST_BUILD_DIR)/bpftool $(HOST_BUILD_DIR)/resolve_btfids \ - $(INCLUDE_DIR)) -$(MAKE_DIRS): - $(call msg,MKDIR,,$@) - $(Q)mkdir -p $@ - -DEFAULT_BPFTOOL :=3D $(HOST_SCRATCH_DIR)/sbin/bpftool - -TEST_GEN_PROGS_EXTENDED +=3D $(DEFAULT_BPFTOOL) - -$(TEST_GEN_PROGS) $(TEST_GEN_PROGS_EXTENDED): $(BPFOBJ) - -BPFTOOL ?=3D $(DEFAULT_BPFTOOL) -$(DEFAULT_BPFTOOL): $(wildcard $(BPFTOOLDIR)/*.[ch] $(BPFTOOLDIR)/Makefile= ) \ - $(HOST_BPFOBJ) | $(HOST_BUILD_DIR)/bpftool - $(Q)$(MAKE) $(submake_extras) -C $(BPFTOOLDIR) \ - ARCH=3D CROSS_COMPILE=3D CC=3D$(HOSTCC) LD=3D$(HOSTLD) \ - EXTRA_CFLAGS=3D'-g -O0' \ - OUTPUT=3D$(HOST_BUILD_DIR)/bpftool/ \ - LIBBPF_OUTPUT=3D$(HOST_BUILD_DIR)/libbpf/ \ - LIBBPF_DESTDIR=3D$(HOST_SCRATCH_DIR)/ \ - prefix=3D DESTDIR=3D$(HOST_SCRATCH_DIR)/ install-bin - -$(BPFOBJ): $(wildcard $(BPFDIR)/*.[ch] $(BPFDIR)/Makefile) \ - | $(BUILD_DIR)/libbpf - $(Q)$(MAKE) $(submake_extras) -C $(BPFDIR) OUTPUT=3D$(BUILD_DIR)/libbpf/ \ - EXTRA_CFLAGS=3D'-g -O0' \ - DESTDIR=3D$(SCRATCH_DIR) prefix=3D all install_headers - -ifneq ($(BPFOBJ),$(HOST_BPFOBJ)) -$(HOST_BPFOBJ): $(wildcard $(BPFDIR)/*.[ch] $(BPFDIR)/Makefile) \ - | $(HOST_BUILD_DIR)/libbpf - $(Q)$(MAKE) $(submake_extras) -C $(BPFDIR) \ - EXTRA_CFLAGS=3D'-g -O0' ARCH=3D CROSS_COMPILE=3D \ - OUTPUT=3D$(HOST_BUILD_DIR)/libbpf/ CC=3D$(HOSTCC) LD=3D$(HOSTLD) \ - DESTDIR=3D$(HOST_SCRATCH_DIR)/ prefix=3D all install_headers -endif - -$(INCLUDE_DIR)/vmlinux.h: $(VMLINUX_BTF) $(BPFTOOL) | $(INCLUDE_DIR) -ifeq ($(VMLINUX_H),) - $(call msg,GEN,,$@) - $(Q)$(BPFTOOL) btf dump file $(VMLINUX_BTF) format c > $@ -else - $(call msg,CP,,$@) - $(Q)cp "$(VMLINUX_H)" $@ -endif - -$(RESOLVE_BTFIDS): $(HOST_BPFOBJ) | $(HOST_BUILD_DIR)/resolve_btfids \ - $(TOOLSDIR)/bpf/resolve_btfids/main.c \ - $(TOOLSDIR)/lib/rbtree.c \ - $(TOOLSDIR)/lib/zalloc.c \ - $(TOOLSDIR)/lib/string.c \ - $(TOOLSDIR)/lib/ctype.c \ - $(TOOLSDIR)/lib/str_error_r.c - $(Q)$(MAKE) $(submake_extras) -C $(TOOLSDIR)/bpf/resolve_btfids \ - CC=3D$(HOSTCC) LD=3D$(HOSTLD) AR=3D$(HOSTAR) \ - LIBBPF_INCLUDE=3D$(HOST_INCLUDE_DIR) \ - OUTPUT=3D$(HOST_BUILD_DIR)/resolve_btfids/ BPFOBJ=3D$(HOST_BPFOBJ) - -# Get Clang's default includes on this system, as opposed to those seen by -# '--target=3Dbpf'. This fixes "missing" files on some architectures/distr= os, -# such as asm/byteorder.h, asm/socket.h, asm/sockios.h, sys/cdefs.h etc. -# -# Use '-idirafter': Don't interfere with include mechanics except where the -# build would have failed anyways. -define get_sys_includes -$(shell $(1) -v -E - &1 \ - | sed -n '/<...> search starts here:/,/End of search list./{ s| \(/.*\)|-= idirafter \1|p }') \ -$(shell $(1) -dM -E - $@ +# Each test binary links against the in-tree static libbpf that lib.bpf.mk +# built (lib.mk has already prefixed TEST_GEN_PROGS with $(OUTPUT)/). +$(TEST_GEN_PROGS): $(BPFOBJ) =20 $(OUTPUT)/%.o: %.c $(BPF_SKELS) hid_common.h $(call msg,CC,,$@) @@ -233,5 +92,4 @@ $(OUTPUT)/%: $(OUTPUT)/%.o $(call msg,BINARY,,$@) $(Q)$(LINK.c) $^ $(LDLIBS) -o $@ =20 -EXTRA_CLEAN :=3D $(SCRATCH_DIR) $(HOST_SCRATCH_DIR) feature bpftool \ - $(addprefix $(OUTPUT)/,*.o *.skel.h no_alu32) +EXTRA_CLEAN +=3D feature bpftool $(addprefix $(OUTPUT)/,*.o no_alu32) --=20 2.53.0-Meta From nobody Sat Jul 25 00:11:12 2026 Received: from mail-pz2-f1.google.com (mail-pz2-f1.google.com [74.125.228.1]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 92D66415F34 for ; Tue, 21 Jul 2026 17:48:46 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.228.1 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784656130; cv=none; b=d0QQr/K4krff3NVbxbexsttxq7k3iPWrlQxjt+CNAVxYe/kPiBjpn9FZ5WDQXwH3hA1U1vA/cXkAWKnIObmKF346mCTNHbNLy/WDeDaGcNhXjPnMdoTQ15a9i6TCUzwDla3+Q2CrEM7Jai2HHHA3T3v2YESkCCf5uiALwL0JVIM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784656130; c=relaxed/simple; bh=+fRWz032Ue1fMDnjt8nBJt0lr8qqkpM0+XI+1mNdwjA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=XJjnsLPW6DvWyhjzsVULmqmyluME3mYRRVE9RbA3jJRSbC9q/E8oDBqdFl4GJMaEEyU2Eyub8N9pTgANVOvSM2+7PaZc9UFmSgsHz91TMB9PFibaNrE/awWgikh9Jw17HzFoQZ3FsXD55vWJga35P+bfOgXg28JA6+jqQpELuO4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=nJ2OqJln; arc=none smtp.client-ip=74.125.228.1 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="nJ2OqJln" Received: by mail-pz2-f1.google.com with SMTP id 41be03b00d2f7-ca00ea47337so3324740a12.0 for ; Tue, 21 Jul 2026 10:48:46 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784656126; x=1785260926; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=V1HVbnWwjz/zPGwxKegdY4p1KeImHVO8vqQ7kQJ99fQ=; b=nJ2OqJlnR05TlUwDW37eUv3kZb69nltnK2CXWUrV3Rhx+PzAqOQhiuBayNxjj56G2P 31FW+VRsJSTsvoxBhEO3mYDTZZcbNsq/reNj3OYaukN1hkCqIK2H3UIrsUhSM1mN0vGG 2uFeV7VmGNROOfqbV6HXLo0JCzS1bqPbnIiZXss3sgAgAsZNVCL9SGN2REZjlZjqXSHx D89uz2F+OS3Sw5xEkkpMy8hSAcm2P//09sagAcoqNaeNW+7Nmmlkj/E9yq1xSe17LX/N bN5C5fm7NcM19qEn1Zlb+WkJTv4jDrgUkPuab600Plsab+Crd/lq+9ou+0K7JOfE2nQA 8Fwg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784656126; x=1785260926; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=V1HVbnWwjz/zPGwxKegdY4p1KeImHVO8vqQ7kQJ99fQ=; b=qWs8InCSbTqXUNUv5LclCGkYJ/Dp2ArguZW7QoSOEnkH4HxEhgFX9bN3+l/mz1G9zF v/0xFBB7TVpGo3DSzrWfIBaczabxTx1YPupu5/sE4AJr5toU6T3Zdof0iKkl2Svbv4TL JM4tl08cnalw44BfkwNhRKiQlN9rqnTTbQIR+U8OK5j/yDLeQaSio/+Vonoz3tbJgC7x s6CUrb0rCXgUikSiMbAUVmqjqkCve8CL/9CEH3CN5SXzcwGlsZUhadFvutGtt+7Bl7z4 4N+oFX3uRda/x5/J1KzUETebwi6s7my60evZjv/72QZwFx0hpb8tbfKbcAJydXZeG7kp H0Dw== X-Forwarded-Encrypted: i=1; AHgh+RoCkaLBv6EAk/FHxVhAJ5exIv/TMVAQXl2MlKCmwdl8s7ppAqUtsgbZGhY4ZLaa91eLNjQ6uP751KQTi+g=@vger.kernel.org X-Gm-Message-State: AOJu0YypEon70c9CzVGDpeFvxeIxHG6QetIo/hPAO0vszLRlXDcLMwtS hDX/BcuvSaSRTDmKRPokfsl8cTeBH5nYacvbyzteQyZsTx+PDp4Y8hqp X-Gm-Gg: AR+sD10ynV7MH/eRR+3M8+jk+adlG/L9CwlvVcvlimM3h/fHqBIXg9pUqnFECknxW20 VOG2MwtwaUuqrCu3RkCjULWlz5d/jy02aRqON4Hm7g8y15rtn9IocDU/mAd8PvI5w0bDwedW6di H3LZ4G0M6G4rnqsqxqMb0WVqimWcsz+Tg0EhR3LyALKAiU3nMaLtV46t9R91HPFaMYB4Y9Qj8ir +izghqhzD6jfnTGdQiiQVzTROGDBHWnorWPedzY5iErDI+4ROzkFt8wfx7s7n2ZhCvCjNMZRxmS wZ6cW9xZpMF4BH0BSR+w/DuVywbZdhEJN+7iiQTLNcY/AWSwSN7gPmYJR53IdBcTB28rkZFloD7 J1JizG+WhkrQI/+wkYDJqVKG8tcoQ94Fkoa63sg601kF7+xgpBL3x2W8yFBoP96cHPcO+8hiIUx kSZQbUR8U= X-Received: by 2002:a05:6a21:114c:b0:3bf:a489:1483 with SMTP id adf61e73a8af0-3c3ad80c23dmr20107991637.33.1784656125438; Tue, 21 Jul 2026 10:48:45 -0700 (PDT) Received: from localhost ([2a03:2880:9ff:65::]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-3147e1b6a45sm872370eec.28.2026.07.21.10.48.44 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 21 Jul 2026 10:48:44 -0700 (PDT) From: Ziyang Men To: Shuah Khan , Tejun Heo , Johannes Weiner , =?UTF-8?q?Michal=20Koutn=C3=BD?= , Jiri Kosina , Benjamin Tissoires , David Vernet , Eduard Zingerman Cc: Andrea Righi , Changwoo Min , Michal Hocko , Roman Gushchin , Shakeel Butt , Muchun Song , Andrew Morton , JP Kobryn , Mykola Lysenko , Nathan Chancellor , linux-kselftest@vger.kernel.org, cgroups@vger.kernel.org, linux-input@vger.kernel.org, sched-ext@lists.linux.dev, linux-mm@kvack.org, kernel-team@meta.com, bpf@vger.kernel.org, llvm@lists.linux.dev, linux-kernel@vger.kernel.org, Ziyang Men Subject: [PATCH v2 4/4] selftests/sched_ext: build BPF schedulers via the shared lib.bpf.mk Date: Tue, 21 Jul 2026 10:48:33 -0700 Message-ID: <20260721174833.1232771-5-ziyang.meme@gmail.com> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260721174833.1232771-1-ziyang.meme@gmail.com> References: <20260721174833.1232771-1-ziyang.meme@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" sched_ext carried its own ~130 lines of libbpf + bpftool + vmlinux.h + BPF-object + skeleton build machinery. Replace it with an include of the shared tools/testing/selftests/lib.bpf.mk, making sched_ext the third in-tree consumer of that fragment, after selftests/cgroup and selftests/hid. sched_ext emits a skeleton and a subskeleton per scheduler with a .bpf.skel.h suffix, so it sets BPF_SKEL_EXT :=3D .bpf.skel.h and BPF_GEN_SUBSKEL :=3D 1. Its BPF programs need scheduler-specific include paths and are not -Werror clean, so it overrides BPF_CFLAGS wholesale after the include; the userspace runner/testcase rules are kept and repointed at the fragment's flat $(OUTPUT) skeletons. No functional change: all 28 skeletons and 28 subskeletons keep a byte-identical public API before and after, the runner builds and links the same way, and the folder builds cleanly under both plain make and LLVM=3D1. Suggested-by: Shakeel Butt Suggested-by: Eduard Zingerman Suggested-by: Mykola Lysenko Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Ziyang Men --- tools/testing/selftests/sched_ext/Makefile | 153 +++++---------------- 1 file changed, 35 insertions(+), 118 deletions(-) diff --git a/tools/testing/selftests/sched_ext/Makefile b/tools/testing/sel= ftests/sched_ext/Makefile index 5d2dffca0e91..67ff96bf34e8 100644 --- a/tools/testing/selftests/sched_ext/Makefile +++ b/tools/testing/selftests/sched_ext/Makefile @@ -14,48 +14,36 @@ CURDIR :=3D $(abspath .) REPOROOT :=3D $(abspath ../../../..) TOOLSDIR :=3D $(REPOROOT)/tools LIBDIR :=3D $(TOOLSDIR)/lib -BPFDIR :=3D $(LIBDIR)/bpf TOOLSINCDIR :=3D $(TOOLSDIR)/include -BPFTOOLDIR :=3D $(TOOLSDIR)/bpf/bpftool APIDIR :=3D $(TOOLSINCDIR)/uapi GENDIR :=3D $(REPOROOT)/include/generated GENHDR :=3D $(GENDIR)/autoconf.h -SCXTOOLSDIR :=3D $(TOOLSDIR)/sched_ext SCXTOOLSINCDIR :=3D $(TOOLSDIR)/sched_ext/include =20 -OUTPUT_DIR :=3D $(OUTPUT)/build -OBJ_DIR :=3D $(OUTPUT_DIR)/obj -INCLUDE_DIR :=3D $(OUTPUT_DIR)/include -BPFOBJ_DIR :=3D $(OBJ_DIR)/libbpf -SCXOBJ_DIR :=3D $(OBJ_DIR)/sched_ext -BPFOBJ :=3D $(BPFOBJ_DIR)/libbpf.a -LIBBPF_OUTPUT :=3D $(OBJ_DIR)/libbpf/libbpf.a - -DEFAULT_BPFTOOL :=3D $(OUTPUT_DIR)/host/sbin/bpftool -HOST_OBJ_DIR :=3D $(OBJ_DIR)/host/bpftool -HOST_LIBBPF_OUTPUT :=3D $(OBJ_DIR)/host/libbpf/ -HOST_LIBBPF_DESTDIR :=3D $(OUTPUT_DIR)/host/ -HOST_DESTDIR :=3D $(OUTPUT_DIR)/host/ - -VMLINUX_BTF_PATHS ?=3D $(if $(O),$(O)/vmlinux) \ - $(if $(KBUILD_OUTPUT),$(KBUILD_OUTPUT)/vmlinux) \ - ../../../../vmlinux \ - /sys/kernel/btf/vmlinux \ - /boot/vmlinux-$(shell uname -r) -VMLINUX_BTF ?=3D $(abspath $(firstword $(wildcard $(VMLINUX_BTF_PATHS)))) -ifeq ($(VMLINUX_BTF),) -$(error Cannot find a vmlinux for VMLINUX_BTF at any of "$(VMLINUX_BTF_PAT= HS)") -endif - -BPFTOOL ?=3D $(DEFAULT_BPFTOOL) +# The BPF schedulers are compiled into skeletons + subskeletons by the sha= red +# lib.bpf.mk fragment. They use the modern *.bpf.c layout with a .bpf.ske= l.h +# suffix and need subskeletons. +BPF_SRCS :=3D $(wildcard *.bpf.c) +BPF_SKEL_EXT :=3D .bpf.skel.h +BPF_GEN_SUBSKEL :=3D 1 +BPF_EXTRA_HDRS =3D $(wildcard $(CURDIR)/include/scx/*.h) \ + $(wildcard $(CURDIR)/include/*.h) +# Keep generated files under build/ as before (skeletons in build/include, +# objects in build/obj) rather than flat in the source dir. lib.bpf.mk cr= eates +# these dirs; SCXOBJ_DIR (userspace objects) reuses the BPF object dir. +BPF_OBJ_DIR :=3D $(OUTPUT)/build/obj/sched_ext +BPF_SKEL_DIR :=3D $(OUTPUT)/build/include +SCXOBJ_DIR :=3D $(BPF_OBJ_DIR) + +include ../lib.bpf.mk =20 ifneq ($(wildcard $(GENHDR)),) GENFLAGS :=3D -DHAVE_GENHDR endif =20 CFLAGS +=3D -g -O2 -rdynamic -pthread -Wall -Werror $(GENFLAGS) \ - -I$(INCLUDE_DIR) -I$(GENDIR) -I$(LIBDIR) \ - -I$(TOOLSINCDIR) -I$(APIDIR) -I$(CURDIR)/include -I$(SCXTOOLSINCDIR) + -I$(GENDIR) -I$(LIBDIR) -I$(TOOLSINCDIR) -I$(APIDIR) \ + -I$(CURDIR)/include -I$(SCXTOOLSINCDIR) =20 # Silence some warnings when compiled with clang ifneq ($(LLVM),) @@ -64,102 +52,29 @@ endif =20 LDFLAGS =3D -lelf -lz -lpthread -lzstd =20 -IS_LITTLE_ENDIAN =3D $(shell $(CC) -dM -E - &1 \ - | sed -n '/<...> search starts here:/,/End of search list./{ s| \(/.*\)|-= idirafter \1|p }') \ -$(shell $(1) $(2) -dM -E - $@ -else - $(call msg,CP,,$@) - $(Q)cp "$(VMLINUX_H)" $@ -endif + -fms-extensions =20 -$(SCXOBJ_DIR)/%.bpf.o: %.bpf.c $(INCLUDE_DIR)/vmlinux.h | $(BPFOBJ) $(SCXO= BJ_DIR) - $(call msg,CLNG-BPF,,$(notdir $@)) - $(Q)$(CLANG) $(BPF_CFLAGS) -target bpf -c $< -o $@ - -$(INCLUDE_DIR)/%.bpf.skel.h: $(SCXOBJ_DIR)/%.bpf.o $(INCLUDE_DIR)/vmlinux.= h $(BPFTOOL) | $(INCLUDE_DIR) - $(eval sched=3D$(notdir $@)) - $(call msg,GEN-SKEL,,$(sched)) - $(Q)$(BPFTOOL) gen object $(<:.o=3D.linked1.o) $< - $(Q)$(BPFTOOL) gen object $(<:.o=3D.linked2.o) $(<:.o=3D.linked1.o) - $(Q)$(BPFTOOL) gen object $(<:.o=3D.linked3.o) $(<:.o=3D.linked2.o) - $(Q)diff $(<:.o=3D.linked2.o) $(<:.o=3D.linked3.o) - $(Q)$(BPFTOOL) gen skeleton $(<:.o=3D.linked3.o) name $(subst .bpf.skel.h= ,,$(sched)) > $@ - $(Q)$(BPFTOOL) gen subskeleton $(<:.o=3D.linked3.o) name $(subst .bpf.ske= l.h,,$(sched)) > $(@:.skel.h=3D.subskel.h) +EXTRA_CLEAN +=3D $(OUTPUT)/build =20 ################ # C schedulers # ################ =20 -override define CLEAN - rm -rf $(OUTPUT_DIR) - rm -f $(TEST_GEN_PROGS) -endef - -# Every testcase takes all of the BPF progs are dependencies by default. T= his +# Every testcase takes all of the BPF progs as dependencies by default. Th= is # allows testcases to load any BPF scheduler, which is useful for testcases # that don't need their own prog to run their test. -all_test_bpfprogs :=3D $(foreach prog,$(wildcard *.bpf.c),$(INCLUDE_DIR)/$= (patsubst %.c,%.skel.h,$(prog))) +all_test_bpfprogs :=3D $(BPF_SKELS) =20 auto-test-targets :=3D \ create_dsq \ @@ -195,7 +110,8 @@ auto-test-targets :=3D \ testcase-targets :=3D $(addsuffix .o,$(addprefix $(SCXOBJ_DIR)/,$(auto-tes= t-targets))) =20 $(SCXOBJ_DIR)/runner.o: runner.c | $(SCXOBJ_DIR) $(BPFOBJ) - $(CC) $(CFLAGS) -c $< -o $@ + $(call msg,CC,,$@) + $(Q)$(CC) $(CFLAGS) -c $< -o $@ =20 # Create all of the test targets object files, whose testcase objects will= be # registered into the runner in ELF constructors. @@ -204,15 +120,16 @@ $(SCXOBJ_DIR)/runner.o: runner.c | $(SCXOBJ_DIR) $(BP= FOBJ) # compiling BPF object files only if one is present, as the wildcard Make # function doesn't support using implicit rules otherwise. $(testcase-targets): $(SCXOBJ_DIR)/%.o: %.c $(SCXOBJ_DIR)/runner.o $(all_t= est_bpfprogs) | $(SCXOBJ_DIR) - $(eval test=3D$(patsubst %.o,%.c,$(notdir $@))) - $(CC) $(CFLAGS) -c $< -o $@ + $(call msg,CC,,$@) + $(Q)$(CC) $(CFLAGS) -c $< -o $@ =20 $(SCXOBJ_DIR)/util.o: util.c | $(SCXOBJ_DIR) - $(CC) $(CFLAGS) -c $< -o $@ + $(call msg,CC,,$@) + $(Q)$(CC) $(CFLAGS) -c $< -o $@ =20 $(OUTPUT)/runner: $(SCXOBJ_DIR)/runner.o $(SCXOBJ_DIR)/util.o $(BPFOBJ) $(= testcase-targets) - @echo "$(testcase-targets)" - $(CC) $(CFLAGS) -o $@ $^ $(LDFLAGS) + $(call msg,BINARY,,$@) + $(Q)$(CC) $(CFLAGS) -o $@ $^ $(LDFLAGS) =20 .DEFAULT_GOAL :=3D all =20 --=20 2.53.0-Meta