From nobody Mon Sep 28 19:34:00 2026 Received: from mta0.migadu.com (out-169.mta0.migadu.com [91.218.175.169]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 01CDD3CCFC3 for ; Tue, 18 Aug 2026 08:36:48 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.169 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787042211; cv=none; b=tNW+hYLi3Y4eIxAB34FzBN74MUY4p7gkAU6c0z7AcKkjswviBkwXvPt9WvOt8VrGQ8ORNmd9Qa7KAxoLMn6DQce0oDvcoevaV1S5n7bFu5VfyxyORuFy1Tll9RxH65ReGhlEUJ6ScQBw9k1At/ZDqn1cS4cYlzeUKTzehS2FDxs= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787042211; c=relaxed/simple; bh=GReinsz2jYAS8LuTZNxhDAjdV+gzaBJE35QjgXUsYGo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=S4e+ZuZJ22KQSoxpCcKas8YpwBiRT8xM+gn3FN1h+fUCdILG+twtnQ7xGiXfoHPvC8GfVQS6ZwV430ZfRzGwJB759e/dKe2WOwwJjA2hzDnHLgXUh1dFKQwcWS8o3+1EC0I0dekkjhpBDp9W5FudZY2Fg7uTnRZuMnogTgiJG5g= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=AxS2nSI5; arc=none smtp.client-ip=91.218.175.169 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="AxS2nSI5" X-Envelope-To: linux-kernel@vger.kernel.org DKIM-Signature: a=rsa-sha256; bh=GReinsz2jYAS8LuTZNxhDAjdV+gzaBJE35QjgXUsYGo=; c=simple/simple; d=linux.dev; h=from:to:subject:date:message-id:mime-version:content-type; s=key1; t=1787042207; v=1; x=1787647007; b=AxS2nSI5T8MCMdlY4J0xF1udO87vv2cvYoZxK572h8blAIFgw+OVEhqQaZdFZNuL8pmbXHHq GlGxa9gerpVRAkodzse4ToWFbEsFfxN5sqT7IVVzWAcYa8oL0JQKOTYjmqtkk2XRlYpwVVgvCoQ I/7cOIyFN7v/h94oNiLdPA2M= X-Envelope-To: linux-kernel@vger.kernel.org Received: from teawater-KVM-Virtual-Machine (39.156.73.13) by smtp.migadu.com with ESMTPS id f63cb70c360e837f; Tue, 18 Aug 2026 08:36:47 +0000 X-Migadu-Flow: FLOW_OUT From: "Hui Zhu" To: Roman Gushchin , JP Kobryn , Shakeel Butt , Andrew Morton , Andrii Nakryiko , Eduard Zingerman , Ihor Solodrai , Alexei Starovoitov , Daniel Borkmann , Kumar Kartikeya Dwivedi , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Shuah Khan , Barry Song , Geliang Tang , linux-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-mm@kvack.org, linux-kselftest@vger.kernel.org Cc: Hui Zhu Subject: [PATCH bpf-next v2 1/2] mm/bpf: Add bpf_proactive_reclaim kfuncs Date: Tue, 18 Aug 2026 16:36:09 +0800 Message-ID: X-Mailer: git-send-email 2.53.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Hui Zhu Expose memcg proactive reclaim to sleepable BPF programs: unsigned long bpf_proactive_reclaim(memcg, size); unsigned long bpf_proactive_reclaim_swappiness(memcg, size, swappiness); They perform one reclaim pass on @memcg, like a write to memory.reclaim: swap is allowed, and the anon/file balance follows the cgroup's swappiness or an explicit override in [MIN_SWAPPINESS, MAX_SWAPPINESS] plus SWAPPINESS_ANON_ONLY. Both delegate to try_to_free_mem_cgroup_pages() with GFP_KERNEL and MEMCG_RECLAIM_MAY_SWAP | MEMCG_RECLAIM_PROACTIVE, the same parameters user_proactive_reclaim() uses, and unlike memory.reclaim they do not retry until @size is reached. Reclaim must not recurse: try_to_free_mem_cgroup_pages() overwrites current->reclaim_state on entry and NULLs it on exit, so a nested call from an in-flight reclaim would corrupt the outer reclaim state (e.g. MGLRU dereferences current->reclaim_state->mm_walk). Both kfuncs therefore refuse to reclaim when PF_MEMALLOC is set, mirroring the guards in the memcg charging path and node_reclaim(). Signed-off-by: Hui Zhu --- mm/bpf_memcontrol.c | 95 +++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 95 insertions(+) diff --git a/mm/bpf_memcontrol.c b/mm/bpf_memcontrol.c index 716df49d7647..92272f9a5825 100644 --- a/mm/bpf_memcontrol.c +++ b/mm/bpf_memcontrol.c @@ -6,6 +6,7 @@ */ =20 #include +#include #include =20 __bpf_kfunc_start_defs(); @@ -159,6 +160,97 @@ __bpf_kfunc void bpf_mem_cgroup_flush_stats(struct mem= _cgroup *memcg) mem_cgroup_flush_stats(memcg); } =20 +/* + * Reclaim must not recurse. try_to_free_mem_cgroup_pages() unconditionally + * overwrites current->reclaim_state on entry and resets it to NULL on exi= t. + * So invoking it from an in-flight reclaim would clobber the outer reclaim + * state and corrupt its accounting. + * + * The guard is PF_MEMALLOC. Every reclaim entry point marks the current + * task with it for the whole reclaim window: try_to_free_mem_cgroup_pages= () + * and __perform_reclaim() do so via memalloc_noreclaim_save(), and kswapd + * keeps it set for its entire lifetime. A hook inside the reclaim path + * (shrink_node, shrink_slab, ...) executes in the context of the + * reclaiming task, where current->flags already carries the flag. The page + * allocator, the memcg charging path and node_reclaim() rely on the same + * flag to avoid reclaim recursion. + * + * In try_to_free_mem_cgroup_pages(), reclaim_state is set slightly before + * PF_MEMALLOC, with only a tracepoint in between, which a sleepable BPF + * program cannot attach to. + * Also, PF_MEMALLOC is set in some non-reclaim contexts (e.g. direct comp= action + * and vmalloc), where the kfunc conservatively refuses to reclaim as well. + */ +static bool bpf_in_reclaim_context(void) +{ + return current->flags & PF_MEMALLOC; +} + +/** + * bpf_proactive_reclaim - proactively reclaim memory from a memory + * cgroup + * @memcg: the target memory cgroup to reclaim from + * @size: the amount of memory to reclaim, in bytes + * + * Trigger one proactive reclaim pass on @memcg, similar to a write to + * the memory.reclaim cgroup file: pages are reclaimed according to the + * cgroup's own swappiness setting and swap is allowed. Note that, + * unlike memory.reclaim, this does not retry until @size is reached; + * callers can invoke it again if needed. + * + * Return: + * The number of pages actually reclaimed, or 0 if @size is smaller + * than a page or the calling task is already in a reclaim/freeing + * context (PF_MEMALLOC). + */ +__bpf_kfunc unsigned long bpf_proactive_reclaim(struct mem_cgroup *memcg, + unsigned long size) +{ + unsigned long nr_pages =3D size / PAGE_SIZE; + + if (!nr_pages || unlikely(bpf_in_reclaim_context())) + return 0; + + return try_to_free_mem_cgroup_pages(memcg, nr_pages, GFP_KERNEL, + MEMCG_RECLAIM_MAY_SWAP | + MEMCG_RECLAIM_PROACTIVE, NULL); +} + +/** + * bpf_proactive_reclaim_swappiness - proactively reclaim memory from a + * memory cgroup with an explicit + * swappiness + * @memcg: the target memory cgroup to reclaim from + * @size: the amount of memory to reclaim, in bytes + * @swappiness: swappiness override for this reclaim pass + * + * Same as bpf_proactive_reclaim(), except that the anon/file reclaim + * balance is controlled by @swappiness instead of the cgroup's + * swappiness setting. Valid values are [MIN_SWAPPINESS, MAX_SWAPPINESS] + * and SWAPPINESS_ANON_ONLY, which restricts reclaim to anon folios. + * + * Return: + * The number of pages actually reclaimed, or 0 if @size is smaller + * than a page, @swappiness is out of range, or the calling task is + * already in a reclaim/freeing context (PF_MEMALLOC). + */ +__bpf_kfunc unsigned long +bpf_proactive_reclaim_swappiness(struct mem_cgroup *memcg, unsigned long s= ize, + int swappiness) +{ + unsigned long nr_pages =3D size / PAGE_SIZE; + + if (!nr_pages || swappiness < MIN_SWAPPINESS || + swappiness > SWAPPINESS_ANON_ONLY || + unlikely(bpf_in_reclaim_context())) + return 0; + + return try_to_free_mem_cgroup_pages(memcg, nr_pages, GFP_KERNEL, + MEMCG_RECLAIM_MAY_SWAP | + MEMCG_RECLAIM_PROACTIVE, + &swappiness); +} + __bpf_kfunc_end_defs(); =20 BTF_KFUNCS_START(bpf_memcontrol_kfuncs) @@ -172,6 +264,9 @@ BTF_ID_FLAGS(func, bpf_mem_cgroup_usage) BTF_ID_FLAGS(func, bpf_mem_cgroup_page_state) BTF_ID_FLAGS(func, bpf_mem_cgroup_flush_stats, KF_SLEEPABLE) =20 +BTF_ID_FLAGS(func, bpf_proactive_reclaim, KF_SLEEPABLE) +BTF_ID_FLAGS(func, bpf_proactive_reclaim_swappiness, KF_SLEEPABLE) + BTF_KFUNCS_END(bpf_memcontrol_kfuncs) =20 static const struct btf_kfunc_id_set bpf_memcontrol_kfunc_set =3D { --=20 2.53.0 From nobody Mon Sep 28 19:34:00 2026 Received: from mta0.migadu.com (out-176.mta0.migadu.com [91.218.175.176]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 97640422E25 for ; Tue, 18 Aug 2026 08:37:00 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.176 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787042223; cv=none; b=RIzO9mccdjKbT605Wha+jgtGvgVeEAA0DW2JeJ00w75ECvxzMP4DclMmsTHgj43X2YKxJkPn8WfzCvPR6vXlogyE5PIhE7RLdFqSdqAyLZZq9mXxK8QIU9LZzI51LqN2kP9Hw/3PiHOhrupKsn2rNabhPOq4/MEk2kR/DAy8L+Q= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787042223; c=relaxed/simple; bh=5FSc6tDpL9C78TLg7xeqAWyEH2FMc3NcZtq41DpkHD0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=lERwGOylNLoPdeNbbJeZgCnWQ5caTaUuRfSfyOOkd23eWX5vGYNVYYFt8WEDiV9H0S+xt62lQvqSa8ZEIDFymmIW6mhTx5XHP3hsOHSzkjbsrX+Ajq6P+no/GtCoxebg03MG5Rtg8h4rmqdeoXkFZ4FunvlRZKDVB5D69B4lx+k= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=FNLXIEFk; arc=none smtp.client-ip=91.218.175.176 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="FNLXIEFk" X-Envelope-To: linux-kernel@vger.kernel.org DKIM-Signature: a=rsa-sha256; bh=5FSc6tDpL9C78TLg7xeqAWyEH2FMc3NcZtq41DpkHD0=; c=simple/simple; d=linux.dev; h=from:to:subject:date:message-id:mime-version:content-type; s=key1; t=1787042218; v=1; x=1787647018; b=FNLXIEFkqoJlOCb32hpGOi0nq6/ZpT/Zt/jBlENiuyChIC5vCojSmUQgQNu+06CxvtIL0+8d iWHm6oX7KDU5nrIVL0ElC8q+v6zGlvsV3QgEjg/g5K5/MGXB8KAdquzDVSlvuYnKvDm9evPmRmq HQDJDRGQYfNMUtQVntRX4VOw= X-Envelope-To: linux-kernel@vger.kernel.org Received: from teawater-KVM-Virtual-Machine (39.156.73.13) by smtp.migadu.com with ESMTPS id 4b2f68af1fbe152b; Tue, 18 Aug 2026 08:36:58 +0000 X-Migadu-Flow: FLOW_OUT From: "Hui Zhu" To: Roman Gushchin , JP Kobryn , Shakeel Butt , Andrew Morton , Andrii Nakryiko , Eduard Zingerman , Ihor Solodrai , Alexei Starovoitov , Daniel Borkmann , Kumar Kartikeya Dwivedi , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Shuah Khan , Barry Song , Geliang Tang , linux-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-mm@kvack.org, linux-kselftest@vger.kernel.org Cc: Hui Zhu Subject: [PATCH bpf-next v2 2/2] selftests/bpf: add memcg async reclaim test Date: Tue, 18 Aug 2026 16:36:10 +0800 Message-ID: X-Mailer: git-send-email 2.53.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Hui Zhu Add memcg_async_reclaim selftest that verifies BPF-driven async proactive reclaim can mitigate refault-induced slowdown under memory pressure. The test creates a parent cgroup with a fixed memory.max, and two child cgroups (high/low) under it. Both children concurrently write and repeatedly read-fault a file larger than the shared limit. A BPF program monitors the "high" cgroup's WORKINGSET_REFAULT_FILE stat via a periodic timer, and when it detects refault growth beyond a threshold, triggers async reclaim on the "low" cgroup using bpf_proactive_reclaim(), expecting the "high" cgroup's workload to finish faster than without such reclaim. The reclaim work is queued asynchronously via bpf_wq. Signed-off-by: Hui Zhu --- .../bpf/prog_tests/memcg_async_reclaim.c | 382 ++++++++++++++++++ .../selftests/bpf/progs/memcg_async_reclaim.c | 167 ++++++++ 2 files changed, 549 insertions(+) create mode 100644 tools/testing/selftests/bpf/prog_tests/memcg_async_recl= aim.c create mode 100644 tools/testing/selftests/bpf/progs/memcg_async_reclaim.c diff --git a/tools/testing/selftests/bpf/prog_tests/memcg_async_reclaim.c b= /tools/testing/selftests/bpf/prog_tests/memcg_async_reclaim.c new file mode 100644 index 000000000000..6fab88203e7d --- /dev/null +++ b/tools/testing/selftests/bpf/prog_tests/memcg_async_reclaim.c @@ -0,0 +1,382 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Memory controller eBPF async reclaim test + */ + +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "cgroup_helpers.h" + +struct bpf_args_s { + u64 high_cgroup_id; + u64 low_cgroup_id; + u64 event_delta_threshold; + u64 check_ns; +}; + +#include "memcg_async_reclaim.skel.h" + +#define FILE_SIZE (32 * 1024 * 1024ul) +#define BUFFER_SIZE (4096) +#define CG_LIMIT (32 * 1024 * 1024ul) +#define READ_TIMES 50 + +#define CG_DIR "/memcg_async_reclaim" +#define CG_HIGH_DIR CG_DIR "/high" +#define CG_LOW_DIR CG_DIR "/low" + +#define CHECK_PERIOD_NS (2 * 1000 * 1000ull) +#define EVENT_DELTA_THRESHOLD 1 + +static int setup_high_low_cgroups(u64 *high_cgroup_id, u64 *low_cgroup_id) +{ + int ret; + char limit_buf[20]; + + ret =3D setup_cgroup_environment(); + if (!ASSERT_OK(ret, "setup_cgroup_environment")) + goto cleanup; + + ret =3D create_and_get_cgroup(CG_DIR); + if (!ASSERT_GE(ret, 0, "create_and_get_cgroup " CG_DIR)) + goto cleanup; + close(ret); + + ret =3D enable_controllers(CG_DIR, "memory"); + if (!ASSERT_OK(ret, "enable_controllers")) + goto cleanup; + + snprintf(limit_buf, sizeof(limit_buf), "%lu", CG_LIMIT); + ret =3D write_cgroup_file(CG_DIR, "memory.max", limit_buf); + if (!ASSERT_OK(ret, "write_cgroup_file memory.max")) + goto cleanup; + + ret =3D write_cgroup_file(CG_DIR, "memory.swap.max", "0"); + if (!ASSERT_OK(ret, "write_cgroup_file memory.swap.max")) + goto cleanup; + + ret =3D create_and_get_cgroup(CG_HIGH_DIR); + if (!ASSERT_GE(ret, 0, "create_and_get_cgroup " CG_HIGH_DIR)) + goto cleanup; + close(ret); + + *high_cgroup_id =3D get_cgroup_id(CG_HIGH_DIR); + if (!ASSERT_GT(*high_cgroup_id, 0, "get_cgroup_id")) + goto cleanup; + + ret =3D create_and_get_cgroup(CG_LOW_DIR); + if (!ASSERT_GE(ret, 0, "create_and_get_cgroup " CG_LOW_DIR)) + goto cleanup; + close(ret); + + *low_cgroup_id =3D get_cgroup_id(CG_LOW_DIR); + if (!ASSERT_GT(*low_cgroup_id, 0, "get_cgroup_id")) + goto cleanup; + + return 0; + +cleanup: + cleanup_cgroup_environment(); + return -1; +} + +static int write_file(const char *filename) +{ + int ret =3D -1; + size_t written =3D 0; + char *buffer; + FILE *fp; + + fp =3D fopen(filename, "wb"); + if (!fp) + goto out; + + buffer =3D malloc(BUFFER_SIZE); + if (!buffer) + goto cleanup_fp; + + memset(buffer, 'A', BUFFER_SIZE); + + while (written < FILE_SIZE) { + size_t to_write =3D FILE_SIZE - written < BUFFER_SIZE ? + FILE_SIZE - written : BUFFER_SIZE; + + if (fwrite(buffer, 1, to_write, fp) !=3D to_write) + goto cleanup; + written +=3D to_write; + } + + ret =3D 0; +cleanup: + free(buffer); +cleanup_fp: + fclose(fp); +out: + return ret; +} + +static int read_file(const char *filename, int iterations) +{ + int ret =3D -1; + long page_size =3D sysconf(_SC_PAGESIZE); + char *map; + size_t i; + int fd; + struct stat sb; + + fd =3D open(filename, O_RDONLY); + if (fd =3D=3D -1) + goto out; + + if (fstat(fd, &sb) =3D=3D -1) + goto cleanup_fd; + + if (sb.st_size !=3D FILE_SIZE) { + fprintf(stderr, "File size mismatch: expected %lu, got %lu\n", + (unsigned long)FILE_SIZE, (unsigned long)sb.st_size); + goto cleanup_fd; + } + + map =3D mmap(NULL, FILE_SIZE, PROT_READ, MAP_PRIVATE, fd, 0); + if (map =3D=3D MAP_FAILED) + goto cleanup_fd; + + for (int iter =3D 0; iter < iterations; iter++) { + for (i =3D 0; i < FILE_SIZE; i +=3D page_size) { + /* access a byte to trigger page fault */ + volatile char v =3D map[i]; + (void)v; + } + } + + if (munmap(map, FILE_SIZE) =3D=3D -1) + goto cleanup_fd; + + ret =3D 0; + +cleanup_fd: + close(fd); +out: + return ret; +} + +static int real_test_child_work(const char *cgroup_path, char *data_filena= me, + char *time_filename, int read_times) +{ + struct timeval start, end; + double elapsed; + FILE *fp; + + if (!ASSERT_OK(join_parent_cgroup(cgroup_path), "join_parent_cgroup")) + return -1; + + gettimeofday(&start, NULL); + + if (!ASSERT_OK(write_file(data_filename), "write_file")) + return -1; + + if (!ASSERT_OK(read_file(data_filename, read_times), "read_file")) + return -1; + + gettimeofday(&end, NULL); + + if (!time_filename) + return 0; + + elapsed =3D (end.tv_sec - start.tv_sec) + + (end.tv_usec - start.tv_usec) / 1000000.0; + printf("%.6f\n", elapsed); + + fp =3D fopen(time_filename, "w"); + if (!ASSERT_OK_PTR(fp, "fopen")) + return -1; + fprintf(fp, "%.6f", elapsed); + fclose(fp); + + return 0; +} + +static int get_time(char *time_filename, double *time) +{ + int ret =3D -1; + FILE *fp; + char buf[64]; + + fp =3D fopen(time_filename, "r"); + if (!ASSERT_OK_PTR(fp, "fopen")) + goto out; + + if (!ASSERT_OK_PTR(fgets(buf, sizeof(buf), fp), "fgets")) + goto cleanup; + + if (sscanf(buf, "%lf", time) !=3D 1) { + PRINT_FAIL("sscanf %s", buf); + goto cleanup; + } + + ret =3D 0; +cleanup: + fclose(fp); +out: + return ret; +} + +static int +run_high_low_workload(double *high_elapsed, double *low_elapsed, int read_= times) +{ + char high_data_file[] =3D "/tmp/memcg_async_high_data_XXXXXX"; + char low_data_file[] =3D "/tmp/memcg_async_low_data_XXXXXX"; + char high_time_file[] =3D "/tmp/memcg_async_high_time_XXXXXX"; + char low_time_file[] =3D "/tmp/memcg_async_low_time_XXXXXX"; + pid_t high_pid =3D -1, low_pid =3D -1; + int fd, status; + int ret =3D -1; + + fd =3D mkstemp(high_data_file); + if (!ASSERT_GE(fd, 0, "mkstemp")) + goto cleanup; + close(fd); + + fd =3D mkstemp(low_data_file); + if (!ASSERT_GE(fd, 0, "mkstemp")) + goto cleanup; + close(fd); + + fd =3D mkstemp(high_time_file); + if (!ASSERT_GE(fd, 0, "mkstemp")) + goto cleanup; + close(fd); + + fd =3D mkstemp(low_time_file); + if (!ASSERT_GE(fd, 0, "mkstemp")) + goto cleanup; + close(fd); + + low_pid =3D fork(); + if (!ASSERT_GE(low_pid, 0, "fork low")) + goto cleanup; + if (low_pid =3D=3D 0) + exit(real_test_child_work(CG_LOW_DIR, low_data_file, + low_time_file, read_times)); + + high_pid =3D fork(); + if (!ASSERT_GE(high_pid, 0, "fork high")) + goto cleanup; + if (high_pid =3D=3D 0) + exit(real_test_child_work(CG_HIGH_DIR, high_data_file, + high_time_file, read_times)); + + if (!ASSERT_GT(waitpid(low_pid, &status, 0), 0, "low waitpid")) + goto cleanup; + if (!ASSERT_TRUE(WIFEXITED(status), "low exited")) + goto cleanup; + if (!ASSERT_EQ(WEXITSTATUS(status), 0, "low exit status")) + goto cleanup; + + if (!ASSERT_GT(waitpid(high_pid, &status, 0), 0, "high waitpid")) + goto cleanup; + if (!ASSERT_TRUE(WIFEXITED(status), "high exited")) + goto cleanup; + if (!ASSERT_EQ(WEXITSTATUS(status), 0, "high exit status")) + goto cleanup; + + if (get_time(high_time_file, high_elapsed)) + goto cleanup; + if (get_time(low_time_file, low_elapsed)) + goto cleanup; + + ret =3D 0; + +cleanup: + /* On failure, make sure no child process is left behind */ + if (ret) { + if (high_pid > 0) { + kill(high_pid, SIGKILL); + (void)waitpid(high_pid, NULL, 0); + } + if (low_pid > 0) { + kill(low_pid, SIGKILL); + (void)waitpid(low_pid, NULL, 0); + } + } + unlink(low_time_file); + unlink(high_time_file); + unlink(low_data_file); + unlink(high_data_file); + return ret; +} + +static int +setup_bpf(u64 high_cgroup_id, u64 low_cgroup_id, + struct memcg_async_reclaim **skel_ptr) +{ + struct memcg_async_reclaim *skel; + struct bpf_args_s bpf_args =3D { + .high_cgroup_id =3D high_cgroup_id, + .low_cgroup_id =3D low_cgroup_id, + .event_delta_threshold =3D EVENT_DELTA_THRESHOLD, + .check_ns =3D CHECK_PERIOD_NS, + }; + LIBBPF_OPTS(bpf_test_run_opts, run_opts, + .ctx_in =3D &bpf_args, + .ctx_size_in =3D sizeof(bpf_args)); + int prog_init_fd, err; + + skel =3D memcg_async_reclaim__open_and_load(); + if (!ASSERT_OK_PTR(skel, "memcg_async_reclaim__open_and_load")) + return -1; + + prog_init_fd =3D bpf_program__fd(skel->progs.wq_prog_init); + + err =3D bpf_prog_test_run_opts(prog_init_fd, &run_opts); + if (!ASSERT_OK(err, "bpf_prog_test_run_opts")) + goto error_out; + if (!ASSERT_EQ(run_opts.retval, 0, "prog_init retval")) + goto error_out; + + *skel_ptr =3D skel; + return 0; + +error_out: + memcg_async_reclaim__destroy(skel); + return -1; +} + +void test_memcg_wq_async_reclaim(void) +{ + u64 high_cgroup_id, low_cgroup_id; + int err; + double high_time =3D 0.0, low_time =3D 0.0; + struct memcg_async_reclaim *skel =3D NULL; + + err =3D setup_high_low_cgroups(&high_cgroup_id, &low_cgroup_id); + if (!ASSERT_OK(err, "setup_high_low_cgroups reclaim")) + return; + + err =3D setup_bpf(high_cgroup_id, low_cgroup_id, &skel); + if (!ASSERT_OK(err, "setup_bpf")) + goto out; + + err =3D run_high_low_workload(&high_time, &low_time, READ_TIMES); + if (!ASSERT_OK(err, "run_high_low_workload reclaim")) + goto out; + + if (high_time >=3D low_time) + PRINT_FAIL("high cgroup not improved with async reclaim: high_time=3D%f = low_time=3D%f", + high_time, low_time); + +out: + if (skel) + memcg_async_reclaim__destroy(skel); + cleanup_cgroup_environment(); +} diff --git a/tools/testing/selftests/bpf/progs/memcg_async_reclaim.c b/tool= s/testing/selftests/bpf/progs/memcg_async_reclaim.c new file mode 100644 index 000000000000..62c2bb7e037b --- /dev/null +++ b/tools/testing/selftests/bpf/progs/memcg_async_reclaim.c @@ -0,0 +1,167 @@ +// SPDX-License-Identifier: GPL-2.0 + +#include "vmlinux.h" +#include "bpf_experimental.h" +#include +#include + +#define CLOCK_MONOTONIC_ID 1 +#define PAGE_SIZE 4096UL +#define RECLAIM_SIZE (32 * PAGE_SIZE) +#define RECLAIM_MAX_ITER 32 + +struct bpf_args_s { + u64 high_cgroup_id; + u64 low_cgroup_id; + u64 event_delta_threshold; + u64 check_ns; +}; + +struct cgroup_memcg { + struct cgroup *cgrp; + struct mem_cgroup *memcg; +}; + +static u64 wq_high_cgroup_id; +static u64 wq_low_cgroup_id; + +static int get_cgroup_memcg_from_id(u64 cgroup_id, struct cgroup_memcg *cm) +{ + cm->cgrp =3D bpf_cgroup_from_id(cgroup_id); + if (!cm->cgrp) + return -1; + + cm->memcg =3D bpf_get_mem_cgroup(&cm->cgrp->self); + if (!cm->memcg) { + bpf_cgroup_release(cm->cgrp); + return -1; + } + + return 0; +} + +static void put_cgroup_memcg(struct cgroup_memcg *cm) +{ + bpf_put_mem_cgroup(cm->memcg); + bpf_cgroup_release(cm->cgrp); +} + +static int get_cgroup_event(u64 cgroup_id, u64 *val) +{ + struct cgroup_memcg cm; + + if (get_cgroup_memcg_from_id(cgroup_id, &cm)) + return -1; + bpf_mem_cgroup_flush_stats(cm.memcg); + *val =3D bpf_mem_cgroup_page_state(cm.memcg, WORKINGSET_REFAULT_FILE); + put_cgroup_memcg(&cm); + + return 0; +} + +static bool +should_reclaim_cgroup(u64 cgroup_id, u64 *prev_event, u64 event_delta_thre= shold) +{ + u64 cur, delta; + + if (get_cgroup_event(cgroup_id, &cur)) + return false; + + delta =3D cur - *prev_event; + *prev_event =3D cur; + + return delta >=3D event_delta_threshold; +} + +static int reclaim_cgroup(u64 cgroup_id) +{ + struct cgroup_memcg cm; + int i; + + if (get_cgroup_memcg_from_id(cgroup_id, &cm)) + return 0; + + for (i =3D 0; i < RECLAIM_MAX_ITER; i++) { + if (!bpf_proactive_reclaim(cm.memcg, RECLAIM_SIZE)) + break; + } + + put_cgroup_memcg(&cm); + + return 0; +} + +struct wq_elem { + struct bpf_timer timer; + struct bpf_wq work; + u64 prev_event; + u64 event_delta_threshold; + u64 check_ns; +}; + +struct { + __uint(type, BPF_MAP_TYPE_ARRAY); + __uint(max_entries, 1); + __type(key, __u32); + __type(value, struct wq_elem); +} wq_map SEC(".maps"); + +static int async_free(void *map, int *key, void *value) +{ + struct wq_elem *elem =3D value; + + if (should_reclaim_cgroup(wq_high_cgroup_id, &elem->prev_event, + elem->event_delta_threshold)) { + reclaim_cgroup(wq_low_cgroup_id); + bpf_wq_start(&elem->work, 0); + } + + return 0; +} + +static int wq_timer_cb(void *map, int *key, struct wq_elem *elem) +{ + bpf_wq_start(&elem->work, 0); + bpf_timer_start(&elem->timer, elem->check_ns, 0); + + return 0; +} + +SEC("syscall") +int wq_prog_init(struct bpf_args_s *ctx) +{ + struct wq_elem *elem; + __u32 key =3D 0; + int ret; + + elem =3D bpf_map_lookup_elem(&wq_map, &key); + if (!elem) + return -1; + + ret =3D bpf_wq_init(&elem->work, &wq_map, 0); + if (ret) + return ret; + + ret =3D bpf_wq_set_callback(&elem->work, async_free, 0); + if (ret) + return ret; + + ret =3D bpf_timer_init(&elem->timer, &wq_map, CLOCK_MONOTONIC_ID); + if (ret) + return ret; + + ret =3D bpf_timer_set_callback(&elem->timer, wq_timer_cb); + if (ret) + return ret; + + elem->prev_event =3D 0; + elem->event_delta_threshold =3D ctx->event_delta_threshold; + elem->check_ns =3D ctx->check_ns; + + wq_high_cgroup_id =3D ctx->high_cgroup_id; + wq_low_cgroup_id =3D ctx->low_cgroup_id; + + return bpf_timer_start(&elem->timer, elem->check_ns, 0); +} + +char LICENSE[] SEC("license") =3D "GPL"; --=20 2.53.0