From nobody Fri Oct 2 03:40:51 2026 Received: from out-173.mta1.migadu.com (mta1.migadu.com [37.59.57.117]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8DFC53F44FC for ; Wed, 5 Aug 2026 09:20:13 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=37.59.57.117 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785921617; cv=none; b=Sqba85Okk2aigrlZReCvQWL+mxaQOcu+uvMNrsuzADxJSxLJAhfSjHpSxPMdzTmS9jo6DVFHAmjJjEZ6cD7a0vLT4eXShbpts+ImgfmhnsMIR5PMl7cV8y/6PT6Q9kwvQMAdx+3K2r5I1559aCk1jZh69zFEnOC8E7cGKlfFQtI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785921617; c=relaxed/simple; bh=ijsnVrTrJAr9iIFufGidF9T2LTUTcg75qzLv7bwVVP8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=MDvaaKlq5Dmcuq09JZosF+sb25ps7CM14nyCRXrZgXmRTAePlfNEwdjMS+k8uKQ1KaRikEvin6xi9VK5R2vu25umZMWw1tLAgPtEiHuHT+fRwEecEVPQjsibMAG5YcQ8374vjCppVBEu4HGmVLG9zf1iSie/cBhCbhNTsSx5/yI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=ArIN6sN/; arc=none smtp.client-ip=37.59.57.117 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="ArIN6sN/" X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers. DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1; t=1785921611; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=nI2x4Jw4QeDH+z0w8stHMu3B2x9kMF1+3hiOMr55Tzo=; b=ArIN6sN/8ojY0V9jMtDPvYjzaY+pzmBJlFZRF64c4OxST+5INcT9DOa5Qsq/KaRA8YsDmT XbD3U1Dmh4eIWbRvatGb0WWwB21vi4Hi1tLDWsJ5vYLqn3lhxdYlcY6CHZOHe1qM7EcBBR 0NE+eI64Ivk/Dsif8WrSBtomjYu0b6E= From: Jiayuan Chen To: bpf@vger.kernel.org Cc: Jiayuan Chen , Emil Tsalapatis , Alexei Starovoitov , Daniel Borkmann , John Fastabend , Andrii Nakryiko , Eduard Zingerman , Kumar Kartikeya Dwivedi , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Ihor Solodrai , Shuah Khan , Sebastian Andrzej Siewior , Clark Williams , Steven Rostedt , linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-rt-devel@lists.linux.dev Subject: [PATCH bpf-next v2 1/4] bpf: Add a sleepable page allocator for map memory Date: Wed, 5 Aug 2026 17:15:54 +0800 Message-ID: <20260805091720.139924-2-jiayuan.chen@linux.dev> In-Reply-To: <20260805091720.139924-1-jiayuan.chen@linux.dev> References: <20260805091720.139924-1-jiayuan.chen@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Migadu-Flow: FLOW_OUT Content-Type: text/plain; charset="utf-8" bpf_map_alloc_pages() picks the allocator via can_alloc_pages(), a conservative guess for BPF program context that is always false under PREEMPT_RT. So even a caller that really is sleepable gets the non-blocking allocator, which never reclaims and never engages the OOM machinery. Add bpf_map_alloc_page_sleepable() for callers that know they are sleepable. It allocates from the map's numa_node, like the other map allocators. The next patch uses it from the arena page fault handler. Signed-off-by: Jiayuan Chen Reviewed-by: Emil Tsalapatis --- include/linux/bpf.h | 1 + kernel/bpf/syscall.c | 21 +++++++++++++++++---- 2 files changed, 18 insertions(+), 4 deletions(-) diff --git a/include/linux/bpf.h b/include/linux/bpf.h index 73bacfc6444d..0dc51c37c25c 100644 --- a/include/linux/bpf.h +++ b/include/linux/bpf.h @@ -2784,6 +2784,7 @@ struct bpf_prog *bpf_prog_get_curr_or_next(u32 *id); =20 int bpf_map_alloc_pages(const struct bpf_map *map, int nid, unsigned long nr_pages, struct page **page_array); +struct page *bpf_map_alloc_page_sleepable(const struct bpf_map *map); #ifdef CONFIG_MEMCG void bpf_map_memcg_enter(const struct bpf_map *map, struct mem_cgroup **ol= d_memcg, struct mem_cgroup **new_memcg); diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c index 8d111da88655..67d8157c6623 100644 --- a/kernel/bpf/syscall.c +++ b/kernel/bpf/syscall.c @@ -602,15 +602,14 @@ static bool can_alloc_pages(void) !IS_ENABLED(CONFIG_PREEMPT_RT); } =20 +#define BPF_PAGE_GFP (GFP_KERNEL | __GFP_ZERO | __GFP_ACCOUNT | __GFP_NOWA= RN) + static struct page *__bpf_alloc_page(int nid) { if (!can_alloc_pages()) return alloc_pages_nolock(__GFP_ACCOUNT, nid, 0); =20 - return alloc_pages_node(nid, - GFP_KERNEL | __GFP_ZERO | __GFP_ACCOUNT - | __GFP_NOWARN, - 0); + return alloc_pages_node(nid, BPF_PAGE_GFP, 0); } =20 int bpf_map_alloc_pages(const struct bpf_map *map, int nid, @@ -636,6 +635,20 @@ int bpf_map_alloc_pages(const struct bpf_map *map, int= nid, return ret; } =20 +/* + * For callers that know they run in a sleepable context, e.g. a user page + * fault handler. can_alloc_pages() is a conservative guess made for BPF + * program context - notably it is always false on PREEMPT_RT - so going + * through bpf_map_alloc_pages() there would needlessly pick the + * non-blocking allocator, which never reclaims and never engages the OOM + * machinery. + */ +struct page *bpf_map_alloc_page_sleepable(const struct bpf_map *map) +{ + might_sleep(); + return alloc_pages_node(map->numa_node, BPF_PAGE_GFP, 0); +} + =20 static int btf_field_cmp(const void *a, const void *b) { --=20 2.43.0 From nobody Fri Oct 2 03:40:51 2026 Received: from out-188.mta1.migadu.com (out-188.mta1.migadu.com [95.215.58.188]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 260963EAC84 for ; Wed, 5 Aug 2026 09:20:33 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=95.215.58.188 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785921634; cv=none; b=OQLoUyquxWIGRD8pRGdMOOlGAaPVDc2LU2bwwXaM5mhkOonTXW5LfyYQkIzDbKJaOFJ1/H2El4kmoVJNC4fLM601jfXayC6ibO9JkH1Pido6UZ3r5YREBRs5HGLWAf97HyeHW20giiwGHLsCNmVtKiKBn0ycugExMT4GTh2HRsY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785921634; c=relaxed/simple; bh=4NwivRdVnxpe0hR8QT/ODMd9jJcAcOepgsNTHf5zFbI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=sLQHfGlJRFswxJM8Jb6124G4MY6E1elgsmxSH12q9RkfZcYKSUZjTCC1EWwcD+pCeFjv+g8g3OQv1BAxzzG+hPD8CDAty2ezs39wVm+yENtOhs2e7opMH/Zpj1aZv4P0bixmhU84pn0Wz1tGwvBPj3AvadXvZ1gozs5ssAqq4/Y= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=Ye1xB6ZF; arc=none smtp.client-ip=95.215.58.188 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="Ye1xB6ZF" X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers. DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1; t=1785921631; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=KTDziX6imkrixNehHfXUmTeOgD2T/xzdXYvI2wwlPos=; b=Ye1xB6ZFDMApQKe4evSzKJRBe8ZMMN0AWfiP4HTSXRxkHEPhgzDQehQJR6KOljq0hobV3I DY/v/OfptacSkrdwRENdWLq6TBoYQyMQ48CCe6wSp3Tdn3ZIz3DBBzY0vPkeI+6at3jX7+ B9ytavNVKT6ajzCp1e6/sCvC5R/LXsk= From: Jiayuan Chen To: bpf@vger.kernel.org Cc: Jiayuan Chen , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , Eduard Zingerman , Kumar Kartikeya Dwivedi , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Ihor Solodrai , John Fastabend , Shuah Khan , Sebastian Andrzej Siewior , Clark Williams , Steven Rostedt , linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-rt-devel@lists.linux.dev Subject: [PATCH bpf-next v2 2/4] bpf: arena: allocate the fault-in page outside the lock Date: Wed, 5 Aug 2026 17:15:55 +0800 Message-ID: <20260805091720.139924-3-jiayuan.chen@linux.dev> In-Reply-To: <20260805091720.139924-1-jiayuan.chen@linux.dev> References: <20260805091720.139924-1-jiayuan.chen@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Migadu-Flow: FLOW_OUT Content-Type: text/plain; charset="utf-8" arena_vm_fault() allocated the page while holding arena->spinlock, so it could only use the non-blocking allocator. Once the memcg is at memory.max that allocation just fails, the fault turns into VM_FAULT_SIGSEGV, and the process gets a SIGSEGV on a perfectly valid arena address. Hitting memory.max is routine (e.g. page cache from reading a big file), so this kills innocent processes. Rework the fault handler: - Preallocate the page before taking the lock, like do_anonymous_page() does, so it can sleep, reclaim and go through the OOM path, and return VM_FAULT_OOM on failure so the memcg OOM handler runs instead of a fake segfault. - A lockless probe skips that preallocation when a page is already mapped (e.g. allocated by the bpf program), so the common case wastes no allocation. The rare race where such a page is freed before we take the lock falls back to the non-blocking allocator under the lock. - Return VM_FAULT_SIGBUS for the non-recoverable errors (lock failure, range-tree and page-table failures) instead of VM_FAULT_SIGSEGV; only BPF_F_SEGV_ON_FAULT, and a scratch-page hole under that flag, is a real user addressing error and keeps VM_FAULT_SIGSEGV. - Tidy up the error labels. Signed-off-by: Jiayuan Chen --- kernel/bpf/arena.c | 92 +++++++++++++++++++++++++++++++++++----------- 1 file changed, 71 insertions(+), 21 deletions(-) diff --git a/kernel/bpf/arena.c b/kernel/bpf/arena.c index 555ee2531ef9..09a718ca4c8b 100644 --- a/kernel/bpf/arena.c +++ b/kernel/bpf/arena.c @@ -481,7 +481,8 @@ static vm_fault_t arena_vm_fault(struct vm_fault *vmf) struct bpf_map *map =3D vmf->vma->vm_file->private_data; struct bpf_arena *arena =3D container_of(map, struct bpf_arena, map); struct mem_cgroup *new_memcg, *old_memcg; - struct page *page; + struct page *page, *new_page =3D NULL; + vm_fault_t fault_ret; long kbase, kaddr; unsigned long flags; int ret; @@ -489,59 +490,108 @@ static vm_fault_t arena_vm_fault(struct vm_fault *vm= f) kbase =3D bpf_arena_get_kern_vm_start(arena); kaddr =3D kbase + (u32)(vmf->address); =20 - if (raw_res_spin_lock_irqsave(&arena->spinlock, flags)) + page =3D vmalloc_to_page((void *)kaddr); + if (!page && !(arena->map.map_flags & BPF_F_SEGV_ON_FAULT)) { + /* + * We run in process context here, so preallocate the page + * outside the lock with an explicitly sleepable allocator. It + * can then go through reclaim (both memcg and global) and the + * OOM path, the way do_anonymous_page() does; under + * arena->spinlock only the non-blocking allocator is available, + * which never reclaims. That also decides the return value: + * VM_FAULT_OOM below is only meaningful if the OOM machinery was + * actually engaged, which the non-blocking allocator never does. + */ + bpf_map_memcg_enter(&arena->map, &old_memcg, &new_memcg); + new_page =3D bpf_map_alloc_page_sleepable(map); + bpf_map_memcg_exit(old_memcg, new_memcg); + if (!new_page) + return VM_FAULT_OOM; + } + + if (raw_res_spin_lock_irqsave(&arena->spinlock, flags)) { /* * A failed lock means a possible deadlock was detected. Don't * return VM_FAULT_RETRY: this handler never took mmap_lock, but * the fault path would re-take it on retry and deadlock. Fail. */ + if (new_page) + free_pages_nolock(new_page, 0); return VM_FAULT_SIGBUS; + } =20 page =3D vmalloc_to_page((void *)kaddr); if (page) { - if (page =3D=3D arena->scratch_page) - /* BPF triggered scratch here; don't lazy-alloc over it */ - goto out_sigsegv; + if (page =3D=3D arena->scratch_page) { + /* + * A scratch page marks a hole. Segfault only if the user + * asked for it; otherwise we could lazy-allocate but + * choose not to over a hole, so report a bus error. + */ + fault_ret =3D (arena->map.map_flags & BPF_F_SEGV_ON_FAULT) ? + VM_FAULT_SIGSEGV : VM_FAULT_SIGBUS; + goto out_err_locked; + } /* already have a page vmap-ed */ goto out; } =20 + if (arena->map.map_flags & BPF_F_SEGV_ON_FAULT) { + /* User space requested to segfault when page is not allocated by bpf pr= og */ + fault_ret =3D VM_FAULT_SIGSEGV; + goto out_err_locked; + } + bpf_map_memcg_enter(&arena->map, &old_memcg, &new_memcg); =20 - if (arena->map.map_flags & BPF_F_SEGV_ON_FAULT) - /* User space requested to segfault when page is not allocated by bpf pr= og */ - goto out_sigsegv_memcg; + if (!new_page) { + /* + * Very rare race: the bpf program had allocated a page here, so + * the lockless probe saw it and we skipped preallocation, but it + * freed the page before we took the lock. Now we do need one; + * sleeping is not allowed here, so fall back to the non-blocking + * allocator and give up if it fails. + */ + ret =3D bpf_map_alloc_pages(map, map->numa_node, 1, &new_page); + if (ret) { + fault_ret =3D VM_FAULT_SIGBUS; + goto out_err_locked_memcg; + } + } =20 ret =3D range_tree_clear(&arena->rt, vmf->pgoff, 1); - if (ret) - goto out_sigsegv_memcg; - - struct apply_range_data data =3D { .arena =3D arena, .pages =3D &page, .i= =3D 0 }; - /* Account into memcg of the process that created bpf_arena */ - ret =3D bpf_map_alloc_pages(map, NUMA_NO_NODE, 1, &page); if (ret) { - range_tree_set(&arena->rt, vmf->pgoff, 1); - goto out_sigsegv_memcg; + fault_ret =3D VM_FAULT_SIGBUS; + goto out_err_locked_memcg; } + struct apply_range_data data =3D { .arena =3D arena, .pages =3D &new_page= , .i =3D 0 }; =20 ret =3D apply_to_page_range(&init_mm, kaddr, PAGE_SIZE, apply_range_set_c= b, &data); if (ret) { range_tree_set(&arena->rt, vmf->pgoff, 1); - free_pages_nolock(page, 0); - goto out_sigsegv_memcg; + fault_ret =3D VM_FAULT_SIGBUS; + goto out_err_locked_memcg; } flush_vmap_cache(kaddr, PAGE_SIZE); bpf_map_memcg_exit(old_memcg, new_memcg); + /* new_page was consumed */ + page =3D new_page; + new_page =3D NULL; out: page_ref_add(page, 1); raw_res_spin_unlock_irqrestore(&arena->spinlock, flags); + if (new_page) + free_pages_nolock(new_page, 0); vmf->page =3D page; return 0; -out_sigsegv_memcg: + +out_err_locked_memcg: bpf_map_memcg_exit(old_memcg, new_memcg); -out_sigsegv: +out_err_locked: raw_res_spin_unlock_irqrestore(&arena->spinlock, flags); - return VM_FAULT_SIGSEGV; + if (new_page) + free_pages_nolock(new_page, 0); + return fault_ret; } =20 static const struct vm_operations_struct arena_vm_ops =3D { --=20 2.43.0 From nobody Fri Oct 2 03:40:51 2026 Received: from out-181.mta1.migadu.com (mta1.migadu.com [37.59.57.117]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0DAB93EB0FF for ; Wed, 5 Aug 2026 09:20:52 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=37.59.57.117 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785921655; cv=none; b=JT3SZ7YtthoJUVf0YPmg7cOllZl5n64j0tnR4RX1SzqHr8aW0OG1+04iR8B1g4F0GpEAJ3VbdjUd/zcU58+nryXSZ2zv2AlSHeJzoLF6EDbzR9rf4qHiHLndJiHA1qyzkeN81RxaSsUGykgdXktYr7NwmIi3TIp1Wemd8NeO95Q= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785921655; c=relaxed/simple; bh=FP5V+/ygynVfbTL23lPPyV9rBlyavXD/fkc1dieAk/c=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=OZoYhHiBa1PHU5+rToWwEPfke+wVHQmS/C8Hn2gzaf1BwMI3A8JVsuG4Jdlxs7/ODbIcY9YEfeu2WZXnTT9Nra74iEP1rEkNT+46HHXUaDs2OLyPpQG+NnP3kz0JwdSwjHbSNUS/ewdt0KSTW4E2PY9SdKjMOQrNQsO0NuGbDW4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=mz4neGt7; arc=none smtp.client-ip=37.59.57.117 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="mz4neGt7" X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers. DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1; t=1785921651; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=aP+xZvCeYjCPMovU6nFO+tDWTQv8q8MytFmDoJ6Dhdo=; b=mz4neGt7wz1ImmnvDP2k2qZuzSwKoxRlR/kfSeuCBLZl7b5prSOiTtkpI7LzchFW2P5zA+ UQkQvpNvyN11shJ7kDnknweZFCwVC6vj0mY1EMin5AvXw+ipEOZZNFjNI9Kc1WBkeaZcS5 kQGZMMrbE1fC53+HvknVcOZ1ZYEwYyw= From: Jiayuan Chen To: bpf@vger.kernel.org Cc: Jiayuan Chen , Alexei Starovoitov , Daniel Borkmann , John Fastabend , Andrii Nakryiko , Eduard Zingerman , Kumar Kartikeya Dwivedi , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Ihor Solodrai , Shuah Khan , Sebastian Andrzej Siewior , Clark Williams , Steven Rostedt , linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-rt-devel@lists.linux.dev Subject: [PATCH bpf-next v2 3/4] selftests/bpf: Add read_cgroup_file() to cgroup_helpers Date: Wed, 5 Aug 2026 17:15:56 +0800 Message-ID: <20260805091720.139924-4-jiayuan.chen@linux.dev> In-Reply-To: <20260805091720.139924-1-jiayuan.chen@linux.dev> References: <20260805091720.139924-1-jiayuan.chen@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Migadu-Flow: FLOW_OUT Content-Type: text/plain; charset="utf-8" cgroup_helpers has write_cgroup_file()/write_cgroup_file_parent() but no read counterpart. Add read_cgroup_file() and read_cgroup_file_parent() so a forked child can read a cgroup file (e.g. memory.current) from the work dir owned by the parent that set the environment up, without hand-building the /mnt/... path. Signed-off-by: Jiayuan Chen --- tools/testing/selftests/bpf/cgroup_helpers.c | 67 ++++++++++++++++++++ tools/testing/selftests/bpf/cgroup_helpers.h | 4 ++ 2 files changed, 71 insertions(+) diff --git a/tools/testing/selftests/bpf/cgroup_helpers.c b/tools/testing/s= elftests/bpf/cgroup_helpers.c index 45cd0b479fe3..4183ff6150c2 100644 --- a/tools/testing/selftests/bpf/cgroup_helpers.c +++ b/tools/testing/selftests/bpf/cgroup_helpers.c @@ -188,6 +188,73 @@ int write_cgroup_file_parent(const char *relative_path= , const char *file, return __write_cgroup_file(cgroup_path, file, buf); } =20 +static int __read_cgroup_file(const char *cgroup_path, const char *file, + char *buf, size_t len) +{ + char file_path[PATH_MAX + 1]; + ssize_t got; + int fd; + + snprintf(file_path, sizeof(file_path), "%s/%s", cgroup_path, file); + fd =3D open(file_path, O_RDONLY); + if (fd < 0) { + log_err("Opening %s", file_path); + return 1; + } + + got =3D read(fd, buf, len - 1); + if (got < 0) { + log_err("Reading %s", file_path); + close(fd); + return 1; + } + buf[got] =3D '\0'; + close(fd); + return 0; +} + +/** + * read_cgroup_file() - Read from a cgroup file + * @relative_path: The cgroup path, relative to the workdir + * @file: The name of the file in cgroupfs to read from + * @buf: Buffer to read into, NUL-terminated on success + * @len: Size of @buf + * + * Read from a file in the given cgroup's directory. + * + * If successful, 0 is returned. + */ +int read_cgroup_file(const char *relative_path, const char *file, + char *buf, size_t len) +{ + char cgroup_path[PATH_MAX - 24]; + + format_cgroup_path(cgroup_path, relative_path); + return __read_cgroup_file(cgroup_path, file, buf, len); +} + +/** + * read_cgroup_file_parent() - Read from a cgroup file in the parent proce= ss + * workdir + * @relative_path: The cgroup path, relative to the parent process workdir + * @file: The name of the file in cgroupfs to read from + * @buf: Buffer to read into, NUL-terminated on success + * @len: Size of @buf + * + * Read from a file in the given cgroup's directory under the parent proce= ss + * workdir. + * + * If successful, 0 is returned. + */ +int read_cgroup_file_parent(const char *relative_path, const char *file, + char *buf, size_t len) +{ + char cgroup_path[PATH_MAX - 24]; + + format_parent_cgroup_path(cgroup_path, relative_path); + return __read_cgroup_file(cgroup_path, file, buf, len); +} + /** * setup_cgroup_environment() - Setup the cgroup environment * diff --git a/tools/testing/selftests/bpf/cgroup_helpers.h b/tools/testing/s= elftests/bpf/cgroup_helpers.h index 3857304be874..d42d2e13044e 100644 --- a/tools/testing/selftests/bpf/cgroup_helpers.h +++ b/tools/testing/selftests/bpf/cgroup_helpers.h @@ -15,6 +15,10 @@ int write_cgroup_file(const char *relative_path, const c= har *file, const char *buf); int write_cgroup_file_parent(const char *relative_path, const char *file, const char *buf); +int read_cgroup_file(const char *relative_path, const char *file, + char *buf, size_t len); +int read_cgroup_file_parent(const char *relative_path, const char *file, + char *buf, size_t len); int cgroup_setup_and_join(const char *relative_path); int get_root_cgroup(void); int create_and_get_cgroup(const char *relative_path); --=20 2.43.0 From nobody Fri Oct 2 03:40:51 2026 Received: from out-174.mta1.migadu.com (out-174.mta1.migadu.com [95.215.58.174]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0F7093F6C28 for ; Wed, 5 Aug 2026 09:21:17 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=95.215.58.174 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785921684; cv=none; b=MFHg9fIVw4BIYa8DBQoXhUYLMt1ttkS4oDpRTnY+LfsvNinq3HzodLZyOL1RDpPrtmNwtfVhreYPcPOKuGRzH2dI/JsSGfp0mIq4VGy4Il4ccw/x+2DTCrFzYYb2xb0MQaWrQ2ulRV7jwg+bbaplQuyddL2h3KuQ9vjPTz95VuU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785921684; c=relaxed/simple; bh=rihW06zt0jOo7vObDCfKiZAycA+BvyIzlcHZEkPoyUc=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=aTpENMe0TSPfXbGf/R8VO0vP8abRaW4qB3jPNrqm48zkpKnN2xntnShdpVFkBa7GWPpOaa4sMOlWhtm+jG0vcFNPt3qD6akAd9Avj/rd/O1ZM7wD7vxdFgoU5zgefmaRk+MFXcQndtXSpu1e4aH6OWxXJX1lsnO4ac97bovt6TI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=ohTCorAb; arc=none smtp.client-ip=95.215.58.174 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="ohTCorAb" X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers. DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1; t=1785921671; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=vahvZ+wmKcdNUH+2tDukG0ly53D64r43IRpsGL/ZSOo=; b=ohTCorAbGwbGa7IEdo6ou/29iah3yAmnh5kkgYKl3VlAImSs3aZY47HGkrUV4yxt4c9Q7B GYIgkJf/PuguFkXCFtc2PTKPp4k+qXZalPfBGUWvCDyvmyauPlanKGHIjXGcdy5I8f9h42 t0JuAWdKsvBMLUBvthEx+Vsyk6Q0s0Q= From: Jiayuan Chen To: bpf@vger.kernel.org Cc: Jiayuan Chen , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , Eduard Zingerman , Kumar Kartikeya Dwivedi , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Ihor Solodrai , John Fastabend , Shuah Khan , Sebastian Andrzej Siewior , Clark Williams , Steven Rostedt , linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org, linux-rt-devel@lists.linux.dev Subject: [PATCH bpf-next v2 4/4] selftests/bpf: Add a test for arena fault-in under memory.max Date: Wed, 5 Aug 2026 17:15:57 +0800 Message-ID: <20260805091720.139924-5-jiayuan.chen@linux.dev> In-Reply-To: <20260805091720.139924-1-jiayuan.chen@linux.dev> References: <20260805091720.139924-1-jiayuan.chen@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Migadu-Flow: FLOW_OUT Content-Type: text/plain; charset="utf-8" A child joins a memcg capped at 64M and faults an arena in until it runs out of the budget. Without the kernel fix the child dies with SIGSEGV on a valid arena address; with it, the child is killed by the memcg OOM killer. With the fix: serial_test_arena_memcg:PASS:child killed by signal serial_test_arena_memcg:PASS:not killed by SIGSEGV #5 arena_memcg:OK # dmesg arena_vm_fault+0x655/0xa90 Memory cgroup out of memory: Killed process 512, file-rss:67920kB Without the fix: serial_test_arena_memcg:PASS:child killed by signal serial_test_arena_memcg:FAIL:not killed by SIGSEGV: actual 11 #5 arena_memcg:FAIL # dmesg test_progs[508]: segfault at 100004025000 ... Signed-off-by: Jiayuan Chen --- .../selftests/bpf/prog_tests/arena_memcg.c | 139 ++++++++++++++++++ .../testing/selftests/bpf/progs/arena_memcg.c | 24 +++ 2 files changed, 163 insertions(+) create mode 100644 tools/testing/selftests/bpf/prog_tests/arena_memcg.c create mode 100644 tools/testing/selftests/bpf/progs/arena_memcg.c diff --git a/tools/testing/selftests/bpf/prog_tests/arena_memcg.c b/tools/t= esting/selftests/bpf/prog_tests/arena_memcg.c new file mode 100644 index 000000000000..ca039ebd3d67 --- /dev/null +++ b/tools/testing/selftests/bpf/prog_tests/arena_memcg.c @@ -0,0 +1,139 @@ +// SPDX-License-Identifier: GPL-2.0 + +#include +#include +#include +#include +#include +#include +#include +#ifndef PAGE_SIZE /* on some archs it comes in sys/user.h */ +#include +#define PAGE_SIZE getpagesize() +#endif + +#include "cgroup_helpers.h" +#include "arena_memcg.skel.h" + +#define CG_PATH "/arena_memcg" + +/* Budget the arena gets on top of whatever is already charged after load.= */ +#define ARENA_BUDGET (64 * 1024 * 1024) + +static void dump_memcg(int (*rd)(const char *, const char *, char *, size_= t)) +{ + char buf[512]; + + /* + * memory.current reads 0 once the child has left the cgroup, so it only + * carries information when dumped from the live child; memory.peak and + * memory.events survive the child and tell the story either way. + */ + if (!rd(CG_PATH, "memory.current", buf, sizeof(buf))) + fprintf(stderr, "memory.current: %s", buf); + if (!rd(CG_PATH, "memory.max", buf, sizeof(buf))) + fprintf(stderr, "memory.max: %s", buf); + if (!rd(CG_PATH, "memory.peak", buf, sizeof(buf))) + fprintf(stderr, "memory.peak: %s", buf); + if (!rd(CG_PATH, "memory.events", buf, sizeof(buf))) + fprintf(stderr, "memory.events:\n%s", buf); + fflush(NULL); /* _exit() in the child would not flush stdio otherwise */ +} + +void serial_test_arena_memcg(void) +{ + int cgroup_fd =3D -1, status; + const long ps =3D PAGE_SIZE; + char buf[64]; + pid_t pid; + + if (setup_cgroup_environment()) + return; + + cgroup_fd =3D create_and_get_cgroup(CG_PATH); + if (!ASSERT_OK_FD(cgroup_fd, "create_and_get_cgroup")) + goto out; + + /* No memory controller -> nothing to test. */ + if (read_cgroup_file(CG_PATH, "memory.current", buf, sizeof(buf))) { + test__skip(); + goto out; + } + + pid =3D fork(); + if (!ASSERT_GE(pid, 0, "fork")) + goto out; + if (pid =3D=3D 0) { + struct arena_memcg *cskel; + __u32 i, npages; + char *base; + size_t sz; + long cur; + + /* + * Do everything from the child: the arena vma is VM_DONTCOPY so + * it would not survive fork(), only the child should be under the + * limit so that a memcg OOM cannot pick test_progs, and a map is + * charged to the memcg of the task that creates it - so join + * before load. The cgroup work dir belongs to the parent that set + * the environment up, so reach it with the _parent() helpers. + * Errors are reported to the parent through the exit code, since + * ASSERT_* in a forked child does not reach it. + */ + snprintf(buf, sizeof(buf), "%d", getpid()); + if (write_cgroup_file_parent(CG_PATH, "cgroup.procs", buf)) + _exit(2); + + cskel =3D arena_memcg__open_and_load(); + if (!cskel) + _exit(3); + + base =3D bpf_map__initial_value(cskel->maps.arena, &sz); + if (!base) + _exit(4); + npages =3D bpf_map__max_entries(cskel->maps.arena); + + /* + * Cap only now, after load: everything but the fault-in is + * charged, so the arena gets a fixed budget regardless of what + * the load itself cost, and the load can never hit the limit. + */ + if (read_cgroup_file_parent(CG_PATH, "memory.current", buf, sizeof(buf))) + _exit(5); + cur =3D strtol(buf, NULL, 10); + snprintf(buf, sizeof(buf), "%ld", cur + ARENA_BUDGET); + if (write_cgroup_file_parent(CG_PATH, "memory.max", buf)) + _exit(6); + + for (i =3D 0; i < npages; i++) + base[(size_t)i * ps] =3D 1; + /* Faulted everything without dying: no pressure built, dump why. */ + dump_memcg(read_cgroup_file_parent); + _exit(0); + } + + if (!ASSERT_EQ(waitpid(pid, &status, 0), pid, "waitpid")) + goto out; + + /* A non-zero exit means the child failed to set up; the code says where.= */ + if (WIFEXITED(status) && WEXITSTATUS(status)) { + ASSERT_OK(WEXITSTATUS(status), "child setup"); + goto out; + } + + /* + * Faulting a valid arena address until memory.max is hit must not look + * like an invalid access. Without the fix the fault path allocated with + * the non-blocking allocator, turned its -ENOMEM into VM_FAULT_SIGSEGV, + * and the child died with SIGSEGV on a valid address; now it is handled + * by the memcg OOM path and the child is killed by SIGKILL instead. + */ + if (!ASSERT_TRUE(WIFSIGNALED(status), "child killed by signal")) + goto out; + if (!ASSERT_NEQ(WTERMSIG(status), SIGSEGV, "not killed by SIGSEGV")) + dump_memcg(read_cgroup_file); +out: + if (cgroup_fd >=3D 0) + close(cgroup_fd); + cleanup_cgroup_environment(); +} diff --git a/tools/testing/selftests/bpf/progs/arena_memcg.c b/tools/testin= g/selftests/bpf/progs/arena_memcg.c new file mode 100644 index 000000000000..adecd9e8463e --- /dev/null +++ b/tools/testing/selftests/bpf/progs/arena_memcg.c @@ -0,0 +1,24 @@ +// SPDX-License-Identifier: GPL-2.0 + +#include +#include +#include "bpf_arena_common.h" + +struct { + __uint(type, BPF_MAP_TYPE_ARENA); + __uint(map_flags, BPF_F_MMAPABLE); + __uint(max_entries, 100000); /* number of pages */ +#ifdef __TARGET_ARCH_arm64 + __ulong(map_extra, 0x1ull << 32); /* start of mmap() region */ +#else + __ulong(map_extra, 0x1ull << 44); /* start of mmap() region */ +#endif +} arena SEC(".maps"); + +SEC("syscall") +int noop(void *ctx) +{ + return 0; +} + +char _license[] SEC("license") =3D "GPL"; --=20 2.43.0