From nobody Mon Sep 28 07:17:28 2026 Received: from mail-pg1-f172.google.com (mail-pg1-f172.google.com [209.85.215.172]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id ED92C3EEAEF for ; Tue, 25 Aug 2026 09:20:04 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.172 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787649608; cv=none; b=JIkjsykrnJZ4S1lLoAJTCVooEYUHlZopgGLTMjauqRJvMSO9fieGSP+o0ee9dntrv1KKvpFja/buTyZriqKJQ+67qHmF/dWOuU+A9WXjjtBD6HJeIpMWtO0TdyjRaiYLPGDH1sY/aIbpFQKF5YMf8KSDFkH1cj+yrl2Au8YmfG8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787649608; c=relaxed/simple; bh=PyE160OajPSsqQd9ddyOfvbdMpeClWf6VPakHByI8uA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=sdN8PFQTM2oD9GsYk/Se8ntwluESrc3Fn697YzaMampSAcxJB7Mb8LgzuRWACShw6nQgKDUhtQBEzdrYjwC0gLPcjCYXOVW4d0kEMpKp3Wfh+6RRHw6TQC0h4ML0Q0JxUkfFGEdTX5TtUIf40YQicyBvQ3lJPsBTO2Sb+yuTHIY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=hO7OLmw2; arc=none smtp.client-ip=209.85.215.172 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="hO7OLmw2" Received: by mail-pg1-f172.google.com with SMTP id 41be03b00d2f7-c9d1fff21edso3744590a12.1 for ; Tue, 25 Aug 2026 02:20:04 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787649604; x=1788254404; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=W7cSIyxl/i1oxSUvdnacr9V24YcqM8fZNDWiyKdu00s=; b=hO7OLmw2zO6P3bvQ9LbCNoSaj8aX03efx43YDK+GS4L3ZM5cjFcvsaSpANllFuAfy8 bWVMt8A5NyMNcu8hie3uU33KDWUoToQ16K04Rw6kx6nuCDw/zg6e3SzinSiO1NnqEVEp vejLxCKCrwKcp6mAQdkyWeEgtL6a16msKngy3Y+sv0jRYJe90vbfg3cpCO1ZAMBhBBem ALlTkOys0UxXgZpikAD4opGx/kDkeNTIwanQJGnSpChEpGKmVLivRH0qiqgSFQBxaa2C TQIB8F52HULgc6mD6a5+T9rb9Wnf40nwTF65GdfiQ1iRnjU0zX43UZmAcfx4vzmelswb A/Rg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787649604; x=1788254404; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=W7cSIyxl/i1oxSUvdnacr9V24YcqM8fZNDWiyKdu00s=; b=e0/73T/cxBgO8cJRwuyIRzL9bMwAEK7Hs/9Wu3613ne9Hp1+H3xt7C31zIZdtndThs UHMXYm2jk5rz+9zmWIWVPXR1fIY9aSQN9Aepx1o8HBUIbjx10DD4D+VxV03a4W4O3V2u 8i0DxVLjGr5mwOB2yq+enOB12gSE7QZv3133IQjBOILc+myABJhXGs01q0H6SodITxpB nY85RzrqcEni4DGMS2K1Fvnr8B2evEqKHTFZxRxRdYh930w25KtzZEI4mP5hV/2qAybS 7IXOz//VB57yHdOXb4nvzJBTS6QRKQ8mQotUkoLm3pDQSoAa27vGWiGWLSjI18+m0q/S ikFQ== X-Forwarded-Encrypted: i=1; AHgh+RrNuQsgTu7maNALxUYiUZQ3uio0diJOTdqsMvkpQyKHVGRejJDlRCTx7pmvplXa+iYQjaNst+1SvX64XYs=@vger.kernel.org X-Gm-Message-State: AFuF++nDQMtC24qmM7bcxbRqL4aPrp7eipSYF1D6lrWV5UBi5ovrM/mF CXe6XTlEijT31K2N1B3EiPLX0CJ44dQvspwdtj0onUH31WbIpZTbATou X-Gm-Gg: AR+sD13N6j6EPPKupfOIGBCtKgx7W+YWZ0hzL1rvfOUVDKjnLz/Zu6TKQk47LLiSJdB Ar8B8H6JShKXFNH0N+81PqHPJJ4WyBKRWWMLz5s6+IwbJlvMtY+/ABR5s1RwWJAKNU94wEKb8R0 QE8hqPf8KRKAwfDk05AmeyFoI4e1uEsaZy+IJGZE9nIIzDKbKQ87N1lXumz9teryFRnZG0SriEs T1CnuVQoSW+ORxb2p/rxiCCiP4xX9M1m9O6WObqQ0nNfFmhyVQDbM1O/4EWVfYBwifYj+VFy7qS ZY9tNLC13ErcVxB4xrV/66gWsyTNJk0Ee8NZgTW0x5L43skLjzbupAU5acX8Udefsj546XvWZ3G 2mGWgpunBMf4Mq/gImP64aOx1Mh3RVvONDfwZOgYAKfSpX09CdGIAalfBipDg7xRewmwYzWVg5V 4TMea1pp9eO8L/KIPP4JgVcBEjcVWvB7jGLQHyV60cOCr9t9KyJhc5QBj2tzH4DXDDb/5cI/hGE xNDp45QdK29nJAuFnw= X-Received: by 2002:a05:6a21:7482:b0:3c3:b57b:6291 with SMTP id adf61e73a8af0-3cd301b5175mr69489021637.18.1787649603514; Tue, 25 Aug 2026 02:20:03 -0700 (PDT) Received: from localhost.localdomain ([103.120.31.178]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-3282afb407bsm5386163eec.30.2026.08.25.02.19.59 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Tue, 25 Aug 2026 02:20:02 -0700 (PDT) From: Khawar Ahemad To: bpf@vger.kernel.org, linux-kernel@vger.kernel.org Cc: ast@kernel.org, daniel@iogearbox.net, andrii@kernel.org, eddyz87@gmail.com, jiayuan.chen@linux.dev, emil@etsalapatis.com Subject: [PATCH v6 1/4] bpf: Add a sleepable page allocator for map memory Date: Tue, 25 Aug 2026 14:49:49 +0530 Message-ID: <20260825091952.81971-2-ahemadkhawar123@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260825091952.81971-1-ahemadkhawar123@gmail.com> References: <20260825091952.81971-1-ahemadkhawar123@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Jiayuan Chen bpf_map_alloc_pages() picks the allocator via can_alloc_pages(), a conservative guess for BPF program context that is always false under PREEMPT_RT. So even a caller that really is sleepable gets the non-blocking allocator, which never reclaims and never engages the OOM machinery. Add bpf_map_alloc_page_sleepable() for callers that know they are sleepable. Like the other bpf map allocators it places the page on the map's numa_node and does not follow the faulting task's NUMA mempolicy; arena memory is shared, so the map's node is the right placement policy. The next patch uses it from the arena page fault handler. Signed-off-by: Jiayuan Chen Reviewed-by: Emil Tsalapatis Signed-off-by: Khawar Ahemad --- include/linux/bpf.h | 1 + kernel/bpf/syscall.c | 21 +++++++++++++++++---- 2 files changed, 18 insertions(+), 4 deletions(-) diff --git a/include/linux/bpf.h b/include/linux/bpf.h index b3cd28d9e3..c817c99d29 100644 --- a/include/linux/bpf.h +++ b/include/linux/bpf.h @@ -2784,6 +2784,7 @@ struct bpf_prog *bpf_prog_get_curr_or_next(u32 *id); =20 int bpf_map_alloc_pages(const struct bpf_map *map, int nid, unsigned long nr_pages, struct page **page_array); +struct page *bpf_map_alloc_page_sleepable(const struct bpf_map *map); #ifdef CONFIG_MEMCG void bpf_map_memcg_enter(const struct bpf_map *map, struct mem_cgroup **ol= d_memcg, struct mem_cgroup **new_memcg); diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c index 6874ba1424..f9b81638e5 100644 --- a/kernel/bpf/syscall.c +++ b/kernel/bpf/syscall.c @@ -602,15 +602,14 @@ static bool can_alloc_pages(void) !IS_ENABLED(CONFIG_PREEMPT_RT); } =20 +#define BPF_PAGE_GFP (GFP_KERNEL | __GFP_ZERO | __GFP_ACCOUNT | __GFP_NOWA= RN) + static struct page *__bpf_alloc_page(int nid) { if (!can_alloc_pages()) return alloc_pages_nolock(__GFP_ACCOUNT, nid, 0); =20 - return alloc_pages_node(nid, - GFP_KERNEL | __GFP_ZERO | __GFP_ACCOUNT - | __GFP_NOWARN, - 0); + return alloc_pages_node(nid, BPF_PAGE_GFP, 0); } =20 int bpf_map_alloc_pages(const struct bpf_map *map, int nid, @@ -636,6 +635,20 @@ int bpf_map_alloc_pages(const struct bpf_map *map, int= nid, return ret; } =20 +/* + * For callers that know they run in a sleepable context, e.g. a user page + * fault handler. can_alloc_pages() is a conservative guess made for BPF + * program context - notably it is always false on PREEMPT_RT - so going + * through bpf_map_alloc_pages() there would needlessly pick the + * non-blocking allocator, which never reclaims and never engages the OOM + * machinery. + */ +struct page *bpf_map_alloc_page_sleepable(const struct bpf_map *map) +{ + might_sleep(); + return alloc_pages_node(map->numa_node, BPF_PAGE_GFP, 0); +} + static int btf_field_cmp(const void *a, const void *b) { const struct btf_field *f1 =3D a, *f2 =3D b; --=20 2.54.0 (Apple Git-157) From nobody Mon Sep 28 07:17:28 2026 Received: from mail-pg1-f177.google.com (mail-pg1-f177.google.com [209.85.215.177]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DAC123EF0C8 for ; Tue, 25 Aug 2026 09:20:08 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.177 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787649610; cv=none; b=gVVWVcfMLEt56Tt/Fspdw56hZCtvvt/rz/jDc+cvB2TlgEooAfjz4Aszos6Ti9cn0EPQy22HPKtyyUiGdsrEHJd8HGnljuqsaH8d928VJHuIM3ZHRjf2pHUoNLEgsQLXSzb0rB1dnckoUHZL6KGhLlQ5BzrDpj3n9odWXgz7jwQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787649610; c=relaxed/simple; bh=nCIKMze/nePG9edrgFsirG3py5omd/Tnqbe3SwzpmlM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=K7e22B2cig0HFtiq5wGYgro78RKsptbs5Oxx0e8DANt/GI16pOXmhr9H83q03fC1lblN/lJVysMpcCsnbz57j0dxHCyzqAtuvRGObs9U5NCHtP/uABgJScxKPMangkyryr//Puco9qAShPbFRnvgYO3mpbYIfumLX2zjPG71itY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=kcWjVjWz; arc=none smtp.client-ip=209.85.215.177 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="kcWjVjWz" Received: by mail-pg1-f177.google.com with SMTP id 41be03b00d2f7-cbedbaba5fdso2705215a12.0 for ; Tue, 25 Aug 2026 02:20:08 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787649608; x=1788254408; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=XGkExxMyhBiKuGkp4+FT/b5oczqhluEeBce3FhXEQbg=; b=kcWjVjWzA3HGaCKocpjGloK4RirvwRX8mz+rtXE2T/Aa1kt2LGaX+zWUgcfjLpIBov lHidaxPX9EVxzW1U+GMAMcBLHgpsd2oaQFjlJCcwuaOosNzwKp5KyEHPoCYJDxKtIF7N Za9tpHZbC2TxTrV0t9oppaA1IDdLNhWr+XNTsy4v4piD9uLLGm4pyGJXsmEO7RfIja0D ay790K610n+wCiwxw7jHmdS07RVLtLJqPPy68PMqRp6KHowt+rOmG64McaXEbhCf3vjT B5UjuccoPKtpeGuWuQajOtFvjqmEE5fpC9SXQWPFAgd/LBIdnICkUZB2I6U96icvXtn5 MI0Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787649608; x=1788254408; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=XGkExxMyhBiKuGkp4+FT/b5oczqhluEeBce3FhXEQbg=; b=rfMEsr3Snvra1MfbQU20tL12sV8hU9zX917cfsgq5W/g7uukvyuiP2k39/L50aZjlC IvsoFR5h+8uEcpD/xu5B53jDOGF7WPU6IzWBwIid/zXujKUOUzpsKbFAFBm7hzI+UIMd dIeVbjJ06XVHygk+phiu9unfDpyjFHY1CzWMkncXAyS6SWrbPSqDGvu/HaTj0BbjSvAj iDHBcdr86fySeHdfM6lvUCdGADI/+qDG5yumg7YGHPxqwN7/bisq39EHtk6oDinq4cCm wvn5qg/U7churAv5fqbzQnqs6UPKHWYFudCWyuDgSsO2ZFuxh/V9bpONP7n+rFgBS9rn Jq4A== X-Forwarded-Encrypted: i=1; AHgh+RrVTITMSH1IvnLMwwQv8AD+WKo3dcwRtmXBQygmIWBD1DjzMy21BLFmPVTJLmwEg6s1kDEgWCojdgfs5nI=@vger.kernel.org X-Gm-Message-State: AFuF++mA6X9TVut+Zqp+ojqk9bUrq1KVC/S9cXmf6+FNgR/fyzrphxQj 06qupfMSOrbUrL7kFXDmL/+5fvVo/xfnitK07QHTL+zdTpZimbcZ568/Drc03t8bRMM= X-Gm-Gg: AR+sD11PHBGyyvGJZbuj1h6xQW3kNZfzV2G4i8L4z9+0zKuNEJ+kgxPIWvgw1WfTmnR UDUSZcOq6p1lfyzyXd8FZS81d0G5MOI5FdBSlbooPcqRs3aBcAGlFApZYOxlinQTlAcko+ODq1q c4U+CEMRp5345e+8MWyVi0XBG8Bd4OL56Qppx09wNhWUkZxqF2ydfLY0WPU8p5L+6xfXrSdU/XW NPi5fhRKGHZnaBecPoJlt9xGMydPFCtXK4wCRhY0JCi4HQEOUXOH080Xeeq8YcIqYSEPuBhgMzf Jb9n4lUBhsXoM8kMZrqFQidEvT09/aSCsesSwpvbCDvg7vXRJzxRgypOdluGI0JHTpMrVICB1Y6 XyB/sFAkSypx+Q6UREjzJ83hml2PySwqbJUB1sKfeSrs6SlLXTTydSMhW9y/Nk/A0/IL2yhfnPC X9B8j8scq/Q1cCUg4GcKk6V9x3ZPGcQEM66lqa4Mu9YCEW2CDzTTKJyXV+ILxpknm9IPR3AoHIb G6kAtsFVjrLHR9/jm5iJIEQ6i8g/g== X-Received: by 2002:a05:6a20:c91c:b0:3b4:8880:2089 with SMTP id adf61e73a8af0-3cd301784afmr65940319637.16.1787649607933; Tue, 25 Aug 2026 02:20:07 -0700 (PDT) Received: from localhost.localdomain ([103.120.31.178]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-3282afb407bsm5386163eec.30.2026.08.25.02.20.03 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Tue, 25 Aug 2026 02:20:07 -0700 (PDT) From: Khawar Ahemad To: bpf@vger.kernel.org, linux-kernel@vger.kernel.org Cc: ast@kernel.org, daniel@iogearbox.net, andrii@kernel.org, eddyz87@gmail.com, jiayuan.chen@linux.dev, emil@etsalapatis.com Subject: [PATCH v6 2/4] bpf: arena: allocate the fault-in page outside the lock Date: Tue, 25 Aug 2026 14:49:50 +0530 Message-ID: <20260825091952.81971-3-ahemadkhawar123@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260825091952.81971-1-ahemadkhawar123@gmail.com> References: <20260825091952.81971-1-ahemadkhawar123@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Jiayuan Chen arena_vm_fault() allocated the page while holding arena->spinlock, so it could only use the non-blocking allocator. Once the memcg is at memory.max that allocation just fails, the fault turns into VM_FAULT_SIGSEGV, and the process gets a SIGSEGV on a perfectly valid arena address. Hitting memory.max is routine (e.g. page cache from reading a big file), so this kills innocent processes. Rework the fault handler: - Preallocate the page before taking the lock, like do_anonymous_page() does, so it can sleep and go through reclaim and the memcg OOM killer, instead of turning a routine memory.max into a fake segfault. - On allocation failure return VM_FAULT_SIGBUS. The allocation already ran reclaim and the OOM killer, so the failure is non-recoverable. For a task faulting its own arena this changes nothing: the OOM killer already picked it inside the allocation and it dies by SIGKILL, the SIGBUS is shadowed by the pending fatal signal, and the memcg OOM is still reported. VM_FAULT_OOM would instead be retried by the fault path and can livelock when the charged memcg is not the faulting task's (e.g. a shared arena) and its OOM killer cannot reach it. - A lockless probe skips that preallocation when a page is already mapped (e.g. allocated by the bpf program), so the common case wastes no allocation. The rare race where such a page is freed before we take the lock falls back to the non-blocking allocator under the lock. - Return VM_FAULT_SIGBUS for the other non-recoverable errors (lock failure, range-tree and page-table failures) instead of VM_FAULT_SIGSEGV; only BPF_F_SEGV_ON_FAULT, and a scratch-page hole under that flag, is a real user addressing error and keeps VM_FAULT_SIGSEGV. - Tidy up the error labels. Reviewed-by: Emil Tsalapatis Signed-off-by: Jiayuan Chen Signed-off-by: Khawar Ahemad --- kernel/bpf/arena.c | 90 +++++++++++++++++++++++++++++++++++----------- 1 file changed, 69 insertions(+), 21 deletions(-) diff --git a/kernel/bpf/arena.c b/kernel/bpf/arena.c index 7b6847200b..fa462a0ff1 100644 --- a/kernel/bpf/arena.c +++ b/kernel/bpf/arena.c @@ -481,7 +481,8 @@ static vm_fault_t arena_vm_fault(struct vm_fault *vmf) struct bpf_map *map =3D vmf->vma->vm_file->private_data; struct bpf_arena *arena =3D container_of(map, struct bpf_arena, map); struct mem_cgroup *new_memcg, *old_memcg; - struct page *page; + struct page *page, *new_page =3D NULL; + vm_fault_t fault_ret; long kbase, kaddr; unsigned long flags; int ret; @@ -489,59 +490,106 @@ static vm_fault_t arena_vm_fault(struct vm_fault *vm= f) kbase =3D bpf_arena_get_kern_vm_start(arena); kaddr =3D kbase + (u32)(vmf->address); =20 - if (raw_res_spin_lock_irqsave(&arena->spinlock, flags)) + page =3D vmalloc_to_page((void *)kaddr); + if (!page && !(arena->map.map_flags & BPF_F_SEGV_ON_FAULT)) { + /* + * Preallocate outside the lock with a sleepable allocator so it + * can reclaim and run the memcg OOM killer, which the + * non-blocking allocator under arena->spinlock cannot. A NULL + * return is non-recoverable, so fail with VM_FAULT_SIGBUS; + * VM_FAULT_OOM would be retried by the fault path and can + * livelock when the charged memcg is not the faulting task's. + */ + bpf_map_memcg_enter(&arena->map, &old_memcg, &new_memcg); + new_page =3D bpf_map_alloc_page_sleepable(map); + bpf_map_memcg_exit(old_memcg, new_memcg); + if (!new_page) + return VM_FAULT_SIGBUS; + } + + if (raw_res_spin_lock_irqsave(&arena->spinlock, flags)) { /* * A failed lock means a possible deadlock was detected. Don't * return VM_FAULT_RETRY: this handler never took mmap_lock, but * the fault path would re-take it on retry and deadlock. Fail. */ + if (new_page) + free_pages_nolock(new_page, 0); return VM_FAULT_SIGBUS; + } =20 page =3D vmalloc_to_page((void *)kaddr); if (page) { - if (page =3D=3D arena->scratch_page) - /* BPF triggered scratch here; don't lazy-alloc over it */ - goto out_sigsegv; + if (page =3D=3D arena->scratch_page) { + /* + * A scratch page marks a hole. Segfault only if the user + * asked for it; otherwise we could lazy-allocate but + * choose not to over a hole, so report a bus error. + */ + fault_ret =3D (arena->map.map_flags & BPF_F_SEGV_ON_FAULT) ? + VM_FAULT_SIGSEGV : VM_FAULT_SIGBUS; + goto out_err_locked; + } /* already have a page vmap-ed */ goto out; } =20 + if (arena->map.map_flags & BPF_F_SEGV_ON_FAULT) { + /* User space requested to segfault when page is not allocated by bpf pr= og */ + fault_ret =3D VM_FAULT_SIGSEGV; + goto out_err_locked; + } + bpf_map_memcg_enter(&arena->map, &old_memcg, &new_memcg); =20 - if (arena->map.map_flags & BPF_F_SEGV_ON_FAULT) - /* User space requested to segfault when page is not allocated by bpf pr= og */ - goto out_sigsegv_memcg; + if (!new_page) { + /* + * Very rare race: the bpf program had allocated a page here, so + * the lockless probe saw it and we skipped preallocation, but it + * freed the page before we took the lock. Now we do need one; + * sleeping is not allowed here, so fall back to the non-blocking + * allocator and give up if it fails. + */ + ret =3D bpf_map_alloc_pages(map, map->numa_node, 1, &new_page); + if (ret) { + fault_ret =3D VM_FAULT_SIGBUS; + goto out_err_locked_memcg; + } + } =20 ret =3D range_tree_clear(&arena->rt, vmf->pgoff, 1); - if (ret) - goto out_sigsegv_memcg; - - struct apply_range_data data =3D { .arena =3D arena, .pages =3D &page, .i= =3D 0 }; - /* Account into memcg of the process that created bpf_arena */ - ret =3D bpf_map_alloc_pages(map, NUMA_NO_NODE, 1, &page); if (ret) { - range_tree_set(&arena->rt, vmf->pgoff, 1); - goto out_sigsegv_memcg; + fault_ret =3D VM_FAULT_SIGBUS; + goto out_err_locked_memcg; } + struct apply_range_data data =3D { .arena =3D arena, .pages =3D &new_page= , .i =3D 0 }; =20 ret =3D apply_to_page_range(&init_mm, kaddr, PAGE_SIZE, apply_range_set_c= b, &data); if (ret) { range_tree_set(&arena->rt, vmf->pgoff, 1); - free_pages_nolock(page, 0); - goto out_sigsegv_memcg; + fault_ret =3D VM_FAULT_SIGBUS; + goto out_err_locked_memcg; } flush_vmap_cache(kaddr, PAGE_SIZE); bpf_map_memcg_exit(old_memcg, new_memcg); + /* new_page was consumed */ + page =3D new_page; + new_page =3D NULL; out: page_ref_add(page, 1); raw_res_spin_unlock_irqrestore(&arena->spinlock, flags); + if (new_page) + free_pages_nolock(new_page, 0); vmf->page =3D page; return 0; -out_sigsegv_memcg: + +out_err_locked_memcg: bpf_map_memcg_exit(old_memcg, new_memcg); -out_sigsegv: +out_err_locked: raw_res_spin_unlock_irqrestore(&arena->spinlock, flags); - return VM_FAULT_SIGSEGV; + if (new_page) + free_pages_nolock(new_page, 0); + return fault_ret; } =20 static const struct vm_operations_struct arena_vm_ops =3D { --=20 2.54.0 (Apple Git-157) From nobody Mon Sep 28 07:17:28 2026 Received: from mail-pj1-f49.google.com (mail-pj1-f49.google.com [209.85.216.49]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3DDBE3EBF07 for ; Tue, 25 Aug 2026 09:20:13 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.49 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787649614; cv=none; b=MK43qdHP1tUanM1BhwqnhKvo1K2+83jv4KDfcSD9gGNONsjJ82OaQxajR1aI2tadqp4HR7Ee25HtBfUB4KUFU7Y41Q7Wws2K54IuqKhvaD7cwFbmYZEhsfHb84S//nxOZ6sIcIiLUBgy7nzOCUjfjXoM9NtQTd5fqhIKtWX77bk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787649614; c=relaxed/simple; bh=/yvTesosUitvxpBGbkVrq+T3ZwnrBfxMu1Tk0/tPvCI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=oPs0MaYot2Sm432gR2QugyWfOG/pQB7/HDoq4vkeuk/Hhypug2v2qI1mcnT5p2w/w3t9sxc5rHegZbD2t8kRTAsiK6ST/C5ijlYNEgfdQOmH34X1soZ6tfJmku3ti1XIK336gauCggVt/LgRrvdew2zmQ9L0ZeTuXpGh0ar/KRE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=V0iie5TW; arc=none smtp.client-ip=209.85.216.49 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="V0iie5TW" Received: by mail-pj1-f49.google.com with SMTP id 98e67ed59e1d1-3900e39d935so4927837a91.0 for ; Tue, 25 Aug 2026 02:20:12 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787649612; x=1788254412; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=0rBrebsEU++q5oyTmtoojXSPfykwosiXDRrVD5SURpU=; b=V0iie5TW7J64I2H/GdYFujiQlMqRe+ajXUvwGW2hSpngsp8mtQWakparn/E4JuDzxi PXZfwqVOBvggzCELl1nnjMXQw29EZMK65JKrx+e/aO+AANwfqa3j219qupcpU+fl6aH1 6OtIhJe+zwzKbJUNmjo5noMX7pXSoKm0nG2VI4YfQHL6dOwRpiP4yZBTUIZmMFDk5EHV +8R3HVjB4syDvKaOTNILBAWnYfMK7JiAD0d1kUzhaOD6OObIKZ2sHOG3YHGv9ucxAGd8 URe9tdLiIH3FJAJ9fCU7JNWfIXJ743JEfncyVDt8TOZxOKueIBFO4EoJo40GKSJotgGY 14RQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787649612; x=1788254412; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=0rBrebsEU++q5oyTmtoojXSPfykwosiXDRrVD5SURpU=; b=o3G8NjLiq20Rskwnw58VrOZzbyE+4Rldw7LAbyuKtRdpPIMHsMAzU6QtedFrYHQZyv J/0ag7y6xx1I8/QAH3MVtNfPLJdlmoWMfUKzN1zGdzjjB/00CPln2RA4/CnqJwWXA5Db xaZ1DHJXtVEc8nPXxZGANkvyQNtsKYBLxUJfmAnyWdiWsXGv0nEUAbGHkXa32U9J9oPU CUlkOWBtaGf3IOSrqt/G32sgjM6kqaN95RpsHCxRjOm7iYNVnfRc9zfDqysErIseWohV QjAW3S38yBnr4SoED06o+vHk01cpn/iuxEkfE+T0/1G+oEs4hIHteSWxiisn53CFKoh0 c1qg== X-Forwarded-Encrypted: i=1; AHgh+Rq4mTBHfVTY3uRE77AA042Ga/j6qHfEc9bYhUbg7tJIvtOzMF3zHIh9eP5koIfZyL5lSLBNGVrmVRqEUPU=@vger.kernel.org X-Gm-Message-State: AFuF++kzIvaXOdGHD7nuReEzwdR0t4jNupUiDUmKpyMeKv7g7MEKmMoV txFecjdngLX6GVtXd5Y1/lMTFrFi6MahWivrf6c/2H1IWIL/niJXjEla X-Gm-Gg: AR+sD11DvFdyiScINpFuxyq9FSM8KvbFdIvdJN3Rzd3TCOc3ozRccZVkZrDQJaNGw5i O9Z3ZdTYhDyOkjyCgAlgzPIeUNCADcpgmhzZkONX79L5Zc4DBUqCHCWfiUpGJn2FocbUytKxbC4 jugshYnZeZjI9NCOCZloECKPlr7NxSAnsJqgemEcsKwcZ21cIiicEHWuffFPNcdddsRL/RuLku0 Tsp51X9z90q8RmWHjeIYkuIJVr2WvuXyKiGVFD5ocVTUhMiXenmZkHLwCvNJg9xHeLeNf+tDlx6 FAG8px/z6plo3GWsLmfOdxEp4NfdoC2eDOhlE5gE4cDO/qKplqZNs7p82CUnN92LOH+G1Miv6Ib Ua2I/bQzWRnBgxSO1bzrU8PCG+UqmFXOeVuTEiop1yFqyyk0Xy910CF0KZvIKjooZH4ZhdNQXKd TZHOyeFqXmFKc7fWiaP1z8xCJwMeloBUPLEjUj72RNCGJAKOrxExFG1TSkRJaayhJlX2lUnqciL AEKZGzTXG7NI5Ivk2Ydq5yWC2Rwwg== X-Received: by 2002:a17:90b:580d:b0:38f:5869:387b with SMTP id 98e67ed59e1d1-395df24e757mr49635358a91.9.1787649612347; Tue, 25 Aug 2026 02:20:12 -0700 (PDT) Received: from localhost.localdomain ([103.120.31.178]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-3282afb407bsm5386163eec.30.2026.08.25.02.20.08 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Tue, 25 Aug 2026 02:20:11 -0700 (PDT) From: Khawar Ahemad To: bpf@vger.kernel.org, linux-kernel@vger.kernel.org Cc: ast@kernel.org, daniel@iogearbox.net, andrii@kernel.org, eddyz87@gmail.com, jiayuan.chen@linux.dev, emil@etsalapatis.com Subject: [PATCH v6 3/4] selftests/bpf: Add read_cgroup_file() to cgroup_helpers Date: Tue, 25 Aug 2026 14:49:51 +0530 Message-ID: <20260825091952.81971-4-ahemadkhawar123@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260825091952.81971-1-ahemadkhawar123@gmail.com> References: <20260825091952.81971-1-ahemadkhawar123@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Jiayuan Chen cgroup_helpers has write_cgroup_file()/write_cgroup_file_parent() but no read counterpart. Add read_cgroup_file() and read_cgroup_file_parent() so a forked child can read a cgroup file (e.g. memory.current) from the work dir owned by the parent that set the environment up, without hand-building the /mnt/... path. Reviewed-by: Emil Tsalapatis Signed-off-by: Jiayuan Chen Signed-off-by: Khawar Ahemad --- tools/testing/selftests/bpf/cgroup_helpers.c | 67 ++++++++++++++++++++ tools/testing/selftests/bpf/cgroup_helpers.h | 4 ++ 2 files changed, 71 insertions(+) diff --git a/tools/testing/selftests/bpf/cgroup_helpers.c b/tools/testing/s= elftests/bpf/cgroup_helpers.c index 45cd0b479f..4183ff6150 100644 --- a/tools/testing/selftests/bpf/cgroup_helpers.c +++ b/tools/testing/selftests/bpf/cgroup_helpers.c @@ -188,6 +188,73 @@ int write_cgroup_file_parent(const char *relative_path= , const char *file, return __write_cgroup_file(cgroup_path, file, buf); } =20 +static int __read_cgroup_file(const char *cgroup_path, const char *file, + char *buf, size_t len) +{ + char file_path[PATH_MAX + 1]; + ssize_t got; + int fd; + + snprintf(file_path, sizeof(file_path), "%s/%s", cgroup_path, file); + fd =3D open(file_path, O_RDONLY); + if (fd < 0) { + log_err("Opening %s", file_path); + return 1; + } + + got =3D read(fd, buf, len - 1); + if (got < 0) { + log_err("Reading %s", file_path); + close(fd); + return 1; + } + buf[got] =3D '\0'; + close(fd); + return 0; +} + +/** + * read_cgroup_file() - Read from a cgroup file + * @relative_path: The cgroup path, relative to the workdir + * @file: The name of the file in cgroupfs to read from + * @buf: Buffer to read into, NUL-terminated on success + * @len: Size of @buf + * + * Read from a file in the given cgroup's directory. + * + * If successful, 0 is returned. + */ +int read_cgroup_file(const char *relative_path, const char *file, + char *buf, size_t len) +{ + char cgroup_path[PATH_MAX - 24]; + + format_cgroup_path(cgroup_path, relative_path); + return __read_cgroup_file(cgroup_path, file, buf, len); +} + +/** + * read_cgroup_file_parent() - Read from a cgroup file in the parent proce= ss + * workdir + * @relative_path: The cgroup path, relative to the parent process workdir + * @file: The name of the file in cgroupfs to read from + * @buf: Buffer to read into, NUL-terminated on success + * @len: Size of @buf + * + * Read from a file in the given cgroup's directory under the parent proce= ss + * workdir. + * + * If successful, 0 is returned. + */ +int read_cgroup_file_parent(const char *relative_path, const char *file, + char *buf, size_t len) +{ + char cgroup_path[PATH_MAX - 24]; + + format_parent_cgroup_path(cgroup_path, relative_path); + return __read_cgroup_file(cgroup_path, file, buf, len); +} + /** * setup_cgroup_environment() - Setup the cgroup environment * diff --git a/tools/testing/selftests/bpf/cgroup_helpers.h b/tools/testing/s= elftests/bpf/cgroup_helpers.h index 3857304be8..d42d2e1304 100644 --- a/tools/testing/selftests/bpf/cgroup_helpers.h +++ b/tools/testing/selftests/bpf/cgroup_helpers.h @@ -15,6 +15,10 @@ int write_cgroup_file(const char *relative_path, const c= har *file, const char *buf); int write_cgroup_file_parent(const char *relative_path, const char *file, const char *buf); +int read_cgroup_file(const char *relative_path, const char *file, + char *buf, size_t len); +int read_cgroup_file_parent(const char *relative_path, const char *file, + char *buf, size_t len); int cgroup_setup_and_join(const char *relative_path); int get_root_cgroup(void); int create_and_get_cgroup(const char *relative_path); --=20 2.54.0 (Apple Git-157) From nobody Mon Sep 28 07:17:28 2026 Received: from mail-pl1-f170.google.com (mail-pl1-f170.google.com [209.85.214.170]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8D8C13F0757 for ; Tue, 25 Aug 2026 09:20:17 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.170 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787649619; cv=none; b=T9A1dp8xHMYSEs7qZLG6hF95wwon0DTlwhM8qp9F8ZPMyrNS+XffQBt16vLmPG5/hOcdmn4fEef/0pf/+tIzbcbwXg4CEe0Q/19zgZeiNXuUzo/aCAoMdQgNEQzeBcw0eM27284bCAHlrA+pRJIprz2Kq89cCINqBe1uNp9LI5Q= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787649619; c=relaxed/simple; bh=NnD+lhD9+A+o9OkKxEE5WkpWKL8Rx0Ut6UVR7teDV8U=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=ODP+rwVHLcwY7aAa+J4VC7cyOLCNelstWeKQ3RjkbBQNtjfS0IdDbSet3PlNuxirSUMBgfBebD18gyyuh/zGDL5tfp7mLWinK89c1WGiP+Y3KA3oHivsBCVsaFBi6Z0xLIJIJGO7ETaZvr0WtvX8GdCaiz+r+7pUqqx7Pefwpnw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=d2cub8tt; arc=none smtp.client-ip=209.85.214.170 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="d2cub8tt" Received: by mail-pl1-f170.google.com with SMTP id d9443c01a7336-2caed617615so44519505ad.3 for ; Tue, 25 Aug 2026 02:20:17 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787649617; x=1788254417; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=XNBcnjuWFLDFOhwY/VldMF42J5EyV88DFjmxmI2ZQhM=; b=d2cub8ttToy8NBWMprf0BmAlauvwpMeh/JqGIi0HCfY5csuLKEhn8nWdU44jOW7hPK nLDj+z2HksyXFPqqAhIl9P6T0SokF/yXP5hsfnU3JY+2o304cOPS2PwTMGDeYK34N+cT Py3Px0Kii/GP9Mz2Sl4Jf7+tm0zTVPfOB7XSFWUvrOfFTgLQbvBjqjJN986TRxSAo85i ieMVwQtl4gwBlGD0SGmuX8OdoaIWGAtdbVTq17sXhwbEIRaTUfRl7KBBxIfArUD9RgAD pNGXYIQuLSDEt3yaHFD+/p1GEN7wgsBwOzv9JyUXDoA0vbY/yhjRkPk9GXzSUVbbfCD5 1KGA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787649617; x=1788254417; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=XNBcnjuWFLDFOhwY/VldMF42J5EyV88DFjmxmI2ZQhM=; b=kRHUIOxjb0Nzwn4efajx70fgaWAeBO4oVr5cjtxdFQKN24FCh4qEPcdpTOxi3jJWIC h60UawsD3Z1Q9dn245b9HRQw//QdTkJJoKCEhJMJHRmuBt8qHUjmNd2nqEdH7c4egwqT zjHR1LKvJZMiq6TTu02ixi6vumxxUQjuOzqmGGa4BasvBDzFUcv7bsf4hXMMMSWVFkdC lc2YjMHujZrlbt3t4oI/5VFrkn2ivBg/8qCIGOLz+Omz+cynupBy1ObwdsG+hvEMLkaJ mYwuG6V284tMYEZZ2QjQXoJ5VJZroqV8TguEvkpfxh/AOLX+V2uKtBq6Xv4kxYCKhoxn pezA== X-Forwarded-Encrypted: i=1; AHgh+RqDUqYms5ocQ5jgu8uF0LrE92Nw4g+EHUp65+AL3o8n8Mo0chVK33D9MicdG/BB4v9GYRlZaoqG7BNuodo=@vger.kernel.org X-Gm-Message-State: AFuF++nnQqPdR2ShgYPL+4U4YUCEFG83wDE9G9nr0L80vJ/mdaXAybhG 3eYF6WRXbLT7AP78DU/rE3e783aosE0k7fK7y+8fychjFJCDpbxNHJDx X-Gm-Gg: AR+sD13FZvI9ySKUtzxDjIJgZMH9EufHA4/kPgtjQGhHyCk3tkNAW/YHviD37LAScHn M+HDjR5gRm6++v/ye4jSPcy6EvqxcnmzGSUo5l9m5Xs4udirhhNKWmOjpKlFQg7YGzFQn36TkLu 29rqXiR5437V+udOjtcfruky0CVwE5rO2q0KsKuVVsdARWE49cmzjoDOB0YJGD/ogx4PhjD6ejV a1BchzlyRBZcbWsn0/hYe/+VW8vQtCl2upnd4EfL3RMoxjbGFbszLTlKL8l0sl7Yg1QAGiGq8Ve CXJv42sRYgNbORU1yk67yXaGG2MZ2ARx/ElRSIm/Oc5SakxDWrqiMHZm19DwcDJnySPSXvIlMSF KAt5Rl72tBRnDdscP0Mwstoy+AZ9l8TubDxdMPyjZ96mJxIG0jGLbDifmbmCGVNDsl07rhY1sjc yqHWAje9aMBigKdAY1vBY6tqAofD++HVv7MgRObLbcb3RS2KDokN3SBfxUidt5YqQPrpPQsSbz+ g8y7lU0X215o+TxAJM= X-Received: by 2002:a17:903:40d1:b0:2d6:f6ba:263d with SMTP id d9443c01a7336-2d6f6ba2701mr37235115ad.7.1787649616699; Tue, 25 Aug 2026 02:20:16 -0700 (PDT) Received: from localhost.localdomain ([103.120.31.178]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-3282afb407bsm5386163eec.30.2026.08.25.02.20.12 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Tue, 25 Aug 2026 02:20:15 -0700 (PDT) From: Khawar Ahemad To: bpf@vger.kernel.org, linux-kernel@vger.kernel.org Cc: ast@kernel.org, daniel@iogearbox.net, andrii@kernel.org, eddyz87@gmail.com, jiayuan.chen@linux.dev, emil@etsalapatis.com Subject: [PATCH v6 4/4] selftests/bpf: Add a test for arena fault-in under memory.max Date: Tue, 25 Aug 2026 14:49:52 +0530 Message-ID: <20260825091952.81971-5-ahemadkhawar123@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260825091952.81971-1-ahemadkhawar123@gmail.com> References: <20260825091952.81971-1-ahemadkhawar123@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Jiayuan Chen A child joins a memcg capped 64M above its post-load usage and faults an arena in until it runs out of that budget. With the fix the arena page comes from the sleepable allocator, so hitting memory.max goes through the memcg OOM path and the child is OOM-killed, which the test checks via memory.events "oom_kill". Without the fix the test may still pass, because a concurrent blocking allocation in the child (e.g. a COW fault on an inherited page) can hit memory.max and OOM-kill it first. The goal is only that the fixed kernel passes reliably. # test_progs -v -t arena_memcg serial_test_arena_memcg:PASS:child killed by signal serial_test_arena_memcg:PASS:memcg oom_kill #5 arena_memcg:OK # dmesg (the OOM comes from the arena sleepable allocation) test_progs invoked oom-killer: gfp_mask=3DGFP_KERNEL_ACCOUNT|__GFP_ZERO arena_vm_fault+0x4bc/0xad0 Memory cgroup out of memory: Killed process 473 (test_progs) Reviewed-by: Emil Tsalapatis Signed-off-by: Jiayuan Chen Signed-off-by: Khawar Ahemad --- .../selftests/bpf/prog_tests/arena_memcg.c | 158 ++++++++++++++++++ .../testing/selftests/bpf/progs/arena_memcg.c | 24 +++ 2 files changed, 182 insertions(+) create mode 100644 tools/testing/selftests/bpf/prog_tests/arena_memcg.c create mode 100644 tools/testing/selftests/bpf/progs/arena_memcg.c diff --git a/tools/testing/selftests/bpf/prog_tests/arena_memcg.c b/tools/t= esting/selftests/bpf/prog_tests/arena_memcg.c new file mode 100644 index 0000000000..c57b98494c --- /dev/null +++ b/tools/testing/selftests/bpf/prog_tests/arena_memcg.c @@ -0,0 +1,158 @@ +// SPDX-License-Identifier: GPL-2.0 + +#include +#include +#include +#include +#include +#include +#include +#ifndef PAGE_SIZE /* on some archs it comes in sys/user.h */ +#include +#define PAGE_SIZE getpagesize() +#endif + +#include "cgroup_helpers.h" +#include "arena_memcg.skel.h" + +#define CG_PATH "/arena_memcg" + +/* Budget the arena gets on top of whatever is already charged after load.= */ +#define ARENA_BUDGET (64 * 1024 * 1024) + +static void dump_memcg(int (*rd)(const char *, const char *, char *, size_= t)) +{ + char buf[512]; + + /* + * memory.current reads 0 once the child has left the cgroup, so it only + * carries information when dumped from the live child; memory.peak and + * memory.events survive the child and tell the story either way. + */ + if (!rd(CG_PATH, "memory.current", buf, sizeof(buf))) + fprintf(stderr, "memory.current: %s", buf); + if (!rd(CG_PATH, "memory.max", buf, sizeof(buf))) + fprintf(stderr, "memory.max: %s", buf); + if (!rd(CG_PATH, "memory.peak", buf, sizeof(buf))) + fprintf(stderr, "memory.peak: %s", buf); + if (!rd(CG_PATH, "memory.events", buf, sizeof(buf))) + fprintf(stderr, "memory.events:\n%s", buf); + fflush(stderr); +} + +/* Read one key from a flat keyed cgroup file, e.g. "oom_kill" in memory.e= vents. */ +static long cg_read_key(const char *cg, const char *file, const char *key) +{ + char buf[512], *p; + + if (read_cgroup_file(cg, file, buf, sizeof(buf))) + return -1; + p =3D strstr(buf, key); + if (!p) + return -1; + return strtol(p + strlen(key), NULL, 10); +} + +void serial_test_arena_memcg(void) +{ + int cgroup_fd =3D -1, status, err; + const long ps =3D PAGE_SIZE; + char buf[64]; + pid_t pid; + + err =3D setup_cgroup_environment(); + if (!ASSERT_OK(err, "setup_cgroup_environment")) + return; + + cgroup_fd =3D create_and_get_cgroup(CG_PATH); + if (!ASSERT_OK_FD(cgroup_fd, "create_and_get_cgroup")) + goto out; + + /* No memory controller -> nothing to test. */ + if (read_cgroup_file(CG_PATH, "memory.current", buf, sizeof(buf))) { + fprintf(stderr, "%s:SKIP:no memory controller\n", __func__); + test__skip(); + goto out; + } + + pid =3D fork(); + if (!ASSERT_GE(pid, 0, "fork")) + goto out; + if (pid =3D=3D 0) { + struct arena_memcg *cskel; + __u32 i, npages; + char *base; + size_t sz; + long cur; + + /* + * Do everything from the child: the arena vma is VM_DONTCOPY so + * it would not survive fork(), only the child should be under the + * limit so that a memcg OOM cannot pick test_progs, and a map is + * charged to the memcg of the task that creates it - so join + * before load. The cgroup work dir belongs to the parent that set + * the environment up, so reach it with the _parent() helpers. + * Errors are reported to the parent through the exit code, since + * ASSERT_* in a forked child does not reach it. + */ + snprintf(buf, sizeof(buf), "%d", getpid()); + if (write_cgroup_file_parent(CG_PATH, "cgroup.procs", buf)) + _exit(2); + + cskel =3D arena_memcg__open_and_load(); + if (!cskel) + _exit(3); + + base =3D bpf_map__initial_value(cskel->maps.arena, &sz); + if (!base) + _exit(4); + npages =3D bpf_map__max_entries(cskel->maps.arena); + + /* + * Cap only now, after load: everything but the fault-in is + * charged, so the arena gets a fixed budget regardless of what + * the load itself cost, and the load can never hit the limit. + */ + if (read_cgroup_file_parent(CG_PATH, "memory.current", buf, sizeof(buf))) + _exit(5); + cur =3D strtol(buf, NULL, 10); + snprintf(buf, sizeof(buf), "%ld", cur + ARENA_BUDGET); + if (write_cgroup_file_parent(CG_PATH, "memory.max", buf)) + _exit(6); + + for (i =3D 0; i < npages; i++) + base[(size_t)i * ps] =3D 1; + /* Faulted everything without dying: dump why (only under -v). */ + dump_memcg(read_cgroup_file_parent); + _exit(0); + } + + if (!ASSERT_EQ(waitpid(pid, &status, 0), pid, "waitpid")) + goto out; + + /* A non-zero exit means the child failed to set up; the code says where.= */ + if (WIFEXITED(status) && WEXITSTATUS(status)) { + ASSERT_OK(WEXITSTATUS(status), "child setup"); + goto out; + } + + /* + * Faulting a valid arena address until memory.max is hit must not look + * like an invalid access. Without the fix the fault path allocated with + * the non-blocking allocator, turned its -ENOMEM into VM_FAULT_SIGSEGV, + * and the child died with SIGSEGV on a valid address; now it is handled + * by the memcg OOM path instead. A SIGKILL alone would not prove the + * memcg OOM killer did it (a global OOM or an unrelated crash could also + * kill the child), so check memory.events.oom_kill, which records the + * memcg OOM and survives the child. + */ + if (!ASSERT_TRUE(WIFSIGNALED(status), "child killed by signal")) + goto out; + if (!ASSERT_GE(cg_read_key(CG_PATH, "memory.events", "oom_kill"), 1, + "memcg oom_kill")) + dump_memcg(read_cgroup_file); +out: + if (cgroup_fd >=3D 0) + close(cgroup_fd); + cleanup_cgroup_environment(); +} diff --git a/tools/testing/selftests/bpf/progs/arena_memcg.c b/tools/testin= g/selftests/bpf/progs/arena_memcg.c new file mode 100644 index 0000000000..88259cfea0 --- /dev/null +++ b/tools/testing/selftests/bpf/progs/arena_memcg.c @@ -0,0 +1,24 @@ +// SPDX-License-Identifier: GPL-2.0 + +#include +#include +#include "bpf_arena_common.h" + +struct { + __uint(type, BPF_MAP_TYPE_ARENA); + __uint(map_flags, BPF_F_MMAPABLE); + __uint(max_entries, 50000); /* number of pages */ +#ifdef __TARGET_ARCH_arm64 + __ulong(map_extra, 0x1ull << 32); /* start of mmap() region */ +#else + __ulong(map_extra, 0x1ull << 44); /* start of mmap() region */ +#endif +} arena SEC(".maps"); + +SEC("syscall") +int noop(void *ctx) +{ + return 0; +} + +char _license[] SEC("license") =3D "GPL"; --=20 2.54.0 (Apple Git-157)