From nobody Fri Jul 24 22:17:47 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 266A7381E98; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763680; cv=none; b=GMC1LK7KYfiWZa4gMvzP/27+CSObfEKdOEiYN0Fc9Z9wGgcUxe57U7xv/yP6D2Kdcy8Efpg6/G3Kx1pYGpa6SvrVVafNJOGxthWt4OzZQfTLDh1PgIFpkL9hBk3g50En2IdhbOCNJe96R9+mxOrJYhOVbyQs1Dvpfj+yLgu5fOA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763680; c=relaxed/simple; bh=QXp1uGYftwwR09EE5b5f4pGTnNpsJYm35tb0ykFlE3s=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=lf2SnVovTc8GI1rS2CytNO3E2iRDUc6+dDlikeFj0bLwftKLWbgImDYGa4WcCJ0bAOFsuLHFlVK0zu4aNuMi2Ghka7maSozqfk1TBspYxdZ0rKp1kP+BXiDixzOy6Rz8s2wBxKrUPwngdS/ZeTpTHfT6GFXqvv2xjn8zRgZUIeU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=UbXvg1Bt; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="UbXvg1Bt" Received: by smtp.kernel.org (Postfix) with ESMTPS id B4967C2BCF6; Wed, 22 Jul 2026 23:41:19 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1784763679; bh=QXp1uGYftwwR09EE5b5f4pGTnNpsJYm35tb0ykFlE3s=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=UbXvg1Bt1s145jeF4rWovqS58MES8avtXkEjqqPO38Xythxqu58Nh6N/44p3uyqEY cQTaCXOzR0vElWX08dLVLGKGAZhCtybDMnz+41NqkUhjmGje9/Sa5WzVlvBK5pOI2t kUtmNTfgcQQUtasT1FS3ZzrPqfbdT35lQoIRdiJV72UBpNtw1kBF2K0H/Uk12OlLb6 dIjgIjT9aXfPbPTp7wvfGKVr6ojBqHfQvBHXXsvG1RwRXAk6NLL1V5LD88wjAbUweZ orFJUj+7BSCuD+h+J9vcZ/FNheljdDQ/eYrR3yXOKsUpt+Ea4jL49RRqp9EtE6AfXt KpSkSbMvgt7Xg== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 92846C4453A; Wed, 22 Jul 2026 23:41:19 +0000 (UTC) From: Ackerley Tng via B4 Relay Date: Wed, 22 Jul 2026 16:41:09 -0700 Subject: [PATCH v4 01/16] mm: hugetlb: Track used_hpages when getting/putting pages from subpool Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260722-hugetlb-alloc-failure-fixes-v4-1-88e8b81970dc@google.com> References: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> In-Reply-To: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> To: Muchun Song , Oscar Salvador , David Hildenbrand , Joshua Hahn , Shakeel Butt , Nhat Pham , Andrew Morton , Peter Xu , Wupeng Ma , fvdl@google.com, rientjes@google.com, jthoughton@google.com, Mike Kravetz , Johannes Weiner , Michal Hocko , Roman Gushchin , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jonathan Corbet , Shuah Khan , Alex Shi , Yanteng Si , Dongliang Mu , Hongxiang Lou , Miaohe Lin Cc: vannapurve@google.com, erdemaktas@google.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, linux-doc@vger.kernel.org, Ackerley Tng , stable@vger.kernel.org X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=ed25519-sha256; t=1784763678; l=15585; i=ackerleytng@google.com; s=20260225; h=from:subject:message-id; bh=LjOzgnGx28MPv1Yhb1uCUnMcPrMOx9BKzeq+bcSee/E=; b=hd6+xuUmlxAzNJra6Q4jn1V7MPjPDpyeeckJSVaSuGFNE9zNfsKAY595pUL3NHsqR1cTzF+9x eo25kJ5PB+TCw5zBRMO2mJxkQy2yf2JlZADQOwOU9DhmaLFpCU6fN4N X-Developer-Key: i=ackerleytng@google.com; a=ed25519; pk=sAZDYXdm6Iz8FHitpHeFlCMXwabodTm7p8/3/8xUxuU= X-Endpoint-Received: by B4 Relay for ackerleytng@google.com/20260225 with auth_id=649 X-Original-From: Ackerley Tng Reply-To: ackerleytng@google.com From: Ackerley Tng hugepage_subpool_put_pages() currently has two distinct responsibilities that conflict: 1. When size is specified for the mount, max_hpages !=3D -1: Keep track of total active pages (allocated + reserved) and decrement this count (used_hpages) when a page is freed or allocation fails. 2. When min_size is specified for the mount, min_hpages !=3D -1: Ensure we don't drop below the guaranteed minimum, and restore a reservation (rsv_hpages) if we do. This causes trouble because when allocation fails (refer to alloc_hugetlb_folio()) if gbl_chg =3D 1 (i.e. no subpool reservation was taken): + To keep used_hpages consistent, HugeTLB needs to call hugepage_subpool_put_pages() to restore undo used_hpages being incremented + But can't call hugepage_subpool_put_pages() if no reservation was consumed. One option would be to conditionally do subpool tracking updates outside of the hugepage_subpool_put_pages() function, but that would spread logic all over. Instead, always track used_hpages, regardless of whether a max_size was requested for the mount, so that the subpool always knows how many pages were allocated through it. Every page allocated through the subpool increments used_hpages, regardless of whether a reservation was taken from it. Conceptually, now, every allocation involving a subpool uses a page from the subpool, which must be returned to the subpool. Every page taken from the subpool tries to use a subpool reservation. Restoring a page to the subpool reservations only if the page was taken from subpool reservations. (If used_hpages >=3D min_hpages, the page must have not have been taken from the reservations.) Always tracking used_hpages provides the subpool with information of both used and reserved counts to make the correct decision for both max_size and min_size correctly. With used_hpages always tracked, + subpool_is_free() can be simplified, such that the subpool can be declared free if there are no more pages in use. + open-coding in hugetlb_reserve_pages() can be removed. Also update the + Documentation for used_hpages in the subpool struct, since it no longer matters whether the used pages count against the maximum. + Docstring for hugepage_subpool_{get,put}_pages + Documentation to use active voice, and remove some details in favor of having details documented in the docstring Also update statfs reporting. Previously, if max_hpages is negative, used_hpages is static at 0, so returning max_hpages - used_hpages returns -1 and is always correct. Now, if the subpool doesn't have a maximum requested size, indicate no limit for free pages (-1). If it does have a maximum size, report the difference between the requested size and the number of used pages. This difference is always positive, because if the mount does have a maximum size, hugepage_subpool_get_pages() ensures that the subpool usage never exceeds the maximum. This fixes a bug in hugetlb_unreserve_pages(), where pages are returned to the subpool regardless of whether it consumed a reservation. The corresponding bug in the failure handling path of alloc_hugetlb_folio() was fixed in a833a693a490e. Fixes: 1c5ecae3a93fa ("hugetlbfs: add minimum size accounting to subpools") Cc: stable@vger.kernel.org Signed-off-by: Ackerley Tng --- Documentation/mm/hugetlbfs_reserv.rst | 17 +-- .../translations/zh_CN/mm/hugetlbfs_reserv.rst | 11 +- fs/hugetlbfs/inode.c | 8 +- include/linux/hugetlb.h | 4 +- mm/hugetlb.c | 118 +++++++++++------= ---- 5 files changed, 75 insertions(+), 83 deletions(-) diff --git a/Documentation/mm/hugetlbfs_reserv.rst b/Documentation/mm/huget= lbfs_reserv.rst index a49115db18c76..d244583fdcbc3 100644 --- a/Documentation/mm/hugetlbfs_reserv.rst +++ b/Documentation/mm/hugetlbfs_reserv.rst @@ -314,21 +314,8 @@ huge pages. If they can not be reserved, the mount fa= ils. The routines hugepage_subpool_get/put_pages() are called when pages are obtained from or released back to a subpool. They perform all subpool accounting, and track any reservations associated with the subpool. -hugepage_subpool_get/put_pages are passed the number of huge pages by which -to adjust the subpool 'used page' count (down for get, up for put). Norma= lly, -they return the same value that was passed or an error if not enough pages -exist in the subpool. - -However, if reserves are associated with the subpool a return value less -than the passed value may be returned. This return value indicates the -number of additional global pool adjustments which must be made. For exam= ple, -suppose a subpool contains 3 reserved huge pages and someone asks for 5. -The 3 reserved pages associated with the subpool can be used to satisfy pa= rt -of the request. But, 2 pages must be obtained from the global pools. To -relay this information to the caller, the value 2 is returned. The caller -is then responsible for attempting to obtain the additional two pages from -the global pools. - +hugepage_subpool_get/put_pages() use the number of huge pages passed to ad= just +the subpool 'used page' count. =20 COW and Reservations =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D diff --git a/Documentation/translations/zh_CN/mm/hugetlbfs_reserv.rst b/Doc= umentation/translations/zh_CN/mm/hugetlbfs_reserv.rst index 20947f8bd0654..ae1f1f31477fc 100644 --- a/Documentation/translations/zh_CN/mm/hugetlbfs_reserv.rst +++ b/Documentation/translations/zh_CN/mm/hugetlbfs_reserv.rst @@ -246,15 +246,8 @@ hugepage_subpool=E7=9A=84min_hpages=E5=AD=97=E6=AE=B5= =E4=B8=AD=E8=A2=AB=E8=B7=9F=E8=B8=AA=E3=80=82=E5=9C=A8=E6=8C=82=E8=BD=BD=E6= =97=B6=EF=BC=8Chugetlb_acct_me =E8=A2=AB=E8=B0=83=E7=94=A8=E4=BB=A5=E9=A2=84=E7=95=99=E6=8C=87=E5=AE=9A= =E6=95=B0=E9=87=8F=E7=9A=84=E5=B7=A8=E9=A1=B5=E3=80=82=E5=A6=82=E6=9E=9C=E5= =AE=83=E4=BB=AC=E4=B8=8D=E8=83=BD=E8=A2=AB=E9=A2=84=E7=95=99=EF=BC=8C=E6=8C= =82=E8=BD=BD=E5=B0=B1=E4=BC=9A=E5=A4=B1=E8=B4=A5=E3=80=82 =20 =E5=BD=93=E4=BB=8E=E5=AD=90=E6=B1=A0=E4=B8=AD=E8=8E=B7=E5=8F=96=E6=88=96= =E9=87=8A=E6=94=BE=E9=A1=B5=E9=9D=A2=E6=97=B6=EF=BC=8C=E4=BC=9A=E8=B0=83=E7= =94=A8hugepage_subpool_get/put_pages()=E5=87=BD=E6=95=B0=E3=80=82 -hugepage_subpool_get/put_pages=E8=A2=AB=E4=BC=A0=E9=80=92=E7=BB=99=E5=B7= =A8=E9=A1=B5=E6=95=B0=E9=87=8F=EF=BC=8C=E4=BB=A5=E6=AD=A4=E6=9D=A5=E8=B0=83= =E6=95=B4=E5=AD=90=E6=B1=A0=E7=9A=84 =E2=80=9C=E5=B7=B2=E7=94=A8=E9=A1=B5= =E9=9D=A2=E2=80=9D =E8=AE=A1=E6=95=B0 -=EF=BC=88get=E4=B8=BA=E4=B8=8B=E9=99=8D=EF=BC=8Cput=E4=B8=BA=E4=B8=8A=E5= =8D=87=EF=BC=89=E3=80=82=E9=80=9A=E5=B8=B8=E6=83=85=E5=86=B5=E4=B8=8B=EF=BC= =8C=E5=A6=82=E6=9E=9C=E5=AD=90=E6=B1=A0=E4=B8=AD=E6=B2=A1=E6=9C=89=E8=B6=B3= =E5=A4=9F=E7=9A=84=E9=A1=B5=E9=9D=A2=EF=BC=8C=E5=AE=83=E4=BB=AC=E4=BC=9A=E8= =BF=94=E5=9B=9E=E4=B8=8E=E4=BC=A0=E9=80=92=E7=9A=84=E7=9B=B8=E5=90=8C=E7=9A= =84=E5=80=BC=E6=88=96 -=E4=B8=80=E4=B8=AA=E9=94=99=E8=AF=AF=E3=80=82 - -=E7=84=B6=E8=80=8C=EF=BC=8C=E5=A6=82=E6=9E=9C=E9=A2=84=E7=95=99=E4=B8=8E= =E5=AD=90=E6=B1=A0=E7=9B=B8=E5=85=B3=E8=81=94=EF=BC=8C=E5=8F=AF=E8=83=BD=E4= =BC=9A=E8=BF=94=E5=9B=9E=E4=B8=80=E4=B8=AA=E5=B0=8F=E4=BA=8E=E4=BC=A0=E9=80= =92=E5=80=BC=E7=9A=84=E8=BF=94=E5=9B=9E=E5=80=BC=E3=80=82=E8=BF=99=E4=B8=AA= =E8=BF=94=E5=9B=9E=E5=80=BC=E8=A1=A8=E7=A4=BA=E5=BF=85=E9=A1=BB=E8=BF=9B=E8= =A1=8C=E7=9A=84=E9=A2=9D=E5=A4=96=E5=85=A8=E5=B1=80 -=E6=B1=A0=E8=B0=83=E6=95=B4=E7=9A=84=E6=95=B0=E9=87=8F=E3=80=82=E4=BE=8B= =E5=A6=82=EF=BC=8C=E5=81=87=E8=AE=BE=E4=B8=80=E4=B8=AA=E5=AD=90=E6=B1=A0=E5= =8C=85=E5=90=AB3=E4=B8=AA=E9=A2=84=E7=95=99=E7=9A=84=E5=B7=A8=E9=A1=B5=EF= =BC=8C=E6=9C=89=E4=BA=BA=E8=A6=81=E6=B1=825=E4=B8=AA=E3=80=82=E4=B8=8E=E5= =AD=90=E6=B1=A0=E7=9B=B8=E5=85=B3=E7=9A=843=E4=B8=AA=E9=A2=84=E7=95=99=E9= =A1=B5=E5=8F=AF=E4=BB=A5=E7=94=A8=E6=9D=A5 -=E6=BB=A1=E8=B6=B3=E9=83=A8=E5=88=86=E8=AF=B7=E6=B1=82=E3=80=82=E4=BD=86= =E6=98=AF=EF=BC=8C=E5=BF=85=E9=A1=BB=E4=BB=8E=E5=85=A8=E5=B1=80=E6=B1=A0=E4= =B8=AD=E8=8E=B7=E5=BE=972=E4=B8=AA=E9=A1=B5=E9=9D=A2=E3=80=82=E4=B8=BA=E4= =BA=86=E5=90=91=E8=B0=83=E7=94=A8=E8=80=85=E8=BD=AC=E8=BE=BE=E8=BF=99=E4=B8= =80=E4=BF=A1=E6=81=AF=EF=BC=8C=E5=B0=86=E8=BF=94=E5=9B=9E=E5=80=BC2=E3=80= =82=E7=84=B6=E5=90=8E=EF=BC=8C=E8=B0=83=E7=94=A8 -=E8=80=85=E8=A6=81=E8=B4=9F=E8=B4=A3=E4=BB=8E=E5=85=A8=E5=B1=80=E6=B1=A0= =E4=B8=AD=E8=8E=B7=E5=8F=96=E5=8F=A6=E5=A4=96=E4=B8=A4=E4=B8=AA=E9=A1=B5=E9= =9D=A2=E3=80=82 - +=E5=AE=83=E4=BB=AC=E8=B4=9F=E8=B4=A3=E6=89=80=E6=9C=89=E5=AD=90=E6=B1=A0= =E7=9A=84=E7=BB=9F=E8=AE=A1=E6=A0=B8=E7=AE=97=EF=BC=8C=E5=B9=B6=E8=B7=9F=E8= =B8=AA=E4=B8=8E=E5=AD=90=E6=B1=A0=E7=9B=B8=E5=85=B3=E8=81=94=E7=9A=84=E9=A2= =84=E7=95=99=E3=80=82 +hugepage_subpool_get/put_pages()=E5=87=BD=E6=95=B0=E4=BD=BF=E7=94=A8=E4=BC= =A0=E5=85=A5=E7=9A=84=E5=B7=A8=E9=A1=B5=E6=95=B0=E9=87=8F=E6=9D=A5=E8=B0=83= =E6=95=B4=E5=AD=90=E6=B1=A0=E7=9A=84=E2=80=9C=E5=B7=B2=E7=94=A8=E9=A1=B5=E9= =9D=A2=E2=80=9D=E8=AE=A1=E6=95=B0=E3=80=82 =20 COW=E5=92=8C=E9=A2=84=E7=95=99 =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D diff --git a/fs/hugetlbfs/inode.c b/fs/hugetlbfs/inode.c index 216e1a0dd0b23..26c0187340636 100644 --- a/fs/hugetlbfs/inode.c +++ b/fs/hugetlbfs/inode.c @@ -1109,8 +1109,12 @@ static int hugetlbfs_statfs(struct dentry *dentry, s= truct kstatfs *buf) =20 spin_lock_irq(&sbinfo->spool->lock); buf->f_blocks =3D sbinfo->spool->max_hpages; - free_pages =3D sbinfo->spool->max_hpages - - sbinfo->spool->used_hpages; + if (sbinfo->spool->max_hpages =3D=3D -1) { + free_pages =3D -1; + } else { + free_pages =3D sbinfo->spool->max_hpages - + sbinfo->spool->used_hpages; + } buf->f_bavail =3D buf->f_bfree =3D free_pages; spin_unlock_irq(&sbinfo->spool->lock); buf->f_files =3D sbinfo->max_inodes; diff --git a/include/linux/hugetlb.h b/include/linux/hugetlb.h index 2abaf99321e90..34b9a3e1be0fa 100644 --- a/include/linux/hugetlb.h +++ b/include/linux/hugetlb.h @@ -38,8 +38,8 @@ struct hugepage_subpool { spinlock_t lock; long count; long max_hpages; /* Maximum huge pages or -1 if no maximum. */ - long used_hpages; /* Used count against maximum, includes */ - /* both allocated and reserved pages. */ + long used_hpages; /* Used page count, includes both */ + /* allocated and reserved pages. */ struct hstate *hstate; long min_hpages; /* Minimum huge pages or -1 if no minimum. */ long rsv_hpages; /* Pages reserved against global pool to */ diff --git a/mm/hugetlb.c b/mm/hugetlb.c index 571212b80835e..36fa3fb3945d8 100644 --- a/mm/hugetlb.c +++ b/mm/hugetlb.c @@ -129,12 +129,8 @@ static inline bool subpool_is_free(struct hugepage_sub= pool *spool) { if (spool->count) return false; - if (spool->max_hpages !=3D -1) - return spool->used_hpages =3D=3D 0; - if (spool->min_hpages !=3D -1) - return spool->rsv_hpages =3D=3D spool->min_hpages; =20 - return true; + return spool->used_hpages =3D=3D 0; } =20 static inline void unlock_or_release_subpool(struct hugepage_subpool *spoo= l, @@ -187,13 +183,18 @@ void hugepage_put_subpool(struct hugepage_subpool *sp= ool) unlock_or_release_subpool(spool, flags); } =20 -/* - * Subpool accounting for allocating and reserving pages. - * Return -ENOMEM if there are not enough resources to satisfy the - * request. Otherwise, return the number of pages by which the - * global pools must be adjusted (upward). The returned value may - * only be different than the passed value (delta) in the case where - * a subpool minimum size must be maintained. +/** + * hugepage_subpool_get_pages - Get pages from a subpool + * @spool: pointer to subpool structure (may be NULL) + * @delta: number of pages to allocate or reserve + * + * Check and update subpool page usage counts when allocating or + * reserving @delta hugepages. + * + * Context: Takes spool->lock using spin_lock_irq(). + * Return: Non-negative number of reservations that cannot be + * satisfied by the subpool, or -ENOMEM if the subpool maximum + * limit would be exceeded. */ static long hugepage_subpool_get_pages(struct hugepage_subpool *spool, long delta) @@ -205,15 +206,14 @@ static long hugepage_subpool_get_pages(struct hugepag= e_subpool *spool, =20 spin_lock_irq(&spool->lock); =20 - if (spool->max_hpages !=3D -1) { /* maximum size accounting */ - if ((spool->used_hpages + delta) <=3D spool->max_hpages) - spool->used_hpages +=3D delta; - else { - ret =3D -ENOMEM; - goto unlock_ret; - } + if (spool->max_hpages !=3D -1 && + spool->used_hpages + delta > spool->max_hpages) { + ret =3D -ENOMEM; + goto unlock_ret; } =20 + spool->used_hpages +=3D delta; + /* minimum size accounting */ if (spool->min_hpages !=3D -1 && spool->rsv_hpages) { if (delta > spool->rsv_hpages) { @@ -234,11 +234,19 @@ static long hugepage_subpool_get_pages(struct hugepag= e_subpool *spool, return ret; } =20 -/* - * Subpool accounting for freeing and unreserving pages. - * Return the number of global page reservations that must be dropped. - * The return value may only be different than the passed value (delta) - * in the case where a subpool minimum size must be maintained. +/** + * hugepage_subpool_put_pages - Release pages back to a subpool + * @spool: pointer to subpool structure (may be NULL) + * @delta: number of pages to free or unreserve + * + * Check and update subpool page usage counts when freeing or + * unreserving @delta hugepages. + * + * Context: Takes spool->lock using spin_lock_irqsave(). May release + * and free @spool if its usage count and references reach + * zero. + * Return: Non-negative number of reservations that the subpool cannot + * absorb. */ static long hugepage_subpool_put_pages(struct hugepage_subpool *spool, long delta) @@ -251,19 +259,24 @@ static long hugepage_subpool_put_pages(struct hugepag= e_subpool *spool, =20 spin_lock_irqsave(&spool->lock, flags); =20 - if (spool->max_hpages !=3D -1) /* maximum size accounting */ - spool->used_hpages -=3D delta; + spool->used_hpages -=3D delta; =20 /* minimum size accounting */ if (spool->min_hpages !=3D -1 && spool->used_hpages < spool->min_hpages) { - if (spool->rsv_hpages + delta <=3D spool->min_hpages) + /* + * limit is the maximum number of reservations that + * can be restored to this subpool. + */ + long limit =3D spool->min_hpages - spool->used_hpages; + + if (spool->rsv_hpages + delta <=3D limit) ret =3D 0; else - ret =3D spool->rsv_hpages + delta - spool->min_hpages; + ret =3D spool->rsv_hpages + delta - limit; =20 spool->rsv_hpages +=3D delta; - if (spool->rsv_hpages > spool->min_hpages) - spool->rsv_hpages =3D spool->min_hpages; + if (spool->rsv_hpages > limit) + spool->rsv_hpages =3D limit; } =20 /* @@ -6542,7 +6555,7 @@ long hugetlb_reserve_pages(struct inode *inode, struct vm_area_struct *vma, vma_flags_t vma_flags) { - long chg =3D -1, add =3D -1, spool_resv, gbl_resv; + long chg =3D -1, add =3D -1, gbl_resv; struct hstate *h =3D hstate_inode(inode); struct hugepage_subpool *spool =3D subpool_inode(inode); struct resv_map *resv_map; @@ -6622,9 +6635,9 @@ long hugetlb_reserve_pages(struct inode *inode, * the subpool has a minimum size, there may be some global * reservations already in place (gbl_reserve). */ - gbl_reserve =3D hugepage_subpool_get_pages(spool, chg); - if (gbl_reserve < 0) { - err =3D gbl_reserve; + gbl_resv =3D hugepage_subpool_get_pages(spool, chg); + if (gbl_resv < 0) { + err =3D gbl_resv; goto out_uncharge_cgroup; } =20 @@ -6632,7 +6645,7 @@ long hugetlb_reserve_pages(struct inode *inode, * Check enough hugepages are available for the reservation. * Hand the pages back to the subpool if there are not */ - err =3D hugetlb_acct_memory(h, gbl_reserve); + err =3D hugetlb_acct_memory(h, gbl_resv); if (err < 0) goto out_put_pages; =20 @@ -6651,7 +6664,7 @@ long hugetlb_reserve_pages(struct inode *inode, add =3D region_add(resv_map, from, to, regions_needed, h, h_cg); =20 if (unlikely(add < 0)) { - hugetlb_acct_memory(h, -gbl_reserve); + hugetlb_acct_memory(h, -gbl_resv); err =3D add; goto out_put_pages; } else if (unlikely(chg > add)) { @@ -6687,26 +6700,21 @@ long hugetlb_reserve_pages(struct inode *inode, } return chg; =20 -out_put_pages: - spool_resv =3D chg - gbl_reserve; - if (spool_resv) { - /* put sub pool's reservation back, chg - gbl_reserve */ - gbl_resv =3D hugepage_subpool_put_pages(spool, spool_resv); - /* - * subpool's reserved pages can not be put back due to race, - * return to hstate. - */ - hugetlb_acct_memory(h, -gbl_resv); - } - /* Restore used_hpages for pages that failed global reservation */ - if (gbl_reserve && spool) { - unsigned long flags; + out_put_pages: + /* + * Return all that was requested from the subpool, let subpool + * tell us the new number of reservations that need to be + * returned to the global pool. + */ + gbl_reserve =3D hugepage_subpool_put_pages(spool, chg); + /* + * There may be a difference between the number of + * reservations to consume and the number to restore now if + * there are multiple threads interacting with the subpool - + * restore the difference. + */ + hugetlb_acct_memory(h, gbl_resv - gbl_reserve); =20 - spin_lock_irqsave(&spool->lock, flags); - if (spool->max_hpages !=3D -1) - spool->used_hpages -=3D gbl_reserve; - unlock_or_release_subpool(spool, flags); - } out_uncharge_cgroup: hugetlb_cgroup_uncharge_cgroup_rsvd(hstate_index(h), chg * pages_per_huge_page(h), h_cg); --=20 2.55.0.229.g6434b31f56-goog From nobody Fri Jul 24 22:17:47 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 26ABA41229C; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763680; cv=none; b=M9bjkrD++wRpw+WSmVdvGKN1gOrHg3CQjbY3yvGgvBamwNpaQjVOPuo4yleKoIA1Zfv1ya1PXWPCta0UUbYABN2nfjJdP9uSowqAcIWIEZHEYVOY7TzLlEXZpb10WFcmIeIOE7dKVSj1E2a5KiIOPw/yBJKKue3Zlc/Qsc7L+Y4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763680; c=relaxed/simple; bh=asDu53szvZ8BxoP3xBFFR39bZX7xQtE305TUvB0bAIE=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=mqtyfU1yj/pCZAyPIwNw3oMrM4fYltsYXVMEHH9snRDoFFV0vnQ9msaQmjDvZ17BTkOygECfuXwtkAwRml5RakOsc8hf5NcUi/VjTSz7j/rcra0z1PCuOvkdbF65Y4P51Jhfwl13wG5NbR3ZEQlh77xYwJgTfuFR+RgayJWVaXY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=lxknJXe7; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="lxknJXe7" Received: by smtp.kernel.org (Postfix) with ESMTPS id C6518C2BCC6; Wed, 22 Jul 2026 23:41:19 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1784763679; bh=asDu53szvZ8BxoP3xBFFR39bZX7xQtE305TUvB0bAIE=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=lxknJXe7l9WWVl1nrdgPiDECtcTR/T0+0QyM5xKGNuyOhOYaOFp6SdHkZle3Cmsoc Fu59LGcdYS7jLw/n7M7E+GtnE/sUYF6xkkhiOcgzVvfurCiYerWm/ECSetMhR0HqYr Vt6Bv0du2qmOM8DxocCyXhex9yRokvfeENUjVAQTSasaITQs72EP77EzOpDWhpXXkR 94vYDqDsQnLa2Sk7FG+RZjISvTQeN0duxCqh3Xsi3DwBYJQLDubjz3HTJwCTv3rUre U5qOfF9SoXIeceLDrvZUJ6fUkWdQNrsvoMBcwtHHgV2nJOwwzGdHCEvowgbkSxYNW0 YyTr1Bi9zTKWw== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id A49A3C44536; Wed, 22 Jul 2026 23:41:19 +0000 (UTC) From: Ackerley Tng via B4 Relay Date: Wed, 22 Jul 2026 16:41:10 -0700 Subject: [PATCH v4 02/16] mm: hugetlb: Return -ENOSPC on memcg charge failure Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260722-hugetlb-alloc-failure-fixes-v4-2-88e8b81970dc@google.com> References: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> In-Reply-To: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> To: Muchun Song , Oscar Salvador , David Hildenbrand , Joshua Hahn , Shakeel Butt , Nhat Pham , Andrew Morton , Peter Xu , Wupeng Ma , fvdl@google.com, rientjes@google.com, jthoughton@google.com, Mike Kravetz , Johannes Weiner , Michal Hocko , Roman Gushchin , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jonathan Corbet , Shuah Khan , Alex Shi , Yanteng Si , Dongliang Mu , Hongxiang Lou , Miaohe Lin Cc: vannapurve@google.com, erdemaktas@google.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, linux-doc@vger.kernel.org, Ackerley Tng , stable@vger.kernel.org X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=ed25519-sha256; t=1784763678; l=2279; i=ackerleytng@google.com; s=20260225; h=from:subject:message-id; bh=wmoXONTSMCE0pZEPoT08cRwobUPLPpiUr/LRWZQPS8U=; b=KDHlw7WdJmh8mcn+h2AhBlYR5bdH3uYOuVSgl+nfxzJXdfEHSQ3u3Zm+FIYqoIRms5b4sL//9 RkxUzY2QUxjAzwB8bEk+AgPBG+by/l/X50vgdac5sR6PTRn695EIGVh X-Developer-Key: i=ackerleytng@google.com; a=ed25519; pk=sAZDYXdm6Iz8FHitpHeFlCMXwabodTm7p8/3/8xUxuU= X-Endpoint-Received: by B4 Relay for ackerleytng@google.com/20260225 with auth_id=649 X-Original-From: Ackerley Tng Reply-To: ackerleytng@google.com From: Ackerley Tng When mem_cgroup_charge_hugetlb() fails with -ENOMEM, alloc_hugetlb_folio() currently propagates this error. This results in the page fault handler returning VM_FAULT_OOM. Because HugeTLB allocations are high-order and use __GFP_RETRY_MAYFAIL, they bypass the OOM killer. Returning VM_FAULT_OOM to the #PF handler without triggering the OOM killer (or having it make progress) leads to an infinite loop of retrying the fault. Avoid this loop by returning -ENOSPC when charging fails, which maps to VM_FAULT_SIGBUS, terminating the process cleanly. Make mem_cgroup_charge_hugetlb() fault handling use a common error handling path, the same handling used for hugetlb_cgroup_uncharge_cgroup{,_rsvd}(), which also don't trigger the OOM killer and hence opt to terminate the process with a SIGBUS. Fixes: 991135774c0e0 ("memcg/hugetlb: introduce mem_cgroup_charge_hugetlb") Cc: stable@vger.kernel.org Signed-off-by: Ackerley Tng Reviewed-by: Muchun Song Tested-by: Joshua Hahn Reviewed-by: Joshua Hahn --- mm/hugetlb.c | 12 +++++++++++- 1 file changed, 11 insertions(+), 1 deletion(-) diff --git a/mm/hugetlb.c b/mm/hugetlb.c index 36fa3fb3945d8..15f9c5a9f75e2 100644 --- a/mm/hugetlb.c +++ b/mm/hugetlb.c @@ -3010,7 +3010,7 @@ struct folio *alloc_hugetlb_folio(struct vm_area_stru= ct *vma, =20 if (ret =3D=3D -ENOMEM) { free_huge_folio(folio); - return ERR_PTR(-ENOMEM); + goto err; } =20 return folio; @@ -3035,6 +3035,16 @@ struct folio *alloc_hugetlb_folio(struct vm_area_str= uct *vma, out_end_reservation: if (map_chg !=3D MAP_CHG_ENFORCED) vma_end_reservation(h, vma, addr); +err: + /* + * Return -ENOSPC when this function fails to allocate or charge a huge + * page. If a standard (PAGE_SIZE) page allocation fails, the OOM killer + * is given a chance to run, which may resolve the failure on + * retry. However, for HugeTLB allocations, the OOM killer is not + * triggered. Returning -ENOMEM (or anything resulting in VM_FAULT_OOM) + * would leak to the #PF handler, causing it to loop indefinitely + * retrying the fault. + */ return ERR_PTR(-ENOSPC); } =20 --=20 2.55.0.229.g6434b31f56-goog From nobody Fri Jul 24 22:17:47 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2B5C442B302; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763680; cv=none; b=spxHoVa6Qq50YKeOsiEC4eNp0HtbtZJYnIm9amuqnHDlE80Ma3dHAcGlH9VZViAzNEyvTs8MlJck7iqftspRXHwkPZb96gHoyNzqDp7UI12B/GMLjasJy1kMshwbu1lpu6uJbxo+wIgVpce+GahRfP69qCRAUqWG3c9a4GkuqEw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763680; c=relaxed/simple; bh=gdyXfp6oV2z7yddoLW13M8klm/MDNV0FUwUjUXRLoak=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=OMvlqVALxYS+PMSoRFFPqxiRLH/30iO7jIvR5YxvT7yTQz6udcdFlKqfwNx62EqC9/dHSHVoH6s7qf1QnZrLLtNXt6r5CypFnUvuq+uiQzLldD5JpaxtCFz7qKFkCdpndZrptTvp9HC++xnKVqrIeMyngrNfF7iw8cCLHiIjKgg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=Z/+b1J2j; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="Z/+b1J2j" Received: by smtp.kernel.org (Postfix) with ESMTPS id EBEE0C2BD00; Wed, 22 Jul 2026 23:41:19 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1784763680; bh=gdyXfp6oV2z7yddoLW13M8klm/MDNV0FUwUjUXRLoak=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=Z/+b1J2jr/NLTK6eDcWdxX8SF6KLayk+pvoOrE7IGnih1EV3hEMXO4Zd6HuamfKeF SDifgvho/8/TKQgoDuzA2ZSYkS9pZ25zVo8tzL/ouRtU6pKTksopLgtD3XziJnD7HX IT1RbjcehTU3JJAXnufFjVSCU7dK0htbcOhvHxpXoSIVaku1S7dEPMIO1mEj9m6hyJ imJmB6N4r9Fq7a8ad5ltb+mHj9yGrhass7umXJ23HpzTWUgfr1kCulHvbx5egVAppE hyu1BTeifjBoIVu5wHcp0csuHPXizP8Sa007V90gfP+q5nmiX6L4maSb7SSV81DImV d3LoFfzi0sFeA== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id D6CC6C44532; Wed, 22 Jul 2026 23:41:19 +0000 (UTC) From: Ackerley Tng via B4 Relay Date: Wed, 22 Jul 2026 16:41:11 -0700 Subject: [PATCH v4 03/16] mm: hugetlb: Use try-commit-cancel protocol for memcg charge of folios Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260722-hugetlb-alloc-failure-fixes-v4-3-88e8b81970dc@google.com> References: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> In-Reply-To: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> To: Muchun Song , Oscar Salvador , David Hildenbrand , Joshua Hahn , Shakeel Butt , Nhat Pham , Andrew Morton , Peter Xu , Wupeng Ma , fvdl@google.com, rientjes@google.com, jthoughton@google.com, Mike Kravetz , Johannes Weiner , Michal Hocko , Roman Gushchin , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jonathan Corbet , Shuah Khan , Alex Shi , Yanteng Si , Dongliang Mu , Hongxiang Lou , Miaohe Lin Cc: vannapurve@google.com, erdemaktas@google.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, linux-doc@vger.kernel.org, Ackerley Tng , stable@vger.kernel.org X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=ed25519-sha256; t=1784763678; l=10726; i=ackerleytng@google.com; s=20260225; h=from:subject:message-id; bh=vqQmK3vN70VbiuGHrvYrgKdPuFJc0Ryl2R9yY2WzCMc=; b=FrpFLnNeyQ3odoN5g/ynV4evoTr/eoT6mQ2NGQz780SQ4+n3wcukCV/XcDDhN/ShTiyR4BlX1 p8UNLtHCnEMCI5eROYm7u8CY3JSd5dEg9TehHnOB1fSlzXl16S/e3nL X-Developer-Key: i=ackerleytng@google.com; a=ed25519; pk=sAZDYXdm6Iz8FHitpHeFlCMXwabodTm7p8/3/8xUxuU= X-Endpoint-Received: by B4 Relay for ackerleytng@google.com/20260225 with auth_id=649 X-Original-From: Ackerley Tng Reply-To: ackerleytng@google.com From: Ackerley Tng Using mem_cgroup_charge_hugetlb() to charge a HugeTLB folio during page fault creates a reservation leak bug if the task hits its memory cgroup limit. When alloc_hugetlb_folio() commits the VMA reservation, the reserved page is removed from the reserve map. If a subsequent call to mem_cgroup_charge_hugetlb() returns -ENOMEM, the allocation is aborted and the physical folio is disposed of via free_huge_page(). However, because the VMA reservation was already consumed, the reservation count in the reserve map is lost. This causes subsequent faults in the VMA address range to fail with premature reservation exhaustion. Additionally, dropping the use of free_huge_folio() on the failure path fixes an issue where free_huge_folio() was incorrectly invoked on a folio with a refcount of 1, triggering refcount mismatches and kernel warnings. To fix this, introduce a try-commit-cancel protocol for memory cgroup charging of HugeTLB folios, matching the architecture used by the hugetlb cgroup controller. Invoking mem_cgroup_hugetlb_try_charge() before consuming the VMA reservation ensures that if the memory cgroup limit is reached, the allocation is aborted cleanly without leaking the reservation entry or having to dispose of a partially initialized folio. An alternative would be to retain the current usage of mem_cgroup_charge_hugetlb() and free_huge_page(), but freeing the folio performs reservation management for subpools and global hstate, which complicates rollback in alloc_hugetlb_folio(). Using a try-commit-cancel protocol is more consistent with the other charging performed in alloc_hugetlb_folio() and easier to understand. Fixes: 991135774c0e0 ("memcg/hugetlb: introduce mem_cgroup_charge_hugetlb") Cc: stable@vger.kernel.org Signed-off-by: Ackerley Tng --- include/linux/memcontrol.h | 30 ++++++++++++ mm/hugetlb.c | 30 ++++++------ mm/memcontrol.c | 114 +++++++++++++++++++++++++++++++++++++++++= ++++ 3 files changed, 159 insertions(+), 15 deletions(-) diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h index e1f46a0016fcf..2c3e2c62f4176 100644 --- a/include/linux/memcontrol.h +++ b/include/linux/memcontrol.h @@ -642,6 +642,15 @@ static inline int mem_cgroup_charge(struct folio *foli= o, struct mm_struct *mm, } =20 int mem_cgroup_charge_hugetlb(struct folio* folio, gfp_t gfp); +int mem_cgroup_hugetlb_try_charge(unsigned int nr_pages, gfp_t gfp, + struct mem_cgroup **memcg_p, + struct obj_cgroup **objcg_p); +void mem_cgroup_hugetlb_commit_charge(struct folio *folio, + struct mem_cgroup *memcg, + struct obj_cgroup *objcg); +void mem_cgroup_hugetlb_cancel_charge(unsigned int nr_pages, + struct mem_cgroup *memcg, + struct obj_cgroup *objcg); =20 int mem_cgroup_swapin_charge_folio(struct folio *folio, unsigned short id, struct mm_struct *mm, gfp_t gfp); @@ -1133,6 +1142,27 @@ static inline int mem_cgroup_charge_hugetlb(struct f= olio* folio, gfp_t gfp) return 0; } =20 +static inline int mem_cgroup_hugetlb_try_charge(unsigned int nr_pages, gfp= _t gfp, + struct mem_cgroup **memcg_p, + struct obj_cgroup **objcg_p) +{ + *memcg_p =3D NULL; + *objcg_p =3D NULL; + return 0; +} + +static inline void mem_cgroup_hugetlb_commit_charge(struct folio *folio, + struct mem_cgroup *memcg, + struct obj_cgroup *objcg) +{ +} + +static inline void mem_cgroup_hugetlb_cancel_charge(unsigned int nr_pages, + struct mem_cgroup *memcg, + struct obj_cgroup *objcg) +{ +} + static inline int mem_cgroup_swapin_charge_folio(struct folio *folio, unsigned short id, struct mm_struct *mm, gfp_t gfp) { diff --git a/mm/hugetlb.c b/mm/hugetlb.c index 15f9c5a9f75e2..5ee1bc5c00bfe 100644 --- a/mm/hugetlb.c +++ b/mm/hugetlb.c @@ -38,6 +38,7 @@ #include #include #include +#include =20 #include #include @@ -2876,6 +2877,8 @@ struct folio *alloc_hugetlb_folio(struct vm_area_stru= ct *vma, int ret, idx; struct hugetlb_cgroup *h_cg =3D NULL; struct hugetlb_cgroup *h_cg_rsvd =3D NULL; + struct mem_cgroup *mem_cg =3D NULL; + struct obj_cgroup *obj_cg =3D NULL; gfp_t gfp =3D htlb_alloc_mask(h) | __GFP_RETRY_MAYFAIL; =20 idx =3D hstate_index(h); @@ -2935,6 +2938,11 @@ struct folio *alloc_hugetlb_folio(struct vm_area_str= uct *vma, if (ret) goto out_uncharge_cgroup_reservation; =20 + ret =3D mem_cgroup_hugetlb_try_charge(pages_per_huge_page(h), gfp, + &mem_cg, &obj_cg); + if (ret) + goto out_uncharge_cgroup; + spin_lock_irq(&hugetlb_lock); /* * glb_chg is passed to indicate whether or not a page must be taken @@ -2946,7 +2954,7 @@ struct folio *alloc_hugetlb_folio(struct vm_area_stru= ct *vma, spin_unlock_irq(&hugetlb_lock); folio =3D alloc_buddy_hugetlb_folio_with_mpol(h, vma, addr); if (!folio) - goto out_uncharge_cgroup; + goto out_uncharge_cgroup_memcg; spin_lock_irq(&hugetlb_lock); list_add(&folio->lru, &h->hugepage_activelist); folio_ref_unfreeze(folio, 1); @@ -2973,6 +2981,9 @@ struct folio *alloc_hugetlb_folio(struct vm_area_stru= ct *vma, =20 spin_unlock_irq(&hugetlb_lock); =20 + mem_cgroup_hugetlb_commit_charge(folio, mem_cg, obj_cg); + lruvec_stat_mod_folio(folio, NR_HUGETLB, pages_per_huge_page(h)); + hugetlb_set_folio_subpool(folio, spool); =20 if (map_chg !=3D MAP_CHG_ENFORCED) { @@ -3000,21 +3011,10 @@ struct folio *alloc_hugetlb_folio(struct vm_area_st= ruct *vma, } } =20 - ret =3D mem_cgroup_charge_hugetlb(folio, gfp); - /* - * Unconditionally increment NR_HUGETLB here. If it turns out that - * mem_cgroup_charge_hugetlb failed, then immediately free the page and - * decrement NR_HUGETLB. - */ - lruvec_stat_mod_folio(folio, NR_HUGETLB, pages_per_huge_page(h)); - - if (ret =3D=3D -ENOMEM) { - free_huge_folio(folio); - goto err; - } - return folio; =20 +out_uncharge_cgroup_memcg: + mem_cgroup_hugetlb_cancel_charge(pages_per_huge_page(h), mem_cg, obj_cg); out_uncharge_cgroup: hugetlb_cgroup_uncharge_cgroup(idx, pages_per_huge_page(h), h_cg); out_uncharge_cgroup_reservation: @@ -3035,7 +3035,7 @@ struct folio *alloc_hugetlb_folio(struct vm_area_stru= ct *vma, out_end_reservation: if (map_chg !=3D MAP_CHG_ENFORCED) vma_end_reservation(h, vma, addr); -err: + /* * Return -ENOSPC when this function fails to allocate or charge a huge * page. If a standard (PAGE_SIZE) page allocation fails, the OOM killer diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 6dc4888a90f3f..0beee5c0ce93b 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -5180,6 +5180,120 @@ int mem_cgroup_charge_hugetlb(struct folio *folio, = gfp_t gfp) return ret; } =20 +/** + * mem_cgroup_hugetlb_try_charge - Try to charge the memcg for a hugetlb f= olio + * @nr_pages: number of base pages to charge + * @gfp: reclaim mode + * @memcg_p: Output pointer to the charged mem_cgroup (if successful and e= nabled) + * @objcg_p: Output pointer to the charged obj_cgroup (if successful and e= nabled) + * + * Prepares and tries to reserve the memory counter for the folio from the= current + * task's memcg. If successful, both *memcg_p and *objcg_p are populated a= nd their + * references are pinned until a subsequent call to mem_cgroup_hugetlb_com= mit_charge + * or mem_cgroup_hugetlb_cancel_charge. + * + * Returns ENOMEM if the memcg is already full. + * Returns 0 if either the charge was successful, or if we skip charging. + */ +int mem_cgroup_hugetlb_try_charge(unsigned int nr_pages, gfp_t gfp, + struct mem_cgroup **memcg_p, + struct obj_cgroup **objcg_p) +{ + struct mem_cgroup *memcg; + struct obj_cgroup *objcg; + int ret =3D 0; + + *memcg_p =3D NULL; + *objcg_p =3D NULL; + + if (mem_cgroup_disabled() || !memcg_accounts_hugetlb() || + !cgroup_subsys_on_dfl(memory_cgrp_subsys)) + return 0; + + memcg =3D get_mem_cgroup_from_current(); + if (!memcg) + return 0; + + objcg =3D get_obj_cgroup_from_memcg(memcg); + if (!objcg) + goto put_memcg; + + if (!obj_cgroup_is_root(objcg)) { + ret =3D try_charge_memcg(memcg, gfp, nr_pages); + if (ret) + goto put_objcg; + } + + *memcg_p =3D memcg; + *objcg_p =3D objcg; + return 0; + +put_objcg: + obj_cgroup_put(objcg); +put_memcg: + mem_cgroup_put(memcg); + return ret; +} + +/** + * mem_cgroup_hugetlb_commit_charge - Commit the memcg charge for a hugetl= b folio + * @folio: folio being charged + * @memcg: Target mem_cgroup obtained from mem_cgroup_hugetlb_try_charge + * @objcg: Target obj_cgroup obtained from mem_cgroup_hugetlb_try_charge + * + * Finalizes the memory and statistics charging for the folio in the speci= fied memcg. + * Transfers the pinned objcg reference to the folio structure (for automa= tic + * uncharging upon freeing via mem_cgroup_uncharge). Releases the try-comm= it reference + * on memcg. + */ +void mem_cgroup_hugetlb_commit_charge(struct folio *folio, + struct mem_cgroup *memcg, + struct obj_cgroup *objcg) +{ + if (!memcg || !objcg) + return; + + commit_charge(folio, objcg); + memcg1_commit_charge(folio, memcg); + + /* + * Drop our try-commit-cancel protocol reference on memcg. + * The objcg reference is TRANSFERRED to the folio by commit_charge, + * so it will be put automatically by __mem_cgroup_uncharge() when + * the folio is freed. + */ + mem_cgroup_put(memcg); +} + +/** + * mem_cgroup_hugetlb_cancel_charge - Cancel and undo a hugetlb folio memc= g charge + * @nr_pages: number of base pages to uncharge + * @memcg: Target mem_cgroup obtained from mem_cgroup_hugetlb_try_charge + * @objcg: Target obj_cgroup obtained from mem_cgroup_hugetlb_try_charge + * + * Cancels and safely rolls back the prepared memory charge for the folio = in the + * specified memcg. Releases the try-commit pinned references on both memc= g and objcg. + */ +void mem_cgroup_hugetlb_cancel_charge(unsigned int nr_pages, + struct mem_cgroup *memcg, + struct obj_cgroup *objcg) +{ + if (!memcg || !objcg) + return; + + if (!obj_cgroup_is_root(objcg)) + refill_stock(memcg, nr_pages); + + /* + * Drop our try-commit-cancel protocol references on both objcg + * and memcg, since this mapping attempt was aborted and the folio + * was never committed. + */ + obj_cgroup_put(objcg); + mem_cgroup_put(memcg); +} + + /** * mem_cgroup_swapin_charge_folio - Charge a newly allocated folio for swa= pin. * @folio: the folio to charge --=20 2.55.0.229.g6434b31f56-goog From nobody Fri Jul 24 22:17:47 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 372D5437464; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763680; cv=none; b=t7kc2rIWBgXI0G6nMsNXuAuHaWljNRyOF5LTiTUOyEI6NpDH8l7UgpdGvifAUV1aSNAfolKCnmerXRwnd0YTR8mvleFKk7UVkXtguIt1I3vmgHxBhHqnPLiVWfV3wa9Y4e8Et3UdGy1NMZcRhpLZStcECOAroC9m8PCDaFqs0t4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763680; c=relaxed/simple; bh=6/kMJwdQl/TqCpTYlCewyDf5fd6mj5P5o/GrXeUXOXQ=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=iTbhyg763Y7Kr/W71omEqK79Y7jkejz1QLWpl9s7y48+JZmFuo4lqQ/EVCMcSBIhUHv5aFZ3yZDoHvZxquZ0FYH21qEjP3okIIx00rG0SG8Aj5+HPYMaVnCujsv4mU/BRXWTsNTsxOM+qZXVLeGrQIKpHFA9qsqXsyx2yyGWJOo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=H8ms7azd; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="H8ms7azd" Received: by smtp.kernel.org (Postfix) with ESMTPS id 08804C2BCFF; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1784763680; bh=6/kMJwdQl/TqCpTYlCewyDf5fd6mj5P5o/GrXeUXOXQ=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=H8ms7azd8/HQ5RN3MfbnfPoRlWed2T0o6iRmz6iFwuDEwzZbiHQEERWCy15nejg7W D7UYxa5nPIPcxre9UMbeD42BSP3GA40bTII3rliiNQW/SdvbsI8G3tj99arLhMbgF9 mNYmyjWEG4KlC2j0X+320mRNOvAyWKy/IZXrRqfP2iJfrBhgt7a4u19AzCFnoj7gF2 z3W4QyiChA0TJO+T/xNSZxUvyitvV/lyOTubwbaU0+K34+rXWC6EiA8WTBGXPSBNQ4 4niAVgUbJLOdWbOBppwOmth3vioaP0oqdEk06oVXPIyYpa9vtwxPh5owsF/uph+1a7 CV+lTvb/vVsyg== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id E8AB8C4453D; Wed, 22 Jul 2026 23:41:19 +0000 (UTC) From: Ackerley Tng via B4 Relay Date: Wed, 22 Jul 2026 16:41:12 -0700 Subject: [PATCH v4 04/16] mm: hugetlb: Remove unused mem_cgroup_charge_hugetlb function Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260722-hugetlb-alloc-failure-fixes-v4-4-88e8b81970dc@google.com> References: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> In-Reply-To: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> To: Muchun Song , Oscar Salvador , David Hildenbrand , Joshua Hahn , Shakeel Butt , Nhat Pham , Andrew Morton , Peter Xu , Wupeng Ma , fvdl@google.com, rientjes@google.com, jthoughton@google.com, Mike Kravetz , Johannes Weiner , Michal Hocko , Roman Gushchin , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jonathan Corbet , Shuah Khan , Alex Shi , Yanteng Si , Dongliang Mu , Hongxiang Lou , Miaohe Lin Cc: vannapurve@google.com, erdemaktas@google.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, linux-doc@vger.kernel.org, Ackerley Tng X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=ed25519-sha256; t=1784763678; l=2865; i=ackerleytng@google.com; s=20260225; h=from:subject:message-id; bh=+gjdnJUXRCMSVp+BWhjDe3KtAs27/6SLSiAQUKfU/dw=; b=kqwcwPC/GCxB/BUQtS9LtL+7g51a/az0dlTIwy29hV9Q3jlMV9rROOA/Qw2WWy3T+Z0ixUMuD ibAIP93u5+dB/8Db6AodxtWs1Ks5khIBrdF7HEQh5/cBr57/m87R30e X-Developer-Key: i=ackerleytng@google.com; a=ed25519; pk=sAZDYXdm6Iz8FHitpHeFlCMXwabodTm7p8/3/8xUxuU= X-Endpoint-Received: by B4 Relay for ackerleytng@google.com/20260225 with auth_id=649 X-Original-From: Ackerley Tng Reply-To: ackerleytng@google.com From: Ackerley Tng Now that the alloc_hugetlb_folio path has been successfully migrated to the new try-commit-cancel memcg charging protocol, the old mem_cgroup_charge_hugetlb function and its associated header and static inline declarations are completely unused. Remove them to clean up the memory controller's codebase. Signed-off-by: Ackerley Tng --- include/linux/memcontrol.h | 6 ------ mm/memcontrol.c | 34 ---------------------------------- 2 files changed, 40 deletions(-) diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h index 2c3e2c62f4176..3bdf35dc2d88f 100644 --- a/include/linux/memcontrol.h +++ b/include/linux/memcontrol.h @@ -641,7 +641,6 @@ static inline int mem_cgroup_charge(struct folio *folio= , struct mm_struct *mm, return __mem_cgroup_charge(folio, mm, gfp); } =20 -int mem_cgroup_charge_hugetlb(struct folio* folio, gfp_t gfp); int mem_cgroup_hugetlb_try_charge(unsigned int nr_pages, gfp_t gfp, struct mem_cgroup **memcg_p, struct obj_cgroup **objcg_p); @@ -1133,11 +1132,6 @@ static inline bool mem_cgroup_below_min(struct mem_c= group *target, =20 static inline int mem_cgroup_charge(struct folio *folio, struct mm_struct *mm, gfp_t gfp) -{ - return 0; -} - -static inline int mem_cgroup_charge_hugetlb(struct folio* folio, gfp_t gfp) { return 0; } diff --git a/mm/memcontrol.c b/mm/memcontrol.c index 0beee5c0ce93b..6764ff041c196 100644 --- a/mm/memcontrol.c +++ b/mm/memcontrol.c @@ -5146,40 +5146,6 @@ int __mem_cgroup_charge(struct folio *folio, struct = mm_struct *mm, gfp_t gfp) return ret; } =20 -/** - * mem_cgroup_charge_hugetlb - charge the memcg for a hugetlb folio - * @folio: folio being charged - * @gfp: reclaim mode - * - * This function is called when allocating a huge page folio, after the pa= ge has - * already been obtained and charged to the appropriate hugetlb cgroup - * controller (if it is enabled). - * - * Returns ENOMEM if the memcg is already full. - * Returns 0 if either the charge was successful, or if we skip the chargi= ng. - */ -int mem_cgroup_charge_hugetlb(struct folio *folio, gfp_t gfp) -{ - struct mem_cgroup *memcg =3D get_mem_cgroup_from_current(); - int ret =3D 0; - - /* - * Even memcg does not account for hugetlb, we still want to update - * system-level stats via lruvec_stat_mod_folio. Return 0, and skip - * charging the memcg. - */ - if (mem_cgroup_disabled() || !memcg_accounts_hugetlb() || - !memcg || !cgroup_subsys_on_dfl(memory_cgrp_subsys)) - goto out; - - if (charge_memcg(folio, memcg, gfp)) - ret =3D -ENOMEM; - -out: - mem_cgroup_put(memcg); - return ret; -} - /** * mem_cgroup_hugetlb_try_charge - Try to charge the memcg for a hugetlb f= olio * @nr_pages: number of base pages to charge --=20 2.55.0.229.g6434b31f56-goog From nobody Fri Jul 24 22:17:47 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5D7EC459AF4; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763680; cv=none; b=bW0t2KDbkKJIxB41Wt3blZJ6dSMA6QolLjQ6sxxltj+boymXR3IJ/GAaoOXIKYsdKGZQ9H3hz2iS9W+ttx1CBITvdW5ZRMV2pI0abH5Kfora7BHAsu2XciyxxhJEUIZrp371ZhVItDcbV+NrsUPRJaXk8d+f6d0Z6gXYvxyDe9c= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763680; c=relaxed/simple; bh=v842uTAMaAywtG7oyLP8iPjSSWfoAFBKwZA8emE/Hno=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=ZUWOzPZ0LNQqaxby3Gf7qGMzRm6It8xcO0UH52YOFtIw45weZtFHZye1bJA2xVxRmZS9S9Q32egg8Lv35LnU9HZetfIirEM/74lZRgoz5TkiW+DgrqHKUm9w1N+rS4mcGyipMrddhc0hxiM/SjLfIAocFGA+68mKbBb/y/S0HMs= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=SYOW2JrO; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="SYOW2JrO" Received: by smtp.kernel.org (Postfix) with ESMTPS id 1A708C2BD04; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1784763680; bh=v842uTAMaAywtG7oyLP8iPjSSWfoAFBKwZA8emE/Hno=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=SYOW2JrOWetwvA2auCJtx5rX+NMEj8Shzph9S9DnmswH+JqUzRxXseW+OGjBOzayv XEPm9ewTfziBv9Ftg0n8CKdQ8EVzU535UWxwerrcpG3ySn8PTVfJN/dTxaDLStsY/K 9no8aN4qTHc+CGlg/LWll/ynR9E0mx2+7P2EUf+NhvlZ8L6+TRVBPHlcqb8vVba6+Y 1/G4QspWHW9SPax4fTW79WZpCBMQ1OK4ymihO6PF63IdixgHm/eCepa2kXU4NRKUJW +1nGXHho5KPL1LqY5ZZGI80sacshSLyd6MJvWjaD5WBlKffRQXJuCcrwRJDCKoefE4 ug3Y7cppIoQZw== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 07A7EC4453F; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) From: Ackerley Tng via B4 Relay Date: Wed, 22 Jul 2026 16:41:13 -0700 Subject: [PATCH v4 05/16] mm: hugetlb: Fix subpool usage leak on allocation failure Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260722-hugetlb-alloc-failure-fixes-v4-5-88e8b81970dc@google.com> References: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> In-Reply-To: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> To: Muchun Song , Oscar Salvador , David Hildenbrand , Joshua Hahn , Shakeel Butt , Nhat Pham , Andrew Morton , Peter Xu , Wupeng Ma , fvdl@google.com, rientjes@google.com, jthoughton@google.com, Mike Kravetz , Johannes Weiner , Michal Hocko , Roman Gushchin , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jonathan Corbet , Shuah Khan , Alex Shi , Yanteng Si , Dongliang Mu , Hongxiang Lou , Miaohe Lin Cc: vannapurve@google.com, erdemaktas@google.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, linux-doc@vger.kernel.org, Ackerley Tng , stable@vger.kernel.org X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=ed25519-sha256; t=1784763678; l=4429; i=ackerleytng@google.com; s=20260225; h=from:subject:message-id; bh=aioziCwNl1YhUSW9QJsopoiM0FNET1b5J61jDqmRBgE=; b=EgFMKfYcEA2o7aDdaYSsd6UI7TTvR5GAm70ONUg9KMMDn45zAWhKmQCqMLfU8tl73gxpDjjB4 l7LDr/Y+kTGDj3eEZeDlBJW0M2vKZYYopIAGmS2bhziu7J6FHpt5CR0 X-Developer-Key: i=ackerleytng@google.com; a=ed25519; pk=sAZDYXdm6Iz8FHitpHeFlCMXwabodTm7p8/3/8xUxuU= X-Endpoint-Received: by B4 Relay for ackerleytng@google.com/20260225 with auth_id=649 X-Original-From: Ackerley Tng Reply-To: ackerleytng@google.com From: Ackerley Tng When alloc_hugetlb_folio() fails early (e.g. buddy allocation failure or hugetlb cgroup charging failure) and gbl_chg =3D=3D 1 (meaning a reservation was not used, but a global page was allocated instead), the subpool page acquired via hugepage_subpool_get_pages() must still be returned. Currently, the error path out_subpool_put: only calls hugepage_subpool_put_pages() if !gbl_chg is true. If gbl_chg is 1, it skips it, permanently leaking the subpool's used_hpages counter. With the earlier patch to always track used_hpages in the subpool, always call hugepage_subpool_put_pages() if map_chg is true to consistently restore the page to the subpool. Condition on map_chg because map_chg also gates hugepage_subpool_get_pages(). Opportunistically rename gbl_chg to gbl_resv_get, which is the number of global pages needed (because neither resv_map nor subpool reservations could be used). If the number of global pages needed is 0, this allocation uses a reservation somewhere, hence proceed to consume a reservation by decrementing h->resv_huge_pages. If the number of global pages needed is 1, reservations are neither created nor consumed. Also rename gbl_reserve to gbl_resv_put, which is the number of pages the subpool could not absorb into its reservations. Adjust global reservations using hugetlb_acct_memory() with the difference between gbl_resv_get and gbl_resv_put to take care of possible races where another thread might have performed get or put with the same subpool, hence perhaps requiring updates to global reservation counts. (If gbl_resv_get = =3D=3D 0 because a resv_map reservation was used, map_chg =3D=3D 0 so the entire subpool returning is skipped - still correct.) Fixes: a833a693a490e ("mm: hugetlb: fix incorrect fallback for subpool") Cc: stable@vger.kernel.org Signed-off-by: Ackerley Tng --- mm/hugetlb.c | 23 ++++++++++------------- 1 file changed, 10 insertions(+), 13 deletions(-) diff --git a/mm/hugetlb.c b/mm/hugetlb.c index 5ee1bc5c00bfe..879e4640dc50d 100644 --- a/mm/hugetlb.c +++ b/mm/hugetlb.c @@ -2872,7 +2872,7 @@ struct folio *alloc_hugetlb_folio(struct vm_area_stru= ct *vma, struct hugepage_subpool *spool =3D subpool_vma(vma); struct hstate *h =3D hstate_vma(vma); struct folio *folio; - long retval, gbl_chg, gbl_reserve; + long retval, gbl_resv_get; map_chg_state map_chg; int ret, idx; struct hugetlb_cgroup *h_cg =3D NULL; @@ -2912,15 +2912,15 @@ struct folio *alloc_hugetlb_folio(struct vm_area_st= ruct *vma, * Or if it can get one from the pool reservation directly. */ if (map_chg) { - gbl_chg =3D hugepage_subpool_get_pages(spool, 1); - if (gbl_chg < 0) + gbl_resv_get =3D hugepage_subpool_get_pages(spool, 1); + if (gbl_resv_get < 0) goto out_end_reservation; } else { /* * If we have the vma reservation ready, no need for extra * global reservation. */ - gbl_chg =3D 0; + gbl_resv_get =3D 0; } =20 /* @@ -2949,7 +2949,7 @@ struct folio *alloc_hugetlb_folio(struct vm_area_stru= ct *vma, * from the global free pool (global change). gbl_chg =3D=3D 0 indicates * a reservation exists for the allocation. */ - folio =3D dequeue_hugetlb_folio_vma(h, vma, addr, gbl_chg); + folio =3D dequeue_hugetlb_folio_vma(h, vma, addr, gbl_resv_get); if (!folio) { spin_unlock_irq(&hugetlb_lock); folio =3D alloc_buddy_hugetlb_folio_with_mpol(h, vma, addr); @@ -2965,7 +2965,7 @@ struct folio *alloc_hugetlb_folio(struct vm_area_stru= ct *vma, * Either dequeued or buddy-allocated folio needs to add special * mark to the folio when it consumes a global reservation. */ - if (!gbl_chg) { + if (!gbl_resv_get) { folio_set_hugetlb_restore_reserve(folio); h->resv_huge_pages--; } @@ -3022,13 +3022,10 @@ struct folio *alloc_hugetlb_folio(struct vm_area_st= ruct *vma, hugetlb_cgroup_uncharge_cgroup_rsvd(idx, pages_per_huge_page(h), h_cg_rsvd); out_subpool_put: - /* - * put page to subpool iff the quota of subpool's rsv_hpages is used - * during hugepage_subpool_get_pages. - */ - if (map_chg && !gbl_chg) { - gbl_reserve =3D hugepage_subpool_put_pages(spool, 1); - hugetlb_acct_memory(h, -gbl_reserve); + if (map_chg) { + long gbl_resv_put =3D hugepage_subpool_put_pages(spool, 1); + + hugetlb_acct_memory(h, gbl_resv_get - gbl_resv_put); } =20 =20 --=20 2.55.0.229.g6434b31f56-goog From nobody Fri Jul 24 22:17:47 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5F4D6459AF6; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763680; cv=none; b=SKnEVpQTy5dpPTMY5ghaGHebLtTcmyBUPi8YwjzrKf9OTNwO7wSjA+mLbFsn9dRCdLs4I9OMFIlaRTa+qcFJvjMSywrmPBI74CI7f+83CafmqK2TCsN0NnCd8M501+bUi7UIeo7GlKOBum1A7tzS5aZXvIOyNLgurMGmRUavuvM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763680; c=relaxed/simple; bh=aGbga1PkXCpZ8ED5LqrkeQ+Fq62YN785kMo+rpYCbc8=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=Vy739bPt4DX5CVb2QZWWCv/SOQDUR1Ogi0sHWzL+9Jlwh5VrbQjVHh0TpW0bq2+H/ASqvNM2fllsn3EXUaHXUU6OMaDvhmnrEruC6akqNZNd/8N4zacxSmRj3fAUKPrF2SMloHAwtAc4TaaEiNLNANDnuloi8YJ+tyctSOUklhY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=LZxa0L4P; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="LZxa0L4P" Received: by smtp.kernel.org (Postfix) with ESMTPS id 2B23DC4AF0E; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1784763680; bh=aGbga1PkXCpZ8ED5LqrkeQ+Fq62YN785kMo+rpYCbc8=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=LZxa0L4POGfxU5e7iZ1UcLss+95VxKOFxH8x1G3dKJjbIRCBFNLJVK+UjPN7SWRvp tDCS99qneiWqx5MNNwIwGEwlNh83mxLAsd4eIbiPF1mnOXakLLMnUCzAV3pdYLzVL4 TmhHGFW+R8YrQ0HM56Cr0LwgnI2Cl1++cRPBhL1+zkFNhfnMvTU/LU8i3QXugmBOVk +zH16pmplCubUvwrByZin0CbgdFKMSIuiRI0iYqGJ15PjQKtx1Rua9zM/CDVC/+a2h Vc7J71jFNi/67z33Rv6uFtUxRmJKe/vREwtEm/PZaEjcdAgAiaIwOPns7EOvuREfqo zpDcHEyzLfrMg== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 1852AC44532; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) From: Ackerley Tng via B4 Relay Date: Wed, 22 Jul 2026 16:41:14 -0700 Subject: [PATCH v4 06/16] mm: hugetlb: Rename local variables for clarity in hugetlb_reserve_pages() Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260722-hugetlb-alloc-failure-fixes-v4-6-88e8b81970dc@google.com> References: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> In-Reply-To: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> To: Muchun Song , Oscar Salvador , David Hildenbrand , Joshua Hahn , Shakeel Butt , Nhat Pham , Andrew Morton , Peter Xu , Wupeng Ma , fvdl@google.com, rientjes@google.com, jthoughton@google.com, Mike Kravetz , Johannes Weiner , Michal Hocko , Roman Gushchin , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jonathan Corbet , Shuah Khan , Alex Shi , Yanteng Si , Dongliang Mu , Hongxiang Lou , Miaohe Lin Cc: vannapurve@google.com, erdemaktas@google.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, linux-doc@vger.kernel.org, Ackerley Tng X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=ed25519-sha256; t=1784763678; l=4405; i=ackerleytng@google.com; s=20260225; h=from:subject:message-id; bh=YN6Ec4dryGn+lNCiGktVw2z+GkC+1qiWsI1wgmafaTw=; b=2bILnGfpQYBpRme36koKemun0e04kP1hxT7XZpV0o3yJRaZSrtupWWZeQRckJ1ePGfhGs4qPM 8BCX7oBBkyyAdb4/2/vTHjmS0gJZT/YU/XnuJHwl3Zt1scXLcXt/uJ9 X-Developer-Key: i=ackerleytng@google.com; a=ed25519; pk=sAZDYXdm6Iz8FHitpHeFlCMXwabodTm7p8/3/8xUxuU= X-Endpoint-Received: by B4 Relay for ackerleytng@google.com/20260225 with auth_id=649 X-Original-From: Ackerley Tng Reply-To: ackerleytng@google.com From: Ackerley Tng In hugetlb_reserve_pages(), the return value of hugepage_subpool_get_pages() was named gbl_resv, while the return value of hugepage_subpool_put_pages() was named gbl_reserve. These variable names were inconsistent with the naming conventions established in alloc_hugetlb_folio(). Rename local variables in hugetlb_reserve_pages() to match naming in alloc_hugetlb_folio(). + Rename gbl_resv to gbl_resv_get, which is the number of global reservations required from the global pool (because neither reservation maps nor pre-existing subpool reservations could cover the request). If gbl_resv_get is 0, the reservation request is fully covered by pre-existing subpool reservations, so no additional global reservations are requested. If gbl_resv_get > 0, it represents the number of additional global reservations required, which are charged to global memory via hugetlb_acct_memory(h, gbl_resv_get). Note: gbl_resv_get in hugetlb_reserve_pages() is conceptually slightly different from gbl_resv_get in alloc_hugetlb_folio(): + In hugetlb_reserve_pages(): additional _reservations_ required, accounted using hugetlb_acct_memory() + In alloc_hugetlb_folio(): additional _pages_ required, no change to h->resv_huge_pages if gbl_resv_get > 0 + Rename gbl_reserve to gbl_resv_put, which is the number of pages the subpool could not absorb into its reservations during rollback (and are thus returned to the global pool). In error rollback paths, hugetlb_acct_memory(h, gbl_resv_get - gbl_resv_put) adjusts global reservations using the difference between global pages requested (gbl_resv_get) and global pages returned (gbl_resv_put). Signed-off-by: Ackerley Tng --- mm/hugetlb.c | 19 ++++++++++--------- 1 file changed, 10 insertions(+), 9 deletions(-) diff --git a/mm/hugetlb.c b/mm/hugetlb.c index 879e4640dc50d..8431c00d48267 100644 --- a/mm/hugetlb.c +++ b/mm/hugetlb.c @@ -6562,12 +6562,13 @@ long hugetlb_reserve_pages(struct inode *inode, struct vm_area_struct *vma, vma_flags_t vma_flags) { - long chg =3D -1, add =3D -1, gbl_resv; + long chg =3D -1, add =3D -1; struct hstate *h =3D hstate_inode(inode); struct hugepage_subpool *spool =3D subpool_inode(inode); struct resv_map *resv_map; struct hugetlb_cgroup *h_cg =3D NULL; - long gbl_reserve, regions_needed =3D 0; + long gbl_resv_get, gbl_resv_put; + long regions_needed =3D 0; int err; =20 /* This should never happen */ @@ -6642,9 +6643,9 @@ long hugetlb_reserve_pages(struct inode *inode, * the subpool has a minimum size, there may be some global * reservations already in place (gbl_reserve). */ - gbl_resv =3D hugepage_subpool_get_pages(spool, chg); - if (gbl_resv < 0) { - err =3D gbl_resv; + gbl_resv_get =3D hugepage_subpool_get_pages(spool, chg); + if (gbl_resv_get < 0) { + err =3D gbl_resv_get; goto out_uncharge_cgroup; } =20 @@ -6652,7 +6653,7 @@ long hugetlb_reserve_pages(struct inode *inode, * Check enough hugepages are available for the reservation. * Hand the pages back to the subpool if there are not */ - err =3D hugetlb_acct_memory(h, gbl_resv); + err =3D hugetlb_acct_memory(h, gbl_resv_get); if (err < 0) goto out_put_pages; =20 @@ -6671,7 +6672,7 @@ long hugetlb_reserve_pages(struct inode *inode, add =3D region_add(resv_map, from, to, regions_needed, h, h_cg); =20 if (unlikely(add < 0)) { - hugetlb_acct_memory(h, -gbl_resv); + hugetlb_acct_memory(h, -gbl_resv_get); err =3D add; goto out_put_pages; } else if (unlikely(chg > add)) { @@ -6713,14 +6714,14 @@ long hugetlb_reserve_pages(struct inode *inode, * tell us the new number of reservations that need to be * returned to the global pool. */ - gbl_reserve =3D hugepage_subpool_put_pages(spool, chg); + gbl_resv_put =3D hugepage_subpool_put_pages(spool, chg); /* * There may be a difference between the number of * reservations to consume and the number to restore now if * there are multiple threads interacting with the subpool - * restore the difference. */ - hugetlb_acct_memory(h, gbl_resv - gbl_reserve); + hugetlb_acct_memory(h, gbl_resv_get - gbl_resv_put); =20 out_uncharge_cgroup: hugetlb_cgroup_uncharge_cgroup_rsvd(hstate_index(h), --=20 2.55.0.229.g6434b31f56-goog From nobody Fri Jul 24 22:17:47 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 748D845A298; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763680; cv=none; b=KYX6blh/Oo3WyRRCMVQypesM9c2sE92XE1NicfzvWvcdp76Mzx4aksLhzl941ech4qle+GIQQlwP0PcQcPWI6Af1N64Pguo/3WQYYqN+mF6TLLgAVqXH6LEii/sk7DdV6iwHQdMh495LImdjjITGr53Ij2joJsuzElorufAQKmQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763680; c=relaxed/simple; bh=5jfm2MbqP71BsjfPjrd0Br42vvUHXqP+s/GdXagyPXY=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=nMPWMPSgNZW3rgn3dQxeNEqPCiv5xpHR9h6K6UhMnF8zeFQmGVBcX0F6xZDA7Lyo1zU+wNj86/WXUVdoT/xI/H8L3k8irgbwb9Lf9PmkbM1okyeDtvGDok6aV1uuF5BDGo1UGg2ysTqqbDtJPNNS35zX6br+cTBL5WIy6jA8XXI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=TFzoNq87; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="TFzoNq87" Received: by smtp.kernel.org (Postfix) with ESMTPS id 4A4D3C19425; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1784763680; bh=5jfm2MbqP71BsjfPjrd0Br42vvUHXqP+s/GdXagyPXY=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=TFzoNq87S38GGo6Ssm7lOwhxQZMCtuARuGEPFPG4IS5BpH2vy5rIMSFiaSyYpl4mJ Xm0VJdnaW6UqgomQ7nAzsoeX2WAIQnZ8lT26S5m7JpYb1ESNUWOYt1V8cx4zYeqQLY sED7WuyAx3Q/WXzc1/6v6MYjVuPRnBB9ahGBSYFR90eqZ1Bps6JovLDKWQ2XPcuEjy h9tq65Q6RqTBhyMp3VmgKPnOLD1k9HsiVZpHD6nlbAwyV+X5dUnm/GHgtuNoqeYTZb pCqBsZimkvBmjF9jR1I9VI6amHsMAUhQ436dSFTfZ79Qe3D41HpWxv80/Yo3E4sdx5 FVhF9YAYEb2OQ== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 35563C4453A; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) From: Ackerley Tng via B4 Relay Date: Wed, 22 Jul 2026 16:41:15 -0700 Subject: [PATCH v4 07/16] mm: hugetlb: Fix Use-After-Free in unlock_or_release_subpool() Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260722-hugetlb-alloc-failure-fixes-v4-7-88e8b81970dc@google.com> References: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> In-Reply-To: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> To: Muchun Song , Oscar Salvador , David Hildenbrand , Joshua Hahn , Shakeel Butt , Nhat Pham , Andrew Morton , Peter Xu , Wupeng Ma , fvdl@google.com, rientjes@google.com, jthoughton@google.com, Mike Kravetz , Johannes Weiner , Michal Hocko , Roman Gushchin , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jonathan Corbet , Shuah Khan , Alex Shi , Yanteng Si , Dongliang Mu , Hongxiang Lou , Miaohe Lin Cc: vannapurve@google.com, erdemaktas@google.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, linux-doc@vger.kernel.org, Ackerley Tng , stable@vger.kernel.org X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=ed25519-sha256; t=1784763678; l=1786; i=ackerleytng@google.com; s=20260225; h=from:subject:message-id; bh=+8BKlqcw0osweyRZX4DVMuCuKwenNXcfBjXVW5HQiFI=; b=1QJ1v6HnKCN4SUgfZ0zCwSM1fabe4JN3I0E4aRntYBp8Vk2QcabaSyeAnsKBGGeTC3Y76lQen YmRVqUCR2SyD0f88goOEuJb9brctkXRJonc3/dbLVzTo9GlGyTpg6do X-Developer-Key: i=ackerleytng@google.com; a=ed25519; pk=sAZDYXdm6Iz8FHitpHeFlCMXwabodTm7p8/3/8xUxuU= X-Endpoint-Received: by B4 Relay for ackerleytng@google.com/20260225 with auth_id=649 X-Original-From: Ackerley Tng Reply-To: ackerleytng@google.com From: Ackerley Tng In unlock_or_release_subpool(), spool->lock was released before calling subpool_is_free(spool). Because subpool_is_free() accesses spool->count and spool->used_hpages locklessly, concurrent threads calling hugepage_subpool_put_pages() or hugepage_put_subpool() could race when spool->count is 0. If Thread 1 drops spool->lock and is preempted before subpool_is_free(), Thread 2 can acquire spool->lock, decrement spool->used_hpages to 0, evaluate subpool_is_free() as true, and free spool via kfree(). When Thread 1 resumes, subpool_is_free() dereferences the freed spool pointer, causing a Use-After-Free (UAF) or double free. Fix this race by evaluating subpool_is_free() while holding spool->lock prior to unlocking. Fixes: 1d88433bb0085 ("mm/hugetlb: fix use after free when subpool max_hpag= es accounting is not enabled") Cc: stable@vger.kernel.org Signed-off-by: Ackerley Tng --- mm/hugetlb.c | 7 +++---- 1 file changed, 3 insertions(+), 4 deletions(-) diff --git a/mm/hugetlb.c b/mm/hugetlb.c index 8431c00d48267..b759748468734 100644 --- a/mm/hugetlb.c +++ b/mm/hugetlb.c @@ -137,12 +137,11 @@ static inline bool subpool_is_free(struct hugepage_su= bpool *spool) static inline void unlock_or_release_subpool(struct hugepage_subpool *spoo= l, unsigned long irq_flags) { + bool is_free =3D subpool_is_free(spool); + spin_unlock_irqrestore(&spool->lock, irq_flags); =20 - /* If no pages are used, and no other handles to the subpool - * remain, give up any reservations based on minimum size and - * free the subpool */ - if (subpool_is_free(spool)) { + if (is_free) { if (spool->min_hpages !=3D -1) hugetlb_acct_memory(spool->hstate, -spool->min_hpages); --=20 2.55.0.229.g6434b31f56-goog From nobody Fri Jul 24 22:17:47 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8379A45D5C1; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763680; cv=none; b=VjPD8+M71HNM/YbGb5kh1Sn3tIpczAEHM0dFXKAQsGu8fZ6lu3cXyOT8w7xAzX+SsnLXEFhYOFxKiCcizBJvQQOtICRPPW95DdEpGT0Tq4QnTF8PlK0QDouKpP7yFSxSHplAhxabEzhWht04cmyy9sSOLIcgmTSm/T7WoGWpwOE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763680; c=relaxed/simple; bh=pZEyi0J8om2ghDb6qME7uBKDwU9wZvWkyYWuA+nq1e8=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=tV12YoVInPYRpWVXahkvJ5Hhiima1Uc8WvA4HNaODIrV9jWvBm0+6p+YoZDWz2AsBqO0C/i4ySOaP1N8nanEvKN4juGerzQPaFlchdPIm7BRUB0x9r+2w4RvqDJhMS4Pu4ndFoZZdxcYLzMyMT4V9GyyH7tZR8iY60u1xAOX3Lo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=W82BRnUd; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="W82BRnUd" Received: by smtp.kernel.org (Postfix) with ESMTPS id 58061C2BCC6; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1784763680; bh=pZEyi0J8om2ghDb6qME7uBKDwU9wZvWkyYWuA+nq1e8=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=W82BRnUdxyWAQ4f+cvMo4Mzo/wV2rmi5lhO2VF23HZfYcpQlM1lcZFumECqhiTapo DDQcSUd/6W9pwIAL8tUYYQaHBDjrQqJ7zsAw0sWyLEfpPyQurqZksyNlpq9aJ9LN6C 46g6yYXKoivcLALO7nTHcCJ94zTmAajL4Y+ilay0BhY7B0glNXu2KAclYcYK+fuE+U M5YeUICXkGKJ0dst8euM3PQgtEBvQNQzh+9kuz+zSY2ZY1NA60vH5cXn6bXWDl8SZZ ZmE7WLlNM4AaVq8ytPb6gMjSj4OFrvYII3K4E8XEWl9D+Kj/QAeD0cvG/8xQbqvOLW G8Du92zuS5plQ== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 459EFC4453D; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) From: Ackerley Tng via B4 Relay Date: Wed, 22 Jul 2026 16:41:16 -0700 Subject: [PATCH v4 08/16] fs: hugetlbfs: Fix global reservation leak in hugetlbfs_fill_super() Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260722-hugetlb-alloc-failure-fixes-v4-8-88e8b81970dc@google.com> References: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> In-Reply-To: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> To: Muchun Song , Oscar Salvador , David Hildenbrand , Joshua Hahn , Shakeel Butt , Nhat Pham , Andrew Morton , Peter Xu , Wupeng Ma , fvdl@google.com, rientjes@google.com, jthoughton@google.com, Mike Kravetz , Johannes Weiner , Michal Hocko , Roman Gushchin , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jonathan Corbet , Shuah Khan , Alex Shi , Yanteng Si , Dongliang Mu , Hongxiang Lou , Miaohe Lin Cc: vannapurve@google.com, erdemaktas@google.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, linux-doc@vger.kernel.org, Ackerley Tng , stable@vger.kernel.org X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=ed25519-sha256; t=1784763678; l=2642; i=ackerleytng@google.com; s=20260225; h=from:subject:message-id; bh=xpwOn4o/8dwkPGcKx/D+08zZJl/SEz5ntK6oD42SVfo=; b=GfSo+082yV0XV2379mD54lWB1yFERnvv5vJcHzxcpA+vOTbZSFPHC1fJRtNvwaK1ItH4SHmMa P9nAWPJRDfrDUTI9SSn5NKpQmgKTijYg6y0r21Dm4bc+LzKU2zgTomV X-Developer-Key: i=ackerleytng@google.com; a=ed25519; pk=sAZDYXdm6Iz8FHitpHeFlCMXwabodTm7p8/3/8xUxuU= X-Endpoint-Received: by B4 Relay for ackerleytng@google.com/20260225 with auth_id=649 X-Original-From: Ackerley Tng Reply-To: ackerleytng@google.com From: Ackerley Tng In hugetlbfs_fill_super(), if hugepage_new_subpool() succeeds in allocating a subpool (which reserves global huge pages when min_hpages !=3D -1), but a subsequent initialization step like d_make_root() fails, the error path directly invoked kfree(sbinfo->spool). Directly freeing the subpool structure with kfree() bypasses hugepage_put_subpool() and its underlying unlock_or_release_subpool() destructor. Consequently, hugetlb_acct_memory() is never called to release the global page reservations allocated for min_hpages, permanently leaking global huge page reservations. Fix this leak by calling hugepage_put_subpool(sbinfo->spool) on the error cleanup path instead of kfree(sbinfo->spool). sbinfo->spool must be checked before dereferencing the subpool in hugepage_put_subpool(). Add a NULL guard to hugepage_put_subpool() instead of checking it in the caller to align it with other subpool helpers like hugepage_subpool_get_pages() and hugepage_subpool_put_pages() that gracefully handle NULL subpools. This allows callers to safely invoke hugepage_put_subpool() without requiring explicit NULL checks. This is also aligned with how kfree() can be called on NULL. With the NULL guard in hugepage_put_subpool(), the NULL check in the only other caller can also be removed. Fixes: 7ca02d0ae586f ("hugetlbfs: accept subpool min_size mount option and = setup accordingly") Cc: stable@vger.kernel.org Signed-off-by: Ackerley Tng --- fs/hugetlbfs/inode.c | 6 ++---- mm/hugetlb.c | 3 +++ 2 files changed, 5 insertions(+), 4 deletions(-) diff --git a/fs/hugetlbfs/inode.c b/fs/hugetlbfs/inode.c index 26c0187340636..e5d86f31eba5b 100644 --- a/fs/hugetlbfs/inode.c +++ b/fs/hugetlbfs/inode.c @@ -1133,9 +1133,7 @@ static void hugetlbfs_put_super(struct super_block *s= b) if (sbi) { sb->s_fs_info =3D NULL; =20 - if (sbi->spool) - hugepage_put_subpool(sbi->spool); - + hugepage_put_subpool(sbi->spool); kfree(sbi); } } @@ -1423,7 +1421,7 @@ hugetlbfs_fill_super(struct super_block *sb, struct f= s_context *fc) goto out_free; return 0; out_free: - kfree(sbinfo->spool); + hugepage_put_subpool(sbinfo->spool); kfree(sbinfo); return -ENOMEM; } diff --git a/mm/hugetlb.c b/mm/hugetlb.c index b759748468734..90ec015a11181 100644 --- a/mm/hugetlb.c +++ b/mm/hugetlb.c @@ -177,6 +177,9 @@ void hugepage_put_subpool(struct hugepage_subpool *spoo= l) { unsigned long flags; =20 + if (!spool) + return; + spin_lock_irqsave(&spool->lock, flags); BUG_ON(!spool->count); spool->count--; --=20 2.55.0.229.g6434b31f56-goog From nobody Fri Jul 24 22:17:47 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8BE1F45D5C5; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763680; cv=none; b=QXR+SOk4aW+0xfWQFY8mC8kEla+hKtn9wvMBzlcQSu+FOgZl8K2DpZC3Vb82CsYzXGOLWFKh2qExY5vPyKfEGIUbyyFl5Ng5/AAo0MhDfaKmNUPuJ8lzikWKO1h6dAHQp1ZdGHX1TTvcaN3589velOOiQUyjlxJMDBY+XKAox1I= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763680; c=relaxed/simple; bh=MilV9fZJGpzCk1sO4io/oFZaBmMxYeYIbPlBrtI2loQ=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=RhY/W76BlRB/nSOPOUWpfAH2Ro/LwAnSbYlRLJ5qFzYyQDry1ql+Wdy5mOeGDStQu+kvVHQeQr11pgdUNAZn/Bz2CvCkokIQZA+RzQx85mnja9dwy3oxP9qtctPHkHno083E9FZpprH0y3kF3G3deSuF+F4VWFel3Jb4TrBw/+Y= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=Hy7Ze70k; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="Hy7Ze70k" Received: by smtp.kernel.org (Postfix) with ESMTPS id 6B6C0C2BCFA; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1784763680; bh=MilV9fZJGpzCk1sO4io/oFZaBmMxYeYIbPlBrtI2loQ=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=Hy7Ze70keZSOblWcDNHKSB1QMiRhBvw2OW8lMbPuPpfzf/JEFPzupHNDwhMNxd/be vPh+IYCvdEIJHXK6GV52PLXhAOiy7gdRcx3s31KrC/2UAm/InxVPMroJmvP5Adf3we HM2mvn2zXr1aZspnfQxHshSSO21o6LYqy+s8jJZ2I67GF6xMpmHRU7vW5q7uyfMtgZ uTB9EPiiBI7R2cP/RAXHMtWjycx3zc70kT6+Tg3KlsOjGNnuRWzRy+JjXBoAyslaSQ lLCKnkdxwhvuaN9y/w3E0CGS2/q/+5obAGa04+8WOtyIOTbb9SOiUMi5H00BwP81ge T013GpNMF+TqQ== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 572E6C4453F; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) From: Ackerley Tng via B4 Relay Date: Wed, 22 Jul 2026 16:41:17 -0700 Subject: [PATCH v4 09/16] WIP: mm: hugetlb: Move subpool functions to hugetlb_subpool.c Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260722-hugetlb-alloc-failure-fixes-v4-9-88e8b81970dc@google.com> References: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> In-Reply-To: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> To: Muchun Song , Oscar Salvador , David Hildenbrand , Joshua Hahn , Shakeel Butt , Nhat Pham , Andrew Morton , Peter Xu , Wupeng Ma , fvdl@google.com, rientjes@google.com, jthoughton@google.com, Mike Kravetz , Johannes Weiner , Michal Hocko , Roman Gushchin , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jonathan Corbet , Shuah Khan , Alex Shi , Yanteng Si , Dongliang Mu , Hongxiang Lou , Miaohe Lin Cc: vannapurve@google.com, erdemaktas@google.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, linux-doc@vger.kernel.org, Ackerley Tng X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=ed25519-sha256; t=1784763678; l=12908; i=ackerleytng@google.com; s=20260225; h=from:subject:message-id; bh=2bVmicF4jARvFucIodaWUi9UAIbjHPAM78oVUl0zi9M=; b=WyTseG1sFPnebaIiiv+7bEgnrhySfmRN5qmhMCio1eCAO5c2bvWJb+tCGB/+zz4bKjC8h+nJE hndmaqutKDPBDn/Tx8v4ymim5itrUBAqU935xcjRPsCkCwLmhNDs6lP X-Developer-Key: i=ackerleytng@google.com; a=ed25519; pk=sAZDYXdm6Iz8FHitpHeFlCMXwabodTm7p8/3/8xUxuU= X-Endpoint-Received: by B4 Relay for ackerleytng@google.com/20260225 with auth_id=649 X-Original-From: Ackerley Tng Reply-To: ackerleytng@google.com From: Ackerley Tng Move all HugeTLB subpool lifecycle, page reservation, and accounting routines out of mm/hugetlb.c and into their own dedicated, encapsulated translation unit at mm/hugetlb_subpool.c. Also introduces the internal mm/hugetlb_subpool.h header for holding the subpool-local APIs, allowing fs/hugetlbfs and mm/ to access the subpool functions cleanly. The subpool internal layout structures remain in include/linux/hugetlb.h until getters are introduced. Signed-off-by: Ackerley Tng --- fs/hugetlbfs/inode.c | 1 + include/linux/hugetlb.h | 4 +- mm/Makefile | 2 +- mm/hugetlb.c | 166 +------------------------------------------- mm/hugetlb_subpool.c | 181 ++++++++++++++++++++++++++++++++++++++++++++= ++++ mm/hugetlb_subpool.h | 17 +++++ 6 files changed, 202 insertions(+), 169 deletions(-) diff --git a/fs/hugetlbfs/inode.c b/fs/hugetlbfs/inode.c index e5d86f31eba5b..86c21f8272470 100644 --- a/fs/hugetlbfs/inode.c +++ b/fs/hugetlbfs/inode.c @@ -25,6 +25,7 @@ #include #include #include +#include "../../mm/hugetlb_subpool.h" #include #include #include diff --git a/include/linux/hugetlb.h b/include/linux/hugetlb.h index 34b9a3e1be0fa..f36be371c6e88 100644 --- a/include/linux/hugetlb.h +++ b/include/linux/hugetlb.h @@ -114,9 +114,7 @@ extern int hugetlb_max_hstate __read_mostly; #define for_each_hstate(h) \ for ((h) =3D hstates; (h) < &hstates[hugetlb_max_hstate]; (h)++) =20 -struct hugepage_subpool *hugepage_new_subpool(struct hstate *h, long max_h= pages, - long min_hpages); -void hugepage_put_subpool(struct hugepage_subpool *spool); +int hugetlb_acct_memory(struct hstate *h, long delta); =20 void hugetlb_dup_vma_private(struct vm_area_struct *vma); void clear_vma_resv_huge_pages(struct vm_area_struct *vma); diff --git a/mm/Makefile b/mm/Makefile index eff9f9e7e061c..3965c959e5099 100644 --- a/mm/Makefile +++ b/mm/Makefile @@ -78,7 +78,7 @@ endif obj-$(CONFIG_SWAP) +=3D page_io.o swap_state.o swapfile.o obj-$(CONFIG_ZSWAP) +=3D zswap.o obj-$(CONFIG_HAS_DMA) +=3D dmapool.o -obj-$(CONFIG_HUGETLBFS) +=3D hugetlb.o hugetlb_sysfs.o hugetlb_sysctl.o +obj-$(CONFIG_HUGETLBFS) +=3D hugetlb.o hugetlb_subpool.o hugetlb_sysfs.o h= ugetlb_sysctl.o ifdef CONFIG_CMA obj-$(CONFIG_HUGETLBFS) +=3D hugetlb_cma.o endif diff --git a/mm/hugetlb.c b/mm/hugetlb.c index 90ec015a11181..4d44a9720a971 100644 --- a/mm/hugetlb.c +++ b/mm/hugetlb.c @@ -51,6 +51,7 @@ #include "hugetlb_vmemmap.h" #include "hugetlb_cma.h" #include "hugetlb_internal.h" +#include "hugetlb_subpool.h" #include =20 int hugetlb_max_hstate __read_mostly; @@ -126,171 +127,6 @@ static void hugetlb_unshare_pmds(struct vm_area_struc= t *vma, unsigned long start, unsigned long end, bool take_locks); static struct resv_map *vma_resv_map(struct vm_area_struct *vma); =20 -static inline bool subpool_is_free(struct hugepage_subpool *spool) -{ - if (spool->count) - return false; - - return spool->used_hpages =3D=3D 0; -} - -static inline void unlock_or_release_subpool(struct hugepage_subpool *spoo= l, - unsigned long irq_flags) -{ - bool is_free =3D subpool_is_free(spool); - - spin_unlock_irqrestore(&spool->lock, irq_flags); - - if (is_free) { - if (spool->min_hpages !=3D -1) - hugetlb_acct_memory(spool->hstate, - -spool->min_hpages); - kfree(spool); - } -} - -struct hugepage_subpool *hugepage_new_subpool(struct hstate *h, long max_h= pages, - long min_hpages) -{ - struct hugepage_subpool *spool; - - spool =3D kzalloc_obj(*spool); - if (!spool) - return NULL; - - spin_lock_init(&spool->lock); - spool->count =3D 1; - spool->max_hpages =3D max_hpages; - spool->hstate =3D h; - spool->min_hpages =3D min_hpages; - - if (min_hpages !=3D -1 && hugetlb_acct_memory(h, min_hpages)) { - kfree(spool); - return NULL; - } - spool->rsv_hpages =3D min_hpages; - - return spool; -} - -void hugepage_put_subpool(struct hugepage_subpool *spool) -{ - unsigned long flags; - - if (!spool) - return; - - spin_lock_irqsave(&spool->lock, flags); - BUG_ON(!spool->count); - spool->count--; - unlock_or_release_subpool(spool, flags); -} - -/** - * hugepage_subpool_get_pages - Get pages from a subpool - * @spool: pointer to subpool structure (may be NULL) - * @delta: number of pages to allocate or reserve - * - * Check and update subpool page usage counts when allocating or - * reserving @delta hugepages. - * - * Context: Takes spool->lock using spin_lock_irq(). - * Return: Non-negative number of reservations that cannot be - * satisfied by the subpool, or -ENOMEM if the subpool maximum - * limit would be exceeded. - */ -static long hugepage_subpool_get_pages(struct hugepage_subpool *spool, - long delta) -{ - long ret =3D delta; - - if (!spool) - return ret; - - spin_lock_irq(&spool->lock); - - if (spool->max_hpages !=3D -1 && - spool->used_hpages + delta > spool->max_hpages) { - ret =3D -ENOMEM; - goto unlock_ret; - } - - spool->used_hpages +=3D delta; - - /* minimum size accounting */ - if (spool->min_hpages !=3D -1 && spool->rsv_hpages) { - if (delta > spool->rsv_hpages) { - /* - * Asking for more reserves than those already taken on - * behalf of subpool. Return difference. - */ - ret =3D delta - spool->rsv_hpages; - spool->rsv_hpages =3D 0; - } else { - ret =3D 0; /* reserves already accounted for */ - spool->rsv_hpages -=3D delta; - } - } - -unlock_ret: - spin_unlock_irq(&spool->lock); - return ret; -} - -/** - * hugepage_subpool_put_pages - Release pages back to a subpool - * @spool: pointer to subpool structure (may be NULL) - * @delta: number of pages to free or unreserve - * - * Check and update subpool page usage counts when freeing or - * unreserving @delta hugepages. - * - * Context: Takes spool->lock using spin_lock_irqsave(). May release - * and free @spool if its usage count and references reach - * zero. - * Return: Non-negative number of reservations that the subpool cannot - * absorb. - */ -static long hugepage_subpool_put_pages(struct hugepage_subpool *spool, - long delta) -{ - long ret =3D delta; - unsigned long flags; - - if (!spool) - return delta; - - spin_lock_irqsave(&spool->lock, flags); - - spool->used_hpages -=3D delta; - - /* minimum size accounting */ - if (spool->min_hpages !=3D -1 && spool->used_hpages < spool->min_hpages) { - /* - * limit is the maximum number of reservations that - * can be restored to this subpool. - */ - long limit =3D spool->min_hpages - spool->used_hpages; - - if (spool->rsv_hpages + delta <=3D limit) - ret =3D 0; - else - ret =3D spool->rsv_hpages + delta - limit; - - spool->rsv_hpages +=3D delta; - if (spool->rsv_hpages > limit) - spool->rsv_hpages =3D limit; - } - - /* - * If hugetlbfs_put_super couldn't free spool due to an outstanding - * quota reference, free it now. - */ - unlock_or_release_subpool(spool, flags); - - return ret; -} - static inline struct hugepage_subpool *subpool_vma(struct vm_area_struct *= vma) { return subpool_inode(file_inode(vma->vm_file)); diff --git a/mm/hugetlb_subpool.c b/mm/hugetlb_subpool.c new file mode 100644 index 0000000000000..6184860ed7374 --- /dev/null +++ b/mm/hugetlb_subpool.c @@ -0,0 +1,181 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Subpool and reserve accounting for HugeTLB folios. + * Extracted from mm/hugetlb.c + */ + +#include +#include +#include +#ifdef __KERNEL__ +#include +#endif +#include +#include + +#include "hugetlb_subpool.h" + +static inline bool subpool_is_free(struct hugepage_subpool *spool) +{ + if (spool->count) + return false; + + return spool->used_hpages =3D=3D 0; +} + +static inline void unlock_or_release_subpool(struct hugepage_subpool *spoo= l, + unsigned long irq_flags) +{ + bool is_free =3D subpool_is_free(spool); + + spin_unlock_irqrestore(&spool->lock, irq_flags); + + if (is_free) { + if (spool->min_hpages !=3D -1) + hugetlb_acct_memory(spool->hstate, + -spool->min_hpages); + kfree(spool); + } +} + +struct hugepage_subpool *hugepage_new_subpool(struct hstate *h, long max_h= pages, + long min_hpages) +{ + struct hugepage_subpool *spool; + + spool =3D kzalloc_obj(*spool); + if (!spool) + return NULL; + + spin_lock_init(&spool->lock); + spool->count =3D 1; + spool->max_hpages =3D max_hpages; + spool->hstate =3D h; + spool->min_hpages =3D min_hpages; + + if (min_hpages !=3D -1 && hugetlb_acct_memory(h, min_hpages)) { + kfree(spool); + return NULL; + } + spool->rsv_hpages =3D min_hpages; + + return spool; +} + +void hugepage_put_subpool(struct hugepage_subpool *spool) +{ + unsigned long flags; + + if (!spool) + return; + + spin_lock_irqsave(&spool->lock, flags); + BUG_ON(!spool->count); + spool->count--; + unlock_or_release_subpool(spool, flags); +} + +/** + * hugepage_subpool_get_pages - Get pages from a subpool + * @spool: pointer to subpool structure (may be NULL) + * @delta: number of pages to allocate or reserve + * + * Check and update subpool page usage counts when allocating or + * reserving @delta hugepages. + * + * Context: Takes spool->lock using spin_lock_irq(). + * Return: Non-negative number of reservations that cannot be + * satisfied by the subpool, or -ENOMEM if the subpool maximum + * limit would be exceeded. + */ +long hugepage_subpool_get_pages(struct hugepage_subpool *spool, + long delta) +{ + long ret =3D delta; + + if (!spool) + return ret; + + spin_lock_irq(&spool->lock); + + if (spool->max_hpages !=3D -1 && + spool->used_hpages + delta > spool->max_hpages) { + ret =3D -ENOMEM; + goto unlock_ret; + } + + spool->used_hpages +=3D delta; + + /* minimum size accounting */ + if (spool->min_hpages !=3D -1 && spool->rsv_hpages) { + if (delta > spool->rsv_hpages) { + /* + * Asking for more reserves than those already taken on + * behalf of subpool. Return difference. + */ + ret =3D delta - spool->rsv_hpages; + spool->rsv_hpages =3D 0; + } else { + ret =3D 0; /* reserves already accounted for */ + spool->rsv_hpages -=3D delta; + } + } + +unlock_ret: + spin_unlock_irq(&spool->lock); + return ret; +} + +/** + * hugepage_subpool_put_pages - Release pages back to a subpool + * @spool: pointer to subpool structure (may be NULL) + * @delta: number of pages to free or unreserve + * + * Check and update subpool page usage counts when freeing or + * unreserving @delta hugepages. + * + * Context: Takes spool->lock using spin_lock_irqsave(). May release + * and free @spool if its usage count and references reach + * zero. + * Return: Non-negative number of reservations that the subpool cannot + * absorb. + */ +long hugepage_subpool_put_pages(struct hugepage_subpool *spool, + long delta) +{ + long ret =3D delta; + unsigned long flags; + + if (!spool) + return delta; + + spin_lock_irqsave(&spool->lock, flags); + + spool->used_hpages -=3D delta; + + /* minimum size accounting */ + if (spool->min_hpages !=3D -1 && spool->used_hpages < spool->min_hpages) { + /* + * limit is the maximum number of reservations that + * can be restored to this subpool. + */ + long limit =3D spool->min_hpages - spool->used_hpages; + + if (spool->rsv_hpages + delta <=3D limit) + ret =3D 0; + else + ret =3D spool->rsv_hpages + delta - limit; + + spool->rsv_hpages +=3D delta; + if (spool->rsv_hpages > limit) + spool->rsv_hpages =3D limit; + } + + /* + * If hugetlbfs_put_super couldn't free spool due to an outstanding + * quota reference, free it now. + */ + unlock_or_release_subpool(spool, flags); + + return ret; +} diff --git a/mm/hugetlb_subpool.h b/mm/hugetlb_subpool.h new file mode 100644 index 0000000000000..be1f1cf012c9c --- /dev/null +++ b/mm/hugetlb_subpool.h @@ -0,0 +1,17 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +#ifndef _MM_HUGETLB_SUBPOOL_H +#define _MM_HUGETLB_SUBPOOL_H + +#include +#include + +struct hstate; +struct hugepage_subpool; + +struct hugepage_subpool *hugepage_new_subpool(struct hstate *h, long max_h= pages, + long min_hpages); +void hugepage_put_subpool(struct hugepage_subpool *spool); +long hugepage_subpool_get_pages(struct hugepage_subpool *spool, long delta= ); +long hugepage_subpool_put_pages(struct hugepage_subpool *spool, long delta= ); + +#endif /* _MM_HUGETLB_SUBPOOL_H */ --=20 2.55.0.229.g6434b31f56-goog From nobody Fri Jul 24 22:17:47 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9A34245D5D3; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763680; cv=none; b=XkzESrRCXnAdDPHGhoAUi7N4HoQgQkwLQRGJOZq5vxzlL99Jl9ZF9bQxyUx0zkV/HgHigcBmlxvVEI1VP6WEMDzW5aAheKNQD5RY8ggPZ5XokAD761fAeEKo6SyqoWPVTRi9U7XvrFdchUXklduTbOju1c1VUst4UfkgTvPgF04= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763680; c=relaxed/simple; bh=5/iRF3xNqMsJsR53Yn485u1SJJR2JsiIiVvY2SQNe+s=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=HX9WUdja1h9seS7Rbl1gmmYsQ21FOwIt0Z+fSOzAwaoa860MkVf6XCQM76dlPYz1mfjWK4G64sDJrkBCvPPJvwjaqCUbECdHWDNq9I9HksjJ9ZFucEeAdb5AZoLsijAVHybV8CnbHP9noX3plAgpT43CRDx1gL0vO5GB4jcKnuU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=c+SxcELL; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="c+SxcELL" Received: by smtp.kernel.org (Postfix) with ESMTPS id 79314C2BCFB; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1784763680; bh=5/iRF3xNqMsJsR53Yn485u1SJJR2JsiIiVvY2SQNe+s=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=c+SxcELLjBmZdtQTreWKXPNovxnW3u7q1JQ10iW2wkifkQmxpZYDgskte5XeqvBNm PNrXcMFlRbxCeOEeollYUTgXaOl3GLV7G2H8Fow4XuQcOhzJunxXRuO6KG0AagMrXu U/2yO8njKtb97sVW68BEkxcMQtnKs5m60gsRABH4Tt5jmGeZq81q95djQnIbeZRAoJ iP38iUplPnQxQP+CCyoZYAOGksNwUFjs4DQEBDWTEh8YJ9W+q1VXmjCyoZOF7fGcfu Uf6Txk5Errh4Ckax4kGdHZ3MrnORlhH94SCSjIxC5/EUzMu3Tfn7SPOkyIKT+WrEhK NdtmrjAE0NuEw== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 67147C44536; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) From: Ackerley Tng via B4 Relay Date: Wed, 22 Jul 2026 16:41:18 -0700 Subject: [PATCH v4 10/16] WIP: fs: hugetlbfs: Refactor subpool getters and integrate with hugetlb_subpool API Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260722-hugetlb-alloc-failure-fixes-v4-10-88e8b81970dc@google.com> References: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> In-Reply-To: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> To: Muchun Song , Oscar Salvador , David Hildenbrand , Joshua Hahn , Shakeel Butt , Nhat Pham , Andrew Morton , Peter Xu , Wupeng Ma , fvdl@google.com, rientjes@google.com, jthoughton@google.com, Mike Kravetz , Johannes Weiner , Michal Hocko , Roman Gushchin , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jonathan Corbet , Shuah Khan , Alex Shi , Yanteng Si , Dongliang Mu , Hongxiang Lou , Miaohe Lin Cc: vannapurve@google.com, erdemaktas@google.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, linux-doc@vger.kernel.org, Ackerley Tng X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=ed25519-sha256; t=1784763678; l=5109; i=ackerleytng@google.com; s=20260225; h=from:subject:message-id; bh=Rr0t6FQpLc5GBEHSOpDNGxRyeiM/tB/HQxHFWvewyxI=; b=r8aum4myXhKIuGOvNFNRV2xnP+sRjMi82PJGF7WGGuvumuJvvLXs2oRu08WpAjRrptqeMPkev cpDbibmLIwAB3MMC8TVRLPb2o/yYixAdK/+t9iAz2EISXgaiFArJ1q+ X-Developer-Key: i=ackerleytng@google.com; a=ed25519; pk=sAZDYXdm6Iz8FHitpHeFlCMXwabodTm7p8/3/8xUxuU= X-Endpoint-Received: by B4 Relay for ackerleytng@google.com/20260225 with auth_id=649 X-Original-From: Ackerley Tng Reply-To: ackerleytng@google.com From: Ackerley Tng Refactor the direct field accesses to spool->max_hpages, spool->min_hpages, and calculate subpool properties using getters inside mm/hugetlb_subpool.c. This will allow the definition of struct hugepage_subpool to be encapsulated and private to mm/hugetlb_subpool.c Signed-off-by: Ackerley Tng --- fs/hugetlbfs/inode.c | 28 +++++++++------------------- mm/hugetlb_subpool.c | 39 ++++++++++++++++++++++++++++++++++++++- mm/hugetlb_subpool.h | 4 ++++ 3 files changed, 51 insertions(+), 20 deletions(-) diff --git a/fs/hugetlbfs/inode.c b/fs/hugetlbfs/inode.c index 86c21f8272470..b424afdedb3ee 100644 --- a/fs/hugetlbfs/inode.c +++ b/fs/hugetlbfs/inode.c @@ -1060,7 +1060,6 @@ static int hugetlbfs_show_options(struct seq_file *m,= struct dentry *root) struct hugetlbfs_sb_info *sbinfo =3D HUGETLBFS_SB(root->d_sb); struct hugepage_subpool *spool =3D sbinfo->spool; unsigned long hpage_size =3D huge_page_size(sbinfo->hstate); - unsigned hpage_shift =3D huge_page_shift(sbinfo->hstate); char mod; =20 if (!uid_eq(sbinfo->uid, GLOBAL_ROOT_UID)) @@ -1082,12 +1081,13 @@ static int hugetlbfs_show_options(struct seq_file *= m, struct dentry *root) } seq_printf(m, ",pagesize=3D%lu%c", hpage_size, mod); if (spool) { - if (spool->max_hpages !=3D -1) - seq_printf(m, ",size=3D%llu", - (unsigned long long)spool->max_hpages << hpage_shift); - if (spool->min_hpages !=3D -1) - seq_printf(m, ",min_size=3D%llu", - (unsigned long long)spool->min_hpages << hpage_shift); + unsigned long long max_size =3D hugepage_subpool_max_size(spool); + unsigned long long min_size =3D hugepage_subpool_min_size(spool); + + if (max_size !=3D -1ULL) + seq_printf(m, ",size=3D%llu", max_size); + if (min_size !=3D -1ULL) + seq_printf(m, ",min_size=3D%llu", min_size); } return 0; } @@ -1106,18 +1106,8 @@ static int hugetlbfs_statfs(struct dentry *dentry, s= truct kstatfs *buf) /* If no limits set, just report 0 or -1 for max/free/used * blocks, like simple_statfs() */ if (sbinfo->spool) { - long free_pages; - - spin_lock_irq(&sbinfo->spool->lock); - buf->f_blocks =3D sbinfo->spool->max_hpages; - if (sbinfo->spool->max_hpages =3D=3D -1) { - free_pages =3D -1; - } else { - free_pages =3D sbinfo->spool->max_hpages - - sbinfo->spool->used_hpages; - } - buf->f_bavail =3D buf->f_bfree =3D free_pages; - spin_unlock_irq(&sbinfo->spool->lock); + buf->f_blocks =3D hugepage_subpool_max_hpages(sbinfo->spool); + buf->f_bavail =3D buf->f_bfree =3D hugepage_subpool_free_hpages(sbinfo-= >spool); buf->f_files =3D sbinfo->max_inodes; buf->f_ffree =3D sbinfo->free_inodes; } diff --git a/mm/hugetlb_subpool.c b/mm/hugetlb_subpool.c index 6184860ed7374..ac0f9057b4921 100644 --- a/mm/hugetlb_subpool.c +++ b/mm/hugetlb_subpool.c @@ -12,7 +12,6 @@ #endif #include #include - #include "hugetlb_subpool.h" =20 static inline bool subpool_is_free(struct hugepage_subpool *spool) @@ -38,6 +37,44 @@ static inline void unlock_or_release_subpool(struct huge= page_subpool *spool, } } =20 +long hugepage_subpool_free_hpages(struct hugepage_subpool *spool) +{ + long free_pages; + + spin_lock_irq(&spool->lock); + if (spool->max_hpages =3D=3D -1) + free_pages =3D -1; + else + free_pages =3D spool->max_hpages - spool->used_hpages; + spin_unlock_irq(&spool->lock); + + return free_pages; +} + +static unsigned int hugepage_subpool_hpage_shift(struct hugepage_subpool *= spool) +{ + return huge_page_shift(spool->hstate); +} + +unsigned long long hugepage_subpool_max_size(struct hugepage_subpool *spoo= l) +{ + if (spool->max_hpages =3D=3D -1) + return -1ULL; + return (unsigned long long)spool->max_hpages << hugepage_subpool_hpage_sh= ift(spool); +} + +unsigned long long hugepage_subpool_min_size(struct hugepage_subpool *spoo= l) +{ + if (spool->min_hpages =3D=3D -1) + return -1ULL; + return (unsigned long long)spool->min_hpages << hugepage_subpool_hpage_sh= ift(spool); +} + +long hugepage_subpool_max_hpages(struct hugepage_subpool *spool) +{ + return spool->max_hpages; +} + struct hugepage_subpool *hugepage_new_subpool(struct hstate *h, long max_h= pages, long min_hpages) { diff --git a/mm/hugetlb_subpool.h b/mm/hugetlb_subpool.h index be1f1cf012c9c..41d22239f2c3e 100644 --- a/mm/hugetlb_subpool.h +++ b/mm/hugetlb_subpool.h @@ -13,5 +13,9 @@ struct hugepage_subpool *hugepage_new_subpool(struct hsta= te *h, long max_hpages, void hugepage_put_subpool(struct hugepage_subpool *spool); long hugepage_subpool_get_pages(struct hugepage_subpool *spool, long delta= ); long hugepage_subpool_put_pages(struct hugepage_subpool *spool, long delta= ); +long hugepage_subpool_free_hpages(struct hugepage_subpool *spool); +long hugepage_subpool_max_hpages(struct hugepage_subpool *spool); +unsigned long long hugepage_subpool_max_size(struct hugepage_subpool *spoo= l); +unsigned long long hugepage_subpool_min_size(struct hugepage_subpool *spoo= l); =20 #endif /* _MM_HUGETLB_SUBPOOL_H */ --=20 2.55.0.229.g6434b31f56-goog From nobody Fri Jul 24 22:17:47 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A633645D5D5; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763680; cv=none; b=nKBg86P4qFxIIZALyJjoynJX5ljB/puMCggr0/OaYdUoHncvVOTgpO2YscwiIO6Kj4DYbHsqPa/h39h0EN+5Fxv3QfFWEqZKQQ8IO0jHFgXQqaH+WAQ8yy/l4qIsCBIV7iqp+w0Odt5U5mN2CA16yobCGz2BJYfgv+SZyTHFB6o= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763680; c=relaxed/simple; bh=7XTSZ54VuyxsLQ8EW8ywMrv/5l0zWuujijQrt7Pcvak=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=q6DAEeF2j5qPlCe0hYFPPMoBXw//5bwl11mBFpSUeaXA3V1PCUrz3YaiQQ5/levH+qNRoaGUirEGYrMA7AhggqBBQnZngKpxdf3MpN1exKKf9zBIvRfYvuBFgrNqo6P0c30nRXDRU/A5SxrayXnYoLYaXBeGifzNgXeC86Br2DQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=ugzi/Wco; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="ugzi/Wco" Received: by smtp.kernel.org (Postfix) with ESMTPS id 87CE8C2BCFF; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1784763680; bh=7XTSZ54VuyxsLQ8EW8ywMrv/5l0zWuujijQrt7Pcvak=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=ugzi/WcoShNp/D/DCH+3feoSAi7gx1TxEeOYybBxLwuQx/FAZK1pCng4z66iCEdP+ OAN92X9SHTy1C1Q54xRLkHoGuIrkSGBO9FStW43D//pZl2TfSvf4tfu2g390cLcw50 ADbJ/NrkcLQrCM3AsoL+sdDK0A2ALaTmLObwL4GxAAyonArXR0meYdvxplyww3uAz/ mhmGWrJS0daVCuxvcsoV2p40e7iNe/r198E9EMnymkWrXbiCmQ6LT42ig8OtO0Fxnj MDmUE9nAX+K9GWCS2Jh9vbb67GDnOkFZSugvsHHYEuLF8NJ09e4/c1wNNeY3JOYKZk Ya+nLZY2sQSdQ== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 761C6C4453D; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) From: Ackerley Tng via B4 Relay Date: Wed, 22 Jul 2026 16:41:19 -0700 Subject: [PATCH v4 11/16] WIP: mm: hugetlb: Make struct hugepage_subpool private to hugetlb_subpool.c Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260722-hugetlb-alloc-failure-fixes-v4-11-88e8b81970dc@google.com> References: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> In-Reply-To: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> To: Muchun Song , Oscar Salvador , David Hildenbrand , Joshua Hahn , Shakeel Butt , Nhat Pham , Andrew Morton , Peter Xu , Wupeng Ma , fvdl@google.com, rientjes@google.com, jthoughton@google.com, Mike Kravetz , Johannes Weiner , Michal Hocko , Roman Gushchin , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jonathan Corbet , Shuah Khan , Alex Shi , Yanteng Si , Dongliang Mu , Hongxiang Lou , Miaohe Lin Cc: vannapurve@google.com, erdemaktas@google.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, linux-doc@vger.kernel.org, Ackerley Tng X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=ed25519-sha256; t=1784763678; l=2210; i=ackerleytng@google.com; s=20260225; h=from:subject:message-id; bh=lOjGCQX3/m/ifrV7JiSqPXOvpZ1jml5gEVW64pYajUU=; b=62CKcF92BwyBeZ4+JVII0ser/DgCEfN1PyHCWdGDtFmjFrAuQZH+HIAkkbwAgIhRX0bTFA9mn 4K8I6LvaxRBBops7keMTrhLFMeObPf79bBddWfrNINpTiku9RWX/K7X X-Developer-Key: i=ackerleytng@google.com; a=ed25519; pk=sAZDYXdm6Iz8FHitpHeFlCMXwabodTm7p8/3/8xUxuU= X-Endpoint-Received: by B4 Relay for ackerleytng@google.com/20260225 with auth_id=649 X-Original-From: Ackerley Tng Reply-To: ackerleytng@google.com From: Ackerley Tng Now that all filesystem and memory management subpool functions across fs/hugetlbfs/inode.c and mm/ use getters, transition struct hugepage_subpool out of the public header include/linux/hugetlb.h and encapsulate it privately inside mm/hugetlb_subpool.c. Replace the header definition with a forward declaration. Signed-off-by: Ackerley Tng --- include/linux/hugetlb.h | 13 ++----------- mm/hugetlb_subpool.c | 12 ++++++++++++ 2 files changed, 14 insertions(+), 11 deletions(-) diff --git a/include/linux/hugetlb.h b/include/linux/hugetlb.h index f36be371c6e88..074a45903973e 100644 --- a/include/linux/hugetlb.h +++ b/include/linux/hugetlb.h @@ -18,6 +18,7 @@ =20 struct mmu_gather; struct node; +struct hstate; =20 void free_huge_folio(struct folio *folio); =20 @@ -34,17 +35,7 @@ void free_huge_folio(struct folio *folio); */ #define __NR_USED_SUBPAGE 3 =20 -struct hugepage_subpool { - spinlock_t lock; - long count; - long max_hpages; /* Maximum huge pages or -1 if no maximum. */ - long used_hpages; /* Used page count, includes both */ - /* allocated and reserved pages. */ - struct hstate *hstate; - long min_hpages; /* Minimum huge pages or -1 if no minimum. */ - long rsv_hpages; /* Pages reserved against global pool to */ - /* satisfy minimum size. */ -}; +struct hugepage_subpool; =20 struct resv_map { struct kref refs; diff --git a/mm/hugetlb_subpool.c b/mm/hugetlb_subpool.c index ac0f9057b4921..39413515a2722 100644 --- a/mm/hugetlb_subpool.c +++ b/mm/hugetlb_subpool.c @@ -14,6 +14,18 @@ #include #include "hugetlb_subpool.h" =20 +struct hugepage_subpool { + spinlock_t lock; + long count; + long max_hpages; /* Maximum huge pages or -1 if no maximum. */ + long used_hpages; /* Used page count, includes both */ + /* allocated and reserved pages. */ + struct hstate *hstate; + long min_hpages; /* Minimum huge pages or -1 if no minimum. */ + long rsv_hpages; /* Pages reserved against global pool to */ + /* satisfy minimum size. */ +}; + static inline bool subpool_is_free(struct hugepage_subpool *spool) { if (spool->count) --=20 2.55.0.229.g6434b31f56-goog From nobody Fri Jul 24 22:17:47 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B764445D5DA; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763680; cv=none; b=Ut4GoW1KnM2R1uNVWwqJXoE1/vZHRCpkbWAjKm68000YVXKPFb0Pj5QR3vK/08x6B79QDAb0OMLvONp7R/zFuehGCTtVr/POgyRn4GHqtUKz1uCdKCa+fO2udz5UjJ68shaLUq1npZZp+om730BGL7xesyMuNSSAFhI5tvMFgIc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763680; c=relaxed/simple; bh=Va4m5Sx9fYuEtvTR2h5n8C+dF3X7m38XYQ8knPwStVc=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=m4xQ+4BpZZ8Kvwc56q1xtMpjOp1nw+Iz/shNowvPhe/S+0ZHgqsrCcproaxwQMfwbNMQDQqhjG9X0+hamdgFc6fQFCMywgwtyW+pMFxEHDFX0GiSAQ3XZpfbI6frl7LJr6DdhEHw7YQAVYVhpSWsoFCz0iD+kFQ39VOyCIYRT80= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=PnwuS69l; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="PnwuS69l" Received: by smtp.kernel.org (Postfix) with ESMTPS id 999D1C2BCFC; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1784763680; bh=Va4m5Sx9fYuEtvTR2h5n8C+dF3X7m38XYQ8knPwStVc=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=PnwuS69lJKLV13AjZe1OBVkwxTjEICmt5ZlCaILdYj3It+aN2xHPMMQh5kVcfczg4 J2+PgB0oP+qt+0wpTmFUkgpWuV0XqjbDSBwgmlhReo2ISu8QdmjIn0QIaG6hOmQxEk cu26STNOlSqVkhPth4WSd0QZd5PJy9wFrRpHxvj37hOO8d+/rZQBFEsb8wKL7zFkVW lEgMWXl0AXZe+Mbx8bRCIdvt/P0SDZIuzZTv23reBu+/0X+BjTiraviDgde3OPE2oj igffL2TOrayeoB0WWk9NXi5LM85Ps4LYHfQwyQb2O652QpR1X+Pv10vKEtPDL6Stya +dtfkJyTYwaiA== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 866BCC44532; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) From: Ackerley Tng via B4 Relay Date: Wed, 22 Jul 2026 16:41:20 -0700 Subject: [PATCH v4 12/16] WIP: tools: testing: Add unit tests for HugeTLB subpool functions Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260722-hugetlb-alloc-failure-fixes-v4-12-88e8b81970dc@google.com> References: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> In-Reply-To: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> To: Muchun Song , Oscar Salvador , David Hildenbrand , Joshua Hahn , Shakeel Butt , Nhat Pham , Andrew Morton , Peter Xu , Wupeng Ma , fvdl@google.com, rientjes@google.com, jthoughton@google.com, Mike Kravetz , Johannes Weiner , Michal Hocko , Roman Gushchin , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jonathan Corbet , Shuah Khan , Alex Shi , Yanteng Si , Dongliang Mu , Hongxiang Lou , Miaohe Lin Cc: vannapurve@google.com, erdemaktas@google.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, linux-doc@vger.kernel.org, Ackerley Tng X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=ed25519-sha256; t=1784763678; l=14658; i=ackerleytng@google.com; s=20260225; h=from:subject:message-id; bh=mkZsBkGR1i2iMA6usas88vw7grWu7JqdfT5S4nhfgh0=; b=ktCZhZBR1fq1rX8XzhqdfUzrgnujNlZmpVtJioipANfWI+nvudolTpM42bFhOmozkCbgwr4SA q5+14r4sr5lBkqUY9t/n+Wemw/6PlDBMIAL97tN3cJN7zFLgjjZWrqG X-Developer-Key: i=ackerleytng@google.com; a=ed25519; pk=sAZDYXdm6Iz8FHitpHeFlCMXwabodTm7p8/3/8xUxuU= X-Endpoint-Received: by B4 Relay for ackerleytng@google.com/20260225 with auth_id=649 X-Original-From: Ackerley Tng Reply-To: ackerleytng@google.com From: Ackerley Tng Introduce unit tests for HugeTLB subpool functions to exercise subpool functions. Set testing up so that tests can be run directly from userspace. Reuses the private kernel struct hugepage_subpool struct layout natively by embedding the implementation directly, avoiding structural definition drift between implementation and testing. Signed-off-by: Ackerley Tng --- tools/testing/hugetlb_subpool/.gitignore | 1 + tools/testing/hugetlb_subpool/Makefile | 18 ++ tools/testing/hugetlb_subpool/test_subpool.c | 400 +++++++++++++++++++++++= ++++ 3 files changed, 419 insertions(+) diff --git a/tools/testing/hugetlb_subpool/.gitignore b/tools/testing/huget= lb_subpool/.gitignore new file mode 100644 index 0000000000000..7348c2c72f1e1 --- /dev/null +++ b/tools/testing/hugetlb_subpool/.gitignore @@ -0,0 +1 @@ +test_subpool diff --git a/tools/testing/hugetlb_subpool/Makefile b/tools/testing/hugetlb= _subpool/Makefile new file mode 100644 index 0000000000000..1bdb7e2635614 --- /dev/null +++ b/tools/testing/hugetlb_subpool/Makefile @@ -0,0 +1,18 @@ +# SPDX-License-Identifier: GPL-2.0 +.PHONY: all clean test + +CC =3D gcc +CFLAGS =3D -Wall -O2 -I../shared -I. -I../../include -I../../arch/x86/incl= ude -pthread +KERNEL_SUBPOOL_H =3D ../../../mm/hugetlb_subpool.h +KERNEL_SUBPOOL_C =3D ../../../mm/hugetlb_subpool.c + +all: test + +test_subpool: test_subpool.c $(KERNEL_SUBPOOL_C) $(KERNEL_SUBPOOL_H) + $(CC) $(CFLAGS) test_subpool.c -o test_subpool + +test: test_subpool + ./test_subpool + +clean: + rm -f test_subpool diff --git a/tools/testing/hugetlb_subpool/test_subpool.c b/tools/testing/h= ugetlb_subpool/test_subpool.c new file mode 100644 index 0000000000000..78900274cc264 --- /dev/null +++ b/tools/testing/hugetlb_subpool/test_subpool.c @@ -0,0 +1,400 @@ +// SPDX-License-Identifier: GPL-2.0 +#include +#include +#include +#include +#include + +/* Mocked Userspace implementation for Kernel Subpool allocation dependenc= ies */ +struct hstate { + int dummy; +}; + +#undef kzalloc_obj +#undef kzalloc_objs +#define kzalloc_obj(P, ...) malloc(sizeof(P)) +#define kzalloc_objs(P, COUNT, ...) malloc(sizeof(P) * (COUNT)) + +#define kfree free +#define kmalloc malloc + +#define huge_page_shift(h) (21 + (0 * ((unsigned long)(h) & 0))) +#define huge_page_size(h) (1UL << huge_page_shift(h)) + +static bool hugetlb_acct_memory_called; +static struct hstate *hugetlb_acct_memory_h; +static long hugetlb_acct_memory_delta; + +static int hugetlb_acct_memory(struct hstate *h, long delta) +{ + hugetlb_acct_memory_called =3D true; + hugetlb_acct_memory_h =3D h; + hugetlb_acct_memory_delta =3D delta; + return 0; +} + +static void reset_hugetlb_acct_memory_mock(void) +{ + hugetlb_acct_memory_called =3D false; + hugetlb_acct_memory_h =3D NULL; + hugetlb_acct_memory_delta =3D 0; +} + +static void assert_hugetlb_acct_memory_called(struct hstate *h, long delta) +{ + assert(hugetlb_acct_memory_called); + assert(hugetlb_acct_memory_h =3D=3D h); + assert(hugetlb_acct_memory_delta =3D=3D delta); + + reset_hugetlb_acct_memory_mock(); +} + +static void assert_hugetlb_acct_memory_not_called(void) +{ + assert(!hugetlb_acct_memory_called); +} + +#include "../../../mm/hugetlb_subpool.h" +#include "../../../mm/hugetlb_subpool.c" + +static void test_subpool_new_put_no_min_limit(void) +{ + struct hstate h; + struct hugepage_subpool *spool; + + spool =3D hugepage_new_subpool(&h, 10, -1); + assert(spool !=3D NULL); + assert(spool->max_hpages =3D=3D 10); + assert(spool->min_hpages =3D=3D -1); + assert(spool->rsv_hpages =3D=3D -1); + assert(spool->count =3D=3D 1); + assert_hugetlb_acct_memory_not_called(); + + hugepage_put_subpool(spool); + assert_hugetlb_acct_memory_not_called(); +} + +static void test_subpool_new_put_with_min_limit(void) +{ + struct hstate h; + struct hugepage_subpool *spool; + + spool =3D hugepage_new_subpool(&h, 20, 5); + assert(spool !=3D NULL); + assert(spool->max_hpages =3D=3D 20); + assert(spool->min_hpages =3D=3D 5); + assert(spool->rsv_hpages =3D=3D 5); + assert(spool->count =3D=3D 1); + assert_hugetlb_acct_memory_called(&h, 5); + + hugepage_put_subpool(spool); + assert_hugetlb_acct_memory_called(&h, -5); +} + +static void test_subpool_get_pages_below_min(void) +{ + struct hstate h; + struct hugepage_subpool *spool; + long ret; + + /* Let's initialize: min_hpages =3D 10, used_hpages =3D 9, rsv_hpages =3D= 1 */ + spool =3D hugepage_new_subpool(&h, -1, 10); + assert_hugetlb_acct_memory_called(&h, 10); + + ret =3D hugepage_subpool_get_pages(spool, 9); + assert(ret =3D=3D 0); + assert(spool->used_hpages =3D=3D 9); + assert(spool->rsv_hpages =3D=3D 1); + + /* Invoke Get (Consumes the remaining 1 subpool reserve!) */ + ret =3D hugepage_subpool_get_pages(spool, 1); + assert(ret =3D=3D 0); /* Covered by subpool reserve! */ + assert(spool->used_hpages =3D=3D 10); + assert(spool->rsv_hpages =3D=3D 0); + + /* Invoke Put (Replenishes the subpool reserve!) */ + ret =3D hugepage_subpool_put_pages(spool, 1); + assert(ret =3D=3D 0); /* Kept by subpool reserve! */ + assert(spool->used_hpages =3D=3D 9); + assert(spool->rsv_hpages =3D=3D 1); + + /* Cleanup: Return used_hpages to 0 so the subpool frees symmetrically! */ + hugepage_subpool_put_pages(spool, 9); + hugepage_put_subpool(spool); + assert_hugetlb_acct_memory_called(&h, -10); +} + +static void test_subpool_get_pages_crossing_min(void) +{ + struct hstate h; + struct hugepage_subpool *spool; + long ret; + + /* Let's initialize: min_hpages =3D 10, used_hpages =3D 10, rsv_hpages = =3D 0 */ + spool =3D hugepage_new_subpool(&h, -1, 10); + assert_hugetlb_acct_memory_called(&h, 10); + + hugepage_subpool_get_pages(spool, 10); + assert(spool->used_hpages =3D=3D 10); + assert(spool->rsv_hpages =3D=3D 0); + + /* Invoke Get (Triggers a request for a Global Buddy/Surplus page!) */ + ret =3D hugepage_subpool_get_pages(spool, 1); + assert(ret =3D=3D 1); /* Requires global page! */ + assert(spool->used_hpages =3D=3D 11); + assert(spool->rsv_hpages =3D=3D 0); + + /* Invoke Put (Above minimum, so it releases the page to the Global Pool!= ) */ + ret =3D hugepage_subpool_put_pages(spool, 1); + assert(ret =3D=3D 1); /* Dropped to global pool! */ + assert(spool->used_hpages =3D=3D 10); + assert(spool->rsv_hpages =3D=3D 0); + + /* Cleanup */ + hugepage_subpool_put_pages(spool, 10); + hugepage_put_subpool(spool); + assert_hugetlb_acct_memory_called(&h, -10); +} + +static void test_subpool_get_pages_crossing_min_multi(void) +{ + struct hstate h; + struct hugepage_subpool *spool; + long ret; + + /* Scenario 1: Crossing entirely into surplus territory by a delta > 1 */ + /* Let's initialize: min_hpages =3D 10, used_hpages =3D 8, rsv_hpages =3D= 2 */ + spool =3D hugepage_new_subpool(&h, -1, 10); + assert_hugetlb_acct_memory_called(&h, 10); + + ret =3D hugepage_subpool_get_pages(spool, 8); + assert(ret =3D=3D 0); + assert(spool->used_hpages =3D=3D 8); + assert(spool->rsv_hpages =3D=3D 2); + + /* Invoke Get with delta =3D 5 (Crosses min limit of 10 up to 13) */ + ret =3D hugepage_subpool_get_pages(spool, 5); + assert(ret =3D=3D 3); /* (8 + 5) - 10 =3D 3 global pages required! */ + assert(spool->used_hpages =3D=3D 13); + assert(spool->rsv_hpages =3D=3D 0); + + /* Invoke Put with delta =3D 5 (Drops from 13 down to 8) */ + ret =3D hugepage_subpool_put_pages(spool, 5); + assert(ret =3D=3D 3); /* 3 surplus pages released to the global pool! */ + assert(spool->used_hpages =3D=3D 8); + assert(spool->rsv_hpages =3D=3D 2); /* 2 subpool reserves perfectly resto= red! */ + + /* Scenario 2: Landing exactly on the min_hpages boundary with delta > 1 = */ + ret =3D hugepage_subpool_get_pages(spool, 2); + assert(ret =3D=3D 0); /* Perfectly covered by remaining 2 subpool reserve= s! */ + assert(spool->used_hpages =3D=3D 10); + assert(spool->rsv_hpages =3D=3D 0); + + ret =3D hugepage_subpool_put_pages(spool, 2); + assert(ret =3D=3D 0); /* Swallowed perfectly to replenish the 2 subpool r= eserves! */ + assert(spool->used_hpages =3D=3D 8); + assert(spool->rsv_hpages =3D=3D 2); + + /* Cleanup */ + hugepage_subpool_put_pages(spool, 8); + hugepage_put_subpool(spool); + assert_hugetlb_acct_memory_called(&h, -10); +} + +static void test_subpool_get_pages_max_limit(void) +{ + struct hstate h; + struct hugepage_subpool *spool; + long ret; + + spool =3D hugepage_new_subpool(&h, 5, -1); + assert_hugetlb_acct_memory_not_called(); + + ret =3D hugepage_subpool_get_pages(spool, 5); + assert(ret =3D=3D 5); + assert(spool->used_hpages =3D=3D 5); + assert(spool->rsv_hpages =3D=3D -1); + + /* Invoke Get (Should trigger -ENOMEM due to max cap limit exceeded!) */ + ret =3D hugepage_subpool_get_pages(spool, 1); + assert(ret =3D=3D -ENOMEM); + assert(spool->used_hpages =3D=3D 5); /* Unchanged */ + + /* Cleanup */ + hugepage_subpool_put_pages(spool, 5); + hugepage_put_subpool(spool); + assert_hugetlb_acct_memory_not_called(); +} + +static void test_subpool_get_pages_no_limits(void) +{ + struct hstate h; + struct hugepage_subpool *spool; + long ret; + + spool =3D hugepage_new_subpool(&h, -1, -1); + assert_hugetlb_acct_memory_not_called(); + + hugepage_subpool_get_pages(spool, 5); + assert(spool->used_hpages =3D=3D 5); + + /* Invoke Get (Surplus Global Territory) */ + ret =3D hugepage_subpool_get_pages(spool, 2); + assert(ret =3D=3D 2); + assert(spool->used_hpages =3D=3D 7); + + /* Invoke Put */ + ret =3D hugepage_subpool_put_pages(spool, 2); + assert(ret =3D=3D 2); + assert(spool->used_hpages =3D=3D 5); + + /* Cleanup */ + hugepage_subpool_put_pages(spool, 5); + hugepage_put_subpool(spool); + assert_hugetlb_acct_memory_not_called(); +} + +static void test_subpool_free_hpages(void) +{ + struct hstate h; + struct hugepage_subpool *spool; + + /* Test that free_hpages with NO min_size works perfectly */ + spool =3D hugepage_new_subpool(&h, 15, -1); + hugepage_subpool_get_pages(spool, 3); + assert(hugepage_subpool_free_hpages(spool) =3D=3D 12); + hugepage_subpool_put_pages(spool, 3); + hugepage_put_subpool(spool); + + /* Test that free_hpages with a min_size configured is COMPLETELY UNAFFEC= TED by it */ + spool =3D hugepage_new_subpool(&h, 15, 5); + assert_hugetlb_acct_memory_called(&h, 5); + hugepage_subpool_get_pages(spool, 3); + assert(hugepage_subpool_free_hpages(spool) =3D=3D 12); /* Should still be= 15 - 3 =3D 12! */ + hugepage_subpool_put_pages(spool, 3); + hugepage_put_subpool(spool); + assert_hugetlb_acct_memory_called(&h, -5); + + spool =3D hugepage_new_subpool(&h, -1, -1); + hugepage_subpool_get_pages(spool, 3); + assert(hugepage_subpool_free_hpages(spool) =3D=3D -1); + hugepage_subpool_put_pages(spool, 3); + hugepage_put_subpool(spool); + + /* Test that free_hpages with a min_size configured and NO max size retur= ns -1 */ + spool =3D hugepage_new_subpool(&h, -1, 5); + assert_hugetlb_acct_memory_called(&h, 5); + hugepage_subpool_get_pages(spool, 3); + assert(hugepage_subpool_free_hpages(spool) =3D=3D -1); + hugepage_subpool_put_pages(spool, 3); + hugepage_put_subpool(spool); + assert_hugetlb_acct_memory_called(&h, -5); + + spool =3D hugepage_new_subpool(&h, 3, -1); + hugepage_subpool_get_pages(spool, 3); + assert(hugepage_subpool_free_hpages(spool) =3D=3D 0); + hugepage_subpool_put_pages(spool, 3); + hugepage_put_subpool(spool); +} + +static void test_subpool_max_hpages(void) +{ + struct hstate h; + struct hugepage_subpool *spool; + + spool =3D hugepage_new_subpool(&h, 123, -1); + assert(hugepage_subpool_max_hpages(spool) =3D=3D 123); + hugepage_put_subpool(spool); + + /* Test that max_hpages with a min_size configured is COMPLETELY UNAFFECT= ED by it */ + spool =3D hugepage_new_subpool(&h, 123, 5); + assert_hugetlb_acct_memory_called(&h, 5); + assert(hugepage_subpool_max_hpages(spool) =3D=3D 123); + hugepage_put_subpool(spool); + assert_hugetlb_acct_memory_called(&h, -5); + + spool =3D hugepage_new_subpool(&h, -1, -1); + assert(hugepage_subpool_max_hpages(spool) =3D=3D -1); + hugepage_put_subpool(spool); + + spool =3D hugepage_new_subpool(&h, -1, 5); + assert_hugetlb_acct_memory_called(&h, 5); + assert(hugepage_subpool_max_hpages(spool) =3D=3D -1); + hugepage_put_subpool(spool); + assert_hugetlb_acct_memory_called(&h, -5); + + spool =3D hugepage_new_subpool(&h, 0, -1); + assert(hugepage_subpool_max_hpages(spool) =3D=3D 0); + hugepage_put_subpool(spool); +} + +static void test_subpool_max_size(void) +{ + struct hstate h; + struct hugepage_subpool *spool; + + spool =3D hugepage_new_subpool(&h, 10, -1); + assert(hugepage_subpool_max_size(spool) =3D=3D (10ULL << 21)); + hugepage_put_subpool(spool); + + /* Test that max_size with a min_size configured is COMPLETELY UNAFFECTED= by it */ + spool =3D hugepage_new_subpool(&h, 10, 5); + assert_hugetlb_acct_memory_called(&h, 5); + assert(hugepage_subpool_max_size(spool) =3D=3D (10ULL << 21)); + hugepage_put_subpool(spool); + assert_hugetlb_acct_memory_called(&h, -5); + + spool =3D hugepage_new_subpool(&h, -1, -1); + assert(hugepage_subpool_max_size(spool) =3D=3D -1ULL); + hugepage_put_subpool(spool); + + spool =3D hugepage_new_subpool(&h, 0, -1); + assert(hugepage_subpool_max_size(spool) =3D=3D 0ULL); + hugepage_put_subpool(spool); +} + +static void test_subpool_min_size(void) +{ + struct hstate h; + struct hugepage_subpool *spool; + + spool =3D hugepage_new_subpool(&h, -1, 5); + assert_hugetlb_acct_memory_called(&h, 5); + assert(hugepage_subpool_min_size(spool) =3D=3D (5ULL << 21)); + hugepage_put_subpool(spool); + assert_hugetlb_acct_memory_called(&h, -5); + + /* Test that min_size with a max_size configured is COMPLETELY UNAFFECTED= by it */ + spool =3D hugepage_new_subpool(&h, 20, 5); + assert_hugetlb_acct_memory_called(&h, 5); + assert(hugepage_subpool_min_size(spool) =3D=3D (5ULL << 21)); + hugepage_put_subpool(spool); + assert_hugetlb_acct_memory_called(&h, -5); + + spool =3D hugepage_new_subpool(&h, -1, -1); + assert(hugepage_subpool_min_size(spool) =3D=3D -1ULL); + hugepage_put_subpool(spool); + + spool =3D hugepage_new_subpool(&h, -1, 0); + assert_hugetlb_acct_memory_called(&h, 0); + assert(hugepage_subpool_min_size(spool) =3D=3D 0ULL); + hugepage_put_subpool(spool); + assert_hugetlb_acct_memory_called(&h, 0); +} + +int main(void) +{ + test_subpool_new_put_no_min_limit(); + test_subpool_new_put_with_min_limit(); + test_subpool_get_pages_below_min(); + test_subpool_get_pages_crossing_min(); + test_subpool_get_pages_crossing_min_multi(); + test_subpool_get_pages_max_limit(); + test_subpool_get_pages_no_limits(); + test_subpool_free_hpages(); + test_subpool_max_hpages(); + test_subpool_max_size(); + test_subpool_min_size(); + + return 0; +} --=20 2.55.0.229.g6434b31f56-goog From nobody Fri Jul 24 22:17:47 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D210D45D5F1; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763680; cv=none; b=uEQuDGY8fOqeAp/pYGchKwhg6rlDojcpf111GiMvQYndQN3tyDf57mhLF4jEjtCn4G4iVO1oUqb+bzyKLs+bq3ZIZmWwgvkOoYTz+s02GOXX0UyYC15kbsERNBaHXoOpX+qyfmkaQzcZFJzJ3I0tIGxrSUOz4HhiosjIXL4MBcc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763680; c=relaxed/simple; bh=tQO4lhRXKQ6uEvJGNjKAll868IKX9d/or000tzu8DAA=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=merzugLiuRSSJEtRye4nfSkEY0jC3Xz9yU+lFCyxjh5Qjv/mWNmdEExY07JL7LVoKhJ9Z5rctA9KDzBn6JIIvWyWQz5rEEdULM15Uo28SHsv/sLEmnDrepSz10WImcLCD/WOz3lRE+IYrCKX3lYWr8AXqtOzbwdT8Vt2U2Z6EeY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=Da/RE/zo; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="Da/RE/zo" Received: by smtp.kernel.org (Postfix) with ESMTPS id ADAE6C19425; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1784763680; bh=tQO4lhRXKQ6uEvJGNjKAll868IKX9d/or000tzu8DAA=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=Da/RE/zoJl4bsjzKC3/H6wBhxJjhv27xrkhGf2e6DeYX9aq6zIh+n6ytW5Gp6GUll ec3Y3rU+Amik1YEt1lejx7UURMqrjTazilzUhYdGgWMKY9/ZJIUN0xPOctaL6JwcSe KnAZyPXxm6Nb9oGSP0ap2pMOrlBKq6jvUtnEKc8YW/omG8kzGSvdj89nppnrvSATBA jU5l+YcJFIA3ooIyyd4Fi4fQARsM5hg+hUek35mgd8ozVwjLqWesvg5WNaiwE2ItoU qOjtW/Fl4hCjl3LsF1sG0Oivg+ODaW0T/kECmSYG+7Ln163AXtbZNgzJY4z44n0y0K 724cJdjfF3u3g== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 99666C4453A; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) From: Ackerley Tng via B4 Relay Date: Wed, 22 Jul 2026 16:41:21 -0700 Subject: [PATCH v4 13/16] WIP: Reproducer for allocation failure due to cgroup v2 memory limits Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260722-hugetlb-alloc-failure-fixes-v4-13-88e8b81970dc@google.com> References: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> In-Reply-To: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> To: Muchun Song , Oscar Salvador , David Hildenbrand , Joshua Hahn , Shakeel Butt , Nhat Pham , Andrew Morton , Peter Xu , Wupeng Ma , fvdl@google.com, rientjes@google.com, jthoughton@google.com, Mike Kravetz , Johannes Weiner , Michal Hocko , Roman Gushchin , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jonathan Corbet , Shuah Khan , Alex Shi , Yanteng Si , Dongliang Mu , Hongxiang Lou , Miaohe Lin Cc: vannapurve@google.com, erdemaktas@google.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, linux-doc@vger.kernel.org, Ackerley Tng X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=ed25519-sha256; t=1784763678; l=6966; i=ackerleytng@google.com; s=20260225; h=from:subject:message-id; bh=JlTcWmu4YylYJqF+iA2BMg6mM33pft18C72PLT5bBdc=; b=CsPLTCLuBSVGQbXs8wDx4crR6XxCc1FrWqzQJTSbrryBKU8/Kqta9FxI5P4Zfib63Gi8X9Vyh BLsF9HAKYoDCxA69jvJVKW35iNPeQTIvHSYzk03E2p3reohwsAxU8Hk X-Developer-Key: i=ackerleytng@google.com; a=ed25519; pk=sAZDYXdm6Iz8FHitpHeFlCMXwabodTm7p8/3/8xUxuU= X-Endpoint-Received: by B4 Relay for ackerleytng@google.com/20260225 with auth_id=649 X-Original-From: Ackerley Tng Reply-To: ackerleytng@google.com From: Ackerley Tng (This reproducer was hacked up and not meant to be merged.) cgroup_v2_allocation_failure.c triggers HugeTLB allocation failure by explo= iting cgroup v2 memory limits. This allows testing the error paths in the kernel = when memory control charging fails, even when physical huge pages are available. The program performs the following steps to trigger the failure: 1. Enable hugetlb accounting in cgroup v2. + The program checks if memory_hugetlb_accounting is enabled in the cg= roup2 mount options. If not, it remounts /sys/fs/cgroup with this option enabled. This ensures that HugeTLB allocations are charged against t= he cgroup memory limits. 2. Create a test cgroup and set limits. + The program creates a new cgroup subdirectory named test_reproducer = under /sys/fs/cgroup. + It sets the memory.max limit of this cgroup to 1MB (which is less th= an the 2MB huge page size). 3. Fork a child process and move it to the test cgroup. + The program forks a child process. + The child process moves itself into the test_reproducer cgroup by wr= iting its PID (using 0 for current process) to cgroup.procs in the test cg= roup directory. 4. Attempt to allocate and touch a 2MB huge page. + The child process maps a 2MB anonymous huge page using mmap with MAP_PRIVATE, MAP_ANONYMOUS, and MAP_HUGETLB. + The child process writes to the mapped address, triggering a page fa= ult. 5. Triggering the kernel bugs. + The page fault handler calls alloc_hugetlb_folio to allocate the huge page. + The allocation of the physical page from buddy allocator succeeds (assuming nr_hugepages is sufficient). + The kernel then attempts to charge this allocation to the child proc= ess's cgroup by calling mem_cgroup_charge_hugetlb. + Since the child's cgroup memory limit is 1MB and the page is 2MB, the charge fails and mem_cgroup_charge_hugetlb returns -ENOMEM. + This triggers the error path in alloc_hugetlb_folio where the bugs (= folio refcount mismatch, infinite loop on ENOMEM, and reservation leaks) a= re handled. Signed-off-by: Ackerley Tng --- cgroup_v2_allocation_failure.c | 169 +++++++++++++++++++++++++++++++++++++= ++++ 1 file changed, 169 insertions(+) diff --git a/cgroup_v2_allocation_failure.c b/cgroup_v2_allocation_failure.c new file mode 100644 index 0000000000000..1a813678d8cd2 --- /dev/null +++ b/cgroup_v2_allocation_failure.c @@ -0,0 +1,169 @@ +// SPDX-License-Identifier: GPL-2.0 +#define _GNU_SOURCE +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#define CGROUP_PATH "/sys/fs/cgroup" +#define TEST_CGROUP "test_reproducer" +#define TEST_CGROUP_PATH CGROUP_PATH "/" TEST_CGROUP + +static void write_file_val(const char *path, const char *val) +{ + int fd =3D open(path, O_WRONLY); + + if (fd < 0) { + fprintf(stderr, "Failed to open %s: %s\n", path, strerror(errno)); + exit(1); + } + if (write(fd, val, strlen(val)) < 0) { + fprintf(stderr, "Failed to write %s to %s: %s\n", val, path, strerror(er= rno)); + close(fd); + exit(1); + } + close(fd); +} + +static int is_hugetlb_accounting_enabled(void) +{ + char spec[256], file[256], type[256], opts[512]; + char line[1024]; + int enabled =3D 0; + FILE *fp; + + fp =3D fopen("/proc/mounts", "r"); + if (!fp) { + perror("fopen /proc/mounts"); + return -1; + } + + while (fgets(line, sizeof(line), fp)) { + if (sscanf(line, "%255s %255s %255s %511s", spec, file, type, opts) =3D= =3D 4) { + if (strcmp(file, CGROUP_PATH) =3D=3D 0 && strcmp(type, "cgroup2") =3D= =3D 0) { + if (strstr(opts, "memory_hugetlb_accounting") !=3D NULL) + enabled =3D 1; + break; + } + } + } + fclose(fp); + return enabled; +} + +static int enable_hugetlb_accounting(void) +{ + int ret; + + printf("Attempting to remount cgroup2 with memory_hugetlb_accounting...\n= "); + ret =3D system("mount -o remount,memory_hugetlb_accounting " CGROUP_PATH); + if (ret !=3D 0) { + fprintf(stderr, "Failed to remount: system() returned %d\n", ret); + return -1; + } + return 0; +} + +int main(int argc, char **argv) +{ + struct stat st; + size_t size; + void *addr; + pid_t pid; + int enabled; + int status; + int fd; + + if (stat(CGROUP_PATH, &st) !=3D 0 || !S_ISDIR(st.st_mode)) { + fprintf(stderr, "cgroup v2 not mounted at %s\n", CGROUP_PATH); + return 1; + } + + enabled =3D is_hugetlb_accounting_enabled(); + if (enabled < 0) + return 1; + + if (!enabled) { + if (enable_hugetlb_accounting() !=3D 0) { + fprintf(stderr, "Could not enable memory_hugetlb_accounting\n"); + return 1; + } + /* Re-check */ + enabled =3D is_hugetlb_accounting_enabled(); + if (enabled <=3D 0) { + fprintf(stderr, "Failed to enable memory_hugetlb_accounting (re-check f= ailed)\n"); + return 1; + } + printf("Successfully enabled memory_hugetlb_accounting\n"); + } else { + printf("memory_hugetlb_accounting is already enabled\n"); + } + + /* Enable memory controller in subtree */ + fd =3D open(CGROUP_PATH "/cgroup.subtree_control", O_WRONLY); + if (fd >=3D 0) { + (void)write(fd, "+memory", 7); + close(fd); + } + + if (mkdir(TEST_CGROUP_PATH, 0755) !=3D 0) { + if (errno !=3D EEXIST) { + perror("mkdir test_reproducer"); + return 1; + } + } + + /* Set memory limit to 1MB (less than 2MB hugepage) */ + write_file_val(TEST_CGROUP_PATH "/memory.max", "1M"); + + pid =3D fork(); + if (pid < 0) { + perror("fork"); + return 1; + } + + if (pid =3D=3D 0) { + /* Child: Move to cgroup */ + write_file_val(TEST_CGROUP_PATH "/cgroup.procs", "0"); + + printf("Child: Attempting to allocate and touch 2MB hugepage...\n"); + /* Allocate 2MB hugepage */ + size =3D 2 * 1024 * 1024; + addr =3D mmap(NULL, size, PROT_READ | PROT_WRITE, + MAP_PRIVATE | MAP_ANONYMOUS | MAP_HUGETLB, -1, 0); + if (addr =3D=3D MAP_FAILED) { + perror("Child: mmap MAP_HUGETLB"); + exit(1); + } + + printf("Child: mmap succeeded at %p, touching it now...\n", addr); + *(char *)addr =3D 1; + + printf("Child: Successfully touched page (bug not triggered?).\n"); + munmap(addr, size); + exit(0); + } + + /* Parent */ + waitpid(pid, &status, 0); + + printf("Parent: Child exited. Cleaning up.\n"); + rmdir(TEST_CGROUP_PATH); + + if (WIFSIGNALED(status)) { + printf("Parent: Child killed by signal %d (%s)\n", + WTERMSIG(status), strsignal(WTERMSIG(status))); + if (WTERMSIG(status) =3D=3D SIGBUS) + printf("Parent: Child got SIGBUS as expected.\n"); + } else if (WIFEXITED(status)) { + printf("Parent: Child exited with status %d\n", WEXITSTATUS(status)); + } + + return 0; +} --=20 2.55.0.229.g6434b31f56-goog From nobody Fri Jul 24 22:17:47 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E17D445D5FB; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763681; cv=none; b=EYJ2fz3SfPTq4CFVF63T0Burs2pvFVyNjvENgrNJV3o+L9XWph9wEpnGf5xLyM8sgLTsXd6VlJM/eYtwmfMDMTwdNaqYUqQYcw9roZAMG3TGneN7vAJc6Ins4tLJ1OJ4E04lcRV/kqbZVrSljX2/9PsUhva3VRyEErTDYc1w6k0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763681; c=relaxed/simple; bh=IaSkp3fdmtv7NRhfe3SroO5xQIMOPLdrrRcj/2IEtlE=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=J1cMrIk41ufCYI3lT4/73H//eb+bo5twIlz2rVrqUFQI0MnlbPbthw+SkrjlD5dEt2Si9KT0VQ91FQkyrLjNYFvE3Ksdnb6lDf4YayMZorIUuE8Y4PrE7r+FCoGPcZinNXCiAUMYVVD6vyWPfXSql0dYToLJ6XkyjhDiGn3mBRI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=XlQib/ow; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="XlQib/ow" Received: by smtp.kernel.org (Postfix) with ESMTPS id BF269C2BCF7; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1784763680; bh=IaSkp3fdmtv7NRhfe3SroO5xQIMOPLdrrRcj/2IEtlE=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=XlQib/owPAVh3H7KfaNS4PyZkmlCbiHxJjSJbi29AcA2PUhaNCMdYbn3nFoRKC9xf 9z7uj6VoT2q8h7zOKugYJorEvmpU3Ogas/oag/J0kR4AO2H+LU63b7FxgJ4Gr1HRvm XgUXx2mabgUEG1SbqIrZybG/YFb+q/b8g/7wgFb08ytsJltdsj4feTnw9aEAejYHuR E9YKl0Fbd/Wm5TdI5yiu0eV1JkPbKkcLxscGACQJPCWJW1kfX8kMWHjTV0SZP4ji1m zTuPDeKmBWVqyzfVfxxz5YIWSu284D3KT0zYL+zxGSRlK87q7wKxeFgXEQCE/u3/cC j7G2h1Adnf0Mw== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id AAA00C4453F; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) From: Ackerley Tng via B4 Relay Date: Wed, 22 Jul 2026 16:41:22 -0700 Subject: [PATCH v4 14/16] WIP: Reproducer for subpool usage leak Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260722-hugetlb-alloc-failure-fixes-v4-14-88e8b81970dc@google.com> References: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> In-Reply-To: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> To: Muchun Song , Oscar Salvador , David Hildenbrand , Joshua Hahn , Shakeel Butt , Nhat Pham , Andrew Morton , Peter Xu , Wupeng Ma , fvdl@google.com, rientjes@google.com, jthoughton@google.com, Mike Kravetz , Johannes Weiner , Michal Hocko , Roman Gushchin , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jonathan Corbet , Shuah Khan , Alex Shi , Yanteng Si , Dongliang Mu , Hongxiang Lou , Miaohe Lin Cc: vannapurve@google.com, erdemaktas@google.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, linux-doc@vger.kernel.org, Ackerley Tng X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=ed25519-sha256; t=1784763678; l=4644; i=ackerleytng@google.com; s=20260225; h=from:subject:message-id; bh=E4drmq1mZkjUTuE98aQslvvt9cikJbWZghcgFkli30k=; b=HiSFLR0mil6K38uaer8wl9Fy+05WtiROHkCptxCr9umgO7pDRf//LRN/q++nLpLWmSKUyE1xs 3Zjk/paa3R0Cx5A8PWmokVRUBC6lXEwTZfCv4jw5HxO6WGQpOvLFmfx X-Developer-Key: i=ackerleytng@google.com; a=ed25519; pk=sAZDYXdm6Iz8FHitpHeFlCMXwabodTm7p8/3/8xUxuU= X-Endpoint-Received: by B4 Relay for ackerleytng@google.com/20260225 with auth_id=649 X-Original-From: Ackerley Tng Reply-To: ackerleytng@google.com From: Ackerley Tng (This reproducer was hacked up and not meant to be merged.) The kernel leaks subpool usage and the subpool structure itself if a HugeTL= Bfs mount specifying size (which sets max_hpages on the subpool) is created. subpool_leak_max_size.sh reproduces this with the following steps: 1. Create mount, specifying size=3D2M (1 page). This sets max_hpages =3D 1 = on the subpool, but does not reserve any pages. 2. Set nr_hugepages =3D 0 and nr_overcommit_hugepages =3D 0 so that physical allocations will fail. 3. Run fallocate -l 2M on a file in the mount. + This calls hugetlbfs_fallocate, which attempts to allocate a page by calling alloc_hugetlb_folio. + alloc_hugetlb_folio calls hugepage_subpool_get_pages to track the allocation against the subpool limit. This increments used_hpages to= 1. + Physical allocation fails because nr_hugepages is 0. + Before patch (Buggy): + The error path in alloc_hugetlb_folio sees gbl_chg is 1 (indicat= ing we tried to allocate a global page) and incorrectly skips calling hugepage_subpool_put_pages. + fallocate fails and returns to userspace, but the subpool used_h= pages counter remains leaked at 1. + After patch: + The error path always calls hugepage_subpool_put_pages if map_ch= g is true, restoring used_hpages to 0. 4. Unmount the filesystem. + During unmount, the kernel calls unlock_or_release_subpool to clean = up the subpool. + It checks if the subpool is free using subpool_is_free, which returns whether used_hpages is 0. + Before patch (Buggy): + Since used_hpages leaked and is 1, subpool_is_free returns false. + The kernel skips freeing the subpool structure, leaking the hugepage_subpool structure in kernel memory. + After patch: + Since used_hpages is 0, subpool_is_free returns true, and the su= bpool structure is correctly freed. Signed-off-by: Ackerley Tng --- subpool_leak_max_size.sh | 72 ++++++++++++++++++++++++++++++++++++++++++++= ++++ 1 file changed, 72 insertions(+) diff --git a/subpool_leak_max_size.sh b/subpool_leak_max_size.sh new file mode 100755 index 0000000000000..226fd2766c4d0 --- /dev/null +++ b/subpool_leak_max_size.sh @@ -0,0 +1,72 @@ +#!/bin/bash +# SPDX-License-Identifier: GPL-2.0 + +if [ "$EUID" -ne 0 ]; then + echo "Please run as root" + exit 1 +fi + +MNT_PATH=3D"/tmp/mnt_hugetlb" +FILE_PATH=3D"$MNT_PATH/test_file" + +# Save original values +orig_nr=3D$(cat /sys/kernel/mm/hugepages/hugepages-2048kB/nr_hugepages) +orig_overcommit=3D$(cat /sys/kernel/mm/hugepages/hugepages-2048kB/nr_overc= ommit_hugepages) + +cleanup() { + echo "Cleaning up..." + rm -f "$FILE_PATH" + umount "$MNT_PATH" 2>/dev/null + rmdir "$MNT_PATH" 2>/dev/null + echo "$orig_nr" > /sys/kernel/mm/hugepages/hugepages-2048kB/nr_hugepag= es + echo "$orig_overcommit" > /sys/kernel/mm/hugepages/hugepages-2048kB/nr= _overcommit_hugepages + echo "Cleanup done." +} +trap cleanup EXIT + +# 1. Mount hugetlbfs with size=3D2M (1 page) +mkdir -p "$MNT_PATH" +if ! mount -t hugetlbfs -o size=3D2M none "$MNT_PATH"; then + echo "Failed to mount hugetlbfs" + exit 1 +fi + +# 2. Set nr_hugepages to 0, overcommit to 0 +echo 0 > /sys/kernel/mm/hugepages/hugepages-2048kB/nr_hugepages +echo 0 > /sys/kernel/mm/hugepages/hugepages-2048kB/nr_overcommit_hugepages + +# Check subpool usage before running +read total free < <(stat -f -c "%b %f" "$MNT_PATH") +used_before=3D$((total - free)) +echo "Before test - Subpool total blocks: $total" +echo "Before test - Subpool free blocks: $free" +echo "Before test - Subpool used blocks: $used_before" +if [ "$used_before" -ne 0 ]; then + echo "ERROR: Subpool is not clean before test starts!" + exit 1 +fi + +# Run fallocate (expecting failure) +echo "Running fallocate (expecting failure)..." +if fallocate -l 2M "$FILE_PATH" 2>/dev/null; then + echo "ERROR: fallocate succeeded but should have failed (nr_hugepages = is 0)" + exit 1 +fi + +# Check subpool usage via statfs +# %b: Total blocks +# %f: Free blocks +read total free < <(stat -f -c "%b %f" "$MNT_PATH") +used=3D$((total - free)) + +echo "Subpool total blocks: $total" +echo "Subpool free blocks: $free" +echo "Subpool used blocks (leaked if > 0): $used" + +if [ "$used" -gt 0 ]; then + echo "RESULT: LEAK DETECTED (FAIL)" + exit 1 +else + echo "RESULT: NO LEAK (PASS)" + exit 0 +fi --=20 2.55.0.229.g6434b31f56-goog From nobody Fri Jul 24 22:17:47 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id F055C4611CE; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763681; cv=none; b=fFQph7ixRCfWknJhjVHeQwMk64c8NjzJAvKlWh1azl5hs0K79de1R07JQkYBHFq0fiucZeAbH9h6dpZRtPRLuqC97irfbSdy/7+KMmVvcnDfpsiNLZIOMN+9C2ZRhcCmdTq0UP1M9KdBQ+sVfX+VbzGyVMoXWzcKWlyvDoe7I5I= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763681; c=relaxed/simple; bh=IvYXLwNv9XVwtXfs6wx9Pgfw6OcM69QiA8/wA9jWf4E=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=YH9hSuKOu0R+b9Rj54huXAWSgEU74pa1XOiyBkwT4RZ+ULZEgf1KRjbEbioMzHj2PdwL9hcdXssrAC+dYuAax9j7lW8auNll+smGCDhkzoENHhMO/5DObU2+NY37MwGxRHGDWTCwX5iDWL4TqWFLvmMId7Tj4nG8gDuybXFSQ0k= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=n5XrAiJt; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="n5XrAiJt" Received: by smtp.kernel.org (Postfix) with ESMTPS id D09D0C2BD00; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1784763680; bh=IvYXLwNv9XVwtXfs6wx9Pgfw6OcM69QiA8/wA9jWf4E=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=n5XrAiJtdeU8ncx5O9FD8WiOOKHr22PVoJUwVt3dU/yMA1mPX6KOnGz8xs8JN0WM9 IL/Hak3kVoSwWAj4nOQX44Ey67veSST8wrTKm7NVlTCnfPXk/FtDgGQ2S9RwzDrevv tf8ilUHi2LPn3JFSip6Egi+3SfeBzGji1+zSeGJL3dGM+4MhxKLsJ+ITQAomH9zK6Q XEIx2GcO5EHiJew3LMnftTZhcw3G4uQqkj4beRoRLygHJHucRuE3VLUUY9qWyD7az9 bHy+n3S+FCVmjFWW/dlg5Tx69EvnRlvzACYlb3axPi6tdcNggnHXR3lCaM+wqnU3el pOYPjoDQgo0OA== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id BCA0DC44536; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) From: Ackerley Tng via B4 Relay Date: Wed, 22 Jul 2026 16:41:23 -0700 Subject: [PATCH v4 15/16] WIP: Reproducer for false restoration on shared HugeTLB mappings Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260722-hugetlb-alloc-failure-fixes-v4-15-88e8b81970dc@google.com> References: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> In-Reply-To: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> To: Muchun Song , Oscar Salvador , David Hildenbrand , Joshua Hahn , Shakeel Butt , Nhat Pham , Andrew Morton , Peter Xu , Wupeng Ma , fvdl@google.com, rientjes@google.com, jthoughton@google.com, Mike Kravetz , Johannes Weiner , Michal Hocko , Roman Gushchin , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jonathan Corbet , Shuah Khan , Alex Shi , Yanteng Si , Dongliang Mu , Hongxiang Lou , Miaohe Lin Cc: vannapurve@google.com, erdemaktas@google.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, linux-doc@vger.kernel.org, Ackerley Tng X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=ed25519-sha256; t=1784763678; l=7594; i=ackerleytng@google.com; s=20260225; h=from:subject:message-id; bh=nOc8hfbJ6YqAx29c86ZvsVDaWXvf89VuywJtSWZG1aU=; b=Ygf0ukiFfgWAOGov+mR9/g6iSz7wYPk1CkPXmJ7JASb36TG2/VHmKkIqNjbjhnVjLSzBvjNWm YF93xaa9IcrC4rhyvTWK3PH5804ubxkt3tuOoT9BAJXzFCKMO4dUDtv X-Developer-Key: i=ackerleytng@google.com; a=ed25519; pk=sAZDYXdm6Iz8FHitpHeFlCMXwabodTm7p8/3/8xUxuU= X-Endpoint-Received: by B4 Relay for ackerleytng@google.com/20260225 with auth_id=649 X-Original-From: Ackerley Tng Reply-To: ackerleytng@google.com From: Ackerley Tng (This reproducer was hacked up and not meant to be merged.) hugetlb_unreserve_pages() unconditionally returns reservations to the subpo= ol (via hugepage_subpool_put_pages()). This means that regardless of whether a subpool reservation was actually used, the reservation is processed by the subpool structure. To create a false restoration, the reproducer performs these steps: 1. Mount with min_size=3D2M (1 page). Global resv_hugepages becomes 1. 2. The program maps 4MB (2 pages) shared (which also grows the file to 4MB). Global resv_hugepages becomes 2 (1 from the mount, 1 new global reservation). 3. The program populates only the first page. Global resv_hugepages decrem= ents to 1 (reservation consumed by allocation). 4. The program exits (closing VMAs/fds). For shared mappings, reservations= are associated with the file inode, so they remain active. Global resv_huge= pages remains 1. 5. The script truncates the file to 2MB (truncate -s 2M). + This synchronously triggers hugetlb_unreserve_pages() to release the reservation of the truncated range (the unallocated 2nd page). + It calls hugepage_subpool_put_pages(spool, 1). + On Vanilla Kernel (Buggy): + used_hpages is 0 (not tracked). + used_hpages (0) < min_hpages (1) is TRUE. + The subpool incorrectly restores the reservation (spool->rsv_hp= ages becomes 1), even though Page 0 is still allocated and satisfies= the mount's minimum guarantee. + hugepage_subpool_put_pages() returns 0, skipping hugetlb_acct_memory(h, -1). + Result: Global resv_hugepages remains stuck at 1 (Leak). + On Fixed Kernel: + used_hpages is tracked and is initially 2. + hugepage_subpool_put_pages(1) decrements used_hpages to 1. + used_hpages (1) < min_hpages (1) is FALSE. + The subpool does not restore the reservation. + hugepage_subpool_put_pages() returns 1. + hugetlb_acct_memory(h, -1) is called. + Result: Global resv_hugepages decrements to 0 (No leak). When the filesystem is unmounted, hugetlbfs_put_super drops the subpool reference. Since the filesystem is being unmounted, the reference count dro= ps to 0, triggering unlock_or_release_subpool. Inside unlock_or_release_subpool, the kernel checks if the subpool is free = using subpool_is_free. + On the buggy kernel, subpool_is_free checks if spool->rsv_hpages is equal= to spool->min_hpages. Because of the phantom reservation, spool->rsv_hpages = was restored to 1. Since min_hpages is 1, the check (1 =3D=3D 1) returns true. + Since the subpool is considered free, the kernel releases the initial mount-time reservation by calling hugetlb_acct_memory to decrement resv_huge_pages by spool->min_hpages (which is 1). + This decrement reduces resv_huge_pages from 1 (the leaked state) to 0. As a result, the leaked reservation is cleaned up during unmount and does n= ot persist afterward. Signed-off-by: Ackerley Tng --- subpool_shared_leak.c | 43 +++++++++++++++++++++++++ subpool_shared_leak.sh | 87 ++++++++++++++++++++++++++++++++++++++++++++++= ++++ 2 files changed, 130 insertions(+) diff --git a/subpool_shared_leak.c b/subpool_shared_leak.c new file mode 100644 index 0000000000000..80b972af1c5ce --- /dev/null +++ b/subpool_shared_leak.c @@ -0,0 +1,43 @@ +// SPDX-License-Identifier: GPL-2.0 +#include +#include +#include +#include +#include +#include + +#define HPAGE_SIZE (2 * 1024 * 1024) + +int main(int argc, char **argv) +{ + const char *file_path; + void *addr; + int fd; + + if (argc < 2) { + fprintf(stderr, "Usage: %s \n", argv[0]); + return 1; + } + file_path =3D argv[1]; + + fd =3D open(file_path, O_CREAT | O_RDWR, 0666); + if (fd < 0) { + perror("open"); + return 1; + } + + addr =3D mmap(NULL, 2 * HPAGE_SIZE, PROT_READ | PROT_WRITE, MAP_SHARED, f= d, 0); + if (addr =3D=3D MAP_FAILED) { + perror("mmap"); + close(fd); + return 1; + } + + /* Allocate 1st page only. 2nd page remains unallocated (but reserved). */ + *(char *)addr =3D 1; + + munmap(addr, 2 * HPAGE_SIZE); + close(fd); + + return 0; +} diff --git a/subpool_shared_leak.sh b/subpool_shared_leak.sh new file mode 100755 index 0000000000000..619036145c523 --- /dev/null +++ b/subpool_shared_leak.sh @@ -0,0 +1,87 @@ +#!/bin/bash +# SPDX-License-Identifier: GPL-2.0 + +if [ "$EUID" -ne 0 ]; then + echo "Please run as root" + exit 1 +fi + +MNT_PATH=3D"/tmp/mnt_hugetlb_shared_leak" +FILE_PATH=3D"$MNT_PATH/test_file" + +# Save original values +orig_nr=3D$(cat /sys/kernel/mm/hugepages/hugepages-2048kB/nr_hugepages) + +cleanup() { + echo "Cleaning up..." + rm -f "$FILE_PATH" + umount "$MNT_PATH" 2>/dev/null + rmdir "$MNT_PATH" 2>/dev/null + echo "$orig_nr" > /sys/kernel/mm/hugepages/hugepages-2048kB/nr_hugepag= es + echo "Cleanup done." +} +trap cleanup EXIT + +# 1. Set nr_hugepages to 2 +echo 2 > /sys/kernel/mm/hugepages/hugepages-2048kB/nr_hugepages + +# 2. Mount hugetlbfs with min_size=3D2M (1 page) +mkdir -p "$MNT_PATH" +if ! mount -t hugetlbfs -o min_size=3D2M none "$MNT_PATH"; then + echo "Failed to mount hugetlbfs" + exit 1 +fi + +# Check resv_hugepages after mount (should be 1) +initial_resv=3D$(cat /sys/kernel/mm/hugepages/hugepages-2048kB/resv_hugepa= ges) +echo "Initial resv_hugepages (after mount): $initial_resv" +if [ "$initial_resv" -ne 1 ]; then + echo "ERROR: Initial resv_hugepages is not 1!" + exit 1 +fi + +# Verify reproducer binary exists +if [ ! -x ./subpool_shared_leak ]; then + echo "reproducer binary './subpool_shared_leak' not found or not execu= table." + echo "Please compile it first: gcc -static -o subpool_shared_leak subp= ool_shared_leak.c" + exit 1 +fi + +# 3. Run helper to map 4MB, allocate 2MB, and close. +# This creates 2 reservations, consumes 1 (by allocating Page 0). +# The unallocated Page 1 reservation remains active in the inode's resv_ma= p. +echo "Running helper..." +./subpool_shared_leak "$FILE_PATH" + +resv_after_helper=3D$(cat /sys/kernel/mm/hugepages/hugepages-2048kB/resv_h= ugepages) +echo "resv_hugepages after helper (should be 1): $resv_after_helper" +# Page 0 is allocated (no longer reserved). Page 1 is reserved. +# So resv_hugepages should be 1. +if [ "$resv_after_helper" -ne 1 ]; then + echo "ERROR: resv_hugepages is not 1 after helper run!" + exit 1 +fi + +# 4. Truncate file to 2MB (releases Page 1 reservation) +echo "Truncating file to 2MB (releasing 1 page reservation)..." +truncate -s 2M "$FILE_PATH" + +# Check resv_hugepages after truncate. +# Since Page 0 is still allocated (and in page cache), and satisfies the +# min_size=3D2M guarantee, we should have 0 reservations remaining. +# If the bug is present, the truncate path will incorrectly restore the +# reservation to the subpool and skip releasing it globally, leaving +# resv_hugepages at 1. +final_resv=3D$(cat /sys/kernel/mm/hugepages/hugepages-2048kB/resv_hugepage= s) +echo "Final resv_hugepages (after 2MB truncate): $final_resv" + +if [ "$final_resv" -eq 1 ]; then + echo "RESULT: LEAK DETECTED (FAIL)" + exit 1 +elif [ "$final_resv" -eq 0 ]; then + echo "RESULT: NO LEAK (PASS)" + exit 0 +else + echo "RESULT: UNEXPECTED STATE ($final_resv)" + exit 2 +fi --=20 2.55.0.229.g6434b31f56-goog From nobody Fri Jul 24 22:17:47 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0E5984657F4; Wed, 22 Jul 2026 23:41:21 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763681; cv=none; b=F/zmBXE7iC4YTT1TTb1fY2sZhVTrXT3w7qNFc1oRi9XizLFUDvG8c1m2807ZOhaOeDvaypo8hMEm1nG9CNJEg4qq2LnNLRkv2oXfvgj29OTiLjIZ8jrU3blZd44WYOlQVtz4HkNC6h9kDlPrELYBsNxSU5snuOrBoCJZI0jlFZ0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784763681; c=relaxed/simple; bh=JpV22Go91sm1l5ZmXhWEkSXKUuaRnYmpd8p8KT1aw60=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=S8+bUw401T+1HxCObtVOvXCvX2yE9+VbwCxTOCPNwUurhRBWl6c6qF3e4+9W3IshyN31dgeZQrULRmB/YWkK0d+1OpSxq73FH+GM95MrHMzfD8MBIs3E/LA2KbPxL+x1Vu/gs7AE6eL4O8KYJ+4pGGIpvhkzlt9Ntmk7OfRFKCQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=P5pONj5x; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="P5pONj5x" Received: by smtp.kernel.org (Postfix) with ESMTPS id E3B5CC2BCF6; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1784763680; bh=JpV22Go91sm1l5ZmXhWEkSXKUuaRnYmpd8p8KT1aw60=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=P5pONj5xGeyu106b1q17XIEyoT1RAq/C73XmAEdk4j8QmpiPk+M6PqsgRsJJLs82B N7lEBWCUKHJ93Kr5JJRGJqPCg7R8z0abyWjyHqGfj6RXRa/5nOmY4Mn0S61+xdwSxJ CYB3nrqzjZKDxEyucLbdkSB1JYysJjctP/QxdkYVMAhJ62UQZ9MFydyvb1dXbixLhJ i6Gm28HnED6PXjKrrQRYh2rbxiESMgJAkc0GGlqBPRcPG87HExLeJJj4JqUIx3l5Yb QKGf0DnCpUFx3KYFxxAj+PPoZIA4RR9MAG7av6mQ+kkp8BHHl5aand/40mdHeFSljE n1q0Fd0Jy44vQ== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id D0752C44532; Wed, 22 Jul 2026 23:41:20 +0000 (UTC) From: Ackerley Tng via B4 Relay Date: Wed, 22 Jul 2026 16:41:24 -0700 Subject: [PATCH v4 16/16] WIP: Reproducer for out_put_pages subpool reserve leakage Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260722-hugetlb-alloc-failure-fixes-v4-16-88e8b81970dc@google.com> References: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> In-Reply-To: <20260722-hugetlb-alloc-failure-fixes-v4-0-88e8b81970dc@google.com> To: Muchun Song , Oscar Salvador , David Hildenbrand , Joshua Hahn , Shakeel Butt , Nhat Pham , Andrew Morton , Peter Xu , Wupeng Ma , fvdl@google.com, rientjes@google.com, jthoughton@google.com, Mike Kravetz , Johannes Weiner , Michal Hocko , Roman Gushchin , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jonathan Corbet , Shuah Khan , Alex Shi , Yanteng Si , Dongliang Mu , Hongxiang Lou , Miaohe Lin Cc: vannapurve@google.com, erdemaktas@google.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, cgroups@vger.kernel.org, linux-doc@vger.kernel.org, Ackerley Tng X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=ed25519-sha256; t=1784763678; l=9148; i=ackerleytng@google.com; s=20260225; h=from:subject:message-id; bh=C8+RXWPxhEZ0iEc4milJz3Gk4NI1eHcFg/eirCB7vkk=; b=FRykE6OsJj6rPG7Phjv7msN60z3TEeSAhyZioTGAlFZtw0k8YLFxH2S4I/+sdh6xff0dBk5Tf ghANT0I23YjCosFVyKBGVj4L7p83snFB5riN4HkN1XQMb0K95tBSTR9 X-Developer-Key: i=ackerleytng@google.com; a=ed25519; pk=sAZDYXdm6Iz8FHitpHeFlCMXwabodTm7p8/3/8xUxuU= X-Endpoint-Received: by B4 Relay for ackerleytng@google.com/20260225 with auth_id=649 X-Original-From: Ackerley Tng Reply-To: ackerleytng@google.com From: Ackerley Tng Add a highly precise C reproducer and accompanying bash execution script to exercise, validate, and stress-test the `out_put_pages` error path rollback semantics in `hugetlb_reserve_pages()`. How it works: 1. The bash script sets the system-wide HugeTLB pool to a highly constrained baseline of exactly `nr_hugepages =3D 1` and `nr_overcommit_hugepages = =3D 0`. 2. It mounts a `hugetlbfs` instance with `-o pagesize=3D2M,min_size=3D2M,si= ze=3D4M`, which causes the kernel to immediately consume the 1 available global pa= ge as the subpool's mount-time minimum size reserve (`rsv_hugepages` become= s 1). 3. The C reproducer then attempts a shared `mmap()` for `4M` (2 pages). - `hugepage_subpool_get_pages()` requests 2 pages, sees 1 reserved, and requests 1 additional global page. - `hugetlb_acct_memory()` attempts to secure that global page but immedi= ately fails with `-ENOMEM` because the pool is exhausted. - The kernel jumps to the `out_put_pages` error path, calling `hugepage_subpool_put_pages()` to symmetrically roll back the reservat= ion. 4. The script verifies that the subpool successfully retains its 1reserved = page during the failure and returns it cleanly to the global pool upon unmoun= t, proving that no underflow, double-free, or reserve leakage occurs in the `out_put_pages` boundary path. Signed-off-by: Ackerley Tng --- hugetlb_reserve_pages_out_put_pages.c | 49 +++++++++++ hugetlb_reserve_pages_out_put_pages.sh | 153 +++++++++++++++++++++++++++++= ++++ 2 files changed, 202 insertions(+) diff --git a/hugetlb_reserve_pages_out_put_pages.c b/hugetlb_reserve_pages_= out_put_pages.c new file mode 100644 index 0000000000000..9e63fc8997d57 --- /dev/null +++ b/hugetlb_reserve_pages_out_put_pages.c @@ -0,0 +1,49 @@ +// SPDX-License-Identifier: GPL-2.0 +#define _GNU_SOURCE +#include +#include +#include +#include +#include +#include +#include + +int main(int argc, char **argv) +{ + const char *file_path; + size_t size; + int fd; + void *addr; + + if (argc < 3) { + fprintf(stderr, "Usage: %s \n", + argv[0]); + return 1; + } + + file_path =3D argv[1]; + size =3D strtoull(argv[2], NULL, 0); + + fd =3D open(file_path, O_CREAT | O_RDWR, 0666); + if (fd < 0) + err(1, "open"); + + printf("Attempting to mmap %zu bytes shared on %s...\n", size, + file_path); + addr =3D mmap(NULL, size, PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0); + if (addr =3D=3D MAP_FAILED) { + if (errno =3D=3D ENOMEM) { + printf("mmap failed with ENOMEM as expected.\n"); + close(fd); + return 0; + } + perror("mmap failed with unexpected error"); + close(fd); + return 1; + } + + printf("ERROR: mmap SUCCEEDED unexpectedly at %p\n", addr); + munmap(addr, size); + close(fd); + return 1; +} diff --git a/hugetlb_reserve_pages_out_put_pages.sh b/hugetlb_reserve_pages= _out_put_pages.sh new file mode 100755 index 0000000000000..030e1915539b4 --- /dev/null +++ b/hugetlb_reserve_pages_out_put_pages.sh @@ -0,0 +1,153 @@ +#!/bin/bash +# SPDX-License-Identifier: GPL-2.0 +set -e + +if [ "$EUID" -ne 0 ]; then + echo "Please run as root" + exit 1 +fi + +SCRIPT_DIR=3D"$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +cd "$SCRIPT_DIR" + +# Detect default hugepage size to support both 2MB and 1GB pages robustly +hpz=3D$(grep -i hugepagesize /proc/meminfo | awk '{print $2}') +kb=3D$hpz +mb=3D$((kb / 1024)) +hpage_size_bytes=3D$((kb * 1024)) + +hpage_dir=3D"hugepages-${kb}kB" +SYSFS_PATH=3D"/sys/kernel/mm/hugepages/$hpage_dir" + +MNT_PATH=3D"/tmp/mnt_hugetlb_repro" +FILE_PATH=3D"$MNT_PATH/test_file" + +# Save original values for safe restoration +orig_nr=3D$(cat "$SYSFS_PATH/nr_hugepages") +orig_overcommit=3D$(cat "$SYSFS_PATH/nr_overcommit_hugepages") + +cleanup() { + echo "Cleaning up..." + rm -f "$FILE_PATH" + umount "$MNT_PATH" 2>/dev/null + rmdir "$MNT_PATH" 2>/dev/null + echo "$orig_nr" > "$SYSFS_PATH/nr_hugepages" + echo "$orig_overcommit" > "$SYSFS_PATH/nr_overcommit_hugepages" + echo "Cleanup done." +} +trap cleanup EXIT + +# Verify reproducer binary exists +if [ ! -x ./hugetlb_reserve_pages_out_put_pages ]; then + echo "reproducer binary './hugetlb_reserve_pages_out_put_pages' not fo= und or not executable." + echo "Please compile it first: gcc -static -o hugetlb_reserve_pages_ou= t_put_pages hugetlb_reserve_pages_out_put_pages.c" + exit 1 +fi + +# 1. Set global pool such that only the mount-time reservation can succeed +echo 1 > "$SYSFS_PATH/nr_hugepages" +echo 0 > "$SYSFS_PATH/nr_overcommit_hugepages" + +initial_resv=3D$(cat "$SYSFS_PATH/resv_hugepages") +echo "Initial resv_hugepages (before mount): $initial_resv" + +# 2. Mount with min_size =3D 1 page, max size =3D 2 pages +min_size_str=3D"${mb}M" +max_size_str=3D"$((mb * 2))M" +mmap_size_bytes=3D$((hpage_size_bytes * 2)) + +mkdir -p "$MNT_PATH" +echo "Mounting hugetlbfs with pagesize=3D${mb}M, min_size=3D$min_size_str,= size=3D$max_size_str..." +if ! mount -t hugetlbfs -o "pagesize=3D${mb}M,min_size=3D$min_size_str,siz= e=3D$max_size_str" none "$MNT_PATH"; then + echo "Failed to mount hugetlbfs" + exit 1 +fi + +resv_after_mount=3D$(cat "$SYSFS_PATH/resv_hugepages") +echo "resv_hugepages after mount: $resv_after_mount" +expected_after_mount=3D$((initial_resv + 1)) +if [ "$resv_after_mount" !=3D "$expected_after_mount" ]; then + echo "ERROR: resv_hugepages is not $expected_after_mount after mount (= actual: $resv_after_mount)!" + exit 1 +fi + +# Check mount stats after mount +expected_bsize=3D$hpage_size_bytes +bsize_S=3D$(stat -f -c "%S" "$MNT_PATH") +bsize_s=3D$(stat -f -c "%s" "$MNT_PATH") +echo "Mount block size after mount: $bsize_S / $bsize_s (expected: $expect= ed_bsize)" +if [ "$bsize_S" !=3D "$expected_bsize" ] && [ "$bsize_s" !=3D "$expected_b= size" ]; then + echo "ERROR: Unexpected mount block size after mount (actual S:$bsize_= S s:$bsize_s, expected: $expected_bsize)" + exit 1 +fi + +actual_stats_mount=3D$(stat -f -c "%b %f %a" "$MNT_PATH") +expected_stats_mount=3D"2 2 2" +echo "Mount stats after mount (total free avail): $actual_stats_mount (exp= ected: $expected_stats_mount)" +if [ "$actual_stats_mount" !=3D "$expected_stats_mount" ]; then + echo "ERROR: Unexpected mount stats after mount: $actual_stats_mount (= expected: $expected_stats_mount)" + exit 1 +fi + +# 3. Run the reproducer to trigger the out_put_pages failure path +echo "Running reproducer (expecting mmap failure with ENOMEM)..." +if ./hugetlb_reserve_pages_out_put_pages "$FILE_PATH" "$mmap_size_bytes"; = then + echo "Reproducer finished successfully." + resv_after_mmap=3D$(cat "$SYSFS_PATH/resv_hugepages") + echo "resv_hugepages after failed mmap: $resv_after_mmap" + expected_after_mmap=3D$expected_after_mount + if [ "$resv_after_mmap" =3D "$expected_after_mmap" ]; then + echo "RESULT: out_put_pages EXERCISED (resv_hugepages preserved at= $expected_after_mmap as expected)" + + # Check mount stats + expected_bsize=3D$hpage_size_bytes + bsize_S=3D$(stat -f -c "%S" "$MNT_PATH") + bsize_s=3D$(stat -f -c "%s" "$MNT_PATH") + echo "Mount block size: $bsize_S / $bsize_s (expected: $expected_b= size)" + if [ "$bsize_S" !=3D "$expected_bsize" ] && [ "$bsize_s" !=3D "$ex= pected_bsize" ]; then + echo "ERROR: Unexpected mount block size (actual S:$bsize_S s:= $bsize_s, expected: $expected_bsize)" + exit 1 + fi + + actual_stats=3D$(stat -f -c "%b %f %a" "$MNT_PATH") + expected_stats=3D"2 2 2" + echo "Mount stats (total free avail): $actual_stats (expected: $ex= pected_stats)" + if [ "$actual_stats" !=3D "$expected_stats" ]; then + echo "RESULT: Unexpected mount stats after failed mmap (FAIL)" + exit 1 + else + echo "RESULT: Mount stats restored to $expected_stats as expec= ted (PASS)" + fi + else + echo "RESULT: Unexpected resv_hugepages value: $resv_after_mmap (e= xpected: $expected_after_mmap)" + exit 1 + fi +else + echo "FAIL: Reproducer returned non-zero (mmap didn't fail with ENOMEM= )" + exit 1 +fi + +# 4. Disable trap and do manual cleanup to check for final unmount underfl= ow +trap - EXIT + +echo "Unmounting..." +umount "$MNT_PATH" +rmdir "$MNT_PATH" + +final_resv=3D$(cat "$SYSFS_PATH/resv_hugepages") +echo "Final resv_hugepages (after unmount): $final_resv" + +# Restore original values +echo "Restoring original hugepage settings..." +echo "$orig_nr" > "$SYSFS_PATH/nr_hugepages" +echo "$orig_overcommit" > "$SYSFS_PATH/nr_overcommit_hugepages" + +if [ "$final_resv" =3D "$initial_resv" ]; then + echo "RESULT: State restored to $initial_resv (or cleaned up if fixed)" + echo "ALL DONE." + exit 0 +else + echo "RESULT: Underflow/Leak/Incorrect state detected! (final_resv =3D= $final_resv, expected =3D $initial_resv)" + echo "ALL DONE." + exit 1 +fi --=20 2.55.0.229.g6434b31f56-goog