From nobody Fri Sep 25 17:45:56 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AEEB545C6FF; Wed, 9 Sep 2026 21:49:38 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788990578; cv=none; b=pXpoYTK3npElItAMUqv+suwEy//4TH2ur/F60UWZBSgcW9qFdUCo3VhTtJRzWayUiPMATqngjwPoJpG0YuvzvnutnM+McvzY4wI7ZdiqKYidXnO4X2YCT4+uRJX9tkZi37MngZXEHqGixRVCeTie23ij3brpig1BVEmkMXDMkOI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788990578; c=relaxed/simple; bh=C/qLou+bdSzowJRVddbVn5Ugw92s7DEQSzLndpwgEvg=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=fKppHTpECwcuWHRy7AIk+U2sEYClP1zG1MoLrKOrjLPLiFDXCvx1kn/1F+JjhOwQSL1sviQxBJHzxpqnVRw19d+HFKxVRuwazIgMcHOMe8jTdj2zq0sn8QPg7+X9UMgraJhVmOyIl3gQebllBwsSeMxbT1iIE1mJMvc2qg3rq9o= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=AePG5bg7; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="AePG5bg7" Received: by smtp.kernel.org (Postfix) with ESMTPS id CC12DC2BCF5; Wed, 9 Sep 2026 21:49:37 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1788990577; bh=C/qLou+bdSzowJRVddbVn5Ugw92s7DEQSzLndpwgEvg=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=AePG5bg7bQ9Kcg7m1I5zN2FvPjPZocMuf5r8Xx1ZZxvn0bNqhTKFKqjE7zgEjg3/R 14ErZWITLSfQkjBHnB0eiysmMIEZai4QJEPmpjuwDI+q+ejkhteXFb0k03Zg3NJymz LVoFbC+hjC3uOXEZEm7tOkUC0ekqgi4dTx3SzL51kiV4Mqif6KllEKbYtp/nCcmeSE cww3LxVvENGdFRW4NNhaFVlaaHYweMhpAp4z19xpG0h5FxGgd8yofCk8yjbje3YE8G DXrRK8NntOwsJ8ztqCqwERco3POaAl6pJDBY4JK8XicijTEw3wbO40YOn7I76KI/vH bd4w/FanbUz2g== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id A9A69C79FBB; Wed, 9 Sep 2026 21:49:37 +0000 (UTC) From: Ackerley Tng via B4 Relay Date: Wed, 09 Sep 2026 14:49:26 -0700 Subject: [PATCH v2 1/4] mm: hugetlb: Track used_hpages when getting/putting pages from subpool Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260909-hugetlb-subpool-always-track-used-v2-1-30c5d83b572a@google.com> References: <20260909-hugetlb-subpool-always-track-used-v2-0-30c5d83b572a@google.com> In-Reply-To: <20260909-hugetlb-subpool-always-track-used-v2-0-30c5d83b572a@google.com> To: Alex Shi , Andrew Morton , David Hildenbrand , Dongliang Mu , Hongxiang Lou , Johannes Weiner , Jonathan Corbet , Joshua Hahn , "Liam R. Howlett" , Lorenzo Stoakes , Miaohe Lin , Michal Hocko , Mike Rapoport , Muchun Song , Nhat Pham , Oscar Salvador , Peter Xu , Randy Dunlap , Roman Gushchin , Shakeel Butt , Shuah Khan , Suren Baghdasaryan , Usama Arif , Vlastimil Babka , Wupeng Ma , Yanteng Si , fvdl@google.com, jthoughton@google.com, rientjes@google.com, vannapurve@google.com Cc: linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, Ackerley Tng , stable@vger.kernel.org X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=ed25519-sha256; t=1788990577; l=12004; i=ackerleytng@google.com; s=20260225; h=from:subject:message-id; bh=4OB6fqpvzzGnqc/A00KP3+bVbl3dmq+Xdyw7zsryO5I=; b=Xa5BJGhK24KwXiumolIPlGn3y/ZkVDE8cQvJZdE5pQ1YaOAopwCi1GKcKXHgILgztP9Q3Jb9+ cMYzQtAyn8zBPBM8jFJYc6zbzH7OYV3YuWZq4Xdb+G7wiAogVmMUsx0 X-Developer-Key: i=ackerleytng@google.com; a=ed25519; pk=sAZDYXdm6Iz8FHitpHeFlCMXwabodTm7p8/3/8xUxuU= X-Endpoint-Received: by B4 Relay for ackerleytng@google.com/20260225 with auth_id=649 X-Original-From: Ackerley Tng Reply-To: ackerleytng@google.com From: Ackerley Tng HugeTLB subpools currently only track used_hpages when the user configures a size limit. This is buggy since when there are existing allocations from the subpool that would have satisfied the minimum reservations, hugepage_subpool_put_pages() will still restore a reservation to the subpool. See below for an example of a false reservation. In addition, the subpool is considered free prematurely, is freed, and this ends up causing a use-after-free. The fix is to always track used_hpages within subpools, which is also beneficial in general because with that information, reservation tracking is also fully managed within hugepage_subpool_put_pages(). The subpool always knows how many pages were allocated through it. Every page allocated through the subpool increments used_hpages, regardless of whether a reservation was taken from it. Conceptually, now, every allocation involving a subpool uses a page from the subpool, which must be returned to the subpool. Every page taken from the subpool tries to use a subpool reservation. Restoring a page to the subpool reservations only if the page was taken from subpool reservations. (If used_hpages >=3D min_hpages, the page must have not have been taken from the reservations.) Always tracking used_hpages provides the subpool with information of both used and reserved counts to make the correct decision for both max_size and min_size correctly. With used_hpages always tracked, subpool_is_free() can be simplified, such that the subpool can be declared free if there are no more pages in use. Also update the + Documentation for used_hpages in the subpool struct, since it no longer matters whether the used pages count against the maximum. + Docstring for hugepage_subpool_{get,put}_pages + Documentation to use active voice, and remove some details in favor of having details documented in the docstring Also update statfs reporting. Previously, if max_hpages is negative, used_hpages is static at 0, so returning max_hpages - used_hpages returns -1 and is always correct. Now, if the subpool doesn't have a maximum requested size, indicate no limit for free pages (-1). If it does have a maximum size, report the difference between the requested size and the number of used pages. This difference is always positive, because if the mount does have a maximum size, hugepage_subpool_get_pages() ensures that the subpool usage never exceeds the maximum. Fixes: 1c5ecae3a93fa ("hugetlbfs: add minimum size accounting to subpools") Signed-off-by: Ackerley Tng Cc: stable@vger.kernel.org Reviewed-by: Joshua Hahn --- Documentation/mm/hugetlbfs_reserv.rst | 17 +---- .../translations/zh_CN/mm/hugetlbfs_reserv.rst | 11 +--- fs/hugetlbfs/inode.c | 8 ++- include/linux/hugetlb.h | 4 +- mm/hugetlb.c | 73 +++++++++++++-----= ---- 5 files changed, 55 insertions(+), 58 deletions(-) diff --git a/Documentation/mm/hugetlbfs_reserv.rst b/Documentation/mm/huget= lbfs_reserv.rst index a49115db18c76..d244583fdcbc3 100644 --- a/Documentation/mm/hugetlbfs_reserv.rst +++ b/Documentation/mm/hugetlbfs_reserv.rst @@ -314,21 +314,8 @@ huge pages. If they can not be reserved, the mount fa= ils. The routines hugepage_subpool_get/put_pages() are called when pages are obtained from or released back to a subpool. They perform all subpool accounting, and track any reservations associated with the subpool. -hugepage_subpool_get/put_pages are passed the number of huge pages by which -to adjust the subpool 'used page' count (down for get, up for put). Norma= lly, -they return the same value that was passed or an error if not enough pages -exist in the subpool. - -However, if reserves are associated with the subpool a return value less -than the passed value may be returned. This return value indicates the -number of additional global pool adjustments which must be made. For exam= ple, -suppose a subpool contains 3 reserved huge pages and someone asks for 5. -The 3 reserved pages associated with the subpool can be used to satisfy pa= rt -of the request. But, 2 pages must be obtained from the global pools. To -relay this information to the caller, the value 2 is returned. The caller -is then responsible for attempting to obtain the additional two pages from -the global pools. - +hugepage_subpool_get/put_pages() use the number of huge pages passed to ad= just +the subpool 'used page' count. =20 COW and Reservations =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D=3D diff --git a/Documentation/translations/zh_CN/mm/hugetlbfs_reserv.rst b/Doc= umentation/translations/zh_CN/mm/hugetlbfs_reserv.rst index 20947f8bd0654..ae1f1f31477fc 100644 --- a/Documentation/translations/zh_CN/mm/hugetlbfs_reserv.rst +++ b/Documentation/translations/zh_CN/mm/hugetlbfs_reserv.rst @@ -246,15 +246,8 @@ hugepage_subpool=E7=9A=84min_hpages=E5=AD=97=E6=AE=B5= =E4=B8=AD=E8=A2=AB=E8=B7=9F=E8=B8=AA=E3=80=82=E5=9C=A8=E6=8C=82=E8=BD=BD=E6= =97=B6=EF=BC=8Chugetlb_acct_me =E8=A2=AB=E8=B0=83=E7=94=A8=E4=BB=A5=E9=A2=84=E7=95=99=E6=8C=87=E5=AE=9A= =E6=95=B0=E9=87=8F=E7=9A=84=E5=B7=A8=E9=A1=B5=E3=80=82=E5=A6=82=E6=9E=9C=E5= =AE=83=E4=BB=AC=E4=B8=8D=E8=83=BD=E8=A2=AB=E9=A2=84=E7=95=99=EF=BC=8C=E6=8C= =82=E8=BD=BD=E5=B0=B1=E4=BC=9A=E5=A4=B1=E8=B4=A5=E3=80=82 =20 =E5=BD=93=E4=BB=8E=E5=AD=90=E6=B1=A0=E4=B8=AD=E8=8E=B7=E5=8F=96=E6=88=96= =E9=87=8A=E6=94=BE=E9=A1=B5=E9=9D=A2=E6=97=B6=EF=BC=8C=E4=BC=9A=E8=B0=83=E7= =94=A8hugepage_subpool_get/put_pages()=E5=87=BD=E6=95=B0=E3=80=82 -hugepage_subpool_get/put_pages=E8=A2=AB=E4=BC=A0=E9=80=92=E7=BB=99=E5=B7= =A8=E9=A1=B5=E6=95=B0=E9=87=8F=EF=BC=8C=E4=BB=A5=E6=AD=A4=E6=9D=A5=E8=B0=83= =E6=95=B4=E5=AD=90=E6=B1=A0=E7=9A=84 =E2=80=9C=E5=B7=B2=E7=94=A8=E9=A1=B5= =E9=9D=A2=E2=80=9D =E8=AE=A1=E6=95=B0 -=EF=BC=88get=E4=B8=BA=E4=B8=8B=E9=99=8D=EF=BC=8Cput=E4=B8=BA=E4=B8=8A=E5= =8D=87=EF=BC=89=E3=80=82=E9=80=9A=E5=B8=B8=E6=83=85=E5=86=B5=E4=B8=8B=EF=BC= =8C=E5=A6=82=E6=9E=9C=E5=AD=90=E6=B1=A0=E4=B8=AD=E6=B2=A1=E6=9C=89=E8=B6=B3= =E5=A4=9F=E7=9A=84=E9=A1=B5=E9=9D=A2=EF=BC=8C=E5=AE=83=E4=BB=AC=E4=BC=9A=E8= =BF=94=E5=9B=9E=E4=B8=8E=E4=BC=A0=E9=80=92=E7=9A=84=E7=9B=B8=E5=90=8C=E7=9A= =84=E5=80=BC=E6=88=96 -=E4=B8=80=E4=B8=AA=E9=94=99=E8=AF=AF=E3=80=82 - -=E7=84=B6=E8=80=8C=EF=BC=8C=E5=A6=82=E6=9E=9C=E9=A2=84=E7=95=99=E4=B8=8E= =E5=AD=90=E6=B1=A0=E7=9B=B8=E5=85=B3=E8=81=94=EF=BC=8C=E5=8F=AF=E8=83=BD=E4= =BC=9A=E8=BF=94=E5=9B=9E=E4=B8=80=E4=B8=AA=E5=B0=8F=E4=BA=8E=E4=BC=A0=E9=80= =92=E5=80=BC=E7=9A=84=E8=BF=94=E5=9B=9E=E5=80=BC=E3=80=82=E8=BF=99=E4=B8=AA= =E8=BF=94=E5=9B=9E=E5=80=BC=E8=A1=A8=E7=A4=BA=E5=BF=85=E9=A1=BB=E8=BF=9B=E8= =A1=8C=E7=9A=84=E9=A2=9D=E5=A4=96=E5=85=A8=E5=B1=80 -=E6=B1=A0=E8=B0=83=E6=95=B4=E7=9A=84=E6=95=B0=E9=87=8F=E3=80=82=E4=BE=8B= =E5=A6=82=EF=BC=8C=E5=81=87=E8=AE=BE=E4=B8=80=E4=B8=AA=E5=AD=90=E6=B1=A0=E5= =8C=85=E5=90=AB3=E4=B8=AA=E9=A2=84=E7=95=99=E7=9A=84=E5=B7=A8=E9=A1=B5=EF= =BC=8C=E6=9C=89=E4=BA=BA=E8=A6=81=E6=B1=825=E4=B8=AA=E3=80=82=E4=B8=8E=E5= =AD=90=E6=B1=A0=E7=9B=B8=E5=85=B3=E7=9A=843=E4=B8=AA=E9=A2=84=E7=95=99=E9= =A1=B5=E5=8F=AF=E4=BB=A5=E7=94=A8=E6=9D=A5 -=E6=BB=A1=E8=B6=B3=E9=83=A8=E5=88=86=E8=AF=B7=E6=B1=82=E3=80=82=E4=BD=86= =E6=98=AF=EF=BC=8C=E5=BF=85=E9=A1=BB=E4=BB=8E=E5=85=A8=E5=B1=80=E6=B1=A0=E4= =B8=AD=E8=8E=B7=E5=BE=972=E4=B8=AA=E9=A1=B5=E9=9D=A2=E3=80=82=E4=B8=BA=E4= =BA=86=E5=90=91=E8=B0=83=E7=94=A8=E8=80=85=E8=BD=AC=E8=BE=BE=E8=BF=99=E4=B8= =80=E4=BF=A1=E6=81=AF=EF=BC=8C=E5=B0=86=E8=BF=94=E5=9B=9E=E5=80=BC2=E3=80= =82=E7=84=B6=E5=90=8E=EF=BC=8C=E8=B0=83=E7=94=A8 -=E8=80=85=E8=A6=81=E8=B4=9F=E8=B4=A3=E4=BB=8E=E5=85=A8=E5=B1=80=E6=B1=A0= =E4=B8=AD=E8=8E=B7=E5=8F=96=E5=8F=A6=E5=A4=96=E4=B8=A4=E4=B8=AA=E9=A1=B5=E9= =9D=A2=E3=80=82 - +=E5=AE=83=E4=BB=AC=E8=B4=9F=E8=B4=A3=E6=89=80=E6=9C=89=E5=AD=90=E6=B1=A0= =E7=9A=84=E7=BB=9F=E8=AE=A1=E6=A0=B8=E7=AE=97=EF=BC=8C=E5=B9=B6=E8=B7=9F=E8= =B8=AA=E4=B8=8E=E5=AD=90=E6=B1=A0=E7=9B=B8=E5=85=B3=E8=81=94=E7=9A=84=E9=A2= =84=E7=95=99=E3=80=82 +hugepage_subpool_get/put_pages()=E5=87=BD=E6=95=B0=E4=BD=BF=E7=94=A8=E4=BC= =A0=E5=85=A5=E7=9A=84=E5=B7=A8=E9=A1=B5=E6=95=B0=E9=87=8F=E6=9D=A5=E8=B0=83= =E6=95=B4=E5=AD=90=E6=B1=A0=E7=9A=84=E2=80=9C=E5=B7=B2=E7=94=A8=E9=A1=B5=E9= =9D=A2=E2=80=9D=E8=AE=A1=E6=95=B0=E3=80=82 =20 COW=E5=92=8C=E9=A2=84=E7=95=99 =3D=3D=3D=3D=3D=3D=3D=3D=3D=3D diff --git a/fs/hugetlbfs/inode.c b/fs/hugetlbfs/inode.c index 7611a8470ea26..5113f743f6fc7 100644 --- a/fs/hugetlbfs/inode.c +++ b/fs/hugetlbfs/inode.c @@ -1109,8 +1109,12 @@ static int hugetlbfs_statfs(struct dentry *dentry, s= truct kstatfs *buf) =20 spin_lock_irq(&sbinfo->spool->lock); buf->f_blocks =3D sbinfo->spool->max_hpages; - free_pages =3D sbinfo->spool->max_hpages - - sbinfo->spool->used_hpages; + if (sbinfo->spool->max_hpages =3D=3D -1) { + free_pages =3D -1; + } else { + free_pages =3D sbinfo->spool->max_hpages - + sbinfo->spool->used_hpages; + } buf->f_bavail =3D buf->f_bfree =3D free_pages; spin_unlock_irq(&sbinfo->spool->lock); buf->f_files =3D sbinfo->max_inodes; diff --git a/include/linux/hugetlb.h b/include/linux/hugetlb.h index 16c4c4caa126c..4551ff3023640 100644 --- a/include/linux/hugetlb.h +++ b/include/linux/hugetlb.h @@ -39,8 +39,8 @@ struct hugepage_subpool { spinlock_t lock; long count; long max_hpages; /* Maximum huge pages or -1 if no maximum. */ - long used_hpages; /* Used count against maximum, includes */ - /* both allocated and reserved pages. */ + long used_hpages; /* Used page count, includes both */ + /* allocated and reserved pages. */ struct hstate *hstate; long min_hpages; /* Minimum huge pages or -1 if no minimum. */ long rsv_hpages; /* Pages reserved against global pool to */ diff --git a/mm/hugetlb.c b/mm/hugetlb.c index 4f6f58bf3db6c..e72e22f887478 100644 --- a/mm/hugetlb.c +++ b/mm/hugetlb.c @@ -130,12 +130,8 @@ static inline bool subpool_is_free(struct hugepage_sub= pool *spool) { if (spool->count) return false; - if (spool->max_hpages !=3D -1) - return spool->used_hpages =3D=3D 0; - if (spool->min_hpages !=3D -1) - return spool->rsv_hpages =3D=3D spool->min_hpages; =20 - return true; + return spool->used_hpages =3D=3D 0; } =20 static inline void unlock_or_release_subpool(struct hugepage_subpool *spoo= l, @@ -193,13 +189,18 @@ void hugepage_put_subpool(struct hugepage_subpool *sp= ool) unlock_or_release_subpool(spool, flags); } =20 -/* - * Subpool accounting for allocating and reserving pages. - * Return -ENOMEM if there are not enough resources to satisfy the - * request. Otherwise, return the number of pages by which the - * global pools must be adjusted (upward). The returned value may - * only be different than the passed value (delta) in the case where - * a subpool minimum size must be maintained. +/** + * hugepage_subpool_get_pages - Get pages from a subpool + * @spool: pointer to subpool structure (may be NULL) + * @delta: number of pages to allocate or reserve + * + * Check and update subpool page usage counts when allocating or + * reserving @delta hugepages. + * + * Context: Takes spool->lock using spin_lock_irq(). + * Return: Non-negative number of reservations that cannot be + * satisfied by the subpool, or -ENOMEM if the subpool maximum + * limit would be exceeded. */ static long hugepage_subpool_get_pages(struct hugepage_subpool *spool, long delta) @@ -211,15 +212,14 @@ static long hugepage_subpool_get_pages(struct hugepag= e_subpool *spool, =20 spin_lock_irq(&spool->lock); =20 - if (spool->max_hpages !=3D -1) { /* maximum size accounting */ - if ((spool->used_hpages + delta) <=3D spool->max_hpages) - spool->used_hpages +=3D delta; - else { - ret =3D -ENOMEM; - goto unlock_ret; - } + if (spool->max_hpages !=3D -1 && + spool->used_hpages + delta > spool->max_hpages) { + ret =3D -ENOMEM; + goto unlock_ret; } =20 + spool->used_hpages +=3D delta; + /* minimum size accounting */ if (spool->min_hpages !=3D -1 && spool->rsv_hpages) { if (delta > spool->rsv_hpages) { @@ -240,11 +240,19 @@ static long hugepage_subpool_get_pages(struct hugepag= e_subpool *spool, return ret; } =20 -/* - * Subpool accounting for freeing and unreserving pages. - * Return the number of global page reservations that must be dropped. - * The return value may only be different than the passed value (delta) - * in the case where a subpool minimum size must be maintained. +/** + * hugepage_subpool_put_pages - Release pages back to a subpool + * @spool: pointer to subpool structure (may be NULL) + * @delta: number of pages to free or unreserve + * + * Check and update subpool page usage counts when freeing or + * unreserving @delta hugepages. + * + * Context: Takes spool->lock using spin_lock_irqsave(). May release + * and free @spool if its usage count and references reach + * zero. + * Return: Non-negative number of reservations that the subpool cannot + * absorb. */ static long hugepage_subpool_put_pages(struct hugepage_subpool *spool, long delta) @@ -257,19 +265,24 @@ static long hugepage_subpool_put_pages(struct hugepag= e_subpool *spool, =20 spin_lock_irqsave(&spool->lock, flags); =20 - if (spool->max_hpages !=3D -1) /* maximum size accounting */ - spool->used_hpages -=3D delta; + spool->used_hpages -=3D delta; =20 /* minimum size accounting */ if (spool->min_hpages !=3D -1 && spool->used_hpages < spool->min_hpages) { - if (spool->rsv_hpages + delta <=3D spool->min_hpages) + /* + * limit is the maximum number of reservations that + * can be restored to this subpool. + */ + long limit =3D spool->min_hpages - spool->used_hpages; + + if (spool->rsv_hpages + delta <=3D limit) ret =3D 0; else - ret =3D spool->rsv_hpages + delta - spool->min_hpages; + ret =3D spool->rsv_hpages + delta - limit; =20 spool->rsv_hpages +=3D delta; - if (spool->rsv_hpages > spool->min_hpages) - spool->rsv_hpages =3D spool->min_hpages; + if (spool->rsv_hpages > limit) + spool->rsv_hpages =3D limit; } =20 /* --=20 2.55.0.1007.g17ff1f9808-goog From nobody Fri Sep 25 17:45:56 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9ED283D3CF4; Wed, 9 Sep 2026 21:49:38 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788990578; cv=none; b=Npm6p/VYjNi/6MHmWwxnIi0cP+iRowr2prcoRDZ258+35qPjwI5BChO2FTKxdazLqwzdsA3hLJuknqP5q4720wViTPdEMLRsr4KI9Q6toGQ64xw3vPSodwOYcVAoPQAj37qfmUg3FAapYwCAT+kel+1w2pwqogWsMBErADrluh8= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788990578; c=relaxed/simple; bh=95oqh24GgfcuS7voUTlH+bgV8KCd93u3lATunNdh7lc=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=doN3qIOgkmuyIof4xuj6rmySnqyKag2gZJ6/lqXgyozt0KbqT4eLasDNJU86qIjO8+DiVvgsvcziXpOa498tlZTEbZ2ZGBOi9gtsa789ATn/l43zt8Td8VhdW3nIHBRNohqUnfFojqeggNt3iBYY+gy6MyEA8Jpq08YhcisNJ/0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=WAxtSjUX; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="WAxtSjUX" Received: by smtp.kernel.org (Postfix) with ESMTPS id DF104C2BCF4; Wed, 9 Sep 2026 21:49:37 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1788990578; bh=95oqh24GgfcuS7voUTlH+bgV8KCd93u3lATunNdh7lc=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=WAxtSjUXMBGV4ctuJag3LGh+Vpl+1bPhqBR6L3h2Rn6KUP3s0/hK6iek6h0Oeu7ql FsTt3GgxPMbp2I/sjNCLovy+Umal2gqut6WseSq+kMASHJ6qVIPZI0sGg4Ca38RCph 7vWeLObZyDMDexhloBG+P1O1oAirnbs3BQaaaAhxSGQIAKUwGMtYMOclC+LOVbcKZw 5BjBVlMWIP3f8dUqqjeUdy8fybatLgsI2d5nye/xCG5td1BFOy5PWYNJk+2Wj3i4bw m+YD5O1WZaUgnsxlp7c1pFx4ixm4q7LzK3aj5bOKh+YHQX+UP5Gab3/TWAcmHGRUQj LAs6cJe+vBP0A== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id BC182C79FB7; Wed, 9 Sep 2026 21:49:37 +0000 (UTC) From: Ackerley Tng via B4 Relay Date: Wed, 09 Sep 2026 14:49:27 -0700 Subject: [PATCH v2 2/4] mm: hugetlb: Fix out_put_pages subpool reserve calculation Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260909-hugetlb-subpool-always-track-used-v2-2-30c5d83b572a@google.com> References: <20260909-hugetlb-subpool-always-track-used-v2-0-30c5d83b572a@google.com> In-Reply-To: <20260909-hugetlb-subpool-always-track-used-v2-0-30c5d83b572a@google.com> To: Alex Shi , Andrew Morton , David Hildenbrand , Dongliang Mu , Hongxiang Lou , Johannes Weiner , Jonathan Corbet , Joshua Hahn , "Liam R. Howlett" , Lorenzo Stoakes , Miaohe Lin , Michal Hocko , Mike Rapoport , Muchun Song , Nhat Pham , Oscar Salvador , Peter Xu , Randy Dunlap , Roman Gushchin , Shakeel Butt , Shuah Khan , Suren Baghdasaryan , Usama Arif , Vlastimil Babka , Wupeng Ma , Yanteng Si , fvdl@google.com, jthoughton@google.com, rientjes@google.com, vannapurve@google.com Cc: linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, Ackerley Tng , stable@vger.kernel.org X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=ed25519-sha256; t=1788990577; l=5067; i=ackerleytng@google.com; s=20260225; h=from:subject:message-id; bh=4Uw5lcnuyxNZEYs1i4bs32YYkWYGwj+suwGu5wWpSr0=; b=AG1QU4HeP/mW1//FjtrLg1NUFFPPsKl+Vw49MmQk1MolvuvY4jhtcxZpJTM6OV/lcwiexL1Lu pV7YB87lkkXDTk8QtwhmGVwJn5ojAoN+EKTpOnaMeJoOxaWgD3NrtYu X-Developer-Key: i=ackerleytng@google.com; a=ed25519; pk=sAZDYXdm6Iz8FHitpHeFlCMXwabodTm7p8/3/8xUxuU= X-Endpoint-Received: by B4 Relay for ackerleytng@google.com/20260225 with auth_id=649 X-Original-From: Ackerley Tng Reply-To: ackerleytng@google.com From: Ackerley Tng When reserving pages for a mapping fails during global accounting, the error path rolls back the adjustments made to the subpool. Currently, this rollback was performed in two separate steps: 1. Returning only the portion of reservations originally satisfied from the subpool 2. Separately adjusting the subpool used pages counter for the portion that was requested from the global pool. In (1.), because the used pages counter had not yet been decremented for the global portion, the subpool observed an inflated used pages count. If the mount was configured with both a minimum size and a maximum size, this inflated count prevented the subpool from recognizing that usage fell below the minimum size guarantee. As a result, the subpool failed to restore its reserved pages counter and instead returned that a global reservation should be dropped. The mount-time reservation is permanently destroyed, leaving global reservation counts depleted and causing an underflow when the filesystem is eventually unmounted. Additionally, if concurrent threads modified subpool usage during the reservation attempt, calculating the rollback amount using stale local variables could cause global reservation counts to diverge. Now that used pages are always tracked within the subpool, return the entire requested page count to the subpool in a single call. Global reservations are then adjusted using the difference between the reservations originally requested and those returned, fixing the issues described above. Fixes: 1d3f9bb4c8af ("mm/hugetlb: restore failed global reservations to sub= pool") Signed-off-by: Ackerley Tng Cc: stable@vger.kernel.org Reviewed-by: Joshua Hahn --- mm/hugetlb.c | 49 +++++++++++++++++++++++-------------------------- 1 file changed, 23 insertions(+), 26 deletions(-) diff --git a/mm/hugetlb.c b/mm/hugetlb.c index e72e22f887478..9eb9f3442574c 100644 --- a/mm/hugetlb.c +++ b/mm/hugetlb.c @@ -6676,12 +6676,14 @@ long hugetlb_reserve_pages(struct inode *inode, struct vm_area_struct *vma, vma_flags_t vma_flags) { - long chg =3D -1, add =3D -1, spool_resv, gbl_resv; + long chg =3D -1, add =3D -1; struct hstate *h =3D hstate_inode(inode); struct hugepage_subpool *spool =3D subpool_inode(inode); struct resv_map *resv_map; struct hugetlb_cgroup *h_cg =3D NULL; - long gbl_reserve, regions_needed =3D 0; + long regions_needed =3D 0; + long gbl_resv_get; + long gbl_resv_put; int err; =20 /* This should never happen */ @@ -6756,9 +6758,9 @@ long hugetlb_reserve_pages(struct inode *inode, * the subpool has a minimum size, there may be some global * reservations already in place (gbl_reserve). */ - gbl_reserve =3D hugepage_subpool_get_pages(spool, chg); - if (gbl_reserve < 0) { - err =3D gbl_reserve; + gbl_resv_get =3D hugepage_subpool_get_pages(spool, chg); + if (gbl_resv_get < 0) { + err =3D gbl_resv_get; goto out_uncharge_cgroup; } =20 @@ -6766,7 +6768,7 @@ long hugetlb_reserve_pages(struct inode *inode, * Check enough hugepages are available for the reservation. * Hand the pages back to the subpool if there are not */ - err =3D hugetlb_acct_memory(h, gbl_reserve); + err =3D hugetlb_acct_memory(h, gbl_resv_get); if (err < 0) goto out_put_pages; =20 @@ -6785,7 +6787,7 @@ long hugetlb_reserve_pages(struct inode *inode, add =3D region_add(resv_map, from, to, regions_needed, h, h_cg); =20 if (unlikely(add < 0)) { - hugetlb_acct_memory(h, -gbl_reserve); + hugetlb_acct_memory(h, -gbl_resv_get); err =3D add; goto out_put_pages; } else if (unlikely(chg > add)) { @@ -6821,26 +6823,21 @@ long hugetlb_reserve_pages(struct inode *inode, } return chg; =20 -out_put_pages: - spool_resv =3D chg - gbl_reserve; - if (spool_resv) { - /* put sub pool's reservation back, chg - gbl_reserve */ - gbl_resv =3D hugepage_subpool_put_pages(spool, spool_resv); - /* - * subpool's reserved pages can not be put back due to race, - * return to hstate. - */ - hugetlb_acct_memory(h, -gbl_resv); - } - /* Restore used_hpages for pages that failed global reservation */ - if (gbl_reserve && spool) { - unsigned long flags; + out_put_pages: + /* + * Return all that was requested from the subpool, let subpool + * tell us the new number of reservations that need to be + * returned to the global pool. + */ + gbl_resv_put =3D hugepage_subpool_put_pages(spool, chg); + /* + * There may be a difference between the number of + * reservations to consume and the number to restore now if + * there are multiple threads interacting with the subpool - + * restore the difference. + */ + hugetlb_acct_memory(h, gbl_resv_get - gbl_resv_put); =20 - spin_lock_irqsave(&spool->lock, flags); - if (spool->max_hpages !=3D -1) - spool->used_hpages -=3D gbl_reserve; - unlock_or_release_subpool(spool, flags); - } out_uncharge_cgroup: hugetlb_cgroup_uncharge_cgroup_rsvd(hstate_index(h), chg * pages_per_huge_page(h), h_cg); --=20 2.55.0.1007.g17ff1f9808-goog From nobody Fri Sep 25 17:45:56 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id ADE224570C2; Wed, 9 Sep 2026 21:49:38 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788990578; cv=none; b=vCxc+65hLtPRmCUeA8lYm10rRNZ6qm2PmtyICej8t4bHtQTARGbQLkviq0C8bgpibejpGwOxb1W8cYLK4zWp0MKJRZamRppkGbVj+AMfOcIewQcbQQfTdfneBIbzvyACO63BEmVTif14qWepDqPW6ajXButxMSyrqGLprxAtIKY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788990578; c=relaxed/simple; bh=IpuMmCTHCH6IIMy98L5rBZSj/YgZOfEudC6ymLrW5rQ=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=fRWYyrR1w3zEBgaZsoaxjfmSkD8oR9DqtejspP+bgIkLSsFlvls2xka8tkNRrtr7jBe23+q04HyGdAqM56AMPNuYCjWkSWSP4vBtDP89CzQKZU17AMIX6tUUe9RyC/ovI49Fxf5t1Ofz9c+PR21mYyldYY/dDmZ7b7mqxODV654= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=bNHx/EH0; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="bNHx/EH0" Received: by smtp.kernel.org (Postfix) with ESMTPS id ED555C2BCFB; Wed, 9 Sep 2026 21:49:37 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1788990578; bh=IpuMmCTHCH6IIMy98L5rBZSj/YgZOfEudC6ymLrW5rQ=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=bNHx/EH0Cb/hdSNbAr+oD0OQovE5QEj0dxlpPO3XinstfZhmjO9wICxnPHSmm/dJ+ cc7ia0kzDRkO3zzhv8nHqzIR2gGA6KAngFpMyRaqFmoDIrt+n6Me4leX6j0dsS+S9p IZa4dKmY2Hs2LSmxqzXoO9eROJc55i+jgu7z0RwF6yDTD5Tzd9xaYCmjSi8Rs5sX9d F4dWuQ9ZeqgM9WZIRiiUFuW0w1l+lls0hHWzOBzbFCNTVas2wybWzv0eBch9NsrP03 Nomi8NeA0qdb9bJmNXVc+KPR7Ue8j+kIIt3LCwYkx57cQZyzB1dQirLfqdXcJeAMyj XoFwWIsgbXEWQ== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id CEE74C88E41; Wed, 9 Sep 2026 21:49:37 +0000 (UTC) From: Ackerley Tng via B4 Relay Date: Wed, 09 Sep 2026 14:49:28 -0700 Subject: [PATCH v2 3/4] mm: hugetlb: Fix subpool usage leak on allocation failure Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260909-hugetlb-subpool-always-track-used-v2-3-30c5d83b572a@google.com> References: <20260909-hugetlb-subpool-always-track-used-v2-0-30c5d83b572a@google.com> In-Reply-To: <20260909-hugetlb-subpool-always-track-used-v2-0-30c5d83b572a@google.com> To: Alex Shi , Andrew Morton , David Hildenbrand , Dongliang Mu , Hongxiang Lou , Johannes Weiner , Jonathan Corbet , Joshua Hahn , "Liam R. Howlett" , Lorenzo Stoakes , Miaohe Lin , Michal Hocko , Mike Rapoport , Muchun Song , Nhat Pham , Oscar Salvador , Peter Xu , Randy Dunlap , Roman Gushchin , Shakeel Butt , Shuah Khan , Suren Baghdasaryan , Usama Arif , Vlastimil Babka , Wupeng Ma , Yanteng Si , fvdl@google.com, jthoughton@google.com, rientjes@google.com, vannapurve@google.com Cc: linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, Ackerley Tng , stable@vger.kernel.org X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=ed25519-sha256; t=1788990577; l=3569; i=ackerleytng@google.com; s=20260225; h=from:subject:message-id; bh=cYJ89cnWWYewpfV2xJr1JZku8Dzt8E3hqm4GB5Ws6/g=; b=saJolngd03Eu0+scYnu3x47RfGTgFmOprIS5lh7xp/uDq2/hCD/CKZYbPTtBY4n9Hn7uqo2rO V8dm6FOIsbRBfyagUeDDamv2dWK3Q7IFZscOJPx3Tg8l77vta/SfUMo X-Developer-Key: i=ackerleytng@google.com; a=ed25519; pk=sAZDYXdm6Iz8FHitpHeFlCMXwabodTm7p8/3/8xUxuU= X-Endpoint-Received: by B4 Relay for ackerleytng@google.com/20260225 with auth_id=649 X-Original-From: Ackerley Tng Reply-To: ackerleytng@google.com From: Ackerley Tng When folio allocation fails early (e.g. buddy allocation failure or cgroup charging failure) and a reservation was not used (meaning an unreserved global page was needed), the subpool page acquired during the allocation attempt must still be returned. Currently, the subpool cleanup error path only returns the page to the subpool if a reservation was used. If no reservation was used, it skips releasing the page back to the subpool, permanently leaking the subpool's used pages counter. With subpools now always tracking used pages, always release the page back to the subpool whenever a subpool page was acquired. Opportunistically rename the local variables tracking global reservations needed and global reservations returned. This clarifies the accounting: a value of zero for needed global reservations indicates an existing reservation satisfies the allocation, while a non-zero value indicates new global pages are required. Adjust global reservations using the difference between reservations needed and reservations returned to properly handle races where concurrent threads interact with the same subpool. Fixes: a833a693a490 ("mm: hugetlb: fix incorrect fallback for subpool") Signed-off-by: Ackerley Tng Cc: stable@vger.kernel.org Reviewed-by: Joshua Hahn --- mm/hugetlb.c | 23 ++++++++++------------- 1 file changed, 10 insertions(+), 13 deletions(-) diff --git a/mm/hugetlb.c b/mm/hugetlb.c index 9eb9f3442574c..652cfb55c6e6e 100644 --- a/mm/hugetlb.c +++ b/mm/hugetlb.c @@ -2957,7 +2957,7 @@ struct folio *alloc_hugetlb_folio(struct vm_area_stru= ct *vma, struct hugepage_subpool *spool =3D subpool_vma(vma); struct hstate *h =3D hstate_vma(vma); struct folio *folio; - long retval, gbl_chg, gbl_reserve; + long retval, gbl_resv_get; map_chg_state map_chg; struct mempolicy_interpreted mpoli; gfp_t gfp =3D htlb_alloc_mask(h); @@ -2996,8 +2996,8 @@ struct folio *alloc_hugetlb_folio(struct vm_area_stru= ct *vma, * Or if it can get one from the pool reservation directly. */ if (map_chg) { - gbl_chg =3D hugepage_subpool_get_pages(spool, 1); - if (gbl_chg < 0) { + gbl_resv_get =3D hugepage_subpool_get_pages(spool, 1); + if (gbl_resv_get < 0) { ret =3D -ENOSPC; goto out_end_reservation; } @@ -3006,7 +3006,7 @@ struct folio *alloc_hugetlb_folio(struct vm_area_stru= ct *vma, * If we have the vma reservation ready, no need for extra * global reservation. */ - gbl_chg =3D 0; + gbl_resv_get =3D 0; } =20 /* @@ -3017,10 +3017,10 @@ struct folio *alloc_hugetlb_folio(struct vm_area_st= ruct *vma, alloc_flags |=3D HUGETLB_ALLOC_CHARG_CGROUP_RSVD; =20 /* - * gbl_chg =3D=3D 0 indicates a reservation exists for this + * gbl_resv_get =3D=3D 0 indicates a reservation exists for this * allocation, so try to use it. */ - if (gbl_chg =3D=3D 0) + if (gbl_resv_get =3D=3D 0) alloc_flags |=3D HUGETLB_ALLOC_USE_GLOBAL_RESERVATIONS; =20 /* Takes reference on mpol. */ @@ -3074,13 +3074,10 @@ struct folio *alloc_hugetlb_folio(struct vm_area_st= ruct *vma, return folio; =20 out_subpool_put: - /* - * put page to subpool iff the quota of subpool's rsv_hpages is used - * during hugepage_subpool_get_pages. - */ - if (map_chg && !gbl_chg) { - gbl_reserve =3D hugepage_subpool_put_pages(spool, 1); - hugetlb_acct_memory(h, -gbl_reserve); + if (map_chg) { + long gbl_resv_put =3D hugepage_subpool_put_pages(spool, 1); + + hugetlb_acct_memory(h, gbl_resv_get - gbl_resv_put); } =20 out_end_reservation: --=20 2.55.0.1007.g17ff1f9808-goog From nobody Fri Sep 25 17:45:56 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AE6A045A2A8; Wed, 9 Sep 2026 21:49:38 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788990578; cv=none; b=Mea08hyYtl7zxFUtoapzUsDyqlUEoi0/LWRsrWjOVrxbD4kZhJX5P2kUm1Z6j0DS0BUP+u06P4CwTCs6syV3Ctx63mg9gXXK/QFCZwLUC7ppkG2S5+PlAnP230NFvymI4yymrM3kQBBzUp1AT94XdE0ZIAdawy2gbU3QltILPpA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788990578; c=relaxed/simple; bh=6HuWvLoHb9MNO+PA9l9yt6M3/Q2rdK1d/7u7hiJun5I=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=pvUIDB51ExnrhLYz2JUClRlKmrf1+97xGdKBSU+FgEa0daVya2x3PrI1C+lKNfz0kUnj6i3AnI9mu5YL3hqabwd6B/3phJ8EdeuOzP1A/iIuaaukBlN5+vWVragjAUFsyReqC9mAUJLn2mLXCUp2ouCVrTZS3whbfSnLRLyjDVY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=bgXRonsL; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="bgXRonsL" Received: by smtp.kernel.org (Postfix) with ESMTPS id 030B1C4AF10; Wed, 9 Sep 2026 21:49:38 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1788990578; bh=6HuWvLoHb9MNO+PA9l9yt6M3/Q2rdK1d/7u7hiJun5I=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=bgXRonsLBmo+3r3edVzKWawVGmlc46sX9f5rRe6VaUxuAApHtQq4YARMQUu59fslm OUUg7lMUmYZKZ8+LrcvRzlUtezkOB3flJDaqyepF5PJCyx25ZGUqRZ1gFMG13trCcp EmeMLwYyfq7zFNDo2TvvYIsvpO2RGRpoA4LUWHD/nBSl27rkwZ+h9WXuj3M8CZ63Gf SXT0ipvQ01PE3nGf2i/ivU5ks0elO1UNYzDXEl8051LggWIyNqLyzC/3/bqidlONGF dggYVCh131PTayhFHaxPefN3P8KMBLSObm7S+T9XirCdBDZyhl1uIwK4W7sceWrVtS snAbk43WZoEKA== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id E1F01C88E40; Wed, 9 Sep 2026 21:49:37 +0000 (UTC) From: Ackerley Tng via B4 Relay Date: Wed, 09 Sep 2026 14:49:29 -0700 Subject: [PATCH v2 4/4] mm: hugetlb: Avoid re-allocating global reservations on region add failure Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260909-hugetlb-subpool-always-track-used-v2-4-30c5d83b572a@google.com> References: <20260909-hugetlb-subpool-always-track-used-v2-0-30c5d83b572a@google.com> In-Reply-To: <20260909-hugetlb-subpool-always-track-used-v2-0-30c5d83b572a@google.com> To: Alex Shi , Andrew Morton , David Hildenbrand , Dongliang Mu , Hongxiang Lou , Johannes Weiner , Jonathan Corbet , Joshua Hahn , "Liam R. Howlett" , Lorenzo Stoakes , Miaohe Lin , Michal Hocko , Mike Rapoport , Muchun Song , Nhat Pham , Oscar Salvador , Peter Xu , Randy Dunlap , Roman Gushchin , Shakeel Butt , Shuah Khan , Suren Baghdasaryan , Usama Arif , Vlastimil Babka , Wupeng Ma , Yanteng Si , fvdl@google.com, jthoughton@google.com, rientjes@google.com, vannapurve@google.com Cc: linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, Ackerley Tng X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=ed25519-sha256; t=1788990577; l=3199; i=ackerleytng@google.com; s=20260225; h=from:subject:message-id; bh=EoLqecRMxYnQilsAlZ6TRmi6EOz6fqVlkxDRkYXZEGk=; b=pn54b1X1S+6OXfUVQLwJl5j4qFJzOlrJTC968pEjDav5b3JyIY5oXTBnAPkDoWOqMP4ghK560 lOxRYC1UM8FCni/NJJLHPTOLuS739e0OBT7fH7w+J2lRsFmvY3WNYZC X-Developer-Key: i=ackerleytng@google.com; a=ed25519; pk=sAZDYXdm6Iz8FHitpHeFlCMXwabodTm7p8/3/8xUxuU= X-Endpoint-Received: by B4 Relay for ackerleytng@google.com/20260225 with auth_id=649 X-Original-From: Ackerley Tng Reply-To: ackerleytng@google.com From: Ackerley Tng When reserving huge pages for a shared mapping, reservations are first requested from the subpool, and any remainder is accounted in global reservations. When adding the file region entries fails later in the process, the reservation attempt must be rolled back. Previously, this error path explicitly dropped the global reservations that were just acquired before jumping to the cleanup label. The cleanup label then returned the pages to the subpool. If concurrent activity in the subpool allowed the subpool to absorb more reservations upon return than it supplied initially, the cleanup label calculated a positive difference and attempted to allocate new global reservations from scratch. This premature release was completely unnecessary because all requested pages were already backed globally: partly by the mount guarantee and partly by the global reservations just acquired. Prematurely dissolving those reservations forced the cleanup path to attempt fresh buddy allocations that could fail under memory pressure. Instead, track the number of global reservations actually accounted so far. In the cleanup label, subtract the already-accounted amount from the difference between requested and returned reservations. This ensures that when global reservations were already acquired, the adjustment is purely non-positive, dropping excess reservations without ever attempting fresh allocations. Signed-off-by: Ackerley Tng --- mm/hugetlb.c | 8 +++++--- 1 file changed, 5 insertions(+), 3 deletions(-) diff --git a/mm/hugetlb.c b/mm/hugetlb.c index 652cfb55c6e6e..1151ad959ffd5 100644 --- a/mm/hugetlb.c +++ b/mm/hugetlb.c @@ -6678,6 +6678,7 @@ long hugetlb_reserve_pages(struct inode *inode, struct hugepage_subpool *spool =3D subpool_inode(inode); struct resv_map *resv_map; struct hugetlb_cgroup *h_cg =3D NULL; + long gbl_resv_accted =3D 0; long regions_needed =3D 0; long gbl_resv_get; long gbl_resv_put; @@ -6768,6 +6769,7 @@ long hugetlb_reserve_pages(struct inode *inode, err =3D hugetlb_acct_memory(h, gbl_resv_get); if (err < 0) goto out_put_pages; + gbl_resv_accted =3D gbl_resv_get; =20 /* * Account for the reservations made. Shared mappings record regions @@ -6784,7 +6786,6 @@ long hugetlb_reserve_pages(struct inode *inode, add =3D region_add(resv_map, from, to, regions_needed, h, h_cg); =20 if (unlikely(add < 0)) { - hugetlb_acct_memory(h, -gbl_resv_get); err =3D add; goto out_put_pages; } else if (unlikely(chg > add)) { @@ -6831,9 +6832,10 @@ long hugetlb_reserve_pages(struct inode *inode, * There may be a difference between the number of * reservations to consume and the number to restore now if * there are multiple threads interacting with the subpool - - * restore the difference. + * restore the difference, taking into account any global + * reservations already acquired. */ - hugetlb_acct_memory(h, gbl_resv_get - gbl_resv_put); + hugetlb_acct_memory(h, gbl_resv_get - gbl_resv_put - gbl_resv_accted); =20 out_uncharge_cgroup: hugetlb_cgroup_uncharge_cgroup_rsvd(hstate_index(h), --=20 2.55.0.1007.g17ff1f9808-goog