From nobody Tue Sep 29 07:39:19 2026 Received: from shelob.surriel.com (shelob.surriel.com [96.67.55.147]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5676338D401 for ; Tue, 11 Aug 2026 02:52:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=96.67.55.147 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786416781; cv=none; b=WbRleWwSw4KVyncjUyjdJcYXrEmR34X+QdqieYQi23WIwplkxsWNngb8gcN+VMYxo8xc12n/7Dnombwy9X1/EM43SoO5R8LzMLQa9UFHJUoLCYdqTlUv7ZToTNhh0QCQhYNTDUMV9rrBF5/rj/AQgPxP38Wb0aHUA/my1TBBW2U= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786416781; c=relaxed/simple; bh=mfJ7PcOBwsnB6aSIscE9ZAxiaKATiOnevE9RU8JcccY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=fUfMyhBAHmEekOXnIaUn2yI/niT9Vo3NEXPHCd1PxjjWVG48K2vs+zpAg80o/fhox7C7LVHOrv4hDXurmx/tKtywRhswpmhvAbbox9Oai9n4sVJQJNEiYcm2Romj27gO8uWiznrvbF+gTnBexfre1Py7vnt3MhIRCKzxrpYitEg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=surriel.com; spf=pass smtp.mailfrom=surriel.com; dkim=pass (2048-bit key) header.d=surriel.com header.i=@surriel.com header.b=JwZAbhVR; arc=none smtp.client-ip=96.67.55.147 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=surriel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=surriel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=surriel.com header.i=@surriel.com header.b="JwZAbhVR" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=surriel.com ; s=mail; h=Content-Transfer-Encoding:MIME-Version:References:In-Reply-To: Message-ID:Date:Subject:Cc:To:From:Sender:Reply-To:Content-Type:Content-ID: Content-Description:Resent-Date:Resent-From:Resent-Sender:Resent-To:Resent-Cc :Resent-Message-ID:List-Id:List-Help:List-Unsubscribe:List-Subscribe: List-Post:List-Owner:List-Archive; bh=5WGoRZehvZShX8ntlBWBCd1g1DPrfcbm3SKIWuxii1M=; b=JwZAbhVR3ay33saSROg6T4Ppwb 25v0RbqOWQ1fh9DD+Qun9DwVwfJw0AM5zZDsgi7/zunwBNgPAq56cnWR+ekB98kkO9VsZ7W7oPpzW NPiNauO41x/GDj8amqVOCqknIHkh16v4VSuXymdJMZe9RP8nPe+tJdtIl1wvwL2S+Doe3P9Xlv3Ly sO5dyRUQc+q8QQ1y5VO4Aa2jkq6KYYJf85A60ytF0GJTYrUSj9IoP4WzjFPxJN/4B1l4noBjE9z8S GIvJB9mp8pyC2N0XYEFU+Xm7i57pzrWO2eEAuy3Onqgj0XF3IDRd2tn9FnEfQ0yYFdUQV8lH0eanZ PGUdJviw==; Received: from [96.67.55.146] (helo=fangorn.surriel.com) by shelob.surriel.com with esmtpsa (TLS1.2) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.97.1) (envelope-from ) id 1wtcb6-000000007kZ-2JhS; Mon, 10 Aug 2026 22:52:12 -0400 From: Rik van Riel To: linux-kernel@vger.kernel.org Cc: kernel-team@meta.com, Rik van Riel , Andrew Morton , David Hildenbrand , Jason Gunthorpe , John Hubbard , Peter Xu , linux-mm@kvack.org Subject: [RFC PATCH v3 1/8] mm/gup: break out gup_fill_pages() helper Date: Mon, 10 Aug 2026 22:51:50 -0400 Message-ID: <20260811025157.1632867-2-riel@surriel.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260811025157.1632867-1-riel@surriel.com> References: <20260811025157.1632867-1-riel@surriel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" __get_user_pages() fills pages[] and flushes each page's caches in an open-coded loop. Move it into a gup_fill_pages() helper, which the follow_page_mask() call chain can then use to fill its own pages[] slots. No functional changes intended. Suggested-by: David Hildenbrand Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Rik van Riel Reviewed-by: Suren Baghdasaryan --- mm/gup.c | 27 ++++++++++++++++++--------- 1 file changed, 18 insertions(+), 9 deletions(-) diff --git a/mm/gup.c b/mm/gup.c index 0692119b7904..7bb40be89529 100644 --- a/mm/gup.c +++ b/mm/gup.c @@ -633,6 +633,23 @@ static struct page *no_page_table(struct vm_area_struc= t *vma, return NULL; } =20 +static void gup_fill_pages(struct vm_area_struct *vma, unsigned long addre= ss, + struct page *page, unsigned long nr, struct page **pages) +{ + unsigned long i; + + if (!pages) + return; + + for (i =3D 0; i < nr; i++) { + struct page *subpage =3D page + i; + + pages[i] =3D subpage; + flush_anon_page(vma, subpage, address + i * PAGE_SIZE); + flush_dcache_page(subpage); + } +} + #ifdef CONFIG_PGTABLE_HAS_HUGE_LEAVES /* FOLL_FORCE can write to even unwritable PUDs in COW mappings. */ static inline bool can_follow_write_pud(pud_t pud, struct page *page, @@ -1461,9 +1478,6 @@ static long __get_user_pages(struct mm_struct *mm, page_increm =3D nr_pages; =20 if (pages) { - struct page *subpage; - unsigned int j; - /* * This must be a large folio (and doesn't need to * be the whole folio; it can be part of it), do @@ -1493,12 +1507,7 @@ static long __get_user_pages(struct mm_struct *mm, } } =20 - for (j =3D 0; j < page_increm; j++) { - subpage =3D page + j; - pages[i + j] =3D subpage; - flush_anon_page(vma, subpage, start + j * PAGE_SIZE); - flush_dcache_page(subpage); - } + gup_fill_pages(vma, start, page, page_increm, pages + i); } =20 i +=3D page_increm; --=20 2.55.0 From nobody Tue Sep 29 07:39:19 2026 Received: from shelob.surriel.com (shelob.surriel.com [96.67.55.147]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8C4E11FBEA8 for ; Tue, 11 Aug 2026 02:52:29 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=96.67.55.147 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786416751; cv=none; b=bByPfWOAVms4N8zt4D3NEOYMwmee89Qh9MeURSd8Zj3N2Hcv48raXA7KuvpWDSBh8e3spbiR/DeUyXF8g5CDamE6M2fk5w2rKZQKo+WL0XVDpLjBkPdBHM7PWPMVGXbVrOAOKiRP8PpMvFa6dS7JiUuqT6nYfVa/997eQYIh3j4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786416751; c=relaxed/simple; bh=tz0EF4d/nGhRh7ziotw8dyLLNmzn6blmiFxgb+QXu4I=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=jiTmDBdJM6EMp+6OLgvTIvTnBvocf8rF5xzQwfbhOyGDBmYN9XB9A+NP1oV3ValBw3dBbNJXfo6pw+8R5RkPxNcUAyfYKqsrrXTHjf87pxG1M6uV9W2DBjgAonUkRIGRUF8WXoTuBz3/xuP3ztIxAIrTAh17fSCfxLmHG3FtmrQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=surriel.com; spf=pass smtp.mailfrom=surriel.com; dkim=pass (2048-bit key) header.d=surriel.com header.i=@surriel.com header.b=OC69/QQK; arc=none smtp.client-ip=96.67.55.147 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=surriel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=surriel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=surriel.com header.i=@surriel.com header.b="OC69/QQK" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=surriel.com ; s=mail; h=Content-Transfer-Encoding:MIME-Version:References:In-Reply-To: Message-ID:Date:Subject:Cc:To:From:Sender:Reply-To:Content-Type:Content-ID: Content-Description:Resent-Date:Resent-From:Resent-Sender:Resent-To:Resent-Cc :Resent-Message-ID:List-Id:List-Help:List-Unsubscribe:List-Subscribe: List-Post:List-Owner:List-Archive; bh=dq+Hb/y548oYe6VtyyengDHPC9XNOyWn3sciWRU0e2A=; b=OC69/QQKWLq0kkbCLSHVA8exPU 9CdibJwt+gubMWC35uedcgdz31JY6f8w8MWr7qjMBrqXiiHvOCXGIU0T3xTWiWdZWbi4FwYkewPqu 1NMhwoJURgBJr3KDdL27NmE7K9zsYmYKge6xo9fiqZVFQ3qlrouH25pEcNpgwl+3xpYbEPzFSJqqP VAL1QbKpewwc30SVninB082VSQaMm0yC2EeASRsrqR4tyO/ZBTjCgCe6i6/bA3Zn8X5TvHhXhXTSN 5DjaLlPv7dlockBbfpZ/p40U12XSaFXe5LrmdgDqe+L51TcMfVNt1x1Gj+DUB+/O//r0BJze/acFm 6XLsuXqw==; Received: from [96.67.55.146] (helo=fangorn.surriel.com) by shelob.surriel.com with esmtpsa (TLS1.2) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.97.1) (envelope-from ) id 1wtcb6-000000007kZ-2Q5S; Mon, 10 Aug 2026 22:52:12 -0400 From: Rik van Riel To: linux-kernel@vger.kernel.org Cc: kernel-team@meta.com, Rik van Riel , Andrew Morton , David Hildenbrand , Jason Gunthorpe , John Hubbard , Peter Xu , linux-mm@kvack.org Subject: [RFC PATCH v3 2/8] mm/gup: convert follow_page_mask() to return a long Date: Mon, 10 Aug 2026 22:51:51 -0400 Message-ID: <20260811025157.1632867-3-riel@surriel.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260811025157.1632867-1-riel@surriel.com> References: <20260811025157.1632867-1-riel@surriel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" follow_page_mask() and its helpers return a struct page pointer: NULL, ERR_PTR(), or the page found. Change the return type to long instead: 0, a negative errno, or 1 with the page stored in a new pages[0] slot. This lets the return value carry a page count rather than a single struct page pointer. follow_huge_pud(), follow_huge_pmd() and follow_page_pte() now fill their slot with gup_fill_pages(); __get_user_pages() reads pages[i] back to still expand a large folio's remaining subpages itself. The vsyscall gate area, which bypasses follow_page_mask(), fills its own slot the same way. *page_mask and __get_user_pages()'s handling of a large folio's remaining subpages are untouched, and mm/gup_test.c (PIN_LONGTERM_BENCHMARK) shows no measurable difference for 4 kB, 64 kB mTHP, or 2 MB THP. No functional changes intended. Suggested-by: David Hildenbrand Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Rik van Riel Reviewed-by: Suren Baghdasaryan --- mm/gup.c | 254 +++++++++++++++++++++++++++++-------------------------- 1 file changed, 134 insertions(+), 120 deletions(-) diff --git a/mm/gup.c b/mm/gup.c index 7bb40be89529..e4e6d0993424 100644 --- a/mm/gup.c +++ b/mm/gup.c @@ -608,15 +608,15 @@ static inline bool can_follow_write_common(struct pag= e *page, return page && PageAnon(page) && PageAnonExclusive(page); } =20 -static struct page *no_page_table(struct vm_area_struct *vma, - unsigned int flags, unsigned long address) +static long no_page_table(struct vm_area_struct *vma, + unsigned int flags, unsigned long address) { if (!(flags & FOLL_DUMP)) - return NULL; + return 0; =20 /* * When core dumping, we don't want to allocate unnecessary pages or - * page tables. Return error instead of NULL to skip handle_mm_fault, + * page tables. Return error instead of 0 to skip handle_mm_fault, * then get_dump_page() will return NULL to leave a hole in the dump. * But we can only make this optimization where a hole would surely * be zero-filled if handle_mm_fault() actually did handle it. @@ -625,12 +625,12 @@ static struct page *no_page_table(struct vm_area_stru= ct *vma, struct hstate *h =3D hstate_vma(vma); =20 if (!hugetlbfs_pagecache_present(h, vma, address)) - return ERR_PTR(-EFAULT); + return -EFAULT; } else if ((vma_is_anonymous(vma) || !vma->vm_ops->fault)) { - return ERR_PTR(-EFAULT); + return -EFAULT; } =20 - return NULL; + return 0; } =20 static void gup_fill_pages(struct vm_area_struct *vma, unsigned long addre= ss, @@ -663,9 +663,10 @@ static inline bool can_follow_write_pud(pud_t pud, str= uct page *page, return can_follow_write_common(page, vma, flags); } =20 -static struct page *follow_huge_pud(struct vm_area_struct *vma, - unsigned long addr, pud_t *pudp, - int flags, unsigned long *page_mask) +static long follow_huge_pud(struct vm_area_struct *vma, + unsigned long addr, pud_t *pudp, + unsigned int flags, unsigned long *page_mask, + struct page **pages) { struct mm_struct *mm =3D vma->vm_mm; struct page *page; @@ -676,25 +677,27 @@ static struct page *follow_huge_pud(struct vm_area_st= ruct *vma, assert_spin_locked(pud_lockptr(mm, pudp)); =20 if (!pud_present(pud)) - return NULL; + return 0; =20 if ((flags & FOLL_WRITE) && !can_follow_write_pud(pud, pfn_to_page(pfn), vma, flags)) - return NULL; + return 0; =20 pfn +=3D (addr & ~PUD_MASK) >> PAGE_SHIFT; page =3D pfn_to_page(pfn); =20 if (!pud_write(pud) && gup_must_unshare(vma, flags, page)) - return ERR_PTR(-EMLINK); + return -EMLINK; =20 ret =3D try_grab_folio(page_folio(page), 1, flags); if (ret) - page =3D ERR_PTR(ret); - else - *page_mask =3D HPAGE_PUD_NR - 1; + return ret; =20 - return page; + *page_mask =3D HPAGE_PUD_NR - 1; + + gup_fill_pages(vma, addr, page, 1, pages); + + return 1; } =20 /* FOLL_FORCE can write to even unwritable PMDs in COW mappings. */ @@ -715,10 +718,10 @@ static inline bool can_follow_write_pmd(pmd_t pmd, st= ruct page *page, return !userfaultfd_huge_pmd_wp(vma, pmd); } =20 -static struct page *follow_huge_pmd(struct vm_area_struct *vma, - unsigned long addr, pmd_t *pmd, - unsigned int flags, - unsigned long *page_mask) +static long follow_huge_pmd(struct vm_area_struct *vma, + unsigned long addr, pmd_t *pmd, + unsigned int flags, unsigned long *page_mask, + struct page **pages) { struct mm_struct *mm =3D vma->vm_mm; pmd_t pmdval =3D *pmd; @@ -730,24 +733,24 @@ static struct page *follow_huge_pmd(struct vm_area_st= ruct *vma, page =3D pmd_page(pmdval); if ((flags & FOLL_WRITE) && !can_follow_write_pmd(pmdval, page, vma, flags)) - return NULL; + return 0; =20 /* Avoid dumping huge zero page */ if ((flags & FOLL_DUMP) && is_huge_zero_pmd(pmdval)) - return ERR_PTR(-EFAULT); + return -EFAULT; =20 if (pmd_protnone(*pmd) && !gup_can_follow_protnone(vma, flags)) - return NULL; + return 0; =20 if (!pmd_write(pmdval) && gup_must_unshare(vma, flags, page)) - return ERR_PTR(-EMLINK); + return -EMLINK; =20 VM_WARN_ON_ONCE_PAGE((flags & FOLL_PIN) && PageAnon(page) && !PageAnonExclusive(page), page); =20 ret =3D try_grab_folio(page_folio(page), 1, flags); if (ret) - return ERR_PTR(ret); + return ret; =20 #ifdef CONFIG_TRANSPARENT_HUGEPAGE if (pmd_trans_huge(pmdval) && (flags & FOLL_TOUCH)) @@ -757,23 +760,26 @@ static struct page *follow_huge_pmd(struct vm_area_st= ruct *vma, page +=3D (addr & ~HPAGE_PMD_MASK) >> PAGE_SHIFT; *page_mask =3D HPAGE_PMD_NR - 1; =20 - return page; + gup_fill_pages(vma, addr, page, 1, pages); + + return 1; } =20 #else /* CONFIG_PGTABLE_HAS_HUGE_LEAVES */ -static struct page *follow_huge_pud(struct vm_area_struct *vma, - unsigned long addr, pud_t *pudp, - int flags, unsigned long *page_mask) +static long follow_huge_pud(struct vm_area_struct *vma, + unsigned long addr, pud_t *pudp, + unsigned int flags, unsigned long *page_mask, + struct page **pages) { - return NULL; + return 0; } =20 -static struct page *follow_huge_pmd(struct vm_area_struct *vma, - unsigned long addr, pmd_t *pmd, - unsigned int flags, - unsigned long *page_mask) +static long follow_huge_pmd(struct vm_area_struct *vma, + unsigned long addr, pmd_t *pmd, + unsigned int flags, unsigned long *page_mask, + struct page **pages) { - return NULL; + return 0; } #endif /* CONFIG_PGTABLE_HAS_HUGE_LEAVES */ =20 @@ -816,15 +822,16 @@ static inline bool can_follow_write_pte(pte_t pte, st= ruct page *page, return !userfaultfd_pte_wp(vma, pte); } =20 -static struct page *follow_page_pte(struct vm_area_struct *vma, - unsigned long address, pmd_t *pmd, unsigned int flags) +static long follow_page_pte(struct vm_area_struct *vma, + unsigned long address, pmd_t *pmd, unsigned int flags, + struct page **pages) { struct mm_struct *mm =3D vma->vm_mm; struct folio *folio; struct page *page; spinlock_t *ptl; pte_t *ptep, pte; - int ret; + long ret; =20 ptep =3D pte_offset_map_lock(mm, pmd, address, &ptl); if (!ptep) @@ -842,14 +849,14 @@ static struct page *follow_page_pte(struct vm_area_st= ruct *vma, */ if ((flags & FOLL_WRITE) && !can_follow_write_pte(pte, page, vma, flags)) { - page =3D NULL; + ret =3D 0; goto out; } =20 if (unlikely(!page)) { if (flags & FOLL_DUMP) { /* Avoid special (like zero) pages in core dumps */ - page =3D ERR_PTR(-EFAULT); + ret =3D -EFAULT; goto out; } =20 @@ -857,14 +864,13 @@ static struct page *follow_page_pte(struct vm_area_st= ruct *vma, page =3D pte_page(pte); } else { ret =3D follow_pfn_pte(vma, address, ptep, flags); - page =3D ERR_PTR(ret); goto out; } } folio =3D page_folio(page); =20 if (!pte_write(pte) && gup_must_unshare(vma, flags, page)) { - page =3D ERR_PTR(-EMLINK); + ret =3D -EMLINK; goto out; } =20 @@ -873,10 +879,8 @@ static struct page *follow_page_pte(struct vm_area_str= uct *vma, =20 /* try_grab_folio() does nothing unless FOLL_GET or FOLL_PIN is set. */ ret =3D try_grab_folio(folio, 1, flags); - if (unlikely(ret)) { - page =3D ERR_PTR(ret); + if (unlikely(ret)) goto out; - } =20 /* * We need to make the page accessible if and only if we are going @@ -886,8 +890,7 @@ static struct page *follow_page_pte(struct vm_area_stru= ct *vma, if (flags & FOLL_PIN) { ret =3D arch_make_folio_accessible(folio); if (ret) { - unpin_user_page(page); - page =3D ERR_PTR(ret); + gup_put_folio(folio, 1, flags); goto out; } } @@ -902,24 +905,27 @@ static struct page *follow_page_pte(struct vm_area_st= ruct *vma, */ folio_mark_accessed(folio); } + + gup_fill_pages(vma, address, page, 1, pages); + ret =3D 1; out: pte_unmap_unlock(ptep, ptl); - return page; + return ret; no_page: pte_unmap_unlock(ptep, ptl); if (!pte_none(pte)) - return NULL; + return 0; return no_page_table(vma, flags, address); } =20 -static struct page *follow_pmd_mask(struct vm_area_struct *vma, - unsigned long address, pud_t *pudp, - unsigned int flags, - unsigned long *page_mask) +static long follow_pmd_mask(struct vm_area_struct *vma, + unsigned long address, pud_t *pudp, + unsigned int flags, unsigned long *page_mask, + struct page **pages) { pmd_t *pmd, pmdval; spinlock_t *ptl; - struct page *page; + long ret; struct mm_struct *mm =3D vma->vm_mm; =20 pmd =3D pmd_offset(pudp, address); @@ -929,7 +935,7 @@ static struct page *follow_pmd_mask(struct vm_area_stru= ct *vma, if (!pmd_present(pmdval)) return no_page_table(vma, flags, address); if (likely(!pmd_leaf(pmdval))) - return follow_page_pte(vma, address, pmd, flags); + return follow_page_pte(vma, address, pmd, flags, pages); =20 if (pmd_protnone(pmdval) && !gup_can_follow_protnone(vma, flags)) return no_page_table(vma, flags, address); @@ -942,28 +948,28 @@ static struct page *follow_pmd_mask(struct vm_area_st= ruct *vma, } if (unlikely(!pmd_leaf(pmdval))) { spin_unlock(ptl); - return follow_page_pte(vma, address, pmd, flags); + return follow_page_pte(vma, address, pmd, flags, pages); } if (pmd_trans_huge(pmdval) && (flags & FOLL_SPLIT_PMD)) { spin_unlock(ptl); split_huge_pmd(vma, pmd, address); /* If pmd was left empty, stuff a page table in there quickly */ - return pte_alloc(mm, pmd) ? ERR_PTR(-ENOMEM) : - follow_page_pte(vma, address, pmd, flags); + return pte_alloc(mm, pmd) ? -ENOMEM : + follow_page_pte(vma, address, pmd, flags, pages); } - page =3D follow_huge_pmd(vma, address, pmd, flags, page_mask); + ret =3D follow_huge_pmd(vma, address, pmd, flags, page_mask, pages); spin_unlock(ptl); - return page; + return ret; } =20 -static struct page *follow_pud_mask(struct vm_area_struct *vma, - unsigned long address, p4d_t *p4dp, - unsigned int flags, - unsigned long *page_mask) +static long follow_pud_mask(struct vm_area_struct *vma, + unsigned long address, p4d_t *p4dp, + unsigned int flags, unsigned long *page_mask, + struct page **pages) { pud_t *pudp, pud; spinlock_t *ptl; - struct page *page; + long ret; struct mm_struct *mm =3D vma->vm_mm; =20 pudp =3D pud_offset(p4dp, address); @@ -972,22 +978,22 @@ static struct page *follow_pud_mask(struct vm_area_st= ruct *vma, return no_page_table(vma, flags, address); if (pud_leaf(pud)) { ptl =3D pud_lock(mm, pudp); - page =3D follow_huge_pud(vma, address, pudp, flags, page_mask); + ret =3D follow_huge_pud(vma, address, pudp, flags, page_mask, pages); spin_unlock(ptl); - if (page) - return page; + if (ret) + return ret; return no_page_table(vma, flags, address); } if (unlikely(pud_bad(pud))) return no_page_table(vma, flags, address); =20 - return follow_pmd_mask(vma, address, pudp, flags, page_mask); + return follow_pmd_mask(vma, address, pudp, flags, page_mask, pages); } =20 -static struct page *follow_p4d_mask(struct vm_area_struct *vma, - unsigned long address, pgd_t *pgdp, - unsigned int flags, - unsigned long *page_mask) +static long follow_p4d_mask(struct vm_area_struct *vma, + unsigned long address, pgd_t *pgdp, + unsigned int flags, unsigned long *page_mask, + struct page **pages) { p4d_t *p4dp, p4d; =20 @@ -998,7 +1004,7 @@ static struct page *follow_p4d_mask(struct vm_area_str= uct *vma, if (!p4d_present(p4d) || p4d_bad(p4d)) return no_page_table(vma, flags, address); =20 - return follow_pud_mask(vma, address, p4dp, flags, page_mask); + return follow_pud_mask(vma, address, p4dp, flags, page_mask, pages); } =20 /** @@ -1007,6 +1013,9 @@ static struct page *follow_p4d_mask(struct vm_area_st= ruct *vma, * @address: virtual address to look up * @flags: flags modifying lookup behaviour * @page_mask: a pointer to output page_mask + * @pages: array to receive the page found, refcounted per @flags, or NULL + * to walk the page tables (e.g. to fault pages in) without + * collecting or refcounting them * * @flags can have FOLL_ flags set, defined in * @@ -1017,17 +1026,17 @@ static struct page *follow_p4d_mask(struct vm_area_= struct *vma, * * On output, @page_mask is set according to the size of the page. * - * Return: the mapped (struct page *), %NULL if no mapping exists, or - * an error pointer if there is a mapping to something not represented - * by a page descriptor (see also vm_normal_page()). + * Return: 1 with @pages[0] filled in if a page was found, 0 if no mapping + * exists at @address, or a negative errno for a mapping to something not + * represented by a page descriptor (see also vm_normal_page()). */ -static struct page *follow_page_mask(struct vm_area_struct *vma, - unsigned long address, unsigned int flags, - unsigned long *page_mask) +static long follow_page_mask(struct vm_area_struct *vma, + unsigned long address, unsigned int flags, + unsigned long *page_mask, struct page **pages) { pgd_t *pgd; struct mm_struct *mm =3D vma->vm_mm; - struct page *page; + long ret; =20 vma_pgtable_walk_begin(vma); =20 @@ -1035,13 +1044,13 @@ static struct page *follow_page_mask(struct vm_area= _struct *vma, pgd =3D pgd_offset(mm, address); =20 if (pgd_none(*pgd) || unlikely(pgd_bad(*pgd))) - page =3D no_page_table(vma, flags, address); + ret =3D no_page_table(vma, flags, address); else - page =3D follow_p4d_mask(vma, address, pgd, flags, page_mask); + ret =3D follow_p4d_mask(vma, address, pgd, flags, page_mask, pages); =20 vma_pgtable_walk_end(vma); =20 - return page; + return ret; } =20 static int get_gate_page(struct mm_struct *mm, unsigned long address, @@ -1391,6 +1400,7 @@ static long __get_user_pages(struct mm_struct *mm, do { struct page *page; unsigned int page_increm; + long nr; =20 /* first iteration or cross vma bound */ if (!vma || start >=3D vma->vm_end) { @@ -1417,8 +1427,12 @@ static long __get_user_pages(struct mm_struct *mm, pages ? &page : NULL); if (ret) goto out; - page_mask =3D 0; - goto next_page; + gup_fill_pages(vma, start, page, 1, + pages ? pages + i : NULL); + i++; + start +=3D PAGE_SIZE; + nr_pages--; + continue; } =20 if (!vma) { @@ -1440,10 +1454,11 @@ static long __get_user_pages(struct mm_struct *mm, } cond_resched(); =20 - page =3D follow_page_mask(vma, start, gup_flags, &page_mask); - if (!page || PTR_ERR(page) =3D=3D -EMLINK) { + nr =3D follow_page_mask(vma, start, gup_flags, &page_mask, + pages ? &pages[i] : NULL); + if (!nr || nr =3D=3D -EMLINK) { ret =3D faultin_page(vma, start, gup_flags, - PTR_ERR(page) =3D=3D -EMLINK, locked); + nr =3D=3D -EMLINK, locked); switch (ret) { case 0: goto retry; @@ -1457,7 +1472,7 @@ static long __get_user_pages(struct mm_struct *mm, goto out; } BUG(); - } else if (PTR_ERR(page) =3D=3D -EEXIST) { + } else if (nr =3D=3D -EEXIST) { /* * Proper page table entry exists, but no corresponding * struct page. If the caller expects **pages to be @@ -1465,49 +1480,48 @@ static long __get_user_pages(struct mm_struct *mm, * for this page. */ if (pages) { - ret =3D PTR_ERR(page); + ret =3D nr; goto out; } - } else if (IS_ERR(page)) { - ret =3D PTR_ERR(page); + } else if (nr < 0) { + ret =3D nr; goto out; } -next_page: + page_increm =3D 1 + (~(start >> PAGE_SHIFT) & page_mask); if (page_increm > nr_pages) page_increm =3D nr_pages; =20 - if (pages) { + /* + * This must be a large folio (and doesn't need to + * be the whole folio; it can be part of it), do + * the refcount work for all the subpages too. + * + * NOTE: here the page may not be the head page + * e.g. when start addr is not thp-size aligned. + * try_grab_folio() should have taken care of tail + * pages. + */ + if (pages && page_increm > 1) { + struct folio *folio =3D page_folio(pages[i]); + /* - * This must be a large folio (and doesn't need to - * be the whole folio; it can be part of it), do - * the refcount work for all the subpages too. - * - * NOTE: here the page may not be the head page - * e.g. when start addr is not thp-size aligned. - * try_grab_folio() should have taken care of tail - * pages. + * Since we already hold refcount on the + * large folio, this should never fail. */ - if (page_increm > 1) { - struct folio *folio =3D page_folio(page); - + if (try_grab_folio(folio, page_increm - 1, + gup_flags)) { /* - * Since we already hold refcount on the - * large folio, this should never fail. + * Release the 1st page ref if the + * folio is problematic, fail hard. */ - if (try_grab_folio(folio, page_increm - 1, - gup_flags)) { - /* - * Release the 1st page ref if the - * folio is problematic, fail hard. - */ - gup_put_folio(folio, 1, gup_flags); - ret =3D -EFAULT; - goto out; - } + gup_put_folio(folio, 1, gup_flags); + ret =3D -EFAULT; + goto out; } =20 - gup_fill_pages(vma, start, page, page_increm, pages + i); + gup_fill_pages(vma, start + PAGE_SIZE, pages[i] + 1, + page_increm - 1, pages + i + 1); } =20 i +=3D page_increm; --=20 2.55.0 From nobody Tue Sep 29 07:39:19 2026 Received: from shelob.surriel.com (shelob.surriel.com [96.67.55.147]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7479B3911A9 for ; Tue, 11 Aug 2026 03:07:14 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=96.67.55.147 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786417635; cv=none; b=C7nsE+fO9CkKROD44q/WEa9hFHswyjVE2RWQXimzL5xZqglSG86nPFxKbuXoKq75vgpZ3jb/IRU+9rAFPBVdjTlxG5+rJClCFzPbz3jHHU5K8OFHVft0R3L/ysxN8NLYDb9qKNXtd/FW5dS7Ux815ZMMHAvAW+cDRvAE/0Ib2iY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786417635; c=relaxed/simple; bh=6k1vvZ2QCIkSSPB4LwmfDbY1MsTRlItzPMSL+ZJDfYM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=jTGi9sLTqD3I3x+bQubjRci/cYTNww/9LyM1yGrSB779bj9i8Xn8W7HgS0aZCv8jcYQ6/k7k6qm3Ffn/0KsalTBhoQW38Q3RPS+hSacZkOdnc3KnNO1iSuVz+obP4qKCRqp8TT8cEkJTEZzDXs8rXW+1NNsCyTF/hrhUZWw4EQk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=surriel.com; spf=pass smtp.mailfrom=surriel.com; dkim=pass (2048-bit key) header.d=surriel.com header.i=@surriel.com header.b=cfi4HImM; arc=none smtp.client-ip=96.67.55.147 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=surriel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=surriel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=surriel.com header.i=@surriel.com header.b="cfi4HImM" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=surriel.com ; s=mail; h=Content-Transfer-Encoding:MIME-Version:References:In-Reply-To: Message-ID:Date:Subject:Cc:To:From:Sender:Reply-To:Content-Type:Content-ID: Content-Description:Resent-Date:Resent-From:Resent-Sender:Resent-To:Resent-Cc :Resent-Message-ID:List-Id:List-Help:List-Unsubscribe:List-Subscribe: List-Post:List-Owner:List-Archive; bh=3RZbPAYAvl32QwZ/NBytlP1RvgJPCM8BP5+kfIW6nbE=; b=cfi4HImMhV4qUUW5bIGJpJ4mnR 1PShb1UDO+ftIPtt6/K82ZZU3h6dbBWlKUKmBcLeJEUes3hkHsg7lKqdRz+5TCrs2jD/3MxpW6f7p 8wuKx3bFxJqDgVTIcaZbSJFT3/SuX2lT1sET1WOh0FQWmitcQqj+STzd8xwbPk5qLNPVIYXDPdoUn JELJvxGnXpEnyq2xGcROLAMLyppdays83XuaOfcoAxw1ahS73RLe1j2k4rlUTiGUYf6zusRVCPaOY AH8W6UKFdm8zcZi0wVJujTMMpjcU/t59Qv0RW9ilGCWLqPlOQQhNUh3XmhaJsJTDJ+sU8bULblfFO xys5upwQ==; Received: from [96.67.55.146] (helo=fangorn.surriel.com) by shelob.surriel.com with esmtpsa (TLS1.2) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.97.1) (envelope-from ) id 1wtcb6-000000007kZ-2X8W; Mon, 10 Aug 2026 22:52:12 -0400 From: Rik van Riel To: linux-kernel@vger.kernel.org Cc: kernel-team@meta.com, Rik van Riel , Andrew Morton , David Hildenbrand , Jason Gunthorpe , John Hubbard , Peter Xu , linux-mm@kvack.org Subject: [RFC PATCH v3 3/8] mm/gup: split follow_page_pte_commit() out of follow_page_pte() Date: Mon, 10 Aug 2026 22:51:52 -0400 Message-ID: <20260811025157.1632867-4-riel@surriel.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260811025157.1632867-1-riel@surriel.com> References: <20260811025157.1632867-1-riel@surriel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" follow_page_pte() does two things once it has resolved a present PTE to a page: run the per-PTE safety checks (write-fault, unshare), then commit to that page: grab a ref, fault it in if pinning, mark it dirty/accessed, and hand it back to the caller. Split the second part into its own follow_page_pte_commit(), unchanged except for taking its inputs as parameters instead of local variables, so the checks and the commit can be applied at different granularities. No functional changes intended. Suggested-by: David Hildenbrand Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Rik van Riel --- mm/gup.c | 78 +++++++++++++++++++++++++++++++++++--------------------- 1 file changed, 49 insertions(+), 29 deletions(-) diff --git a/mm/gup.c b/mm/gup.c index e4e6d0993424..b755ceaac0f5 100644 --- a/mm/gup.c +++ b/mm/gup.c @@ -822,6 +822,52 @@ static inline bool can_follow_write_pte(pte_t pte, str= uct page *page, return !userfaultfd_pte_wp(vma, pte); } =20 +/* + * The caller has already run every per-PTE safety check (present, + * write-fault, gup_must_unshare()) on the PTE, so this only does the + * per-folio work: the refcount grab, the FOLL_PIN accessibility fault-in, + * dirty/accessed marking, and the array fill with the cache flush. + */ +static long follow_page_pte_commit(struct vm_area_struct *vma, + unsigned long address, struct folio *folio, struct page *page, + pte_t pte, unsigned int flags, struct page **pages) +{ + long ret; + + /* try_grab_folio() does nothing unless FOLL_GET or FOLL_PIN is set. */ + ret =3D try_grab_folio(folio, 1, flags); + if (unlikely(ret)) + return ret; + + /* + * We need to make the page accessible if and only if we are going + * to access its content (the FOLL_PIN case). Please see + * Documentation/core-api/pin_user_pages.rst for details. + */ + if (flags & FOLL_PIN) { + ret =3D arch_make_folio_accessible(folio); + if (ret) { + gup_put_folio(folio, 1, flags); + return ret; + } + } + if (flags & FOLL_TOUCH) { + if ((flags & FOLL_WRITE) && + !pte_dirty(pte) && !folio_test_dirty(folio)) + folio_mark_dirty(folio); + /* + * pte_mkyoung() would be more correct here, but atomic care + * is needed to avoid losing the dirty bit: it is easier to use + * folio_mark_accessed(). + */ + folio_mark_accessed(folio); + } + + gup_fill_pages(vma, address, page, 1, pages); + + return 0; +} + static long follow_page_pte(struct vm_area_struct *vma, unsigned long address, pmd_t *pmd, unsigned int flags, struct page **pages) @@ -877,36 +923,10 @@ static long follow_page_pte(struct vm_area_struct *vm= a, VM_WARN_ON_ONCE_PAGE((flags & FOLL_PIN) && PageAnon(page) && !PageAnonExclusive(page), page); =20 - /* try_grab_folio() does nothing unless FOLL_GET or FOLL_PIN is set. */ - ret =3D try_grab_folio(folio, 1, flags); - if (unlikely(ret)) + ret =3D follow_page_pte_commit(vma, address, folio, page, pte, flags, + pages); + if (ret) goto out; - - /* - * We need to make the page accessible if and only if we are going - * to access its content (the FOLL_PIN case). Please see - * Documentation/core-api/pin_user_pages.rst for details. - */ - if (flags & FOLL_PIN) { - ret =3D arch_make_folio_accessible(folio); - if (ret) { - gup_put_folio(folio, 1, flags); - goto out; - } - } - if (flags & FOLL_TOUCH) { - if ((flags & FOLL_WRITE) && - !pte_dirty(pte) && !folio_test_dirty(folio)) - folio_mark_dirty(folio); - /* - * pte_mkyoung() would be more correct here, but atomic care - * is needed to avoid losing the dirty bit: it is easier to use - * folio_mark_accessed(). - */ - folio_mark_accessed(folio); - } - - gup_fill_pages(vma, address, page, 1, pages); ret =3D 1; out: pte_unmap_unlock(ptep, ptl); --=20 2.55.0 From nobody Tue Sep 29 07:39:19 2026 Received: from shelob.surriel.com (shelob.surriel.com [96.67.55.147]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 014DD2DC798 for ; Tue, 11 Aug 2026 03:07:13 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=96.67.55.147 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786417635; cv=none; b=M6Mh3K7B9V1MzSUTp4mZydYgnqBgn0pt6GIvv/hztRQAcc84ijLnxOeBBtPmtWDgYkUT+336Ys+v+Os9b9oSpgYPi+mvn4NkNUN1nopvUUkYaZnx2prYFBMX++5rbwg9EK3n5yicx/qT1EObU1e/U4bJRMsEuJ2CYoej8gtgfZ4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786417635; c=relaxed/simple; bh=0FlWuYknOh3GgSqK/YQrNa+XApS+rXvH5zFyj97YRCc=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=cCYeYDsCicWFHC4V2FTvSGedrg/zwVLqAw6UtMsxr5WAtpKo44C4d7BMsdoPl0/xjyn7BmGdNmrk+Ns2dS/USj4P0sA0yPEsrOJfgS2WcLrJbkksqD6CqOUL+fjk8Gpkj3QWJfheFFJ/yIFfELUJrAmTlDuvgXUUM1F7K2gynUo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=surriel.com; spf=pass smtp.mailfrom=surriel.com; dkim=pass (2048-bit key) header.d=surriel.com header.i=@surriel.com header.b=Fe3qezAr; arc=none smtp.client-ip=96.67.55.147 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=surriel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=surriel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=surriel.com header.i=@surriel.com header.b="Fe3qezAr" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=surriel.com ; s=mail; h=Content-Transfer-Encoding:MIME-Version:References:In-Reply-To: Message-ID:Date:Subject:Cc:To:From:Sender:Reply-To:Content-Type:Content-ID: Content-Description:Resent-Date:Resent-From:Resent-Sender:Resent-To:Resent-Cc :Resent-Message-ID:List-Id:List-Help:List-Unsubscribe:List-Subscribe: List-Post:List-Owner:List-Archive; bh=cng7GWmuH0afMBNJaOnPoRqJa9xaiwRqJYVnLxI+0+w=; b=Fe3qezArnzVz1Pb+5pmpOAIgcB UW4WK1nDwE9bLqjp7C3Jj5yV5sXselN+PwuUh0zTTGkmfN6t4kesgCEezvcAsF8IYx4uLZ6OSpGU3 GLrDNS4XVzS4vR7Td68MCOqoOXPryHSP9IWYuQ7oQRVQemVnJycOy1DTPYVd5fx7tbAwNPIdK84Em ztZG/gBdtb41/ZzHoNEcRbKNafwo1897lrhZYSOJNO9gAv10xQJw25N/QbQzhDqvGw/3YNdwwwcJu ef7QHar4jrMtViZK9BuCZdT5REoBM/bAOn2ZnsFIziLSRo5aF6CFwgZ2VhIpesOb+DYeR0MhFoEU5 zDLAwYyg==; Received: from [96.67.55.146] (helo=fangorn.surriel.com) by shelob.surriel.com with esmtpsa (TLS1.2) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.97.1) (envelope-from ) id 1wtcb6-000000007kZ-2dKV; Mon, 10 Aug 2026 22:52:12 -0400 From: Rik van Riel To: linux-kernel@vger.kernel.org Cc: kernel-team@meta.com, Rik van Riel , Andrew Morton , David Hildenbrand , Jason Gunthorpe , John Hubbard , Peter Xu , linux-mm@kvack.org Subject: [RFC PATCH v3 4/8] mm/gup: break out follow_one_pte() helper Date: Mon, 10 Aug 2026 22:51:53 -0400 Message-ID: <20260811025157.1632867-5-riel@surriel.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260811025157.1632867-1-riel@surriel.com> References: <20260811025157.1632867-1-riel@surriel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" follow_page_pte() is 92 lines and does two separate things: work out which page a PTE maps, if any, and commit to the page it found. The first half reaches the second through five exit paths, two of which unwind the PTE lock in different ways. Split the resolve half into follow_one_pte(), which returns the page it resolved, NULL when the PTE cannot be followed, or the errno the caller must report. follow_page_pte() is left with one unlock and one exit. no_page_table() can look up the page cache, so the case that needs it is recorded and the call made after dropping the PTE lock, as before. No functional changes intended. Suggested-by: David Hildenbrand Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Rik van Riel --- mm/gup.c | 100 +++++++++++++++++++++++++++++++------------------------ 1 file changed, 56 insertions(+), 44 deletions(-) diff --git a/mm/gup.c b/mm/gup.c index b755ceaac0f5..5af6a23285de 100644 --- a/mm/gup.c +++ b/mm/gup.c @@ -868,74 +868,86 @@ static long follow_page_pte_commit(struct vm_area_str= uct *vma, return 0; } =20 -static long follow_page_pte(struct vm_area_struct *vma, - unsigned long address, pmd_t *pmd, unsigned int flags, - struct page **pages) +/* + * Resolve one present PTE to the page it maps. Returns no page and no err= or + * when the PTE cannot be followed but the caller may fault it in, and a + * negative errno when the caller must report the failure. + */ +static long follow_one_pte(struct vm_area_struct *vma, unsigned long addre= ss, + pte_t *ptep, pte_t pte, unsigned int flags, struct page **pagep) { - struct mm_struct *mm =3D vma->vm_mm; - struct folio *folio; struct page *page; - spinlock_t *ptl; - pte_t *ptep, pte; - long ret; =20 - ptep =3D pte_offset_map_lock(mm, pmd, address, &ptl); - if (!ptep) - return no_page_table(vma, flags, address); - pte =3D ptep_get(ptep); + *pagep =3D NULL; + if (!pte_present(pte)) - goto no_page; + return 0; if (pte_protnone(pte) && !gup_can_follow_protnone(vma, flags)) - goto no_page; + return 0; =20 page =3D vm_normal_page(vma, address, pte); =20 /* * We only care about anon pages in can_follow_write_pte(). */ - if ((flags & FOLL_WRITE) && - !can_follow_write_pte(pte, page, vma, flags)) { - ret =3D 0; - goto out; - } + if ((flags & FOLL_WRITE) && !can_follow_write_pte(pte, page, vma, flags)) + return 0; =20 if (unlikely(!page)) { if (flags & FOLL_DUMP) { /* Avoid special (like zero) pages in core dumps */ - ret =3D -EFAULT; - goto out; - } - - if (is_zero_pfn(pte_pfn(pte))) { - page =3D pte_page(pte); - } else { - ret =3D follow_pfn_pte(vma, address, ptep, flags); - goto out; + return -EFAULT; } + if (!is_zero_pfn(pte_pfn(pte))) + return follow_pfn_pte(vma, address, ptep, flags); + page =3D pte_page(pte); } - folio =3D page_folio(page); =20 - if (!pte_write(pte) && gup_must_unshare(vma, flags, page)) { - ret =3D -EMLINK; - goto out; - } + if (!pte_write(pte) && gup_must_unshare(vma, flags, page)) + return -EMLINK; =20 VM_WARN_ON_ONCE_PAGE((flags & FOLL_PIN) && PageAnon(page) && !PageAnonExclusive(page), page); =20 - ret =3D follow_page_pte_commit(vma, address, folio, page, pte, flags, - pages); - if (ret) - goto out; - ret =3D 1; -out: + *pagep =3D page; + return 0; +} + +static long follow_page_pte(struct vm_area_struct *vma, + unsigned long address, pmd_t *pmd, unsigned int flags, + struct page **pages) +{ + struct mm_struct *mm =3D vma->vm_mm; + bool need_no_page_table =3D false; + struct page *page; + spinlock_t *ptl; + pte_t *ptep, pte; + long ret; + + ptep =3D pte_offset_map_lock(mm, pmd, address, &ptl); + if (!ptep) + return no_page_table(vma, flags, address); + pte =3D ptep_get(ptep); + + ret =3D follow_one_pte(vma, address, ptep, pte, flags, &page); + if (!ret && page) { + ret =3D follow_page_pte_commit(vma, address, page_folio(page), + page, pte, flags, pages); + if (!ret) + ret =3D 1; + } else if (!ret && pte_none(pte)) { + /* + * no_page_table() may look up the page cache, so it cannot run + * under the PTE lock. + */ + need_no_page_table =3D true; + } + pte_unmap_unlock(ptep, ptl); + + if (need_no_page_table) + return no_page_table(vma, flags, address); return ret; -no_page: - pte_unmap_unlock(ptep, ptl); - if (!pte_none(pte)) - return 0; - return no_page_table(vma, flags, address); } =20 static long follow_pmd_mask(struct vm_area_struct *vma, --=20 2.55.0 From nobody Tue Sep 29 07:39:19 2026 Received: from shelob.surriel.com (shelob.surriel.com [96.67.55.147]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2FBA529D26E for ; Tue, 11 Aug 2026 03:06:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=96.67.55.147 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786417618; cv=none; b=HiFMEYxCtvd9u+AlpbpKILj/v2nJIhUxS3P63a+LiM07ySYKqbUJEE+pApPowySGw+b8DWlVxpWbje4NpVXARktz/dt7hGRnp9xlFH3J6wNbwiyRb58gLksXz2RYHfYVeIbcK0eViMNpvBaEPRdME+wbSmydxWurnQg4IQ4V87I= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786417618; c=relaxed/simple; bh=p5xlxtIn7PtXM3ZLVKijqA4Ww+a09LWL3JdL4fMgiVM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=ESipnH3gv81dDJVRMJnkkJakKbP1Qg3gf372s7zPut4yLwN7/boOcDDcmAWo8fdnYrKmJhdWDYYdqMQxa4zCzQ6Q+hDUMc2MIj2NDBpgg1YbNLakSYac1WbfidqfJt9XguVcD50LOeIOjYIn+RlZ1WSl01dUC7K2uMrcJ9LUI5Y= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=surriel.com; spf=pass smtp.mailfrom=surriel.com; dkim=pass (2048-bit key) header.d=surriel.com header.i=@surriel.com header.b=motWnnzm; arc=none smtp.client-ip=96.67.55.147 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=surriel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=surriel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=surriel.com header.i=@surriel.com header.b="motWnnzm" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=surriel.com ; s=mail; h=Content-Transfer-Encoding:MIME-Version:References:In-Reply-To: Message-ID:Date:Subject:Cc:To:From:Sender:Reply-To:Content-Type:Content-ID: Content-Description:Resent-Date:Resent-From:Resent-Sender:Resent-To:Resent-Cc :Resent-Message-ID:List-Id:List-Help:List-Unsubscribe:List-Subscribe: List-Post:List-Owner:List-Archive; bh=r6D600qISDM+QQRVU2p7DCVxL1IDlJvuL+ELgqy5iEU=; b=motWnnzmr2o6TRv+iCXEnjqKHX HLOaQ8VhFNTgoIunHfm4W00/hOSLu5hCbKq1ic/tDm/sPKI8TDIIVhnDfCaMwCTmb+RiXkyDv6N6b pwouy4q9lCveHsVrBNToskkxLIEnnvsX/4KZ4NDKlQYzBT1FoZOOKtpWvVr4hzeIm/0v3I1RUYPL3 YYuVBkxyry2g1AWm0cRqgiCLnMHG+z7QqYdpwesaGR3fojqLcc99SxUcvuzAca5VPQq7ihx82eblv 3ljsk9hAz4sJjNMIPOs7AsxBFpZS2Z5uzXUpYoVDzXKarvQbgdTlvC+hnwgXgqlXID/+zhTHaCEv7 ggh58I3Q==; Received: from [96.67.55.146] (helo=fangorn.surriel.com) by shelob.surriel.com with esmtpsa (TLS1.2) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.97.1) (envelope-from ) id 1wtcb6-000000007kZ-2ixu; Mon, 10 Aug 2026 22:52:12 -0400 From: Rik van Riel To: linux-kernel@vger.kernel.org Cc: kernel-team@meta.com, Rik van Riel , Andrew Morton , David Hildenbrand , Jason Gunthorpe , John Hubbard , Peter Xu , linux-mm@kvack.org Subject: [RFC PATCH v3 5/8] mm/gup: fill the pages array outside the pud/pmd lock Date: Mon, 10 Aug 2026 22:51:54 -0400 Message-ID: <20260811025157.1632867-6-riel@surriel.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260811025157.1632867-1-riel@surriel.com> References: <20260811025157.1632867-1-riel@surriel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" follow_huge_pud() and follow_huge_pmd() fill pages[] and flush the page's caches while still holding the pud or pmd lock. Neither flush_anon_page() nor flush_dcache_page() needs that lock. Have the huge paths store the page and let follow_pud_mask() and follow_pmd_mask() do the fill after they unlock, so the flushes happen outside the critical section. This should be safe because try_grab_folio() has already taken a folio reference before the unlock, so nothing can free the page while the fill runs, and the fill itself touches neither the page tables nor the pud or pmd entry it was reached through. No functional changes intended. Suggested-by: David Hildenbrand Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Rik van Riel Reviewed-by: Suren Baghdasaryan --- mm/gup.c | 20 ++++++++++++++++++-- 1 file changed, 18 insertions(+), 2 deletions(-) diff --git a/mm/gup.c b/mm/gup.c index 5af6a23285de..4036d3dc27df 100644 --- a/mm/gup.c +++ b/mm/gup.c @@ -695,7 +695,8 @@ static long follow_huge_pud(struct vm_area_struct *vma, =20 *page_mask =3D HPAGE_PUD_NR - 1; =20 - gup_fill_pages(vma, addr, page, 1, pages); + if (pages) + pages[0] =3D page; =20 return 1; } @@ -760,7 +761,8 @@ static long follow_huge_pmd(struct vm_area_struct *vma, page +=3D (addr & ~HPAGE_PMD_MASK) >> PAGE_SHIFT; *page_mask =3D HPAGE_PMD_NR - 1; =20 - gup_fill_pages(vma, addr, page, 1, pages); + if (pages) + pages[0] =3D page; =20 return 1; } @@ -991,6 +993,14 @@ static long follow_pmd_mask(struct vm_area_struct *vma, } ret =3D follow_huge_pmd(vma, address, pmd, flags, page_mask, pages); spin_unlock(ptl); + + /* + * The ref is already held, so the page cannot go away: fill the + * array and flush caches without the pmd lock. + */ + if (ret > 0 && pages) + gup_fill_pages(vma, address, pages[0], ret, pages); + return ret; } =20 @@ -1012,6 +1022,12 @@ static long follow_pud_mask(struct vm_area_struct *v= ma, ptl =3D pud_lock(mm, pudp); ret =3D follow_huge_pud(vma, address, pudp, flags, page_mask, pages); spin_unlock(ptl); + /* + * The ref is already held, so the page cannot go away: fill + * the array and flush caches without the lock. + */ + if (ret > 0 && pages) + gup_fill_pages(vma, address, pages[0], ret, pages); if (ret) return ret; return no_page_table(vma, flags, address); --=20 2.55.0 From nobody Tue Sep 29 07:39:19 2026 Received: from shelob.surriel.com (shelob.surriel.com [96.67.55.147]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2FB2B1D5146 for ; Tue, 11 Aug 2026 03:06:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=96.67.55.147 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786417619; cv=none; b=ECRjTkUHqtwuumbd9dRTp0tAfWxTPlbEw/Nhtad+8apVrOaiDbUDDb4sxMh3v61NCYFyUsPa7dxFRh7pdC8ZSZxp9ceh3gU90N5Y5oVa7UMCbLZK6dl3YdzN35yfIfn2pXtBumIdeqQMrJV9Xl9Tm/87AwyGif7s65FLEZ7zG6k= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786417619; c=relaxed/simple; bh=uNXTgSwIiY69QCRVLMrPl0PQmdKhkSYUGPG8AXFv7YI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=vEgxJD0scHwy7Xy2LWShTshxuTHqdqeabcBIV3YqN94ZuJkuDCewoXLRcAUpVLC772iI0gNAS01PixFoJcq5F8X3I2vz8xZepsgVvH27dn3+aiIdImSyVRp2GequeyzOH3Mic5rR2422qJNQfYGSz89SRa1g3qtParm5kM3rPeA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=surriel.com; spf=pass smtp.mailfrom=surriel.com; dkim=pass (2048-bit key) header.d=surriel.com header.i=@surriel.com header.b=eVYMgUEH; arc=none smtp.client-ip=96.67.55.147 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=surriel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=surriel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=surriel.com header.i=@surriel.com header.b="eVYMgUEH" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=surriel.com ; s=mail; h=Content-Transfer-Encoding:MIME-Version:References:In-Reply-To: Message-ID:Date:Subject:Cc:To:From:Sender:Reply-To:Content-Type:Content-ID: Content-Description:Resent-Date:Resent-From:Resent-Sender:Resent-To:Resent-Cc :Resent-Message-ID:List-Id:List-Help:List-Unsubscribe:List-Subscribe: List-Post:List-Owner:List-Archive; bh=o4/cELZKpN9Jscfu/G3Szd5piq+pvPYj8i+M2wz0YJg=; b=eVYMgUEH65xNmDXoq7mZCQ6Eg0 Z66EcMWnUJ4SSa8XU3Tagr4plr2E3bI6utDinRBnVSlhAwJaQqrUwcqlNvWV40OckRjrxohMjhcAQ s2q3booJTy0w5BfLdr9AODWKSGe++cTlTc+yNrUZerw0/UmD774DewWnG2A8lNerWaj669wJe1xeR z4liytRtQzyg2bAhd0BOPjhbzCQBe/MNLmjV9lo86N8v0P5V2VWlVUzEGaBWRslCo57PGZPe+XuLY Q+pIJ9UpB1bgWeejKM2F5wRDJA+cd7HjDM2IofmvEb3Cfo2EPz1IfFAR6DTta1GQZoMDn/ml+sCi2 dWObgaIg==; Received: from [96.67.55.146] (helo=fangorn.surriel.com) by shelob.surriel.com with esmtpsa (TLS1.2) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.97.1) (envelope-from ) id 1wtcb6-000000007kZ-2oev; Mon, 10 Aug 2026 22:52:12 -0400 From: Rik van Riel To: linux-kernel@vger.kernel.org Cc: kernel-team@meta.com, Rik van Riel , Andrew Morton , David Hildenbrand , Jason Gunthorpe , John Hubbard , Peter Xu , linux-mm@kvack.org Subject: [RFC PATCH v3 6/8] mm/gup: return a huge page's full count from follow_page_mask() Date: Mon, 10 Aug 2026 22:51:55 -0400 Message-ID: <20260811025157.1632867-7-riel@surriel.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260811025157.1632867-1-riel@surriel.com> References: <20260811025157.1632867-1-riel@surriel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" follow_huge_pud()/follow_huge_pmd() already know the huge page's full size but report it via a separate *page_mask output; __get_user_pages() does a second try_grab_folio() call and subpage loop for everything past the first page. Have the huge paths report their count as the return value instead, clamped to the huge page's size and @end. The merged grab returns whatever error try_grab_folio() gives, instead of forcing -EFAULT on the second call's failure. *page_mask and __get_user_pages()'s second-grab/subpage loop are now dead; remove them. The old silent page_increm clamp becomes a VM_WARN_ON_ONCE, since @end already bounds the count and refs/pages[] are already committed by the time the caller sees it -- truncating here would leak references, not just waste a comparison. follow_page_pte() is unaffected, still returning at most 1 page. The -EEXIST path needs an explicit nr =3D 1 when pages =3D=3D NULL: it used= to get that from page_mask staying 0 for a PFN-special PTE, but nr holds -EEXIST there, which would grow nr_pages instead of shrinking it. mm/gup_test.c (PIN_LONGTERM_BENCHMARK) shows no measurable change; the lock-hold-time improvement is by inspection, not measurement. Suggested-by: David Hildenbrand Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Rik van Riel --- mm/gup.c | 150 +++++++++++++++++++++---------------------------------- 1 file changed, 58 insertions(+), 92 deletions(-) diff --git a/mm/gup.c b/mm/gup.c index 4036d3dc27df..ea2bb379183e 100644 --- a/mm/gup.c +++ b/mm/gup.c @@ -664,14 +664,14 @@ static inline bool can_follow_write_pud(pud_t pud, st= ruct page *page, } =20 static long follow_huge_pud(struct vm_area_struct *vma, - unsigned long addr, pud_t *pudp, - unsigned int flags, unsigned long *page_mask, - struct page **pages) + unsigned long addr, unsigned long end, pud_t *pudp, + unsigned int flags, struct page **pages) { struct mm_struct *mm =3D vma->vm_mm; struct page *page; pud_t pud =3D *pudp; unsigned long pfn =3D pud_pfn(pud); + unsigned long off, nr; int ret; =20 assert_spin_locked(pud_lockptr(mm, pudp)); @@ -683,22 +683,23 @@ static long follow_huge_pud(struct vm_area_struct *vm= a, !can_follow_write_pud(pud, pfn_to_page(pfn), vma, flags)) return 0; =20 - pfn +=3D (addr & ~PUD_MASK) >> PAGE_SHIFT; + off =3D PFN_DOWN(addr & ~PUD_MASK); + pfn +=3D off; page =3D pfn_to_page(pfn); =20 if (!pud_write(pud) && gup_must_unshare(vma, flags, page)) return -EMLINK; =20 - ret =3D try_grab_folio(page_folio(page), 1, flags); + nr =3D min(HPAGE_PUD_NR - off, PFN_DOWN(end - addr)); + + ret =3D try_grab_folio(page_folio(page), nr, flags); if (ret) return ret; =20 - *page_mask =3D HPAGE_PUD_NR - 1; - if (pages) pages[0] =3D page; =20 - return 1; + return nr; } =20 /* FOLL_FORCE can write to even unwritable PMDs in COW mappings. */ @@ -720,13 +721,13 @@ static inline bool can_follow_write_pmd(pmd_t pmd, st= ruct page *page, } =20 static long follow_huge_pmd(struct vm_area_struct *vma, - unsigned long addr, pmd_t *pmd, - unsigned int flags, unsigned long *page_mask, - struct page **pages) + unsigned long addr, unsigned long end, pmd_t *pmd, + unsigned int flags, struct page **pages) { struct mm_struct *mm =3D vma->vm_mm; pmd_t pmdval =3D *pmd; struct page *page; + unsigned long off, nr; int ret; =20 assert_spin_locked(pmd_lockptr(mm, pmd)); @@ -749,7 +750,10 @@ static long follow_huge_pmd(struct vm_area_struct *vma, VM_WARN_ON_ONCE_PAGE((flags & FOLL_PIN) && PageAnon(page) && !PageAnonExclusive(page), page); =20 - ret =3D try_grab_folio(page_folio(page), 1, flags); + off =3D PFN_DOWN(addr & ~HPAGE_PMD_MASK); + nr =3D min(HPAGE_PMD_NR - off, PFN_DOWN(end - addr)); + + ret =3D try_grab_folio(page_folio(page), nr, flags); if (ret) return ret; =20 @@ -758,28 +762,25 @@ static long follow_huge_pmd(struct vm_area_struct *vm= a, touch_pmd(vma, addr, pmd, flags & FOLL_WRITE); #endif /* CONFIG_TRANSPARENT_HUGEPAGE */ =20 - page +=3D (addr & ~HPAGE_PMD_MASK) >> PAGE_SHIFT; - *page_mask =3D HPAGE_PMD_NR - 1; + page +=3D off; =20 if (pages) pages[0] =3D page; =20 - return 1; + return nr; } =20 #else /* CONFIG_PGTABLE_HAS_HUGE_LEAVES */ static long follow_huge_pud(struct vm_area_struct *vma, - unsigned long addr, pud_t *pudp, - unsigned int flags, unsigned long *page_mask, - struct page **pages) + unsigned long addr, unsigned long end, pud_t *pudp, + unsigned int flags, struct page **pages) { return 0; } =20 static long follow_huge_pmd(struct vm_area_struct *vma, - unsigned long addr, pmd_t *pmd, - unsigned int flags, unsigned long *page_mask, - struct page **pages) + unsigned long addr, unsigned long end, pmd_t *pmd, + unsigned int flags, struct page **pages) { return 0; } @@ -953,9 +954,8 @@ static long follow_page_pte(struct vm_area_struct *vma, } =20 static long follow_pmd_mask(struct vm_area_struct *vma, - unsigned long address, pud_t *pudp, - unsigned int flags, unsigned long *page_mask, - struct page **pages) + unsigned long address, unsigned long end, pud_t *pudp, + unsigned int flags, struct page **pages) { pmd_t *pmd, pmdval; spinlock_t *ptl; @@ -991,7 +991,7 @@ static long follow_pmd_mask(struct vm_area_struct *vma, return pte_alloc(mm, pmd) ? -ENOMEM : follow_page_pte(vma, address, pmd, flags, pages); } - ret =3D follow_huge_pmd(vma, address, pmd, flags, page_mask, pages); + ret =3D follow_huge_pmd(vma, address, end, pmd, flags, pages); spin_unlock(ptl); =20 /* @@ -1005,9 +1005,8 @@ static long follow_pmd_mask(struct vm_area_struct *vm= a, } =20 static long follow_pud_mask(struct vm_area_struct *vma, - unsigned long address, p4d_t *p4dp, - unsigned int flags, unsigned long *page_mask, - struct page **pages) + unsigned long address, unsigned long end, p4d_t *p4dp, + unsigned int flags, struct page **pages) { pud_t *pudp, pud; spinlock_t *ptl; @@ -1020,11 +1019,13 @@ static long follow_pud_mask(struct vm_area_struct *= vma, return no_page_table(vma, flags, address); if (pud_leaf(pud)) { ptl =3D pud_lock(mm, pudp); - ret =3D follow_huge_pud(vma, address, pudp, flags, page_mask, pages); + ret =3D follow_huge_pud(vma, address, end, pudp, flags, pages); spin_unlock(ptl); /* * The ref is already held, so the page cannot go away: fill - * the array and flush caches without the lock. + * the array and flush caches without the lock. A 1 GB folio + * can be up to HPAGE_PUD_NR pages, too long to flush under a + * spinlock. */ if (ret > 0 && pages) gup_fill_pages(vma, address, pages[0], ret, pages); @@ -1035,13 +1036,12 @@ static long follow_pud_mask(struct vm_area_struct *= vma, if (unlikely(pud_bad(pud))) return no_page_table(vma, flags, address); =20 - return follow_pmd_mask(vma, address, pudp, flags, page_mask, pages); + return follow_pmd_mask(vma, address, end, pudp, flags, pages); } =20 static long follow_p4d_mask(struct vm_area_struct *vma, - unsigned long address, pgd_t *pgdp, - unsigned int flags, unsigned long *page_mask, - struct page **pages) + unsigned long address, unsigned long end, pgd_t *pgdp, + unsigned int flags, struct page **pages) { p4d_t *p4dp, p4d; =20 @@ -1052,18 +1052,18 @@ static long follow_p4d_mask(struct vm_area_struct *= vma, if (!p4d_present(p4d) || p4d_bad(p4d)) return no_page_table(vma, flags, address); =20 - return follow_pud_mask(vma, address, p4dp, flags, page_mask, pages); + return follow_pud_mask(vma, address, end, p4dp, flags, pages); } =20 /** - * follow_page_mask - look up a page descriptor from a user-virtual address + * follow_page_mask - look up pages at a user-virtual address * @vma: vm_area_struct mapping @address * @address: virtual address to look up + * @end: virtual address at which to stop batching contiguous pages * @flags: flags modifying lookup behaviour - * @page_mask: a pointer to output page_mask - * @pages: array to receive the page found, refcounted per @flags, or NULL - * to walk the page tables (e.g. to fault pages in) without - * collecting or refcounting them + * @pages: array to receive the pages, refcounted per @flags, or NULL to + * walk the page tables (e.g. to fault pages in) without collecting + * or refcounting them * * @flags can have FOLL_ flags set, defined in * @@ -1072,15 +1072,15 @@ static long follow_p4d_mask(struct vm_area_struct *= vma, * trigger a fault with FAULT_FLAG_UNSHARE set. Note that unsharing is only * relevant with FOLL_PIN and !FOLL_WRITE. * - * On output, @page_mask is set according to the size of the page. - * - * Return: 1 with @pages[0] filled in if a page was found, 0 if no mapping - * exists at @address, or a negative errno for a mapping to something not - * represented by a page descriptor (see also vm_normal_page()). + * Return: the number of contiguous pages starting at @address that were + * placed into @pages (if non-NULL), which may be fewer than the pages + * requested via @end; 0 if no mapping exists at @address; or a negative + * errno for a mapping to something not represented by a page descriptor + * (see also vm_normal_page()). */ static long follow_page_mask(struct vm_area_struct *vma, - unsigned long address, unsigned int flags, - unsigned long *page_mask, struct page **pages) + unsigned long address, unsigned long end, + unsigned int flags, struct page **pages) { pgd_t *pgd; struct mm_struct *mm =3D vma->vm_mm; @@ -1088,13 +1088,12 @@ static long follow_page_mask(struct vm_area_struct = *vma, =20 vma_pgtable_walk_begin(vma); =20 - *page_mask =3D 0; pgd =3D pgd_offset(mm, address); =20 if (pgd_none(*pgd) || unlikely(pgd_bad(*pgd))) ret =3D no_page_table(vma, flags, address); else - ret =3D follow_p4d_mask(vma, address, pgd, flags, page_mask, pages); + ret =3D follow_p4d_mask(vma, address, end, pgd, flags, pages); =20 vma_pgtable_walk_end(vma); =20 @@ -1432,7 +1431,6 @@ static long __get_user_pages(struct mm_struct *mm, { long ret =3D 0, i =3D 0; struct vm_area_struct *vma =3D NULL; - unsigned long page_mask =3D 0; =20 if (!nr_pages) return 0; @@ -1447,7 +1445,6 @@ static long __get_user_pages(struct mm_struct *mm, =20 do { struct page *page; - unsigned int page_increm; long nr; =20 /* first iteration or cross vma bound */ @@ -1502,8 +1499,8 @@ static long __get_user_pages(struct mm_struct *mm, } cond_resched(); =20 - nr =3D follow_page_mask(vma, start, gup_flags, &page_mask, - pages ? &pages[i] : NULL); + nr =3D follow_page_mask(vma, start, start + nr_pages * PAGE_SIZE, + gup_flags, pages ? &pages[i] : NULL); if (!nr || nr =3D=3D -EMLINK) { ret =3D faultin_page(vma, start, gup_flags, nr =3D=3D -EMLINK, locked); @@ -1525,56 +1522,25 @@ static long __get_user_pages(struct mm_struct *mm, * Proper page table entry exists, but no corresponding * struct page. If the caller expects **pages to be * filled in, bail out now, because that can't be done - * for this page. + * for this page. Otherwise advance by the one page + * follow_page_mask() looked at. */ if (pages) { ret =3D nr; goto out; } + nr =3D 1; } else if (nr < 0) { ret =3D nr; goto out; } =20 - page_increm =3D 1 + (~(start >> PAGE_SHIFT) & page_mask); - if (page_increm > nr_pages) - page_increm =3D nr_pages; - - /* - * This must be a large folio (and doesn't need to - * be the whole folio; it can be part of it), do - * the refcount work for all the subpages too. - * - * NOTE: here the page may not be the head page - * e.g. when start addr is not thp-size aligned. - * try_grab_folio() should have taken care of tail - * pages. - */ - if (pages && page_increm > 1) { - struct folio *folio =3D page_folio(pages[i]); - - /* - * Since we already hold refcount on the - * large folio, this should never fail. - */ - if (try_grab_folio(folio, page_increm - 1, - gup_flags)) { - /* - * Release the 1st page ref if the - * folio is problematic, fail hard. - */ - gup_put_folio(folio, 1, gup_flags); - ret =3D -EFAULT; - goto out; - } - - gup_fill_pages(vma, start + PAGE_SIZE, pages[i] + 1, - page_increm - 1, pages + i + 1); - } + /* Check that we didn't pin more pages than the caller will free. */ + VM_WARN_ON_ONCE(nr > nr_pages); =20 - i +=3D page_increm; - start +=3D page_increm * PAGE_SIZE; - nr_pages -=3D page_increm; + i +=3D nr; + start +=3D nr * PAGE_SIZE; + nr_pages -=3D nr; } while (nr_pages); out: return i ? i : ret; --=20 2.55.0 From nobody Tue Sep 29 07:39:19 2026 Received: from shelob.surriel.com (shelob.surriel.com [96.67.55.147]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 073D23C1961 for ; Tue, 11 Aug 2026 02:53:31 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=96.67.55.147 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786416814; cv=none; b=fGWupbdEfbzMnKxVoXOW6R5opCNpMAG9RGpUdeIlataR7fRzVnbj5DTIVWG90744E8CKviWa1sTTvQ2D11P5RO1VSoWftTF8dCz7M6hvbeUkcwyJWfMZ+XNkluF9Yebx+MmFqN2GxyvJaQVlIKx/qHR52WQAYMxwfRj1/ebcTDU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786416814; c=relaxed/simple; bh=WTV7/ZWfY6pOREFz5wEYPIh/TvdH40JY/J575UHbjEg=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=bcFth7kNt99swOBrqHVpymMg6cHmbNGT3HhQqbQeYsN3S1gQsWR+ybOTPHtqFpUHPKuavZwgG9NOCCmhiBZwn4Ld8+HuAYM74ficyaGhahTHdYmTOz7IodyjFdONABFQaQVGp2PdeDaUZgPkvt4UVX8ks577oHjGO22nCt0RQKI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=surriel.com; spf=pass smtp.mailfrom=surriel.com; dkim=pass (2048-bit key) header.d=surriel.com header.i=@surriel.com header.b=kEOSdFEl; arc=none smtp.client-ip=96.67.55.147 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=surriel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=surriel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=surriel.com header.i=@surriel.com header.b="kEOSdFEl" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=surriel.com ; s=mail; h=Content-Transfer-Encoding:MIME-Version:References:In-Reply-To: Message-ID:Date:Subject:Cc:To:From:Sender:Reply-To:Content-Type:Content-ID: Content-Description:Resent-Date:Resent-From:Resent-Sender:Resent-To:Resent-Cc :Resent-Message-ID:List-Id:List-Help:List-Unsubscribe:List-Subscribe: List-Post:List-Owner:List-Archive; bh=eZ/rjKRcYKxx8KUHZb4cFiPQxKG7HSt16TQtpyQm7Yk=; b=kEOSdFEluEJOxNHhfym9xqH9nc pQLLzVDfYv4xko/kenTv63WmjX7h+eRwy5uib9secxUzrKuCHnQj5QLFPURhUiQe32Az9Xe9MaUF9 ujvJ18g7t7KzgoJsYML9R8z/zanxAxZWDTSAxLlQmpNWaxA2l2IBvXdl5zfdvrlAtUxExHuNTXqo8 67zOIHV7yNVoLs3fHIlDiO3Ny5Fk8on0gl3tBRxuUtPnmgeSxcWsHHkyY7k15vOGqqsl6i8N7h5Nq hH1As3jBby+MCCUn86l+a5Ol2wmDTa/EUHpo9Ht1z+4Cms+cfwBqZUPsI05dmpbDeqYteLFo3ZNL2 dAGKEAhA==; Received: from [96.67.55.146] (helo=fangorn.surriel.com) by shelob.surriel.com with esmtpsa (TLS1.2) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.97.1) (envelope-from ) id 1wtcb6-000000007kZ-2wGb; Mon, 10 Aug 2026 22:52:12 -0400 From: Rik van Riel To: linux-kernel@vger.kernel.org Cc: kernel-team@meta.com, Rik van Riel , Andrew Morton , David Hildenbrand , Jason Gunthorpe , John Hubbard , Peter Xu , linux-mm@kvack.org Subject: [RFC PATCH v3 7/8] mm/gup: walk multiple PTEs per follow_page_pte() call Date: Mon, 10 Aug 2026 22:51:56 -0400 Message-ID: <20260811025157.1632867-8-riel@surriel.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260811025157.1632867-1-riel@surriel.com> References: <20260811025157.1632867-1-riel@surriel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" follow_page_pte() looks at one PTE per call, so __get_user_pages() calls it once per page, restarting the pgd/p4d/pud/pmd descent and retaking the PTE lock each time. Walk every PTE from @address to the page-table/VMA/@end boundary in one call instead. Adjacent pages from different folios, or plain base pages with no folio relationship at all, are covered by the same call under one lock. A failure on the first PTE is returned as before. A failure after that ends the walk, and __get_user_pages() retrying the read gets the error. Measured with mm/gup_test.c, median get time over 16 iterations on a 256 MB MADV_HUGEPAGE region in a 4 CPU VM. Folio size was confirmed through the per-size anon_fault_alloc counters, 4100 folios for 64 kB and 128 for 2 MB: gup_test -L -m 256 -n 65536 -r 16 -t before after 4 kB base pages 2753 us 1178 us (2.3x) 64 kB mTHP 2946 us 1361 us (2.2x) 2 MB THP (control) 72 us 70 us Base pages and mTHP gain about the same amount, because what is saved is the page-table descent that no longer restarts per page, not anything folio-specific. 2 MB THP does not reach follow_page_pte(), so it stays flat. Suggested-by: David Hildenbrand Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Rik van Riel --- mm/gup.c | 62 ++++++++++++++++++++++++++++++++++++++------------------ 1 file changed, 42 insertions(+), 20 deletions(-) diff --git a/mm/gup.c b/mm/gup.c index ea2bb379183e..c4233a7b8a48 100644 --- a/mm/gup.c +++ b/mm/gup.c @@ -916,38 +916,60 @@ static long follow_one_pte(struct vm_area_struct *vma= , unsigned long address, return 0; } =20 +/* + * Walk the PTEs from the start address to the end of this page table or V= MA, + * whichever comes first, and commit every page found. + * + * A failure on the first PTE is returned to the caller. A failure after t= hat + * is a short read; __get_user_pages() retrying the read will get the erro= r. + */ static long follow_page_pte(struct vm_area_struct *vma, - unsigned long address, pmd_t *pmd, unsigned int flags, - struct page **pages) + unsigned long address, unsigned long end, pmd_t *pmd, + unsigned int flags, struct page **pages) { struct mm_struct *mm =3D vma->vm_mm; bool need_no_page_table =3D false; - struct page *page; + pte_t *ptep, *orig_ptep; + unsigned long walk_end; + unsigned long nr =3D 0; spinlock_t *ptl; - pte_t *ptep, pte; - long ret; + long ret =3D 0; =20 - ptep =3D pte_offset_map_lock(mm, pmd, address, &ptl); + orig_ptep =3D ptep =3D pte_offset_map_lock(mm, pmd, address, &ptl); if (!ptep) return no_page_table(vma, flags, address); - pte =3D ptep_get(ptep); - - ret =3D follow_one_pte(vma, address, ptep, pte, flags, &page); - if (!ret && page) { - ret =3D follow_page_pte_commit(vma, address, page_folio(page), - page, pte, flags, pages); - if (!ret) - ret =3D 1; - } else if (!ret && pte_none(pte)) { + + walk_end =3D min(pmd_addr_end(address, end), vma->vm_end); + + for (; address < walk_end; address +=3D PAGE_SIZE, ptep++) { + pte_t pte =3D ptep_get(ptep); + struct page *page; + + ret =3D follow_one_pte(vma, address, ptep, pte, flags, &page); + if (!ret && page) { + ret =3D follow_page_pte_commit(vma, address, + page_folio(page), page, + pte, flags, + pages ? pages + nr : NULL); + if (!ret) { + nr++; + continue; + } + } + /* * no_page_table() may look up the page cache, so it cannot run * under the PTE lock. */ - need_no_page_table =3D true; + if (!ret && pte_none(pte)) + need_no_page_table =3D true; + break; } =20 - pte_unmap_unlock(ptep, ptl); + pte_unmap_unlock(orig_ptep, ptl); =20 + if (nr) + return nr; if (need_no_page_table) return no_page_table(vma, flags, address); return ret; @@ -969,7 +991,7 @@ static long follow_pmd_mask(struct vm_area_struct *vma, if (!pmd_present(pmdval)) return no_page_table(vma, flags, address); if (likely(!pmd_leaf(pmdval))) - return follow_page_pte(vma, address, pmd, flags, pages); + return follow_page_pte(vma, address, end, pmd, flags, pages); =20 if (pmd_protnone(pmdval) && !gup_can_follow_protnone(vma, flags)) return no_page_table(vma, flags, address); @@ -982,14 +1004,14 @@ static long follow_pmd_mask(struct vm_area_struct *v= ma, } if (unlikely(!pmd_leaf(pmdval))) { spin_unlock(ptl); - return follow_page_pte(vma, address, pmd, flags, pages); + return follow_page_pte(vma, address, end, pmd, flags, pages); } if (pmd_trans_huge(pmdval) && (flags & FOLL_SPLIT_PMD)) { spin_unlock(ptl); split_huge_pmd(vma, pmd, address); /* If pmd was left empty, stuff a page table in there quickly */ return pte_alloc(mm, pmd) ? -ENOMEM : - follow_page_pte(vma, address, pmd, flags, pages); + follow_page_pte(vma, address, end, pmd, flags, pages); } ret =3D follow_huge_pmd(vma, address, end, pmd, flags, pages); spin_unlock(ptl); --=20 2.55.0 From nobody Tue Sep 29 07:39:19 2026 Received: from shelob.surriel.com (shelob.surriel.com [96.67.55.147]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6EBE8394496 for ; Tue, 11 Aug 2026 03:29:47 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=96.67.55.147 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786418988; cv=none; b=XUZ1JCGi1wSHnYPmzsL0PHgHBAMZDxiqCGoAbIdvb7r31TO3KeIr9T+DlkWt/KxVva2VP01JcXhnJ09ccN0/n5H1F7oYVhyBsI5AoV3/twhBej8oeQkRxC7eMA3up1mRfKypr+r2R74uV1acVxSp7Vu2w4k/Vv2DtpPJeYnPFQY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786418988; c=relaxed/simple; bh=zOonfQvvV97ilnJo6WRr1j5v0xjNc4ynPsX849cxh5s=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=nKesV2FoG/VFfNdc1IVFqgIhYj3OflSlsCmKjOXVJpMnkH8qzvn1HAz5B3/0EkVJ1BLOZNRaLel8lThXX5+GWeIlZ+lRZ7XqS74LxhLWSzS1XCV3uAm8YPyn0V8i7GF9YKmu1rrbB+K7Po9bsTzR8bmWcHiDJ3kxobzHgT9aLVU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=surriel.com; spf=pass smtp.mailfrom=surriel.com; dkim=pass (2048-bit key) header.d=surriel.com header.i=@surriel.com header.b=B1KNBknq; arc=none smtp.client-ip=96.67.55.147 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=surriel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=surriel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=surriel.com header.i=@surriel.com header.b="B1KNBknq" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=surriel.com ; s=mail; h=Content-Transfer-Encoding:MIME-Version:References:In-Reply-To: Message-ID:Date:Subject:Cc:To:From:Sender:Reply-To:Content-Type:Content-ID: Content-Description:Resent-Date:Resent-From:Resent-Sender:Resent-To:Resent-Cc :Resent-Message-ID:List-Id:List-Help:List-Unsubscribe:List-Subscribe: List-Post:List-Owner:List-Archive; bh=YYSV7ewB+nMetdXPY8LjvmotcaHNCvlEgZ2HICuq4Os=; b=B1KNBknqtNr4D8k8c+8D19ORi4 bZCvNF99HN347La92JvlZe+SdQjR6lNY2OGg0oqAVXL1UsaI4ybJkbrOOHcQlIRIKQ/mvHT9XDwI6 ZvjMAPZRwfe7U94/SSU9H1Tcz1qS3eybTAxbuShMZW6PenRkiJtUvxo7+pDK3JcX3pcpjJ8obQe1X /SnGj6U68Prz7w7o22f8+Cin5/EyrbdX6YIBHgWUFb6gj1yqRJAbymxF488EsBN8I+eCjPFhd7cqh C6kWlRHCZTccIGh2GgxBdoeUqLDd7iNYzp0/zEh0bTZytzrqLMT9T+6wAGd/XKwIZqSekFMBO2ifG AVq8hzkA==; Received: from [96.67.55.146] (helo=fangorn.surriel.com) by shelob.surriel.com with esmtpsa (TLS1.2) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.97.1) (envelope-from ) id 1wtcb6-000000007kZ-321P; Mon, 10 Aug 2026 22:52:12 -0400 From: Rik van Riel To: linux-kernel@vger.kernel.org Cc: kernel-team@meta.com, Rik van Riel , Andrew Morton , David Hildenbrand , Jason Gunthorpe , John Hubbard , Peter Xu , linux-mm@kvack.org Subject: [RFC PATCH v3 8/8] mm/gup: batch contiguous same-folio PTEs into one refcount grab Date: Mon, 10 Aug 2026 22:51:57 -0400 Message-ID: <20260811025157.1632867-9-riel@surriel.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260811025157.1632867-1-riel@surriel.com> References: <20260811025157.1632867-1-riel@surriel.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" follow_page_pte() now walks a whole page table in one call, but still resolves and commits each PTE on its own, so a PTE-mapped large folio (mTHP) pays one try_grab_folio() per subpage. Add follow_pte_batch() and give follow_page_pte_commit() a run length. Each run starts with a full per-PTE resolve, then follow_pte_batch() finds how many more PTEs extend it with a pte_same() scan rather than a re-derivation of each page, and the run is committed with one refcount grab. Runs stop at folio boundaries, so plain base pages are unaffected. Gather the dirty bits from all the PTEs in a batch, in order to mark the folio dirty for FOLL_WRITE. Same benchmark as the previous change, with before again taken at the base of the series: gup_test -L -m 256 -n 65536 -r 16 -t before after 64 kB mTHP 2946 us 231 us (12.8x) 4 kB base (control) 2753 us 1198 us (2.3x) 2 MB THP (control) 72 us 71 us The 4 kB column is the previous change's 2.3x, unchanged: base pages share no folio, so there is nothing to batch. 64 kB mTHP goes from that change's 2.2x to 12.8x, a further 5.9x from committing a run with one refcount grab instead of one per subpage. Suggested-by: David Hildenbrand Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Rik van Riel --- mm/gup.c | 58 ++++++++++++++++++++++++++++++++++++++++++++++++-------- 1 file changed, 50 insertions(+), 8 deletions(-) diff --git a/mm/gup.c b/mm/gup.c index c4233a7b8a48..106806634b3c 100644 --- a/mm/gup.c +++ b/mm/gup.c @@ -833,12 +833,13 @@ static inline bool can_follow_write_pte(pte_t pte, st= ruct page *page, */ static long follow_page_pte_commit(struct vm_area_struct *vma, unsigned long address, struct folio *folio, struct page *page, - pte_t pte, unsigned int flags, struct page **pages) + pte_t pte, unsigned long nr, unsigned int flags, + struct page **pages) { long ret; =20 /* try_grab_folio() does nothing unless FOLL_GET or FOLL_PIN is set. */ - ret =3D try_grab_folio(folio, 1, flags); + ret =3D try_grab_folio(folio, nr, flags); if (unlikely(ret)) return ret; =20 @@ -850,7 +851,7 @@ static long follow_page_pte_commit(struct vm_area_struc= t *vma, if (flags & FOLL_PIN) { ret =3D arch_make_folio_accessible(folio); if (ret) { - gup_put_folio(folio, 1, flags); + gup_put_folio(folio, nr, flags); return ret; } } @@ -866,7 +867,7 @@ static long follow_page_pte_commit(struct vm_area_struc= t *vma, folio_mark_accessed(folio); } =20 - gup_fill_pages(vma, address, page, 1, pages); + gup_fill_pages(vma, address, page, nr, pages); =20 return 0; } @@ -916,6 +917,35 @@ static long follow_one_pte(struct vm_area_struct *vma,= unsigned long address, return 0; } =20 +/* + * Return how many PTEs map consecutive pages of the same folio and can be + * committed as one run. Always at least 1. + * + * The write-fault and unshare checks in follow_one_pte() are per PTE, but= a + * writable run needs no repeat: a writable anon page is exclusive. A read= -only + * run under FOLL_WRITE or FOLL_PIN does need the per-page check, so it st= ays + * one page at a time. + */ +static unsigned long follow_pte_batch(struct vm_area_struct *vma, + unsigned long address, unsigned long walk_end, + struct folio *folio, pte_t *ptep, pte_t *batch_pte, unsigned int flags) +{ + unsigned long max; + + if (!folio_test_large(folio)) + return 1; + if (!pte_write(*batch_pte) && (flags & (FOLL_WRITE | FOLL_PIN))) + return 1; + + max =3D (walk_end - address) >> PAGE_SHIFT; + if (max <=3D 1) + return 1; + + /* Merge young/dirty across batch so folio_mark_dirty sees any dirty. */ + return folio_pte_batch_flags(folio, vma, ptep, batch_pte, max, + FPB_RESPECT_WRITE | FPB_MERGE_YOUNG_DIRTY); +} + /* * Walk the PTEs from the start address to the end of this page table or V= MA, * whichever comes first, and commit every page found. @@ -947,12 +977,24 @@ static long follow_page_pte(struct vm_area_struct *vm= a, =20 ret =3D follow_one_pte(vma, address, ptep, pte, flags, &page); if (!ret && page) { - ret =3D follow_page_pte_commit(vma, address, - page_folio(page), page, - pte, flags, + struct folio *folio =3D page_folio(page); + unsigned long batch; + + pte_t batch_pte =3D pte; + + batch =3D follow_pte_batch(vma, address, walk_end, folio, + ptep, &batch_pte, flags); + ret =3D follow_page_pte_commit(vma, address, folio, page, + batch_pte, batch, flags, pages ? pages + nr : NULL); if (!ret) { - nr++; + nr +=3D batch; + /* + * The loop's own increment covers one PTE; skip + * the rest of the batch. + */ + ptep +=3D batch - 1; + address +=3D (batch - 1) * PAGE_SIZE; continue; } } --=20 2.55.0