mm/gup.c | 515 +++++++++++++++++++++++++++++++++---------------------- 1 file changed, 308 insertions(+), 207 deletions(-)
follow_page_mask() walks the page tables one page at a time, even when the
caller asked for a whole run of contiguous pages. Every page of a large folio
re-enters the pmd/pud/pte walk and re-takes the page table lock.
This series changes follow_page_mask() to return a page count and a fill an
array of pages, instead of a single struct page, so a walker can hand back
more than one page per call.
Patches 1 to 4 are preparation, no functional change:
1: move __get_user_pages()'s open-coded pages[] fill and cache flush into a
gup_fill_pages() helper, which the rest of the series reuses.
2: convert the follow_page_mask()/follow_p4d_mask()/follow_pud_mask()/
follow_pmd_mask()/follow_page_pte() call chain to return a long instead of
a struct page pointer or ERR_PTR(). Every path still handles one page.
3: split the "commit to a resolved page" tail of follow_page_pte() into
follow_page_pte_commit().
4: split the "work out which page this PTE maps" half of follow_page_pte()
into follow_one_pte(), leaving one unlock and one exit.
Patch 5 has the huge page paths store the page and leave the array fill to
follow_pud_mask()/follow_pmd_mask() after they unlock, so the cache flushes
happen outside the pud/pmd critical section.
Patch 6 has follow_huge_pud()/follow_huge_pmd() report the huge page's real
subpage count instead of a separate *page_mask output, and retires *page_mask
and __get_user_pages()'s try_grab_folio() call and subpage loop.
Patch 7 walks every PTE in a page table in one follow_page_pte() call instead
of one per page.
Patch 8 adds follow_pte_batch() so a contiguous same-folio run is committed
with one refcount grab. This is the only patch whose benefit depends on folio
size; patch 7 alone covers plain base pages.
Patches 7 and 8 carry their own benchmark tables, both measured against the
base of the series, so the split between the two mechanisms is visible: the
single-call walk is worth 2.3x on base pages and 2.2x on 64 kB mTHP, and
refcount batching adds a further 5.9x on the mTHP case.
v3:
- split up the series into 8 much smaller patches (David & Lorenzo)
- shorten changelogs where they were too long (Lorenzo)
- fix FOLL_WRITE folio dirtying by gathering dirty bits from all PTEs
Link: https://lore.kernel.org/r/20260730035350.1fc95dd8@fangorn/ [RFC]
Link: https://lore.kernel.org/r/20260801031540.2742891-1-riel@surriel.com/ [RFC v2]
Suggested-by: David Hildenbrand <david@kernel.org>
Rik van Riel (8):
mm/gup: break out gup_fill_pages() helper
mm/gup: convert follow_page_mask() to return a long
mm/gup: split follow_page_pte_commit() out of follow_page_pte()
mm/gup: break out follow_one_pte() helper
mm/gup: fill the pages array outside the pud/pmd lock
mm/gup: return a huge page's full count from follow_page_mask()
mm/gup: walk multiple PTEs per follow_page_pte() call
mm/gup: batch contiguous same-folio PTEs into one refcount grab
mm/gup.c | 515 +++++++++++++++++++++++++++++++++----------------------
1 file changed, 308 insertions(+), 207 deletions(-)
--
2.55.0
On Mon, Aug 10, 2026 at 10:51:49PM -0400, Rik van Riel wrote: > follow_page_mask() walks the page tables one page at a time, even when the > caller asked for a whole run of contiguous pages. Every page of a large folio > re-enters the pmd/pud/pte walk and re-takes the page table lock. > > This series changes follow_page_mask() to return a page count and a fill an > array of pages, instead of a single struct page, so a walker can hand back > more than one page per call. And what many including the most performance critical callers want instead is really a single bio_vec. Maybe we can go the extra step for that, as it would reduce the number of calls into gup significantly.
On Tue, 2026-08-25 at 00:52 -0700, Christoph Hellwig wrote: > On Mon, Aug 10, 2026 at 10:51:49PM -0400, Rik van Riel wrote: > > follow_page_mask() walks the page tables one page at a time, even > > when the > > caller asked for a whole run of contiguous pages. Every page of a > > large folio > > re-enters the pmd/pud/pte walk and re-takes the page table lock. > > > > This series changes follow_page_mask() to return a page count and a > > fill an > > array of pages, instead of a single struct page, so a walker can > > hand back > > more than one page per call. > > And what many including the most performance critical callers want > instead is really a single bio_vec. Maybe we can go the extra step > for > that, as it would reduce the number of calls into gup significantly. > Which code are you referring to here? blk_rq_map_user() seems to end up calling pin_user_pages_fast(), via iov_iter_extract_user_pages, and it directly fills in the bio_vec's pages array. Is there another performance critical path that goes through the slower get_user_pages() path? -- All Rights Reversed.
On Tue, Aug 25, 2026 at 09:37:09AM -0400, Rik van Riel wrote: > > And what many including the most performance critical callers want > > instead is really a single bio_vec. Maybe we can go the extra step > > for > > that, as it would reduce the number of calls into gup significantly. > > > Which code are you referring to here? > > blk_rq_map_user() seems to end up calling pin_user_pages_fast(), > via iov_iter_extract_user_pages, and it directly fills in the > bio_vec's pages array. > > Is there another performance critical path that goes > through the slower get_user_pages() path? The other callers of iov_iter_extract_bvecs matter more, but this is the main user. And iov_iter_extract_bvecs right now is very inefficient when used on larger folios, as each call to iov_iter_extract_bvecs and thus iov_iter_extract_pages and iov_iter_extract_user_pages can only fill in up to 256 pages. On x86 that is a single PMD mapped folio. Passing down the bio_vec array means that for PMD-mapped ranges we reduce the calls into pin_user_pages_fast by a factor of 256.
On Tue, 2026-08-25 at 21:39 -0700, Christoph Hellwig wrote: > On Tue, Aug 25, 2026 at 09:37:09AM -0400, Rik van Riel wrote: > > > And what many including the most performance critical callers > > > want > > > instead is really a single bio_vec. Maybe we can go the extra > > > step > > > for > > > that, as it would reduce the number of calls into gup > > > significantly. > > > > > Which code are you referring to here? > > > > blk_rq_map_user() seems to end up calling pin_user_pages_fast(), > > via iov_iter_extract_user_pages, and it directly fills in the > > bio_vec's pages array. > > > > Is there another performance critical path that goes > > through the slower get_user_pages() path? > > The other callers of iov_iter_extract_bvecs matter more, but this is > the main user. > > And iov_iter_extract_bvecs right now is very inefficient when used > on larger folios, I think it would be possible to create a path into follow_page_mask() that passes a bio_vec, and where we fill in the bio_vec fields the same way bvec_set_page() does. Then iov_iter_extract_bvecs would no longer need to iterate over the pages returned by get_user_pages_fast. That seems like a follow-up series though, and we need to think carefully about whether we would also want to unify part of the gup_fast and gup code. Currently the gup code depends on locking to keep page tables from going away, while the gup_fast code depends on disabling irqs. -- All Rights Reversed.
On Tue, Aug 25, 2026 at 09:39:14PM -0700, Christoph Hellwig wrote: > The other callers of iov_iter_extract_bvecs matter more, but this is > the main user. > > And iov_iter_extract_bvecs right now is very inefficient when used > on larger folios, as each call to iov_iter_extract_bvecs and thus > iov_iter_extract_pages and iov_iter_extract_user_pages can only > fill in up to 256 pages. On x86 that is a single PMD mapped > folio. Passing down the bio_vec array means that for PMD-mapped > ranges we reduce the calls into pin_user_pages_fast by a factor > of 256. Also in addition to reducing the calls to pin_user_pages_fast, it also removes a fair amount of work to recalculate the bio_len in the get_contig_folio_len loop in iov_iter_extract_bvecs, which has shown up in profiles (although only marginally so).
On 8/26/26 09:35, Christoph Hellwig wrote: > On Tue, Aug 25, 2026 at 09:39:14PM -0700, Christoph Hellwig wrote: >> The other callers of iov_iter_extract_bvecs matter more, but this is >> the main user. >> >> And iov_iter_extract_bvecs right now is very inefficient when used >> on larger folios, as each call to iov_iter_extract_bvecs and thus >> iov_iter_extract_pages and iov_iter_extract_user_pages can only >> fill in up to 256 pages. On x86 that is a single PMD mapped >> folio. Passing down the bio_vec array means that for PMD-mapped >> ranges we reduce the calls into pin_user_pages_fast by a factor >> of 256. > > Also in addition to reducing the calls to pin_user_pages_fast, it also > removes a fair amount of work to recalculate the bio_len in the > get_contig_folio_len loop in iov_iter_extract_bvecs, which has shown up > in profiles (although only marginally so). We could think of a mechanism that fills a given structure in a better way, and I think we recently discussed returning folio ranges, and might to that through something like phyr (or however that is called :) ). But that's really not something is would mix into this work here. -- Cheers, David
On 8/11/26 04:51, Rik van Riel wrote: > follow_page_mask() walks the page tables one page at a time, even when the > caller asked for a whole run of contiguous pages. Every page of a large folio > re-enters the pmd/pud/pte walk and re-takes the page table lock. > > This series changes follow_page_mask() to return a page count and a fill an > array of pages, instead of a single struct page, so a walker can hand back > more than one page per call. > > Patches 1 to 4 are preparation, no functional change: > > 1: move __get_user_pages()'s open-coded pages[] fill and cache flush into a > gup_fill_pages() helper, which the rest of the series reuses. > 2: convert the follow_page_mask()/follow_p4d_mask()/follow_pud_mask()/ > follow_pmd_mask()/follow_page_pte() call chain to return a long instead of > a struct page pointer or ERR_PTR(). Every path still handles one page. > 3: split the "commit to a resolved page" tail of follow_page_pte() into > follow_page_pte_commit(). > 4: split the "work out which page this PTE maps" half of follow_page_pte() > into follow_one_pte(), leaving one unlock and one exit. > > Patch 5 has the huge page paths store the page and leave the array fill to > follow_pud_mask()/follow_pmd_mask() after they unlock, so the cache flushes > happen outside the pud/pmd critical section. > > Patch 6 has follow_huge_pud()/follow_huge_pmd() report the huge page's real > subpage count instead of a separate *page_mask output, and retires *page_mask > and __get_user_pages()'s try_grab_folio() call and subpage loop. > > Patch 7 walks every PTE in a page table in one follow_page_pte() call instead > of one per page. > > Patch 8 adds follow_pte_batch() so a contiguous same-folio run is committed > with one refcount grab. This is the only patch whose benefit depends on folio > size; patch 7 alone covers plain base pages. > > Patches 7 and 8 carry their own benchmark tables, both measured against the > base of the series, so the split between the two mechanisms is visible: the > single-call walk is worth 2.3x on base pages and 2.2x on 64 kB mTHP, and > refcount batching adds a further 5.9x on the mTHP case. > > v3: > - split up the series into 8 much smaller patches (David & Lorenzo) > - shorten changelogs where they were too long (Lorenzo) > - fix FOLL_WRITE folio dirtying by gathering dirty bits from all PTEs I'll try getting to this soon. As I raised previously, the whole follow_page_* terminology is just stale, and likely we should just not add new functions that use this terminology. I.e., follow_page_pte_commit() -- Cheers, David
© 2016 - 2026 Red Hat, Inc.