[RFC PATCH v3 0/8] batch lookups in follow_page_mask()

Rik van Riel posted 8 patches 1 month, 2 weeks ago
mm/gup.c | 515 +++++++++++++++++++++++++++++++++----------------------
1 file changed, 308 insertions(+), 207 deletions(-)
[RFC PATCH v3 0/8] batch lookups in follow_page_mask()
Posted by Rik van Riel 1 month, 2 weeks ago
follow_page_mask() walks the page tables one page at a time, even when the
caller asked for a whole run of contiguous pages. Every page of a large folio
re-enters the pmd/pud/pte walk and re-takes the page table lock.

This series changes follow_page_mask() to return a page count and a fill an
array of pages, instead of a single struct page, so a walker can hand back
more than one page per call.

Patches 1 to 4 are preparation, no functional change:

  1: move __get_user_pages()'s open-coded pages[] fill and cache flush into a
     gup_fill_pages() helper, which the rest of the series reuses.
  2: convert the follow_page_mask()/follow_p4d_mask()/follow_pud_mask()/
     follow_pmd_mask()/follow_page_pte() call chain to return a long instead of
     a struct page pointer or ERR_PTR(). Every path still handles one page.
  3: split the "commit to a resolved page" tail of follow_page_pte() into
     follow_page_pte_commit().
  4: split the "work out which page this PTE maps" half of follow_page_pte()
     into follow_one_pte(), leaving one unlock and one exit.

Patch 5 has the huge page paths store the page and leave the array fill to
follow_pud_mask()/follow_pmd_mask() after they unlock, so the cache flushes
happen outside the pud/pmd critical section.

Patch 6 has follow_huge_pud()/follow_huge_pmd() report the huge page's real
subpage count instead of a separate *page_mask output, and retires *page_mask
and __get_user_pages()'s try_grab_folio() call and subpage loop.

Patch 7 walks every PTE in a page table in one follow_page_pte() call instead
of one per page.

Patch 8 adds follow_pte_batch() so a contiguous same-folio run is committed
with one refcount grab. This is the only patch whose benefit depends on folio
size; patch 7 alone covers plain base pages.

Patches 7 and 8 carry their own benchmark tables, both measured against the
base of the series, so the split between the two mechanisms is visible: the
single-call walk is worth 2.3x on base pages and 2.2x on 64 kB mTHP, and
refcount batching adds a further 5.9x on the mTHP case.

v3:
 - split up the series into 8 much smaller patches (David & Lorenzo)
 - shorten changelogs where they were too long (Lorenzo)
 - fix FOLL_WRITE folio dirtying by gathering dirty bits from all PTEs

Link: https://lore.kernel.org/r/20260730035350.1fc95dd8@fangorn/ [RFC]
Link: https://lore.kernel.org/r/20260801031540.2742891-1-riel@surriel.com/ [RFC v2]
Suggested-by: David Hildenbrand <david@kernel.org>

Rik van Riel (8):
  mm/gup: break out gup_fill_pages() helper
  mm/gup: convert follow_page_mask() to return a long
  mm/gup: split follow_page_pte_commit() out of follow_page_pte()
  mm/gup: break out follow_one_pte() helper
  mm/gup: fill the pages array outside the pud/pmd lock
  mm/gup: return a huge page's full count from follow_page_mask()
  mm/gup: walk multiple PTEs per follow_page_pte() call
  mm/gup: batch contiguous same-folio PTEs into one refcount grab

 mm/gup.c | 515 +++++++++++++++++++++++++++++++++----------------------
 1 file changed, 308 insertions(+), 207 deletions(-)

-- 
2.55.0
Re: [RFC PATCH v3 0/8] batch lookups in follow_page_mask()
Posted by Christoph Hellwig 1 month ago
On Mon, Aug 10, 2026 at 10:51:49PM -0400, Rik van Riel wrote:
> follow_page_mask() walks the page tables one page at a time, even when the
> caller asked for a whole run of contiguous pages. Every page of a large folio
> re-enters the pmd/pud/pte walk and re-takes the page table lock.
> 
> This series changes follow_page_mask() to return a page count and a fill an
> array of pages, instead of a single struct page, so a walker can hand back
> more than one page per call.

And what many including the most performance critical callers want
instead is really a single bio_vec.  Maybe we can go the extra step for
that, as it would reduce the number of calls into gup significantly.
Re: [RFC PATCH v3 0/8] batch lookups in follow_page_mask()
Posted by Rik van Riel 1 month ago
On Tue, 2026-08-25 at 00:52 -0700, Christoph Hellwig wrote:
> On Mon, Aug 10, 2026 at 10:51:49PM -0400, Rik van Riel wrote:
> > follow_page_mask() walks the page tables one page at a time, even
> > when the
> > caller asked for a whole run of contiguous pages. Every page of a
> > large folio
> > re-enters the pmd/pud/pte walk and re-takes the page table lock.
> > 
> > This series changes follow_page_mask() to return a page count and a
> > fill an
> > array of pages, instead of a single struct page, so a walker can
> > hand back
> > more than one page per call.
> 
> And what many including the most performance critical callers want
> instead is really a single bio_vec.  Maybe we can go the extra step
> for
> that, as it would reduce the number of calls into gup significantly.
> 
Which code are you referring to here?

blk_rq_map_user() seems to end up calling pin_user_pages_fast(),
via iov_iter_extract_user_pages, and it directly fills in the
bio_vec's pages array.

Is there another performance critical path that goes
through the slower get_user_pages() path?

-- 
All Rights Reversed.
Re: [RFC PATCH v3 0/8] batch lookups in follow_page_mask()
Posted by Christoph Hellwig 1 month ago
On Tue, Aug 25, 2026 at 09:37:09AM -0400, Rik van Riel wrote:
> > And what many including the most performance critical callers want
> > instead is really a single bio_vec.  Maybe we can go the extra step
> > for
> > that, as it would reduce the number of calls into gup significantly.
> > 
> Which code are you referring to here?
> 
> blk_rq_map_user() seems to end up calling pin_user_pages_fast(),
> via iov_iter_extract_user_pages, and it directly fills in the
> bio_vec's pages array.
> 
> Is there another performance critical path that goes
> through the slower get_user_pages() path?

The other callers of iov_iter_extract_bvecs matter more, but this is
the main user.

And iov_iter_extract_bvecs right now is very inefficient when used
on larger folios, as each call to iov_iter_extract_bvecs and thus
iov_iter_extract_pages and iov_iter_extract_user_pages can only
fill in up to 256 pages.  On x86 that is a single PMD mapped
folio.  Passing down the bio_vec array means that for PMD-mapped
ranges we reduce the calls into pin_user_pages_fast by a factor
of 256.
Re: [RFC PATCH v3 0/8] batch lookups in follow_page_mask()
Posted by Rik van Riel 1 month ago
On Tue, 2026-08-25 at 21:39 -0700, Christoph Hellwig wrote:
> On Tue, Aug 25, 2026 at 09:37:09AM -0400, Rik van Riel wrote:
> > > And what many including the most performance critical callers
> > > want
> > > instead is really a single bio_vec.  Maybe we can go the extra
> > > step
> > > for
> > > that, as it would reduce the number of calls into gup
> > > significantly.
> > > 
> > Which code are you referring to here?
> > 
> > blk_rq_map_user() seems to end up calling pin_user_pages_fast(),
> > via iov_iter_extract_user_pages, and it directly fills in the
> > bio_vec's pages array.
> > 
> > Is there another performance critical path that goes
> > through the slower get_user_pages() path?
> 
> The other callers of iov_iter_extract_bvecs matter more, but this is
> the main user.
> 
> And iov_iter_extract_bvecs right now is very inefficient when used
> on larger folios,

I think it would be possible to create a path
into follow_page_mask() that passes a bio_vec,
and where we fill in the bio_vec fields the
same way bvec_set_page() does.

Then iov_iter_extract_bvecs would no longer
need to iterate over the pages returned by
get_user_pages_fast.

That seems like a follow-up series though,
and we need to think carefully about whether
we would also want to unify part of the 
gup_fast and gup code.

Currently the gup code depends on locking
to keep page tables from going away, while
the gup_fast code depends on disabling irqs.

-- 
All Rights Reversed.
Re: [RFC PATCH v3 0/8] batch lookups in follow_page_mask()
Posted by Christoph Hellwig 1 month ago
On Tue, Aug 25, 2026 at 09:39:14PM -0700, Christoph Hellwig wrote:
> The other callers of iov_iter_extract_bvecs matter more, but this is
> the main user.
> 
> And iov_iter_extract_bvecs right now is very inefficient when used
> on larger folios, as each call to iov_iter_extract_bvecs and thus
> iov_iter_extract_pages and iov_iter_extract_user_pages can only
> fill in up to 256 pages.  On x86 that is a single PMD mapped
> folio.  Passing down the bio_vec array means that for PMD-mapped
> ranges we reduce the calls into pin_user_pages_fast by a factor
> of 256.

Also in addition to reducing the calls to pin_user_pages_fast, it also
removes a fair amount of work to recalculate the bio_len in the
get_contig_folio_len loop in iov_iter_extract_bvecs, which has shown up
in profiles (although only marginally so).
Re: [RFC PATCH v3 0/8] batch lookups in follow_page_mask()
Posted by David Hildenbrand (Arm) 1 month ago
On 8/26/26 09:35, Christoph Hellwig wrote:
> On Tue, Aug 25, 2026 at 09:39:14PM -0700, Christoph Hellwig wrote:
>> The other callers of iov_iter_extract_bvecs matter more, but this is
>> the main user.
>>
>> And iov_iter_extract_bvecs right now is very inefficient when used
>> on larger folios, as each call to iov_iter_extract_bvecs and thus
>> iov_iter_extract_pages and iov_iter_extract_user_pages can only
>> fill in up to 256 pages.  On x86 that is a single PMD mapped
>> folio.  Passing down the bio_vec array means that for PMD-mapped
>> ranges we reduce the calls into pin_user_pages_fast by a factor
>> of 256.
> 
> Also in addition to reducing the calls to pin_user_pages_fast, it also
> removes a fair amount of work to recalculate the bio_len in the
> get_contig_folio_len loop in iov_iter_extract_bvecs, which has shown up
> in profiles (although only marginally so).

We could think of a mechanism that fills a given structure in a better way, and
I think we recently discussed returning folio ranges, and might to that through
something like phyr (or however that is called :) ).

But that's really not something is would mix into this work here.

-- 
Cheers,

David
Re: [RFC PATCH v3 0/8] batch lookups in follow_page_mask()
Posted by David Hildenbrand (Arm) 1 month, 2 weeks ago
On 8/11/26 04:51, Rik van Riel wrote:
> follow_page_mask() walks the page tables one page at a time, even when the
> caller asked for a whole run of contiguous pages. Every page of a large folio
> re-enters the pmd/pud/pte walk and re-takes the page table lock.
> 
> This series changes follow_page_mask() to return a page count and a fill an
> array of pages, instead of a single struct page, so a walker can hand back
> more than one page per call.
> 
> Patches 1 to 4 are preparation, no functional change:
> 
>   1: move __get_user_pages()'s open-coded pages[] fill and cache flush into a
>      gup_fill_pages() helper, which the rest of the series reuses.
>   2: convert the follow_page_mask()/follow_p4d_mask()/follow_pud_mask()/
>      follow_pmd_mask()/follow_page_pte() call chain to return a long instead of
>      a struct page pointer or ERR_PTR(). Every path still handles one page.
>   3: split the "commit to a resolved page" tail of follow_page_pte() into
>      follow_page_pte_commit().
>   4: split the "work out which page this PTE maps" half of follow_page_pte()
>      into follow_one_pte(), leaving one unlock and one exit.
> 
> Patch 5 has the huge page paths store the page and leave the array fill to
> follow_pud_mask()/follow_pmd_mask() after they unlock, so the cache flushes
> happen outside the pud/pmd critical section.
> 
> Patch 6 has follow_huge_pud()/follow_huge_pmd() report the huge page's real
> subpage count instead of a separate *page_mask output, and retires *page_mask
> and __get_user_pages()'s try_grab_folio() call and subpage loop.
> 
> Patch 7 walks every PTE in a page table in one follow_page_pte() call instead
> of one per page.
> 
> Patch 8 adds follow_pte_batch() so a contiguous same-folio run is committed
> with one refcount grab. This is the only patch whose benefit depends on folio
> size; patch 7 alone covers plain base pages.
> 
> Patches 7 and 8 carry their own benchmark tables, both measured against the
> base of the series, so the split between the two mechanisms is visible: the
> single-call walk is worth 2.3x on base pages and 2.2x on 64 kB mTHP, and
> refcount batching adds a further 5.9x on the mTHP case.
> 
> v3:
>  - split up the series into 8 much smaller patches (David & Lorenzo)
>  - shorten changelogs where they were too long (Lorenzo)
>  - fix FOLL_WRITE folio dirtying by gathering dirty bits from all PTEs

I'll try getting to this soon. As I raised previously, the whole follow_page_*
terminology is just stale, and likely we should just not add new functions that
use this terminology.

I.e., follow_page_pte_commit()

-- 
Cheers,

David