MAINTAINERS | 1 + include/linux/huge_mm.h | 9 - mm/collapse.h | 153 ++++++++++++ mm/khugepaged.c | 527 +++++++++++++++------------------------- mm/madvise.c | 169 ++++++++++++- 5 files changed, 521 insertions(+), 338 deletions(-) create mode 100644 mm/collapse.h
From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
[ This is the first of the cleanups I said I would front-load ]
There is no line between the collapse engine and the callers that ask for
a collapse. khugepaged.c holds both, and they reach into each other.
- Sixteen tests through the collapse path read cc->is_khugepaged to work
out what they are allowed to do, when every one of those decisions was
made by the caller before it asked.
- collapse_single_pmd() does both halves of a collapse behind one call and
drops mmap_lock somewhere in the middle. Which of its paths dropped it
is not something a caller can see, so it hands back a bool and the
caller keeps track.
- MADV_COLLAPSE's implementation -- the walk over the user's range, the
per-PMD loop, the errno translation -- sits in khugepaged.c, which is
the daemon's file.
So: draw the line. State what a caller allows in a policy, split the call
in two with the lock as the boundary, and move the syscall to madvise.c.
What the engine offers is then four calls, with the lock state written
down against each, and a policy the caller fills for itself:
collapse_control_init(cc) once, before the first table
collapse_policy_*(&cc->policy) what this caller allows
collapse_scan_pmd(vma, addr, ...) per table, under mmap_lock
collapse_run_pmd(mm, addr, cc) when a scan found work, no mmap_lock
collapse_control_release(cc) once, when done
The engine stays in khugepaged.c for now; what changes is that it has an
interface, and that neither half has to ask about the other. madvise.c
gains the operation it should have had all along.
Changes since v1
================
https://lore.kernel.org/all/cover.1788533997.git.kas@kernel.org/
- Rebased onto mm-new with Vernon Yang's tracepoint fixes in it. Patch 8
no longer merges the two calls to each scan tracepoint, since the base
already has one; its changelog now says what the status field reports.
- Patch 3: nr_occupied_ptes is nr_eligible_ptes, and the mthp_collapse()
comment counts eligible PTEs too (Zi, Baolin).
- Patch 4: no comments on the two constants (Baolin).
- Patch 5: one line per policy field (Baolin).
- Patch 8: the file side is split like the anonymous one (Zi).
collapse_scan_file() runs under mmap_lock in the scan and only judges;
collapse_file() runs in the run. See Behaviour below.
- Reviewed-by from Zi Yan and Baolin Wang on 1-4, 6 and 7.
Patches
=======
The first three stand alone and can be taken separately:
1 drop the mmgrab() MADV_COLLAPSE has held since 7d8faaf15545
2 count collapses in khugepaged's own walk, where the daemon's
bookkeeping belongs
3 rename cc->mthp_present_ptes to eligible_ptes, which is what a set
bit means
Then the interface, in order:
4 add collapse.h, and move enum scan_result, struct collapse_control
and the two constants into it
5 struct collapse_policy, filled by the caller; the is_khugepaged
tests become field reads, and the flag goes
6 drop collapse_possible(), a wrapper that only turns a mask into a bool
7 collapse_control_init_scan() is a per-table reset, so name it
collapse_scan_reset()
8 give the scan and the collapse a function each: collapse_scan_pmd()
and collapse_run_pmd()
9 open-code the entry point that joined them, so each caller owns the
lock across the boundary and the bool goes
10 work out the orders a VMA allows once per VMA, not once per table
11 declare the four calls in collapse.h, with the lock rules
12 MADV_COLLAPSE moves to madvise.c
Behaviour
=========
No functional change is intended. Nothing here changes which tables get
collapsed, into what, or what MADV_COLLAPSE returns. The tracepoints are
the one place a change can be seen from outside; the rest is where work
happens, not what it does.
- Patch 8: mm_khugepaged_scan_pmd and mm_khugepaged_scan_file fire before
the collapse rather than after it. For an accepted table their status
field reads SCAN_SUCCEED, where it used to carry what the collapse made
of the table; that is now for mm_collapse_huge_page and
mm_khugepaged_collapse_file to report.
Three things move that a reader should not have to find in the diff:
- Patch 5: khugepaged fills its policy once per scan pass, so the
max_ptes_* limits and the defrag setting behind the allocation mask are
sampled once per pass rather than once per table. A knob written
mid-pass takes effect on the next pass instead of the next table.
- Patch 8: the file scan runs under mmap_lock, where before the lock was
given up first. A file table the scan refuses no longer ends
khugepaged's pass over that mm; only a table it goes on to collapse
does. A PMD folio the scan finds already in the page cache sends the
run straight to retracting the PTE table, and the writeback retry
re-runs collapse_file() alone.
- Patch 10: which orders a table is scanned for is sampled once per VMA
rather than once per table. It cannot widen what a collapse does; the
order is tested again under the lock the collapse retakes.
For an anonymous table, and for a file table that gets collapsed, the lock
is given up and taken again at exactly the points it was before; the only
difference is that the caller is the one doing it.
selftests/mm khugepaged passes on x86-64 with a KASAN, lockdep and
DEBUG_VM config, and every patch builds, CONFIG_TRANSPARENT_HUGEPAGE=n
included.
Kiryl Shutsemau (Meta) (12):
mm/khugepaged: drop redundant mm_struct pin in madvise_collapse()
mm/khugepaged: count collapses where khugepaged makes them
mm/khugepaged: rename mthp_present_ptes bitmap to eligible_ptes
mm/collapse: add collapse.h for the collapse interface
mm/collapse: state what a collapse may do in the policy
mm/collapse: drop the collapse_possible() wrapper
mm/collapse: name the per-table scan reset for what it resets
mm/collapse: separate scanning a PTE table from collapsing it
mm/collapse: open-code collapse_single_pmd() in its two callers
mm/collapse: work out the orders a VMA allows once per VMA
mm/collapse: declare the collapse interface in collapse.h
mm/collapse: implement MADV_COLLAPSE in madvise.c
MAINTAINERS | 1 +
include/linux/huge_mm.h | 9 -
mm/collapse.h | 153 ++++++++++++
mm/khugepaged.c | 527 +++++++++++++++-------------------------
mm/madvise.c | 169 ++++++++++++-
5 files changed, 521 insertions(+), 338 deletions(-)
create mode 100644 mm/collapse.h
base-commit: cf558a250cf4475a8936902b8978fbb6c61016f8
--
2.54.0
On 9/10/26 14:02, Kiryl Shutsemau wrote:
> From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
>
> [ This is the first of the cleanups I said I would front-load ]
>
> There is no line between the collapse engine and the callers that ask for
> a collapse. khugepaged.c holds both, and they reach into each other.
>
> - Sixteen tests through the collapse path read cc->is_khugepaged to work
> out what they are allowed to do, when every one of those decisions was
> made by the caller before it asked.
>
> - collapse_single_pmd() does both halves of a collapse behind one call and
> drops mmap_lock somewhere in the middle. Which of its paths dropped it
> is not something a caller can see, so it hands back a bool and the
> caller keeps track.
>
> - MADV_COLLAPSE's implementation -- the walk over the user's range, the
> per-PMD loop, the errno translation -- sits in khugepaged.c, which is
> the daemon's file.
>
> So: draw the line. State what a caller allows in a policy, split the call
> in two with the lock as the boundary, and move the syscall to madvise.c.
> What the engine offers is then four calls, with the lock state written
> down against each, and a policy the caller fills for itself:
>
> collapse_control_init(cc) once, before the first table
> collapse_policy_*(&cc->policy) what this caller allows
> collapse_scan_pmd(vma, addr, ...) per table, under mmap_lock
> collapse_run_pmd(mm, addr, cc) when a scan found work, no mmap_lock
> collapse_control_release(cc) once, when done
>
> The engine stays in khugepaged.c for now; what changes is that it has an
> interface, and that neither half has to ask about the other. madvise.c
> gains the operation it should have had all along.
>
> Changes since v1
> ================
>
> https://lore.kernel.org/all/cover.1788533997.git.kas@kernel.org/
>
> - Rebased onto mm-new with Vernon Yang's tracepoint fixes in it. Patch 8
> no longer merges the two calls to each scan tracepoint, since the base
> already has one; its changelog now says what the status field reports.
>
> - Patch 3: nr_occupied_ptes is nr_eligible_ptes, and the mthp_collapse()
> comment counts eligible PTEs too (Zi, Baolin).
>
> - Patch 4: no comments on the two constants (Baolin).
>
> - Patch 5: one line per policy field (Baolin).
>
> - Patch 8: the file side is split like the anonymous one (Zi).
> collapse_scan_file() runs under mmap_lock in the scan and only judges;
> collapse_file() runs in the run. See Behaviour below.
>
> - Reviewed-by from Zi Yan and Baolin Wang on 1-4, 6 and 7.
>
I'll hopefully get too look at this soon (after digging through older stuff in
my queue).
Skimming over some patches, a note that we should not be undoing recent
cleanups without a very good reason.
E.g.,:
commit a155d945b73c5b0668e898df5495afe45bb261cd
Author: Nico Pache <nico.pache@linux.dev>
Date: Wed Mar 25 05:40:22 2026 -0600
mm/khugepaged: unify khugepaged and madv_collapse with collapse_single_pmd()
The khugepaged daemon and madvise_collapse have two different
implementations that do almost the same thing. Create collapse_single_pmd
to increase code reuse and create an entry point to these two users.
Refactor madvise_collapse and collapse_scan_mm_slot to use the new
collapse_single_pmd function. To help reduce confusion around the
mmap_locked variable, we rename mmap_locked to lock_dropped in the
collapse_scan_mm_slot() function, and remove the redundant mmap_locked in
madvise_collapse(); this further unifies the code readiblity. the
SCAN_PTE_MAPPED_HUGEPAGE enum is no longer reachable in the
madvise_collapse() function, so we drop it from the list of "continuing"
enums.
This introduces a minor behavioral change that is most likely an
undiscovered bug. The current implementation of khugepaged tests
collapse_test_exit_or_disable() before calling collapse_pte_mapped_thp,
but we weren't doing it in the madvise_collapse case. By unifying these
two callers madvise_collapse now also performs this check. We also modify
the return value to be SCAN_ANY_PROCESS which properly indicates that this
process is no longer valid to operate on.
--
Cheers,
David
On Fri, Sep 11, 2026 at 05:06:58PM +0200, David Hildenbrand (Arm) wrote:
> On 9/10/26 14:02, Kiryl Shutsemau wrote:
> > From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
> >
> > [ This is the first of the cleanups I said I would front-load ]
> >
> > There is no line between the collapse engine and the callers that ask for
> > a collapse. khugepaged.c holds both, and they reach into each other.
> >
> > - Sixteen tests through the collapse path read cc->is_khugepaged to work
> > out what they are allowed to do, when every one of those decisions was
> > made by the caller before it asked.
> >
> > - collapse_single_pmd() does both halves of a collapse behind one call and
> > drops mmap_lock somewhere in the middle. Which of its paths dropped it
> > is not something a caller can see, so it hands back a bool and the
> > caller keeps track.
> >
> > - MADV_COLLAPSE's implementation -- the walk over the user's range, the
> > per-PMD loop, the errno translation -- sits in khugepaged.c, which is
> > the daemon's file.
> >
> > So: draw the line. State what a caller allows in a policy, split the call
> > in two with the lock as the boundary, and move the syscall to madvise.c.
> > What the engine offers is then four calls, with the lock state written
> > down against each, and a policy the caller fills for itself:
> >
> > collapse_control_init(cc) once, before the first table
> > collapse_policy_*(&cc->policy) what this caller allows
> > collapse_scan_pmd(vma, addr, ...) per table, under mmap_lock
> > collapse_run_pmd(mm, addr, cc) when a scan found work, no mmap_lock
> > collapse_control_release(cc) once, when done
> >
> > The engine stays in khugepaged.c for now; what changes is that it has an
> > interface, and that neither half has to ask about the other. madvise.c
> > gains the operation it should have had all along.
> >
> > Changes since v1
> > ================
> >
> > https://lore.kernel.org/all/cover.1788533997.git.kas@kernel.org/
> >
> > - Rebased onto mm-new with Vernon Yang's tracepoint fixes in it. Patch 8
> > no longer merges the two calls to each scan tracepoint, since the base
> > already has one; its changelog now says what the status field reports.
> >
> > - Patch 3: nr_occupied_ptes is nr_eligible_ptes, and the mthp_collapse()
> > comment counts eligible PTEs too (Zi, Baolin).
> >
> > - Patch 4: no comments on the two constants (Baolin).
> >
> > - Patch 5: one line per policy field (Baolin).
> >
> > - Patch 8: the file side is split like the anonymous one (Zi).
> > collapse_scan_file() runs under mmap_lock in the scan and only judges;
> > collapse_file() runs in the run. See Behaviour below.
> >
> > - Reviewed-by from Zi Yan and Baolin Wang on 1-4, 6 and 7.
> >
>
> I'll hopefully get too look at this soon (after digging through older stuff in
> my queue).
>
> Skimming over some patches, a note that we should not be undoing recent
> cleanups without a very good reason.
I don't think we undo it.
Both madvise and khugepaged use the same interface to the collapse
engine. Anon and file paths are handled internally in the engine. What
changed is that we have two calls into the engine instead of one.
Collapse consists of two phases: finding what to collapse and collapsing
the found range. These two phases have vastly different locking
expectations.
The scan reads a PTE table under mmap_lock, fails often and doesn't drop
the lock to move to next range.
The collapse allocates, may sleep in writeback and takes mmap_lock for
write itself. So the lock inherited from scan is no good.
collapse_single_pmd() hid that boundary inside one call. It had to drop
the lock somewhere in the middle, on some paths and not others, and the
only way for the caller to find out was the lock_dropped bool.
With scan and run as separate calls each has one lock rule: scan is
called locked and returns locked, run is called unlocked. There is
nothing left to report, so the ugly lock_dropped goes away.
It is the same move as Nico's da98790891a4 ("require collapse_huge_page
to enter/exit with the lock dropped"), one level up.
--
Kiryl Shutsemau / Kirill A. Shutemov
On 9/11/26 17:56, Kiryl Shutsemau wrote:
> On Fri, Sep 11, 2026 at 05:06:58PM +0200, David Hildenbrand (Arm) wrote:
>> On 9/10/26 14:02, Kiryl Shutsemau wrote:
>>> From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
>>>
>>> [ This is the first of the cleanups I said I would front-load ]
>>>
>>> There is no line between the collapse engine and the callers that ask for
>>> a collapse. khugepaged.c holds both, and they reach into each other.
>>>
>>> - Sixteen tests through the collapse path read cc->is_khugepaged to work
>>> out what they are allowed to do, when every one of those decisions was
>>> made by the caller before it asked.
>>>
>>> - collapse_single_pmd() does both halves of a collapse behind one call and
>>> drops mmap_lock somewhere in the middle. Which of its paths dropped it
>>> is not something a caller can see, so it hands back a bool and the
>>> caller keeps track.
>>>
>>> - MADV_COLLAPSE's implementation -- the walk over the user's range, the
>>> per-PMD loop, the errno translation -- sits in khugepaged.c, which is
>>> the daemon's file.
>>>
>>> So: draw the line. State what a caller allows in a policy, split the call
>>> in two with the lock as the boundary, and move the syscall to madvise.c.
>>> What the engine offers is then four calls, with the lock state written
>>> down against each, and a policy the caller fills for itself:
>>>
>>> collapse_control_init(cc) once, before the first table
>>> collapse_policy_*(&cc->policy) what this caller allows
>>> collapse_scan_pmd(vma, addr, ...) per table, under mmap_lock
>>> collapse_run_pmd(mm, addr, cc) when a scan found work, no mmap_lock
>>> collapse_control_release(cc) once, when done
>>>
>>> The engine stays in khugepaged.c for now; what changes is that it has an
>>> interface, and that neither half has to ask about the other. madvise.c
>>> gains the operation it should have had all along.
>>>
>>> Changes since v1
>>> ================
>>>
>>> https://lore.kernel.org/all/cover.1788533997.git.kas@kernel.org/
>>>
>>> - Rebased onto mm-new with Vernon Yang's tracepoint fixes in it. Patch 8
>>> no longer merges the two calls to each scan tracepoint, since the base
>>> already has one; its changelog now says what the status field reports.
>>>
>>> - Patch 3: nr_occupied_ptes is nr_eligible_ptes, and the mthp_collapse()
>>> comment counts eligible PTEs too (Zi, Baolin).
>>>
>>> - Patch 4: no comments on the two constants (Baolin).
>>>
>>> - Patch 5: one line per policy field (Baolin).
>>>
>>> - Patch 8: the file side is split like the anonymous one (Zi).
>>> collapse_scan_file() runs under mmap_lock in the scan and only judges;
>>> collapse_file() runs in the run. See Behaviour below.
>>>
>>> - Reviewed-by from Zi Yan and Baolin Wang on 1-4, 6 and 7.
>>>
>>
>> I'll hopefully get too look at this soon (after digging through older stuff in
>> my queue).
>>
>> Skimming over some patches, a note that we should not be undoing recent
>> cleanups without a very good reason.
>
> I don't think we undo it.
Good, I only skimmed it and read "[PATCH v2 09/12] mm/collapse: open-code
collapse_single_pmd() in its two callers".
>
> Both madvise and khugepaged use the same interface to the collapse
> engine. Anon and file paths are handled internally in the engine. What
> changed is that we have two calls into the engine instead of one.
>
> Collapse consists of two phases: finding what to collapse and collapsing
> the found range. These two phases have vastly different locking
> expectations.
>
> The scan reads a PTE table under mmap_lock, fails often and doesn't drop
> the lock to move to next range.
>
> The collapse allocates, may sleep in writeback and takes mmap_lock for
> write itself. So the lock inherited from scan is no good.
>
> collapse_single_pmd() hid that boundary inside one call. It had to drop
> the lock somewhere in the middle, on some paths and not others, and the
> only way for the caller to find out was the lock_dropped bool.
>
> With scan and run as separate calls each has one lock rule: scan is
> called locked and returns locked, run is called unlocked. There is
> nothing left to report, so the ugly lock_dropped goes away.
>
> It is the same move as Nico's da98790891a4 ("require collapse_huge_page
> to enter/exit with the lock dropped"), one level up.
>
Makes sense. I'll get to this next week!
--
Cheers,
David
On Fri, Sep 11, 2026 at 04:56:06PM +0100, Kiryl Shutsemau wrote:
> On Fri, Sep 11, 2026 at 05:06:58PM +0200, David Hildenbrand (Arm) wrote:
> > On 9/10/26 14:02, Kiryl Shutsemau wrote:
> > > From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
> > >
> > > [ This is the first of the cleanups I said I would front-load ]
> > >
> > > There is no line between the collapse engine and the callers that ask for
> > > a collapse. khugepaged.c holds both, and they reach into each other.
> > >
> > > - Sixteen tests through the collapse path read cc->is_khugepaged to work
> > > out what they are allowed to do, when every one of those decisions was
> > > made by the caller before it asked.
> > >
> > > - collapse_single_pmd() does both halves of a collapse behind one call and
> > > drops mmap_lock somewhere in the middle. Which of its paths dropped it
> > > is not something a caller can see, so it hands back a bool and the
> > > caller keeps track.
> > >
> > > - MADV_COLLAPSE's implementation -- the walk over the user's range, the
> > > per-PMD loop, the errno translation -- sits in khugepaged.c, which is
> > > the daemon's file.
> > >
> > > So: draw the line. State what a caller allows in a policy, split the call
> > > in two with the lock as the boundary, and move the syscall to madvise.c.
> > > What the engine offers is then four calls, with the lock state written
> > > down against each, and a policy the caller fills for itself:
> > >
> > > collapse_control_init(cc) once, before the first table
> > > collapse_policy_*(&cc->policy) what this caller allows
> > > collapse_scan_pmd(vma, addr, ...) per table, under mmap_lock
> > > collapse_run_pmd(mm, addr, cc) when a scan found work, no mmap_lock
> > > collapse_control_release(cc) once, when done
> > >
> > > The engine stays in khugepaged.c for now; what changes is that it has an
> > > interface, and that neither half has to ask about the other. madvise.c
> > > gains the operation it should have had all along.
> > >
> > > Changes since v1
> > > ================
> > >
> > > https://lore.kernel.org/all/cover.1788533997.git.kas@kernel.org/
> > >
> > > - Rebased onto mm-new with Vernon Yang's tracepoint fixes in it. Patch 8
> > > no longer merges the two calls to each scan tracepoint, since the base
> > > already has one; its changelog now says what the status field reports.
> > >
> > > - Patch 3: nr_occupied_ptes is nr_eligible_ptes, and the mthp_collapse()
> > > comment counts eligible PTEs too (Zi, Baolin).
> > >
> > > - Patch 4: no comments on the two constants (Baolin).
> > >
> > > - Patch 5: one line per policy field (Baolin).
> > >
> > > - Patch 8: the file side is split like the anonymous one (Zi).
> > > collapse_scan_file() runs under mmap_lock in the scan and only judges;
> > > collapse_file() runs in the run. See Behaviour below.
> > >
> > > - Reviewed-by from Zi Yan and Baolin Wang on 1-4, 6 and 7.
> > >
> >
> > I'll hopefully get too look at this soon (after digging through older stuff in
> > my queue).
> >
> > Skimming over some patches, a note that we should not be undoing recent
> > cleanups without a very good reason.
>
> I don't think we undo it.
>
> Both madvise and khugepaged use the same interface to the collapse
> engine. Anon and file paths are handled internally in the engine. What
> changed is that we have two calls into the engine instead of one.
>
> Collapse consists of two phases: finding what to collapse and collapsing
> the found range. These two phases have vastly different locking
> expectations.
>
> The scan reads a PTE table under mmap_lock, fails often and doesn't drop
> the lock to move to next range.
>
> The collapse allocates, may sleep in writeback and takes mmap_lock for
> write itself. So the lock inherited from scan is no good.
>
> collapse_single_pmd() hid that boundary inside one call. It had to drop
> the lock somewhere in the middle, on some paths and not others, and the
> only way for the caller to find out was the lock_dropped bool.
>
> With scan and run as separate calls each has one lock rule: scan is
> called locked and returns locked, run is called unlocked. There is
> nothing left to report, so the ugly lock_dropped goes away.
>
> It is the same move as Nico's da98790891a4 ("require collapse_huge_page
> to enter/exit with the lock dropped"), one level up.
Forgot to mention, this kind of split by lock boundary makes it trivial
to switch scan to per-VMA locking.
--
Kiryl Shutsemau / Kirill A. Shutemov
© 2016 - 2026 Red Hat, Inc.