[RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives

Kiryl Shutsemau posted 57 patches 1 month, 1 week ago
MAINTAINERS                                   |   19 +-
fs/proc/task_mmu.c                            |    4 +-
include/linux/huge_mm.h                       |    9 -
include/linux/mm.h                            |   14 +
include/linux/pgtable.h                       |   17 +
.../events/{huge_memory.h => collapse.h}      |  176 +-
kernel/bpf/btf.c                              |    8 +-
mm/Makefile                                   |    2 +-
mm/collapse.c                                 | 3808 +++++++++++++++++
mm/collapse.h                                 |  244 ++
mm/hugetlb.c                                  |    8 +-
mm/khugepaged.c                               | 2711 +-----------
mm/madvise.c                                  |  251 +-
mm/migrate_device.c                           |    9 +-
mm/mremap.c                                   |    2 +-
tools/testing/selftests/mm/khugepaged.c       |  346 ++
tools/testing/selftests/mm/khugepaged_race.c  |   29 +-
.../selftests/mm/khugepaged_sync_check.c      |   65 +-
tools/testing/selftests/mm/vm_util.c          |    2 +-
19 files changed, 5022 insertions(+), 2702 deletions(-)
rename include/trace/events/{huge_memory.h => collapse.h} (60%)
create mode 100644 mm/collapse.c
create mode 100644 mm/collapse.h
[RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives
Posted by Kiryl Shutsemau 1 month, 1 week ago
From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>

Yes, I know, this is a lot of changes. But I'm happy with the overall state
of the patchset and the only reason I tag it as RFC is that it is tricky
to get 57 patches upstream.

I wanted to give a view of the end state first. I will suggest a possible
way to split it below.

I would appreciate any feedback.

TL;DR
=====

This replaces khugepaged's anonymous collapse with an engine that
can collapse sub-PMD ranges. It is built around migration entries and
frozen folios instead of heavy locking and isolation, aiming for better
scalability and less disruption to the workload being collapsed.

Why
===

mTHP collapse landed in khugepaged in 7.2 and I was glad to see it.  We
at Meta run arm64 with 64K base pages, where a PMD is 512M: PMD-order THP
is of limited use at that size, and mTHP is exactly what we want.

It turned out not to help us.

khugepaged only ever looks at PMD-aligned windows, and it is not an easy
limitation to lift.

Fixing the alignment is a one-line change, but what it feeds assumes the
PMD everywhere that matters: collapse_huge_page() clears and flushes the
whole PMD whatever order it is collapsing, installs a PMD leaf because
that is the only thing it can produce, and keeps everyone out with
mmap_write_lock, anon_vma_lock_write() and an IPI broadcast while it
does.

Which is why hugepage_vma_revalidate() demands that the VMA span the
whole PMD even for an mTHP order -- "we'd need to lock all VMAs in the
PMD range to support this", as the comment there puts it.  A PMD-granular
operation is only safe when one VMA owns the PMD, and that is exactly the
restriction in the way.  The alignment is the symptom; the PMD is the
design.

So both roots have to go.

Design
======

The old mechanism holds the address space still because it has nothing
else stopping the sources from moving under the copy.  The new engine
makes the sources themselves inert instead, with the two barriers
migration already uses, raised in that order:

  1. migration entries replace the source PTEs.  Faults and GUP-slow
     now wait on the source folio's lock, which is taken before the
     first entry becomes visible.
  2. the source folio's refcount is frozen to its expected value.
     GUP-fast, pfn walkers, reclaim, compaction and memory-failure all
     fail folio_try_get() and back off.

Between the two, nothing can reach a source, so the copy runs with no
lock held at all -- and the address space is left alone while it does.

What that removes from every collapse path:

  mmap_write_lock              -> mmap_read
  anon_vma_lock_write()        -> nothing: an rmap walk needs the folio
                                  locked, and the engine holds that lock
                                  from freeze to putback
  tlb_remove_table_sync_one()  -> nothing: one ranged flush per round
  LRU isolation                -> nothing: sources are inert in place

Working in windows rather than whole PMDs takes care of the other root.
A sub-PMD window is collapsed under the page table lock, so a collapse
disturbs only the window it collapses, and each candidate is validated
at its own order -- a window need only fit its own VMA.  A PMD-order
candidate still has to own the whole PMD, which is the old rule kept
where it is still needed.

Candidates are carried through the passes a batch at a time rather than
one window at a time, so a round pays for its flush and its lock
acquisitions once.

With the barriers holding the sources still, which read lock the engine
takes stops being part of the design.  A round works inside a single
VMA, so patches 43-49 switch it from mmap_read to per-VMA locking: an
mmap_write elsewhere in the mm then stops waiting for a collapse that
has nothing to do with it.  That block is the only part of the series
that needs per-VMA locking to be unconditional, and it is a separate
dependency (see below); everything before it runs under mmap_read and
does not care.

Patch 7 sketches the engine as a comment naming every pass, what lock it
takes and what it may sleep on; the details are there rather than here.

What falls out beyond the lock diet:

 - mTHP collapse in VMAs smaller than a PMD, which is the arm64 case
   above: a 2M VMA on an arm64/64K machine collapses nothing today at
   any order, and collapses to mTHP here.
 - Hole and zeropage population at every order, so partially populated
   windows collapse to mTHP under the same max_ptes_none policy as PMD.
 - Sources come in spans -- any stretch of consecutive PTEs mapping
   consecutive pages of one folio -- so partially mapped and scrambled
   compound sources (the PTE-mapped-THP re-collapse class) work at
   every order.
 - A table that cannot become one huge page still yields the largest
   windows inside it, where before a single disqualified PTE gave up
   the whole table.

Reading the series
==================

57 patches is a lot to land on a list.  They go in blocks:

  1-6    helpers and shared state: pte_folio(), pte_none_or_zero(),
         mm/collapse.h, and the policy that replaces asking whether
         khugepaged started a collapse
  7-8    the engine's shape: entry points, a call-tree comment naming
         every pass, and the scan filled in
  9-23   the collapse half, top down: the round frame, then each pass
         in turn, then selection and the retry store
  24     per-candidate tracing, before the switch takes the old
         tracepoints away
  25-28  the switch: point the anon path at the engine, widen coverage
         to sub-PMD VMAs, delete the mechanism it replaces
  29-35  move what is left of collapse out of khugepaged.c, and
         MADV_COLLAPSE into madvise.c
  36-42  tracing: the engine's own events and trace header
  43-49  per-VMA locking, and the mm reference that makes it safe
  50-56  selftests for what the engine can now do
  57     MAINTAINERS

The two patches worth reading first if you read nothing else are 7 (the
design, as a comment naming the whole call tree) and 16 (the freeze,
which is where the safety argument lives).

A possible split, if that helps:

  1-2    two mm helpers, pte_folio() and pte_none_or_zero().  Both
         convert callers outside collapse and are useful on their own
  3-27   the engine and the switch-over.  This is the smallest unit
         that does anything: stop earlier and the tree carries an
         engine nothing calls
  28     remove the mechanism the engine replaces
  29-42  moving what is left of collapse out of khugepaged.c, and the
         engine's own tracepoints
  43-49  per-VMA locking
  50-57  selftests and MAINTAINERS

Keeping the removal separate leaves both engines in the tree with only
the new one reachable, so the switch can be reverted on its own if
something turns up.  The old mechanism is already carried that way for
three patches inside the series, so this costs nothing but 975 lines of
unreferenced code until 28 lands.  That safety net only lasts until the
blocks after it land, though: once collapse has moved out of
khugepaged.c and the locking has changed, reverting the switch no longer
gives back a working old engine.

Base and dependencies
=====================

This applies on the selftests series, not on plain mm-new:

  [PATCH v4 00/19] selftests/mm: improve khugepaged coverage
  https://lore.kernel.org/all/20260815015901.1236937-1-kirill@shutemov.name/

which is on mm-new 33f61b12d297.

Patches 43-49 depend on Suren's unconditional per-VMA locks:

  [PATCH v6 0/5] mm: Unconditional per-VMA locks and cleanups
  https://lore.kernel.org/all/20260813193433.3318288-1-surenb@google.com/

That series is not in mm-new yet, and with patch 46 applied SMP=n does
not build without it: lock_next_vma() is behind CONFIG_PER_VMA_LOCK in
mmap_lock.h.  Everything up to patch 42 builds and runs on mm-new as it
stands.  There is no fallback path by choice -- adding one would mean
carrying two locking models through every pass.

Both branches are available at

  git://git.kernel.org/pub/scm/linux/kernel/git/kas/linux.git collapse/rfc-v1

and the benchmark used for the numbers below, which is unposted and not a
dependency, at

  git://git.kernel.org/pub/scm/linux/kernel/git/kas/linux.git perf/bench-usemem

Performance
===========

Measuring khugepaged is awkward.  It is a background daemon, so what
matters is what a workload feels while it runs, not what the daemon
reports about itself -- and the usual coverage instrument is no help
below the PMD: smaps AnonHugePages only counts PMD-order folios, so it
reads zero however much mTHP has been collapsed.

So I wrote "perf bench mem usemem" for this.  It touches a region while
khugepaged works on it and reports the workload's own latency
percentiles and throughput, against per-size counters that can see
sub-PMD folios.  The branch is above; it is unposted and not a
dependency.

x86-64, production configs (no KASAN, no lockdep, no DEBUG_VM, no
PAGE_TABLE_CHECK), interleaved rounds on an idle host, equal work on
every arm.

I measured three kernels, so the two halves of the series can be told
apart in the numbers below:

  A   the base
  B   the new engine, still under mmap_read
  C   B plus per-VMA locking -- what this series ends up with

The engine: a sub-PMD collapse stops blanking the surrounding 2M
-----------------------------------------------------------------

base routes sub-PMD collapse through collapse_huge_page(), whose
pmdp_collapse_flush() and tlb_remove_table_sync_one() are not gated on
order: to collapse an order-4 window of 16 pages it clears and flushes
the whole 512-page PMD and IPIs, then repopulates.  The engine does the
window under the PTL.

A thread reading and writing a 32G region while it is collapsed at
order-4, 16384 collapses on every arm:

                         A        B        C
  read p99 (ns)       3071     1023     1023   -66.7%
  write p99 (ns)      3071      927      927   -69.8%

and the workload's read rate rises by 68% on both engine kernels.

B == C, so this is the engine, not the locking.

At PMD order the same workload is flat, and that is expected rather
than disappointing: it is the one configuration where both mechanisms
disturb exactly the same 2M.  Read it as no regression at PMD order.

Per-VMA locking: address-space operations stop waiting on the scan
-------------------------------------------------------------------

MADV_HUGEPAGE/MADV_NOHUGEPAGE toggling against a scanning mm, which is
what jemalloc does with its arenas.  4096 collapses on every arm:

                         A        B        C
  ops/sec           340656   340820   603305   +77%
  p99 (ns)           77823    86015     3327   -96%
  p99.9 (ns)         86015    94207     9215   -89%

B is about 11% worse than base at p99 here, consistently across runs:
the engine alone slightly worsens hint-toggle latency, and per-VMA
locking is what turns it into a win.  Both halves are in this series, so
C is what a reviewer gets, but the middle column is the honest one.

The trade is real in the other direction too.  On settled memory with
nothing to collapse and scan_sleep_millisecs=0, per-VMA locking costs
about 47% of scan throughput against one mmap_read for the whole walk.
That is a synthetic worst case -- the daemon wraps 8000 times a second
there, where production defaults to 10s between passes -- and it buys
mmap/munmap p99 of 56us against 1.4us.

Collapse itself is not slower
-----------------------------

One complete pass over a 32G region, 16384 collapses, khugepaged CPU
from /proc/<pid>/stat, 7 repeats:

  A base    median 5760 ms   spread 12.7%
  B engine  median 4990 ms   spread  3.8%
  C pervma  median 5020 ms   spread 12.4%

The base arm is bimodal, so its median moves with sampling and the
percentage is soft.  The distribution-free statement is better: every
engine run used less CPU than every base run.

The engine also allocates one destination per folio installed, where
base allocates 5.25 and frees the rest again: nothing is allocated until
the sources are frozen and the collapse can no longer be refused.

A measurement note, since an earlier version of this series quoted worse
figures.  khugepaged CPU has to be measured per collapse or per
completed pass, never over a fixed window with scan_sleep_millisecs=0:
the daemon never sleeps, so whichever kernel finishes the work sooner
spends the rest of the window scanning settled memory and is charged for
it.  Measured that way the engine appeared to cost 10% more CPU;
measured per unit of work it costs less.

Costs
=====

At PMD order the engine issues two TLB flushes per collapse where the
old mechanism issues one: the freeze's ranged flush plus the terminal
layer's pmdp_collapse_flush().  A PMD candidate is alone in its round,
so nothing amortizes the first.  Dropping the old per-collapse
tlb_remove_table_sync_one() IPI presumably pays for it, but that was not
measured and is not claimed here.

There may be a way out -- a PMD migration entry over the table during
the window, so the CPU never caches a walk to shoot down -- but that
means teaching every pmd-level walker a new kind of entry, and I have
not tried it.

PMD collapse deposits a freshly allocated page table instead of
redepositing the detached one.  Whoever withdraws a deposited table
frees it immediately, with nothing to hold a lockless walker off first,
and under a read lock the detached table may still be traversed by
GUP-fast or an RCU pte walk.  It goes to pte_free_defer() instead,
exactly as retract_page_tables() does.  One transient table page per PMD
collapse buys the IPI's absence.

That cost goes away if zap_deposited_table() -- the only site that frees
a deposited table outright, the others redeposit it or repopulate the
PMD with it -- used pte_free_defer().  The deposit would no longer have
to be quiescent and the detached table could go straight back.  It would
defer every THP zap's table free, and I have not tried it.

A shared source now costs an extra copy.  The freeze needs every page
exclusive to this mm, so the fault-in pass breaks CoW first -- an
allocation and a copy -- and the collapse then copies that page into the
destination; the old mechanism copied a shared page straight into the
new folio and broke the sharing that way.  It is bounded by
max_ptes_shared, which khugepaged holds at zero below the PMD order, so
in practice this is PMD-order collapse and MADV_COLLAPSE.

Size
====

mm/ grows by 1915 lines net: 4475 added against 2560 deleted.

That is not a claim that this is less code, but it is less than it
looks.  khugepaged.c goes from 3283 lines to 908.  The new engine is
4052 lines across mm/collapse.c and mm/collapse.h, of which 1344 --
about a third -- are comments, which is where the pipeline's invariants
are written down.  What replaces three install paths with their own
isolate/copy/rollback is one engine and one contract.

Testing
=======

Both matrices run the mm selftests plus a race harness, on the
validation config: KASAN, lockdep, PROVE_LOCKING, DEBUG_VM and
PAGE_TABLE_CHECK, 16G of guest memory, swap active so the swap-in
prepass is exercised rather than skipped.

  x86-64        433 pass, 0 fail, 12 skip
  arm64/64K     581 pass, 0 fail, 18 skip

dmesg clean on both.  The arm64 skips are a pre-existing shmem
MADV_COLLAPSE -EINVAL on 64K pages, confirmed against the base by A/B.

Every one of the 57 patches builds with no new warnings; !NUMA and !MMU
(arm nommu) build clean.  SMP=n does not build, for the reason in the
dependencies section above.

The race harness also gets longer soaks -- 1800s per driver mode, with
memory pressure and swap -- and the engine is fuzzed with syzkaller on a
KCOV+KASAN build.  That found two bugs the selftests could not reach: a
teardown that dropped rmap while the source was still frozen, where
removing an mlocked mapping munlocks and munlock_folio() takes a
reference a frozen folio forbids; and a whole-table MADV_DONTNEED racing
the copy window under CONFIG_PT_RECLAIM, which freed the table and left
the sources frozen and locked.  Both are fixed, and both gained coverage
-- the mlocked case is patch 53.

Kiryl Shutsemau (Meta) (57):
  mm: add pte_folio()
  mm: add pte_none_or_zero()
  mm/collapse: add collapse.h for the shared collapse state
  mm/collapse: rename mthp_present_ptes to eligible_ptes
  mm/collapse: state what a collapse may do in the policy
  mm/collapse: move the smallest collapse order to collapse.h
  mm/collapse: sketch the new anonymous collapse engine
  mm/collapse: scan a table for what a collapse could use
  mm/collapse: collect candidate windows into a round
  mm/collapse: run a round and feed the outcomes back
  mm/collapse: sketch the passes of a round
  mm/collapse: allocate a destination per candidate
  mm/collapse: revalidate a round against the VMA
  mm/collapse: fault the sources in before the freeze
  mm/collapse: check what a candidate would freeze
  mm/collapse: freeze the sources behind migration entries
  mm/collapse: copy the sources into the destinations
  mm/collapse: install the destinations at PTE level
  mm/collapse: install a PMD leaf as the terminal layer
  mm/collapse: put the sources back
  mm/collapse: settle whatever the round reached
  mm/collapse: walk a table with a selection cursor
  mm/collapse: give a refused region a second chance
  mm/collapse: report each candidate's outcome to tracing
  mm/collapse: collapse anonymous memory with the new engine
  mm/collapse: give collapse_single_pmd() the range to work on
  mm/collapse: scan the windows a VMA can hold
  mm/collapse: remove the mechanism the engine replaces
  mm/collapse: move what a collapse is judged on into collapse.c
  mm/collapse: name the max_ptes ceiling after collapse
  mm/khugepaged: count collapses where khugepaged makes them
  mm/collapse: move the file collapse into collapse.c
  mm/collapse: split collapse into a scan and a run
  mm/collapse: implement MADV_COLLAPSE in madvise.c
  mm/madvise: drop MADV_COLLAPSE's redundant mm reference
  mm/collapse: report what the scan found
  mm/collapse: report what the fault-in pass paid
  mm/collapse: report the round, and what it made faulters wait
  mm/collapse: name the file collapse's tracepoints after collapse
  mm/collapse: remove the tracepoints of the mechanism that is gone
  mm/collapse: give collapse its own trace header
  mm/collapse: allow error injection into the freeze
  mm/khugepaged: check the scan budget before the work, not after
  mm/khugepaged: hold the address space open across a scan
  mm/collapse: take a per-VMA read lock for the round
  mm/khugepaged: scan under a per-VMA read lock
  mm/madvise: collapse under a per-VMA read lock
  mm/collapse: assert the mm reference the engine relies on
  mm/khugepaged: drop the mmap_lock barrier from __khugepaged_exit()
  selftests/mm: attribute collapses by candidate event alone
  selftests/mm: cover collapse inside a sub-PMD VMA
  selftests/mm: cover a hole-y window in a sub-PMD VMA
  selftests/mm: cover collapse of mlocked ranges
  selftests/mm: cover collapse beside a MADV_FREE'd page
  selftests/mm: cover collapse beside a pinned page
  selftests/mm: cover the scaled max_ptes_shared limit
  MAINTAINERS: add an entry for collapse

 MAINTAINERS                                   |   19 +-
 fs/proc/task_mmu.c                            |    4 +-
 include/linux/huge_mm.h                       |    9 -
 include/linux/mm.h                            |   14 +
 include/linux/pgtable.h                       |   17 +
 .../events/{huge_memory.h => collapse.h}      |  176 +-
 kernel/bpf/btf.c                              |    8 +-
 mm/Makefile                                   |    2 +-
 mm/collapse.c                                 | 3808 +++++++++++++++++
 mm/collapse.h                                 |  244 ++
 mm/hugetlb.c                                  |    8 +-
 mm/khugepaged.c                               | 2711 +-----------
 mm/madvise.c                                  |  251 +-
 mm/migrate_device.c                           |    9 +-
 mm/mremap.c                                   |    2 +-
 tools/testing/selftests/mm/khugepaged.c       |  346 ++
 tools/testing/selftests/mm/khugepaged_race.c  |   29 +-
 .../selftests/mm/khugepaged_sync_check.c      |   65 +-
 tools/testing/selftests/mm/vm_util.c          |    2 +-
 19 files changed, 5022 insertions(+), 2702 deletions(-)
 rename include/trace/events/{huge_memory.h => collapse.h} (60%)
 create mode 100644 mm/collapse.c
 create mode 100644 mm/collapse.h


base-commit: 8b76faf42c5d342b3bf0b1fd97bdaf603ee57354
-- 
2.54.0
Re: [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives
Posted by David Hildenbrand (Arm) 1 month, 1 week ago
On 8/17/26 00:45, Kiryl Shutsemau wrote:
> From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
> 
> Yes, I know, this is a lot of changes. But I'm happy with the overall state
> of the patchset and the only reason I tag it as RFC is that it is tricky
> to get 57 patches upstream.
> 
> I wanted to give a view of the end state first. I will suggest a possible
> way to split it below.
> 
> I would appreciate any feedback.

Replying here on the overall design first before reading into the other
discussions (and process related topics).

> 
> TL;DR
> =====
> 
> This replaces khugepaged's anonymous collapse with an engine that
> can collapse sub-PMD ranges. It is built around migration entries and
> frozen folios instead of heavy locking and isolation, aiming for better
> scalability and less disruption to the workload being collapsed.

I recall us discussing something around using some PTE/PMD markers (e.g.,
migration entries) in the past.

One thing that needed care is handling concurrent MADV_DONTNEED + faultin after
dropping relevant locks.

> 
> Why
> ===
> 
> mTHP collapse landed in khugepaged in 7.2 and I was glad to see it.  We
> at Meta run arm64 with 64K base pages, where a PMD is 512M: PMD-order THP
> is of limited use at that size, and mTHP is exactly what we want.

Right, as the first step, we decided to go for the simpler and minimally
intrusive approach of using the existing mechanism that always operates on PMD
ranges, keeping using the existing pmdp_collapse_flush()-based mechanism and
locking in place.

That's why anything more elaborate will require significantly more LOC :)

> 
> It turned out not to help us.
> 
> khugepaged only ever looks at PMD-aligned windows, and it is not an easy
> limitation to lift.


I recall we discussed some simpler way to make this work with VMAs that don't
fully span PMDs: I think write-locking VMAs (+ rmap) that cover the PMD was
discussed as a low-hanging fruit, such that other page table walkers would not
suddenly stumble over the temporarily removed page table.

So the mmap_write_lock() + vma write-locks prevents concurrent mmap+page faults
and the rmap locks prevent concurrent rmap walks.

There were discussions on the impact when a PMD spans many VMAs, and I think one
conclusion was that such scenarios are likely not worth considering (e.g., 512
VMAs in a single PMD, all with different rmap locks; your workload sucks already).

> 
> Fixing the alignment is a one-line change, but what it feeds assumes the
> PMD everywhere that matters: collapse_huge_page() clears and flushes the
> whole PMD whatever order it is collapsing, installs a PMD leaf because
> that is the only thing it can produce, and keeps everyone out with
> mmap_write_lock, anon_vma_lock_write() and an IPI broadcast while it
> does.
> 
> Which is why hugepage_vma_revalidate() demands that the VMA span the
> whole PMD even for an mTHP order -- "we'd need to lock all VMAs in the
> PMD range to support this", as the comment there puts it.  A PMD-granular
> operation is only safe when one VMA owns the PMD, and that is exactly the
> restriction in the way.  The alignment is the symptom; the PMD is the
> design.

I disagree with "A PMD-granular operation is only safe when one VMA owns the
PMD". It's safe when all page table walkers can be stopped (see above).

> 
> So both roots have to go.
> 
> Design
> ======
> 
> The old mechanism holds the address space still because it has nothing
> else stopping the sources from moving under the copy.  The new engine
> makes the sources themselves inert instead, with the two barriers
> migration already uses, raised in that order:
> 
>   1. migration entries replace the source PTEs.  Faults and GUP-slow
>      now wait on the source folio's lock, which is taken before the
>      first entry becomes visible.
>   2. the source folio's refcount is frozen to its expected value.
>      GUP-fast, pfn walkers, reclaim, compaction and memory-failure all
>      fail folio_try_get() and back off.
> 
> Between the two, nothing can reach a source, so the copy runs with no
> lock held at all -- and the address space is left alone while it does.

Right. Concurrent MADV_DONTNEED can zap migration PTEs and other faults even
re-fault fresh anon folios. So that must be detected before replacing migration
entries again I guess.

> 
> What that removes from every collapse path:
> 
>   mmap_write_lock              -> mmap_read
>   anon_vma_lock_write()        -> nothing: an rmap walk needs the folio
>                                   locked, and the engine holds that lock
>                                   from freeze to putback
>   tlb_remove_table_sync_one()  -> nothing: one ranged flush per round
>   LRU isolation                -> nothing: sources are inert in place
> 

Not sure how you handle PMD collapse. I recall problems with migration entries
on the PMD level for non-folio things (we discussed something along these lines
also in the past).

> Working in windows rather than whole PMDs takes care of the other root.
> A sub-PMD window is collapsed under the page table lock, so a collapse
> disturbs only the window it collapses, and each candidate is validated

I recall us discussing that holding the PT lock for a longer collapse operation
(especially on 64k) is problematic. But I don't get all the details from your
description here.

> at its own order -- a window need only fit its own VMA.  A PMD-order
> candidate still has to own the whole PMD, which is the old rule kept
> where it is still needed.
> 
> Candidates are carried through the passes a batch at a time rather than
> one window at a time, so a round pays for its flush and its lock
> acquisitions once.

Now I am starting to feel that there are too many changes packed in a single
series :)

> 
> With the barriers holding the sources still, which read lock the engine
> takes stops being part of the design.  A round works inside a single
> VMA, so patches 43-49 switch it from mmap_read to per-VMA locking: an
> mmap_write elsewhere in the mm then stops waiting for a collapse that
> has nothing to do with it.  That block is the only part of the series
> that needs per-VMA locking to be unconditional, and it is a separate
> dependency (see below); everything before it runs under mmap_read and
> does not care.
> 
> Patch 7 sketches the engine as a comment naming every pass, what lock it
> takes and what it may sleep on; the details are there rather than here.
> 
> What falls out beyond the lock diet:
> 
>  - mTHP collapse in VMAs smaller than a PMD, which is the arm64 case
>    above: a 2M VMA on an arm64/64K machine collapses nothing today at
>    any order, and collapses to mTHP here.
>  - Hole and zeropage population at every order, so partially populated
>    windows collapse to mTHP under the same max_ptes_none policy as PMD.
>  - Sources come in spans -- any stretch of consecutive PTEs mapping
>    consecutive pages of one folio -- so partially mapped and scrambled
>    compound sources (the PTE-mapped-THP re-collapse class) work at
>    every order.
>  - A table that cannot become one huge page still yields the largest
>    windows inside it, where before a single disqualified PTE gave up
>    the whole table.
> 
> Reading the series
> ==================
> 
> 57 patches is a lot to land on a list.  They go in blocks:
> 
>   1-6    helpers and shared state: pte_folio(), pte_none_or_zero(),
>          mm/collapse.h, and the policy that replaces asking whether
>          khugepaged started a collapse
>   7-8    the engine's shape: entry points, a call-tree comment naming
>          every pass, and the scan filled in
>   9-23   the collapse half, top down: the round frame, then each pass
>          in turn, then selection and the retry store

I fail to parse this sentence.

>   24     per-candidate tracing, before the switch takes the old
>          tracepoints away
>   25-28  the switch: point the anon path at the engine, widen coverage
>          to sub-PMD VMAs, delete the mechanism it replaces
>   29-35  move what is left of collapse out of khugepaged.c, and
>          MADV_COLLAPSE into madvise.c
>   36-42  tracing: the engine's own events and trace header
>   43-49  per-VMA locking, and the mm reference that makes it safe
>   50-56  selftests for what the engine can now do
>   57     MAINTAINERS
> 
> The two patches worth reading first if you read nothing else are 7 (the
> design, as a comment naming the whole call tree) and 16 (the freeze,
> which is where the safety argument lives).
> 
> A possible split, if that helps:
> 
>   1-2    two mm helpers, pte_folio() and pte_none_or_zero().  Both
>          convert callers outside collapse and are useful on their own
>   3-27   the engine and the switch-over.  This is the smallest unit
>          that does anything: stop earlier and the tree carries an
>          engine nothing calls
>   28     remove the mechanism the engine replaces
>   29-42  moving what is left of collapse out of khugepaged.c, and the
>          engine's own tracepoints
>   43-49  per-VMA locking
>   50-57  selftests and MAINTAINERS
> 
> Keeping the removal separate leaves both engines in the tree with only
> the new one reachable, so the switch can be reverted on its own if
> something turns up.  The old mechanism is already carried that way for
> three patches inside the series, so this costs nothing but 975 lines of
> unreferenced code until 28 lands.  That safety net only lasts until the
> blocks after it land, though: once collapse has moved out of
> khugepaged.c and the locking has changed, reverting the switch no longer
> gives back a working old engine.

[...]

> Performance
> ===========
> 
> Measuring khugepaged is awkward.  It is a background daemon, so what
> matters is what a workload feels while it runs, not what the daemon
> reports about itself -- and the usual coverage instrument is no help
> below the PMD: smaps AnonHugePages only counts PMD-order folios, so it
> reads zero however much mTHP has been collapsed.
> 
> So I wrote "perf bench mem usemem" for this.  It touches a region while
> khugepaged works on it and reports the workload's own latency
> percentiles and throughput, against per-size counters that can see
> sub-PMD folios.  The branch is above; it is unposted and not a
> dependency.

Something more realistic might be running some workload in a VM whereby the VM
is getting collapsed by khugepaged.

[...]

> There may be a way out -- a PMD migration entry over the table during
> the window, so the CPU never caches a walk to shoot down -- but that
> means teaching every pmd-level walker a new kind of entry, and I have
> not tried it.

I think we discussed that in the past and it's absolutely nasty.

[...]

> Size
> ====
> 
> mm/ grows by 1915 lines net: 4475 added against 2560 deleted.
> 
That's quite a lot for something that reads like a cleanup at first.


Okay, let me read the other discussions.

-- 
Cheers,

David
Re: [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives
Posted by Kiryl Shutsemau 1 month, 1 week ago
On Tue, Aug 18, 2026 at 03:55:55PM +0200, David Hildenbrand (Arm) wrote:
> > This replaces khugepaged's anonymous collapse with an engine that
> > can collapse sub-PMD ranges. It is built around migration entries and
> > frozen folios instead of heavy locking and isolation, aiming for better
> > scalability and less disruption to the workload being collapsed.
> 
> I recall us discussing something around using some PTE/PMD markers (e.g.,
> migration entries) in the past.
> 
> One thing that needed care is handling concurrent MADV_DONTNEED + faultin after
> dropping relevant locks.

Handled at install time.

I drop the PTL after the freeze to allow allocation and copy, but sample
the PTE values (see saved_ptes) at freeze time. If something changed
under us by the time we install the new page table entries, we give up
on that candidate and roll it back; the rest of the round still
installs. We allow harmless transitions: zero page to none.

> > Which is why hugepage_vma_revalidate() demands that the VMA span the
> > whole PMD even for an mTHP order -- "we'd need to lock all VMAs in the
> > PMD range to support this", as the comment there puts it.  A PMD-granular
> > operation is only safe when one VMA owns the PMD, and that is exactly the
> > restriction in the way.  The alignment is the symptom; the PMD is the
> > design.
> 
> I disagree with "A PMD-granular operation is only safe when one VMA owns the
> PMD". It's safe when all page table walkers can be stopped (see above).

Fair, the sentence is too strong.  mmap_write_lock plus a VMA write lock on
every VMA the PMD covers, plus their rmap locks, would make it safe.

But the mechanism still clears the whole PMD, flushes, IPIs and
repopulates it to collapse each 16-page window. That's very noisy to the
workload.

And I am not sure how to deal with rmap locking here.  Nothing in mm
holds two unrelated anon_vma rwsems: vma_prepare() takes one for both VMAs
it touches, because a merge requires them to share the anon_vma, and
anon_vma_clone() takes one because "all anon_vma's share the same root".
The lock ordering in mm/rmap.c has a single anon_vma->rwsem level, so a PMD
spanning unrelated mappings would need an ordering rule that does not exist
today.

> > Between the two, nothing can reach a source, so the copy runs with no
> > lock held at all -- and the address space is left alone while it does.
> 
> Right. Concurrent MADV_DONTNEED can zap migration PTEs and other faults even
> re-fault fresh anon folios. So that must be detected before replacing migration
> entries again I guess.

Yes, that is the install-time check above.

The copy itself is safe: a zap of a migration entry only clears the slot
and adjusts rss -- see zap_nonpresent_ptes(), which neither puts the
folio nor drops its rmap -- so a frozen, locked source cannot go away
under the copy.

The rollback does the rest: it leaves a foreign slot as found and drops
the rmap the freeze kept, which is the half of the teardown the zapper
could not do.

> 
> > 
> > What that removes from every collapse path:
> > 
> >   mmap_write_lock              -> mmap_read
> >   anon_vma_lock_write()        -> nothing: an rmap walk needs the folio
> >                                   locked, and the engine holds that lock
> >                                   from freeze to putback
> >   tlb_remove_table_sync_one()  -> nothing: one ranged flush per round
> >   LRU isolation                -> nothing: sources are inert in place
> > 
> 
> Not sure how you handle PMD collapse. I recall problems with migration entries
> on the PMD level for non-folio things (we discussed something along these lines
> also in the past).

There are no PMD-level migration entries here.

The freeze is always at PTE level, so the pmd keeps pointing at the
table until the last step, and the PMD leaf goes in as the terminal
layer: verify, pmdp_collapse_flush(), deposit a fresh table, set the
leaf, all in one section under the pmd lock with the pte ptl nested
inside.

A pmd-level walker sees the old table or the leaf and never pmd_none,
and faults stay held at pte level by the migration entries throughout,
which is what lets PMD collapse run under a VMA read lock like
everything else.

> 
> > Working in windows rather than whole PMDs takes care of the other root.
> > A sub-PMD window is collapsed under the page table lock, so a collapse
> > disturbs only the window it collapses, and each candidate is validated
> 
> I recall us discussing that holding the PT lock for a longer collapse operation
> (especially on 64k) is problematic. But I don't get all the details from your
> description here.

As I mentioned above, we drop the ptl after the freeze. And take it a
second time for the install.  Allocation and the copy run in between with
no lock held -- on 64K the copy at PMD order is 512M of it, which is why
it cannot sit under either lock.

> 
> > at its own order -- a window need only fit its own VMA.  A PMD-order
> > candidate still has to own the whole PMD, which is the old rule kept
> > where it is still needed.
> > 
> > Candidates are carried through the passes a batch at a time rather than
> > one window at a time, so a round pays for its flush and its lock
> > acquisitions once.
> 
> Now I am starting to feel that there are too many changes packed in a single
> series :)

Batching is not an extra here, it is the only way collapse works at mTHP
sizes.  What a round pays once -- lru_add_drain(), the mmu_notifier
invalidate window, the two ptl acquisitions and the TLB flush -- would
otherwise be paid per collapse.  A 2M PMD is 32 collapses at order-4, and
a gigabyte collapsed into 64K mTHP is 16384 of them: the serialization
would dominate, and the workload would feel every one.

> > Reading the series
> > ==================
> > 
> > 57 patches is a lot to land on a list.  They go in blocks:
> > 
> >   1-6    helpers and shared state: pte_folio(), pte_none_or_zero(),
> >          mm/collapse.h, and the policy that replaces asking whether
> >          khugepaged started a collapse
> >   7-8    the engine's shape: entry points, a call-tree comment naming
> >          every pass, and the scan filled in
> >   9-23   the collapse half, top down: the round frame, then each pass
> >          in turn, then selection and the retry store
> 
> I fail to parse this sentence.

Bad wording.  The engine is introduced in a top-down manner:

  - patch 7 adds the external interface and a comment naming every pass
  - 8 the scan
  - 9-11 how candidates become a round and which passes the round has
  - 12-21 fill in one pass per patch in the order they run
  - 22-23 add the cursor that picks the windows and the second chance
    for the refused ones.

I'll spell that out in the next iteration.

> > So I wrote "perf bench mem usemem" for this.  It touches a region while
> > khugepaged works on it and reports the workload's own latency
> > percentiles and throughput, against per-size counters that can see
> > sub-PMD folios.  The branch is above; it is unposted and not a
> > dependency.
> 
> Something more realistic might be running some workload in a VM whereby the VM
> is getting collapsed by khugepaged.

Fair, I'll run that: a guest touching its memory while the host collapses
the backing mapping, measured from inside the guest.

> 
> [...]
> 
> > There may be a way out -- a PMD migration entry over the table during
> > the window, so the CPU never caches a walk to shoot down -- but that
> > means teaching every pmd-level walker a new kind of entry, and I have
> > not tried it.
> 
> I think we discussed that in the past and it's absolutely nasty.

Let's leave this for now. I will give it a try when it gets its turn.

> 
> [...]
> 
> > Size
> > ====
> > 
> > mm/ grows by 1915 lines net: 4475 added against 2560 deleted.
> > 
> That's quite a lot for something that reads like a cleanup at first.

About half of that is comments. :)

-- 
  Kiryl Shutsemau / Kirill A. Shutemov
Re: [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives
Posted by David Hildenbrand (Arm) 2 weeks ago
On 8/19/26 19:09, Kiryl Shutsemau wrote:
> On Tue, Aug 18, 2026 at 03:55:55PM +0200, David Hildenbrand (Arm) wrote:
>>> This replaces khugepaged's anonymous collapse with an engine that
>>> can collapse sub-PMD ranges. It is built around migration entries and
>>> frozen folios instead of heavy locking and isolation, aiming for better
>>> scalability and less disruption to the workload being collapsed.
>>
>> I recall us discussing something around using some PTE/PMD markers (e.g.,
>> migration entries) in the past.
>>
>> One thing that needed care is handling concurrent MADV_DONTNEED + faultin after
>> dropping relevant locks.
> 
> Handled at install time.
> 
> I drop the PTL after the freeze to allow allocation and copy, but sample
> the PTE values (see saved_ptes) at freeze time. If something changed
> under us by the time we install the new page table entries, we give up
> on that candidate and roll it back; the rest of the round still
> installs. We allow harmless transitions: zero page to none.

I think there was more to it, and Jann also hinted at some examples in his
reply. But we'll get to that part once it's no longer buried in 39 patches ;)

> 
>>> Which is why hugepage_vma_revalidate() demands that the VMA span the
>>> whole PMD even for an mTHP order -- "we'd need to lock all VMAs in the
>>> PMD range to support this", as the comment there puts it.  A PMD-granular
>>> operation is only safe when one VMA owns the PMD, and that is exactly the
>>> restriction in the way.  The alignment is the symptom; the PMD is the
>>> design.
>>
>> I disagree with "A PMD-granular operation is only safe when one VMA owns the
>> PMD". It's safe when all page table walkers can be stopped (see above).
> 
> Fair, the sentence is too strong.  mmap_write_lock plus a VMA write lock on
> every VMA the PMD covers, plus their rmap locks, would make it safe.

Ack.

> 
> But the mechanism still clears the whole PMD, flushes, IPIs and
> repopulates it to collapse each 16-page window. That's very noisy to the
> workload.
> 

Right. khugepaged itself is pretty noise already, though. So one would have to
understand "how much more noisy and who cares".

I can understand why one would want to make khugepaged less noisy, though.

> And I am not sure how to deal with rmap locking here.  Nothing in mm
> holds two unrelated anon_vma rwsems: vma_prepare() takes one for both VMAs
> it touches, because a merge requires them to share the anon_vma, and
> anon_vma_clone() takes one because "all anon_vma's share the same root".
> The lock ordering in mm/rmap.c has a single anon_vma->rwsem level, so a PMD
> spanning unrelated mappings would need an ordering rule that does not exist
> today.

That's a good point. Try-locking would likely work but might have other effects.
Nobody tried this so far.

Lorenzo is on his way to simplify a lot on that anon_vma front (scalable cow),
and IIRC it would also simplify that case.

> 
>>> Between the two, nothing can reach a source, so the copy runs with no
>>> lock held at all -- and the address space is left alone while it does.
>>
>> Right. Concurrent MADV_DONTNEED can zap migration PTEs and other faults even
>> re-fault fresh anon folios. So that must be detected before replacing migration
>> entries again I guess.
> 
> Yes, that is the install-time check above.
> 
> The copy itself is safe: a zap of a migration entry only clears the slot
> and adjusts rss -- see zap_nonpresent_ptes(), which neither puts the
> folio nor drops its rmap -- so a frozen, locked source cannot go away
> under the copy.

I'll have to think about the impact of having these folios frozen for a longer
time, instead of only very briefly during migration.

E.g., these folios will then be unmovable for the entirety of the collapse
operation, because folio_try_get() by memory offlining/cma/compaction will just
fail.

[...]

> 
> There are no PMD-level migration entries here.

That's good.

> 
> The freeze is always at PTE level, so the pmd keeps pointing at the
> table until the last step, and the PMD leaf goes in as the terminal
> layer: verify, pmdp_collapse_flush(), deposit a fresh table, set the
> leaf, all in one section under the pmd lock with the pte ptl nested
> inside.
> 
> A pmd-level walker sees the old table or the leaf and never pmd_none,
> and faults stay held at pte level by the migration entries throughout,
> which is what lets PMD collapse run under a VMA read lock like
> everything else.
> 
>>
>>> Working in windows rather than whole PMDs takes care of the other root.
>>> A sub-PMD window is collapsed under the page table lock, so a collapse
>>> disturbs only the window it collapses, and each candidate is validated
>>
>> I recall us discussing that holding the PT lock for a longer collapse operation
>> (especially on 64k) is problematic. But I don't get all the details from your
>> description here.
> 
> As I mentioned above, we drop the ptl after the freeze. And take it a
> second time for the install.  Allocation and the copy run in between with
> no lock held -- on 64K the copy at PMD order is 512M of it, which is why
> it cannot sit under either lock.

Hold on, are freezing all folios to collapse? That cannot possibly work with
COW-shared folios that are mapped into other address spaces.

We must only freeze a folio if we are sure that no frozen reference can go away
concurrently.

But maybe I misunderstood or you handle this in a special way elsewhere.

-- 
Cheers,

David
Re: [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives
Posted by Kiryl Shutsemau 1 week, 6 days ago
On Mon, Sep 14, 2026 at 05:07:37PM +0200, David Hildenbrand (Arm) wrote:
> > As I mentioned above, we drop the ptl after the freeze. And take it a
> > second time for the install.  Allocation and the copy run in between with
> > no lock held -- on 64K the copy at PMD order is 512M of it, which is why
> > it cannot sit under either lock.
> 
> Hold on, are freezing all folios to collapse? That cannot possibly work with
> COW-shared folios that are mapped into other address spaces.

I break COW for worthy candidates in the faultin prepass before freeze,
alongside swapin. It simplifies the picture substantially. No rmap mess.

-- 
  Kiryl Shutsemau / Kirill A. Shutemov
Re: [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives
Posted by David Hildenbrand (Arm) 1 week, 6 days ago
On 9/15/26 12:46, Kiryl Shutsemau wrote:
> On Mon, Sep 14, 2026 at 05:07:37PM +0200, David Hildenbrand (Arm) wrote:
>>> As I mentioned above, we drop the ptl after the freeze. And take it a
>>> second time for the install.  Allocation and the copy run in between with
>>> no lock held -- on 64K the copy at PMD order is 512M of it, which is why
>>> it cannot sit under either lock.
>>
>> Hold on, are freezing all folios to collapse? That cannot possibly work with
>> COW-shared folios that are mapped into other address spaces.
> 
> I break COW for worthy candidates in the faultin prepass before freeze,
> alongside swapin. It simplifies the picture substantially. No rmap mess.

Hm, okay. So temporary memory allocations and page copying.

-- 
Cheers,

David
Re: [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives
Posted by Lorenzo Stoakes (ARM) 1 month, 1 week ago
+cc Pedro for discussion about perf numbers

On Sun, Aug 16, 2026 at 11:45:12PM +0100, Kiryl Shutsemau wrote:
> From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
>
> Yes, I know, this is a lot of changes. But I'm happy with the overall state
> of the patchset and the only reason I tag it as RFC is that it is tricky
> to get 57 patches upstream.
>
> I wanted to give a view of the end state first. I will suggest a possible
> way to split it below.
>
> I would appreciate any feedback.

Oh Kiryl :)

We have a THP cabal meeting every couple of weeks where it would have been
useful for you to raise this first.

In any case - this series is not something we'd consider at the moment,
even broken into parts.

David and I have put THP into feature freeze - until the codebase is
subtantially improved we're not really interested in seeing significant
development.

The technical debt is substantial and has to be paid down first.

See [0] for a rough list of TODOs in this regard.

(I also intend to do at least 1 series that improves things shortly
myself.)

mTHP khugepaged was allowed in _despite_ my being reticent given this
technical debt, but with a proviso of a feature freeze afterwards.

It's still fine obviously to send a theoretical RFC for early feedback. But
there can be no un-RFC'ing any time soon.

I'd suggest looking at cleanups that could act as foundations for your
changes.

I'll try to find some time to look through the actual patches to make
concrete suggestions/see if we could break some of this out like that.

Some more comments below.

>
> TL;DR
> =====
>
> This replaces khugepaged's anonymous collapse with an engine that
> can collapse sub-PMD ranges. It is built around migration entries and
> frozen folios instead of heavy locking and isolation, aiming for better
> scalability and less disruption to the workload being collapsed.

OK interesting :)

>
> Why
> ===
>
> mTHP collapse landed in khugepaged in 7.2 and I was glad to see it.  We
> at Meta run arm64 with 64K base pages, where a PMD is 512M: PMD-order THP
> is of limited use at that size, and mTHP is exactly what we want.

Do you have some numbers that indicate to what degree mTHP khugepaged is
beneficial?

I was discussing this with Pedro recently and he pointed out that we have
very poor coverage of the actual observed benefits of mTHP khugepaged
(there are reasons to have merged it anyway but still).

Some detailed analysis on how it benefits Meta would be really beneficial,
with specifics and what settings you find useful.

Is it only useful for larger page size? What orders have you found to be
most effective?

Is it useful in itself or do you view it as a stepping stone to further
work, like this?

>
> It turned out not to help us.

I mean maybe this answers my question :)

>
> khugepaged only ever looks at PMD-aligned windows, and it is not an easy
> limitation to lift.

Yes. This assumption is very much baked in.

I guess this is coming from the perspective of having ranges that are
neither PMD-aligned nor sized (far harder to achieve with 512 MiB PMD size
obviously).

>
> Fixing the alignment is a one-line change, but what it feeds assumes the

Hmm not so sure about that... especially given how baked in these
assumptions are.

Which by the way, all speaks to the need for rework.

The first stage in my view would be to improve the code to the point that
these kinds of assumptions fall out of it, which then lays the foundations
for future changes to eliminate the assumptions.

> PMD everywhere that matters: collapse_huge_page() clears and flushes the
> whole PMD whatever order it is collapsing, installs a PMD leaf because
> that is the only thing it can produce, and keeps everyone out with
> mmap_write_lock, anon_vma_lock_write() and an IPI broadcast while it
> does.

To be clear - the anon path. I think important to clarify :)

And yeah it does IPI for any sensible arch (with
CONFIG_MMU_GATHER_RCU_TABLE_FREE) via tlb_remove_table_sync_one(). The
other arches IPI anyway on TLB invalidation.

[Though I intend to make all page table freeing RCU relatively soon which
should? Eliminate the need for this, possibly?]

It doesn't insert a PMD leaf entry for mTHP collapse though, that's
incorrect.

It does, as you say, clear down the PMD even on mTHP PTE range
installation, and has to hold the rmap lock for a longer period to account
for adjacent ranges that share a PMD entry.

All in all it's a mess yes and clearly can be improved.

>
> Which is why hugepage_vma_revalidate() demands that the VMA span the
> whole PMD even for an mTHP order -- "we'd need to lock all VMAs in the
> PMD range to support this", as the comment there puts it.  A PMD-granular
> operation is only safe when one VMA owns the PMD, and that is exactly the
> restriction in the way.  The alignment is the symptom; the PMD is the
> design.
>
> So both roots have to go.

Yeah this is pretty radical stuff so it's just an absolute no until
technical debt is paid down properly.

That means signficant rework to make this code something far from the
overly-coupled spaghetti mess that it is now.

Efforts have been ongoing with this, Nico and others have submitted
cleanups and obviously more is required here.

But yes I think decoupling the ossified PMD assumption is important.

However we have to remember that a lot of the user-visible API assumes PMD
sizing and so the code has to clearly reflect this and make it clear that

>
> Design
> ======
>
> The old mechanism holds the address space still because it has nothing
> else stopping the sources from moving under the copy.  The new engine
> makes the sources themselves inert instead, with the two barriers
> migration already uses, raised in that order:
>
>   1. migration entries replace the source PTEs.  Faults and GUP-slow
>      now wait on the source folio's lock, which is taken before the
>      first entry becomes visible.
>   2. the source folio's refcount is frozen to its expected value.
>      GUP-fast, pfn walkers, reclaim, compaction and memory-failure all
>      fail folio_try_get() and back off.
>
> Between the two, nothing can reach a source, so the copy runs with no
> lock held at all -- and the address space is left alone while it does.

Hmm, are migration entries the right mechnanism here? Are you actually
migrating the pages to a large folio here, or using them to get the
behaviour you want on fault/GUP?

Same question in general for the freezing.

I'm not suggesting quite that these are the wrong mechanisms, just querying
how they are being used here.

I do like the idea of using the folio lock though. All rmap operations now
take the folio lock when traversing the rmap.

>
> What that removes from every collapse path:
>
>   mmap_write_lock              -> mmap_read
>   anon_vma_lock_write()        -> nothing: an rmap walk needs the folio
>                                   locked, and the engine holds that lock
>                                   from freeze to putback

I do like the idea of eliminating uses of the rmap lock like this, not only
for contention's sake but also for scalable CoW purposes which introduces
challenges with regards to holding these.

In fact, migration and huge memory collapse are the really problematic
areas.

>   tlb_remove_table_sync_one()  -> nothing: one ranged flush per round

Aren't we reliant upon this synchronisation for correctness?

>   LRU isolation                -> nothing: sources are inert in place
>
> Working in windows rather than whole PMDs takes care of the other root.
> A sub-PMD window is collapsed under the page table lock, so a collapse
> disturbs only the window it collapses, and each candidate is validated
> at its own order -- a window need only fit its own VMA.  A PMD-order
> candidate still has to own the whole PMD, which is the old rule kept
> where it is still needed.
>
> Candidates are carried through the passes a batch at a time rather than
> one window at a time, so a round pays for its flush and its lock
> acquisitions once.
>
> With the barriers holding the sources still, which read lock the engine
> takes stops being part of the design.  A round works inside a single
> VMA, so patches 43-49 switch it from mmap_read to per-VMA locking: an
> mmap_write elsewhere in the mm then stops waiting for a collapse that
> has nothing to do with it.  That block is the only part of the series
> that needs per-VMA locking to be unconditional, and it is a separate
> dependency (see below); everything before it runs under mmap_read and
> does not care.
>
> Patch 7 sketches the engine as a comment naming every pass, what lock it
> takes and what it may sleep on; the details are there rather than here.
>
> What falls out beyond the lock diet:
>
>  - mTHP collapse in VMAs smaller than a PMD, which is the arm64 case
>    above: a 2M VMA on an arm64/64K machine collapses nothing today at
>    any order, and collapses to mTHP here.

I do worry about knock-ons from this. Has to be carefully checked.

>  - Hole and zeropage population at every order, so partially populated
>    windows collapse to mTHP under the same max_ptes_none policy as PMD.
>  - Sources come in spans -- any stretch of consecutive PTEs mapping
>    consecutive pages of one folio -- so partially mapped and scrambled
>    compound sources (the PTE-mapped-THP re-collapse class) work at
>    every order.
>  - A table that cannot become one huge page still yields the largest
>    windows inside it, where before a single disqualified PTE gave up
>    the whole table.

Are you permitting collapse of ranges that straddle PTEs?

Though in general I'm confused by the single disqualified PTE here -

>
> Reading the series
> ==================
>
> 57 patches is a lot to land on a list.  They go in blocks:
>
>   1-6    helpers and shared state: pte_folio(), pte_none_or_zero(),
>          mm/collapse.h, and the policy that replaces asking whether
>          khugepaged started a collapse
>   7-8    the engine's shape: entry points, a call-tree comment naming
>          every pass, and the scan filled in
>   9-23   the collapse half, top down: the round frame, then each pass
>          in turn, then selection and the retry store
>   24     per-candidate tracing, before the switch takes the old
>          tracepoints away
>   25-28  the switch: point the anon path at the engine, widen coverage
>          to sub-PMD VMAs, delete the mechanism it replaces
>   29-35  move what is left of collapse out of khugepaged.c, and
>          MADV_COLLAPSE into madvise.c
>   36-42  tracing: the engine's own events and trace header
>   43-49  per-VMA locking, and the mm reference that makes it safe
>   50-56  selftests for what the engine can now do
>   57     MAINTAINERS

Yeah as I replied there the MAINTAINERS change is totally unacceptable :)

This is THP code and belongs to the THP maintainers/reviewers.

No coup d'etat please :)

Also reviewership/maintainership in this area is predicated on significant
review contributions which are badly needed in THP. So effort in this
respect is appreciated too :)

And come to a THP cabal meeting to discuss your changes with us please!
We're friendly :)

>
> The two patches worth reading first if you read nothing else are 7 (the
> design, as a comment naming the whole call tree) and 16 (the freeze,
> which is where the safety argument lives).
>
> A possible split, if that helps:
>
>   1-2    two mm helpers, pte_folio() and pte_none_or_zero().  Both
>          convert callers outside collapse and are useful on their own
>   3-27   the engine and the switch-over.  This is the smallest unit
>          that does anything: stop earlier and the tree carries an
>          engine nothing calls
>   28     remove the mechanism the engine replaces
>   29-42  moving what is left of collapse out of khugepaged.c, and the
>          engine's own tracepoints
>   43-49  per-VMA locking
>   50-57  selftests and MAINTAINERS
>
> Keeping the removal separate leaves both engines in the tree with only
> the new one reachable, so the switch can be reverted on its own if
> something turns up.  The old mechanism is already carried that way for
> three patches inside the series, so this costs nothing but 975 lines of
> unreferenced code until 28 lands.  That safety net only lasts until the
> blocks after it land, though: once collapse has moved out of
> khugepaged.c and the locking has changed, reverting the switch no longer
> gives back a working old engine.

Again as above, no to these changes in any form at the moment.

The techical debt must be paid down first.

>
> Base and dependencies
> =====================
>
> This applies on the selftests series, not on plain mm-new:
>
>   [PATCH v4 00/19] selftests/mm: improve khugepaged coverage
>   https://lore.kernel.org/all/20260815015901.1236937-1-kirill@shutemov.name/
>
> which is on mm-new 33f61b12d297.
>
> Patches 43-49 depend on Suren's unconditional per-VMA locks:
>
>   [PATCH v6 0/5] mm: Unconditional per-VMA locks and cleanups
>   https://lore.kernel.org/all/20260813193433.3318288-1-surenb@google.com/
>
> That series is not in mm-new yet, and with patch 46 applied SMP=n does
> not build without it: lock_next_vma() is behind CONFIG_PER_VMA_LOCK in
> mmap_lock.h.  Everything up to patch 42 builds and runs on mm-new as it
> stands.  There is no fallback path by choice -- adding one would mean
> carrying two locking models through every pass.
>
> Both branches are available at
>
>   git://git.kernel.org/pub/scm/linux/kernel/git/kas/linux.git collapse/rfc-v1
>
> and the benchmark used for the numbers below, which is unposted and not a
> dependency, at
>
>   git://git.kernel.org/pub/scm/linux/kernel/git/kas/linux.git perf/bench-usemem

As above.

>
> Performance
> ===========
>
> Measuring khugepaged is awkward.  It is a background daemon, so what
> matters is what a workload feels while it runs, not what the daemon
> reports about itself -- and the usual coverage instrument is no help
> below the PMD: smaps AnonHugePages only counts PMD-order folios, so it
> reads zero however much mTHP has been collapsed.
>
> So I wrote "perf bench mem usemem" for this.  It touches a region while
> khugepaged works on it and reports the workload's own latency
> percentiles and throughput, against per-size counters that can see
> sub-PMD folios.  The branch is above; it is unposted and not a
> dependency.
>
> x86-64, production configs (no KASAN, no lockdep, no DEBUG_VM, no
> PAGE_TABLE_CHECK), interleaved rounds on an idle host, equal work on
> every arm.
>
> I measured three kernels, so the two halves of the series can be told
> apart in the numbers below:
>
>   A   the base
>   B   the new engine, still under mmap_read
>   C   B plus per-VMA locking -- what this series ends up with
>
> The engine: a sub-PMD collapse stops blanking the surrounding 2M
> -----------------------------------------------------------------
>
> base routes sub-PMD collapse through collapse_huge_page(), whose
> pmdp_collapse_flush() and tlb_remove_table_sync_one() are not gated on
> order: to collapse an order-4 window of 16 pages it clears and flushes
> the whole 512-page PMD and IPIs, then repopulates.  The engine does the
> window under the PTL.
>
> A thread reading and writing a 32G region while it is collapsed at
> order-4, 16384 collapses on every arm:
>
>                          A        B        C
>   read p99 (ns)       3071     1023     1023   -66.7%
>   write p99 (ns)      3071      927      927   -69.8%
>
> and the workload's read rate rises by 68% on both engine kernels.
>
> B == C, so this is the engine, not the locking.
>
> At PMD order the same workload is flat, and that is expected rather
> than disappointing: it is the one configuration where both mechanisms
> disturb exactly the same 2M.  Read it as no regression at PMD order.
>
> Per-VMA locking: address-space operations stop waiting on the scan
> -------------------------------------------------------------------
>
> MADV_HUGEPAGE/MADV_NOHUGEPAGE toggling against a scanning mm, which is
> what jemalloc does with its arenas.  4096 collapses on every arm:
>
>                          A        B        C
>   ops/sec           340656   340820   603305   +77%
>   p99 (ns)           77823    86015     3327   -96%
>   p99.9 (ns)         86015    94207     9215   -89%
>
> B is about 11% worse than base at p99 here, consistently across runs:
> the engine alone slightly worsens hint-toggle latency, and per-VMA
> locking is what turns it into a win.  Both halves are in this series, so
> C is what a reviewer gets, but the middle column is the honest one.
>
> The trade is real in the other direction too.  On settled memory with
> nothing to collapse and scan_sleep_millisecs=0, per-VMA locking costs
> about 47% of scan throughput against one mmap_read for the whole walk.
> That is a synthetic worst case -- the daemon wraps 8000 times a second
> there, where production defaults to 10s between passes -- and it buys
> mmap/munmap p99 of 56us against 1.4us.
>
> Collapse itself is not slower
> -----------------------------
>
> One complete pass over a 32G region, 16384 collapses, khugepaged CPU
> from /proc/<pid>/stat, 7 repeats:
>
>   A base    median 5760 ms   spread 12.7%
>   B engine  median 4990 ms   spread  3.8%
>   C pervma  median 5020 ms   spread 12.4%
>
> The base arm is bimodal, so its median moves with sampling and the
> percentage is soft.  The distribution-free statement is better: every
> engine run used less CPU than every base run.
>
> The engine also allocates one destination per folio installed, where
> base allocates 5.25 and frees the rest again: nothing is allocated until
> the sources are frozen and the collapse can no longer be refused.
>
> A measurement note, since an earlier version of this series quoted worse
> figures.  khugepaged CPU has to be measured per collapse or per
> completed pass, never over a fixed window with scan_sleep_millisecs=0:
> the daemon never sleeps, so whichever kernel finishes the work sooner
> spends the rest of the window scanning settled memory and is charged for
> it.  Measured that way the engine appeared to cost 10% more CPU;
> measured per unit of work it costs less.
>
> Costs
> =====
>
> At PMD order the engine issues two TLB flushes per collapse where the
> old mechanism issues one: the freeze's ranged flush plus the terminal
> layer's pmdp_collapse_flush().  A PMD candidate is alone in its round,
> so nothing amortizes the first.  Dropping the old per-collapse
> tlb_remove_table_sync_one() IPI presumably pays for it, but that was not
> measured and is not claimed here.
>
> There may be a way out -- a PMD migration entry over the table during
> the window, so the CPU never caches a walk to shoot down -- but that
> means teaching every pmd-level walker a new kind of entry, and I have
> not tried it.

Hmm this seems like complexity on top of complexity...

>
> PMD collapse deposits a freshly allocated page table instead of
> redepositing the detached one.  Whoever withdraws a deposited table
> frees it immediately, with nothing to hold a lockless walker off first,
> and under a read lock the detached table may still be traversed by
> GUP-fast or an RCU pte walk.  It goes to pte_free_defer() instead,
> exactly as retract_page_tables() does.  One transient table page per PMD
> collapse buys the IPI's absence.
>
> That cost goes away if zap_deposited_table() -- the only site that frees
> a deposited table outright, the others redeposit it or repopulate the
> PMD with it -- used pte_free_defer().  The deposit would no longer have
> to be quiescent and the detached table could go straight back.  It would
> defer every THP zap's table free, and I have not tried it.
>
> A shared source now costs an extra copy.  The freeze needs every page
> exclusive to this mm, so the fault-in pass breaks CoW first -- an
> allocation and a copy -- and the collapse then copies that page into the
> destination; the old mechanism copied a shared page straight into the
> new folio and broke the sharing that way.  It is bounded by
> max_ptes_shared, which khugepaged holds at zero below the PMD order, so
> in practice this is PMD-order collapse and MADV_COLLAPSE.

Hmm I guess better to be explicit.

>
> Size
> ====
>
> mm/ grows by 1915 lines net: 4475 added against 2560 deleted.
>
> That is not a claim that this is less code, but it is less than it
> looks.  khugepaged.c goes from 3283 lines to 908.  The new engine is
> 4052 lines across mm/collapse.c and mm/collapse.h, of which 1344 --
> about a third -- are comments, which is where the pipeline's invariants
> are written down.  What replaces three install paths with their own
> isolate/copy/rollback is one engine and one contract.

As above. The existing codebase is a mess and must be cleaned up first,
this isn't optional.

>
> Testing
> =======
>
> Both matrices run the mm selftests plus a race harness, on the
> validation config: KASAN, lockdep, PROVE_LOCKING, DEBUG_VM and
> PAGE_TABLE_CHECK, 16G of guest memory, swap active so the swap-in
> prepass is exercised rather than skipped.
>
>   x86-64        433 pass, 0 fail, 12 skip
>   arm64/64K     581 pass, 0 fail, 18 skip
>
> dmesg clean on both.  The arm64 skips are a pre-existing shmem
> MADV_COLLAPSE -EINVAL on 64K pages, confirmed against the base by A/B.
>
> Every one of the 57 patches builds with no new warnings; !NUMA and !MMU
> (arm nommu) build clean.  SMP=n does not build, for the reason in the
> dependencies section above.
>
> The race harness also gets longer soaks -- 1800s per driver mode, with
> memory pressure and swap -- and the engine is fuzzed with syzkaller on a
> KCOV+KASAN build.  That found two bugs the selftests could not reach: a
> teardown that dropped rmap while the source was still frozen, where
> removing an mlocked mapping munlocks and munlock_folio() takes a
> reference a frozen folio forbids; and a whole-table MADV_DONTNEED racing
> the copy window under CONFIG_PT_RECLAIM, which freed the table and left
> the sources frozen and locked.  Both are fixed, and both gained coverage
> -- the mlocked case is patch 53.

Thanks for the detailed explanation.

>
> Kiryl Shutsemau (Meta) (57):
>   mm: add pte_folio()
>   mm: add pte_none_or_zero()
>   mm/collapse: add collapse.h for the shared collapse state
>   mm/collapse: rename mthp_present_ptes to eligible_ptes
>   mm/collapse: state what a collapse may do in the policy
>   mm/collapse: move the smallest collapse order to collapse.h
>   mm/collapse: sketch the new anonymous collapse engine
>   mm/collapse: scan a table for what a collapse could use
>   mm/collapse: collect candidate windows into a round
>   mm/collapse: run a round and feed the outcomes back
>   mm/collapse: sketch the passes of a round
>   mm/collapse: allocate a destination per candidate
>   mm/collapse: revalidate a round against the VMA
>   mm/collapse: fault the sources in before the freeze
>   mm/collapse: check what a candidate would freeze
>   mm/collapse: freeze the sources behind migration entries
>   mm/collapse: copy the sources into the destinations
>   mm/collapse: install the destinations at PTE level
>   mm/collapse: install a PMD leaf as the terminal layer
>   mm/collapse: put the sources back
>   mm/collapse: settle whatever the round reached
>   mm/collapse: walk a table with a selection cursor
>   mm/collapse: give a refused region a second chance
>   mm/collapse: report each candidate's outcome to tracing
>   mm/collapse: collapse anonymous memory with the new engine
>   mm/collapse: give collapse_single_pmd() the range to work on
>   mm/collapse: scan the windows a VMA can hold
>   mm/collapse: remove the mechanism the engine replaces
>   mm/collapse: move what a collapse is judged on into collapse.c
>   mm/collapse: name the max_ptes ceiling after collapse
>   mm/khugepaged: count collapses where khugepaged makes them
>   mm/collapse: move the file collapse into collapse.c
>   mm/collapse: split collapse into a scan and a run
>   mm/collapse: implement MADV_COLLAPSE in madvise.c
>   mm/madvise: drop MADV_COLLAPSE's redundant mm reference
>   mm/collapse: report what the scan found
>   mm/collapse: report what the fault-in pass paid
>   mm/collapse: report the round, and what it made faulters wait
>   mm/collapse: name the file collapse's tracepoints after collapse
>   mm/collapse: remove the tracepoints of the mechanism that is gone
>   mm/collapse: give collapse its own trace header
>   mm/collapse: allow error injection into the freeze
>   mm/khugepaged: check the scan budget before the work, not after
>   mm/khugepaged: hold the address space open across a scan
>   mm/collapse: take a per-VMA read lock for the round
>   mm/khugepaged: scan under a per-VMA read lock
>   mm/madvise: collapse under a per-VMA read lock
>   mm/collapse: assert the mm reference the engine relies on
>   mm/khugepaged: drop the mmap_lock barrier from __khugepaged_exit()
>   selftests/mm: attribute collapses by candidate event alone
>   selftests/mm: cover collapse inside a sub-PMD VMA
>   selftests/mm: cover a hole-y window in a sub-PMD VMA
>   selftests/mm: cover collapse of mlocked ranges
>   selftests/mm: cover collapse beside a MADV_FREE'd page
>   selftests/mm: cover collapse beside a pinned page
>   selftests/mm: cover the scaled max_ptes_shared limit
>   MAINTAINERS: add an entry for collapse
>
>  MAINTAINERS                                   |   19 +-
>  fs/proc/task_mmu.c                            |    4 +-
>  include/linux/huge_mm.h                       |    9 -
>  include/linux/mm.h                            |   14 +
>  include/linux/pgtable.h                       |   17 +
>  .../events/{huge_memory.h => collapse.h}      |  176 +-
>  kernel/bpf/btf.c                              |    8 +-
>  mm/Makefile                                   |    2 +-
>  mm/collapse.c                                 | 3808 +++++++++++++++++
>  mm/collapse.h                                 |  244 ++
>  mm/hugetlb.c                                  |    8 +-
>  mm/khugepaged.c                               | 2711 +-----------
>  mm/madvise.c                                  |  251 +-
>  mm/migrate_device.c                           |    9 +-
>  mm/mremap.c                                   |    2 +-
>  tools/testing/selftests/mm/khugepaged.c       |  346 ++
>  tools/testing/selftests/mm/khugepaged_race.c  |   29 +-
>  .../selftests/mm/khugepaged_sync_check.c      |   65 +-
>  tools/testing/selftests/mm/vm_util.c          |    2 +-
>  19 files changed, 5022 insertions(+), 2702 deletions(-)
>  rename include/trace/events/{huge_memory.h => collapse.h} (60%)
>  create mode 100644 mm/collapse.c
>  create mode 100644 mm/collapse.h
>
>
> base-commit: 8b76faf42c5d342b3bf0b1fd97bdaf603ee57354
> --
> 2.54.0
>

--
Cheers, Lorenzo

[0]:https://docs.google.com/document/d/1dsAXvVtioR2Hj98E_Um6220jamdznWxYsyx8jO0aUGU/edit?usp=sharing
Re: [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives
Posted by Kiryl Shutsemau 1 month, 1 week ago
On Mon, Aug 17, 2026 at 09:52:13AM +0100, Lorenzo Stoakes (ARM) wrote:
> We have a THP cabal meeting every couple of weeks where it would have been
> useful for you to raise this first.

Fair -- though my invite is on my old @linux.intel.com address.  Could
you forward it to kas@kernel.org?

> In any case - this series is not something we'd consider at the moment,
> even broken into parts.
> 
> David and I have put THP into feature freeze - until the codebase is
> subtantially improved we're not really interested in seeing significant
> development.
> 
> The technical debt is substantial and has to be paid down first.
> 
> See [0] for a rough list of TODOs in this regard.

I read the TODO list and I'll pick from it -- though I notice the
technical debt section includes "Literally all of the code in
mm/huge_memory.c and mm/khugepaged.c", which I'd argue this series is a
fairly committed attempt at :)

One clean up I wanted to do is consolidate code by functionality, not by
the THP/non-THP split.  Move all page fault handler code into mm/fault.c,
unmap code into mm/zap.c, fork's copying into mm/fork.c -- mirroring
kernel/fork.c, so the mm half of a subsystem sits under the same name.
Large folios are an integral part of mm nowadays and I don't think we
benefit from keeping THP in a separate file.  It is also an opportunity to
shift away from mm/memory.c being a kitchen sink.

David and I talked about this at LSF/MM.  I can give it a try if it fits
your idea of "feature freeze" -- and if it doesn't collide with the series
you have in flight, in which case I'd rather go after yours than around it.

> > Why
> > ===
> >
> > mTHP collapse landed in khugepaged in 7.2 and I was glad to see it.  We
> > at Meta run arm64 with 64K base pages, where a PMD is 512M: PMD-order THP
> > is of limited use at that size, and mTHP is exactly what we want.
> 
> Do you have some numbers that indicate to what degree mTHP khugepaged is
> beneficial?

Not from the fleet yet -- that experiment is still ahead of me, so I can't
give you order-by-order numbers.

What I can say is that on x86 we lean on khugepaged heavily to get THPs in
place; it is not a marginal contributor for us.  On arm64 with 64K base
pages we get nowhere near the x86 numbers without khugepaged being able to
produce mTHP at all.

I'll grant the other half of it, and more strongly than you put it: for a
64K mTHP on a 4K base page the TLB win is modest, and with today's
mechanism -- which clears and flushes the whole 2M PMD to install it -- I
can believe the disruption exceeds the gain and the net effect on a
workload is negative.  That is what I measured: a thread reading and
writing a region while khugepaged collapses it at order-4, read p99 3071ns
against 1023ns.  It shows the disruption is real and that it comes down;
whether the collapse pays for itself at that order is a separate question.

> > khugepaged only ever looks at PMD-aligned windows, and it is not an easy
> > limitation to lift.
> 
> Yes. This assumption is very much baked in.
> 
> I guess this is coming from the perspective of having ranges that are
> neither PMD-aligned nor sized (far harder to achieve with 512 MiB PMD size
> obviously).

Right, and it is the size rather than the alignment that bites.  The real
requirement is that a VMA contain a whole PMD-aligned, PMD-sized range,
which at 2M most anonymous mappings of any size manage and at 512M almost
none do.  That is why the limitation was easy to miss until the PMD got
big.

> > Fixing the alignment is a one-line change, but what it feeds assumes the
> 
> Hmm not so sure about that... especially given how baked in these
> assumptions are.

I think we agree -- that was the setup, not the claim.  Dropping the
ALIGN() is the one line; the point of the sentence is that it buys nothing
on its own, because what you then hand a sub-PMD range to still clears the
whole PMD and still demands the VMA span it.

> Which by the way, all speaks to the need for rework.
> 
> The first stage in my view would be to improve the code to the point that
> these kinds of assumptions fall out of it, which then lays the foundations
> for future changes to eliminate the assumptions.

For the plumbing, yes -- policy, file layout, the scan/run split all
improve by refactoring in place.

I don't think the locking model gets there that way, though, which is why
I built a second engine rather than morphing the first.  The old safety
argument is "hold the address space still"; the new one is "make the
sources inert".  They are not two points on a line -- mmap_write cannot go
before something else holds the sources still, and doing that inside
collapse_huge_page() means replacing the copy, the install and the rollback
at the same time, which is the whole function.  Every halfway state has
neither argument in full.

What is gradual here is the review rather than the mechanism: the engine
arrives one pass at a time, each reviewable alone, with the old one live
until one patch switches over.

> > PMD everywhere that matters: collapse_huge_page() clears and flushes the
> > whole PMD whatever order it is collapsing, installs a PMD leaf because
> > that is the only thing it can produce, and keeps everyone out with
> > mmap_write_lock, anon_vma_lock_write() and an IPI broadcast while it
> > does.
> 
> To be clear - the anon path. I think important to clarify :)

Yes, anon only.  The file path moves into collapse.c and picks up the
scan/run split, but its mechanism is untouched.

And you are right that "installs a PMD leaf" is wrong above: only a
PMD-order collapse installs one, a sub-PMD collapse repopulates the table
with PTEs.  What is order-blind is everything around it -- the PMD is
still cleared and flushed first.

I would like to bring it into the engine as well, and with it the private
copies in MAP_PRIVATE file mappings, which neither path collapses today --
the anon side requires vma_is_anonymous() and the file side works on the
page cache.

I stopped because I wanted to keep the patch count in double digits. :P

> And yeah it does IPI for any sensible arch (with
> CONFIG_MMU_GATHER_RCU_TABLE_FREE) via tlb_remove_table_sync_one(). The
> other arches IPI anyway on TLB invalidation.
> 
> [Though I intend to make all page table freeing RCU relatively soon which
> should? Eliminate the need for this, possibly?]

That would suit this well, and it is worth covering deposited page tables
in it if they are not already in scope.  Today PMD collapse has to deposit
a freshly allocated table rather than redepositing the one it detached,
because a deposited table must be safe for zap_deposited_table() to free
immediately.  Make that free RCU-deferred and the detached table can go
straight back -- one allocation less per PMD collapse.

> However we have to remember that a lot of the user-visible API assumes PMD
> sizing and so the code has to clearly reflect this and make it clear that

Your sentence got cut off, but if this is about the tunables then let me
flag what the series does with them.  max_ptes_none, _swap and _shared are
counts out of a PMD, and a range smaller than one has fewer PTEs than the
budget, so a raw comparison can never refuse it -- a 64K range on 4K pages
is 16 PTEs against a max_ptes_shared default of 256.  For swap and shared
the engine therefore compares fractions: count * HPAGE_PMD_NR against
max * nr_scanned.

max_ptes_none stays as mTHP collapse has it, all-or-nothing: 0, or
everything at that order.  That is deliberate -- allowing holes at an order
below the largest enabled one lets khugepaged fill them and collapse the
result at the next order up, which is the ratchet max_ptes_none exists to
bound.  There is room to scale it at the terminal order, where there is no
larger order to creep into, but I have not done that here.

> > Design
> > ======
> >
> > The old mechanism holds the address space still because it has nothing
> > else stopping the sources from moving under the copy.  The new engine
> > makes the sources themselves inert instead, with the two barriers
> > migration already uses, raised in that order:
> >
> >   1. migration entries replace the source PTEs.  Faults and GUP-slow
> >      now wait on the source folio's lock, which is taken before the
> >      first entry becomes visible.
> >   2. the source folio's refcount is frozen to its expected value.
> >      GUP-fast, pfn walkers, reclaim, compaction and memory-failure all
> >      fail folio_try_get() and back off.
> >
> > Between the two, nothing can reach a source, so the copy runs with no
> > lock held at all -- and the address space is left alone while it does.
> 
> Hmm, are migration entries the right mechnanism here? Are you actually
> migrating the pages to a large folio here, or using them to get the
> behaviour you want on fault/GUP?

Both, and I would argue the behaviour is not a side effect: what a
migration entry means to a waiter -- this page is going away, sleep on its
folio lock and look again -- is exactly true of a source under collapse.
Fault, GUP-slow and rmap then all do the right thing with no new code,
which is the case for reusing the entry rather than inventing a marker
every waiter would have to learn.

What is not reused is mm/migrate.c.  A migration entry encodes one PFN, so
it cannot name an N:1 destination of a different order; the engine takes
the hold-still half and does the remap itself at install.  It is a
migration in substance -- contents move to another folio, the old mappings
are replaced -- but not one migrate.c could drive.

> Same question in general for the freezing.

folio_ref_freeze() means nobody may take a new reference, which is the
property the copy needs, and is why migration and split use it too.

> > What that removes from every collapse path:
> >
> >   mmap_write_lock              -> mmap_read
> >   anon_vma_lock_write()        -> nothing: an rmap walk needs the folio
> >                                   locked, and the engine holds that lock
> >                                   from freeze to putback
> 
> I do like the idea of eliminating uses of the rmap lock like this, not only
> for contention's sake but also for scalable CoW purposes which introduces
> challenges with regards to holding these.
> 
> In fact, migration and huge memory collapse are the really problematic
> areas.

If that is about their rmap complexity, collapse gets easier here rather
than harder.

The engine takes no rmap lock and walks no rmap.  A page shared with
another process is unshared before anything is frozen -- the fault-in
pass breaks CoW, exactly as a write would -- so by freeze time a source
is exclusive to this mm and its only live mappings are the ones being
replaced.  Migration has to cope with a folio mapped from many mms; this
never sees one.

If what scalable CoW needs is that collapse stops messing with rmap,
that is what this does.

> 
> >   tlb_remove_table_sync_one()  -> nothing: one ranged flush per round
> 
> Aren't we reliant upon this synchronisation for correctness?

Not in the new engine.

In collapse_huge_page() the IPI is what makes the *detached* table safe
to use: it pmdp_collapse_flush()es the PMD and then copies out of the
table it just detached, so it has to know no lockless walker is still
inside it.

The new engine never does that -- a sub-PMD window is collapsed in place
under the page table lock, so nothing is detached, and at PMD order the
table is detached, never touched again, and freed with pte_free_defer().
The synchronisation is still there, it is RCU rather than an IPI; and on
the arches without RCU table free the flush itself IPIs, as you say.

> >  - A table that cannot become one huge page still yields the largest
> >    windows inside it, where before a single disqualified PTE gave up
> >    the whole table.
> 
> Are you permitting collapse of ranges that straddle PTEs?

No -- a candidate never crosses a page table or a VMA.  A round works
within one table, and each candidate is validated against the VMA at its
own order.

> Though in general I'm confused by the single disqualified PTE here -

Taking that literally is how I meant it: collapse_scan_pmd() goto
out_unmap's on the first PTE that fails any of its checks -- uffd, non-anon,
clean lazyfree, off the LRU, unexpected refcount -- and mthp_collapse() only
runs if the verdict came back SCAN_SUCCEED.  So one
such PTE anywhere in the table means nothing in that table collapses, at
any order, even at an order whose windows avoid it entirely.

> > There may be a way out -- a PMD migration entry over the table during
> > the window, so the CPU never caches a walk to shoot down -- but that
> > means teaching every pmd-level walker a new kind of entry, and I have
> > not tried it.
> 
> Hmm this seems like complexity on top of complexity...

I find it rather elegant, and expect it to be minimally intrusive: let such
a PMD be walkable exactly as a present one is, so software descends through
it as usual while the CPU sees a non-present entry and caches nothing.
Transparent to software, opaque to the CPU.

Out of scope for this patchset either way.

-- 
  Kiryl Shutsemau / Kirill A. Shutemov
Re: [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives
Posted by David Hildenbrand (Arm) 1 month, 1 week ago
On 8/17/26 15:38, Kiryl Shutsemau wrote:
> On Mon, Aug 17, 2026 at 09:52:13AM +0100, Lorenzo Stoakes (ARM) wrote:
>> We have a THP cabal meeting every couple of weeks where it would have been
>> useful for you to raise this first.
> 
> Fair -- though my invite is on my old @linux.intel.com address.  Could
> you forward it to kas@kernel.org?
> 
>> In any case - this series is not something we'd consider at the moment,
>> even broken into parts.
>>
>> David and I have put THP into feature freeze - until the codebase is
>> subtantially improved we're not really interested in seeing significant
>> development.
>>
>> The technical debt is substantial and has to be paid down first.
>>
>> See [0] for a rough list of TODOs in this regard.
> 
> I read the TODO list and I'll pick from it -- though I notice the
> technical debt section includes "Literally all of the code in
> mm/huge_memory.c and mm/khugepaged.c", which I'd argue this series is a
> fairly committed attempt at :)
> 
> One clean up I wanted to do is consolidate code by functionality, not by
> the THP/non-THP split.  Move all page fault handler code into mm/fault.c,
> unmap code into mm/zap.c, fork's copying into mm/fork.c -- mirroring
> kernel/fork.c, so the mm half of a subsystem sits under the same name.
> Large folios are an integral part of mm nowadays and I don't think we
> benefit from keeping THP in a separate file.  It is also an opportunity to
> shift away from mm/memory.c being a kitchen sink.
> 
> David and I talked about this at LSF/MM. 

Ah, I missed the context in my other reply. Lorenzo already had some patches at
some point to split up mm/memory.c into better chunks that will also better help
our subcomponent maintenance model.

I think the challenge is how to handle huge_memory.c, because ideally, we'd not
have these stupid callbacks into huge_memory.c once we make PMDs just a
first-class citizen.

This is, unfortunately, also something that needs more thought, because we don't
want to end up moving stuff back and forth.

-- 
Cheers,

David
Re: [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives
Posted by Kiryl Shutsemau 1 month, 1 week ago
On Tue, Aug 18, 2026 at 04:15:27PM +0200, David Hildenbrand (Arm) wrote:
> > David and I talked about this at LSF/MM. 
> 
> Ah, I missed the context in my other reply. Lorenzo already had some patches at
> some point to split up mm/memory.c into better chunks that will also better help
> our subcomponent maintenance model.

Lorenzo, if you can dig those out, I would rather build on them than start
over.  Happy to do the legwork.

> I think the challenge is how to handle huge_memory.c, because ideally, we'd not
> have these stupid callbacks into huge_memory.c once we make PMDs just a
> first-class citizen.

Agreed on settling the target shape before moving anything.

Splitting by operation answers most of huge_memory.c on its own, though.

Its entry points are named after the operation they implement, and each has
one destination: do_huge_pmd_anonymous_page(), do_huge_pmd_wp_page() and
do_huge_pmd_numa_page() belong next to do_anonymous_page() in
mm/fault.c.

zap_huge_pmd() next to zap_pte_range(), copy_huge_pmd() next to
copy_pte_range().  change_huge_pmd(), move_huge_pmd() and follow_huge_pmd()
have their callers in the right file already -- mprotect.c, mremap.c, gup.c
-- so they move to the caller.  The callback goes away because the PMD case
becomes a branch beside the PTE case instead of a call into another file.

-- 
  Kiryl Shutsemau / Kirill A. Shutemov
Re: [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives
Posted by Lorenzo Stoakes (ARM) 1 month, 1 week ago
On Tue, Aug 18, 2026 at 04:15:27PM +0200, David Hildenbrand (Arm) wrote:
> On 8/17/26 15:38, Kiryl Shutsemau wrote:
> > On Mon, Aug 17, 2026 at 09:52:13AM +0100, Lorenzo Stoakes (ARM) wrote:
> >> We have a THP cabal meeting every couple of weeks where it would have been
> >> useful for you to raise this first.
> >
> > Fair -- though my invite is on my old @linux.intel.com address.  Could
> > you forward it to kas@kernel.org?
> >
> >> In any case - this series is not something we'd consider at the moment,
> >> even broken into parts.
> >>
> >> David and I have put THP into feature freeze - until the codebase is
> >> subtantially improved we're not really interested in seeing significant
> >> development.
> >>
> >> The technical debt is substantial and has to be paid down first.
> >>
> >> See [0] for a rough list of TODOs in this regard.
> >
> > I read the TODO list and I'll pick from it -- though I notice the
> > technical debt section includes "Literally all of the code in
> > mm/huge_memory.c and mm/khugepaged.c", which I'd argue this series is a
> > fairly committed attempt at :)
> >
> > One clean up I wanted to do is consolidate code by functionality, not by
> > the THP/non-THP split.  Move all page fault handler code into mm/fault.c,
> > unmap code into mm/zap.c, fork's copying into mm/fork.c -- mirroring
> > kernel/fork.c, so the mm half of a subsystem sits under the same name.
> > Large folios are an integral part of mm nowadays and I don't think we
> > benefit from keeping THP in a separate file.  It is also an opportunity to
> > shift away from mm/memory.c being a kitchen sink.
> >
> > David and I talked about this at LSF/MM.
>
> Ah, I missed the context in my other reply. Lorenzo already had some patches at
> some point to split up mm/memory.c into better chunks that will also better help
> our subcomponent maintenance model.

Ah yeah I kinda lost those but indeed doing this is a good idea.

I should try to dig those out again or look again when I have a chance...

>
> I think the challenge is how to handle huge_memory.c, because ideally, we'd not
> have these stupid callbacks into huge_memory.c once we make PMDs just a
> first-class citizen.

Yes and another point and I think it's one that Kiryl also gets at is - can we
_please_ stop pretending huge folios == THP == what the page cache does == page
special cases like DAX? :)

I think Matthew had ideas (TM) about separating out CONFIG_THP from that stuff
but that's another way in which we can unwend some of the horrors.

>
> This is, unfortunately, also something that needs more thought, because we don't
> want to end up moving stuff back and forth.

Yeah annoyingly it's not all straightforward. I wonder if actually gauging off
of what series often touch in patches might be a good way of figuring out how to
separate, funnily enough (certainly for mm-next conflict resolution purposes
anyway).

>
> --
> Cheers,
>
> David

--
Cheers, Lorenzo
Re: [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives
Posted by Kiryl Shutsemau 1 month, 1 week ago
On Tue, Aug 18, 2026 at 03:41:40PM +0100, Lorenzo Stoakes (ARM) wrote:
> > Ah, I missed the context in my other reply. Lorenzo already had some patches at
> > some point to split up mm/memory.c into better chunks that will also better help
> > our subcomponent maintenance model.
> 
> Ah yeah I kinda lost those but indeed doing this is a good idea.

Oh, well, I can give it a try from scratch.

> I should try to dig those out again or look again when I have a chance...
> 
> >
> > I think the challenge is how to handle huge_memory.c, because ideally, we'd not
> > have these stupid callbacks into huge_memory.c once we make PMDs just a
> > first-class citizen.
> 
> Yes and another point and I think it's one that Kiryl also gets at is - can we
> _please_ stop pretending huge folios == THP == what the page cache does == page
> special cases like DAX? :)

Yes.  That is also why I want the split by operation rather than by
THP-ness: a PMD case belongs next to the PTE case for the same operation,
not in a file that collects everything huge.

-- 
  Kiryl Shutsemau / Kirill A. Shutemov
Re: [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives
Posted by Lorenzo Stoakes (ARM) 1 month, 1 week ago
(I'll reply to the rest of this later)

On Mon, Aug 17, 2026 at 02:38:44PM +0100, Kiryl Shutsemau wrote:
> David and I talked about this at LSF/MM.  I can give it a try if it fits
> your idea of "feature freeze" -- and if it doesn't collide with the series
> you have in flight, in which case I'd rather go after yours than around it.

Whether intended or not you sound rather like you are trying to override me
in favour of my co-maintainer here and it's... not helpful.

David and I co-maintain THP together, are in constant communication, and
have a great working relationship :)

IOW - if one of us states a position on the sub(sub?)system - then take
that to be the actual position.

I have poured what must be hundreds of hours now into THP maintainership -
it's by far my biggest workload on the maintenance front, by far the most
painful and by far the most thankless.

I do it because I care about mm a great deal and am, frankly, driven by a
desire to see THP turn from a flaming trash pile of a code base with
confusing semantics and many, many broken parts into something that serves
the community's needs with far less maintenance burden.

Looking over your series it seems some of the patches works in this
direction (great!), but much else of it fundamentally changes key
behaviour.

So it's just a question of deferring the latter until we get to a sane
point with the former.

I keep talking about this stuff because companies are motivated by wanting
to solve their problems (understandably) but if nobody pushes back then THP
will continue to be a series of changes lumped on top of one another adding
more and more technical debt.

--
Cheers, Lorenzo
Re: [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives
Posted by David Hildenbrand (Arm) 1 month, 1 week ago
On 8/18/26 15:06, Lorenzo Stoakes (ARM) wrote:
> (I'll reply to the rest of this later)
> 
> On Mon, Aug 17, 2026 at 02:38:44PM +0100, Kiryl Shutsemau wrote:
>> David and I talked about this at LSF/MM.  I can give it a try if it fits
>> your idea of "feature freeze" -- and if it doesn't collide with the series
>> you have in flight, in which case I'd rather go after yours than around it.
> 
> Whether intended or not you sound rather like you are trying to override me
> in favour of my co-maintainer here and it's... not helpful.

I think Kiryl tried to say that we discussed at LSF/MM which areas of MM scream
for an improvement, and we discussed that khugepaged is just horrible code.

I think we all agree that there is a lot of room for improvement, but the big
question is:

(a) When does it stop being a cleanup and is a new feature in disguise that
    makes the code more complicated and even harder to maintain.

(b) Can it just naturally be made looking like a cleanup.

Ideally, we'd get b), in small, nice-to-review chunks that incrementally improve
the code without inflating it heavily or moving everything around.

The current locking is nasty, so anything that moves us one step closer into
something that is not only simpler but also more scalable is nice. I am a bit
concerned with the churn in the series as is.

After this series, mm/collapse.c itself is way larger than just mm/khugepaged.c
originally, which raises some eyebrows.

We should also be aware that people are proposing file/shmem mTHP collapse, so
ideally what we refactor would naturally unify some of these code paths.

I am wondering whether shmem mTHP collapse should come first. (I'm hoping that
shmem mTHP collapse can unify some of the anon+file collapse code in a nice way,
to similarly just look like a cleanup while enabling a new scenario. Which is
really what I am hoping for because the current code is A MESS with weirdly
named functions all over the place. I hope it can be unified somehow ... and
that needs some proper thought)

> 
> David and I co-maintain THP together, are in constant communication, and
> have a great working relationship :)

Yes! :)

> 
> IOW - if one of us states a position on the sub(sub?)system - then take
> that to be the actual position.
> 
> I have poured what must be hundreds of hours now into THP maintainership -
> it's by far my biggest workload on the maintenance front, by far the most
> painful and by far the most thankless.
> 
> I do it because I care about mm a great deal and am, frankly, driven by a
> desire to see THP turn from a flaming trash pile of a code base with
> confusing semantics and many, many broken parts into something that serves
> the community's needs with far less maintenance burden.
> 
> Looking over your series it seems some of the patches works in this
> direction (great!), but much else of it fundamentally changes key
> behaviour.

Agreed, I think we really should unify+cleanup the existing code first before
doing more drastic changes.

Having a series that throws all of khugepaged.c into a mixer and pours something
new into collapse.c is ... concerning :)

But I am sure there is a way to incrementally improve the code? At least that's
what I hope.

> 
> So it's just a question of deferring the latter until we get to a sane
> point with the former.
Thanks Lorenzo.

-- 
Cheers,

David
Re: [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives
Posted by Kiryl Shutsemau 1 month, 1 week ago
On Tue, Aug 18, 2026 at 04:12:17PM +0200, David Hildenbrand (Arm) wrote:
> I think we all agree that there is a lot of room for improvement, but the big
> question is:
> 
> (a) When does it stop being a cleanup and is a new feature in disguise that
>     makes the code more complicated and even harder to maintain.
> 
> (b) Can it just naturally be made looking like a cleanup.
> 
> Ideally, we'd get b), in small, nice-to-review chunks that incrementally improve
> the code without inflating it heavily or moving everything around.

It is not a cleanup and I would rather not sell it as one.  It replaces a
mechanism, so judged as (b) it fails by construction.

I believe the end result is much cleaner. But I might be biased. :)

> The current locking is nasty, so anything that moves us one step closer into
> something that is not only simpler but also more scalable is nice. I am a bit
> concerned with the churn in the series as is.
> 
> After this series, mm/collapse.c itself is way larger than just mm/khugepaged.c
> originally, which raises some eyebrows.

Line count is a poor proxy for simplicity or scalability.  What the
engine changes is the serialization model, and that is the part collapse
needs changed: the PMD granularity and the exclusion both come out of
the locking.

Incremental does not reach it, though.  The old mechanism is correct
because it holds mmap_write_lock, the anon_vma write lock and a reference
from the LRU; the engine is correct because the sources are frozen behind
migration entries.  There is no halfway state that is correct under both,
so the switch lands as one patch.

What can be incremental is everything around it: the engine goes in beside
the old mechanism, patch 25 points the anon path at it, and 28 removes what
it replaces.  Until 28 both are in the tree with only one of them
reachable, so the switch can be reverted on its own.

> We should also be aware that people are proposing file/shmem mTHP collapse, so
> ideally what we refactor would naturally unify some of these code paths.
> 
> I am wondering whether shmem mTHP collapse should come first. (I'm hoping that
> shmem mTHP collapse can unify some of the anon+file collapse code in a nice way,
> to similarly just look like a cleanup while enabling a new scenario.

mTHP collapse as it stands has limited usability: PMD-aligned windows only,
and one VMA has to own the PMD.  Bolting file collapse onto the same
structure adds to the debt instead of paying it down.

It would fit the new design.  The frame -- scan, candidate selection, the
round and its passes -- has nothing anon-specific in it; what is
anon-specific sits in the freeze (folio_test_anon(), PageAnonExclusive())
and the unshare in the fault-in pass.  A file source would bring its own
check, freeze, copy and install.

I am not sure it should, though.

Do we want to find file collapse candidates by walking the virtual
address space at all?

collapse_file() already works on the mapping -- it builds the folio in
the page cache and then repairs every mapping through
retract_page_tables() -- so the VMA walk only picks which inode range to
try, and it reaches only what a registered mm maps right now.  Large
folios buy more than TLB reach: fewer page cache entries, cheaper
writeback, natural locking batch, etc.  Those apply whether the file is
mapped or not, and going at the inode directly would reach them.

> Agreed, I think we really should unify+cleanup the existing code first before
> doing more drastic changes.
> 
> Having a series that throws all of khugepaged.c into a mixer and pours something
> new into collapse.c is ... concerning :)

The moving around is patches 29-35 and the tracing after them.  None of it
is needed for the engine: 1-28 add it, switch the anon path over and delete
the old mechanism, without moving anything else out of khugepaged.c.  If
the churn is the problem, v2 can stop there and the moves can come later as
their own series.

-- 
  Kiryl Shutsemau / Kirill A. Shutemov
Re: [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives
Posted by David Hildenbrand (Arm) 2 weeks ago
Following up on some older threads (sorry, the last month was insane)

[...]

>> Ideally, we'd get b), in small, nice-to-review chunks that incrementally improve
>> the code without inflating it heavily or moving everything around.
> 
> It is not a cleanup and I would rather not sell it as one.  It replaces a
> mechanism, so judged as (b) it fails by construction.
> 
> I believe the end result is much cleaner. But I might be biased. :)
> 

It should start with cleanups, though, and not moving stuff around and rewriting
it in a rather uncontrolled manner.

I saw that you started sending out cleanups, good. :)

>> The current locking is nasty, so anything that moves us one step closer into
>> something that is not only simpler but also more scalable is nice. I am a bit
>> concerned with the churn in the series as is.
>>
>> After this series, mm/collapse.c itself is way larger than just mm/khugepaged.c
>> originally, which raises some eyebrows.
> 
> Line count is a poor proxy for simplicity or scalability.  What the
> engine changes is the serialization model, and that is the part collapse
> needs changed: the PMD granularity and the exclusion both come out of
> the locking.
> 
> Incremental does not reach it, though.  The old mechanism is correct
> because it holds mmap_write_lock, the anon_vma write lock and a reference
> from the LRU; the engine is correct because the sources are frozen behind
> migration entries.  There is no halfway state that is correct under both,
> so the switch lands as one patch.
> 
> What can be incremental is everything around it: the engine goes in beside
> the old mechanism, patch 25 points the anon path at it, and 28 removes what
> it replaces.  Until 28 both are in the tree with only one of them
> reachable, so the switch can be reverted on its own.

What I am saying is that there must be a way where we slowly move into that
direction.

We try to move in small, controlled increments where possible.

Huge rewrites are not helpful, neither for reviewers nor for maintainers nor for
consumers.

Sure, at some point there will be a switch, and bigger changes but what has been
presented in this RFC was way too much and way too radical for one series.

Nobody can really review 39 patches of RFC.

Often it's more helpful to propose an overall design idea, and then discuss if
and how to get there.

(and the design idea should not read like schlop otherwise nobody will have an
interest in engaging with it)

> 
>> We should also be aware that people are proposing file/shmem mTHP collapse, so
>> ideally what we refactor would naturally unify some of these code paths.
>>
>> I am wondering whether shmem mTHP collapse should come first. (I'm hoping that
>> shmem mTHP collapse can unify some of the anon+file collapse code in a nice way,
>> to similarly just look like a cleanup while enabling a new scenario.
> 
> mTHP collapse as it stands has limited usability: PMD-aligned windows only,
> and one VMA has to own the PMD.  Bolting file collapse onto the same
> structure adds to the debt instead of paying it down.
> 
> It would fit the new design.  The frame -- scan, candidate selection, the
> round and its passes -- has nothing anon-specific in it; what is
> anon-specific sits in the freeze (folio_test_anon(), PageAnonExclusive())
> and the unshare in the fault-in pass.  A file source would bring its own
> check, freeze, copy and install.
> 
> I am not sure it should, though.
> 
> Do we want to find file collapse candidates by walking the virtual
> address space at all?

In guest_memfd, we recently discussed that we actually would want a mechanism to
collapse even without walking the VA space .... but maybe guest_memfd is just
too special, not sure.

You still need a mmap to identify the candidate file IIRC.

> 
> collapse_file() already works on the mapping -- it builds the folio in
> the page cache and then repairs every mapping through
> retract_page_tables() -- so the VMA walk only picks which inode range to
> try, and it reaches only what a registered mm maps right now.  Large
> folios buy more than TLB reach: fewer page cache entries, cheaper
> writeback, natural locking batch, etc.  Those apply whether the file is
> mapped or not, and going at the inode directly would reach them.
> 
>> Agreed, I think we really should unify+cleanup the existing code first before
>> doing more drastic changes.
>>
>> Having a series that throws all of khugepaged.c into a mixer and pours something
>> new into collapse.c is ... concerning :)
> 
> The moving around is patches 29-35 and the tracing after them.  None of it
> is needed for the engine: 1-28 add it, switch the anon path over and delete
> the old mechanism, without moving anything else out of khugepaged.c.  If
> the churn is the problem, v2 can stop there and the moves can come later as
> their own series.
> 

The churn is always the problem.

Small, controlled increments please.

-- 
Cheers,

David
Re: [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives
Posted by Kiryl Shutsemau 1 week, 6 days ago
On Mon, Sep 14, 2026 at 04:54:06PM +0200, David Hildenbrand (Arm) wrote:
> I saw that you started sending out cleanups, good. :)

Right. I have few more cleanup ideas in the queue, before substantive
changes.

> Nobody can really review 39 patches of RFC.
> 
> Often it's more helpful to propose an overall design idea, and then discuss if
> and how to get there.

Fair.

The overall design one-pager is in comment in mm/collapse.c in 07/57.

I can pull it out, expand and post as RFD if it helpful.

> > Do we want to find file collapse candidates by walking the virtual
> > address space at all?
> 
> In guest_memfd, we recently discussed that we actually would want a mechanism to
> collapse even without walking the VA space .... but maybe guest_memfd is just
> too special, not sure.
> 
> You still need a mmap to identify the candidate file IIRC.

At the moment, yes. But it is not the only option. We can start from
walking inodes for superblocks that opted in for collapse.

> The churn is always the problem.
> 
> Small, controlled increments please.

Ack.

Just to re-iterate what I told Lorenzo: I never intended to push this as
one big patchset. The purpose of the RFC was to demonstrate the end
state I pursue. I am flexible on how we get there.

-- 
  Kiryl Shutsemau / Kirill A. Shutemov
Re: [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives
Posted by Lorenzo Stoakes (ARM) 1 month ago
On Wed, Aug 19, 2026 at 07:08:07PM +0100, Kiryl Shutsemau wrote:
> On Tue, Aug 18, 2026 at 04:12:17PM +0200, David Hildenbrand (Arm) wrote:
> > I think we all agree that there is a lot of room for improvement, but the big
> > question is:
> >
> > (a) When does it stop being a cleanup and is a new feature in disguise that
> >     makes the code more complicated and even harder to maintain.
> >
> > (b) Can it just naturally be made looking like a cleanup.
> >
> > Ideally, we'd get b), in small, nice-to-review chunks that incrementally improve
> > the code without inflating it heavily or moving everything around.
>
> It is not a cleanup and I would rather not sell it as one.  It replaces a
> mechanism, so judged as (b) it fails by construction.
>
> I believe the end result is much cleaner. But I might be biased. :)
>
> > The current locking is nasty, so anything that moves us one step closer into
> > something that is not only simpler but also more scalable is nice. I am a bit
> > concerned with the churn in the series as is.
> >
> > After this series, mm/collapse.c itself is way larger than just mm/khugepaged.c
> > originally, which raises some eyebrows.
>
> Line count is a poor proxy for simplicity or scalability.  What the
> engine changes is the serialization model, and that is the part collapse
> needs changed: the PMD granularity and the exclusion both come out of
> the locking.
>
> Incremental does not reach it, though.  The old mechanism is correct
> because it holds mmap_write_lock, the anon_vma write lock and a reference
> from the LRU; the engine is correct because the sources are frozen behind
> migration entries.  There is no halfway state that is correct under both,
> so the switch lands as one patch.
>
> What can be incremental is everything around it: the engine goes in beside
> the old mechanism, patch 25 points the anon path at it, and 28 removes what
> it replaces.  Until 28 both are in the tree with only one of them
> reachable, so the switch can be reverted on its own.
>
> > We should also be aware that people are proposing file/shmem mTHP collapse, so
> > ideally what we refactor would naturally unify some of these code paths.
> >
> > I am wondering whether shmem mTHP collapse should come first. (I'm hoping that
> > shmem mTHP collapse can unify some of the anon+file collapse code in a nice way,
> > to similarly just look like a cleanup while enabling a new scenario.
>
> mTHP collapse as it stands has limited usability: PMD-aligned windows only,
> and one VMA has to own the PMD.  Bolting file collapse onto the same
> structure adds to the debt instead of paying it down.
>
> It would fit the new design.  The frame -- scan, candidate selection, the
> round and its passes -- has nothing anon-specific in it; what is
> anon-specific sits in the freeze (folio_test_anon(), PageAnonExclusive())
> and the unshare in the fault-in pass.  A file source would bring its own
> check, freeze, copy and install.
>
> I am not sure it should, though.
>
> Do we want to find file collapse candidates by walking the virtual
> address space at all?
>
> collapse_file() already works on the mapping -- it builds the folio in
> the page cache and then repairs every mapping through
> retract_page_tables() -- so the VMA walk only picks which inode range to
> try, and it reaches only what a registered mm maps right now.  Large
> folios buy more than TLB reach: fewer page cache entries, cheaper
> writeback, natural locking batch, etc.  Those apply whether the file is
> mapped or not, and going at the inode directly would reach them.
>
> > Agreed, I think we really should unify+cleanup the existing code first before
> > doing more drastic changes.
> >
> > Having a series that throws all of khugepaged.c into a mixer and pours something
> > new into collapse.c is ... concerning :)
>
> The moving around is patches 29-35 and the tracing after them.  None of it
> is needed for the engine: 1-28 add it, switch the anon path over and delete
> the old mechanism, without moving anything else out of khugepaged.c.  If
> the churn is the problem, v2 can stop there and the moves can come later as
> their own series.

This whole reply seems AI-generated...

You replying only to David twice in this sub-thread which isn't exactly giving
me warm fuzzy feelings about the working-around-me concerns I raised here.

So simple feedback - send a relatively small, no-functional-change series that
improves THP code and lays foundations for future changes. After the merge
window.

Can you explicitly ack this please?

>
> --
>   Kiryl Shutsemau / Kirill A. Shutemov

--
Cheers, Lorenzo
Re: [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives
Posted by Kiryl Shutsemau 1 month ago
On Mon, Aug 24, 2026 at 03:22:02PM +0100, Lorenzo Stoakes (ARM) wrote:
> This whole reply seems AI-generated...

It is not.

I heard David's feedback. Will adjust accordingly.

> 
> You replying only to David twice in this sub-thread which isn't exactly giving
> me warm fuzzy feelings about the working-around-me concerns I raised here.
> 
> So simple feedback - send a relatively small, no-functional-change series that
> improves THP code and lays foundations for future changes. After the merge
> window.
> 
> Can you explicitly ack this please?

Ack.

-- 
  Kiryl Shutsemau / Kirill A. Shutemov
Re: [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives
Posted by Lorenzo Stoakes (ARM) 1 month ago
On Mon, Aug 24, 2026 at 03:59:50PM +0100, Kiryl Shutsemau wrote:
> On Mon, Aug 24, 2026 at 03:22:02PM +0100, Lorenzo Stoakes (ARM) wrote:
> > This whole reply seems AI-generated...
>
> It is not.
>
> I heard David's feedback. Will adjust accordingly.
>
> >
> > You replying only to David twice in this sub-thread which isn't exactly giving
> > me warm fuzzy feelings about the working-around-me concerns I raised here.
> >
> > So simple feedback - send a relatively small, no-functional-change series that
> > improves THP code and lays foundations for future changes. After the merge
> > window.
> >
> > Can you explicitly ack this please?
>
> Ack.

Just to follow up - Kiryl and I spoke directly (much better medium than text :)
and it's clear  that there's just been some miscommunication here.

So - in general all's good and we are it seems in violent _agreement_ about
moving forwards (split out to reasonable foundational series that build upon one
another -> eventual switch on).

Let peace reign :)

>
> --
>   Kiryl Shutsemau / Kirill A. Shutemov

--
Cheers, Lorenzo
Re: [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives
Posted by Lorenzo Stoakes (ARM) 1 month, 1 week ago
On Tue, Aug 18, 2026 at 04:12:17PM +0200, David Hildenbrand (Arm) wrote:
> On 8/18/26 15:06, Lorenzo Stoakes (ARM) wrote:
> > (I'll reply to the rest of this later)
> >
> > On Mon, Aug 17, 2026 at 02:38:44PM +0100, Kiryl Shutsemau wrote:
> >> David and I talked about this at LSF/MM.  I can give it a try if it fits
> >> your idea of "feature freeze" -- and if it doesn't collide with the series
> >> you have in flight, in which case I'd rather go after yours than around it.
> >
> > Whether intended or not you sound rather like you are trying to override me
> > in favour of my co-maintainer here and it's... not helpful.
>
> I think Kiryl tried to say that we discussed at LSF/MM which areas of MM scream
> for an improvement, and we discussed that khugepaged is just horrible code.

Yes, and obviously I agree with that very much!

>
> I think we all agree that there is a lot of room for improvement, but the big
> question is:
>
> (a) When does it stop being a cleanup and is a new feature in disguise that
>     makes the code more complicated and even harder to maintain.
>
> (b) Can it just naturally be made looking like a cleanup.

Right and point (a) is exactly my pushback here.

And it's not always so easy to separate.

>
> Ideally, we'd get b), in small, nice-to-review chunks that incrementally improve
> the code without inflating it heavily or moving everything around.

Yep.

>
> The current locking is nasty, so anything that moves us one step closer into
> something that is not only simpler but also more scalable is nice. I am a bit
> concerned with the churn in the series as is.

Yes.

What I'm saying is, essentially, go read the code. Go see how coupled things
are. Go read the functions that require you to keep a giant stack of
state to even know what's going on.

Look at the bug rate, and how subtle the bugs are - essentially - 'behold the
horrors' :)

And rather than being opposed to fundamental reworks (actually - it's the exact
opposite - I think THP needs changing from top-to-bottom):

Based on experience of seeing work done in THP - we are _making it worse_ when
we add features without paying down this debt.

I think it's nuanced, because as part of reworking things (you say this below
too), patterns and approaches can fall out.

And you can naturally lead things towards a sensible rework.

>
> After this series, mm/collapse.c itself is way larger than just mm/khugepaged.c
> originally, which raises some eyebrows.

Yeah exactly.

>
> We should also be aware that people are proposing file/shmem mTHP collapse, so
> ideally what we refactor would naturally unify some of these code paths.

Right yes. There's no harm in _laying the foundations_ for future changes.

In fact a lot of reworks are about doing exactly that - you can often go one of
2 paths:

	a. push the feature in as some tacked-on thing that works but adds
	   complexity/maintainership overhhead/etc. or

	b. Change the architecture to suit the feature you intend.

So my opposition is to a, not b.

>
> I am wondering whether shmem mTHP collapse should come first. (I'm hoping that
> shmem mTHP collapse can unify some of the anon+file collapse code in a nice way,
> to similarly just look like a cleanup while enabling a new scenario. Which is
> really what I am hoping for because the current code is A MESS with weirdly
> named functions all over the place. I hope it can be unified somehow ... and
> that needs some proper thought)

Yes.

And to be clear and I am going to say this in my (proper) reply to Kiryl - I do
think moving away from the presumption of PMD collapse is _key_ to a more
general rework.

But it's about how we get there, how the rest of the code looks, what other work
we do around it.

>
> >
> > David and I co-maintain THP together, are in constant communication, and
> > have a great working relationship :)
>
> Yes! :)

:)

>
> >
> > IOW - if one of us states a position on the sub(sub?)system - then take
> > that to be the actual position.
> >
> > I have poured what must be hundreds of hours now into THP maintainership -
> > it's by far my biggest workload on the maintenance front, by far the most
> > painful and by far the most thankless.
> >
> > I do it because I care about mm a great deal and am, frankly, driven by a
> > desire to see THP turn from a flaming trash pile of a code base with
> > confusing semantics and many, many broken parts into something that serves
> > the community's needs with far less maintenance burden.
> >
> > Looking over your series it seems some of the patches works in this
> > direction (great!), but much else of it fundamentally changes key
> > behaviour.
>
> Agreed, I think we really should unify+cleanup the existing code first before
> doing more drastic changes.
>
> Having a series that throws all of khugepaged.c into a mixer and pours something
> new into collapse.c is ... concerning :)
>
> But I am sure there is a way to incrementally improve the code? At least that's
> what I hope.

Yes. And it does look like a lot of the early patches are along the right road.

I do plan to do some proper feedback on this series along the lines of figuring
out how we move this forwards.

>
> >
> > So it's just a question of deferring the latter until we get to a sane
> > point with the former.
> Thanks Lorenzo.
>
> --
> Cheers,
>
> David

--
Cheers, Lorenzo
Re: [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives
Posted by Zi Yan 1 month, 1 week ago
On Sun Aug 16, 2026 at 6:45 PM EDT, Kiryl Shutsemau wrote:
> From: "Kiryl Shutsemau (Meta)" <kas@kernel.org>
>
> Yes, I know, this is a lot of changes. But I'm happy with the overall state
> of the patchset and the only reason I tag it as RFC is that it is tricky
> to get 57 patches upstream.
>
> I wanted to give a view of the end state first. I will suggest a possible
> way to split it below.
>
> I would appreciate any feedback.
>
> TL;DR
> =====
>
> This replaces khugepaged's anonymous collapse with an engine that
> can collapse sub-PMD ranges. It is built around migration entries and
> frozen folios instead of heavy locking and isolation, aiming for better
> scalability and less disruption to the workload being collapsed.
>
> Why
> ===
>
> mTHP collapse landed in khugepaged in 7.2 and I was glad to see it.  We
> at Meta run arm64 with 64K base pages, where a PMD is 512M: PMD-order THP
> is of limited use at that size, and mTHP is exactly what we want.
>
> It turned out not to help us.
>
> khugepaged only ever looks at PMD-aligned windows, and it is not an easy
> limitation to lift.
>
> Fixing the alignment is a one-line change, but what it feeds assumes the
> PMD everywhere that matters: collapse_huge_page() clears and flushes the
> whole PMD whatever order it is collapsing, installs a PMD leaf because
> that is the only thing it can produce, and keeps everyone out with
> mmap_write_lock, anon_vma_lock_write() and an IPI broadcast while it
> does.
>
> Which is why hugepage_vma_revalidate() demands that the VMA span the
> whole PMD even for an mTHP order -- "we'd need to lock all VMAs in the
> PMD range to support this", as the comment there puts it.  A PMD-granular
> operation is only safe when one VMA owns the PMD, and that is exactly the
> restriction in the way.  The alignment is the symptom; the PMD is the
> design.
>
> So both roots have to go.

I agree that khugepaged is designed for PMD-aligned collapse and this is
a limitation we want to get rid of. It is great you are looking at them.

>
> Design
> ======
>
> The old mechanism holds the address space still because it has nothing
> else stopping the sources from moving under the copy.  The new engine
> makes the sources themselves inert instead, with the two barriers
> migration already uses, raised in that order:
>
>   1. migration entries replace the source PTEs.  Faults and GUP-slow
>      now wait on the source folio's lock, which is taken before the
>      first entry becomes visible.
>   2. the source folio's refcount is frozen to its expected value.
>      GUP-fast, pfn walkers, reclaim, compaction and memory-failure all
>      fail folio_try_get() and back off.
>
> Between the two, nothing can reach a source, so the copy runs with no
> lock held at all -- and the address space is left alone while it does.
>
> What that removes from every collapse path:
>
>   mmap_write_lock              -> mmap_read
>   anon_vma_lock_write()        -> nothing: an rmap walk needs the folio
>                                   locked, and the engine holds that lock
>                                   from freeze to putback
>   tlb_remove_table_sync_one()  -> nothing: one ranged flush per round
>   LRU isolation                -> nothing: sources are inert in place

I remember we were discussing using migration entry and the issue with
mmap_write_lock() in the context of in-place THP promotion and the
conclusion was that because MADV_DONTNEED (maybe MADV_REMOVE or
MADV_PAGEOUT) works on page table and does not change VMAs,
mmap_write_lock() is needed to prevent things being changed under
khugepaged. Anything different in normal khugepaged collapse process, so
that it is OK to use mmap_read_lock? Let me know if I misremember it.

Thanks.

-- 
Best Regards,
Yan, Zi
Re: [RFC PATCH 00/57] mm/collapse: rebuild collapse on migration primitives
Posted by Kiryl Shutsemau 1 month, 1 week ago
On Sun, Aug 16, 2026 at 10:02:53PM -0400, Zi Yan wrote:
> > What that removes from every collapse path:
> >
> >   mmap_write_lock              -> mmap_read
> >   anon_vma_lock_write()        -> nothing: an rmap walk needs the folio
> >                                   locked, and the engine holds that lock
> >                                   from freeze to putback
> >   tlb_remove_table_sync_one()  -> nothing: one ranged flush per round
> >   LRU isolation                -> nothing: sources are inert in place
> 
> I remember we were discussing using migration entry and the issue with
> mmap_write_lock() in the context of in-place THP promotion and the
> conclusion was that because MADV_DONTNEED (maybe MADV_REMOVE or
> MADV_PAGEOUT) works on page table and does not change VMAs,
> mmap_write_lock() is needed to prevent things being changed under
> khugepaged. Anything different in normal khugepaged collapse process, so
> that it is OK to use mmap_read_lock? Let me know if I misremember it.

IIRC that discussion predates Hugh's pte_offset_map() rework -- since
0d940a9b270b the helper takes rcu_read_lock() and fails if the pmd is
none, !present or huge, so mmap_write is no longer what keeps a pte
walker out.  MADV_DONTNEED is still not excluded, and the engine does not
try to: the install re-reads every slot under the ptl and publishes only
if it still holds the migration entry this round put there, leaving a
zapped slot alone and dropping the rmap the frozen source still held.

-- 
  Kiryl Shutsemau / Kirill A. Shutemov