[RFC PATCH 0/4] mm: compaction: mTHP-friendly memory compaction

Bo Zhang posted 4 patches 1 month ago
mm/compaction.c | 83 ++++++++++++++++++++++++++++++++++++++-----------
mm/internal.h   |  3 ++
mm/vmscan.c     | 23 +++-----------
3 files changed, 72 insertions(+), 37 deletions(-)
[RFC PATCH 0/4] mm: compaction: mTHP-friendly memory compaction
Posted by Bo Zhang 1 month ago
Hi all,

This series improves memory compaction to better serve mTHP (multi-size
THP) allocations, particularly for small orders like order-2 (16KB).
The changes cover the proactive compaction and kswapd-triggered compaction
paths.

Problem:

The current proactive compaction targets COMPACTION_HPAGE_ORDER (order-9,
2MB), which is unfriendly to mTHP in two ways:

1. It migrates folios that already satisfy mTHP allocation needs. For
   example, an order-2 folio is a valid mTHP page, yet compaction still
   moves it around trying to form order-9 blocks. This is unnecessary
   work and wastes energy.

2. Compaction designed for 2MB huge pages is too heavyweight for mTHP.
   mTHP allocations are frequent and only require small contiguous blocks
   (e.g., 4 pages for order-2). A lighter-weight, mTHP-aware compaction
   strategy is needed to reduce overhead and power consumption.

Approach:

This series makes four changes:

1. Generalize the fragmentation score functions to accept an order parameter
   and use the minimum always-enabled mTHP order as the compaction target.

2. During proactive compaction (compact_memory), skip isolating folios that
   already satisfy mTHP requirements, avoiding unnecessary migration overhead.

3. Allow proactive compaction to proceed concurrently with kswapd for
   non-costly mTHP orders, since kswapd reclaim alone may not produce the
   contiguous blocks needed for these allocations.

4. Introduce zone_effective_free_pages() that provides mTHP-aware free page
   accounting for watermark checks, counting only buddy blocks that can
   actually satisfy mTHP allocations.

Test setup:

Boot Ubuntu with 2700 MB of memory, with mTHP disabled
initially and the defrag mode set to defer+madvise.

Then set the 16 KB mTHP size to always and run a kernel
build with -j20.

For both cases below, we run a background script to
proactively trigger compaction as follows:
 #!/bin/bash

 while true; do
 	echo 50 > /proc/sys/vm/compaction_proactiveness
 	sleep 0.1
 done

W/o patch:

*** Executing round 0 ***

real	2m1.312s
user	25m38.968s
sys	5m16.934s
anon_fault_alloc: 6328991
anon_fault_fallback: 214553

*** Executing round 1 ***

real	1m55.370s
user	25m24.482s
sys	3m52.512s
anon_fault_alloc: 6355263
anon_fault_fallback: 108692

*** Executing round 2 ***

real	2m7.579s
user	25m11.530s
sys	3m45.456s
anon_fault_alloc: 6355816
anon_fault_fallback: 107852

*** Executing round 3 ***

real	1m53.824s
user	25m26.774s
sys	3m42.160s
anon_fault_alloc: 6355457
anon_fault_fallback: 107705

W/patch:

*** Executing round 0 ***

real	1m55.906s
user	25m16.985s
sys	4m24.480s
anon_fault_alloc: 6354486
anon_fault_fallback: 109845

*** Executing round 1 ***

real	1m51.303s
user	25m16.456s
sys	3m26.629s
anon_fault_alloc: 6392515
anon_fault_fallback: 69797

*** Executing round 2 ***

real	1m51.495s
user	25m13.075s
sys	3m28.501s
anon_fault_alloc: 6395096
anon_fault_fallback: 67510

*** Executing round 3 ***

real	1m52.980s
user	25m8.217s
sys	3m37.506s
anon_fault_alloc: 6389556
anon_fault_fallback: 72743

Before "Executing round 0", mTHP is not enabled. Therefore, both
cases show a higher anon_fault_fallback in Round 0 than in the
other rounds. With the patch, however, memory can be compacted faster
into an mTHP-friendly state, resulting in a much lower fallback rate
in Round 0. In the other rounds, the patch also consistently shows
a lower anon_fault_fallback, as well as lower sys and wall time for
the kernel build.

Open questions:

1. When multiple mTHP orders are enabled (e.g., order-2 and order-4 both
   "always"), this series only targets the minimum order. Should proactive
   compaction also independently evaluate and serve higher orders?

2. In skip_isolation_on_order(), the filter order ideally should come from
   the compaction control path. However, during proactive compaction
   target_order is always -1 (via compact_memory), and there is no clean
   way to pass the mTHP order down from upper layers. Currently we read
   huge_anon_orders_always directly, but this variable can be changed by
   userspace at any time, making the semantic fragile (the compaction may
   start with one order target and finish with another). Ideas on how to
   plumb the target order through the proactive compaction path cleanly
   are welcome.

3. In __compact_finished(), the original code skips proactive compaction
   when kswapd is running to avoid interference. Patch 3 removes this
   skip for non-costly mTHP orders (< PAGE_ALLOC_COSTLY_ORDER). The
   reason is that small-order compaction is lightweight and likely to
   succeed quickly even while kswapd is reclaiming, forming an order-2
   block requires migrating very few pages. Does this approach make sense,
   or is there a better way to coordinate proactive compaction with kswapd
   in the mTHP scenario?

4. In the direct reclaim path (__alloc_pages_slowpath), compact_first is
   only set for costly orders or non-movable allocations. For mTHP always-
   enabled non-costly orders (e.g., order-2 MIGRATE_MOVABLE), when free
   memory is sufficient (watermarks met) but fragmentation is high,
   compaction is more appropriate than reclaim. Should we also set
   compact_first for this case to avoid unnecessary reclaim?

Bo Zhang (4):
  mm: compaction: make proactive compaction mTHP-aware
  mm: compaction: skip isolating large folios that satisfy the mTHP order
  mm: compaction: don't skip proactive compaction for non-costly mTHP
  mm: adjust free_pages to make __zone_watermark_ok() mTHP-aware

 mm/compaction.c | 83 ++++++++++++++++++++++++++++++++++++++-----------
 mm/internal.h   |  3 ++
 mm/vmscan.c     | 23 +++-----------
 3 files changed, 72 insertions(+), 37 deletions(-)

--
2.34.1
Re: [RFC PATCH 0/4] mm: compaction: mTHP-friendly memory compaction
Posted by Lorenzo Stoakes (ARM) 1 week, 2 days ago
This is the friendly patch-bot of Lorenzo Stoakes.

You have sent him a patch/series that has triggered this response.

He used to manually respond to these common problems, but in order to save
his sanity (he kept writing the same thing over and over, yet to different
people), I was created.

Hopefully you will not take offence and will fix the problem in your patch
and resubmit it so that it can be accepted into the Linux kernel tree.

When sending emails to mm:

For one of several possible reasons this mail has triggered an AI
detector script.

Note that, while we are fine with AI assistance, it is kernel
policy that you must disclose this with an Assisted-by tag like:

Assisted-by: LLM

See https://docs.kernel.org/process/coding-assistants.html

It is also kernel policy that you must fully understand and take
responsibility for every patch that you send.

See https://docs.kernel.org/process/generated-content.html most
notably:

        If tools permit you to generate a contribution automatically, expect
        additional scrutiny in proportion to how much of it was generated.

        As with the output of any tooling, the result may be incorrect or
        inappropriate. You are expected to understand and to be able to
        defend everything you submit. If you are unable to do so, then do
        not submit the resulting changes.

        If you do so anyway, maintainers are entitled to reject your series
        without detailed review.

In general, if you are a newcomer to mm, we expect you to do smaller work
before moving on to larger changes, so you build understanding of both the
technical aspects of mm and how we do things.

If you wish to discuss this problem further, or you have questions about
how to resolve this issue, please feel free to respond to this email and
Lorenzo will reply once he has dug out from the pending patches received
from other developers.

thanks,

Lorenzo's patch email bot

[ Idea shamelessly stolen from greg-kh ]

-- 
Cheers, Lorenzo
Re: [RFC PATCH 0/4] mm: compaction: mTHP-friendly memory compaction
Posted by Bo Zhang 6 days, 2 hours ago
Hi Lorenzo,

Thanks for the note. Let me clarify the background of this series, both
regarding the AI-assistance question and regarding the scope of the change.

This work started as part of our exploration of 16KB mTHP for the Android
scenario. The overall direction of making compaction lighter-weight and
mTHP-aware came from offline discussions with David, who suggested this
lightweight-compaction angle for small mTHP orders. So the size and scope
of the series is not me diving in unguided as a newcomer: it follows his
directional guidance. Barry then provided reviews and change suggestions
throughout the iterations.

On the AI-assistance side:

- All the code was written by me, by hand. I understand each change and
  take full responsibility for the patches.

- The cover letter and some commit messages were polished with an LLM
  after I wrote the initial drafts (to smooth the English). The technical
  content, the design (jointly with Barry), and the analysis are mine.

Please let me know whether an Assisted-by: LLM tag is expected for the
cover / commit-message polishing; I'm happy to add it in v2 if that's the
convention.

Bo
Re: [RFC PATCH 0/4] mm: compaction: mTHP-friendly memory compaction
Posted by Barry Song 5 days, 9 hours ago
On Tue, Sep 22, 2026 at 12:15 PM Bo Zhang <zhangbo0325@gmail.com> wrote:
>
> Hi Lorenzo,
>
> Thanks for the note. Let me clarify the background of this series, both
> regarding the AI-assistance question and regarding the scope of the change.
>
> This work started as part of our exploration of 16KB mTHP for the Android
> scenario. The overall direction of making compaction lighter-weight and
> mTHP-aware came from offline discussions with David, who suggested this
> lightweight-compaction angle for small mTHP orders. So the size and scope
> of the series is not me diving in unguided as a newcomer: it follows his
> directional guidance. Barry then provided reviews and change suggestions
> throughout the iterations.
>
> On the AI-assistance side:
>
> - All the code was written by me, by hand. I understand each change and
>   take full responsibility for the patches.
>
> - The cover letter and some commit messages were polished with an LLM
>   after I wrote the initial drafts (to smooth the English). The technical
>   content, the design (jointly with Barry), and the analysis are mine.

Maybe Lorenzo can adjust the bot a little so that it doesn't trigger
an AI warning just because the changelog or commit message looks
AI-polished? Perhaps the AI detector could focus more on the code
changes instead?

Please have a little mercy on us non-native English speakers. :-)

Thanks
Barry
Re: [RFC PATCH 0/4] mm: compaction: mTHP-friendly memory compaction
Posted by Zi Yan 3 weeks ago
On Tue Aug 25, 2026 at 12:38 AM EDT, Bo Zhang wrote:
> Hi all,
>
> This series improves memory compaction to better serve mTHP (multi-size
> THP) allocations, particularly for small orders like order-2 (16KB).
> The changes cover the proactive compaction and kswapd-triggered compaction
> paths.
>
> Problem:
>
> The current proactive compaction targets COMPACTION_HPAGE_ORDER (order-9,
> 2MB), which is unfriendly to mTHP in two ways:
>
> 1. It migrates folios that already satisfy mTHP allocation needs. For
>    example, an order-2 folio is a valid mTHP page, yet compaction still
>    moves it around trying to form order-9 blocks. This is unnecessary
>    work and wastes energy.

But skip_isolation_on_order() skips a folio with an order >= target
order.

>
> 2. Compaction designed for 2MB huge pages is too heavyweight for mTHP.

Yes.

>    mTHP allocations are frequent and only require small contiguous blocks
>    (e.g., 4 pages for order-2). A lighter-weight, mTHP-aware compaction
>    strategy is needed to reduce overhead and power consumption.
>
> Approach:
>
> This series makes four changes:
>
> 1. Generalize the fragmentation score functions to accept an order parameter
>    and use the minimum always-enabled mTHP order as the compaction target.

Makes sense.

>
> 2. During proactive compaction (compact_memory), skip isolating folios that
>    already satisfy mTHP requirements, avoiding unnecessary migration overhead.

Oh, you are targeting proactive compaction, where
skip_isolation_on_order() does not apply.

>
> 3. Allow proactive compaction to proceed concurrently with kswapd for
>    non-costly mTHP orders, since kswapd reclaim alone may not produce the
>    contiguous blocks needed for these allocations.
>
> 4. Introduce zone_effective_free_pages() that provides mTHP-aware free page
>    accounting for watermark checks, counting only buddy blocks that can
>    actually satisfy mTHP allocations.
>
> Test setup:
>
> Boot Ubuntu with 2700 MB of memory, with mTHP disabled
> initially and the defrag mode set to defer+madvise.
>
> Then set the 16 KB mTHP size to always and run a kernel
> build with -j20.
>
> For both cases below, we run a background script to
> proactively trigger compaction as follows:
>  #!/bin/bash
>
>  while true; do
>  	echo 50 > /proc/sys/vm/compaction_proactiveness
>  	sleep 0.1
>  done
>
> W/o patch:
>
> *** Executing round 0 ***
>
> real	2m1.312s
> user	25m38.968s
> sys	5m16.934s
> anon_fault_alloc: 6328991
> anon_fault_fallback: 214553
>
> *** Executing round 1 ***
>
> real	1m55.370s
> user	25m24.482s
> sys	3m52.512s
> anon_fault_alloc: 6355263
> anon_fault_fallback: 108692
>
> *** Executing round 2 ***
>
> real	2m7.579s
> user	25m11.530s
> sys	3m45.456s
> anon_fault_alloc: 6355816
> anon_fault_fallback: 107852
>
> *** Executing round 3 ***
>
> real	1m53.824s
> user	25m26.774s
> sys	3m42.160s
> anon_fault_alloc: 6355457
> anon_fault_fallback: 107705
>
> W/patch:
>
> *** Executing round 0 ***
>
> real	1m55.906s
> user	25m16.985s
> sys	4m24.480s
> anon_fault_alloc: 6354486
> anon_fault_fallback: 109845
>
> *** Executing round 1 ***
>
> real	1m51.303s
> user	25m16.456s
> sys	3m26.629s
> anon_fault_alloc: 6392515
> anon_fault_fallback: 69797
>
> *** Executing round 2 ***
>
> real	1m51.495s
> user	25m13.075s
> sys	3m28.501s
> anon_fault_alloc: 6395096
> anon_fault_fallback: 67510
>
> *** Executing round 3 ***
>
> real	1m52.980s
> user	25m8.217s
> sys	3m37.506s
> anon_fault_alloc: 6389556
> anon_fault_fallback: 72743
>
> Before "Executing round 0", mTHP is not enabled. Therefore, both
> cases show a higher anon_fault_fallback in Round 0 than in the
> other rounds. With the patch, however, memory can be compacted faster
> into an mTHP-friendly state, resulting in a much lower fallback rate
> in Round 0. In the other rounds, the patch also consistently shows
> a lower anon_fault_fallback, as well as lower sys and wall time for
> the kernel build.

What about the impact on THP compaction? How does it affect direct
compaction for both mTHP and THP?

It sounds to me that this patch series target proactive compaction. Am I
getting right?

Thanks.

>
> Open questions:
>
> 1. When multiple mTHP orders are enabled (e.g., order-2 and order-4 both
>    "always"), this series only targets the minimum order. Should proactive
>    compaction also independently evaluate and serve higher orders?
>
> 2. In skip_isolation_on_order(), the filter order ideally should come from
>    the compaction control path. However, during proactive compaction
>    target_order is always -1 (via compact_memory), and there is no clean
>    way to pass the mTHP order down from upper layers. Currently we read
>    huge_anon_orders_always directly, but this variable can be changed by
>    userspace at any time, making the semantic fragile (the compaction may
>    start with one order target and finish with another). Ideas on how to
>    plumb the target order through the proactive compaction path cleanly
>    are welcome.
>
> 3. In __compact_finished(), the original code skips proactive compaction
>    when kswapd is running to avoid interference. Patch 3 removes this
>    skip for non-costly mTHP orders (< PAGE_ALLOC_COSTLY_ORDER). The
>    reason is that small-order compaction is lightweight and likely to
>    succeed quickly even while kswapd is reclaiming, forming an order-2
>    block requires migrating very few pages. Does this approach make sense,
>    or is there a better way to coordinate proactive compaction with kswapd
>    in the mTHP scenario?
>
> 4. In the direct reclaim path (__alloc_pages_slowpath), compact_first is
>    only set for costly orders or non-movable allocations. For mTHP always-
>    enabled non-costly orders (e.g., order-2 MIGRATE_MOVABLE), when free
>    memory is sufficient (watermarks met) but fragmentation is high,
>    compaction is more appropriate than reclaim. Should we also set
>    compact_first for this case to avoid unnecessary reclaim?
>
> Bo Zhang (4):
>   mm: compaction: make proactive compaction mTHP-aware
>   mm: compaction: skip isolating large folios that satisfy the mTHP order
>   mm: compaction: don't skip proactive compaction for non-costly mTHP
>   mm: adjust free_pages to make __zone_watermark_ok() mTHP-aware
>
>  mm/compaction.c | 83 ++++++++++++++++++++++++++++++++++++++-----------
>  mm/internal.h   |  3 ++
>  mm/vmscan.c     | 23 +++-----------
>  3 files changed, 72 insertions(+), 37 deletions(-)
>
> --
> 2.34.1




-- 
Best Regards,
Yan, Zi
Re: [RFC PATCH 0/4] mm: compaction: mTHP-friendly memory compaction
Posted by Bo Zhang 2 weeks, 6 days ago
On Sun Sep 07, 2026 at 10:51 PM EDT, Zi Yan wrote:
> But skip_isolation_on_order() skips a folio with an order >= target
> order.
> ...
> Oh, you are targeting proactive compaction, where
> skip_isolation_on_order() does not apply.

Right. To clarify the cover letter: the "migrating folios that already
satisfy mTHP" concern is specific to proactive compaction, where
target_order is -1 (via compact_memory) and the order >= target_order
check in skip_isolation_on_order() does not apply. For compaction with an
explicit target order that path already handles it.

> What about the impact on THP compaction? How does it affect direct
> compaction for both mTHP and THP?
>
> It sounds to me that this patch series target proactive compaction. Am I
> getting right?

Two things:

1) Traditional THP (order-9) is not affected. The mTHP-aware branch in
zone_effective_free_pages() only triggers when order == compact_hpage_order(),
i.e. the minimum always-enabled mTHP order (e.g. order-2). An order-9 THP
request does not match that, so it keeps its original behavior exactly
(NR_FREE_PAGES_BLOCKS under defrag_mode, NR_FREE_PAGES otherwise). We
didn't change the THP path.

2) The series isn't limited to proactive compaction. Patches 1-3 target
proactive compaction, but patch 4 also covers the kswapd -> kcompactd path
via pgdat_balanced(), so both proactive compaction and kswapd wakeup are
addressed.

For direct compaction: patch 4 does touch compaction_suit_allocation_order(),
which is shared with direct compaction, so order-2 mTHP direct compaction
would also fall into the new accounting. However, the direct compaction case
needs more testing and thought. How direct compaction should behave for mTHP
is something worth discussing together to decide the right approach.

Thanks for the review.

Bo