mm/compaction.c | 83 ++++++++++++++++++++++++++++++++++++++----------- mm/internal.h | 3 ++ mm/vmscan.c | 23 +++----------- 3 files changed, 72 insertions(+), 37 deletions(-)
Hi all, This series improves memory compaction to better serve mTHP (multi-size THP) allocations, particularly for small orders like order-2 (16KB). The changes cover the proactive compaction and kswapd-triggered compaction paths. Problem: The current proactive compaction targets COMPACTION_HPAGE_ORDER (order-9, 2MB), which is unfriendly to mTHP in two ways: 1. It migrates folios that already satisfy mTHP allocation needs. For example, an order-2 folio is a valid mTHP page, yet compaction still moves it around trying to form order-9 blocks. This is unnecessary work and wastes energy. 2. Compaction designed for 2MB huge pages is too heavyweight for mTHP. mTHP allocations are frequent and only require small contiguous blocks (e.g., 4 pages for order-2). A lighter-weight, mTHP-aware compaction strategy is needed to reduce overhead and power consumption. Approach: This series makes four changes: 1. Generalize the fragmentation score functions to accept an order parameter and use the minimum always-enabled mTHP order as the compaction target. 2. During proactive compaction (compact_memory), skip isolating folios that already satisfy mTHP requirements, avoiding unnecessary migration overhead. 3. Allow proactive compaction to proceed concurrently with kswapd for non-costly mTHP orders, since kswapd reclaim alone may not produce the contiguous blocks needed for these allocations. 4. Introduce zone_effective_free_pages() that provides mTHP-aware free page accounting for watermark checks, counting only buddy blocks that can actually satisfy mTHP allocations. Test setup: Boot Ubuntu with 2700 MB of memory, with mTHP disabled initially and the defrag mode set to defer+madvise. Then set the 16 KB mTHP size to always and run a kernel build with -j20. For both cases below, we run a background script to proactively trigger compaction as follows: #!/bin/bash while true; do echo 50 > /proc/sys/vm/compaction_proactiveness sleep 0.1 done W/o patch: *** Executing round 0 *** real 2m1.312s user 25m38.968s sys 5m16.934s anon_fault_alloc: 6328991 anon_fault_fallback: 214553 *** Executing round 1 *** real 1m55.370s user 25m24.482s sys 3m52.512s anon_fault_alloc: 6355263 anon_fault_fallback: 108692 *** Executing round 2 *** real 2m7.579s user 25m11.530s sys 3m45.456s anon_fault_alloc: 6355816 anon_fault_fallback: 107852 *** Executing round 3 *** real 1m53.824s user 25m26.774s sys 3m42.160s anon_fault_alloc: 6355457 anon_fault_fallback: 107705 W/patch: *** Executing round 0 *** real 1m55.906s user 25m16.985s sys 4m24.480s anon_fault_alloc: 6354486 anon_fault_fallback: 109845 *** Executing round 1 *** real 1m51.303s user 25m16.456s sys 3m26.629s anon_fault_alloc: 6392515 anon_fault_fallback: 69797 *** Executing round 2 *** real 1m51.495s user 25m13.075s sys 3m28.501s anon_fault_alloc: 6395096 anon_fault_fallback: 67510 *** Executing round 3 *** real 1m52.980s user 25m8.217s sys 3m37.506s anon_fault_alloc: 6389556 anon_fault_fallback: 72743 Before "Executing round 0", mTHP is not enabled. Therefore, both cases show a higher anon_fault_fallback in Round 0 than in the other rounds. With the patch, however, memory can be compacted faster into an mTHP-friendly state, resulting in a much lower fallback rate in Round 0. In the other rounds, the patch also consistently shows a lower anon_fault_fallback, as well as lower sys and wall time for the kernel build. Open questions: 1. When multiple mTHP orders are enabled (e.g., order-2 and order-4 both "always"), this series only targets the minimum order. Should proactive compaction also independently evaluate and serve higher orders? 2. In skip_isolation_on_order(), the filter order ideally should come from the compaction control path. However, during proactive compaction target_order is always -1 (via compact_memory), and there is no clean way to pass the mTHP order down from upper layers. Currently we read huge_anon_orders_always directly, but this variable can be changed by userspace at any time, making the semantic fragile (the compaction may start with one order target and finish with another). Ideas on how to plumb the target order through the proactive compaction path cleanly are welcome. 3. In __compact_finished(), the original code skips proactive compaction when kswapd is running to avoid interference. Patch 3 removes this skip for non-costly mTHP orders (< PAGE_ALLOC_COSTLY_ORDER). The reason is that small-order compaction is lightweight and likely to succeed quickly even while kswapd is reclaiming, forming an order-2 block requires migrating very few pages. Does this approach make sense, or is there a better way to coordinate proactive compaction with kswapd in the mTHP scenario? 4. In the direct reclaim path (__alloc_pages_slowpath), compact_first is only set for costly orders or non-movable allocations. For mTHP always- enabled non-costly orders (e.g., order-2 MIGRATE_MOVABLE), when free memory is sufficient (watermarks met) but fragmentation is high, compaction is more appropriate than reclaim. Should we also set compact_first for this case to avoid unnecessary reclaim? Bo Zhang (4): mm: compaction: make proactive compaction mTHP-aware mm: compaction: skip isolating large folios that satisfy the mTHP order mm: compaction: don't skip proactive compaction for non-costly mTHP mm: adjust free_pages to make __zone_watermark_ok() mTHP-aware mm/compaction.c | 83 ++++++++++++++++++++++++++++++++++++++----------- mm/internal.h | 3 ++ mm/vmscan.c | 23 +++----------- 3 files changed, 72 insertions(+), 37 deletions(-) -- 2.34.1
This is the friendly patch-bot of Lorenzo Stoakes.
You have sent him a patch/series that has triggered this response.
He used to manually respond to these common problems, but in order to save
his sanity (he kept writing the same thing over and over, yet to different
people), I was created.
Hopefully you will not take offence and will fix the problem in your patch
and resubmit it so that it can be accepted into the Linux kernel tree.
When sending emails to mm:
For one of several possible reasons this mail has triggered an AI
detector script.
Note that, while we are fine with AI assistance, it is kernel
policy that you must disclose this with an Assisted-by tag like:
Assisted-by: LLM
See https://docs.kernel.org/process/coding-assistants.html
It is also kernel policy that you must fully understand and take
responsibility for every patch that you send.
See https://docs.kernel.org/process/generated-content.html most
notably:
If tools permit you to generate a contribution automatically, expect
additional scrutiny in proportion to how much of it was generated.
As with the output of any tooling, the result may be incorrect or
inappropriate. You are expected to understand and to be able to
defend everything you submit. If you are unable to do so, then do
not submit the resulting changes.
If you do so anyway, maintainers are entitled to reject your series
without detailed review.
In general, if you are a newcomer to mm, we expect you to do smaller work
before moving on to larger changes, so you build understanding of both the
technical aspects of mm and how we do things.
If you wish to discuss this problem further, or you have questions about
how to resolve this issue, please feel free to respond to this email and
Lorenzo will reply once he has dug out from the pending patches received
from other developers.
thanks,
Lorenzo's patch email bot
[ Idea shamelessly stolen from greg-kh ]
--
Cheers, Lorenzo
Hi Lorenzo, Thanks for the note. Let me clarify the background of this series, both regarding the AI-assistance question and regarding the scope of the change. This work started as part of our exploration of 16KB mTHP for the Android scenario. The overall direction of making compaction lighter-weight and mTHP-aware came from offline discussions with David, who suggested this lightweight-compaction angle for small mTHP orders. So the size and scope of the series is not me diving in unguided as a newcomer: it follows his directional guidance. Barry then provided reviews and change suggestions throughout the iterations. On the AI-assistance side: - All the code was written by me, by hand. I understand each change and take full responsibility for the patches. - The cover letter and some commit messages were polished with an LLM after I wrote the initial drafts (to smooth the English). The technical content, the design (jointly with Barry), and the analysis are mine. Please let me know whether an Assisted-by: LLM tag is expected for the cover / commit-message polishing; I'm happy to add it in v2 if that's the convention. Bo
On Tue, Sep 22, 2026 at 12:15 PM Bo Zhang <zhangbo0325@gmail.com> wrote: > > Hi Lorenzo, > > Thanks for the note. Let me clarify the background of this series, both > regarding the AI-assistance question and regarding the scope of the change. > > This work started as part of our exploration of 16KB mTHP for the Android > scenario. The overall direction of making compaction lighter-weight and > mTHP-aware came from offline discussions with David, who suggested this > lightweight-compaction angle for small mTHP orders. So the size and scope > of the series is not me diving in unguided as a newcomer: it follows his > directional guidance. Barry then provided reviews and change suggestions > throughout the iterations. > > On the AI-assistance side: > > - All the code was written by me, by hand. I understand each change and > take full responsibility for the patches. > > - The cover letter and some commit messages were polished with an LLM > after I wrote the initial drafts (to smooth the English). The technical > content, the design (jointly with Barry), and the analysis are mine. Maybe Lorenzo can adjust the bot a little so that it doesn't trigger an AI warning just because the changelog or commit message looks AI-polished? Perhaps the AI detector could focus more on the code changes instead? Please have a little mercy on us non-native English speakers. :-) Thanks Barry
On Tue Aug 25, 2026 at 12:38 AM EDT, Bo Zhang wrote: > Hi all, > > This series improves memory compaction to better serve mTHP (multi-size > THP) allocations, particularly for small orders like order-2 (16KB). > The changes cover the proactive compaction and kswapd-triggered compaction > paths. > > Problem: > > The current proactive compaction targets COMPACTION_HPAGE_ORDER (order-9, > 2MB), which is unfriendly to mTHP in two ways: > > 1. It migrates folios that already satisfy mTHP allocation needs. For > example, an order-2 folio is a valid mTHP page, yet compaction still > moves it around trying to form order-9 blocks. This is unnecessary > work and wastes energy. But skip_isolation_on_order() skips a folio with an order >= target order. > > 2. Compaction designed for 2MB huge pages is too heavyweight for mTHP. Yes. > mTHP allocations are frequent and only require small contiguous blocks > (e.g., 4 pages for order-2). A lighter-weight, mTHP-aware compaction > strategy is needed to reduce overhead and power consumption. > > Approach: > > This series makes four changes: > > 1. Generalize the fragmentation score functions to accept an order parameter > and use the minimum always-enabled mTHP order as the compaction target. Makes sense. > > 2. During proactive compaction (compact_memory), skip isolating folios that > already satisfy mTHP requirements, avoiding unnecessary migration overhead. Oh, you are targeting proactive compaction, where skip_isolation_on_order() does not apply. > > 3. Allow proactive compaction to proceed concurrently with kswapd for > non-costly mTHP orders, since kswapd reclaim alone may not produce the > contiguous blocks needed for these allocations. > > 4. Introduce zone_effective_free_pages() that provides mTHP-aware free page > accounting for watermark checks, counting only buddy blocks that can > actually satisfy mTHP allocations. > > Test setup: > > Boot Ubuntu with 2700 MB of memory, with mTHP disabled > initially and the defrag mode set to defer+madvise. > > Then set the 16 KB mTHP size to always and run a kernel > build with -j20. > > For both cases below, we run a background script to > proactively trigger compaction as follows: > #!/bin/bash > > while true; do > echo 50 > /proc/sys/vm/compaction_proactiveness > sleep 0.1 > done > > W/o patch: > > *** Executing round 0 *** > > real 2m1.312s > user 25m38.968s > sys 5m16.934s > anon_fault_alloc: 6328991 > anon_fault_fallback: 214553 > > *** Executing round 1 *** > > real 1m55.370s > user 25m24.482s > sys 3m52.512s > anon_fault_alloc: 6355263 > anon_fault_fallback: 108692 > > *** Executing round 2 *** > > real 2m7.579s > user 25m11.530s > sys 3m45.456s > anon_fault_alloc: 6355816 > anon_fault_fallback: 107852 > > *** Executing round 3 *** > > real 1m53.824s > user 25m26.774s > sys 3m42.160s > anon_fault_alloc: 6355457 > anon_fault_fallback: 107705 > > W/patch: > > *** Executing round 0 *** > > real 1m55.906s > user 25m16.985s > sys 4m24.480s > anon_fault_alloc: 6354486 > anon_fault_fallback: 109845 > > *** Executing round 1 *** > > real 1m51.303s > user 25m16.456s > sys 3m26.629s > anon_fault_alloc: 6392515 > anon_fault_fallback: 69797 > > *** Executing round 2 *** > > real 1m51.495s > user 25m13.075s > sys 3m28.501s > anon_fault_alloc: 6395096 > anon_fault_fallback: 67510 > > *** Executing round 3 *** > > real 1m52.980s > user 25m8.217s > sys 3m37.506s > anon_fault_alloc: 6389556 > anon_fault_fallback: 72743 > > Before "Executing round 0", mTHP is not enabled. Therefore, both > cases show a higher anon_fault_fallback in Round 0 than in the > other rounds. With the patch, however, memory can be compacted faster > into an mTHP-friendly state, resulting in a much lower fallback rate > in Round 0. In the other rounds, the patch also consistently shows > a lower anon_fault_fallback, as well as lower sys and wall time for > the kernel build. What about the impact on THP compaction? How does it affect direct compaction for both mTHP and THP? It sounds to me that this patch series target proactive compaction. Am I getting right? Thanks. > > Open questions: > > 1. When multiple mTHP orders are enabled (e.g., order-2 and order-4 both > "always"), this series only targets the minimum order. Should proactive > compaction also independently evaluate and serve higher orders? > > 2. In skip_isolation_on_order(), the filter order ideally should come from > the compaction control path. However, during proactive compaction > target_order is always -1 (via compact_memory), and there is no clean > way to pass the mTHP order down from upper layers. Currently we read > huge_anon_orders_always directly, but this variable can be changed by > userspace at any time, making the semantic fragile (the compaction may > start with one order target and finish with another). Ideas on how to > plumb the target order through the proactive compaction path cleanly > are welcome. > > 3. In __compact_finished(), the original code skips proactive compaction > when kswapd is running to avoid interference. Patch 3 removes this > skip for non-costly mTHP orders (< PAGE_ALLOC_COSTLY_ORDER). The > reason is that small-order compaction is lightweight and likely to > succeed quickly even while kswapd is reclaiming, forming an order-2 > block requires migrating very few pages. Does this approach make sense, > or is there a better way to coordinate proactive compaction with kswapd > in the mTHP scenario? > > 4. In the direct reclaim path (__alloc_pages_slowpath), compact_first is > only set for costly orders or non-movable allocations. For mTHP always- > enabled non-costly orders (e.g., order-2 MIGRATE_MOVABLE), when free > memory is sufficient (watermarks met) but fragmentation is high, > compaction is more appropriate than reclaim. Should we also set > compact_first for this case to avoid unnecessary reclaim? > > Bo Zhang (4): > mm: compaction: make proactive compaction mTHP-aware > mm: compaction: skip isolating large folios that satisfy the mTHP order > mm: compaction: don't skip proactive compaction for non-costly mTHP > mm: adjust free_pages to make __zone_watermark_ok() mTHP-aware > > mm/compaction.c | 83 ++++++++++++++++++++++++++++++++++++++----------- > mm/internal.h | 3 ++ > mm/vmscan.c | 23 +++----------- > 3 files changed, 72 insertions(+), 37 deletions(-) > > -- > 2.34.1 -- Best Regards, Yan, Zi
On Sun Sep 07, 2026 at 10:51 PM EDT, Zi Yan wrote: > But skip_isolation_on_order() skips a folio with an order >= target > order. > ... > Oh, you are targeting proactive compaction, where > skip_isolation_on_order() does not apply. Right. To clarify the cover letter: the "migrating folios that already satisfy mTHP" concern is specific to proactive compaction, where target_order is -1 (via compact_memory) and the order >= target_order check in skip_isolation_on_order() does not apply. For compaction with an explicit target order that path already handles it. > What about the impact on THP compaction? How does it affect direct > compaction for both mTHP and THP? > > It sounds to me that this patch series target proactive compaction. Am I > getting right? Two things: 1) Traditional THP (order-9) is not affected. The mTHP-aware branch in zone_effective_free_pages() only triggers when order == compact_hpage_order(), i.e. the minimum always-enabled mTHP order (e.g. order-2). An order-9 THP request does not match that, so it keeps its original behavior exactly (NR_FREE_PAGES_BLOCKS under defrag_mode, NR_FREE_PAGES otherwise). We didn't change the THP path. 2) The series isn't limited to proactive compaction. Patches 1-3 target proactive compaction, but patch 4 also covers the kswapd -> kcompactd path via pgdat_balanced(), so both proactive compaction and kswapd wakeup are addressed. For direct compaction: patch 4 does touch compaction_suit_allocation_order(), which is shared with direct compaction, so order-2 mTHP direct compaction would also fall into the new accounting. However, the direct compaction case needs more testing and thought. How direct compaction should behave for mTHP is something worth discussing together to decide the right approach. Thanks for the review. Bo
© 2016 - 2026 Red Hat, Inc.