.../ABI/testing/sysfs-kernel-mm-damon | 7 ++ Documentation/mm/damon/design.rst | 5 + include/linux/damon.h | 15 ++- mm/damon/core.c | 2 + mm/damon/sysfs-schemes.c | 48 +++++++ mm/damon/vaddr.c | 90 +++++++++++++ tools/testing/selftests/damon/Makefile | 1 + tools/testing/selftests/damon/_damon_sysfs.py | 9 +- tools/testing/selftests/damon/damos_split.py | 99 +++++++++++++++ tools/testing/selftests/damon/sysfs.py | 11 +- 10 files changed, 279 insertions(+), 8 deletions(-)
DAMOS_SPLIT splits large folios in a target region down to a
configured target order, using the existing split_folio_to_order().
No new core-mm code or exported symbols are introduced.
Based on mm-unstable at 61cccb8363fc ("mm/swap, PM: hibernate:
atomically replace hibernation pin").
Different addresses within a PMD-mapped folio resolve to the same
PMD Accessed bit. Accesses to a small part of the folio can
therefore coarsen DAMON's observed hot set relative to the actual
working set.
DAMOS already provides promotion actions (HUGEPAGE, COLLAPSE) but
has no corresponding demotion action. DAMOS_SPLIT fills this gap.
It is a mechanism, not a policy -- it does not decide which folios
to split. Selection is left to DAMON's existing access patterns,
filters, and future probe/PMU signals.
target_order selects the split target: 0 for order-0 base pages,
or a supported smaller mTHP order. Both anonymous and file-backed
folios are supported. The locking follows split_huge_pages_in_pid()
in mm/huge_memory.c.
Tests
=====
damos_split.py (VM + Kunpeng 920):
anon THP -> order-0 split: PASS
sangfor_exp.py (Kunpeng 920, tmpfs, 4096 MiB):
- Created a 4 GiB PMD-mapped tmpfs workload.
- Applied DAMOS_SPLIT with target_order=0.
- ShmemPmdMapped dropped from 4194304 KiB to 0 in every round.
- Repeated for five rounds without functional failures.
The functional selftest (damos_split.py) is included in this series.
Additional experiment scripts and raw results are available on
request. Performance characterization using masim [1] and KMB [2]
is in progress.
[1] https://github.com/sjp38/masim
[2] https://gitee.com/OpenCloudOS/kernel-multi-bench
Open questions
==============
- Selection policy: this series keeps folio selection outside the
action and relies on DAMOS access patterns, filters, and quotas.
Is this the appropriate layering for future probe-based signals?
- Hysteresis: khugepaged may re-collapse a just-split folio.
Should cooldown live in DAMON policy or khugepaged?
- File-backed folios: adjust target_order upward to filesystem
minimum, or keep current "fail and skip"?
Beyond the action API itself, feedback on real workloads that need
proactive large-folio demotion is particularly welcome. Follow-up
work will evaluate candidate selection signals, including DAMON
probes and hardware-assisted sampling, as well as target-order
selection and split/collapse hysteresis. Those policies are
intentionally kept outside this series.
Changes since v2 [3]
====================
- Split-only series (collapse deferred).
- Dropped SPE feedback (mechanism/policy separation).
- DAMOS_MTHP_SPLIT -> DAMOS_SPLIT.
- order field in existing union (no struct size increase).
- Added functional selftest (damos_split.py).
- checkpatch: 0 errors, 0 warnings.
[3] https://lore.kernel.org/20260701123000.00000-1-lianux.mm@gmail.com/
Lian Wang (Processmission) (3):
mm/damon: introduce DAMOS_SPLIT action
mm/damon/vaddr: implement DAMOS_SPLIT handler
selftests/damon: add functional test for DAMOS_SPLIT
.../ABI/testing/sysfs-kernel-mm-damon | 7 ++
Documentation/mm/damon/design.rst | 5 +
include/linux/damon.h | 15 ++-
mm/damon/core.c | 2 +
mm/damon/sysfs-schemes.c | 48 +++++++
mm/damon/vaddr.c | 90 +++++++++++++
tools/testing/selftests/damon/Makefile | 1 +
tools/testing/selftests/damon/_damon_sysfs.py | 9 +-
tools/testing/selftests/damon/damos_split.py | 99 +++++++++++++++
tools/testing/selftests/damon/sysfs.py | 11 +-
10 files changed, 279 insertions(+), 8 deletions(-)
You were warned at not Cc-ing THP developers in the previous revision, since it
was modifying THP source code. Since this version is modifying only DAMON
source code, you don't really need to Cc more than DAMON developers. Asking
wider inputs is good practice. And it worked very well. We got great inputs
from Asier, David and Zi. But some people don't really like having too much
mails in their inbox. My personal rule of thumb is just running
get_maintainer.pl via 'hkml patch format' [1]. You must have your own rule,
though :)
On Mon, 20 Jul 2026 11:03:24 +0800 Lian Wang <lianux.mm@gmail.com> wrote:
> DAMOS_SPLIT splits large folios in a target region down to a
> configured target order, using the existing split_folio_to_order().
> No new core-mm code or exported symbols are introduced.
The last sentence may better to go to changelog. If you want to highlight, you
can put changelog at the top of the cover letter.
>
> Based on mm-unstable at 61cccb8363fc ("mm/swap, PM: hibernate:
> atomically replace hibernation pin").
This is also not feasible to be the cover letter main content.
>
> Different addresses within a PMD-mapped folio resolve to the same
> PMD Accessed bit. Accesses to a small part of the folio can
> therefore coarsen DAMON's observed hot set relative to the actual
> working set.
>
> DAMOS already provides promotion actions (HUGEPAGE, COLLAPSE) but
> has no corresponding demotion action. DAMOS_SPLIT fills this gap.
> It is a mechanism, not a policy -- it does not decide which folios
> to split. Selection is left to DAMON's existing access patterns,
> filters, and future probe/PMU signals.
As I commented to the previous version [2], this sounds like you are saying two
very different things. Monitoring quality degradation issue and better THP
handling of DAMOS. This makes understanding the motivation of this series
difficult, as other people also pointed out.
Based on your replies to others, now I understand you are proposing DAMOS_SPLIT
as a way for improving the monitoring results. I'm waiting for your more
clarification of the issue, to better assess if this makes sense or not, as I
commented on the reply to Zi's reply.
>
> target_order selects the split target: 0 for order-0 base pages,
> or a supported smaller mTHP order. Both anonymous and file-backed
> folios are supported. The locking follows split_huge_pages_in_pid()
> in mm/huge_memory.c.
>
> Tests
> =====
>
> damos_split.py (VM + Kunpeng 920):
> anon THP -> order-0 split: PASS
>
> sangfor_exp.py (Kunpeng 920, tmpfs, 4096 MiB):
>
> - Created a 4 GiB PMD-mapped tmpfs workload.
> - Applied DAMOS_SPLIT with target_order=0.
> - ShmemPmdMapped dropped from 4194304 KiB to 0 in every round.
> - Repeated for five rounds without functional failures.
>
> The functional selftest (damos_split.py) is included in this series.
> Additional experiment scripts and raw results are available on
> request. Performance characterization using masim [1] and KMB [2]
> is in progress.
It is completely fine to keep having tests in progress. But, please make the
story complete. What damos_split.py and sangfor_exp.py do? What the results
mean? What the performance tests will do with what expectation?
>
> [1] https://github.com/sjp38/masim
> [2] https://gitee.com/OpenCloudOS/kernel-multi-bench
>
> Open questions
> ==============
>
> - Selection policy: this series keeps folio selection outside the
> action and relies on DAMOS access patterns, filters, and quotas.
> Is this the appropriate layering for future probe-based signals?
You mentioned this series is for monitoring quality improvement. If so,
shouldn't you just apply it to all THPs, regardless of the access pattern? I'm
again being confused. More clarification of the motivation would be useful.
>
> - Hysteresis: khugepaged may re-collapse a just-split folio.
> Should cooldown live in DAMON policy or khugepaged?
Ditto.
>
> - File-backed folios: adjust target_order upward to filesystem
> minimum, or keep current "fail and skip"?
I don't fully understand the question. Could you please elaborate more?
>
> Beyond the action API itself, feedback on real workloads that need
> proactive large-folio demotion is particularly welcome. Follow-up
> work will evaluate candidate selection signals, including DAMON
> probes and hardware-assisted sampling, as well as target-order
> selection and split/collapse hysteresis. Those policies are
> intentionally kept outside this series.
You mentioned this work is for monitoring quality improvement. Now you are
saying somewhat followup. I'm again being confused.
>
> Changes since v2 [3]
> ====================
>
> - Split-only series (collapse deferred).
> - Dropped SPE feedback (mechanism/policy separation).
> - DAMOS_MTHP_SPLIT -> DAMOS_SPLIT.
> - order field in existing union (no struct size increase).
> - Added functional selftest (damos_split.py).
> - checkpatch: 0 errors, 0 warnings.
>
> [3] https://lore.kernel.org/20260701123000.00000-1-lianux.mm@gmail.com/
>
> Lian Wang (Processmission) (3):
> mm/damon: introduce DAMOS_SPLIT action
> mm/damon/vaddr: implement DAMOS_SPLIT handler
> selftests/damon: add functional test for DAMOS_SPLIT
>
> .../ABI/testing/sysfs-kernel-mm-damon | 7 ++
> Documentation/mm/damon/design.rst | 5 +
> include/linux/damon.h | 15 ++-
> mm/damon/core.c | 2 +
> mm/damon/sysfs-schemes.c | 48 +++++++
> mm/damon/vaddr.c | 90 +++++++++++++
> tools/testing/selftests/damon/Makefile | 1 +
> tools/testing/selftests/damon/_damon_sysfs.py | 9 +-
> tools/testing/selftests/damon/damos_split.py | 99 +++++++++++++++
> tools/testing/selftests/damon/sysfs.py | 11 +-
> 10 files changed, 279 insertions(+), 8 deletions(-)
[1] https://github.com/sjp38/hackermail/blob/master/USAGE.md#formatting-patches
[2] https://lore.kernel.org/20260702183551.91007-1-sj@kernel.org
Thanks,
SJ
On 7/20/26 04:03, Lian Wang wrote:
> DAMOS_SPLIT splits large folios in a target region down to a
> configured target order, using the existing split_folio_to_order().
> No new core-mm code or exported symbols are introduced.
>
> Based on mm-unstable at 61cccb8363fc ("mm/swap, PM: hibernate:
> atomically replace hibernation pin").
>
> Different addresses within a PMD-mapped folio resolve to the same
> PMD Accessed bit. Accesses to a small part of the folio can
> therefore coarsen DAMON's observed hot set relative to the actual
> working set.
>
> DAMOS already provides promotion actions (HUGEPAGE, COLLAPSE) but
> has no corresponding demotion action. DAMOS_SPLIT fills this gap.
> It is a mechanism, not a policy -- it does not decide which folios
> to split. Selection is left to DAMON's existing access patterns,
> filters, and future probe/PMU signals.
>
> target_order selects the split target: 0 for order-0 base pages,
> or a supported smaller mTHP order. Both anonymous and file-backed
> folios are supported. The locking follows split_huge_pages_in_pid()
> in mm/huge_memory.c.
>
> Tests
> =====
>
> damos_split.py (VM + Kunpeng 920):
> anon THP -> order-0 split: PASS
>
> sangfor_exp.py (Kunpeng 920, tmpfs, 4096 MiB):
>
> - Created a 4 GiB PMD-mapped tmpfs workload.
> - Applied DAMOS_SPLIT with target_order=0.
> - ShmemPmdMapped dropped from 4194304 KiB to 0 in every round.
> - Repeated for five rounds without functional failures.
>
> The functional selftest (damos_split.py) is included in this series.
> Additional experiment scripts and raw results are available on
> request. Performance characterization using masim [1] and KMB [2]
> is in progress.
>
> [1] https://github.com/sjp38/masim
> [2] https://gitee.com/OpenCloudOS/kernel-multi-bench
Hi,
you give no real motivation and evaluation why this is required or why this
gives the user any benefit.
HUGEPAGE + COLLAPSE is clear: give me THPs in a size not controlled by user
space, because the expectation is that this memory will be performance sensitive.
A SPLIT with an explicit order is not really want we want and it does not fit
the existing primitives.
--
Cheers,
David
Hi David,
On 7/20/2026 10:44 AM, David Hildenbrand (Arm) wrote:
> you give no real motivation and evaluation why this is required or
> why this gives the user any benefit.
> A SPLIT with an explicit order is not really want we want and it
> does not fit the existing primitives.
Thank you for the direct feedback. Let me explain where this came
from -- the cover letter should have included this context.
This started from a real problem at Sangfor. The scenario is:
KVM-QEMU virtualization on Kunpeng 920, with KVM guest memory
backed by tmpfs shared mappings (THP=always on the host). An
Oracle database runs inside the VM. DAMON monitors the KVM
process on the host to measure the hot-memory ratio.
The KVM process allocates and uses a large amount of memory.
Under the same workload, DAMON reports a significantly higher
hot-memory ratio with THP enabled versus THP disabled. Direct
tmpfs write tests inside the VM -- touching at 4K and 2M
strides -- show a clear gap between the two cases.
DAMON parameters used:
operations=vaddr
monitoring_attrs/nr_regions/min=500
monitoring_attrs/nr_regions/max=2000
monitoring_attrs/intervals/sample_us=500000
monitoring_attrs/intervals/aggr_us=20000000
monitoring_attrs/intervals/update_us=60000000
schemes/0/action=stat
schemes/0/access_pattern/nr_accesses/min=1
schemes/0/access_pattern/nr_accesses/max=max
The underlying issue is that under PMD-mapped THP, DAMON's monitoring
granularity is coarser than the actual working set -- a single
Accessed bit covers 512 base pages. Before SJ's probe infrastructure
arrives, there is a gap: DAMON cannot distinguish hot sub-pages from
cold ones within a single THP.
Split is one possible mechanism to bridge that gap -- by dismantling
the PMD mapping, each base page gets its own PTE Accessed bit and
DAMON recovers fine-grain monitoring. It is not intended to be a
permanent API, and certainly not "the opposite of collapse".
I did not write this scenario into the cover letter because our test
results do not yet show a clear quantitative benefit worth claiming,
and I did not want to oversell. Without the context, I understand it
looks like I randomly proposed a new primitive -- that was not the
intention.
SJ acknowledged [1] that the monitoring problem under THP is real.
My RFC is a concrete proposal to start the discussion. If split with
an explicit order is not the right primitive, I would appreciate your
thoughts on what the correct DAMOS abstraction for this should be.
[1] https://lore.kernel.org/20260620203915.82947-1-sj@kernel.org/
Thanks,
Lian Wang
On Mon Jul 20, 2026 at 5:56 AM EDT, Lian Wang wrote: > Hi David, > > On 7/20/2026 10:44 AM, David Hildenbrand (Arm) wrote: >> you give no real motivation and evaluation why this is required or >> why this gives the user any benefit. >> A SPLIT with an explicit order is not really want we want and it >> does not fit the existing primitives. > > Thank you for the direct feedback. Let me explain where this came > from -- the cover letter should have included this context. > > This started from a real problem at Sangfor. The scenario is: > > KVM-QEMU virtualization on Kunpeng 920, with KVM guest memory > backed by tmpfs shared mappings (THP=always on the host). An > Oracle database runs inside the VM. DAMON monitors the KVM > process on the host to measure the hot-memory ratio. > > The KVM process allocates and uses a large amount of memory. > Under the same workload, DAMON reports a significantly higher > hot-memory ratio with THP enabled versus THP disabled. Direct > tmpfs write tests inside the VM -- touching at 4K and 2M > strides -- show a clear gap between the two cases. > > DAMON parameters used: > > operations=vaddr > monitoring_attrs/nr_regions/min=500 > monitoring_attrs/nr_regions/max=2000 > monitoring_attrs/intervals/sample_us=500000 > monitoring_attrs/intervals/aggr_us=20000000 > monitoring_attrs/intervals/update_us=60000000 > schemes/0/action=stat > schemes/0/access_pattern/nr_accesses/min=1 > schemes/0/access_pattern/nr_accesses/max=max > > The underlying issue is that under PMD-mapped THP, DAMON's monitoring > granularity is coarser than the actual working set -- a single > Accessed bit covers 512 base pages. Before SJ's probe infrastructure > arrives, there is a gap: DAMON cannot distinguish hot sub-pages from > cold ones within a single THP. > > Split is one possible mechanism to bridge that gap -- by dismantling > the PMD mapping, each base page gets its own PTE Accessed bit and > DAMON recovers fine-grain monitoring. It is not intended to be a > permanent API, and certainly not "the opposite of collapse". If you just want PTE level access bit information, why not split PMD mapping instead of the THP itself? In addition, the issue is about access monitoring granularity in DAMON, why should user care and know about THP split operations? I would expect DAMON detects the inability of getting fine grain access information and split the PMD mapping itself instead of a user initiated DAMON_SPLIT. If that is not possible with DAMON, an alternative is to provide something more generic like DAMON_SAMPLE, which does the split under the hood, instead of exposing MM internal operations. > > I did not write this scenario into the cover letter because our test > results do not yet show a clear quantitative benefit worth claiming, > and I did not want to oversell. Without the context, I understand it > looks like I randomly proposed a new primitive -- that was not the > intention. > > SJ acknowledged [1] that the monitoring problem under THP is real. > My RFC is a concrete proposal to start the discussion. If split with > an explicit order is not the right primitive, I would appreciate your > thoughts on what the correct DAMOS abstraction for this should be. > > [1] https://lore.kernel.org/20260620203915.82947-1-sj@kernel.org/ > > Thanks, > Lian Wang -- Best Regards, Yan, Zi
Hello, On Mon, 20 Jul 2026 15:12:50 -0400 "Zi Yan" <ziy@nvidia.com> wrote: > On Mon Jul 20, 2026 at 5:56 AM EDT, Lian Wang wrote: > > Hi David, > > > > On 7/20/2026 10:44 AM, David Hildenbrand (Arm) wrote: > >> you give no real motivation and evaluation why this is required or > >> why this gives the user any benefit. > >> A SPLIT with an explicit order is not really want we want and it > >> does not fit the existing primitives. > > > > Thank you for the direct feedback. Let me explain where this came > > from -- the cover letter should have included this context. > > > > This started from a real problem at Sangfor. The scenario is: > > > > KVM-QEMU virtualization on Kunpeng 920, with KVM guest memory > > backed by tmpfs shared mappings (THP=always on the host). An > > Oracle database runs inside the VM. DAMON monitors the KVM > > process on the host to measure the hot-memory ratio. > > > > The KVM process allocates and uses a large amount of memory. > > Under the same workload, DAMON reports a significantly higher > > hot-memory ratio with THP enabled versus THP disabled. Direct > > tmpfs write tests inside the VM -- touching at 4K and 2M > > strides -- show a clear gap between the two cases. > > > > DAMON parameters used: > > > > operations=vaddr > > monitoring_attrs/nr_regions/min=500 > > monitoring_attrs/nr_regions/max=2000 > > monitoring_attrs/intervals/sample_us=500000 > > monitoring_attrs/intervals/aggr_us=20000000 > > monitoring_attrs/intervals/update_us=60000000 Thank you for sharing your detailed setup. It is helpful. Btw, have you considered using intervals auto-tuning [1]? > > schemes/0/action=stat > > schemes/0/access_pattern/nr_accesses/min=1 > > schemes/0/access_pattern/nr_accesses/max=max > > > > The underlying issue is that under PMD-mapped THP, DAMON's monitoring > > granularity is coarser than the actual working set -- a single > > Accessed bit covers 512 base pages. Before SJ's probe infrastructure > > arrives, there is a gap: DAMON cannot distinguish hot sub-pages from > > cold ones within a single THP. > > > > Split is one possible mechanism to bridge that gap -- by dismantling > > the PMD mapping, each base page gets its own PTE Accessed bit and > > DAMON recovers fine-grain monitoring. Thank you for clarifying the motivation of this series. To me, it's still unclear what is the real user impact, though. I mean, I can understand DAMON suddenly reporting more hot memory can surprise some people. But, why that matters in what extent for your use case? You may not run DAMON on your system only to read the information. You may run it to do something beneficial using the information. What is that, and how badly degraded DAMON's monitoring results affect it? Overall, unless the real impact is serious, splitting huge pages only for better DAMON monitoring sounds like not a good tradeoff. You will increase DAMON overhead and lose THP benefits in some extent. > > It is not intended to be a > > permanent API, and certainly not "the opposite of collapse". Once it is added to the kernel, we have to support it for long term. Let's not introduce something for only temporal use. > > If you just want PTE level access bit information, why not split PMD > mapping instead of the THP itself? > > In addition, the issue is about access monitoring granularity in DAMON, > why should user care and know about THP split operations? I would expect > DAMON detects the inability of getting fine grain access information and > split the PMD mapping itself instead of a user initiated DAMON_SPLIT. If > that is not possible with DAMON, an alternative is to provide something > more generic like DAMON_SAMPLE, which does the split under the hood, > instead of exposing MM internal operations. Thank you for good opinion, Zi. I agree all the points. That said, I still want to understand the problem first. > > > > > I did not write this scenario into the cover letter because our test > > results do not yet show a clear quantitative benefit worth claiming, > > and I did not want to oversell. Without the context, I understand it > > looks like I randomly proposed a new primitive -- that was not the > > intention. > > > > SJ acknowledged [1] that the monitoring problem under THP is real. Yes, the behavior is real and I agree your theory of how it happens. I don't clearly understand if it is really bad in what situations, though. > > My RFC is a concrete proposal to start the discussion. If split with > > an explicit order is not the right primitive, I would appreciate your > > thoughts on what the correct DAMOS abstraction for this should be. Only after understanding what is the problem and how bad it is, we will be able to think of different approaches and assess those. To me, it is still unclear what is the real problem and how bad it is. I will wait for your further clarifications of those. > > > > [1] https://lore.kernel.org/20260620203915.82947-1-sj@kernel.org/ [1] https://origin.kernel.org/doc/html/latest/mm/damon/design.html#monitoring-intervals-auto-tuning Thanks, SJ [...]
Hi SJ, On 7/20/2026 5:47 PM, SJ Park wrote: > To me, it's still unclear what is the real user impact, though. > ... > Only after understanding what is the problem and how bad it is, we > will be able to think of different approaches and assess those. You are right. I will work with Sangfor to quantify the real impact in their production scenario -- what DAMOS action is driven by the inflated hot-memory readings, and how badly the monitoring error affects the actual outcome. Thank you, David, and Zi Yan for the honest and constructive feedback. It made me realize I should have led with the problem, not the mechanism. Let me take a step back, gather concrete data on the problem severity, and come back with a clearer picture. I will follow up in this thread once I have something solid to share. Thanks, Lian Wang
Hi Lian,
On 7/20/2026 6:03 AM, Lian Wang wrote:
> DAMOS_SPLIT splits large folios in a target region down to a
> configured target order, using the existing split_folio_to_order().
> No new core-mm code or exported symbols are introduced.
>
> Based on mm-unstable at 61cccb8363fc ("mm/swap, PM: hibernate:
> atomically replace hibernation pin").
>
> Different addresses within a PMD-mapped folio resolve to the same
> PMD Accessed bit. Accesses to a small part of the folio can
> therefore coarsen DAMON's observed hot set relative to the actual
> working set.
>
> DAMOS already provides promotion actions (HUGEPAGE, COLLAPSE) but
> has no corresponding demotion action. DAMOS_SPLIT fills this gap.
> It is a mechanism, not a policy -- it does not decide which folios
> to split. Selection is left to DAMON's existing access patterns,
> filters, and future probe/PMU signals.
You should mention why page split is be needed. The fact that page
collapsing exist doesn't necessarily mean that split should exist.
I agree that it is a nice feature, but it should be backed in the
cover letter.
> target_order selects the split target: 0 for order-0 base pages,
> or a supported smaller mTHP order. Both anonymous and file-backed
> folios are supported. The locking follows split_huge_pages_in_pid()
> in mm/huge_memory.c.
>
> Tests
> =====
>
> damos_split.py (VM + Kunpeng 920):
> anon THP -> order-0 split: PASS
>
> sangfor_exp.py (Kunpeng 920, tmpfs, 4096 MiB):
>
> - Created a 4 GiB PMD-mapped tmpfs workload.
> - Applied DAMOS_SPLIT with target_order=0.
> - ShmemPmdMapped dropped from 4194304 KiB to 0 in every round.
> - Repeated for five rounds without functional failures.
>
> The functional selftest (damos_split.py) is included in this series.
> Additional experiment scripts and raw results are available on
> request. Performance characterization using masim [1] and KMB [2]
> is in progress.
>
> [1] https://github.com/sjp38/masim
> [2] https://gitee.com/OpenCloudOS/kernel-multi-bench
>
> Open questions
> ==============
>
> - Selection policy: this series keeps folio selection outside the
> action and relies on DAMOS access patterns, filters, and quotas.
> Is this the appropriate layering for future probe-based signals?
>
> - Hysteresis: khugepaged may re-collapse a just-split folio.
> Should cooldown live in DAMON policy or khugepaged?
>
> - File-backed folios: adjust target_order upward to filesystem
> minimum, or keep current "fail and skip"?
>
> Beyond the action API itself, feedback on real workloads that need
> proactive large-folio demotion is particularly welcome. Follow-up
> work will evaluate candidate selection signals, including DAMON
> probes and hardware-assisted sampling, as well as target-order
> selection and split/collapse hysteresis. Those policies are
> intentionally kept outside this series.
> Changes since v2 [3]
> ====================
>
> - Split-only series (collapse deferred).
> - Dropped SPE feedback (mechanism/policy separation).
> - DAMOS_MTHP_SPLIT -> DAMOS_SPLIT.
> - order field in existing union (no struct size increase).
> - Added functional selftest (damos_split.py).
> - checkpatch: 0 errors, 0 warnings.
>
> [3] https://lore.kernel.org/20260701123000.00000-1-lianux.mm@gmail.com/
Could you add v1 as well?
>
> Lian Wang (Processmission) (3):
> mm/damon: introduce DAMOS_SPLIT action
> mm/damon/vaddr: implement DAMOS_SPLIT handler
> selftests/damon: add functional test for DAMOS_SPLIT
>
> .../ABI/testing/sysfs-kernel-mm-damon | 7 ++
> Documentation/mm/damon/design.rst | 5 +
> include/linux/damon.h | 15 ++-
> mm/damon/core.c | 2 +
> mm/damon/sysfs-schemes.c | 48 +++++++
> mm/damon/vaddr.c | 90 +++++++++++++
> tools/testing/selftests/damon/Makefile | 1 +
> tools/testing/selftests/damon/_damon_sysfs.py | 9 +-
> tools/testing/selftests/damon/damos_split.py | 99 +++++++++++++++
> tools/testing/selftests/damon/sysfs.py | 11 +-
> 10 files changed, 279 insertions(+), 8 deletions(-)
>
--
Asier Gutierrez
Huawei
Hi Asier, Thanks for the quick feedback. On 7/20/2026 12:28 PM, Gutierrez Asier wrote: > You should mention why page split is be needed. The fact that page > collapsing exist doesn't necessarily mean that split should exist. Fair point. The underlying problem I'm trying to address is that DAMON's vaddr monitoring loses accuracy under PMD-mapped THP: multiple sampled addresses share a single Accessed bit, so the observed hot set is coarser than the true working set. Split is one way to restore fine-grain monitoring -- by dismantling the PMD mapping, each base page gets its own PTE Accessed bit and DAMON can see the real access distribution again. Split is not the only possible approach, and it is certainly not intended to be "the opposite of collapse". It is just one concrete proposal to start the discussion. What I really care about is whether the community agrees that this monitoring granularity problem is worth solving. If there are better ways to address it, I'm very open to that direction. The RFC is as much about the problem as it is about the mechanism. Feedback on real workloads that suffer from this coarsening, and on alternative approaches, is exactly what I'm hoping for. > Could you add v1 as well? Good catch, will add in the next revision. [1] https://lore.kernel.org/20260620203915.82947-1-sj@kernel.org/ Thanks, Lian Wang
© 2016 - 2026 Red Hat, Inc.