[RFC PATCH 0/8] mm/memcontrol: introduce per-tier memory accounting and control

liuqiqi@kylinos.cn posted 8 patches 1 month, 1 week ago
include/linux/cgroup-defs.h  |   5 +
include/linux/memcontrol.h   |  26 ++
include/linux/memory-tiers.h |  12 +
include/linux/swap.h         |   6 +
kernel/cgroup/cgroup.c       |  21 +
mm/memcontrol.c              | 786 ++++++++++++++++++++++++++++++++++-
mm/memory-tiers.c            |  58 +++
mm/vmscan.c                  |  26 +-
8 files changed, 936 insertions(+), 4 deletions(-)
[RFC PATCH 0/8] mm/memcontrol: introduce per-tier memory accounting and control
Posted by liuqiqi@kylinos.cn 1 month, 1 week ago
From: Qiqi Liu <liuqiqi@kylinos.cn>

This RFC introduces per-tier memory cgroup accounting. Each cgroup
tracks its memory usage per memory tier (e.g. DRAM, CXL), exposed
through a new memory.tier control file that reports per-tier usage
and accepts independent high (soft) and max (hard) limits per tier.
By default these limits are auto-derived from memory.high / memory.max
based on per-tier capacity ratios, and can be manually overridden.

The implementation integrates with the existing memory tiering and
demotion infrastructure. Per-tier usage (anonymous and file) is tracked
via dedicated page counters, and cross-tier migrations (e.g. demotion
from DRAM to CXL) correctly re-account charges. When a tier hits its
high limit, async reclaim is triggered within that tier's NUMA nodes;
exceeding max enforces reclaim scoped to the tier's own nodes, or OOM.

The feature is fully opt-in. When disabled, no extra counters or
charge/uncharge paths are created, memory.tier reads empty, and there
is no measurable overhead.

Why per-tier limits?
-------------------

On tiered memory systems, memory.max constrains total usage but cannot
express "keep fast-tier usage under X". Without per-tier limits, a
workload can monopolise DRAM, pushing other cgroups onto slower tiers.
This series gives each cgroup independent high (soft) and max (hard)
limits per tier, exposed and set through a new memory.tier file.

By default those limits auto-derive from memory.high / memory.max by
capacity ratio; writing memory.tier pins a tier.

This series takes a different approach from Joshua Hahn's toptier RFC [1],
tracking a separate page_counter per (memcg, tier) for N-tier support and
exposing writable per-tier limits under a cgroup mount option.

Patch structure
---------------

  1/8  mm/memory-tiers: add node_to_tier_id and tier_id_to_nodemask
  2/8  mm/vmscan: add try_to_free_mem_cgroup_pages_nodemask
  3/8  mm/memcontrol: add per-tier page counter infrastructure and lifecycle
  4/8  mm/memcontrol: add per-tier charge and uncharge
  5/8  mm/memcontrol: add per-cpu stock for tier charge/uncharge
  6/8  mm/memcontrol: add memory.tier control file
  7/8  mm/memcontrol: auto-derive tier high/max from memory.high/max
  8/8  cgroup: add memory_tiered_limits cgroup mount option

Patches 1-2 are infrastructure (helpers in memory-tiers and vmscan).
Patches 3-5 add the core accounting: counter lifecycle (3),
per-page charge/uncharge (4), and stock batching (5).
Patch 6 adds the userspace file. Patch 7 adds auto-derivation. Patch 8
gates everything behind a mount option + kernel cmdline, so the feature
adds no measurable overhead when not opted in.

Usage
-----

Boot with:

  cgroup_memory_tiered_limits=1

Or remount at runtime (affects newly created cgroups only):

  mount -o remount,memory_tiered_limits /sys/fs/cgroup

Per-tier limits and usage can then be read from and written to
memory.tier.

Scope and limitations
---------------------

- Only LRU folios (anonymous and file pages) are tier-accounted. Kernel
  memory and socket buffers are not yet accounted per tier; support for
  these is planned as follow-up work.
- Per-tier memory.min and memory.low protections are not implemented.
  These can be added later by extending the per-tier interface to
  expose and enforce min/low protection.
- The command-line parameter mirrors cgroup_favordynmods; automatic
  enablement via the cgroup mount path is left to userspace.

Testing
-------

Tested on QEMU with fake NUMA (DRAM tier 4 + CXL tier 22), with
cgroup_memory_tiered_limits=1 on the kernel command line and demotion
enabled.

Set up a cgroup, apply per-tier limits, and run a memory-intensive
workload:

  $ mkdir /sys/fs/cgroup/mycgroup
  $ cd /sys/fs/cgroup/mycgroup
  $ echo "tier4.high=200000000" > memory.tier
  $ echo "tier4.max=300000000" > memory.tier
  $ echo 1 > /sys/kernel/mm/numa/demotion_enabled
  $ cgexec -g memory:/mycgroup ~/stream --ntimes 5 --malloc &
  $ cat memory.tier
  tier4.current=296488960
  tier4.high=199999488
  tier4.max=299999232
  tier22.current=1625464832
  tier22.high=max
  tier22.max=max

DRAM (tier4) usage stays under tier4.max (hard limit, no OOM) but exceeds
tier4.high (soft limit, suggesting that async reclaim is in progress);
CXL (tier22) absorbs the overflow via demotion.

Also verified:
  - tierN.current tracks per-tier usage (anon + file).
  - Cross-tier migration (demotion) correctly re-accounts.
  - memory.high / memory.max auto-derives tierN.high / tierN.max.
  - Manual override (writing a number to memory.tier) pins the limit.
  - Tier max enforcement triggers reclaim scoped to the tier's nodes.
  - Feature fully off (no mount option): no counters, no charge/uncharge,
    memory.tier exists but reads empty.

Open questions
--------------

- Should kmem/slab tier accounting be included in this series or deferred
  to a follow-up?
- Should per-tier memory.min and memory.low protection be part of this
  series or left for later?

[1] https://lore.kernel.org/all/20260423203445.2914963-1-joshua.hahnjy@gmail.com/

Signed-off-by: Qiqi Liu <liuqiqi@kylinos.cn>

Qiqi Liu (8):
  mm/memory-tiers: add node_to_tier_id and tier_id_to_nodemask
  mm/vmscan: add try_to_free_mem_cgroup_pages_nodemask
  mm/memcontrol: add per-tier page counter infrastructure and lifecycle
  mm/memcontrol: add per-tier charge and uncharge
  mm/memcontrol: add per-cpu stock for tier charge/uncharge
  mm/memcontrol: add memory.tier control file
  mm/memcontrol: auto-derive tier high/max from memory.high/max
  cgroup: add memory_tiered_limits cgroup mount option

 include/linux/cgroup-defs.h  |   5 +
 include/linux/memcontrol.h   |  26 ++
 include/linux/memory-tiers.h |  12 +
 include/linux/swap.h         |   6 +
 kernel/cgroup/cgroup.c       |  21 +
 mm/memcontrol.c              | 786 ++++++++++++++++++++++++++++++++++-
 mm/memory-tiers.c            |  58 +++
 mm/vmscan.c                  |  26 +-
 8 files changed, 936 insertions(+), 4 deletions(-)

-- 
2.43.0
Re: [RFC PATCH 0/8] mm/memcontrol: introduce per-tier memory accounting and control
Posted by Michal Hocko 1 month, 1 week ago
Are you aware of a similar work in this area by Joshua
https://lore.kernel.org/all/20260807202059.2620949-1-joshua.hahnjy@gmail.com/T/#u?
We owe Joshua review feedback for quite some time but if I have to be
honest the most impeding factor on my end is that I am not really
convinced tier aware controlling is the right direction. I have
expressed some concerns on one of the earlier proposal by Joshua
https://lore.kernel.org/all/aZ2LC0KPF0xsAwAL@tiehlicka/T/#u

In any way it would be great to talk and compare your approaches see
where they align and the discuss further.

On Tue 18-08-26 10:31:13, liuqiqi@kylinos.cn wrote:
> From: Qiqi Liu <liuqiqi@kylinos.cn>
> 
> This RFC introduces per-tier memory cgroup accounting. Each cgroup
> tracks its memory usage per memory tier (e.g. DRAM, CXL), exposed
> through a new memory.tier control file that reports per-tier usage
> and accepts independent high (soft) and max (hard) limits per tier.
> By default these limits are auto-derived from memory.high / memory.max
> based on per-tier capacity ratios, and can be manually overridden.
> 
> The implementation integrates with the existing memory tiering and
> demotion infrastructure. Per-tier usage (anonymous and file) is tracked
> via dedicated page counters, and cross-tier migrations (e.g. demotion
> from DRAM to CXL) correctly re-account charges. When a tier hits its
> high limit, async reclaim is triggered within that tier's NUMA nodes;
> exceeding max enforces reclaim scoped to the tier's own nodes, or OOM.
> 
> The feature is fully opt-in. When disabled, no extra counters or
> charge/uncharge paths are created, memory.tier reads empty, and there
> is no measurable overhead.
> 
> Why per-tier limits?
> -------------------
> 
> On tiered memory systems, memory.max constrains total usage but cannot
> express "keep fast-tier usage under X". Without per-tier limits, a
> workload can monopolise DRAM, pushing other cgroups onto slower tiers.
> This series gives each cgroup independent high (soft) and max (hard)
> limits per tier, exposed and set through a new memory.tier file.
> 
> By default those limits auto-derive from memory.high / memory.max by
> capacity ratio; writing memory.tier pins a tier.
> 
> This series takes a different approach from Joshua Hahn's toptier RFC [1],
> tracking a separate page_counter per (memcg, tier) for N-tier support and
> exposing writable per-tier limits under a cgroup mount option.
> 
> Patch structure
> ---------------
> 
>   1/8  mm/memory-tiers: add node_to_tier_id and tier_id_to_nodemask
>   2/8  mm/vmscan: add try_to_free_mem_cgroup_pages_nodemask
>   3/8  mm/memcontrol: add per-tier page counter infrastructure and lifecycle
>   4/8  mm/memcontrol: add per-tier charge and uncharge
>   5/8  mm/memcontrol: add per-cpu stock for tier charge/uncharge
>   6/8  mm/memcontrol: add memory.tier control file
>   7/8  mm/memcontrol: auto-derive tier high/max from memory.high/max
>   8/8  cgroup: add memory_tiered_limits cgroup mount option
> 
> Patches 1-2 are infrastructure (helpers in memory-tiers and vmscan).
> Patches 3-5 add the core accounting: counter lifecycle (3),
> per-page charge/uncharge (4), and stock batching (5).
> Patch 6 adds the userspace file. Patch 7 adds auto-derivation. Patch 8
> gates everything behind a mount option + kernel cmdline, so the feature
> adds no measurable overhead when not opted in.
> 
> Usage
> -----
> 
> Boot with:
> 
>   cgroup_memory_tiered_limits=1
> 
> Or remount at runtime (affects newly created cgroups only):
> 
>   mount -o remount,memory_tiered_limits /sys/fs/cgroup
> 
> Per-tier limits and usage can then be read from and written to
> memory.tier.
> 
> Scope and limitations
> ---------------------
> 
> - Only LRU folios (anonymous and file pages) are tier-accounted. Kernel
>   memory and socket buffers are not yet accounted per tier; support for
>   these is planned as follow-up work.
> - Per-tier memory.min and memory.low protections are not implemented.
>   These can be added later by extending the per-tier interface to
>   expose and enforce min/low protection.
> - The command-line parameter mirrors cgroup_favordynmods; automatic
>   enablement via the cgroup mount path is left to userspace.
> 
> Testing
> -------
> 
> Tested on QEMU with fake NUMA (DRAM tier 4 + CXL tier 22), with
> cgroup_memory_tiered_limits=1 on the kernel command line and demotion
> enabled.
> 
> Set up a cgroup, apply per-tier limits, and run a memory-intensive
> workload:
> 
>   $ mkdir /sys/fs/cgroup/mycgroup
>   $ cd /sys/fs/cgroup/mycgroup
>   $ echo "tier4.high=200000000" > memory.tier
>   $ echo "tier4.max=300000000" > memory.tier
>   $ echo 1 > /sys/kernel/mm/numa/demotion_enabled
>   $ cgexec -g memory:/mycgroup ~/stream --ntimes 5 --malloc &
>   $ cat memory.tier
>   tier4.current=296488960
>   tier4.high=199999488
>   tier4.max=299999232
>   tier22.current=1625464832
>   tier22.high=max
>   tier22.max=max
> 
> DRAM (tier4) usage stays under tier4.max (hard limit, no OOM) but exceeds
> tier4.high (soft limit, suggesting that async reclaim is in progress);
> CXL (tier22) absorbs the overflow via demotion.
> 
> Also verified:
>   - tierN.current tracks per-tier usage (anon + file).
>   - Cross-tier migration (demotion) correctly re-accounts.
>   - memory.high / memory.max auto-derives tierN.high / tierN.max.
>   - Manual override (writing a number to memory.tier) pins the limit.
>   - Tier max enforcement triggers reclaim scoped to the tier's nodes.
>   - Feature fully off (no mount option): no counters, no charge/uncharge,
>     memory.tier exists but reads empty.
> 
> Open questions
> --------------
> 
> - Should kmem/slab tier accounting be included in this series or deferred
>   to a follow-up?
> - Should per-tier memory.min and memory.low protection be part of this
>   series or left for later?
> 
> [1] https://lore.kernel.org/all/20260423203445.2914963-1-joshua.hahnjy@gmail.com/
> 
> Signed-off-by: Qiqi Liu <liuqiqi@kylinos.cn>
> 
> Qiqi Liu (8):
>   mm/memory-tiers: add node_to_tier_id and tier_id_to_nodemask
>   mm/vmscan: add try_to_free_mem_cgroup_pages_nodemask
>   mm/memcontrol: add per-tier page counter infrastructure and lifecycle
>   mm/memcontrol: add per-tier charge and uncharge
>   mm/memcontrol: add per-cpu stock for tier charge/uncharge
>   mm/memcontrol: add memory.tier control file
>   mm/memcontrol: auto-derive tier high/max from memory.high/max
>   cgroup: add memory_tiered_limits cgroup mount option
> 
>  include/linux/cgroup-defs.h  |   5 +
>  include/linux/memcontrol.h   |  26 ++
>  include/linux/memory-tiers.h |  12 +
>  include/linux/swap.h         |   6 +
>  kernel/cgroup/cgroup.c       |  21 +
>  mm/memcontrol.c              | 786 ++++++++++++++++++++++++++++++++++-
>  mm/memory-tiers.c            |  58 +++
>  mm/vmscan.c                  |  26 +-
>  8 files changed, 936 insertions(+), 4 deletions(-)
> 
> -- 
> 2.43.0

-- 
Michal Hocko
SUSE Labs
Re: [RFC PATCH 0/8] mm/memcontrol: introduce per-tier memory accounting and control
Posted by Joshua Hahn 1 month, 1 week ago
On Tue, 18 Aug 2026 10:12:19 +0200 Michal Hocko <mhocko@suse.com> wrote:

> Are you aware of a similar work in this area by Joshua
> https://lore.kernel.org/all/20260807202059.2620949-1-joshua.hahnjy@gmail.com/T/#u?
> We owe Joshua review feedback for quite some time but if I have to be
> honest the most impeding factor on my end is that I am not really
> convinced tier aware controlling is the right direction. I have
> expressed some concerns on one of the earlier proposal by Joshua
> https://lore.kernel.org/all/aZ2LC0KPF0xsAwAL@tiehlicka/T/#u
> 
> In any way it would be great to talk and compare your approaches see
> where they align and the discuss further.

Hi Michal,

I hope you have been doing well! Maybe this is a good opportunity to
squash concerns about the series : -) I've sent a new version which
addresses a lot of the concerns that David Hildenbrand had previously
brought up [1]. This implementation is as transparent as can be to the
user. I hope that this makes sense to you! 

IIRC you had some concerns about getting the interface right the first
time we introduce this feature. Hopefully with no user-visible features
(and only a mount option) your concerns there are resolved and hopefully
we can work to figure out what remaining concerns there are.

Thank you as always for your time. I hope you have a great day!
Joshua

[1] https://lore.kernel.org/all/20260807202059.2620949-1-joshua.hahnjy@gmail.com/
Re: [RFC PATCH 0/8] mm/memcontrol: introduce per-tier memory accounting and control
Posted by Shakeel Butt 1 month, 1 week ago
On Tue, Aug 18, 2026 at 10:31:13AM +0800, liuqiqi@kylinos.cn wrote:
> From: Qiqi Liu <liuqiqi@kylinos.cn>
> 
[...]
> 
> This series takes a different approach from Joshua Hahn's toptier RFC [1],
> tracking a separate page_counter per (memcg, tier) for N-tier support and
> exposing writable per-tier limits under a cgroup mount option.

Please don't duplicate the effort. If you have concerns regarding Joshua's
series, comment on his series and work with him to push the work through.
Re: [RFC PATCH 0/8] mm/memcontrol: introduce per-tier memory accounting and control
Posted by Joshua Hahn 1 month, 1 week ago
On Tue, 18 Aug 2026 10:31:13 +0800 liuqiqi@kylinos.cn wrote:

> From: Qiqi Liu <liuqiqi@kylinos.cn>
> 
> This RFC introduces per-tier memory cgroup accounting. Each cgroup
> tracks its memory usage per memory tier (e.g. DRAM, CXL), exposed
> through a new memory.tier control file that reports per-tier usage
> and accepts independent high (soft) and max (hard) limits per tier.
> By default these limits are auto-derived from memory.high / memory.max
> based on per-tier capacity ratios, and can be manually overridden.
> 
> The implementation integrates with the existing memory tiering and
> demotion infrastructure. Per-tier usage (anonymous and file) is tracked
> via dedicated page counters, and cross-tier migrations (e.g. demotion
> from DRAM to CXL) correctly re-account charges. When a tier hits its
> high limit, async reclaim is triggered within that tier's NUMA nodes;
> exceeding max enforces reclaim scoped to the tier's own nodes, or OOM.
> 
> The feature is fully opt-in. When disabled, no extra counters or
> charge/uncharge paths are created, memory.tier reads empty, and there
> is no measurable overhead.
> 
> Why per-tier limits?
> -------------------
> 
> On tiered memory systems, memory.max constrains total usage but cannot
> express "keep fast-tier usage under X". Without per-tier limits, a
> workload can monopolise DRAM, pushing other cgroups onto slower tiers.
> This series gives each cgroup independent high (soft) and max (hard)
> limits per tier, exposed and set through a new memory.tier file.
> 
> By default those limits auto-derive from memory.high / memory.max by
> capacity ratio; writing memory.tier pins a tier.
> 
> This series takes a different approach from Joshua Hahn's toptier RFC [1],
> tracking a separate page_counter per (memcg, tier) for N-tier support and
> exposing writable per-tier limits under a cgroup mount option.

Hi Qiqi,

Thanks for sending the series. I'm glad that there is additional interest
in making tiered limits in the system. In this series, I do see a lot
of duplicate work with my work here [2]. It looks like you cited [1]
which is an older version of the series that I sent out. Notably the
new version has N-tier support and a cgroup mount option.

Also consider my series in [3] where I am moving stock to the
page_counter level. It's been a while, but I'm hoping to send out a new
version of that series next week.

With all of that considered, I wanted to know what differences your RFC
here has with my series. From where I stand, the only difference I can
see was making the limits exposed / writable, which was an explicit
design decision that I made to de-clutter the memcg tuning space and
try to make the mechanism as transparent to the user as possible.

The other parts of the series (per-tier reclaim, per-tier-memcg
page_counter accounting, auto-scaling high/max from memory.high/max)
seems to be the same as my series.

Rather than duplicate our effort I think it would be best to foucs all
of our effort and the maintainers' effort into discussing the design
decisions for the series.

Joshua

> Patch structure
> ---------------
> 
>   1/8  mm/memory-tiers: add node_to_tier_id and tier_id_to_nodemask
>   2/8  mm/vmscan: add try_to_free_mem_cgroup_pages_nodemask
>   3/8  mm/memcontrol: add per-tier page counter infrastructure and lifecycle
>   4/8  mm/memcontrol: add per-tier charge and uncharge
>   5/8  mm/memcontrol: add per-cpu stock for tier charge/uncharge
>   6/8  mm/memcontrol: add memory.tier control file
>   7/8  mm/memcontrol: auto-derive tier high/max from memory.high/max
>   8/8  cgroup: add memory_tiered_limits cgroup mount option
> 
> Patches 1-2 are infrastructure (helpers in memory-tiers and vmscan).
> Patches 3-5 add the core accounting: counter lifecycle (3),
> per-page charge/uncharge (4), and stock batching (5).
> Patch 6 adds the userspace file. Patch 7 adds auto-derivation. Patch 8
> gates everything behind a mount option + kernel cmdline, so the feature
> adds no measurable overhead when not opted in.
> 
> Usage
> -----
> 
> Boot with:
> 
>   cgroup_memory_tiered_limits=1
> 
> Or remount at runtime (affects newly created cgroups only):
> 
>   mount -o remount,memory_tiered_limits /sys/fs/cgroup
> 
> Per-tier limits and usage can then be read from and written to
> memory.tier.
> 
> Scope and limitations
> ---------------------
> 
> - Only LRU folios (anonymous and file pages) are tier-accounted. Kernel
>   memory and socket buffers are not yet accounted per tier; support for
>   these is planned as follow-up work.
> - Per-tier memory.min and memory.low protections are not implemented.
>   These can be added later by extending the per-tier interface to
>   expose and enforce min/low protection.
> - The command-line parameter mirrors cgroup_favordynmods; automatic
>   enablement via the cgroup mount path is left to userspace.
> 
> Testing
> -------
> 
> Tested on QEMU with fake NUMA (DRAM tier 4 + CXL tier 22), with
> cgroup_memory_tiered_limits=1 on the kernel command line and demotion
> enabled.
> 
> Set up a cgroup, apply per-tier limits, and run a memory-intensive
> workload:
> 
>   $ mkdir /sys/fs/cgroup/mycgroup
>   $ cd /sys/fs/cgroup/mycgroup
>   $ echo "tier4.high=200000000" > memory.tier
>   $ echo "tier4.max=300000000" > memory.tier
>   $ echo 1 > /sys/kernel/mm/numa/demotion_enabled
>   $ cgexec -g memory:/mycgroup ~/stream --ntimes 5 --malloc &
>   $ cat memory.tier
>   tier4.current=296488960
>   tier4.high=199999488
>   tier4.max=299999232
>   tier22.current=1625464832
>   tier22.high=max
>   tier22.max=max
> 
> DRAM (tier4) usage stays under tier4.max (hard limit, no OOM) but exceeds
> tier4.high (soft limit, suggesting that async reclaim is in progress);
> CXL (tier22) absorbs the overflow via demotion.
> 
> Also verified:
>   - tierN.current tracks per-tier usage (anon + file).
>   - Cross-tier migration (demotion) correctly re-accounts.
>   - memory.high / memory.max auto-derives tierN.high / tierN.max.
>   - Manual override (writing a number to memory.tier) pins the limit.
>   - Tier max enforcement triggers reclaim scoped to the tier's nodes.
>   - Feature fully off (no mount option): no counters, no charge/uncharge,
>     memory.tier exists but reads empty.
> 
> Open questions
> --------------
> 
> - Should kmem/slab tier accounting be included in this series or deferred
>   to a follow-up?
> - Should per-tier memory.min and memory.low protection be part of this
>   series or left for later?
> 
> [1] https://lore.kernel.org/all/20260423203445.2914963-1-joshua.hahnjy@gmail.com/
> 
> Signed-off-by: Qiqi Liu <liuqiqi@kylinos.cn>
> 
> Qiqi Liu (8):
>   mm/memory-tiers: add node_to_tier_id and tier_id_to_nodemask
>   mm/vmscan: add try_to_free_mem_cgroup_pages_nodemask
>   mm/memcontrol: add per-tier page counter infrastructure and lifecycle
>   mm/memcontrol: add per-tier charge and uncharge
>   mm/memcontrol: add per-cpu stock for tier charge/uncharge
>   mm/memcontrol: add memory.tier control file
>   mm/memcontrol: auto-derive tier high/max from memory.high/max
>   cgroup: add memory_tiered_limits cgroup mount option
> 
>  include/linux/cgroup-defs.h  |   5 +
>  include/linux/memcontrol.h   |  26 ++
>  include/linux/memory-tiers.h |  12 +
>  include/linux/swap.h         |   6 +
>  kernel/cgroup/cgroup.c       |  21 +
>  mm/memcontrol.c              | 786 ++++++++++++++++++++++++++++++++++-
>  mm/memory-tiers.c            |  58 +++
>  mm/vmscan.c                  |  26 +-
>  8 files changed, 936 insertions(+), 4 deletions(-)
> 
> -- 
> 2.43.0

[2] https://lore.kernel.org/all/20260807202059.2620949-1-joshua.hahnjy@gmail.com/
[3] https://lore.kernel.org/all/20260623180124.868655-1-joshua.hahnjy@gmail.com/
Re: [RFC PATCH 0/8] mm/memcontrol: introduce per-tier memory accounting and control
Posted by liuqiqi@kylinos.cn 1 month, 1 week ago
From: Qiqi Liu <liuqiqi@kylinos.cn>

Hi all,

Thank you all for your replies. I am not very familiar with the
community's workflow and should have reviewed the mailing list
archives and existing implementations more carefully. I sincerely
apologize for any inconvenience this may have caused.

My work is based on
https://lore.kernel.org/all/20260528134212.240492-1-liuqiqi@kylinos.cn/
aiming to develop memory tiering limits for cgroups. During
development, I referenced Joshua's v2, but failed to notice that
v3 had already been posted when I submitted my series.

I have studied Joshua's v3, and our core mechanisms are largely
consistent. However, there are two differences:

1. Read/write per-tier interface (memory.tier): each cgroup
   tracks its memory usage by tier (e.g., DRAM, CXL), exposed
   via a new memory.tier control file. This file reports
   per-tier usage and accepts per-tier high (soft limit) and
   max (hard limit) settings. By default, these limits are
   automatically derived from memory.high/max based on each
   tier's capacity ratio, but manual overrides are supported,
   allowing administrators to constrain specific tiers on a
   per-cgroup basis.

   Its advantages are:
   - It can express allocations that fixed capacity ratios
     cannot.
   - Latency-sensitive tenants can be given a larger share of
     the fast tier.
   - High-capacity tenants can have their soft limits removed
     for the slow tier.

   Whether or not to constrain a specific tier should be a
   decision made by the administrator on a per-cgroup basis.
   This serves as an answer to Michal's earlier question in
   the thread regarding "do you intend to limit memory
   consumption on particular tier even without an external
   pressure?"

2. Per-tier stock: I have implemented and tested per-tier
   stock -- it batches atomic operations on tier counters from
   per-page to every 32 pages. Joshua's patch (moving the stock
   down to the page_counter level) takes a more generic
   approach, and I believe we should continue with his
   direction here.

Memory hardware is becoming increasingly diverse (DRAM, CXL,
and further remote tiers), and the performance gaps between
tiers are significant. Different workloads have different
requirements -- some demand low latency, while others need
capacity. When the fast tier cannot accommodate the working
sets of all workloads, it should be the administrator's
scheduling decision to determine fast-tier allocations. I
believe tier-aware control is a direction worth pursuing, and
the differences above are where we can contribute.

As Shakeel suggested, and given that Joshua's v3 already
contains the core mechanism, I am dropping my current
standalone patchset. I would like to ask if Joshua would be
willing to collaborate with me on this, treating memory.tier
as an extension to the patch series and proposing it as
follow-up patches based on v3.

PS: I have already subscribed to the mailing list to avoid
such misunderstandings in the future.

Best regards,
Qiqi Liu
Re: [RFC PATCH 0/8] mm/memcontrol: introduce per-tier memory accounting and control
Posted by Gregory Price 1 month, 1 week ago
On Wed, Aug 19, 2026 at 09:22:11PM +0800, liuqiqi@kylinos.cn wrote:
> From: Qiqi Liu <liuqiqi@kylinos.cn>
> 
> Hi all,

Hi! Thank you for following up.  A few things.

> 
> Thank you all for your replies. I am not very familiar with the
> community's workflow and should have reviewed the mailing list
> archives and existing implementations more carefully. I sincerely
> apologize for any inconvenience this may have caused.
> 

Less of an inconvience, we want to save you time as much as we want to
save the larger community's time.  Having multiple interested parties
vet common work - rather than propose differing solutions - does that.

Welcome to the discussion, glad to have more eyes on the problem!

Hopefully I can provide some context on the history here, since I've
been working with Joshua for a while on this in the background.

> My work is based on
> https://lore.kernel.org/all/20260528134212.240492-1-liuqiqi@kylinos.cn/

On this patch, It's not clear why an RCU-protected pointer is
unsuitable.  RCU is hot-path safe, it's just not stable nor
sleep-safe, which should be sufficient for any operation which
may be looking up this particular mapping.

These values are not expected to be aggressively written to, so RCU
essentially becomes a NOP on the reader side - it's extremely cheap.

More ideologically - adding a cached value of an RCU protected value
is somewhat anti-thetical to the entire purpose of using RCU in the
first place - it creates more footguns than it solves.

That aside, getting to the tier-aware memcg limits...

> aiming to develop memory tiering limits for cgroups. During
> development, I referenced Joshua's v2, but failed to notice that
> v3 had already been posted when I submitted my series.
> 
> I have studied Joshua's v3, and our core mechanisms are largely
> consistent. However, there are two differences:
> 
> 1. Read/write per-tier interface (memory.tier): each cgroup
>    tracks its memory usage by tier (e.g., DRAM, CXL), exposed
>    via a new memory.tier control file. This file reports
>    per-tier usage and accepts per-tier high (soft limit) and
>    max (hard limit) settings. By default, these limits are
>    automatically derived from memory.high/max based on each
>    tier's capacity ratio, but manual overrides are supported,
>    allowing administrators to constrain specific tiers on a
>    per-cgroup basis.
> 

There's two levels of operation we need to think about here:

1) What the kernel does by default without tuning
2) What the kernel enables admins to tune

If we don't have a cogent story around how #1 should occur for
this feature - then every knob you expose for #2 is just creating
a mess of tunables no one can possibly understand (let alone maintain).

That's why Joshua's series has no tunable knobs - any such knob is
simply unwarranted at this point.  (This decision was born from both
on-list and in-person feedback).

>    Its advantages are:
>    - It can express allocations that fixed capacity ratios
>      cannot.

Which should come from a use case born out of demonstrating fixed ratios
are actually insufficient and cannot be made to self-tune.

But we don't even have those yet.

>    - Latency-sensitive tenants can be given a larger share of
>      the fast tier.
>    - High-capacity tenants can have their soft limits removed
>      for the slow tier.
> 

These are the same issue as the first bullet, just differently shaped.

>    Whether or not to constrain a specific tier should be a
>    decision made by the administrator on a per-cgroup basis.

This is an opinion, not a fact, and should be based on data that
demonstrates the kernel is incapable of making the (or a) "right"
decision in a sufficiently common scenario.

> When the fast tier cannot accommodate the working sets of all
> workloads, it should be the administrator's scheduling decision to
> determine fast-tier allocations.

There's basically 3 use-cases that have been collected that I've seen
which tier-aware memcg looks to address:


1) Self-policed fairness

   Stiff per-tier limits that cgroups impose on themselves.
   i.e. proactively applying tier(memory.high/max) to ensure no
   container's tier(memory.min) is ever violated.

   This creates reduced variance in exchange for lower throughput.

   This is paradigm essentially does not exist today except via
   cpuset.mems (e.g. putting everything for a task on CXL). This
   is intended for things that want stronger QoS controls.

2) Opportunistic fairness

   While there is sufficient space on a higher tier, cgroups should
   be allowed to "over-use" the upper tier opportunistically to maximize
   thoughput - but when someone's tier(memory.min) is violated because
   another container is over-using, we nudge everyone toward fairness.

   This creates higher throughput in exchange for increase variance.

   This is milder modification to the existing global opportunistic
   behavior.  Think of it like trying to apply a soft memory QoS.

   It's unclear whether this actually has value, but can probably
   be accomplished via existing min/high/max, rather than needing new
   sysfs toggles.

3) Per-cgroup adjustable tier limits.

   A scheduler knows something about the workloads it wants to
   have custom tier limits per-workload.

   This should be seen as an evolution born out of finding where
   1 and 2 are insufficient.  It's putting the cart before the horse
   to go directly to this point.

   Very few of us are convinced such complexity is actually warranted,
   especially because the simpler (and less ABI-permanent) #1 and #2
   haven't even been fully explored.

> As Shakeel suggested, and given that Joshua's v3 already
> contains the core mechanism, I am dropping my current
> standalone patchset. I would like to ask if Joshua would be
> willing to collaborate with me on this, treating memory.tier
> as an extension to the patch series and proposing it as
> follow-up patches based on v3.
> 

Joshua can speak for himself, but more eyes and testing and data is
always welcome.

~Gregory
[syzbot ci] Re: mm/memcontrol: introduce per-tier memory accounting and control
Posted by syzbot ci 1 month, 1 week ago
syzbot ci has tested the following series

[v1] mm/memcontrol: introduce per-tier memory accounting and control
https://lore.kernel.org/all/20260818023121.100613-1-liuqiqi@kylinos.cn
* [RFC PATCH 1/8] mm/memory-tiers: add node_to_tier_id and tier_id_to_nodemask
* [RFC PATCH 2/8] mm/vmscan: add try_to_free_mem_cgroup_pages_nodemask
* [RFC PATCH 3/8] mm/memcontrol: add per-tier page counter infrastructure and lifecycle
* [RFC PATCH 4/8] mm/memcontrol: add per-tier charge and uncharge
* [RFC PATCH 5/8] mm/memcontrol: add per-cpu stock for tier charge/uncharge
* [RFC PATCH 6/8] mm/memcontrol: add memory.tier control file
* [RFC PATCH 7/8] mm/memcontrol: auto-derive tier high/max from memory.high/max
* [RFC PATCH 8/8] cgroup: add memory_tiered_limits cgroup mount option

and found the following issue:
WARNING in uncharge_folio

Full report is available here:
https://ci.syzbot.org/series/9db711d8-d389-45d1-b245-7e8b316a8432

***

WARNING in uncharge_folio

tree:      torvalds
URL:       https://kernel.googlesource.com/pub/scm/linux/kernel/git/torvalds/linux
base:      dcacab904fe78d60840ba947a104993ee9ded887
arch:      amd64
compiler:  Debian clang version 22.1.8 (++20260613092233+e80beda6e255-1~exp1~20260613092250.77), Debian LLD 22.1.8
config:    https://ci.syzbot.org/builds/bf213e6f-1b40-41fa-a938-385807688eec/config

------------[ cut here ]------------
debug_locks && !(rcu_read_lock_held() || lock_is_held(&(&cgroup_mutex)->dep_map))
WARNING: ./include/linux/memcontrol.h:391 at uncharge_folio+0x459/0x5b0, CPU#1: udevd/5681
Modules linked in:
CPU: 1 UID: 0 PID: 5681 Comm: udevd Not tainted syzkaller #0 PREEMPT(full) 
Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.2-debian-1.16.2-1 04/01/2014
RIP: 0010:uncharge_folio+0x459/0x5b0
Code: c4 18 5b 41 5c 41 5d 41 5e 41 5f 5d c3 cc cc cc cc cc 48 8b 3c 24 48 83 c4 18 5b 41 5c 41 5d 41 5e 41 5f 5d e9 88 79 00 00 90 <0f> 0b 90 e9 c9 fe ff ff 89 e9 80 e1 07 80 c1 03 38 c1 0f 8c 49 fc
RSP: 0018:ffffc90004ff7210 EFLAGS: 00010246
RAX: 0000000000000000 RBX: ffffc90004ff72a0 RCX: 0000000080000001
RDX: 0000000000000000 RSI: ffffffff8e4a6d0d RDI: ffffffff8c4bb280
RBP: ffffc90004ff72a8 R08: ffffea0006e7a7c7 R09: 1ffffd4000dcf4f8
R10: dffffc0000000000 R11: fffff94000dcf4f9 R12: ffffea0006e7a7f8
R13: 1ffffd4000dcf4f8 R14: 0000000000000001 R15: ffffea0006e7a7c0
FS:  00007f1ec7ce0c80(0000) GS:ffff8882a8f6a000(0000) knlGS:0000000000000000
CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
CR2: 00007f0cdf670000 CR3: 000000017358a000 CR4: 00000000000006f0
Call Trace:
 <TASK>
 __mem_cgroup_uncharge_folios+0xfd/0x1d0
 folio_batch_move_lru+0x86f/0xa60
 lru_add_drain_cpu+0xbc/0x750
 lru_add_drain+0x121/0x3e0
 __folio_batch_release+0x48/0x90
 shmem_undo_range+0x4e6/0x15d0
 shmem_evict_inode+0x280/0xa80
 evict+0x624/0xb50
 dentry_kill+0x1b9/0x880
 finish_dput+0x1a/0x260
 filename_renameat2+0x61b/0x9a0
 __se_sys_rename+0x55/0x2c0
 do_syscall_64+0x174/0x580
 entry_SYSCALL_64_after_hwframe+0x77/0x7f
RIP: 0033:0x7f1ec789a93b
Code: 48 8b 15 f0 64 15 00 83 c8 ff 64 83 3a 15 75 0e 48 8b 7c 24 08 e8 d5 d4 07 00 f7 d8 19 c0 48 83 c4 18 c3 b8 52 00 00 00 0f 05 <48> 3d 00 f0 ff ff 76 10 48 8b 15 be 64 15 00 f7 d8 64 89 02 48 83
RSP: 002b:00007ffcb336cd08 EFLAGS: 00000202 ORIG_RAX: 0000000000000052
RAX: ffffffffffffffda RBX: 0000000000000000 RCX: 00007f1ec789a93b
RDX: 0000558c622ea032 RSI: 00007ffcb336cd28 RDI: 00007ffcb336d128
RBP: 000055893abc0910 R08: 0000000000000006 R09: f80f8ce6f311cb22
R10: 00000000000001b6 R11: 0000000000000202 R12: 000055893aba2730
R13: 00007ffcb336cd28 R14: 00007ffcb336d128 R15: 0000558926efb160
 </TASK>


***

If these findings have caused you to resend the series or submit a
separate fix, please add the following tag to your commit message:
  Tested-by: syzbot@syzkaller.appspotmail.com

---
This report is generated by a bot. It may contain errors.
syzbot ci engineers can be reached at syzkaller@googlegroups.com.

To test a fix for this bug, please reply with `#syz test`
(on a separate line) and attach the patch to the email.

Notes:
- The patch will be applied on top of the tested series (as an
  incremental fix).
- To test a new version of the whole series, please send it directly
  to syzbot@lists.linux.dev.
- Arguments like custom git repos and branches are not supported.