[PATCH v2 0/4] mm: restore per-memcg reclaim for NONSLAB shrinkers under nokmem

Qinyun Tan posted 4 patches 2 weeks, 4 days ago
There is a newer version of this series
include/linux/memcontrol.h |  8 +++++---
mm/huge_memory.c           |  3 ++-
mm/list_lru.c              | 14 ++++++++------
mm/memcontrol.c            |  6 ------
mm/zswap.c                 |  4 ++--
5 files changed, 17 insertions(+), 18 deletions(-)
[PATCH v2 0/4] mm: restore per-memcg reclaim for NONSLAB shrinkers under nokmem
Posted by Qinyun Tan 2 weeks, 4 days ago
With cgroup.memory=nokmem, the THP deferred split shrinker and the
zswap shrinker are degraded in two ways.

First, both shrinkers are missing the SHRINKER_NONSLAB flag, so
shrinker_memcg_alloc() demotes them to non-memcg-aware shrinkers:
limit-induced reclaim of a cgroup neither splits its partially
unmapped THPs nor writes back its zswapped pages.  v1 of this series
[1] restored the flag to fix that.

However, as Sashiko's review of v1 pointed out [2], the flag alone
is not enough.  __list_lru_init() also collapses every list_lru into
per-node lists under nokmem, so even with the flag restored, the
objects of all cgroups share one list per node: the per-memcg
shrinker bit is only set for whichever memcg happens to repopulate
the empty list, so pressure in other cgroups may not even trigger
the scan, and when it does, the scan walks everyone's objects.

Before commit fafaeceb89a5 ("mm: switch deferred split shrinker to
list_lru"), THP had fully per-memcg deferred split queues embedded
in struct mem_cgroup, working independently of kmem accounting.
nokmem only opts out of kernel slab accounting; THPs and zswapped
pages are user memory and remain charged to their cgroups, so
per-memcg reclaim is still what these shrinkers want.

This series keeps list_lrus backed by SHRINKER_NONSLAB shrinkers
memcg aware under nokmem:

Patch 1 drops the kmemcg_id copy, which is only assigned when kmem
accounting is enabled, and derives the list_lru xarray index from
the memcg ID directly, so the index works independently of kmem
accounting.  It also drops the nokmem early return from
memcg_offline_kmem() so these lrus are reparented on offline.

Patch 2 keeps a list_lru memcg aware under nokmem when its backing
shrinker is registered SHRINKER_NONSLAB.

Patches 3 and 4 restore/add SHRINKER_NONSLAB on the THP deferred
split shrinker and the zswap shrinker, unchanged from v1.

The savings of nokmem are preserved: slab-backed lrus (e.g. the
superblock dentry/inode lrus) still fall back to per-node lists, and
the per-memcg lists are only allocated when a memcg actually holds
such objects.

[1] https://lore.kernel.org/lkml/20260904033503.4067283-1-qinyuntan@linux.alibaba.com/
[2] https://sashiko.dev/#/patchset/20260904033503.4067283-1-qinyuntan@linux.alibaba.com

Changes in v2:
 - Keep the list_lrus behind the two shrinkers per-memcg under
   nokmem, so the shrinker bits are set for the right memcgs and
   scans stay scoped to the target cgroup's objects (patches 1-2,
   new; addresses Sashiko's review of v1).
 - Patches 3-4 unchanged from v1; collected the review tags.

Qinyun Tan (4):
  mm: memcontrol: drop kmemcg_id and use the memcg ID for list_lru
    indexing
  mm: list_lru: keep per-memcg lists with nokmem for NONSLAB-backed lrus
  mm: thp: restore SHRINKER_NONSLAB on the deferred split shrinker
  mm: zswap: mark the zswap shrinker SHRINKER_NONSLAB

 include/linux/memcontrol.h |  8 +++++---
 mm/huge_memory.c           |  3 ++-
 mm/list_lru.c              | 14 ++++++++------
 mm/memcontrol.c            |  6 ------
 mm/zswap.c                 |  4 ++--
 5 files changed, 17 insertions(+), 18 deletions(-)

-- 
2.43.7