fs/super.c | 18 +- include/linux/memcontrol.h | 11 +- include/linux/shmem_fs.h | 12 +- mm/page_owner.c | 2 +- mm/shmem.c | 382 ++++++++++++++++++++++++++++--------- mm/zswap.c | 17 +- 6 files changed, 315 insertions(+), 127 deletions(-)
From: Qi Zheng <zhengqi.arch@bytedance.com>
Changes in v4:
- add [PATCH v4 2/4] to make obj_cgroup_memcg() handle NULL objcg
- store obj_cgroup instead of mem_cgroup in shmem_inode_info to avoid
pinning a dying memcg through a long-lived CSS reference
(pointed by sashiko)
- fix is_shmem_unused_huge_match() to always check the NUMA node for
shrinker reclaim, not only for non-root memcg reclaim
(pointed by sashiko)
- collect Reviewed-by
- rebase onto the next-20260814
Note: [PATCH v4 1/4] should ideally be folded into commit 0ef8faff490be
("fs: push nr_cached_objects memcg gating into individual filesystems") in
linux-next.
Changes in v3:
- add a fix patch to fix missed removal of super_fs_objects_eligible()
- move the original shrinklist addition logic after all checks are completed,
and split it into a separate patch. (suggested by Baolin)
- simplify the shmem_unused_huge_requeue() (suggested by Baolin)
- keep the move_back label in shmem_unused_huge_shrink() (suggested by Baolin)
- rebase onto the next-20260731
Changes in v2:
- temporarily add the dependent patch from Usama to the series for review
convenience
- remove shrinklist_scan and shrinklist_isolated from struct shmem_inode_info,
and re-implement the logic by resuing the same info->shrinklist
(suggested by Baolin)
- add more comments (suggested by Andrew)
- fix missing initialization of info->shrinklist_memcg (pointed by sashiko)
- rebase onto the next-20260717
Hi all,
The shmem unused huge shrinker maintains a per-superblock list of inodes
whose tail huge folio extends beyond i_size. Because this list is not
memcg aware, reclaim triggered by memcg A can scan inodes across the
entire superblock and split huge folios charged to unrelated memcg B,
causing unexpected impact on it.
In the worst case, memcg A has no reclaimable shmem at all, making the
reclaim entirely useless and incurring unnecessary latency. We observed
this in production, where page lock contention during split caused
multi-hundred-millisecond stalls:
tid 11340 comm scanner locked a page for 182264 us! kstack:
unlock_page+1
split_huge_page_to_list+3135
shmem_unused_huge_shrink+767
super_cache_scan+329
do_shrink_slab+291
shrink_slab+533
shrink_node+400
do_try_to_free_pages+206
try_to_free_mem_cgroup_pages+262
try_charge_memcg+591
mem_cgroup_charge+136
__handle_mm_fault+2431
handle_mm_fault+194
do_user_addr_fault+462
__do_page_fault+176
do_page_fault+48
page_fault+62
Usama's recent patch [1] prevents the shmem unused shrinker from being
invoked during memcg-level reclaim altogether, but this is overly
conservative: we can do better by reclaiming only the shmem charged to
the reclaiming memcg.
This series converts the shrinker list to a memcg-aware list_lru, so
that non-root memcg reclaim walks only candidates charged to the
reclaiming memcg. Global reclaim, root memcg reclaim and shmem quota
reclaim retain their existing global semantics.
To avoid pinning a dying memcg through a long-lived CSS reference, each
inode stores an obj_cgroup reference instead of a mem_cgroup reference.
The list_lru add/delete paths resolve the current memcg from the objcg
under RCU, staying consistent with list_lru's own memcg migration on
offline.
Thanks,
Qi
[1]. https://lore.kernel.org/all/20260715103516.2410175-1-usama.arif@linux.dev/
Qi Zheng (4):
fs: fix missed removal of super_fs_objects_eligible()
mm: memcontrol: make obj_cgroup_memcg() handle NULL objcg
mm: shmem: move unused huge shrinklist queuing past the truncation
check
mm: shmem: make unused huge shrinker memcg aware
fs/super.c | 18 +-
include/linux/memcontrol.h | 11 +-
include/linux/shmem_fs.h | 12 +-
mm/page_owner.c | 2 +-
mm/shmem.c | 382 ++++++++++++++++++++++++++++---------
mm/zswap.c | 17 +-
6 files changed, 315 insertions(+), 127 deletions(-)
--
2.54.0
On Mon, 17 Aug 2026 17:03:24 +0800 Qi Zheng <qi.zheng@linux.dev> wrote: > Changes in v4: > Changes in v3: > Changes in v2: Thanks for the diligent versioning info. fyi, it is conventional to maintain this below the --- separator. It's not really the most important part of the [0/N]! > > The shmem unused huge shrinker maintains a per-superblock list of inodes > whose tail huge folio extends beyond i_size. Because this list is not > memcg aware, reclaim triggered by memcg A can scan inodes across the > entire superblock and split huge folios charged to unrelated memcg B, > causing unexpected impact on it. > > In the worst case, memcg A has no reclaimable shmem at all, making the > reclaim entirely useless and incurring unnecessary latency. We observed > this in production, where page lock contention during split caused > multi-hundred-millisecond stalls: Ugh. That's the most important part! > tid 11340 comm scanner locked a page for 182264 us! kstack: > unlock_page+1 > split_huge_page_to_list+3135 > shmem_unused_huge_shrink+767 > super_cache_scan+329 > do_shrink_slab+291 > shrink_slab+533 > shrink_node+400 > do_try_to_free_pages+206 > try_to_free_mem_cgroup_pages+262 > try_charge_memcg+591 > mem_cgroup_charge+136 > __handle_mm_fault+2431 > handle_mm_fault+194 > do_user_addr_fault+462 > __do_page_fault+176 > do_page_fault+48 > page_fault+62 > > Usama's recent patch [1] prevents the shmem unused shrinker from being > invoked during memcg-level reclaim altogether, but this is overly > conservative: we can do better by reclaiming only the shmem charged to > the reclaiming memcg. > > This series converts the shrinker list to a memcg-aware list_lru, so > that non-root memcg reclaim walks only candidates charged to the > reclaiming memcg. Global reclaim, root memcg reclaim and shmem quota > reclaim retain their existing global semantics. > > To avoid pinning a dying memcg through a long-lived CSS reference, each > inode stores an obj_cgroup reference instead of a mem_cgroup reference. > The list_lru add/delete paths resolve the current memcg from the objcg > under RCU, staying consistent with list_lru's own memcg migration on > offline. Sashiko said a few things and they look disturbing-if-true: https://sashiko.dev/#/patchset/cover.1786955972.git.zhengqi.arch@bytedance.com (Apologies if this has already been considered - we don't have ways of tracking all this (yet, I hope) apart from personal memory and personal memorys are quite fried at present)
Hi Andrew, On 8/28/26 6:50 AM, Andrew Morton wrote: > On Mon, 17 Aug 2026 17:03:24 +0800 Qi Zheng <qi.zheng@linux.dev> wrote: > >> Changes in v4: >> Changes in v3: >> Changes in v2: > > Thanks for the diligent versioning info. fyi, it is conventional to > maintain this below the --- separator. It's not really the most > important part of the [0/N]! Got it. > >> >> The shmem unused huge shrinker maintains a per-superblock list of inodes >> whose tail huge folio extends beyond i_size. Because this list is not >> memcg aware, reclaim triggered by memcg A can scan inodes across the >> entire superblock and split huge folios charged to unrelated memcg B, >> causing unexpected impact on it. >> >> In the worst case, memcg A has no reclaimable shmem at all, making the >> reclaim entirely useless and incurring unnecessary latency. We observed >> this in production, where page lock contention during split caused >> multi-hundred-millisecond stalls: > > Ugh. That's the most important part! > >> tid 11340 comm scanner locked a page for 182264 us! kstack: >> unlock_page+1 >> split_huge_page_to_list+3135 >> shmem_unused_huge_shrink+767 >> super_cache_scan+329 >> do_shrink_slab+291 >> shrink_slab+533 >> shrink_node+400 >> do_try_to_free_pages+206 >> try_to_free_mem_cgroup_pages+262 >> try_charge_memcg+591 >> mem_cgroup_charge+136 >> __handle_mm_fault+2431 >> handle_mm_fault+194 >> do_user_addr_fault+462 >> __do_page_fault+176 >> do_page_fault+48 >> page_fault+62 >> >> Usama's recent patch [1] prevents the shmem unused shrinker from being >> invoked during memcg-level reclaim altogether, but this is overly >> conservative: we can do better by reclaiming only the shmem charged to >> the reclaiming memcg. >> >> This series converts the shrinker list to a memcg-aware list_lru, so >> that non-root memcg reclaim walks only candidates charged to the >> reclaiming memcg. Global reclaim, root memcg reclaim and shmem quota >> reclaim retain their existing global semantics. >> >> To avoid pinning a dying memcg through a long-lived CSS reference, each >> inode stores an obj_cgroup reference instead of a mem_cgroup reference. >> The list_lru add/delete paths resolve the current memcg from the objcg >> under RCU, staying consistent with list_lru's own memcg migration on >> offline. > > Sashiko said a few things and they look disturbing-if-true: > > https://sashiko.dev/#/patchset/cover.1786955972.git.zhengqi.arch@bytedance.com > > (Apologies if this has already been considered - we don't have ways of > tracking all this (yet, I hope) apart from personal memory and personal > memorys are quite fried at present) As both Usama and I have pointed out [1][2], [PATCH v4 1/4] is the part that got dropped during the merge. The complete patch [3] was actually reviewed a while ago. [1].https://lore.kernel.org/all/20260810101954.822260-1-usama.arif@linux.dev/ [2]. https://lore.kernel.org/all/9c7efd5f-f8d3-4926-acb4-34c326ffb1c3@linux.dev/ [3]. https://lore.kernel.org/all/20260715103516.2410175-1-usama.arif@linux.dev/ Thanks, Qi >
© 2016 - 2026 Red Hat, Inc.