[PATCH v4 0/4] make unused huge shrinker memcg aware

Qi Zheng posted 4 patches 1 month, 1 week ago
fs/super.c                 |  18 +-
include/linux/memcontrol.h |  11 +-
include/linux/shmem_fs.h   |  12 +-
mm/page_owner.c            |   2 +-
mm/shmem.c                 | 382 ++++++++++++++++++++++++++++---------
mm/zswap.c                 |  17 +-
6 files changed, 315 insertions(+), 127 deletions(-)
[PATCH v4 0/4] make unused huge shrinker memcg aware
Posted by Qi Zheng 1 month, 1 week ago
From: Qi Zheng <zhengqi.arch@bytedance.com>

Changes in v4:
 - add [PATCH v4 2/4] to make obj_cgroup_memcg() handle NULL objcg
 - store obj_cgroup instead of mem_cgroup in shmem_inode_info to avoid
   pinning a dying memcg through a long-lived CSS reference
   (pointed by sashiko)
 - fix is_shmem_unused_huge_match() to always check the NUMA node for
   shrinker reclaim, not only for non-root memcg reclaim
   (pointed by sashiko)
 - collect Reviewed-by
 - rebase onto the next-20260814

Note: [PATCH v4 1/4] should ideally be folded into commit 0ef8faff490be
("fs: push nr_cached_objects memcg gating into individual filesystems") in
linux-next.

Changes in v3:
 - add a fix patch to fix missed removal of super_fs_objects_eligible()
 - move the original shrinklist addition logic after all checks are completed,
   and split it into a separate patch. (suggested by Baolin)
 - simplify the shmem_unused_huge_requeue() (suggested by Baolin)
 - keep the move_back label in shmem_unused_huge_shrink() (suggested by Baolin)
 - rebase onto the next-20260731

Changes in v2:
 - temporarily add the dependent patch from Usama to the series for review
   convenience
 - remove shrinklist_scan and shrinklist_isolated from struct shmem_inode_info,
   and re-implement the logic by resuing the same info->shrinklist
   (suggested by Baolin)
 - add more comments (suggested by Andrew)
 - fix missing initialization of info->shrinklist_memcg (pointed by sashiko)
 - rebase onto the next-20260717

Hi all,

The shmem unused huge shrinker maintains a per-superblock list of inodes
whose tail huge folio extends beyond i_size.  Because this list is not
memcg aware, reclaim triggered by memcg A can scan inodes across the
entire superblock and split huge folios charged to unrelated memcg B,
causing unexpected impact on it.

In the worst case, memcg A has no reclaimable shmem at all, making the
reclaim entirely useless and incurring unnecessary latency.  We observed
this in production, where page lock contention during split caused
multi-hundred-millisecond stalls:

  tid 11340 comm scanner locked a page for 182264 us! kstack:
          unlock_page+1
          split_huge_page_to_list+3135
          shmem_unused_huge_shrink+767
          super_cache_scan+329
          do_shrink_slab+291
          shrink_slab+533
          shrink_node+400
          do_try_to_free_pages+206
          try_to_free_mem_cgroup_pages+262
          try_charge_memcg+591
          mem_cgroup_charge+136
          __handle_mm_fault+2431
          handle_mm_fault+194
          do_user_addr_fault+462
          __do_page_fault+176
          do_page_fault+48
          page_fault+62

Usama's recent patch [1] prevents the shmem unused shrinker from being
invoked during memcg-level reclaim altogether, but this is overly
conservative: we can do better by reclaiming only the shmem charged to
the reclaiming memcg.

This series converts the shrinker list to a memcg-aware list_lru, so
that non-root memcg reclaim walks only candidates charged to the
reclaiming memcg.  Global reclaim, root memcg reclaim and shmem quota
reclaim retain their existing global semantics.

To avoid pinning a dying memcg through a long-lived CSS reference, each
inode stores an obj_cgroup reference instead of a mem_cgroup reference.
The list_lru add/delete paths resolve the current memcg from the objcg
under RCU, staying consistent with list_lru's own memcg migration on
offline.

Thanks,
Qi

[1]. https://lore.kernel.org/all/20260715103516.2410175-1-usama.arif@linux.dev/

Qi Zheng (4):
  fs: fix missed removal of super_fs_objects_eligible()
  mm: memcontrol: make obj_cgroup_memcg() handle NULL objcg
  mm: shmem: move unused huge shrinklist queuing past the truncation
    check
  mm: shmem: make unused huge shrinker memcg aware

 fs/super.c                 |  18 +-
 include/linux/memcontrol.h |  11 +-
 include/linux/shmem_fs.h   |  12 +-
 mm/page_owner.c            |   2 +-
 mm/shmem.c                 | 382 ++++++++++++++++++++++++++++---------
 mm/zswap.c                 |  17 +-
 6 files changed, 315 insertions(+), 127 deletions(-)

-- 
2.54.0
Re: [PATCH v4 0/4] make unused huge shrinker memcg aware
Posted by Andrew Morton 1 month ago
On Mon, 17 Aug 2026 17:03:24 +0800 Qi Zheng <qi.zheng@linux.dev> wrote:

> Changes in v4:
> Changes in v3:
> Changes in v2:

Thanks for the diligent versioning info.  fyi, it is conventional to
maintain this below the --- separator.  It's not really the most
important part of the [0/N]!

> 
> The shmem unused huge shrinker maintains a per-superblock list of inodes
> whose tail huge folio extends beyond i_size.  Because this list is not
> memcg aware, reclaim triggered by memcg A can scan inodes across the
> entire superblock and split huge folios charged to unrelated memcg B,
> causing unexpected impact on it.
> 
> In the worst case, memcg A has no reclaimable shmem at all, making the
> reclaim entirely useless and incurring unnecessary latency.  We observed
> this in production, where page lock contention during split caused
> multi-hundred-millisecond stalls:

Ugh.  That's the most important part!

>   tid 11340 comm scanner locked a page for 182264 us! kstack:
>           unlock_page+1
>           split_huge_page_to_list+3135
>           shmem_unused_huge_shrink+767
>           super_cache_scan+329
>           do_shrink_slab+291
>           shrink_slab+533
>           shrink_node+400
>           do_try_to_free_pages+206
>           try_to_free_mem_cgroup_pages+262
>           try_charge_memcg+591
>           mem_cgroup_charge+136
>           __handle_mm_fault+2431
>           handle_mm_fault+194
>           do_user_addr_fault+462
>           __do_page_fault+176
>           do_page_fault+48
>           page_fault+62
> 
> Usama's recent patch [1] prevents the shmem unused shrinker from being
> invoked during memcg-level reclaim altogether, but this is overly
> conservative: we can do better by reclaiming only the shmem charged to
> the reclaiming memcg.
> 
> This series converts the shrinker list to a memcg-aware list_lru, so
> that non-root memcg reclaim walks only candidates charged to the
> reclaiming memcg.  Global reclaim, root memcg reclaim and shmem quota
> reclaim retain their existing global semantics.
> 
> To avoid pinning a dying memcg through a long-lived CSS reference, each
> inode stores an obj_cgroup reference instead of a mem_cgroup reference.
> The list_lru add/delete paths resolve the current memcg from the objcg
> under RCU, staying consistent with list_lru's own memcg migration on
> offline.

Sashiko said a few things and they look disturbing-if-true:

	https://sashiko.dev/#/patchset/cover.1786955972.git.zhengqi.arch@bytedance.com

(Apologies if this has already been considered - we don't have ways of
tracking all this (yet, I hope) apart from personal memory and personal
memorys are quite fried at present)
Re: [PATCH v4 0/4] make unused huge shrinker memcg aware
Posted by Qi Zheng 1 month ago
Hi Andrew,

On 8/28/26 6:50 AM, Andrew Morton wrote:
> On Mon, 17 Aug 2026 17:03:24 +0800 Qi Zheng <qi.zheng@linux.dev> wrote:
> 
>> Changes in v4:
>> Changes in v3:
>> Changes in v2:
> 
> Thanks for the diligent versioning info.  fyi, it is conventional to
> maintain this below the --- separator.  It's not really the most
> important part of the [0/N]!

Got it.

> 
>>
>> The shmem unused huge shrinker maintains a per-superblock list of inodes
>> whose tail huge folio extends beyond i_size.  Because this list is not
>> memcg aware, reclaim triggered by memcg A can scan inodes across the
>> entire superblock and split huge folios charged to unrelated memcg B,
>> causing unexpected impact on it.
>>
>> In the worst case, memcg A has no reclaimable shmem at all, making the
>> reclaim entirely useless and incurring unnecessary latency.  We observed
>> this in production, where page lock contention during split caused
>> multi-hundred-millisecond stalls:
> 
> Ugh.  That's the most important part!
> 
>>    tid 11340 comm scanner locked a page for 182264 us! kstack:
>>            unlock_page+1
>>            split_huge_page_to_list+3135
>>            shmem_unused_huge_shrink+767
>>            super_cache_scan+329
>>            do_shrink_slab+291
>>            shrink_slab+533
>>            shrink_node+400
>>            do_try_to_free_pages+206
>>            try_to_free_mem_cgroup_pages+262
>>            try_charge_memcg+591
>>            mem_cgroup_charge+136
>>            __handle_mm_fault+2431
>>            handle_mm_fault+194
>>            do_user_addr_fault+462
>>            __do_page_fault+176
>>            do_page_fault+48
>>            page_fault+62
>>
>> Usama's recent patch [1] prevents the shmem unused shrinker from being
>> invoked during memcg-level reclaim altogether, but this is overly
>> conservative: we can do better by reclaiming only the shmem charged to
>> the reclaiming memcg.
>>
>> This series converts the shrinker list to a memcg-aware list_lru, so
>> that non-root memcg reclaim walks only candidates charged to the
>> reclaiming memcg.  Global reclaim, root memcg reclaim and shmem quota
>> reclaim retain their existing global semantics.
>>
>> To avoid pinning a dying memcg through a long-lived CSS reference, each
>> inode stores an obj_cgroup reference instead of a mem_cgroup reference.
>> The list_lru add/delete paths resolve the current memcg from the objcg
>> under RCU, staying consistent with list_lru's own memcg migration on
>> offline.
> 
> Sashiko said a few things and they look disturbing-if-true:
> 
> 	https://sashiko.dev/#/patchset/cover.1786955972.git.zhengqi.arch@bytedance.com
> 
> (Apologies if this has already been considered - we don't have ways of
> tracking all this (yet, I hope) apart from personal memory and personal
> memorys are quite fried at present)

As both Usama and I have pointed out [1][2], [PATCH v4 1/4] is the part
that got dropped during the merge. The complete patch [3] was actually
reviewed a while ago.

[1].https://lore.kernel.org/all/20260810101954.822260-1-usama.arif@linux.dev/
[2]. 
https://lore.kernel.org/all/9c7efd5f-f8d3-4926-acb4-34c326ffb1c3@linux.dev/
[3]. 
https://lore.kernel.org/all/20260715103516.2410175-1-usama.arif@linux.dev/

Thanks,
Qi

>