[PATCH v3 0/3] make unused huge shrinker memcg aware

Qi Zheng posted 3 patches 1 month, 4 weeks ago
There is a newer version of this series
fs/super.c               |  18 +-
include/linux/shmem_fs.h |  12 +-
mm/shmem.c               | 361 +++++++++++++++++++++++++++++----------
3 files changed, 283 insertions(+), 108 deletions(-)
[PATCH v3 0/3] make unused huge shrinker memcg aware
Posted by Qi Zheng 1 month, 4 weeks ago
From: Qi Zheng <zhengqi.arch@bytedance.com>

Changes in v3:
 - add a fix patch to fix missed removal of super_fs_objects_eligible()
 - move the original shrinklist addition logic after all checks are completed,
   and split it into a separate patch. (suggested by Baolin)
 - simplify the shmem_unused_huge_requeue() (suggested by Baolin)
 - keep the move_back label in shmem_unused_huge_shrink() (suggested by Baolin)
 - rebase onto the next-20260731

Changes in v2:
 - temporarily add the dependent patch from Usama to the series for review
   convenience
 - remove shrinklist_scan and shrinklist_isolated from struct shmem_inode_info,
   and re-implement the logic by resuing the same info->shrinklist
   (suggested by Baolin)
 - add more comments (suggested by Andrew)
 - fix missing initialization of info->shrinklist_memcg (pointed by sashiko)
 - rebase onto the next-20260717

Qi Zheng (3):
  fs: fix missed removal of super_fs_objects_eligible()
  mm: shmem: move unused huge shrinklist queuing past the truncation
    check
  mm: shmem: make unused huge shrinker memcg aware

 fs/super.c               |  18 +-
 include/linux/shmem_fs.h |  12 +-
 mm/shmem.c               | 361 +++++++++++++++++++++++++++++----------
 3 files changed, 283 insertions(+), 108 deletions(-)

-- 
2.54.0
Re: [PATCH v3 0/3] make unused huge shrinker memcg aware
Posted by Andrew Morton 1 month, 4 weeks ago
On Mon,  3 Aug 2026 16:46:32 +0800 Qi Zheng <qi.zheng@linux.dev> wrote:

> 

What David said.

Also, AI review might have found a couple of things:
	https://sashiko.dev/#/patchset/cover.1785746588.git.zhengqi.arch@bytedance.com
Re: [PATCH v3 0/3] make unused huge shrinker memcg aware
Posted by Qi Zheng 1 month, 4 weeks ago
Hi Andrew,

On 8/4/26 5:24 AM, Andrew Morton wrote:
> On Mon,  3 Aug 2026 16:46:32 +0800 Qi Zheng <qi.zheng@linux.dev> wrote:
> 
>>
> 
> What David said.
> 
> Also, AI review might have found a couple of things:
> 	https://sashiko.dev/#/patchset/cover.1785746588.git.zhengqi.arch@bytedance.com

I saw it, and it sounds reasonable.

 > > +static struct mem_cgroup *
 > > +shmem_unused_huge_alloc_lru(struct shmem_sb_info *sbinfo, struct 
folio *folio,
 > > +			    gfp_t gfp)
 > > +{
 > > +	struct mem_cgroup *memcg;
 > > +	int ret;
 > > +
 > > +	memcg = get_mem_cgroup_from_folio(folio);
 > > +	if (!memcg)
 > > +		return NULL;
 > Could this leak the mem_cgroup structure if a file is truncated?
 > When shmem_unused_huge_add() calls this function, it stores the memcg
 > reference in info->shrinklist_memcg.
 > If an unprivileged user truncates the file to 0 bytes, 
shmem_truncate_range()
 > frees the folio, but it does not remove the inode from the shrinklist 
or drop
 > this memcg reference.
 > Without memory pressure to trigger the shrinker, the mem_cgroup might 
remain
 > pinned indefinitely until the file is deleted, potentially allowing 
zombie
 > memcgs to accumulate on the system.

Indeed, will also call shmem_unused_huge_del() in shmem_truncate_range()
to take the inode off the shrinklist in time.

 > [ ... ]
 > > +static bool is_shmem_unused_huge_match(struct folio *folio,
 > > +				       struct shrink_control *sc)
 > > +{
 > > +	struct mem_cgroup *memcg = NULL;
 > > +	bool match;
 > > +
 > > +	/*
 > > +	 * Only non-root memcg reclaim needs to match the folio charge 
against
 > > +	 * sc->memcg. Skip the folio memcg check for the following cases:
 > > +	 * 1. shmem quota reclaim (sc == NULL)
 > > +	 * 2. global shrinker reclaim
 > > +	 * 3. root memcg reclaim
 > > +	 */
 > > +	if (!sc || !sc->memcg || mem_cgroup_is_root(sc->memcg))
 > > +		return true;
 > > +
 > > +	if (folio_nid(folio) != sc->nid)
 > > +		return false;
 > Does this early return bypass the NUMA node check for global and root 
memcg
 > reclaim?
 > The comment indicates this is only intended to skip the folio memcg 
check.
 > However, returning true immediately also bypasses the sc->nid check.
 > If a huge folio is migrated to a different NUMA node, kswapd or direct
 > reclaim might incorrectly split folios on the wrong node instead of
 > requeuing the inode to the correct node's list.

Indeed, root memcg reclaim also needs to check if the nid matches.

Will move the nid check earlier in v4:

       if (!sc)
               return true;

       if (folio_nid(folio) != sc->nid)
               return false;

       if (!sc->memcg || mem_cgroup_is_root(sc->memcg))
               return true;

Thanks,
Qi
Re: [PATCH v3 0/3] make unused huge shrinker memcg aware
Posted by David Hildenbrand (Arm) 1 month, 4 weeks ago
On 8/3/26 10:46, Qi Zheng wrote:
> From: Qi Zheng <zhengqi.arch@bytedance.com>

I'm missing description and motivation here.

This is only about shrinking huge pages that span end of shmem files.

Is this really a problem? And if so, why?

-- 
Cheers,

David
Re: [PATCH v3 0/3] make unused huge shrinker memcg aware
Posted by Qi Zheng 1 month, 4 weeks ago
Hi David,

On 8/3/26 8:21 PM, David Hildenbrand (Arm) wrote:
> On 8/3/26 10:46, Qi Zheng wrote:
>> From: Qi Zheng <zhengqi.arch@bytedance.com>
> 
> I'm missing description and motivation here.

My bad, since v1 was just a single patch, I got lazy and didn't bother
adding a cover letter description later on.

> 
> This is only about shrinking huge pages that span end of shmem files.
> 
> Is this really a problem? And if so, why?

Yes, this is a real-world problem that we encountered in production.

The root cause is that shmem unused shrinker used to be non-memcg-aware.
This could lead to a scenario where reclaim triggered by one memcg A
reclaims the shmem of another memcg B, causing unexpected impact on it.

Even worse, memcg A might have no reclaimable shmem at all, making this
completely useless work and incurring some performance overhead.

such as:

tid 11340 comm scanner locked a page for 182264 us! kstack:
         unlock_page+1
         split_huge_page_to_list+3135
         shmem_unused_huge_shrink+767
         super_cache_scan+329
         do_shrink_slab+291
         shrink_slab+533
         shrink_node+400
         do_try_to_free_pages+206
         try_to_free_mem_cgroup_pages+262
         try_charge_memcg+591
         mem_cgroup_charge+136
         __handle_mm_fault+2431
         handle_mm_fault+194
         do_user_addr_fault+462
         __do_page_fault+176
         do_page_fault+48
         page_fault+62

Later, with Usama's patch [1], the shmem unused shrinker is no longer
triggered during memcg-level reclaim. But this actually doesn't make
sense either, since we can clearly just reclaim from memcg A individuallty.

[1]. 
https://lore.kernel.org/all/20260715103516.2410175-1-usama.arif@linux.dev/

Thanks,
Qi


>