fs/super.c | 18 +- include/linux/shmem_fs.h | 12 +- mm/shmem.c | 361 +++++++++++++++++++++++++++++---------- 3 files changed, 283 insertions(+), 108 deletions(-)
From: Qi Zheng <zhengqi.arch@bytedance.com>
Changes in v3:
- add a fix patch to fix missed removal of super_fs_objects_eligible()
- move the original shrinklist addition logic after all checks are completed,
and split it into a separate patch. (suggested by Baolin)
- simplify the shmem_unused_huge_requeue() (suggested by Baolin)
- keep the move_back label in shmem_unused_huge_shrink() (suggested by Baolin)
- rebase onto the next-20260731
Changes in v2:
- temporarily add the dependent patch from Usama to the series for review
convenience
- remove shrinklist_scan and shrinklist_isolated from struct shmem_inode_info,
and re-implement the logic by resuing the same info->shrinklist
(suggested by Baolin)
- add more comments (suggested by Andrew)
- fix missing initialization of info->shrinklist_memcg (pointed by sashiko)
- rebase onto the next-20260717
Qi Zheng (3):
fs: fix missed removal of super_fs_objects_eligible()
mm: shmem: move unused huge shrinklist queuing past the truncation
check
mm: shmem: make unused huge shrinker memcg aware
fs/super.c | 18 +-
include/linux/shmem_fs.h | 12 +-
mm/shmem.c | 361 +++++++++++++++++++++++++++++----------
3 files changed, 283 insertions(+), 108 deletions(-)
--
2.54.0
On Mon, 3 Aug 2026 16:46:32 +0800 Qi Zheng <qi.zheng@linux.dev> wrote: > What David said. Also, AI review might have found a couple of things: https://sashiko.dev/#/patchset/cover.1785746588.git.zhengqi.arch@bytedance.com
Hi Andrew,
On 8/4/26 5:24 AM, Andrew Morton wrote:
> On Mon, 3 Aug 2026 16:46:32 +0800 Qi Zheng <qi.zheng@linux.dev> wrote:
>
>>
>
> What David said.
>
> Also, AI review might have found a couple of things:
> https://sashiko.dev/#/patchset/cover.1785746588.git.zhengqi.arch@bytedance.com
I saw it, and it sounds reasonable.
> > +static struct mem_cgroup *
> > +shmem_unused_huge_alloc_lru(struct shmem_sb_info *sbinfo, struct
folio *folio,
> > + gfp_t gfp)
> > +{
> > + struct mem_cgroup *memcg;
> > + int ret;
> > +
> > + memcg = get_mem_cgroup_from_folio(folio);
> > + if (!memcg)
> > + return NULL;
> Could this leak the mem_cgroup structure if a file is truncated?
> When shmem_unused_huge_add() calls this function, it stores the memcg
> reference in info->shrinklist_memcg.
> If an unprivileged user truncates the file to 0 bytes,
shmem_truncate_range()
> frees the folio, but it does not remove the inode from the shrinklist
or drop
> this memcg reference.
> Without memory pressure to trigger the shrinker, the mem_cgroup might
remain
> pinned indefinitely until the file is deleted, potentially allowing
zombie
> memcgs to accumulate on the system.
Indeed, will also call shmem_unused_huge_del() in shmem_truncate_range()
to take the inode off the shrinklist in time.
> [ ... ]
> > +static bool is_shmem_unused_huge_match(struct folio *folio,
> > + struct shrink_control *sc)
> > +{
> > + struct mem_cgroup *memcg = NULL;
> > + bool match;
> > +
> > + /*
> > + * Only non-root memcg reclaim needs to match the folio charge
against
> > + * sc->memcg. Skip the folio memcg check for the following cases:
> > + * 1. shmem quota reclaim (sc == NULL)
> > + * 2. global shrinker reclaim
> > + * 3. root memcg reclaim
> > + */
> > + if (!sc || !sc->memcg || mem_cgroup_is_root(sc->memcg))
> > + return true;
> > +
> > + if (folio_nid(folio) != sc->nid)
> > + return false;
> Does this early return bypass the NUMA node check for global and root
memcg
> reclaim?
> The comment indicates this is only intended to skip the folio memcg
check.
> However, returning true immediately also bypasses the sc->nid check.
> If a huge folio is migrated to a different NUMA node, kswapd or direct
> reclaim might incorrectly split folios on the wrong node instead of
> requeuing the inode to the correct node's list.
Indeed, root memcg reclaim also needs to check if the nid matches.
Will move the nid check earlier in v4:
if (!sc)
return true;
if (folio_nid(folio) != sc->nid)
return false;
if (!sc->memcg || mem_cgroup_is_root(sc->memcg))
return true;
Thanks,
Qi
On 8/3/26 10:46, Qi Zheng wrote: > From: Qi Zheng <zhengqi.arch@bytedance.com> I'm missing description and motivation here. This is only about shrinking huge pages that span end of shmem files. Is this really a problem? And if so, why? -- Cheers, David
Hi David,
On 8/3/26 8:21 PM, David Hildenbrand (Arm) wrote:
> On 8/3/26 10:46, Qi Zheng wrote:
>> From: Qi Zheng <zhengqi.arch@bytedance.com>
>
> I'm missing description and motivation here.
My bad, since v1 was just a single patch, I got lazy and didn't bother
adding a cover letter description later on.
>
> This is only about shrinking huge pages that span end of shmem files.
>
> Is this really a problem? And if so, why?
Yes, this is a real-world problem that we encountered in production.
The root cause is that shmem unused shrinker used to be non-memcg-aware.
This could lead to a scenario where reclaim triggered by one memcg A
reclaims the shmem of another memcg B, causing unexpected impact on it.
Even worse, memcg A might have no reclaimable shmem at all, making this
completely useless work and incurring some performance overhead.
such as:
tid 11340 comm scanner locked a page for 182264 us! kstack:
unlock_page+1
split_huge_page_to_list+3135
shmem_unused_huge_shrink+767
super_cache_scan+329
do_shrink_slab+291
shrink_slab+533
shrink_node+400
do_try_to_free_pages+206
try_to_free_mem_cgroup_pages+262
try_charge_memcg+591
mem_cgroup_charge+136
__handle_mm_fault+2431
handle_mm_fault+194
do_user_addr_fault+462
__do_page_fault+176
do_page_fault+48
page_fault+62
Later, with Usama's patch [1], the shmem unused shrinker is no longer
triggered during memcg-level reclaim. But this actually doesn't make
sense either, since we can clearly just reclaim from memcg A individuallty.
[1].
https://lore.kernel.org/all/20260715103516.2410175-1-usama.arif@linux.dev/
Thanks,
Qi
>
© 2016 - 2026 Red Hat, Inc.