[PATCH v4 0/4] mm: workingset: fix the shadow node budget under MGLRU

Hui Zhu posted 4 patches 2 weeks, 3 days ago
There is a newer version of this series
mm/memcontrol-v1.h |  5 ++--
mm/memcontrol.c    | 72 +++++++++++++++++++++++++---------------------
mm/workingset.c    |  5 ++--
3 files changed, 45 insertions(+), 37 deletions(-)
[PATCH v4 0/4] mm: workingset: fix the shadow node budget under MGLRU
Posted by Hui Zhu 2 weeks, 3 days ago
From: Hui Zhu <zhuhui@kylinos.cn>

Commit 7404bd37cfbe ("mm: workingset: use lruvec_lru_size() to get the
number of lru pages") broke the workingset shadow node budget under
MGLRU: lruvec_lru_size() reads mz->lru_zone_size, which MGLRU never
maintains, so count_shadow_nodes() sees the evictable LRU lists as
empty and the shadow shrinker reclaims eviction tokens almost as fast
as they are created, losing thrashing protection.

Patch 1 extends the dying-mcg stat redirection (previously cgroup v1
only) to all hierarchies, addressing the reparenting race that motivated
7404bd37cfbe.

Patch 2 then switches count_shadow_nodes() back to
lruvec_page_state_local(), which both classic LRU and MGLRU maintain.

Patch 3 recovers the performance.  Patch 1 added an unconditional
rcu_read_lock() to the stat update fast path; patch 3 checks
memcg_is_dying() first and takes the RCU lock only on the rare dying
path.

Patch 4 closes an accounting gap that patch 2 makes visible: on cgroup
v2, reparent_state_local() never moves the dying memcg's
non-hierarchical lruvec stats to the parent, so the parent receives
the uncharges without the matching charges and its state_local
underflows.  Patch 4 reparents those stats, mirroring cgroup v1.

Performance testing
===================

The test script and the raw results are available at [1].

Environment: 10-vCPU QEMU guest, 8 GiB RAM, cgroup v2; 7 runs per
configuration, medians reported.  Workloads:

  w1-anon-churn: single-threaded anon fault/charge loop in a memcg
                 (MADV_DONTNEED + re-fault, no reclaim).  Every touch
                 is a real fault with charge and memcg stat updates,
                 so it stresses exactly the fast path patch 1 changes.
  w2-file-churn: file read loop under memory.high pressure
                 (reclaim-bound, noisier).
  w3-reparent:   reparent accounting sanity check.

w1-anon-churn (pages/s):

                 classic LRU          MGLRU
base             4393028              4377122
patches 1-2      4385996   (-0.2%)    4352887   (-0.6%)
patches 1-3      4381832   (-0.3%)    4377053   (+0.0%)

w2-file-churn (MB/s):

                 classic LRU          MGLRU
base             8277                 8226
patches 1-2      8226      (-0.6%)    8123      (-1.3%)
patches 1-3      8157      (-1.4%)    8294      (+0.8%)

w3-reparent passed on all kernels.

The small overhead visible with patches 1-2 comes from the redirection
added by patch 1; patch 3 brings w1 back to the base level in both LRU
configurations.  The remaining differences are within run-to-run noise.
Patch 4 only touches the memcg offline path and is not exercised by
these workloads.

[1] https://gist.github.com/teawater/32f373ec41d185d840455eb167321a5a

Changelog:
v4:
According to the comments of Andrew, Fix "follow-up patch" to
"preceding patch" in the commit message to match the reordered series.
According to the comments of Sashiko, add patch 4 to reparent the
non-hierarchical lruvec stats on cgroup v2 to fixing the state_local
underflow.
v3:
According to the comments of Shakeel, reorder the series per review,
add Fixes/Cc stable and the user-visible impact to patch 1.

Hui Zhu (4):
  mm: memcg: redirect stats updates of dying memcgs for all hierarchies
  mm: workingset: use lruvec_page_state_local() to count lru pages
  mm: memcg: skip the RCU lock when the memcg is not dying
  mm: memcg: reparent non-hierarchical lruvec stats on cgroup v2

 mm/memcontrol-v1.h |  5 ++--
 mm/memcontrol.c    | 72 +++++++++++++++++++++++++---------------------
 mm/workingset.c    |  5 ++--
 3 files changed, 45 insertions(+), 37 deletions(-)

-- 
2.43.0
Re: [PATCH v4 0/4] mm: workingset: fix the shadow node budget under MGLRU
Posted by Andrew Morton 2 weeks, 2 days ago
On Tue,  8 Sep 2026 11:41:10 +0800 Hui Zhu <hui.zhu@linux.dev> wrote:

> From: Hui Zhu <zhuhui@kylinos.cn>
> 
> Commit 7404bd37cfbe ("mm: workingset: use lruvec_lru_size() to get the
> number of lru pages") broke the workingset shadow node budget under
> MGLRU: lruvec_lru_size() reads mz->lru_zone_size, which MGLRU never
> maintains, so count_shadow_nodes() sees the evictable LRU lists as
> empty and the shadow shrinker reclaims eviction tokens almost as fast
> as they are created, losing thrashing protection.
> 
> ...
>
> Performance testing
> ===================
> 
> The test script and the raw results are available at [1].
> 
> Environment: 10-vCPU QEMU guest, 8 GiB RAM, cgroup v2; 7 runs per
> configuration, medians reported.  Workloads:
> 
>   w1-anon-churn: single-threaded anon fault/charge loop in a memcg
>                  (MADV_DONTNEED + re-fault, no reclaim).  Every touch
>                  is a real fault with charge and memcg stat updates,
>                  so it stresses exactly the fast path patch 1 changes.
>   w2-file-churn: file read loop under memory.high pressure
>                  (reclaim-bound, noisier).
>   w3-reparent:   reparent accounting sanity check.
> 
> w1-anon-churn (pages/s):
> 
>                  classic LRU          MGLRU
> base             4393028              4377122
> patches 1-2      4385996   (-0.2%)    4352887   (-0.6%)
> patches 1-3      4381832   (-0.3%)    4377053   (+0.0%)
> 
> w2-file-churn (MB/s):
> 
>                  classic LRU          MGLRU
> base             8277                 8226
> patches 1-2      8226      (-0.6%)    8123      (-1.3%)
> patches 1-3      8157      (-1.4%)    8294      (+0.8%)

Am I misinterpreting this?  This difference is probably within
inter-run variability?
Re: [PATCH v4 0/4] mm: workingset: fix the shadow node budget under MGLRU
Posted by Hui Zhu 2 weeks, 1 day ago
> On Tue,  8 Sep 2026 11:41:10 +0800 Hui Zhu <hui.zhu@linux.dev> wrote:
>
>> From: Hui Zhu <zhuhui@kylinos.cn>
>>
>> Commit 7404bd37cfbe ("mm: workingset: use lruvec_lru_size() to get the
>> number of lru pages") broke the workingset shadow node budget under
>> MGLRU: lruvec_lru_size() reads mz->lru_zone_size, which MGLRU never
>> maintains, so count_shadow_nodes() sees the evictable LRU lists as
>> empty and the shadow shrinker reclaims eviction tokens almost as fast
>> as they are created, losing thrashing protection.
>>
>> ...
>>
>> Performance testing
>> ===================
>>
>> The test script and the raw results are available at [1].
>>
>> Environment: 10-vCPU QEMU guest, 8 GiB RAM, cgroup v2; 7 runs per
>> configuration, medians reported.  Workloads:
>>
>>    w1-anon-churn: single-threaded anon fault/charge loop in a memcg
>>                   (MADV_DONTNEED + re-fault, no reclaim).  Every touch
>>                   is a real fault with charge and memcg stat updates,
>>                   so it stresses exactly the fast path patch 1 changes.
>>    w2-file-churn: file read loop under memory.high pressure
>>                   (reclaim-bound, noisier).
>>    w3-reparent:   reparent accounting sanity check.
>>
>> w1-anon-churn (pages/s):
>>
>>                   classic LRU          MGLRU
>> base             4393028              4377122
>> patches 1-2      4385996   (-0.2%)    4352887   (-0.6%)
>> patches 1-3      4381832   (-0.3%)    4377053   (+0.0%)
>>
>> w2-file-churn (MB/s):
>>
>>                   classic LRU          MGLRU
>> base             8277                 8226
>> patches 1-2      8226      (-0.6%)    8123      (-1.3%)
>> patches 1-3      8157      (-1.4%)    8294      (+0.8%)
> Am I misinterpreting this?  This difference is probably within
> inter-run variability?
>
You are reading it correctly.
The w2-file-churn differences are within run-to-run noise: it is a
reclaim-bound workload dominated by reclaim and I/O, which is too noisy
to expose the small fast-path overhead.
That is what the "within run-to-run noise" note in the cover letter
refers to.

The meaningful signal is in w1-anon-churn, which is designed to hit
exactly the fast path patch 1 changes: every iteration is a real fault
with charge and memcg stat updates, no reclaim involved.
There patches 1-2 show a consistent small overhead (-0.2%/-0.6%), and
patch 3 brings both LRU configurations back to the base level.

I can reword the cover letter in the next version to make this clearer
if you think it would help.

Best,
Hui