mm/memcontrol-v1.h | 5 ++-- mm/memcontrol.c | 72 +++++++++++++++++++++++++--------------------- mm/workingset.c | 5 ++-- 3 files changed, 45 insertions(+), 37 deletions(-)
From: Hui Zhu <zhuhui@kylinos.cn>
Commit 7404bd37cfbe ("mm: workingset: use lruvec_lru_size() to get the
number of lru pages") broke the workingset shadow node budget under
MGLRU: lruvec_lru_size() reads mz->lru_zone_size, which MGLRU never
maintains, so count_shadow_nodes() sees the evictable LRU lists as
empty and the shadow shrinker reclaims eviction tokens almost as fast
as they are created, losing thrashing protection.
Patch 1 extends the dying-mcg stat redirection (previously cgroup v1
only) to all hierarchies, addressing the reparenting race that motivated
7404bd37cfbe.
Patch 2 then switches count_shadow_nodes() back to
lruvec_page_state_local(), which both classic LRU and MGLRU maintain.
Patch 3 recovers the performance. Patch 1 added an unconditional
rcu_read_lock() to the stat update fast path; patch 3 checks
memcg_is_dying() first and takes the RCU lock only on the rare dying
path.
Patch 4 closes an accounting gap that patch 2 makes visible: on cgroup
v2, reparent_state_local() never moves the dying memcg's
non-hierarchical lruvec stats to the parent, so the parent receives
the uncharges without the matching charges and its state_local
underflows. Patch 4 reparents those stats, mirroring cgroup v1.
Performance testing
===================
The test script and the raw results are available at [1].
Environment: 10-vCPU QEMU guest, 8 GiB RAM, cgroup v2; 7 runs per
configuration, medians reported. Workloads:
w1-anon-churn: single-threaded anon fault/charge loop in a memcg
(MADV_DONTNEED + re-fault, no reclaim). Every touch
is a real fault with charge and memcg stat updates,
so it stresses exactly the fast path patch 1 changes.
w2-file-churn: file read loop under memory.high pressure
(reclaim-bound, noisier).
w3-reparent: reparent accounting sanity check.
w1-anon-churn (pages/s):
classic LRU MGLRU
base 4393028 4377122
patches 1-2 4385996 (-0.2%) 4352887 (-0.6%)
patches 1-3 4381832 (-0.3%) 4377053 (+0.0%)
w2-file-churn (MB/s):
classic LRU MGLRU
base 8277 8226
patches 1-2 8226 (-0.6%) 8123 (-1.3%)
patches 1-3 8157 (-1.4%) 8294 (+0.8%)
w3-reparent passed on all kernels.
The small overhead visible with patches 1-2 comes from the redirection
added by patch 1; patch 3 brings w1 back to the base level in both LRU
configurations. The remaining differences are within run-to-run noise.
Patch 4 only touches the memcg offline path and is not exercised by
these workloads.
[1] https://gist.github.com/teawater/32f373ec41d185d840455eb167321a5a
Changelog:
v4:
According to the comments of Andrew, Fix "follow-up patch" to
"preceding patch" in the commit message to match the reordered series.
According to the comments of Sashiko, add patch 4 to reparent the
non-hierarchical lruvec stats on cgroup v2 to fixing the state_local
underflow.
v3:
According to the comments of Shakeel, reorder the series per review,
add Fixes/Cc stable and the user-visible impact to patch 1.
Hui Zhu (4):
mm: memcg: redirect stats updates of dying memcgs for all hierarchies
mm: workingset: use lruvec_page_state_local() to count lru pages
mm: memcg: skip the RCU lock when the memcg is not dying
mm: memcg: reparent non-hierarchical lruvec stats on cgroup v2
mm/memcontrol-v1.h | 5 ++--
mm/memcontrol.c | 72 +++++++++++++++++++++++++---------------------
mm/workingset.c | 5 ++--
3 files changed, 45 insertions(+), 37 deletions(-)
--
2.43.0
On Tue, 8 Sep 2026 11:41:10 +0800 Hui Zhu <hui.zhu@linux.dev> wrote:
> From: Hui Zhu <zhuhui@kylinos.cn>
>
> Commit 7404bd37cfbe ("mm: workingset: use lruvec_lru_size() to get the
> number of lru pages") broke the workingset shadow node budget under
> MGLRU: lruvec_lru_size() reads mz->lru_zone_size, which MGLRU never
> maintains, so count_shadow_nodes() sees the evictable LRU lists as
> empty and the shadow shrinker reclaims eviction tokens almost as fast
> as they are created, losing thrashing protection.
>
> ...
>
> Performance testing
> ===================
>
> The test script and the raw results are available at [1].
>
> Environment: 10-vCPU QEMU guest, 8 GiB RAM, cgroup v2; 7 runs per
> configuration, medians reported. Workloads:
>
> w1-anon-churn: single-threaded anon fault/charge loop in a memcg
> (MADV_DONTNEED + re-fault, no reclaim). Every touch
> is a real fault with charge and memcg stat updates,
> so it stresses exactly the fast path patch 1 changes.
> w2-file-churn: file read loop under memory.high pressure
> (reclaim-bound, noisier).
> w3-reparent: reparent accounting sanity check.
>
> w1-anon-churn (pages/s):
>
> classic LRU MGLRU
> base 4393028 4377122
> patches 1-2 4385996 (-0.2%) 4352887 (-0.6%)
> patches 1-3 4381832 (-0.3%) 4377053 (+0.0%)
>
> w2-file-churn (MB/s):
>
> classic LRU MGLRU
> base 8277 8226
> patches 1-2 8226 (-0.6%) 8123 (-1.3%)
> patches 1-3 8157 (-1.4%) 8294 (+0.8%)
Am I misinterpreting this? This difference is probably within
inter-run variability?
> On Tue, 8 Sep 2026 11:41:10 +0800 Hui Zhu <hui.zhu@linux.dev> wrote:
>
>> From: Hui Zhu <zhuhui@kylinos.cn>
>>
>> Commit 7404bd37cfbe ("mm: workingset: use lruvec_lru_size() to get the
>> number of lru pages") broke the workingset shadow node budget under
>> MGLRU: lruvec_lru_size() reads mz->lru_zone_size, which MGLRU never
>> maintains, so count_shadow_nodes() sees the evictable LRU lists as
>> empty and the shadow shrinker reclaims eviction tokens almost as fast
>> as they are created, losing thrashing protection.
>>
>> ...
>>
>> Performance testing
>> ===================
>>
>> The test script and the raw results are available at [1].
>>
>> Environment: 10-vCPU QEMU guest, 8 GiB RAM, cgroup v2; 7 runs per
>> configuration, medians reported. Workloads:
>>
>> w1-anon-churn: single-threaded anon fault/charge loop in a memcg
>> (MADV_DONTNEED + re-fault, no reclaim). Every touch
>> is a real fault with charge and memcg stat updates,
>> so it stresses exactly the fast path patch 1 changes.
>> w2-file-churn: file read loop under memory.high pressure
>> (reclaim-bound, noisier).
>> w3-reparent: reparent accounting sanity check.
>>
>> w1-anon-churn (pages/s):
>>
>> classic LRU MGLRU
>> base 4393028 4377122
>> patches 1-2 4385996 (-0.2%) 4352887 (-0.6%)
>> patches 1-3 4381832 (-0.3%) 4377053 (+0.0%)
>>
>> w2-file-churn (MB/s):
>>
>> classic LRU MGLRU
>> base 8277 8226
>> patches 1-2 8226 (-0.6%) 8123 (-1.3%)
>> patches 1-3 8157 (-1.4%) 8294 (+0.8%)
> Am I misinterpreting this? This difference is probably within
> inter-run variability?
>
You are reading it correctly.
The w2-file-churn differences are within run-to-run noise: it is a
reclaim-bound workload dominated by reclaim and I/O, which is too noisy
to expose the small fast-path overhead.
That is what the "within run-to-run noise" note in the cover letter
refers to.
The meaningful signal is in w1-anon-churn, which is designed to hit
exactly the fast path patch 1 changes: every iteration is a real fault
with charge and memcg stat updates, no reclaim involved.
There patches 1-2 show a consistent small overhead (-0.2%/-0.6%), and
patch 3 brings both LRU configurations back to the base level.
I can reword the cover letter in the next version to make this clearer
if you think it would help.
Best,
Hui
© 2016 - 2026 Red Hat, Inc.