include/linux/mmzone.h | 2 + mm/vmscan.c | 173 +++++++++++++++++++++++++++++++++++++---- 2 files changed, 162 insertions(+), 13 deletions(-)
From: Hui Zhu <zhuhui@kylinos.cn>
Previously the patch was based on mm-unstable, which caused issues during
Sashiko apply.
So I rebased it onto mm-stable and resend.
Thanks Andrew for the heads-up.
The legacy reclaim path has two mechanisms around isolated folios that
MGLRU lacks:
1. NR_ISOLATED_ANON/FILE counters are updated when folios are isolated
from the inactive lists. Compaction's too_many_isolated() relies on
them to decide when to back off. The MGLRU reclaim path never
updates them, so compaction cannot see MGLRU's in-flight isolation.
2. shrink_inactive_list() throttles direct reclaim via
too_many_isolated() when isolated folios pile up. MGLRU's
evict_folios() isolates folios without any such check, so many
concurrent reclaimers can over-isolate the same (oldest) generation,
leading to unnecessary swapping, thrashing and premature memcg OOM.
Patch 1 fixes the counter accounting in the MGLRU isolation path.
Patch 2 adds a per-lruvec throttle, mirroring the legacy behavior but
adapted to MGLRU's per-lruvec contention and its dynamic type
selection/fallback.
Testing
=======
Two test scripts are provided to reproduce the problem and validate
the fix. Both are available at:
https://gist.github.com/teawater/3ef51251f2e91a5a600e3d26bb477e34
Test environment: 10 CPU / 8GB QEMU guest, MGLRU enabled, a 16MB
memory cgroup, anonymous working set, swap backed by dm-delay (50ms
read/write delay) to slow swap-out and lengthen the isolation window.
mglru_iso_repro.sh (64 threads, 48MB working set, 60s):
Drives concurrent direct reclaim inside the memcg and measures scan
efficiency, throttle events, in-flight isolation and throughput.
Neither kernel OOMs at this concurrency; the value of the patch shows
in reclaim quality:
unpatched patched
OOM kills 0 0
mm_vmscan_throttled 0 35645 (all
VMSCAN_THROTTLE_ISOLATED)
nr_isolated peak 0 (invisible) 230
total touches 246,499,132,369 301,456,426,692
scan efficiency 0.0261 0.0194
The patched kernel completes ~22% more work in the same 60s: the
throttle keeps concurrent reclaimers from trampling the same
generation, so less CPU is burned in reclaim. Note nr_isolated is
always 0 on the unpatched kernel - the over-isolation is invisible
there, which is exactly what patch 1 fixes.
mglru_iso_repro_v2.sh (192 threads, 48MB working set, 60s):
Raises concurrency to the point where over-isolation becomes fatal:
unpatched patched
OOM kills 1 (task killed) 0
memcg oom events 51 0
mm_vmscan_throttled 0 512292 (all
VMSCAN_THROTTLE_ISOLATED)
total touches 0 (killed) 682,300,003,972
With 192 threads the unpatched kernel cannot keep reclaim ahead of
allocation and the task is OOM-killed; the patched kernel survives the
full run and keeps reclaim making progress.
Changelog:
v5:
According to the comments of Kairun and feedback from the test scripts,
make the throttle check per lruvec (new nr_isolated counter) and skip
empty types in the allowed mask.
v4:
According to the commens of Baolin and Barry, rework patch 2:
drop the throttle_is_throttled() helper extracted
from shrink_inactive_list() and leave the legacy path untouched.
The new MGLRU-only throttle_evictable_types() only sleeps when all
evictable types are over-isolated - v3 throttled as soon as any of
them was, which unnecessarily blocked the reclaim of the other type
and passes the mask of the remaining types to isolate_folios(),
which restricts both its initial choice and its fallback, so that
isolation never lands on a throttled type. The v3 gate did not
constrain the type actually isolated, so the fallback could still
pick the over-isolated one.
Re-run the tests and update the test log.
v3:
According to the commens of Baolin, remove the redundant nr_isolated
check before restoring the NR_ISOLATED_* counters in evict_folios().
rename the extracted helper to throttle_is_throttled() to avoid
confusion with the existing wake_throttle_isolated() naming space.
Use for_each_evictable_type() in the MGLRU throttle to check each
evictable type's isolation instead of only the type returned by
get_type_to_scan(), since isolate_folios() may fall back to the
other type.
Re-run the tests and update the test log.
v2:
According to the commens of Kairun, Rebased on mm-unstable.
Split into two patches; patch 2 is new and adds the
too_many_isolated() throttling to the MGLRU eviction path, which v1
did not cover.
Add test infomations.
Hui Zhu (2):
mm/vmscan: fix missing NR_ISOLATED counter update in MGLRU reclaim
path
mm/vmscan: throttle MGLRU eviction when isolated folios pile up
include/linux/mmzone.h | 2 +
mm/vmscan.c | 173 +++++++++++++++++++++++++++++++++++++----
2 files changed, 162 insertions(+), 13 deletions(-)
--
2.43.0
On Fri, Sep 18, 2026 at 4:07 PM Hui Zhu <hui.zhu@linux.dev> wrote: > > From: Hui Zhu <zhuhui@kylinos.cn> > > Previously the patch was based on mm-unstable, which caused issues during > Sashiko apply. > So I rebased it onto mm-stable and resend. > Thanks Andrew for the heads-up. > > The legacy reclaim path has two mechanisms around isolated folios that > MGLRU lacks: > > 1. NR_ISOLATED_ANON/FILE counters are updated when folios are isolated > from the inactive lists. Compaction's too_many_isolated() relies on > them to decide when to back off. The MGLRU reclaim path never > updates them, so compaction cannot see MGLRU's in-flight isolation. > > 2. shrink_inactive_list() throttles direct reclaim via > too_many_isolated() when isolated folios pile up. MGLRU's > evict_folios() isolates folios without any such check, so many > concurrent reclaimers can over-isolate the same (oldest) generation, > leading to unnecessary swapping, thrashing and premature memcg OOM. > > Patch 1 fixes the counter accounting in the MGLRU isolation path. > Patch 2 adds a per-lruvec throttle, mirroring the legacy behavior but > adapted to MGLRU's per-lruvec contention and its dynamic type > selection/fallback. > > Testing > ======= > Two test scripts are provided to reproduce the problem and validate > the fix. Both are available at: > https://gist.github.com/teawater/3ef51251f2e91a5a600e3d26bb477e34 > Test environment: 10 CPU / 8GB QEMU guest, MGLRU enabled, a 16MB > memory cgroup, anonymous working set, swap backed by dm-delay (50ms > read/write delay) to slow swap-out and lengthen the isolation window. > > mglru_iso_repro.sh (64 threads, 48MB working set, 60s): > Drives concurrent direct reclaim inside the memcg and measures scan > efficiency, throttle events, in-flight isolation and throughput. > Neither kernel OOMs at this concurrency; the value of the patch shows > in reclaim quality: > unpatched patched > OOM kills 0 0 > mm_vmscan_throttled 0 35645 (all > VMSCAN_THROTTLE_ISOLATED) > nr_isolated peak 0 (invisible) 230 > total touches 246,499,132,369 301,456,426,692 > scan efficiency 0.0261 0.0194 > > The patched kernel completes ~22% more work in the same 60s: the > throttle keeps concurrent reclaimers from trampling the same > generation, so less CPU is burned in reclaim. Note nr_isolated is > always 0 on the unpatched kernel - the over-isolation is invisible > there, which is exactly what patch 1 fixes. > > mglru_iso_repro_v2.sh (192 threads, 48MB working set, 60s): > Raises concurrency to the point where over-isolation becomes fatal: > unpatched patched > OOM kills 1 (task killed) 0 > memcg oom events 51 0 > mm_vmscan_throttled 0 512292 (all > VMSCAN_THROTTLE_ISOLATED) > total touches 0 (killed) 682,300,003,972 > > With 192 threads the unpatched kernel cannot keep reclaim ahead of > allocation and the task is OOM-killed; the patched kernel survives the > full run and keeps reclaim making progress. > Hi Hui, I’m perfectly fine with using a script to create a scenario with a small memcg (e.g. 16 MB or 64 MB) and many threads. This is how we do R&D and build simple reproducers for complex problems. But I’m still struggling to understand how this scenario maps to or affects real-world workloads, other than `iso_stress.c` in your scripts. Do you have any idea of a real workload that suffers from the lack of throttling, which we could present in the cover letter? Best Regards Barry
© 2016 - 2026 Red Hat, Inc.