[PATCH 0/3] mm/mglru: clean up isolate_folios and scan_folios for readability and clarity

Barry Song (Xiaomi) posted 3 patches 1 month, 1 week ago
mm/vmscan.c | 78 ++++++++++++++++++++++++++++++++++++-----------------
1 file changed, 54 insertions(+), 24 deletions(-)
[PATCH 0/3] mm/mglru: clean up isolate_folios and scan_folios for readability and clarity
Posted by Barry Song (Xiaomi) 1 month, 1 week ago
This is a cleanup series split out from the MGLRU swappiness series [1],
with the cleanup changes separated to make them easier to review.

Right now, isolate_folios() is quite difficult to follow:

1. It uses for_each_evictable_type(i, swappiness) to iterate over the
types, but i is not actually used as the type within the loop body.

2. It uses scanned == 0 to detect whether the current reclaim type is
exhausted, but this is not an accurate indication.

3. It has an internal retry when no folios can be isolated after scanning
some folios, but the retry is implemented in a way nobody can understand.

This patchset makes these behaviors explicit and much easier to follow.

Run kernel builds for several rounds in a 1 GB memcg and take the
average build time. The patchset shows almost no performance impact,
with a very small improvement that could simply be noise:

                w/o patch   w/patch      Delta

real             109.149    108.990     -0.15%
sys              223.213    222.592     -0.28%

pgscan_file      498,208    486,994     -2.25%
refault_file     262,696    254,688     -3.05%

Thanks very much to Baolin and Lian for their previous reviews of the
original RFC patchset for this part.

[1] https://lore.kernel.org/linux-mm/20260812121658.69965-1-baohua@kernel.org/

Barry Song (Xiaomi) (2):
  mm/mglru: improve scan_folios() exhaustion detection
  mm/mglru: retry the same type once if isolation fails due to races

Ridong Chen (1):
  mm/mglru: improve readability of isolate_folios()

 mm/vmscan.c | 78 ++++++++++++++++++++++++++++++++++++-----------------
 1 file changed, 54 insertions(+), 24 deletions(-)

-- 
2.34.1
Re: [PATCH 0/3] mm/mglru: clean up isolate_folios and scan_folios for readability and clarity
Posted by Baolin Wang 1 month ago

On 8/20/26 12:56 PM, Barry Song (Xiaomi) wrote:
> This is a cleanup series split out from the MGLRU swappiness series [1],
> with the cleanup changes separated to make them easier to review.
> 
> Right now, isolate_folios() is quite difficult to follow:
> 
> 1. It uses for_each_evictable_type(i, swappiness) to iterate over the
> types, but i is not actually used as the type within the loop body.
> 
> 2. It uses scanned == 0 to detect whether the current reclaim type is
> exhausted, but this is not an accurate indication.
> 
> 3. It has an internal retry when no folios can be isolated after scanning
> some folios, but the retry is implemented in a way nobody can understand.
> 
> This patchset makes these behaviors explicit and much easier to follow.
> 
> Run kernel builds for several rounds in a 1 GB memcg and take the
> average build time. The patchset shows almost no performance impact,
> with a very small improvement that could simply be noise:

Just FYI:

I tested this patchset with a 3G memcg limit and a 10G zram device, 
running 'make -j32' to build kernel on my 32-core Arm machines, and got 
some performance improvement for ths sys time:
	w/o patch	w/patch
sys	1845s		1570s

>                  w/o patch   w/patch      Delta
> 
> real             109.149    108.990     -0.15%
> sys              223.213    222.592     -0.28%
> 
> pgscan_file      498,208    486,994     -2.25%
> refault_file     262,696    254,688     -3.05%
> 
> Thanks very much to Baolin and Lian for their previous reviews of the
> original RFC patchset for this part.
> 
> [1] https://lore.kernel.org/linux-mm/20260812121658.69965-1-baohua@kernel.org/
> 
> Barry Song (Xiaomi) (2):
>    mm/mglru: improve scan_folios() exhaustion detection
>    mm/mglru: retry the same type once if isolation fails due to races
> 
> Ridong Chen (1):
>    mm/mglru: improve readability of isolate_folios()
> 
>   mm/vmscan.c | 78 ++++++++++++++++++++++++++++++++++++-----------------
>   1 file changed, 54 insertions(+), 24 deletions(-)
>
Re: [PATCH 0/3] mm/mglru: clean up isolate_folios and scan_folios for readability and clarity
Posted by Kairui Song 1 month ago
On Mon, Aug 24, 2026 at 3:25 PM Baolin Wang
<baolin.wang@linux.alibaba.com> wrote:
> On 8/20/26 12:56 PM, Barry Song (Xiaomi) wrote:
> > This is a cleanup series split out from the MGLRU swappiness series [1],
> > with the cleanup changes separated to make them easier to review.
> >
> > Right now, isolate_folios() is quite difficult to follow:
> >
> > 1. It uses for_each_evictable_type(i, swappiness) to iterate over the
> > types, but i is not actually used as the type within the loop body.
> >
> > 2. It uses scanned == 0 to detect whether the current reclaim type is
> > exhausted, but this is not an accurate indication.
> >
> > 3. It has an internal retry when no folios can be isolated after scanning
> > some folios, but the retry is implemented in a way nobody can understand.
> >
> > This patchset makes these behaviors explicit and much easier to follow.
> >
> > Run kernel builds for several rounds in a 1 GB memcg and take the
> > average build time. The patchset shows almost no performance impact,
> > with a very small improvement that could simply be noise:
>
> Just FYI:
>
> I tested this patchset with a 3G memcg limit and a 10G zram device,
> running 'make -j32' to build kernel on my 32-core Arm machines, and got
> some performance improvement for ths sys time:
>         w/o patch       w/patch
> sys     1845s           1570s
>

Hi Baoliln

That's a very interesting result, can you share a bit more info about
it? e.g. vmstat? I'm curious how this happens.

I suspect patch 2 or 3 changes the swappiness / reclaim / aging type
selection behavior, or maybe it reduced the reclaim amount?
Re: [PATCH 0/3] mm/mglru: clean up isolate_folios and scan_folios for readability and clarity
Posted by Baolin Wang 1 month ago

On 8/24/26 7:05 PM, Kairui Song wrote:
> On Mon, Aug 24, 2026 at 3:25 PM Baolin Wang
> <baolin.wang@linux.alibaba.com> wrote:
>> On 8/20/26 12:56 PM, Barry Song (Xiaomi) wrote:
>>> This is a cleanup series split out from the MGLRU swappiness series [1],
>>> with the cleanup changes separated to make them easier to review.
>>>
>>> Right now, isolate_folios() is quite difficult to follow:
>>>
>>> 1. It uses for_each_evictable_type(i, swappiness) to iterate over the
>>> types, but i is not actually used as the type within the loop body.
>>>
>>> 2. It uses scanned == 0 to detect whether the current reclaim type is
>>> exhausted, but this is not an accurate indication.
>>>
>>> 3. It has an internal retry when no folios can be isolated after scanning
>>> some folios, but the retry is implemented in a way nobody can understand.
>>>
>>> This patchset makes these behaviors explicit and much easier to follow.
>>>
>>> Run kernel builds for several rounds in a 1 GB memcg and take the
>>> average build time. The patchset shows almost no performance impact,
>>> with a very small improvement that could simply be noise:
>>
>> Just FYI:
>>
>> I tested this patchset with a 3G memcg limit and a 10G zram device,
>> running 'make -j32' to build kernel on my 32-core Arm machines, and got
>> some performance improvement for ths sys time:
>>          w/o patch       w/patch
>> sys     1845s           1570s
>>
> 
> Hi Baoliln
> 
> That's a very interesting result, can you share a bit more info about
> it? e.g. vmstat? I'm curious how this happens.

Sure.

> I suspect patch 2 or 3 changes the swappiness / reclaim / aging type
> selection behavior, or maybe it reduced the reclaim amount?

I gathered the memcg stats as shown below.

It looks like the direct reason for the performance improvement is an 
obvious reduction in anon refaults, which is what this patch aims to 
achieve I think. That is to say, we should respect the type chosen by 
the PID for reclaim, and try to exhaust that type before falling back to 
another.

Before this series:
workingset_refault_anon 59706846
workingset_refault_file 9229283
workingset_activate_anon 15108561
workingset_activate_file 365660
workingset_restore_anon 15108561
workingset_restore_file 1018493
workingset_nodereclaim 0
pgsteal_kswapd 0
pgsteal_direct 92589340
pgsteal_khugepaged 549
pgsteal_proactive 0
pgscan_kswapd 0
pgscan_direct 456629355
pgscan_khugepaged 549

After this series:
workingset_refault_anon 41277086
workingset_refault_file 9131522
workingset_activate_anon 12924656
workingset_activate_file 345901
workingset_restore_anon 12924656
workingset_restore_file 1005959
workingset_nodereclaim 0
pgsteal_kswapd 0
pgsteal_direct 71723800
pgsteal_khugepaged 4013
pgsteal_proactive 0
pgscan_kswapd 0
pgscan_direct 342179445
pgscan_khugepaged 9311
Re: [PATCH 0/3] mm/mglru: clean up isolate_folios and scan_folios for readability and clarity
Posted by Kairui Song 1 month ago
On Tue, Aug 25, 2026 at 5:06 PM Baolin Wang
<baolin.wang@linux.alibaba.com> wrote:
> On 8/24/26 7:05 PM, Kairui Song wrote:
> > On Mon, Aug 24, 2026 at 3:25 PM Baolin Wang
> > <baolin.wang@linux.alibaba.com> wrote:
> >> On 8/20/26 12:56 PM, Barry Song (Xiaomi) wrote:
> >>> This is a cleanup series split out from the MGLRU swappiness series [1],
> >>> with the cleanup changes separated to make them easier to review.
> >>>
> >>> Right now, isolate_folios() is quite difficult to follow:
> >>>
> >>> 1. It uses for_each_evictable_type(i, swappiness) to iterate over the
> >>> types, but i is not actually used as the type within the loop body.
> >>>
> >>> 2. It uses scanned == 0 to detect whether the current reclaim type is
> >>> exhausted, but this is not an accurate indication.
> >>>
> >>> 3. It has an internal retry when no folios can be isolated after scanning
> >>> some folios, but the retry is implemented in a way nobody can understand.
> >>>
> >>> This patchset makes these behaviors explicit and much easier to follow.
> >>>
> >>> Run kernel builds for several rounds in a 1 GB memcg and take the
> >>> average build time. The patchset shows almost no performance impact,
> >>> with a very small improvement that could simply be noise:
> >>
> >> Just FYI:
> >>
> >> I tested this patchset with a 3G memcg limit and a 10G zram device,
> >> running 'make -j32' to build kernel on my 32-core Arm machines, and got
> >> some performance improvement for ths sys time:
> >>          w/o patch       w/patch
> >> sys     1845s           1570s
> >>
> >
> > Hi Baoliln
> >
> > That's a very interesting result, can you share a bit more info about
> > it? e.g. vmstat? I'm curious how this happens.
>
> Sure.
>
> > I suspect patch 2 or 3 changes the swappiness / reclaim / aging type
> > selection behavior, or maybe it reduced the reclaim amount?
>
> I gathered the memcg stats as shown below.
>
> It looks like the direct reason for the performance improvement is an
> obvious reduction in anon refaults, which is what this patch aims to
> achieve I think. That is to say, we should respect the type chosen by
> the PID for reclaim, and try to exhaust that type before falling back to
> another.
>
> Before this series:
> workingset_refault_anon 59706846
> workingset_refault_file 9229283
> workingset_activate_anon 15108561
> workingset_activate_file 365660
> workingset_restore_anon 15108561
> workingset_restore_file 1018493
> workingset_nodereclaim 0
> pgsteal_kswapd 0
> pgsteal_direct 92589340
> pgsteal_khugepaged 549
> pgsteal_proactive 0
> pgscan_kswapd 0
> pgscan_direct 456629355
> pgscan_khugepaged 549
>
> After this series:
> workingset_refault_anon 41277086
> workingset_refault_file 9131522
> workingset_activate_anon 12924656
> workingset_activate_file 345901
> workingset_restore_anon 12924656
> workingset_restore_file 1005959
> workingset_nodereclaim 0
> pgsteal_kswapd 0
> pgsteal_direct 71723800
> pgsteal_khugepaged 4013
> pgsteal_proactive 0
> pgscan_kswapd 0
> pgscan_direct 342179445
> pgscan_khugepaged 9311

Thanks for the info!

I think it might matches what I had in mind: we currently have a very
subtle behavior for MGLRU: in try_to_inc_min_seq, it will refuse to
increase the gen min_seq beyound (max_seq - MIN_NR_GENS) even if the
last gen is empty, if another type still have MAX_NR_GENS left. As a
result, if one type is reclaimed more, and another type have many
folios stuck in the olest gen, the over reclaimed type will have a
very tiny oldest generation (all new folios land on the oldest
generation and are reclaimed immediately), that olest gen will stay at
a near zero size and gets exhausted very frequently, with min_seq not
increased.

This will happen to file type very frequently if swappiness is low, or
happen very frequently to anon if files are frequently reclaimed,
really depends on the workload.

The old "if (!scanned && type_fallback_allowed)" will return false and
raise the priority very frequently. Raising the priority means the
next reclaim cycle with run with larger reclaim ratio, resulting in
over-reclaim on both folios and slab. What's more, it might trigger
aging.

With this series, it will see that as exhausted, and fall back to
other type, which in fact respects the swappiness / PID less, but also
avoid the raised priority and over reclaim in some cases. I did a test
on my machine and it seems matches that model:

Before:
*** Executing swappiness 1 ***
sys: 665.941
refault_file: 945242
refault_anon: 916129
pgscan_anon: 13195435
pgscan_file: 2419874
*** Executing swappiness 60 ***
sys: 650.801
refault_file: 238839
refault_anon: 1026646
pgscan_anon: 13082867
pgscan_file: 809642
*** Executing swappiness 100 ***
sys: 621.135
refault_file: 211113
refault_anon: 931943
pgscan_anon: 12008517
pgscan_file: 778586
*** Executing swappiness 150 ***
sys: 633.591
refault_file: 198239
refault_anon: 987093
pgscan_anon: 12645094
pgscan_file: 709927
*** Executing swappiness 200 ***
sys: 627.747
refault_file: 163649
refault_anon: 1024657
pgscan_anon: 12088189
pgscan_file: 653841

After:
*** Executing swappiness 1 ***
sys: 729.234
refault_file: 962067
refault_anon: 1175614
pgscan_anon: 14588298
pgscan_file: 2396037
*** Executing swappiness 60 ***
sys: 677.685
refault_file: 238877
refault_anon: 1008627
pgscan_anon: 13744350
pgscan_file: 888445
*** Executing swappiness 100 ***
sys: 668.947
refault_file: 250359
refault_anon: 994929
pgscan_anon: 13307019
pgscan_file: 794981
*** Executing swappiness 150 ***
sys: 649.764
refault_file: 180899
refault_anon: 971073
pgscan_anon: 12238885
pgscan_file: 709684
*** Executing swappiness 200 ***
sys: 648.377
refault_file: 228112
refault_anon: 1237758
pgscan_anon: 12653346
pgscan_file: 734499

So I think in theory we might see more ineffective swappiness after
this series, I'm also a bit concerned about the "lrugen->min_seq[type]
+ MIN_NR_GENS == lrugen->max_seq;" check which is strictly related to
the limitation of try_to_inc_min_seq I mentioned here... In the long
term I think we better get rid of that limitation, not sure how the
behavior will change or fit after this.

BTW the above tests were done with ordinary block SWAP, with ZRAM, I
see an improvement instead (it's related to how SYNC SWAP discards the
folio immediately and blocks the thread while ordinary SWAP does IO
asynchronously so there would be less concurrent reclaim, and changes
the dirty/writeback mix that evict_folios()'s
 loop sees, especially nr_reclaimed, and here we will avoid raise of priority):

Before:
*** Executing swappiness 1 ***
sys: 599.135
refault_file: 925843
refault_anon: 1838025
pgscan_anon: 20390687
pgscan_file: 2423955
*** Executing swappiness 60 ***
sys: 591.367
refault_file: 336184
refault_anon: 2350551
pgscan_anon: 21099503
pgscan_file: 1109003
*** Executing swappiness 100 ***
sys: 593.672
refault_file: 328276
refault_anon: 2332940
pgscan_anon: 20894727
pgscan_file: 954910
*** Executing swappiness 150 ***
sys: 585.939
refault_file: 296374
refault_anon: 2063673
pgscan_anon: 20537898
pgscan_file: 935107
*** Executing swappiness 200 ***
sys: 574.283
refault_file: 313356
refault_anon: 2065897
pgscan_anon: 20145043
pgscan_file: 881381

After:
*** Executing swappiness 1 ***
sys: 578.699
refault_file: 844457
refault_anon: 1793171
pgscan_anon: 19307170
pgscan_file: 2364248
*** Executing swappiness 60 ***
sys: 582.053
refault_file: 390580
refault_anon: 2066277
pgscan_anon: 20413030
pgscan_file: 1201373
*** Executing swappiness 100 ***
sys: 567.520
refault_file: 295709
refault_anon: 1923291
pgscan_anon: 19299837
pgscan_file: 978007
*** Executing swappiness 150 ***
sys: 582.983
refault_file: 311833
refault_anon: 2041126
pgscan_anon: 20547029
pgscan_file: 940061
*** Executing swappiness 200 ***
sys: 562.710
refault_file: 301914
refault_anon: 2005846
pgscan_anon: 19257095
pgscan_file: 839522

So in summary I think the gain is avoiding priority increase (by
falling back to another type and satisfying the reclaim target in one
iteration, when one type have a tailing draning gen); the loss is less
effective swappiness. In fact, I think we might fall back to another
type MORE, especially when under pressure, not less, contrary to what
the patch message suggests. Not really against this change, just a
headup that there is a impact on swappiness / aging of this change in
an unexpected way, and things may gets more interesting if we want to
lift that limitation in try_to_inc_min_seq as this is directly related
to that.

Right now I'm checkng if we can archive perfect swappiness as the IO
cost factor for MGLRU, and avoid OOM when under high pressure, so at
least the blocker "get_nr_gens(lruvec, type) == MIN_NR_GENS" for
reclaim will have to gone I think (we are already seeing tons of OOM
causes by that blocker).
Re: [PATCH 0/3] mm/mglru: clean up isolate_folios and scan_folios for readability and clarity
Posted by Barry Song 1 month ago
On Wed, Aug 26, 2026 at 2:10 AM Kairui Song <ryncsn@gmail.com> wrote:
>
> On Tue, Aug 25, 2026 at 5:06 PM Baolin Wang
> <baolin.wang@linux.alibaba.com> wrote:
> > On 8/24/26 7:05 PM, Kairui Song wrote:
> > > On Mon, Aug 24, 2026 at 3:25 PM Baolin Wang
> > > <baolin.wang@linux.alibaba.com> wrote:
> > >> On 8/20/26 12:56 PM, Barry Song (Xiaomi) wrote:
> > >>> This is a cleanup series split out from the MGLRU swappiness series [1],
> > >>> with the cleanup changes separated to make them easier to review.
> > >>>
> > >>> Right now, isolate_folios() is quite difficult to follow:
> > >>>
> > >>> 1. It uses for_each_evictable_type(i, swappiness) to iterate over the
> > >>> types, but i is not actually used as the type within the loop body.
> > >>>
> > >>> 2. It uses scanned == 0 to detect whether the current reclaim type is
> > >>> exhausted, but this is not an accurate indication.
> > >>>
> > >>> 3. It has an internal retry when no folios can be isolated after scanning
> > >>> some folios, but the retry is implemented in a way nobody can understand.
> > >>>
> > >>> This patchset makes these behaviors explicit and much easier to follow.
> > >>>
> > >>> Run kernel builds for several rounds in a 1 GB memcg and take the
> > >>> average build time. The patchset shows almost no performance impact,
> > >>> with a very small improvement that could simply be noise:
> > >>
> > >> Just FYI:
> > >>
> > >> I tested this patchset with a 3G memcg limit and a 10G zram device,
> > >> running 'make -j32' to build kernel on my 32-core Arm machines, and got
> > >> some performance improvement for ths sys time:
> > >>          w/o patch       w/patch
> > >> sys     1845s           1570s
> > >>
> > >
> > > Hi Baoliln
> > >
> > > That's a very interesting result, can you share a bit more info about
> > > it? e.g. vmstat? I'm curious how this happens.
> >
> > Sure.
> >
> > > I suspect patch 2 or 3 changes the swappiness / reclaim / aging type
> > > selection behavior, or maybe it reduced the reclaim amount?
> >
> > I gathered the memcg stats as shown below.
> >
> > It looks like the direct reason for the performance improvement is an
> > obvious reduction in anon refaults, which is what this patch aims to
> > achieve I think. That is to say, we should respect the type chosen by
> > the PID for reclaim, and try to exhaust that type before falling back to
> > another.
> >
> > Before this series:
> > workingset_refault_anon 59706846
> > workingset_refault_file 9229283
> > workingset_activate_anon 15108561
> > workingset_activate_file 365660
> > workingset_restore_anon 15108561
> > workingset_restore_file 1018493
> > workingset_nodereclaim 0
> > pgsteal_kswapd 0
> > pgsteal_direct 92589340
> > pgsteal_khugepaged 549
> > pgsteal_proactive 0
> > pgscan_kswapd 0
> > pgscan_direct 456629355
> > pgscan_khugepaged 549
> >
> > After this series:
> > workingset_refault_anon 41277086
> > workingset_refault_file 9131522
> > workingset_activate_anon 12924656
> > workingset_activate_file 345901
> > workingset_restore_anon 12924656
> > workingset_restore_file 1005959
> > workingset_nodereclaim 0
> > pgsteal_kswapd 0
> > pgsteal_direct 71723800
> > pgsteal_khugepaged 4013
> > pgsteal_proactive 0
> > pgscan_kswapd 0
> > pgscan_direct 342179445
> > pgscan_khugepaged 9311
>
> Thanks for the info!
>
> I think it might matches what I had in mind: we currently have a very
> subtle behavior for MGLRU: in try_to_inc_min_seq, it will refuse to
> increase the gen min_seq beyound (max_seq - MIN_NR_GENS) even if the
> last gen is empty, if another type still have MAX_NR_GENS left. As a
> result, if one type is reclaimed more, and another type have many
> folios stuck in the olest gen, the over reclaimed type will have a
> very tiny oldest generation (all new folios land on the oldest
> generation and are reclaimed immediately), that olest gen will stay at
> a near zero size and gets exhausted very frequently, with min_seq not
> increased.
>
> This will happen to file type very frequently if swappiness is low, or
> happen very frequently to anon if files are frequently reclaimed,
> really depends on the workload.
>
> The old "if (!scanned && type_fallback_allowed)" will return false and
> raise the priority very frequently. Raising the priority means the
> next reclaim cycle with run with larger reclaim ratio, resulting in
> over-reclaim on both folios and slab. What's more, it might trigger
> aging.
>
> With this series, it will see that as exhausted, and fall back to
> other type, which in fact respects the swappiness / PID less, but also

We could have the following cases:

 1. If the reclaimed type has 3 generations left (with its oldest generation
   exhausted), while !type still has 4 generations left, falling back might
   be the best behavior. There is no way to avoid the fallback by increasing
   `sc->priority`, because `should_run_aging()` will not trigger aging while
   !type still has 4 generations.

2. If the reclaimed type has 3 generations left (and its oldest generation is
   exhausted), while !type also has 3 generations left, falling back might
   violate swappiness. If we return and increase `sc->priority`,
   `should_run_aging()` may return true and trigger aging instead:

return evictable_min_seq(min_seq, swappiness) + MIN_NR_GENS == max_seq;

So maybe we could slightly adjust the code as shown below, although I’m
not sure if the added complexity is worthwhile:

diff --git a/mm/vmscan.c b/mm/vmscan.c
index 1f302386d8ab..181285463bbe 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -4878,9 +4878,10 @@ static int isolate_folios(unsigned long
nr_to_scan, struct lruvec *lruvec,

        /*
         * We are running out of the current reclaim type. Fall back to
-        * the other type if allowed.
+        * the other type if it has at least two generations to reclaim.
         */
-       if (exhausted && type_fallback_allowed) {
+       if (exhausted && type_fallback_allowed &&
+           get_nr_gens(lruvec, !type) > MIN_NR_GENS + 1) {
                type = !type;
                type_fallback_allowed = false;
                goto retry;

> avoid the raised priority and over reclaim in some cases. I did a test
> on my machine and it seems matches that model:
>
> Before:
> *** Executing swappiness 1 ***
> sys: 665.941
> refault_file: 945242
> refault_anon: 916129
> pgscan_anon: 13195435
> pgscan_file: 2419874
> *** Executing swappiness 60 ***
> sys: 650.801
> refault_file: 238839
> refault_anon: 1026646
> pgscan_anon: 13082867
> pgscan_file: 809642
> *** Executing swappiness 100 ***
> sys: 621.135
> refault_file: 211113
> refault_anon: 931943
> pgscan_anon: 12008517
> pgscan_file: 778586
> *** Executing swappiness 150 ***
> sys: 633.591
> refault_file: 198239
> refault_anon: 987093
> pgscan_anon: 12645094
> pgscan_file: 709927
> *** Executing swappiness 200 ***
> sys: 627.747
> refault_file: 163649
> refault_anon: 1024657
> pgscan_anon: 12088189
> pgscan_file: 653841
>
> After:
> *** Executing swappiness 1 ***
> sys: 729.234
> refault_file: 962067
> refault_anon: 1175614
> pgscan_anon: 14588298
> pgscan_file: 2396037
> *** Executing swappiness 60 ***
> sys: 677.685
> refault_file: 238877
> refault_anon: 1008627
> pgscan_anon: 13744350
> pgscan_file: 888445
> *** Executing swappiness 100 ***
> sys: 668.947
> refault_file: 250359
> refault_anon: 994929
> pgscan_anon: 13307019
> pgscan_file: 794981
> *** Executing swappiness 150 ***
> sys: 649.764
> refault_file: 180899
> refault_anon: 971073
> pgscan_anon: 12238885
> pgscan_file: 709684
> *** Executing swappiness 200 ***
> sys: 648.377
> refault_file: 228112
> refault_anon: 1237758
> pgscan_anon: 12653346
> pgscan_file: 734499
>
> So I think in theory we might see more ineffective swappiness after
> this series, I'm also a bit concerned about the "lrugen->min_seq[type]
> + MIN_NR_GENS == lrugen->max_seq;" check which is strictly related to
> the limitation of try_to_inc_min_seq I mentioned here... In the long
> term I think we better get rid of that limitation, not sure how the
> behavior will change or fit after this.

In the earlier RFC, I got rid of the generation sync in
`try_to_inc_min_seq()`:

https://lore.kernel.org/linux-mm/20260726122123.7614-4-baohua@kernel.org/

It does help with swappiness, but also slightly increases sys time at
swappiness values around 60, so I dropped it from RFC v4.

Basically, letting file and anon catch up with each other seems to offer
the best performance at normal swappiness values such as 60–100, but it
essentially makes MGLRU's swappiness a toy.

>
> BTW the above tests were done with ordinary block SWAP, with ZRAM, I
> see an improvement instead (it's related to how SYNC SWAP discards the
> folio immediately and blocks the thread while ordinary SWAP does IO
> asynchronously so there would be less concurrent reclaim, and changes
> the dirty/writeback mix that evict_folios()'s
>  loop sees, especially nr_reclaimed, and here we will avoid raise of priority):
>
> Before:
> *** Executing swappiness 1 ***
> sys: 599.135
> refault_file: 925843
> refault_anon: 1838025
> pgscan_anon: 20390687
> pgscan_file: 2423955
> *** Executing swappiness 60 ***
> sys: 591.367
> refault_file: 336184
> refault_anon: 2350551
> pgscan_anon: 21099503
> pgscan_file: 1109003
> *** Executing swappiness 100 ***
> sys: 593.672
> refault_file: 328276
> refault_anon: 2332940
> pgscan_anon: 20894727
> pgscan_file: 954910
> *** Executing swappiness 150 ***
> sys: 585.939
> refault_file: 296374
> refault_anon: 2063673
> pgscan_anon: 20537898
> pgscan_file: 935107
> *** Executing swappiness 200 ***
> sys: 574.283
> refault_file: 313356
> refault_anon: 2065897
> pgscan_anon: 20145043
> pgscan_file: 881381
>
> After:
> *** Executing swappiness 1 ***
> sys: 578.699
> refault_file: 844457
> refault_anon: 1793171
> pgscan_anon: 19307170
> pgscan_file: 2364248
> *** Executing swappiness 60 ***
> sys: 582.053
> refault_file: 390580
> refault_anon: 2066277
> pgscan_anon: 20413030
> pgscan_file: 1201373
> *** Executing swappiness 100 ***
> sys: 567.520
> refault_file: 295709
> refault_anon: 1923291
> pgscan_anon: 19299837
> pgscan_file: 978007
> *** Executing swappiness 150 ***
> sys: 582.983
> refault_file: 311833
> refault_anon: 2041126
> pgscan_anon: 20547029
> pgscan_file: 940061
> *** Executing swappiness 200 ***
> sys: 562.710
> refault_file: 301914
> refault_anon: 2005846
> pgscan_anon: 19257095
> pgscan_file: 839522
>
> So in summary I think the gain is avoiding priority increase (by
> falling back to another type and satisfying the reclaim target in one
> iteration, when one type have a tailing draning gen); the loss is less
> effective swappiness. In fact, I think we might fall back to another
> type MORE, especially when under pressure, not less, contrary to what
> the patch message suggests. Not really against this change, just a
> headup that there is a impact on swappiness / aging of this change in
> an unexpected way, and things may gets more interesting if we want to
> lift that limitation in try_to_inc_min_seq as this is directly related
> to that.

As mentioned above, I’m not sure if we want to increase the code complexity
by distinguishing between the following two cases:

1. The other type has 2+ generations left to reclaim. In this case, even
   after `sc->priority` increases, we still have to fall back because
   `should_run_aging()` won’t trigger aging.

2. The other type has only 1 generation left to reclaim. In this case,
   `should_run_aging()` will trigger aging, so the previously reclaimed type
   might get a new reclaimable generation in the next iteration.

>
> Right now I'm checkng if we can archive perfect swappiness as the IO
> cost factor for MGLRU, and avoid OOM when under high pressure, so at
> least the blocker "get_nr_gens(lruvec, type) == MIN_NR_GENS" for
> reclaim will have to gone I think (we are already seeing tons of OOM
> causes by that blocker).

I tried a couple of ways to remove `get_nr_gens(lruvec, type) == MIN_NR_GENS`
from `scan_folios()`, but all of them increased sys time at normal
swappiness values. If we really want to remove it, we need to figure out
how to do so without impacting performance :-)

Best Regards
Barry
Re: [PATCH 0/3] mm/mglru: clean up isolate_folios and scan_folios for readability and clarity
Posted by Baolin Wang 1 month ago

On 8/26/26 5:11 AM, Barry Song wrote:
> On Wed, Aug 26, 2026 at 2:10 AM Kairui Song <ryncsn@gmail.com> wrote:
>>
>> On Tue, Aug 25, 2026 at 5:06 PM Baolin Wang
>> <baolin.wang@linux.alibaba.com> wrote:
>>> On 8/24/26 7:05 PM, Kairui Song wrote:
>>>> On Mon, Aug 24, 2026 at 3:25 PM Baolin Wang
>>>> <baolin.wang@linux.alibaba.com> wrote:
>>>>> On 8/20/26 12:56 PM, Barry Song (Xiaomi) wrote:
>>>>>> This is a cleanup series split out from the MGLRU swappiness series [1],
>>>>>> with the cleanup changes separated to make them easier to review.
>>>>>>
>>>>>> Right now, isolate_folios() is quite difficult to follow:
>>>>>>
>>>>>> 1. It uses for_each_evictable_type(i, swappiness) to iterate over the
>>>>>> types, but i is not actually used as the type within the loop body.
>>>>>>
>>>>>> 2. It uses scanned == 0 to detect whether the current reclaim type is
>>>>>> exhausted, but this is not an accurate indication.
>>>>>>
>>>>>> 3. It has an internal retry when no folios can be isolated after scanning
>>>>>> some folios, but the retry is implemented in a way nobody can understand.
>>>>>>
>>>>>> This patchset makes these behaviors explicit and much easier to follow.
>>>>>>
>>>>>> Run kernel builds for several rounds in a 1 GB memcg and take the
>>>>>> average build time. The patchset shows almost no performance impact,
>>>>>> with a very small improvement that could simply be noise:
>>>>>
>>>>> Just FYI:
>>>>>
>>>>> I tested this patchset with a 3G memcg limit and a 10G zram device,
>>>>> running 'make -j32' to build kernel on my 32-core Arm machines, and got
>>>>> some performance improvement for ths sys time:
>>>>>           w/o patch       w/patch
>>>>> sys     1845s           1570s
>>>>>
>>>>
>>>> Hi Baoliln
>>>>
>>>> That's a very interesting result, can you share a bit more info about
>>>> it? e.g. vmstat? I'm curious how this happens.
>>>
>>> Sure.
>>>
>>>> I suspect patch 2 or 3 changes the swappiness / reclaim / aging type
>>>> selection behavior, or maybe it reduced the reclaim amount?
>>>
>>> I gathered the memcg stats as shown below.
>>>
>>> It looks like the direct reason for the performance improvement is an
>>> obvious reduction in anon refaults, which is what this patch aims to
>>> achieve I think. That is to say, we should respect the type chosen by
>>> the PID for reclaim, and try to exhaust that type before falling back to
>>> another.
>>>
>>> Before this series:
>>> workingset_refault_anon 59706846
>>> workingset_refault_file 9229283
>>> workingset_activate_anon 15108561
>>> workingset_activate_file 365660
>>> workingset_restore_anon 15108561
>>> workingset_restore_file 1018493
>>> workingset_nodereclaim 0
>>> pgsteal_kswapd 0
>>> pgsteal_direct 92589340
>>> pgsteal_khugepaged 549
>>> pgsteal_proactive 0
>>> pgscan_kswapd 0
>>> pgscan_direct 456629355
>>> pgscan_khugepaged 549
>>>
>>> After this series:
>>> workingset_refault_anon 41277086
>>> workingset_refault_file 9131522
>>> workingset_activate_anon 12924656
>>> workingset_activate_file 345901
>>> workingset_restore_anon 12924656
>>> workingset_restore_file 1005959
>>> workingset_nodereclaim 0
>>> pgsteal_kswapd 0
>>> pgsteal_direct 71723800
>>> pgsteal_khugepaged 4013
>>> pgsteal_proactive 0
>>> pgscan_kswapd 0
>>> pgscan_direct 342179445
>>> pgscan_khugepaged 9311
>>
>> Thanks for the info!
>>
>> I think it might matches what I had in mind: we currently have a very
>> subtle behavior for MGLRU: in try_to_inc_min_seq, it will refuse to
>> increase the gen min_seq beyound (max_seq - MIN_NR_GENS) even if the
>> last gen is empty, if another type still have MAX_NR_GENS left. As a
>> result, if one type is reclaimed more, and another type have many
>> folios stuck in the olest gen, the over reclaimed type will have a
>> very tiny oldest generation (all new folios land on the oldest
>> generation and are reclaimed immediately), that olest gen will stay at
>> a near zero size and gets exhausted very frequently, with min_seq not
>> increased.
>>
>> This will happen to file type very frequently if swappiness is low, or
>> happen very frequently to anon if files are frequently reclaimed,
>> really depends on the workload.
>>
>> The old "if (!scanned && type_fallback_allowed)" will return false and
>> raise the priority very frequently. Raising the priority means the
>> next reclaim cycle with run with larger reclaim ratio, resulting in
>> over-reclaim on both folios and slab. What's more, it might trigger
>> aging.
>>
>> With this series, it will see that as exhausted, and fall back to
>> other type, which in fact respects the swappiness / PID less, but also
> 
> We could have the following cases:
> 
>   1. If the reclaimed type has 3 generations left (with its oldest generation
>     exhausted), while !type still has 4 generations left, falling back might
>     be the best behavior. There is no way to avoid the fallback by increasing
>     `sc->priority`, because `should_run_aging()` will not trigger aging while
>     !type still has 4 generations.
> 
> 2. If the reclaimed type has 3 generations left (and its oldest generation is
>     exhausted), while !type also has 3 generations left, falling back might
>     violate swappiness. If we return and increase `sc->priority`,
>     `should_run_aging()` may return true and trigger aging instead:
> 
> return evictable_min_seq(min_seq, swappiness) + MIN_NR_GENS == max_seq;
> 
> So maybe we could slightly adjust the code as shown below, although I’m
> not sure if the added complexity is worthwhile:
> 
> diff --git a/mm/vmscan.c b/mm/vmscan.c
> index 1f302386d8ab..181285463bbe 100644
> --- a/mm/vmscan.c
> +++ b/mm/vmscan.c
> @@ -4878,9 +4878,10 @@ static int isolate_folios(unsigned long
> nr_to_scan, struct lruvec *lruvec,
> 
>          /*
>           * We are running out of the current reclaim type. Fall back to
> -        * the other type if allowed.
> +        * the other type if it has at least two generations to reclaim.
>           */
> -       if (exhausted && type_fallback_allowed) {
> +       if (exhausted && type_fallback_allowed &&
> +           get_nr_gens(lruvec, !type) > MIN_NR_GENS + 1) {
>                  type = !type;
>                  type_fallback_allowed = false;
>                  goto retry;
> 
>> avoid the raised priority and over reclaim in some cases. I did a test
>> on my machine and it seems matches that model:

With the above changes, I can also reproduce the slight regression with 
NVMe swap (set swappiness = 100) mentioned by Kairui (but the zram case 
still looks good):

Before:
sys: 757
refault_anon 5554715
refault_file 3867998

After:
sys: 765
refault_anon 5791349
refault_file 3939552

>> Before:
>> *** Executing swappiness 1 ***
>> sys: 665.941
>> refault_file: 945242
>> refault_anon: 916129
>> pgscan_anon: 13195435
>> pgscan_file: 2419874
>> *** Executing swappiness 60 ***
>> sys: 650.801
>> refault_file: 238839
>> refault_anon: 1026646
>> pgscan_anon: 13082867
>> pgscan_file: 809642
>> *** Executing swappiness 100 ***
>> sys: 621.135
>> refault_file: 211113
>> refault_anon: 931943
>> pgscan_anon: 12008517
>> pgscan_file: 778586
>> *** Executing swappiness 150 ***
>> sys: 633.591
>> refault_file: 198239
>> refault_anon: 987093
>> pgscan_anon: 12645094
>> pgscan_file: 709927
>> *** Executing swappiness 200 ***
>> sys: 627.747
>> refault_file: 163649
>> refault_anon: 1024657
>> pgscan_anon: 12088189
>> pgscan_file: 653841
>>
>> After:
>> *** Executing swappiness 1 ***
>> sys: 729.234
>> refault_file: 962067
>> refault_anon: 1175614
>> pgscan_anon: 14588298
>> pgscan_file: 2396037
>> *** Executing swappiness 60 ***
>> sys: 677.685
>> refault_file: 238877
>> refault_anon: 1008627
>> pgscan_anon: 13744350
>> pgscan_file: 888445
>> *** Executing swappiness 100 ***
>> sys: 668.947
>> refault_file: 250359
>> refault_anon: 994929
>> pgscan_anon: 13307019
>> pgscan_file: 794981
>> *** Executing swappiness 150 ***
>> sys: 649.764
>> refault_file: 180899
>> refault_anon: 971073
>> pgscan_anon: 12238885
>> pgscan_file: 709684
>> *** Executing swappiness 200 ***
>> sys: 648.377
>> refault_file: 228112
>> refault_anon: 1237758
>> pgscan_anon: 12653346
>> pgscan_file: 734499
>>
>> So I think in theory we might see more ineffective swappiness after
>> this series, I'm also a bit concerned about the "lrugen->min_seq[type]
>> + MIN_NR_GENS == lrugen->max_seq;" check which is strictly related to
>> the limitation of try_to_inc_min_seq I mentioned here... In the long
>> term I think we better get rid of that limitation, not sure how the
>> behavior will change or fit after this.
> 
> In the earlier RFC, I got rid of the generation sync in
> `try_to_inc_min_seq()`:
> 
> https://lore.kernel.org/linux-mm/20260726122123.7614-4-baohua@kernel.org/
> 
> It does help with swappiness, but also slightly increases sys time at
> swappiness values around 60, so I dropped it from RFC v4.
> 
> Basically, letting file and anon catch up with each other seems to offer
> the best performance at normal swappiness values such as 60–100, but it
> essentially makes MGLRU's swappiness a toy.
> 
>>
>> BTW the above tests were done with ordinary block SWAP, with ZRAM, I
>> see an improvement instead (it's related to how SYNC SWAP discards the
>> folio immediately and blocks the thread while ordinary SWAP does IO
>> asynchronously so there would be less concurrent reclaim, and changes
>> the dirty/writeback mix that evict_folios()'s
>>   loop sees, especially nr_reclaimed, and here we will avoid raise of priority):
>>
>> Before:
>> *** Executing swappiness 1 ***
>> sys: 599.135
>> refault_file: 925843
>> refault_anon: 1838025
>> pgscan_anon: 20390687
>> pgscan_file: 2423955
>> *** Executing swappiness 60 ***
>> sys: 591.367
>> refault_file: 336184
>> refault_anon: 2350551
>> pgscan_anon: 21099503
>> pgscan_file: 1109003
>> *** Executing swappiness 100 ***
>> sys: 593.672
>> refault_file: 328276
>> refault_anon: 2332940
>> pgscan_anon: 20894727
>> pgscan_file: 954910
>> *** Executing swappiness 150 ***
>> sys: 585.939
>> refault_file: 296374
>> refault_anon: 2063673
>> pgscan_anon: 20537898
>> pgscan_file: 935107
>> *** Executing swappiness 200 ***
>> sys: 574.283
>> refault_file: 313356
>> refault_anon: 2065897
>> pgscan_anon: 20145043
>> pgscan_file: 881381
>>
>> After:
>> *** Executing swappiness 1 ***
>> sys: 578.699
>> refault_file: 844457
>> refault_anon: 1793171
>> pgscan_anon: 19307170
>> pgscan_file: 2364248
>> *** Executing swappiness 60 ***
>> sys: 582.053
>> refault_file: 390580
>> refault_anon: 2066277
>> pgscan_anon: 20413030
>> pgscan_file: 1201373
>> *** Executing swappiness 100 ***
>> sys: 567.520
>> refault_file: 295709
>> refault_anon: 1923291
>> pgscan_anon: 19299837
>> pgscan_file: 978007
>> *** Executing swappiness 150 ***
>> sys: 582.983
>> refault_file: 311833
>> refault_anon: 2041126
>> pgscan_anon: 20547029
>> pgscan_file: 940061
>> *** Executing swappiness 200 ***
>> sys: 562.710
>> refault_file: 301914
>> refault_anon: 2005846
>> pgscan_anon: 19257095
>> pgscan_file: 839522
>>
>> So in summary I think the gain is avoiding priority increase (by
>> falling back to another type and satisfying the reclaim target in one
>> iteration, when one type have a tailing draning gen); the loss is less
>> effective swappiness. In fact, I think we might fall back to another
>> type MORE, especially when under pressure, not less, contrary to what
>> the patch message suggests. Not really against this change, just a
>> headup that there is a impact on swappiness / aging of this change in
>> an unexpected way, and things may gets more interesting if we want to
>> lift that limitation in try_to_inc_min_seq as this is directly related
>> to that.

Thanks Kairui for your data and anylysis.

I mostly agree. But on whether we should fallback more, I still prefer 
Barry's idea. The current patch might be a bit too aggressive, cause it 
only fallback to another type when the oldest generation is fully 
exhausted. I've always felt that falling back too easily doesn't really 
respect the PID's choice.

Of course, there's also the concern you raised about the limitation in 
try_to_inc_min_seq() when updating the min_seq. I think we can address 
that together later, so that we respect the PID's choice (i.e., try the 
reclaimed type as much as possible) without ending up with a tailing 
draining generation that blocks aging. That may not be simple, though.

Anyway, personally I'd prefer to keep the original code until more 
impact investigation is done.

Just my 2 cents.
Re: [PATCH 0/3] mm/mglru: clean up isolate_folios and scan_folios for readability and clarity
Posted by Kairui Song 1 month ago
On Thu, Aug 27, 2026 at 2:21 PM Baolin Wang
<baolin.wang@linux.alibaba.com> wrote:
>
> With the above changes, I can also reproduce the slight regression with
> NVMe swap (set swappiness = 100) mentioned by Kairui (but the zram case
> still looks good):
>
> Before:
> sys: 757
> refault_anon 5554715
> refault_file 3867998
>
> After:
> sys: 765
> refault_anon 5791349
> refault_file 3939552
>

...

>
> Thanks Kairui for your data and anylysis.
>
> I mostly agree. But on whether we should fallback more, I still prefer
> Barry's idea. The current patch might be a bit too aggressive, cause it
> only fallback to another type when the oldest generation is fully
> exhausted. I've always felt that falling back too easily doesn't really
> respect the PID's choice.
>
> Of course, there's also the concern you raised about the limitation in
> try_to_inc_min_seq() when updating the min_seq. I think we can address
> that together later, so that we respect the PID's choice (i.e., try the
> reclaimed type as much as possible) without ending up with a tailing
> draining generation that blocks aging. That may not be simple, though.

Actually in the long term I think we should get away from the PID
rather than respect it more for eviction choice :), and use a actual
calculation that respects swappiness and IO cost (we can still use
the data collected by it though).

There is another series form Barry (and the clean up here is extracted from it):
https://lore.kernel.org/linux-mm/20260812121658.69965-1-baohua@kernel.org/

To achieve reasonable swappiness following the PID and aging
protection idea, we will inevitably increase overhead, and PID is
really bad at protecting the cache anyway.

PID controllers generally correct long-term linear statistic well, but
they aren't well-suited to reacting to abrupt workload changes. While
MM workloads are full of such bursts and phase changes. What's worse
the design aggregates folios into tiers based on access, and the
limited high tier means folios staying in the same high tier can have
dramatically different access behaviors.

I think PID still makes sense as a long-term reference: it can correct
cross-generation behavior over time, but relying on it as the primary
mechanism for for single type eviction or protection is problematic.
In my view, this is also one of the major limitations that has made
MGLRU difficult to generalize across a wide range of workloads.

See: https://lore.kernel.org/linux-mm/20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com/,
by softening PID protection (still used, just with a much lower
factor) and using FG protection, zipf and real workloads gets a much
higher hit rate (much higher than 10%, and dramatically better in many
cases, the basic 10% gain is end-to-end gain diluted by the overall
cost).

So in the long term I think: For promotion, we will use FG; for
eviction (the issue being patched here and in that Barry's other
patch), we will respect swappiness & IO cost based scan budget
calculation. Currently, the IO cost is just the refault count, which
seems fine, this is similar to classical LRU, but ignores scan cost
because MGLRU's scan is special, full of lazy-promoted folios that
already offset the scan budget well. PID can stay to catch long term
characteristics, at a weaker factor.

(BTW I know that PID have a way to be more adaptive to bursts;
however, that's really hard to tune for MM, and doesn't fix the long
time cold folio stay in same high tier issue).

Maybe I'll just post a series soon for the eviction part.

>
> Anyway, personally I'd prefer to keep the original code until more
> impact investigation is done.
>
> Just my 2 cents.

No problem, I'm OK with this change for now.
Re: [PATCH 0/3] mm/mglru: clean up isolate_folios and scan_folios for readability and clarity
Posted by Barry Song 1 month ago
On Thu, Aug 27, 2026 at 3:25 PM Kairui Song <ryncsn@gmail.com> wrote:
[...]
> > Anyway, personally I'd prefer to keep the original code until more
> > impact investigation is done.
> >
> > Just my 2 cents.
>
> No problem, I'm OK with this change for now.

Thanks, Baolin and Kairui.

If the effect of the exhaustion check is unclear — it improves the
zram case but might slightly regress the SSD/NVMe case — we could
simply drop patch 2 and keep patches 1 and 3 in v2. That would make
the patchset focused entirely on readability.

Best Regards
Barry
Re: [PATCH 0/3] mm/mglru: clean up isolate_folios and scan_folios for readability and clarity
Posted by Lian Wang (ProcessMission) 1 month, 1 week ago
From: Lian Wang <lianux.mm@gmail.com>

Hi Barry,

I went through all three patches. The explicit fallback, exhaustion
reporting, and bounded retry make the control flow much clearer.

In particular, the exhaustion check preserves the early-stop cases and
avoids falling back after only the oldest generation has been drained when
another generation remains reclaimable. The retry is still bounded to one
extra attempt.

No issues from my side.

For the series:

Reviewed-by: Lian Wang <lianux.mm@gmail.com>

Thanks,
Lian

On Thu, 20 Aug 2026 12:56:00 +0800 "Barry Song (Xiaomi)" <baohua@kernel.org> wrote:

> This is a cleanup series split out from the MGLRU swappiness series [1],
> with the cleanup changes separated to make them easier to review.
>
> This patchset makes these behaviors explicit and much easier to follow.
>
> Thanks very much to Baolin and Lian for their previous reviews of the
> original RFC patchset for this part.