mm/vmscan.c | 48 +++++++++++++++++++++++++++++++++--------------- 1 file changed, 33 insertions(+), 15 deletions(-)
shrink_lruvec() drives reclaim in SWAP_CLUSTER_MAX (32) chunks, but
isolate_lru_folios() may scan far more than that per call on a single
LRU. The excess is never charged back, so shrink_lruvec() keeps
rescanning the same folios round after round. When reclaim targets a
lower zone, the same scanner also keeps walking zone-ineligible folios
that can never satisfy the allocation, inflating nr_reclaimed into a
false progress that delays the OOM.
This series fixes both:
[1/2] Charge the isolate overshoot against the scan quota so the
next round skips already-scanned folios.
[2/2] Stop scanning once too many zone-ineligible folios have been
skipped, instead of force-isolating them.
Wupeng Ma (2):
mm: vmscan: charge isolate overshoot against scan quota
mm: vmscan: stop scanning ineligible folios after max_nr_skipped
mm/vmscan.c | 48 +++++++++++++++++++++++++++++++++---------------
1 file changed, 33 insertions(+), 15 deletions(-)
--
2.43.0
On Tue, 11 Aug 2026 15:08:47 +0800 Wupeng Ma <mawupeng1@huawei.com> wrote: > shrink_lruvec() drives reclaim in SWAP_CLUSTER_MAX (32) chunks, but > isolate_lru_folios() may scan far more than that per call on a single > LRU. The excess is never charged back, so shrink_lruvec() keeps > rescanning the same folios round after round. When reclaim targets a > lower zone, the same scanner also keeps walking zone-ineligible folios > that can never satisfy the allocation, inflating nr_reclaimed into a > false progress that delays the OOM. > > This series fixes both: > > [1/2] Charge the isolate overshoot against the scan quota so the > next round skips already-scanned folios. > [2/2] Stop scanning once too many zone-ineligible folios have been > skipped, instead of force-isolating them. This all sounds fairly bad. Is there some report? Real-world test results? Laboratory test results? Something to help us understand the impact of the issues and the situations in which they occur. Thanks.
On 周三 2026-8-12 09:42, Andrew Morton wrote:
> On Tue, 11 Aug 2026 15:08:47 +0800 Wupeng Ma <mawupeng1@huawei.com> wrote:
>
>> shrink_lruvec() drives reclaim in SWAP_CLUSTER_MAX (32) chunks, but
>> isolate_lru_folios() may scan far more than that per call on a single
>> LRU. The excess is never charged back, so shrink_lruvec() keeps
>> rescanning the same folios round after round. When reclaim targets a
>> lower zone, the same scanner also keeps walking zone-ineligible folios
>> that can never satisfy the allocation, inflating nr_reclaimed into a
>> false progress that delays the OOM.
>>
>> This series fixes both:
>>
>> [1/2] Charge the isolate overshoot against the scan quota so the
>> next round skips already-scanned folios.
>> [2/2] Stop scanning once too many zone-ineligible folios have been
>> skipped, instead of force-isolating them.
>
> This all sounds fairly bad. Is there some report? Real-world test
> results? Laboratory test results? Something to help us understand the
> impact of the issues and the situations in which they occur.
>
Sorry — this background belongs in the cover letter and I left it
out. Adding it here.
## Situation
We observed slow, unexpected OOM behavior during extreme stress testing,
and while digging into the reclaim code during that analysis we spotted
these latent risks in isolate_lru_folios() — the overshoot never
being charged back, and the force-isolate path on lower-zone reclaim.
The concerns below are the ones surfaced from reading the code, then
confirmed by constructing the situation deliberately.
The problem only appears when reclaim targets a lower zone while the
LRU holds folios from a higher zone:
- reclaim_idx points at DMA32/DMA (the triggering allocation asked
for a lower zone, e.g. __GFP_DMA32), and
- the LRU holds folios from a higher zone (Normal/Movable).
Normal userspace allocations go to the highest zone, so reclaim_idx
never points below it and the zone-skip branch is never taken. It
needs a real lower-zone allocator to drain that zone below watermark;
that is also why it went unnoticed upstream — the 1c7b17cf hard-lockup
fix that introduced the force-isolate path was found only on a ~1 TB
box running DMA32 module allocations.
## Impact
Reproduced on x86 QEMU, 7.2-rc6, with a kernel module doing
__GFP_DMA32 allocations to drain DMA32 below watermark (LRU folios
sit in Movable):
- A single isolate_lru_folios() call scanned 32794 pages and took
25 (32769 skipped): 99.9% wasted on ineligible folios.
- Across one run, isolate was called 339 times, 140-165 of which
isolated nothing (taken=0, pure empty scans).
- shrink_lruvec() charges only the 32-page quota per round and
never refunds the overshoot, so the same skipped folios are
rescanned round after round. vmstat on a memcg OOM path:
Δpgscan_direct / Δpgsteal_direct = 2836624 / 112719 = 25.2x
(isolate -> shrink_folio_list returns the folio -> isolate again).
Once max_nr_skipped hits SWAP_CLUSTER_MAX_SKIPPED, the current code
force-isolates the remaining ineligible folios. On the inactive LRU
they reach shrink_folio_list() and get reclaimed though they can
never satisfy the allocation, inflating nr_reclaimed and resetting
no_progress_loops in should_reclaim_retry(), delaying the OOM.
## A/B results (same .config, md5-identical)
| metric | baseline | patched | note |
|-------------------------|------------|---------|------|
| empty scans (taken=0) | 140-165 | 0 | stable across runs |
| isolate calls | 339 | 7 | 48x fewer |
| total pages scanned | 343711 | 3140 | 109x fewer |
| single-call max scan | 32794 | 3104 | |
Empty-scan 0 is the stable evidence (holds every run). The
Δpgscan/Δpgsteal ratio is volatile (baseline 56-26675x, patched
65-1324x, ranges overlap) and is not relied on alone.
## Caveats and reproduction
- Triggering needs a lower-zone-pressured box. To confirm the code
analysis, the situation was constructed on a small x86 QEMU VM
(1500M, CONFIG_LRU_GEN=n) with kernel cmdline
`movable_zone=DMA32 kernelcore=256M` (DMA32 small, Normal empty,
Movable large), then a kernel module doing
`alloc_page(GFP_DMA32)` drains DMA32 below watermark. LRU folios
sit in Movable and are zone-skipped while reclaiming for DMA32.
Observed via the `mm_vmscan_lru_isolate` tracepoint and
/proc/vmstat (pgscan_direct, pgsteal_direct).
> Thanks.
Hi, maintainers Kindly ping. On 周三 2026-8-12 10:29, mawupeng wrote: > > > On 周三 2026-8-12 09:42, Andrew Morton wrote: >> On Tue, 11 Aug 2026 15:08:47 +0800 Wupeng Ma <mawupeng1@huawei.com> wrote: >> >>> shrink_lruvec() drives reclaim in SWAP_CLUSTER_MAX (32) chunks, but >>> isolate_lru_folios() may scan far more than that per call on a single >>> LRU. The excess is never charged back, so shrink_lruvec() keeps >>> rescanning the same folios round after round. When reclaim targets a >>> lower zone, the same scanner also keeps walking zone-ineligible folios >>> that can never satisfy the allocation, inflating nr_reclaimed into a >>> false progress that delays the OOM. >>> >>> This series fixes both: >>> >>> [1/2] Charge the isolate overshoot against the scan quota so the >>> next round skips already-scanned folios. >>> [2/2] Stop scanning once too many zone-ineligible folios have been >>> skipped, instead of force-isolating them. >> >> This all sounds fairly bad. Is there some report? Real-world test >> results? Laboratory test results? Something to help us understand the >> impact of the issues and the situations in which they occur. >> > > Sorry — this background belongs in the cover letter and I left it > out. Adding it here. > > ## Situation > > We observed slow, unexpected OOM behavior during extreme stress testing, > and while digging into the reclaim code during that analysis we spotted > these latent risks in isolate_lru_folios() — the overshoot never > being charged back, and the force-isolate path on lower-zone reclaim. > The concerns below are the ones surfaced from reading the code, then > confirmed by constructing the situation deliberately. > > The problem only appears when reclaim targets a lower zone while the > LRU holds folios from a higher zone: > > - reclaim_idx points at DMA32/DMA (the triggering allocation asked > for a lower zone, e.g. __GFP_DMA32), and > - the LRU holds folios from a higher zone (Normal/Movable). > > Normal userspace allocations go to the highest zone, so reclaim_idx > never points below it and the zone-skip branch is never taken. It > needs a real lower-zone allocator to drain that zone below watermark; > that is also why it went unnoticed upstream — the 1c7b17cf hard-lockup > fix that introduced the force-isolate path was found only on a ~1 TB > box running DMA32 module allocations. > > ## Impact > > Reproduced on x86 QEMU, 7.2-rc6, with a kernel module doing > __GFP_DMA32 allocations to drain DMA32 below watermark (LRU folios > sit in Movable): > > - A single isolate_lru_folios() call scanned 32794 pages and took > 25 (32769 skipped): 99.9% wasted on ineligible folios. > - Across one run, isolate was called 339 times, 140-165 of which > isolated nothing (taken=0, pure empty scans). > - shrink_lruvec() charges only the 32-page quota per round and > never refunds the overshoot, so the same skipped folios are > rescanned round after round. vmstat on a memcg OOM path: > Δpgscan_direct / Δpgsteal_direct = 2836624 / 112719 = 25.2x > (isolate -> shrink_folio_list returns the folio -> isolate again). > > Once max_nr_skipped hits SWAP_CLUSTER_MAX_SKIPPED, the current code > force-isolates the remaining ineligible folios. On the inactive LRU > they reach shrink_folio_list() and get reclaimed though they can > never satisfy the allocation, inflating nr_reclaimed and resetting > no_progress_loops in should_reclaim_retry(), delaying the OOM. > > ## A/B results (same .config, md5-identical) > > | metric | baseline | patched | note | > |-------------------------|------------|---------|------| > | empty scans (taken=0) | 140-165 | 0 | stable across runs | > | isolate calls | 339 | 7 | 48x fewer | > | total pages scanned | 343711 | 3140 | 109x fewer | > | single-call max scan | 32794 | 3104 | | > > Empty-scan 0 is the stable evidence (holds every run). The > Δpgscan/Δpgsteal ratio is volatile (baseline 56-26675x, patched > 65-1324x, ranges overlap) and is not relied on alone. > > ## Caveats and reproduction > > - Triggering needs a lower-zone-pressured box. To confirm the code > analysis, the situation was constructed on a small x86 QEMU VM > (1500M, CONFIG_LRU_GEN=n) with kernel cmdline > `movable_zone=DMA32 kernelcore=256M` (DMA32 small, Normal empty, > Movable large), then a kernel module doing > `alloc_page(GFP_DMA32)` drains DMA32 below watermark. LRU folios > sit in Movable and are zone-skipped while reclaiming for DMA32. > Observed via the `mm_vmscan_lru_isolate` tracepoint and > /proc/vmstat (pgscan_direct, pgsteal_direct). > >> Thanks. >
On Tue, 25 Aug 2026 08:59:54 +0800 mawupeng <mawupeng1@huawei.com> wrote: > Hi, maintainers > Kindly ping. It's been quite a while. It would be helpful to refresh/retest/resend, please. While doing this, please update the changelogging to help reviewers understand the end-user impact of the issue. Please also give some though to whether we should backport these fixes into earlier kernels, with a cc:stable. Given the unusual circumstances which are required to hit this, I'm thinking "no", but feel free to disagree!
On 周日 2026-8-30 11:18, Andrew Morton wrote: > On Tue, 25 Aug 2026 08:59:54 +0800 mawupeng <mawupeng1@huawei.com> wrote: > >> Hi, maintainers >> Kindly ping. > > It's been quite a while. It would be helpful to refresh/retest/resend, > please. Thanks, Andrew. I’ll retest it with the background (situation, end-user impact, reproduction, A/B) folded into the cover letter and re-tested on the current -rc. > > While doing this, please update the changelogging to help reviewers > understand the end-user impact of the issue. > > Please also give some though to whether we should backport these fixes > into earlier kernels, with a cc:stable. Given the unusual > circumstances which are required to hit this, I'm thinking "no", but > feel free to disagree! On cc:stable — I agree it doesn’t seem warranted. The conditions needed to trigger this are unusual enough that I don’t think a backport is justified, so I’d lean toward not adding Cc:stable. Happy to revisit if I’m missing something. >
© 2016 - 2026 Red Hat, Inc.