fs/btrfs/compression.c | 3 +- fs/proc/task_mmu.c | 22 ++- include/linux/memcontrol.h | 47 +++++- include/linux/mm_inline.h | 256 ++++++++++++++++++++++++++------ include/linux/mmzone.h | 137 ++++++++++++----- kernel/bounds.c | 2 +- mm/filemap.c | 8 +- mm/folio.c | 49 +----- mm/huge_memory.c | 8 +- mm/khugepaged.c | 6 +- mm/madvise.c | 37 +++-- mm/memcontrol.c | 22 +-- mm/migrate.c | 4 - mm/page_io.c | 3 +- mm/readahead.c | 8 +- mm/vmscan.c | 360 ++++++++++++++++++++++++++++----------------- mm/workingset.c | 66 ++++++--- 17 files changed, 689 insertions(+), 349 deletions(-)
This is the updated RFC following the idea proposed at LSF/MM/BPF [1] this
year. It's very usable, stable, and performing well, but I'll keep it RFC
for V1 as some tests are still ongoing and results can be more accurate
with further auditing.
With this series, I'm seeing an obvious performance gain across all kinds
of tests, and it reduces MGLRU's flag usage by one. It also fixes several
long-standing issues including under-accounted PSI and poor workingset
tracking (especially for page cache).
Some test results (CLRU means classical LRU):
Build kernel test, running make -j48 in a 3G memcg, using disk swap and
holding the kernel and build output on the same NVMe drive, 16 runs using
different swappiness configurations [2]; the patched version is better than
mainline at almost every swappiness value, measuring the total average:
real sys pgpgin pswpin pswpout refault_file refault_anon
CLRU 6m06s 31m01s 50.3M 3.20M 13.8M 10.3M 3.35M
Before 2m56s 11m06s 10.6M 1.59M 5.40M 414k 1.03M
After 2m49s 10m39s 9.0M 1.36M 4.82M 280k 861k
delta -7s -27s -15% -14% -11% -32% -16%
MongoDB YCSB workloadb (recordcount:20000000 operationcount:6000000,
threads:48, in a 16G memcg), 3 runs [3]:
CLRU: 98389.94 ops/s
MGLRU Before: 86421.44 ops/s
MGLRU After: 95378.50 ops/s (+10.3%)
Chromium & Node.js test, using ZRAM as swap, on a 48c96t machine with
128G memory, 64 workers, run for 1 hour [4]:
Total requests:
CLRU: 63822
MGLRU Before: 153029
MGLRU After: 225664 (+47.4%)
(NOTE: It seems some recent change broken MGLRU's fainress guarteen and
also made this test dramatically faster than a few months ago, which isn't
related to this series and reading are even better now, but I'll take a
deeper look later.)
FIO, zipf 0.9 distribution on an NVMe disk, in a 16G cgroup, total file
size 40G; this measures the LRU's theoretical ability to distinguish the
hotter portion:
fio --name=fg --numjobs=16 --nrfiles=1 \
--filename_format="$testdir/rnvmedk.\$jobnum.img" \
--size=${FILE_MIB}M \
--buffered=1 --ioengine=sync --rw=randread \
--random_distribution=zipf:$ZIPF --bs=$BS --time_based \
--ramp_time=45s --runtime=600s \
--group_reporting
CLRU: Avg: 2454.37 MB/s
Before: Avg: 2350.40 MB/s
After: Avg: 2611.37 MB/s (+11.1%)
I also retested the LevelDB benchmark from the cache_ext paper [5].
Interestingly, mainline MGLRU already beats CLRU on this one after a
recent change in lru_gen_folio_seq that bumps new folios with refs == 1
to the second-oldest generation. That change accidentally gave random
reads a higher hotness level while making sequential reads much colder:
sequential reads involve many readahead hits, and readahead folios start
with refs == 0, so they're already in the oldest generation, and
folio_mark_accessed() on a readahead hit has almost no effect on generations
in mainline MGLRU. Meanwhile, all direct-hit (random read) folios start
with refs == 1 in the second-oldest generation. As a result, the scan-get
test natively protects the "get" part and sacrifices the "scan" part.
That's not the best solution though. It's unreliable because it depends
on LRU drain timing, and it hurts workloads where the sequential part is
actually hotter (any workload involving a hotter large file and many
small cold files will be affected).
This series improves on that base: it covers ordinary workloads without
hurting the scan-get workload and without relying on that initial bump.
LevelDB Scan / Get, Throughput Total:
CLRU: 4668.8 ops/s
MGLRU: 5026.9 ops/s (faster than CLRU, but hurts other workloads)
MGLRU After: 5029.7 ops/s (fastest in all cases, and no regression)
The hot-sequential and cold-random workload can be easily reproduced with
SQLite and grep. SQLite continuously scans and looks up a small hot
portion of a DB file, while grep iterates over a set of small files much
larger than RAM [6]:
SQLite scan & lookup time: Grep iterate time:
CLRU: 14.51ms 13281.37ms
MGLRU mainline: 567.05ms 13694.47ms
MGLRU After this series: 10.58ms 12930.43ms
The grep cold portion is larger than RAM and accessed only once per
iteration, so there's no promotion of any of it. CLRU handles
this reasonably; mainline MGLRU has a clear regression; MGLRU-FG now
not only recovers but is able to catch some hot parts from the cold grep
workload. This test is somewhat subjective, but the signal is clear.
Additionally, PSI, smaps, and readahead all benefit from better accuracy
since this series unifies the flag usage between classical LRU and MGLRU.
Other tests such as MySQL are looking fine, with no regressions.
Refault distance is not included yet, so MGLRU may respond more slowly to
workingset shifts. That can be added later, as previously demonstrated
[7], [8].
Extra note about future development: this series is highly compatible with
ideas like workingset reporting [9]. The "gen climbing folio" design may
appear to conflict with workingset reporting's idea of using generations
as access-gap identifiers, but it doesn't — the solution is
straightforward: once we can extend the generation number to a larger
value (e.g. 64 or 128), the refs-driven promotion can stop at a lower
gen (e.g. oldest_gen + 16), leaving the remaining newer generations as
perfectly time-gap-separated bins.
The tier count is not fixed either; we'll need to find a way to tune it
if tiers go beyond 4, but that shouldn't be hard.
More details are in the individual commit messages.
Link: https://lore.kernel.org/linux-mm/CAMgjq7BoekNjg-Ra3C8M7=8=75su38w=HD782T5E_cxyeCeH_g@mail.gmail.com/ [1]
Link: https://lore.kernel.org/linux-mm/CAGsJ_4xre-x0e+qNVm=KLFnO1dbPkPX5RuecqwvTZu-vS+o8yQ@mail.gmail.com/ [2]
Link: https://github.com/brianfrankcooper/YCSB/blob/master/workloads/workloadb [3]
Link: https://lore.kernel.org/all/20221220214923.1229538-1-yuzhao@google.com/ [4]
Link: https://dl.acm.org/doi/10.1145/3731569.3764820 [5]
Link: https://github.com/ryncsn/emm-test-project/tree/master/sqlite-grep [6]
Link: https://lwn.net/Articles/945266/ [7]
Link: https://lore.kernel.org/linux-mm/20260502-mglru-fg-v1-0-913619b014d9@tencent.com/ [8]
Link: https://lwn.net/Articles/976985/ [9]
Signed-off-by: Kairui Song <kasong@tencent.com>
---
Kairui Song (15):
mm/memcontrol: make lru_zone_size atomic and simplify sanity check
mm/memcontrol: allow update of LRU statistic without holding LRU lock
mm/mglru: introduce and always use helpers for manipulating page flags
mm/mglru: make generation page counters atomic
mm/mglru: move max_seq read into walk_update_folio
mm/mglru: use explicit tier range in read_ctrl_pos()
mm/mglru: move refault workingset activation into lru_gen_refault
mm/memcg: add folio-based lruvec live helper
mm/mglru: frequency guided workingset promotion (MGLRU-FG)
mm/mglru: make folio lru referenced times count a generic API
mm/mglru: replace folio workinset check and update with new helper
mm/smap: report workingset folios as referenced
mm/huge_memory: mark file folio as accessed more accurately on split
mm/khugepaged: consider workingset folios as referenced
mm/madvise: convert to new lru refs API and better support for MGLRU
fs/btrfs/compression.c | 3 +-
fs/proc/task_mmu.c | 22 ++-
include/linux/memcontrol.h | 47 +++++-
include/linux/mm_inline.h | 256 ++++++++++++++++++++++++++------
include/linux/mmzone.h | 137 ++++++++++++-----
kernel/bounds.c | 2 +-
mm/filemap.c | 8 +-
mm/folio.c | 49 +-----
mm/huge_memory.c | 8 +-
mm/khugepaged.c | 6 +-
mm/madvise.c | 37 +++--
mm/memcontrol.c | 22 +--
mm/migrate.c | 4 -
mm/page_io.c | 3 +-
mm/readahead.c | 8 +-
mm/vmscan.c | 360 ++++++++++++++++++++++++++++-----------------
mm/workingset.c | 66 ++++++---
17 files changed, 689 insertions(+), 349 deletions(-)
---
base-commit: 94f9b3980dd446b56acf1dfed649e9b32a9f3813
change-id: 20260722-mglru-fg-3a2c8574725b
Best regards,
--
Kairui Song <kasong@tencent.com>
Hi Kairui,
Thanks for the MGLRU-FG series - it is a very interesting and
well-written piece of work. I ported it locally and ran a subset of the
benchmarks from your cover letter on our test VM. The results are
positive and consistent with yours, so I wanted to share them as early
feedback. I would be happy to test the next posted revision directly
and provide a formal Tested-by tag then.
Test environment
----------------
- Host: 96 physical cores / 192 threads
- Storage: VM images on a 4 TB Netac SATA SSD (no NVMe); VM disks
are virtio-backed image files on that SSD
- VM: 96 vCPUs, 128 GiB RAM (KVM)
- Baseline: mm-new @ a032d41a86cb
- MGLRU-FG squashed into local commit 9fc6ff682d08
- NOTE: the port additionally contains a local guard in
folio_inc_lru_refs() for non-LRU folios, so this is not an exact
bit-for-bit test of the posted series.
1. fio, Zipf 0.9 buffered 4 KiB random reads (all IO through the
page cache), 40 GiB total files, 16 jobs, 16 GiB memcg,
45 s ramp + 600 s run:
run 1 run 2 run 3 median
CLRU 797.64 796.74 800.21 797.64 MB/s
MGLRU before 766.88 273.75(*) 769.81 766.88 MB/s
MGLRU-FG 838.16 840.26 367.92(*) 838.16 MB/s
Median improvement over MGLRU before: +9.30% (cover letter: 11.1%).
The two marked runs entered an unexplained low-throughput state;
CLRU stayed at 797-800 MB/s, and a reboot restored MGLRU-before to
769.81 MB/s, so I kept the outliers visible rather than dropping
them. I am investigating whether this is related to MGLRU
history/state across hot switches.
2. MongoDB 4.0.23 / YCSB 0.17.0 workload B
20 M records, 6 M operations, 48 threads, 16 GiB memcg, 8 GiB
WiredTiger cache:
run 1 run 2 run 3 median
CLRU 18695.72 18673.44 18438.05 18673.44 ops/s
MGLRU before 17517.69 16196.08 18180.28 17517.69 ops/s
MGLRU-FG 19871.70 19721.86 19321.00 19721.86 ops/s
MGLRU-FG improved median throughput by +12.58% (cover letter:
10.3%). Median read latency / p95 / p99 improved by 11.35% /
12.30% / 26.06%, respectively.
Caveat: the master DB was loaded with w=0 and stopped too soon, so
each run had about 22.7k READ NOT_FOUND results (~0.4% of reads).
Every variant was cloned from the same XFS reflink master, so I
think the relative result is still valid, but I will redo it with a
fully settled master.
3. SQLite hot lookup + cold grep stream (sqlite-grep test, 90,000 cold
files, 300 MiB memcg); median of run-medians for the hot lookup:
MGLRU before: 277.18 ms
MGLRU-FG: 13.74 ms (-95.0%)
Eight of ten MGLRU-FG samples were about 13-14 ms, whereas most
MGLRU-before samples were 249-322 ms. Grep time itself changed by
only about +1%. Compared with MGLRU before, MGLRU-FG also reduced
file refaults by about 21%, direct scans by about 17.5%, and pgpgin
by about 15.1%. This was the clearest reproduction of the intended
protection against a one-shot cold stream evicting hot data.
General observations
--------------------
- No memcg OOM kills, and no kernel BUG / Oops / Call Trace / panic in
dmesg across all completed runs.
- Confirmed that folio_inc_lru_refs() is exercised at high frequency
under the fio workload.
- make -j48 / 3 GiB memcg / disk-swap (CLRU side only) completed in
2h18m59.9s with 20.9M file refaults and 48.1M anon refaults, no OOM
kill. This VM builds about 13.4k objects and many modules, so its
absolute time is not comparable with the machine in the cover letter.
The before/after MGLRU sides are still running; I can send the
same-machine comparison separately.
One thing I want to flag: my local port adds a guard in
folio_inc_lru_refs() for non-LRU folios. I noticed that syzbot
reported a WARNING in folio_inc_lru_refs() on this series
("!memcg && !mem_cgroup_disabled()" in the exit_mmap path). I am not
sure whether the two are related, but if it helps, I can check whether
my guard's trigger path matches that report.
Because this was a local port plus an additional guard, I am reporting
these results without a Tested-by tag for now. I would be happy to test
the next posted revision directly and provide a formal Tested-by then.
Thanks and best regards,
zhaozhengzhuo
zhaozhengzhuo@uniontech.com
On Fri, Aug 28, 2026 at 6:26 PM zhaozhengzhuo
<zhaozhengzhuo@uniontech.com> wrote:
>
> Hi Kairui,
Hi!
>
> Thanks for the MGLRU-FG series - it is a very interesting and
> well-written piece of work. I ported it locally and ran a subset of the
> benchmarks from your cover letter on our test VM. The results are
> positive and consistent with yours, so I wanted to share them as early
> feedback. I would be happy to test the next posted revision directly
> and provide a formal Tested-by tag then.
Thanks for the testing! Nice to see people testing and help verify the
performance, really good data. MGLRU-FG is supposed to be a generic
improvement based on better modeling and tracking of the access info
of folios, glad to see this can also be verified on your side. I'll
also try attach more test data in V2.
>
> Test environment
> ----------------
> - Host: 96 physical cores / 192 threads
> - Storage: VM images on a 4 TB Netac SATA SSD (no NVMe); VM disks
> are virtio-backed image files on that SSD
> - VM: 96 vCPUs, 128 GiB RAM (KVM)
> - Baseline: mm-new @ a032d41a86cb
> - MGLRU-FG squashed into local commit 9fc6ff682d08
> - NOTE: the port additionally contains a local guard in
> folio_inc_lru_refs() for non-LRU folios, so this is not an exact
> bit-for-bit test of the posted series.
>
> 1. fio, Zipf 0.9 buffered 4 KiB random reads (all IO through the
> page cache), 40 GiB total files, 16 jobs, 16 GiB memcg,
> 45 s ramp + 600 s run:
>
> run 1 run 2 run 3 median
> CLRU 797.64 796.74 800.21 797.64 MB/s
> MGLRU before 766.88 273.75(*) 769.81 766.88 MB/s
> MGLRU-FG 838.16 840.26 367.92(*) 838.16 MB/s
>
> Median improvement over MGLRU before: +9.30% (cover letter: 11.1%).
> The two marked runs entered an unexplained low-throughput state;
> CLRU stayed at 797-800 MB/s, and a reboot restored MGLRU-before to
> 769.81 MB/s, so I kept the outliers visible rather than dropping
> them. I am investigating whether this is related to MGLRU
> history/state across hot switches.
Oh, usually I reboot the system rather than use runtime switch,
runtime switch is known to skew all the shadow and many other metrics
for MGLRU. But there is also another possibility: latency spikes in
MGLRU, e.g. aging, which we will fix later.
> 2. MongoDB 4.0.23 / YCSB 0.17.0 workload B
> 20 M records, 6 M operations, 48 threads, 16 GiB memcg, 8 GiB
> WiredTiger cache:
>
> run 1 run 2 run 3 median
> CLRU 18695.72 18673.44 18438.05 18673.44 ops/s
> MGLRU before 17517.69 16196.08 18180.28 17517.69 ops/s
> MGLRU-FG 19871.70 19721.86 19321.00 19721.86 ops/s
>
> MGLRU-FG improved median throughput by +12.58% (cover letter:
> 10.3%). Median read latency / p95 / p99 improved by 11.35% /
> 12.30% / 26.06%, respectively.
> Caveat: the master DB was loaded with w=0 and stopped too soon, so
> each run had about 22.7k READ NOT_FOUND results (~0.4% of reads).
> Every variant was cloned from the same XFS reflink master, so I
> think the relative result is still valid, but I will redo it with a
> fully settled master.
>
> 3. SQLite hot lookup + cold grep stream (sqlite-grep test, 90,000 cold
> files, 300 MiB memcg); median of run-medians for the hot lookup:
>
> MGLRU before: 277.18 ms
> MGLRU-FG: 13.74 ms (-95.0%)
>
> Eight of ten MGLRU-FG samples were about 13-14 ms, whereas most
> MGLRU-before samples were 249-322 ms. Grep time itself changed by
> only about +1%. Compared with MGLRU before, MGLRU-FG also reduced
> file refaults by about 21%, direct scans by about 17.5%, and pgpgin
> by about 15.1%. This was the clearest reproduction of the intended
> protection against a one-shot cold stream evicting hot data.
Nice data from DB tests. I believe these gain are from the actual
imrpovement of LRU's ability to keeping the workingset and evict
less-frequently used folios.
> General observations
> --------------------
> - No memcg OOM kills, and no kernel BUG / Oops / Call Trace / panic in
> dmesg across all completed runs.
> - Confirmed that folio_inc_lru_refs() is exercised at high frequency
> under the fio workload.
> - make -j48 / 3 GiB memcg / disk-swap (CLRU side only) completed in
> 2h18m59.9s with 20.9M file refaults and 48.1M anon refaults, no OOM
> kill. This VM builds about 13.4k objects and many modules, so its
> absolute time is not comparable with the machine in the cover letter.
> The before/after MGLRU sides are still running; I can send the
> same-machine comparison separately.
>
> One thing I want to flag: my local port adds a guard in
> folio_inc_lru_refs() for non-LRU folios. I noticed that syzbot
> reported a WARNING in folio_inc_lru_refs() on this series
> ("!memcg && !mem_cgroup_disabled()" in the exit_mmap path). I am not
> sure whether the two are related, but if it helps, I can check whether
> my guard's trigger path matches that report.
It can encounter non-LRU folios, and raise an false warning. I'll
handle this better in V2. Non-LRU folios are totally fine and often
seen there, it just need to skip the gen part to avoid the WARN.
> Because this was a local port plus an additional guard, I am reporting
> these results without a Tested-by tag for now. I would be happy to test
> the next posted revision directly and provide a formal Tested-by then.
>
> Thanks and best regards,
> zhaozhengzhuo
> zhaozhengzhuo@uniontech.com
Thanks again! There are some work going on upstream, so I spent some
time on other items; I will post V2 Ccing you.
syzbot ci has tested the following series [v1] mm/mglru: frequency guided promotion (MGLRU-FG) and flag cleanup https://lore.kernel.org/all/20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com * [PATCH RFC 01/15] mm/memcontrol: make lru_zone_size atomic and simplify sanity check * [PATCH RFC 02/15] mm/memcontrol: allow update of LRU statistic without holding LRU lock * [PATCH RFC 03/15] mm/mglru: introduce and always use helpers for manipulating page flags * [PATCH RFC 04/15] mm/mglru: make generation page counters atomic * [PATCH RFC 05/15] mm/mglru: move max_seq read into walk_update_folio * [PATCH RFC 06/15] mm/mglru: use explicit tier range in read_ctrl_pos() * [PATCH RFC 07/15] mm/mglru: move refault workingset activation into lru_gen_refault * [PATCH RFC 08/15] mm/memcg: add folio-based lruvec live helper * [PATCH RFC 09/15] mm/mglru: frequency guided workingset promotion (MGLRU-FG) * [PATCH RFC 10/15] mm/mglru: make folio lru referenced times count a generic API * [PATCH RFC 11/15] mm/mglru: replace folio workinset check and update with new helper * [PATCH RFC 12/15] mm/smap: report workingset folios as referenced * [PATCH RFC 13/15] mm/huge_memory: mark file folio as accessed more accurately on split * [PATCH RFC 14/15] mm/khugepaged: consider workingset folios as referenced * [PATCH RFC 15/15] mm/madvise: convert to new lru refs API and better support for MGLRU and found the following issue: WARNING in folio_inc_lru_refs Full report is available here: https://ci.syzbot.org/series/5db36d1d-9faa-4882-9f0d-1a8f52132274 *** WARNING in folio_inc_lru_refs tree: mm-new URL: https://kernel.googlesource.com/pub/scm/linux/kernel/git/akpm/mm.git base: 94f9b3980dd446b56acf1dfed649e9b32a9f3813 arch: amd64 compiler: Debian clang version 22.1.8 (++20260613092233+e80beda6e255-1~exp1~20260613092250.77), Debian LLD 22.1.8 config: https://ci.syzbot.org/builds/9f324367-2b95-4fd0-9025-fef1ff1f605a/config page: refcount:3 mapcount:2 mapping:0000000000000000 index:0x0 pfn:0xe4ee flags: 0xfff00000002000(reserved|node=0|zone=1|lastcpupid=0x7ff) raw: 00fff00000002000 ffffea0000393b88 ffffea0000393b88 0000000000000000 raw: 0000000000000000 0000000000000000 0000000300000001 0000000000000000 page dumped because: VM_WARN_ON_ONCE_FOLIO(!memcg && !mem_cgroup_disabled()) page_owner info is not present (never set?) ------------[ cut here ]------------ 1 WARNING: ./include/linux/memcontrol.h:745 at folio_inc_lru_refs+0xb4f/0xc10, CPU#0: mount/5025 Modules linked in: CPU: 0 UID: 0 PID: 5025 Comm: mount Not tainted syzkaller #0 PREEMPT(full) Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.2-debian-1.16.2-1 04/01/2014 RIP: 0010:folio_inc_lru_refs+0xb4f/0xc10 Code: ff 4c 89 e7 e8 c2 ae fd ff e9 e9 fc ff ff e8 18 91 ba ff 4c 89 e7 48 c7 c6 20 90 f8 8b e8 99 b0 1b ff c6 05 a6 a9 34 0e 01 90 <0f> 0b 90 e9 85 f6 ff ff e8 f4 90 ba ff e9 29 f8 ff ff 44 89 f1 80 RSP: 0018:ffffc9000324f4c0 EFLAGS: 00010246 RAX: 8e914a919570c700 RBX: 0000000000000000 RCX: 0000000000000001 RDX: 0000000000000000 RSI: ffffffff8e4b4187 RDI: ffff888174f53c00 RBP: ffffc9000324f5d0 R08: 0000000000000003 R09: 0000000000000004 R10: dffffc0000000000 R11: fffffbfff1d3ca24 R12: ffffea0000393b80 R13: 1ffffd4000072770 R14: 1ffff92000649ea8 R15: dffffc0000000000 FS: 0000000000000000(0000) GS:ffff88818d949000(0000) knlGS:0000000000000000 CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 CR2: 00007fb77b215440 CR3: 000000000e946000 CR4: 00000000000006f0 Call Trace: <TASK> __zap_vma_range+0x20f5/0x4f70 unmap_vmas+0x390/0x550 exit_mmap+0x293/0x9f0 __mmput+0x118/0x420 exit_mm+0x221/0x2d0 do_exit+0x6cd/0x2360 do_group_exit+0x22d/0x2f0 __x64_sys_exit_group+0x3f/0x40 x64_sys_call+0x221a/0x2240 do_syscall_64+0x174/0x580 entry_SYSCALL_64_after_hwframe+0x77/0x7f RIP: 0033:0x7fb77b2d3a90 Code: Unable to access opcode bytes at 0x7fb77b2d3a66. RSP: 002b:00007ffe32d58618 EFLAGS: 00000246 ORIG_RAX: 00000000000000e7 RAX: ffffffffffffffda RBX: 00007fb77b3c4860 RCX: 00007fb77b2d3a90 RDX: 00000000000000e7 RSI: 000000000000003c RDI: 0000000000000000 RBP: 00007fb77b3c4860 R08: 00007ffe32d58490 R09: 00007ffe32d58570 R10: 00007ffe32d584d0 R11: 0000000000000246 R12: 0000000000000000 R13: 0000000000000000 R14: 00007fb77b3c8658 R15: 0000000000000001 </TASK> *** If these findings have caused you to resend the series or submit a separate fix, please add the following tag to your commit message: Tested-by: syzbot@syzkaller.appspotmail.com --- This report is generated by a bot. It may contain errors. syzbot ci engineers can be reached at syzkaller@googlegroups.com. To test a patch for this bug, please reply with `#syz test` (should be on a separate line). The patch should be attached to the email. Note: arguments like custom git repos and branches are not supported.
On Mon, Aug 03, 2026 at 10:26:17PM +0800, syzbot ci wrote: > syzbot ci has tested the following series > > [v1] mm/mglru: frequency guided promotion (MGLRU-FG) and flag cleanup > https://lore.kernel.org/all/20260804-mglru-fg-v1-0-4d8dad39dad6@tencent.com > * [PATCH RFC 01/15] mm/memcontrol: make lru_zone_size atomic and simplify sanity check > * [PATCH RFC 02/15] mm/memcontrol: allow update of LRU statistic without holding LRU lock > * [PATCH RFC 03/15] mm/mglru: introduce and always use helpers for manipulating page flags > * [PATCH RFC 04/15] mm/mglru: make generation page counters atomic > * [PATCH RFC 05/15] mm/mglru: move max_seq read into walk_update_folio > * [PATCH RFC 06/15] mm/mglru: use explicit tier range in read_ctrl_pos() > * [PATCH RFC 07/15] mm/mglru: move refault workingset activation into lru_gen_refault > * [PATCH RFC 08/15] mm/memcg: add folio-based lruvec live helper > * [PATCH RFC 09/15] mm/mglru: frequency guided workingset promotion (MGLRU-FG) > * [PATCH RFC 10/15] mm/mglru: make folio lru referenced times count a generic API > * [PATCH RFC 11/15] mm/mglru: replace folio workinset check and update with new helper > * [PATCH RFC 12/15] mm/smap: report workingset folios as referenced > * [PATCH RFC 13/15] mm/huge_memory: mark file folio as accessed more accurately on split > * [PATCH RFC 14/15] mm/khugepaged: consider workingset folios as referenced > * [PATCH RFC 15/15] mm/madvise: convert to new lru refs API and better support for MGLRU > > and found the following issue: > WARNING in folio_inc_lru_refs > > Full report is available here: > https://ci.syzbot.org/series/5db36d1d-9faa-4882-9f0d-1a8f52132274 > > *** > > WARNING in folio_inc_lru_refs > > tree: mm-new > URL: https://kernel.googlesource.com/pub/scm/linux/kernel/git/akpm/mm.git > base: 94f9b3980dd446b56acf1dfed649e9b32a9f3813 > arch: amd64 > compiler: Debian clang version 22.1.8 (++20260613092233+e80beda6e255-1~exp1~20260613092250.77), Debian LLD 22.1.8 > config: https://ci.syzbot.org/builds/9f324367-2b95-4fd0-9025-fef1ff1f605a/config > > page: refcount:3 mapcount:2 mapping:0000000000000000 index:0x0 pfn:0xe4ee > flags: 0xfff00000002000(reserved|node=0|zone=1|lastcpupid=0x7ff) > raw: 00fff00000002000 ffffea0000393b88 ffffea0000393b88 0000000000000000 > raw: 0000000000000000 0000000000000000 0000000300000001 0000000000000000 > page dumped because: VM_WARN_ON_ONCE_FOLIO(!memcg && !mem_cgroup_disabled()) > page_owner info is not present (never set?) > ------------[ cut here ]------------ > 1 > WARNING: ./include/linux/memcontrol.h:745 at folio_inc_lru_refs+0xb4f/0xc10, CPU#0: mount/5025 > Modules linked in: > CPU: 0 UID: 0 PID: 5025 Comm: mount Not tainted syzkaller #0 PREEMPT(full) > Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.2-debian-1.16.2-1 04/01/2014 > RIP: 0010:folio_inc_lru_refs+0xb4f/0xc10 > Code: ff 4c 89 e7 e8 c2 ae fd ff e9 e9 fc ff ff e8 18 91 ba ff 4c 89 e7 48 c7 c6 20 90 f8 8b e8 99 b0 1b ff c6 05 a6 a9 34 0e 01 90 <0f> 0b 90 e9 85 f6 ff ff e8 f4 90 ba ff e9 29 f8 ff ff 44 89 f1 80 > RSP: 0018:ffffc9000324f4c0 EFLAGS: 00010246 > RAX: 8e914a919570c700 RBX: 0000000000000000 RCX: 0000000000000001 > RDX: 0000000000000000 RSI: ffffffff8e4b4187 RDI: ffff888174f53c00 > RBP: ffffc9000324f5d0 R08: 0000000000000003 R09: 0000000000000004 > R10: dffffc0000000000 R11: fffffbfff1d3ca24 R12: ffffea0000393b80 > R13: 1ffffd4000072770 R14: 1ffff92000649ea8 R15: dffffc0000000000 > FS: 0000000000000000(0000) GS:ffff88818d949000(0000) knlGS:0000000000000000 > CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 > CR2: 00007fb77b215440 CR3: 000000000e946000 CR4: 00000000000006f0 > Call Trace: > <TASK> > __zap_vma_range+0x20f5/0x4f70 > unmap_vmas+0x390/0x550 > exit_mmap+0x293/0x9f0 > __mmput+0x118/0x420 > exit_mm+0x221/0x2d0 > do_exit+0x6cd/0x2360 > do_group_exit+0x22d/0x2f0 > __x64_sys_exit_group+0x3f/0x40 > x64_sys_call+0x221a/0x2240 > do_syscall_64+0x174/0x580 > entry_SYSCALL_64_after_hwframe+0x77/0x7f > RIP: 0033:0x7fb77b2d3a90 > Code: Unable to access opcode bytes at 0x7fb77b2d3a66. > RSP: 002b:00007ffe32d58618 EFLAGS: 00000246 ORIG_RAX: 00000000000000e7 > RAX: ffffffffffffffda RBX: 00007fb77b3c4860 RCX: 00007fb77b2d3a90 > RDX: 00000000000000e7 RSI: 000000000000003c RDI: 0000000000000000 > RBP: 00007fb77b3c4860 R08: 00007ffe32d58490 R09: 00007ffe32d58570 > R10: 00007ffe32d584d0 R11: 0000000000000246 R12: 0000000000000000 > R13: 0000000000000000 R14: 00007fb77b3c8658 R15: 0000000000000001 > </TASK> OK, so we might hit a uncharged folio in folio_mark_access, which isn't strange, right now it will just skip the gen bump and work as expected, problem is it's triggering this warning, and doing redundant work for lruvec lookup. To fix that, checking flags first then the lruvec should be good. Will do in V2.
© 2016 - 2026 Red Hat, Inc.