[RFC PATCH v3 0/4] mm: enable lru cache for smaller large folios

Barry Song (Xiaomi) posted 4 patches 1 month, 1 week ago
include/linux/folio_batch.h | 25 +++++++++++++++++++++++++
mm/folio.c                  | 10 +++++++++-
mm/huge_memory.c            | 23 +++++++++++++++++++----
mm/internal.h               |  4 ++--
mm/memory.c                 |  6 ++++++
5 files changed, 61 insertions(+), 7 deletions(-)
[RFC PATCH v3 0/4] mm: enable lru cache for smaller large folios
Posted by Barry Song (Xiaomi) 1 month, 1 week ago
This patchset enables the per-CPU LRU cache for large folios with fewer
than `FOLIO_BATCH_SIZE` (31) pages. It also limits each per-CPU LRU cache
to at most `FOLIO_BATCH_SIZE` pages to avoid negatively affecting
accounting and memory reclamation pressure.

This is particularly beneficial on systems that use relatively small
large folios. For larger folios, the benefit is likely to be smaller
because far fewer folios are expected to contend for the LRU cache.

* Use the following microbenchmark:

 #include <pthread.h>
 #include <sys/mman.h>
 #include <string.h>

 #define NUM_THREADS     20
 #define MEM_SIZE        (16 * 1024 * 1024)
 #define LOOP_COUNT      1000

 void* thread_worker(void* arg) {
        void *addr = mmap(NULL, MEM_SIZE, PROT_READ | PROT_WRITE,
                        MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);

        for (int i = 0; i < LOOP_COUNT; i++) {
                for(int j = 0; j < MEM_SIZE; j += 4096)
                        *(unsigned char *)(addr + j) = 0x55;
                madvise(addr, MEM_SIZE, MADV_DONTNEED);
        }

        munmap(addr, MEM_SIZE);
        pthread_exit(NULL);
 }

 int main() {
        pthread_t threads[NUM_THREADS];

        for (long t = 0; t < NUM_THREADS; t++) {
                pthread_create(&threads[t], NULL, thread_worker, (void*)t);
        }

        for (int t = 0; t < NUM_THREADS; t++) {
                pthread_join(threads[t], NULL);
        }

        return 0;
 }

w/o patch:

$ time ./a.out 

real	0m17.760s
user	0m8.207s
sys	5m46.067s

perf lock report:
                Name   acquired  contended     avg wait   total wait     max wait     min wait 

                       21414857   21414857     30.07 us     10.73 m     101.94 us       880 ns 
                          42890      42890     11.29 us    484.21 ms    123.08 us       923 ns 
                           1678       1678      1.57 us      2.63 ms      5.77 us       985 ns 
                             52         52      1.46 us     75.98 us      2.69 us      1.07 us 
                             18         18     17.82 us    320.69 us     58.92 us      1.62 us 
                             18         18      2.36 ms     42.52 ms      5.31 ms      1.47 us 
           rcu_state         12         12      1.98 us     23.78 us      2.62 us      1.47 us 
           rcu_state          9          9      1.72 us     15.49 us      2.27 us      1.34 us 
                              2          2      2.42 us      4.83 us      2.48 us      2.36 us 

w/ patch:

$ time ./a.out 

real	0m16.292s
user	0m8.587s
sys	5m13.787s

perf lock report

                Name   acquired  contended     avg wait   total wait     max wait     min wait 

                        2641286    2641286     46.87 us      2.06 m     107.92 us      1.01 us 
                         235275     235275     10.57 us      2.49 s     115.32 us      1.02 us 
           rcu_state       1982       1982      7.58 us     15.02 ms     31.72 us      1.26 us 
           rcu_state       1929       1929      7.42 us     14.32 ms     30.50 us      1.43 us 
                             86         86      3.61 us    310.19 us     13.89 us      1.15 us 
                             20         20     19.70 us    394.03 us    108.07 us      2.15 us 
       tasklist_lock          1          1      2.13 us      2.13 us      2.13 us      2.13 us 


* Build the kernel in a 1 GiB memcg by -j20 with zRAM configured as swap:

w/o patch:

Perf lock report:

                Name   acquired  contended     avg wait   total wait     max wait     min wait 

                         782337     782337     17.61 us     13.78 s     402.22 us       896 ns 
                          55459      55459     19.48 us      1.08 s     117.59 us      1.01 us 
                           7826       7826      8.01 us     62.68 ms     19.03 us       887 ns 
                           5324       5324      7.83 us     41.68 ms     37.59 us      1.11 us 
           rcu_state       5144       5144      6.27 us     32.23 ms     25.17 us      1.58 us 
           rcu_state       5142       5142      6.35 us     32.67 ms     30.55 us      1.48 us 
                           3855       3855      8.74 us     33.68 ms     42.32 us       996 ns 
                           2770       2770      9.22 us     25.55 ms     27.83 us       914 ns 
                           2342       2342      5.88 us     13.77 ms    318.75 us      1.00 us 
time:

*** Executing round 0 ***

real    1m46.847s
user    25m10.848s
sys     2m57.282s

*** Executing round 1 ***

real    1m46.423s
user    25m10.072s
sys     2m54.348s

*** Executing round 2 ***

real    1m46.308s
user    25m13.800s
sys     2m58.963s

*** Executing round 3 ***

real    1m46.155s
user    25m18.079s
sys     2m59.721s

*** Executing round 4 ***

real    1m45.980s
user    25m15.493s
sys     2m56.959s

w/ patch:

perf lock report
                Name   acquired  contended     avg wait   total wait     max wait     min wait 

                         202647     202647     34.27 us      6.94 s     467.82 us      1.18 us 
                          51819      51819     16.26 us    842.46 ms    245.55 us       885 ns 
     inode_hash_lock      15169      15169      8.58 us    130.21 ms     23.23 us      1.04 us 
                           5306       5306      7.54 us     40.03 ms     31.53 us      1.17 us 
           rcu_state       4945       4945      6.97 us     34.47 ms     30.00 us      1.76 us 
           rcu_state       4899       4899      7.03 us     34.42 ms     24.65 us      1.42 us 
                           3923       3923      8.56 us     33.57 ms     27.51 us      1.08 us 
                           2212       2212      5.23 us     11.56 ms    222.79 us       965 ns 
                           1412       1412      6.35 us      8.97 ms     23.09 us      1.77 us 

time:

*** Executing round 0 ***

real	1m46.463s
user	25m17.448s
sys	2m49.524s

*** Executing round 1 ***

real	1m46.274s
user	25m12.178s
sys	2m53.522s

*** Executing round 2 ***

real	1m46.362s
user	25m13.115s
sys	2m53.005s

*** Executing round 3 ***

real	1m46.036s
user	25m17.627s
sys	2m53.477s

*** Executing round 4 ***

real	1m46.329s
user	25m15.130s
sys	2m51.508s

-RFC v3:
  * Rather than hard-coding `< COSTLY_ORDER` to enable the lru_cache,
    allow the lru_cache for larger orders as long as the folio contains
    fewer than `FOLIO_BATCH_LRU` pages. Also limit the total number of
    pages held in the lru_cache. Hugh may prefer this approach as well.
  * Clean up the comments and `if` conditions in "mm: improve large folio
    reuse for LRU-cached folios" based on David's feedback. Thanks!
  * Properly drain the lru_cache when splitting folios. Thanks to Sashiko
    and David!

-RFC v2:
 * Make __wp_can_reuse_large_anon_folio() aware of LRU-cached large
   folios. As David pointed out, it currently does not account for
   large folios residing in the per-CPU LRU cache.
   https://lore.kernel.org/all/20260709081536.82768-1-baohua@kernel.org/

Barry Song (Xiaomi) (4):
  mm: allow smaller large folios to use lru_cache
  mm: improve large folio reuse for LRU-cached folios
  mm: drain LRU cache if necessary for splitting large folios
  mm: batch lru_cache draining in deferred_split_scan

 include/linux/folio_batch.h | 25 +++++++++++++++++++++++++
 mm/folio.c                  | 10 +++++++++-
 mm/huge_memory.c            | 23 +++++++++++++++++++----
 mm/internal.h               |  4 ++--
 mm/memory.c                 |  6 ++++++
 5 files changed, 61 insertions(+), 7 deletions(-)

-- 
2.34.1
Re: [RFC PATCH v3 0/4] mm: enable lru cache for smaller large folios
Posted by Garg, Shivank 2 weeks, 3 days ago
On Wed, 2026-08-19 at 06:59 +0800, Barry Song (Xiaomi) wrote:
> This patchset enables the per-CPU LRU cache for large folios with fewer
> than `FOLIO_BATCH_SIZE` (31) pages. It also limits each per-CPU LRU cache
> to at most `FOLIO_BATCH_SIZE` pages to avoid negatively affecting
> accounting and memory reclamation pressure.
> 
> This is particularly beneficial on systems that use relatively small
> large folios. For larger folios, the benefit is likely to be smaller
> because far fewer folios are expected to contend for the LRU cache.
> 
> * Use the following microbenchmark:
> 
>  #include <pthread.h>
>  #include <sys/mman.h>
>  #include <string.h>
> 
>  #define NUM_THREADS     20
>  #define MEM_SIZE        (16 * 1024 * 1024)
>  #define LOOP_COUNT      1000
> 
>  void* thread_worker(void* arg) {
>         void *addr = mmap(NULL, MEM_SIZE, PROT_READ | PROT_WRITE,
>                         MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
> 
>         for (int i = 0; i < LOOP_COUNT; i++) {
>                 for(int j = 0; j < MEM_SIZE; j += 4096)
>                         *(unsigned char *)(addr + j) = 0x55;
>                 madvise(addr, MEM_SIZE, MADV_DONTNEED);
>         }
> 
>         munmap(addr, MEM_SIZE);
>         pthread_exit(NULL);
>  }
> 
>  int main() {
>         pthread_t threads[NUM_THREADS];
> 
>         for (long t = 0; t < NUM_THREADS; t++) {
>                 pthread_create(&threads[t], NULL, thread_worker, (void*)t);
>         }
> 
>         for (int t = 0; t < NUM_THREADS; t++) {
>                 pthread_join(threads[t], NULL);
>         }
> 
>         return 0;
>  }
> 
> w/o patch:
> 
> $ time ./a.out 
> 
> real	0m17.760s
> user	0m8.207s
> sys	5m46.067s
> 
> perf lock report:
>                 Name   acquired  contended     avg wait   total wait     max wait     min wait 
> 
>                        21414857   21414857     30.07 us     10.73 m     101.94 us       880 ns 
>                           42890      42890     11.29 us    484.21 ms    123.08 us       923 ns 
>                            1678       1678      1.57 us      2.63 ms      5.77 us       985 ns 
>                              52         52      1.46 us     75.98 us      2.69 us      1.07 us 
>                              18         18     17.82 us    320.69 us     58.92 us      1.62 us 
>                              18         18      2.36 ms     42.52 ms      5.31 ms      1.47 us 
>            rcu_state         12         12      1.98 us     23.78 us      2.62 us      1.47 us 
>            rcu_state          9          9      1.72 us     15.49 us      2.27 us      1.34 us 
>                               2          2      2.42 us      4.83 us      2.48 us      2.36 us 
> 
> w/ patch:
> 
> $ time ./a.out 
> 
> real	0m16.292s
> user	0m8.587s
> sys	5m13.787s
> 
> perf lock report
> 
>                 Name   acquired  contended     avg wait   total wait     max wait     min wait 
> 
>                         2641286    2641286     46.87 us      2.06 m     107.92 us      1.01 us 
>                          235275     235275     10.57 us      2.49 s     115.32 us      1.02 us 
>            rcu_state       1982       1982      7.58 us     15.02 ms     31.72 us      1.26 us 
>            rcu_state       1929       1929      7.42 us     14.32 ms     30.50 us      1.43 us 
>                              86         86      3.61 us    310.19 us     13.89 us      1.15 us 
>                              20         20     19.70 us    394.03 us    108.07 us      2.15 us 
>        tasklist_lock          1          1      2.13 us      2.13 us      2.13 us      2.13 us 
> 
> 
> * Build the kernel in a 1 GiB memcg by -j20 with zRAM configured as swap:
> 
> w/o patch:
> 
> Perf lock report:
> 
>                 Name   acquired  contended     avg wait   total wait     max wait     min wait 
> 
>                          782337     782337     17.61 us     13.78 s     402.22 us       896 ns 
>                           55459      55459     19.48 us      1.08 s     117.59 us      1.01 us 
>                            7826       7826      8.01 us     62.68 ms     19.03 us       887 ns 
>                            5324       5324      7.83 us     41.68 ms     37.59 us      1.11 us 
>            rcu_state       5144       5144      6.27 us     32.23 ms     25.17 us      1.58 us 
>            rcu_state       5142       5142      6.35 us     32.67 ms     30.55 us      1.48 us 
>                            3855       3855      8.74 us     33.68 ms     42.32 us       996 ns 
>                            2770       2770      9.22 us     25.55 ms     27.83 us       914 ns 
>                            2342       2342      5.88 us     13.77 ms    318.75 us      1.00 us 
> time:
> 
> *** Executing round 0 ***
> 
> real    1m46.847s
> user    25m10.848s
> sys     2m57.282s
> 
> *** Executing round 1 ***
> 
> real    1m46.423s
> user    25m10.072s
> sys     2m54.348s
> 
> *** Executing round 2 ***
> 
> real    1m46.308s
> user    25m13.800s
> sys     2m58.963s
> 
> *** Executing round 3 ***
> 
> real    1m46.155s
> user    25m18.079s
> sys     2m59.721s
> 
> *** Executing round 4 ***
> 
> real    1m45.980s
> user    25m15.493s
> sys     2m56.959s
> 
> w/ patch:
> 
> perf lock report
>                 Name   acquired  contended     avg wait   total wait     max wait     min wait 
> 
>                          202647     202647     34.27 us      6.94 s     467.82 us      1.18 us 
>                           51819      51819     16.26 us    842.46 ms    245.55 us       885 ns 
>      inode_hash_lock      15169      15169      8.58 us    130.21 ms     23.23 us      1.04 us 
>                            5306       5306      7.54 us     40.03 ms     31.53 us      1.17 us 
>            rcu_state       4945       4945      6.97 us     34.47 ms     30.00 us      1.76 us 
>            rcu_state       4899       4899      7.03 us     34.42 ms     24.65 us      1.42 us 
>                            3923       3923      8.56 us     33.57 ms     27.51 us      1.08 us 
>                            2212       2212      5.23 us     11.56 ms    222.79 us       965 ns 
>                            1412       1412      6.35 us      8.97 ms     23.09 us      1.77 us 
> 
> time:
> 
> *** Executing round 0 ***
> 
> real	1m46.463s
> user	25m17.448s
> sys	2m49.524s
> 
> *** Executing round 1 ***
> 
> real	1m46.274s
> user	25m12.178s
> sys	2m53.522s
> 
> *** Executing round 2 ***
> 
> real	1m46.362s
> user	25m13.115s
> sys	2m53.005s
> 
> *** Executing round 3 ***
> 
> real	1m46.036s
> user	25m17.627s
> sys	2m53.477s
> 
> *** Executing round 4 ***
> 
> real	1m46.329s
> user	25m15.130s
> sys	2m51.508s
> 
> -RFC v3:
>   * Rather than hard-coding `< COSTLY_ORDER` to enable the lru_cache,
>     allow the lru_cache for larger orders as long as the folio contains
>     fewer than `FOLIO_BATCH_LRU` pages. Also limit the total number of
>     pages held in the lru_cache. Hugh may prefer this approach as well.
>   * Clean up the comments and `if` conditions in "mm: improve large folio
>     reuse for LRU-cached folios" based on David's feedback. Thanks!
>   * Properly drain the lru_cache when splitting folios. Thanks to Sashiko
>     and David!
> 
> -RFC v2:
>  * Make __wp_can_reuse_large_anon_folio() aware of LRU-cached large
>    folios. As David pointed out, it currently does not account for
>    large folios residing in the per-CPU LRU cache.
>    https://lore.kernel.org/all/20260709081536.82768-1-baohua@kernel.org/
> 
> Barry Song (Xiaomi) (4):
>   mm: allow smaller large folios to use lru_cache
>   mm: improve large folio reuse for LRU-cached folios
>   mm: drain LRU cache if necessary for splitting large folios
>   mm: batch lru_cache draining in deferred_split_scan
> 
>  include/linux/folio_batch.h | 25 +++++++++++++++++++++++++
>  mm/folio.c                  | 10 +++++++++-
>  mm/huge_memory.c            | 23 +++++++++++++++++++----
>  mm/internal.h               |  4 ++--
>  mm/memory.c                 |  6 ++++++
>  5 files changed, 61 insertions(+), 7 deletions(-)


Thanks for this series.
I tested this against a Graph500 regression with 16K mTHP.
The series removes the regression and reduces folio_lruvec_lock_irqsave contention to approximately the 4K rate.


Tested-by: Shivank Garg <shivankg@amd.com>


System and Workload
===================

AMD EPYC Zen 5 System, 2 sockets, 160 cores/socket,
SMT on -> 320 cores / 640 threads
2 NUMA nodes, ~512 GB/node
node 0 CPUs: 0-159,320-479

Base: v7.3-rc2+ (50d05c7c76c9)
Runtime: performance governor, preempt=full (lazy)

Workload: Graph500, Scale 27, edgefactor 16
Metric: median_time in ms (lower is better)

Bind to node 0 with numactl -C 0-159,320-479 -m 0


Performance
===========

                  Base       Patched    delta
4K                109.60     109.11     -0.4%
mTHP-16K          144.10      99.00    -31.3%
mTHP-64K          104.75     104.25     -0.5%

mTHP-16K was ~31% slower than THP-never, due to lock contentions.
With your series, mTHP-16 performs ~9% faster than THP-never.


The dominant folio_lruvec_lock_irqsave entries were:

                   Base                  Patched
               contended  wait         contended  wait
4K               834603   2.41 m         844193   2.61 m
mTHP-16K        7008834   28.88 m        840695   4.87 m
mTHP-64K         188144   752 ms         40874    93.46 ms


Thanks,
Shivank
Re: [RFC PATCH v3 0/4] mm: enable lru cache for smaller large folios
Posted by Barry Song 2 weeks, 2 days ago
On Fri, Sep 11, 2026 at 11:39 PM Garg, Shivank <shivankg@amd.com> wrote:
[...]
>
> Thanks for this series.
> I tested this against a Graph500 regression with 16K mTHP.
> The series removes the regression and reduces folio_lruvec_lock_irqsave contention to approximately the 4K rate.
>
>
> Tested-by: Shivank Garg <shivankg@amd.com>
>

Thanks very much, Shivank. I’m very happy to see that
mTHP LRUCACHE not only resolves your regression but also
improves the performance of your case.

As discussed with David, this series will be respun after Hugh’s
work[1] is merged.

In the meantime, I’d really appreciate it if you could test an
updated version that I haven’t sent out before:

https://git.kernel.org/pub/scm/linux/kernel/git/baohua/linux.git/log/?h=mthp_lrucache

It includes some modifications based on Sashiko’s review comments
on the version you tested. I’d like to know whether it still provides
the same performance gains you saw with that version.

[1] https://lore.kernel.org/linux-mm/e28f9a94-4339-f8ac-8301-6be3c9b5b7ce@google.com/

>
> System and Workload
> ===================
>
> AMD EPYC Zen 5 System, 2 sockets, 160 cores/socket,
> SMT on -> 320 cores / 640 threads
> 2 NUMA nodes, ~512 GB/node
> node 0 CPUs: 0-159,320-479
>
> Base: v7.3-rc2+ (50d05c7c76c9)
> Runtime: performance governor, preempt=full (lazy)
>
> Workload: Graph500, Scale 27, edgefactor 16
> Metric: median_time in ms (lower is better)
>
> Bind to node 0 with numactl -C 0-159,320-479 -m 0
>
>
> Performance
> ===========
>
>                   Base       Patched    delta
> 4K                109.60     109.11     -0.4%
> mTHP-16K          144.10      99.00    -31.3%
> mTHP-64K          104.75     104.25     -0.5%
>
> mTHP-16K was ~31% slower than THP-never, due to lock contentions.
> With your series, mTHP-16 performs ~9% faster than THP-never.
>
>
> The dominant folio_lruvec_lock_irqsave entries were:
>
>                    Base                  Patched
>                contended  wait         contended  wait
> 4K               834603   2.41 m         844193   2.61 m
> mTHP-16K        7008834   28.88 m        840695   4.87 m
> mTHP-64K         188144   752 ms         40874    93.46 ms
>
>

Best Regards
Barry
Re: [RFC PATCH v3 0/4] mm: enable lru cache for smaller large folios
Posted by David Hildenbrand (Arm) 1 month, 1 week ago
On 8/19/26 00:59, Barry Song (Xiaomi) wrote:
> This patchset enables the per-CPU LRU cache for large folios with fewer
> than `FOLIO_BATCH_SIZE` (31) pages. It also limits each per-CPU LRU cache
> to at most `FOLIO_BATCH_SIZE` pages to avoid negatively affecting
> accounting and memory reclamation pressure.
> 
> This is particularly beneficial on systems that use relatively small
> large folios. For larger folios, the benefit is likely to be smaller
> because far fewer folios are expected to contend for the LRU cache.

As raised, there is this problem with collect_longterm_unpinnable_folios()

(see
https://lore.kernel.org/r/20260806-lru_cache_drain_for_folio-v1-1-c6287d295e99@kernel.org
)

whereby we don't know how many refs we actually hold. Certainly not 1.

We might have to wait for Hugh's cleanup to handle that cleanly (and avoid all
the other LRU cache draining).

-- 
Cheers,

David
Re: [RFC PATCH v3 0/4] mm: enable lru cache for smaller large folios
Posted by Barry Song 1 month, 1 week ago
On Fri, Aug 21, 2026 at 2:23 AM David Hildenbrand (Arm)
<david@kernel.org> wrote:
>
> On 8/19/26 00:59, Barry Song (Xiaomi) wrote:
> > This patchset enables the per-CPU LRU cache for large folios with fewer
> > than `FOLIO_BATCH_SIZE` (31) pages. It also limits each per-CPU LRU cache
> > to at most `FOLIO_BATCH_SIZE` pages to avoid negatively affecting
> > accounting and memory reclamation pressure.
> >
> > This is particularly beneficial on systems that use relatively small
> > large folios. For larger folios, the benefit is likely to be smaller
> > because far fewer folios are expected to contend for the LRU cache.
>
> As raised, there is this problem with collect_longterm_unpinnable_folios()
>
> (see
> https://lore.kernel.org/r/20260806-lru_cache_drain_for_folio-v1-1-c6287d295e99@kernel.org
> )
>
> whereby we don't know how many refs we actually hold. Certainly not 1.
>
> We might have to wait for Hugh's cleanup to handle that cleanly (and avoid all
> the other LRU cache draining).

Hi David,
Thanks for raising this.
Yes, I saw your comment and took a closer look at it. I think the
best approach for now is to leave that part untouched until Hugh's
patch lands? We might end up doing some extra draining in that case,
but that's safe—any value greater than 1 might not be. We may just
drain more than necessary?

Best Regards
Barry
Re: [RFC PATCH v3 0/4] mm: enable lru cache for smaller large folios
Posted by David Hildenbrand (Arm) 1 month, 1 week ago
On 8/20/26 22:18, Barry Song wrote:
> On Fri, Aug 21, 2026 at 2:23 AM David Hildenbrand (Arm)
> <david@kernel.org> wrote:
>>
>> On 8/19/26 00:59, Barry Song (Xiaomi) wrote:
>>> This patchset enables the per-CPU LRU cache for large folios with fewer
>>> than `FOLIO_BATCH_SIZE` (31) pages. It also limits each per-CPU LRU cache
>>> to at most `FOLIO_BATCH_SIZE` pages to avoid negatively affecting
>>> accounting and memory reclamation pressure.
>>>
>>> This is particularly beneficial on systems that use relatively small
>>> large folios. For larger folios, the benefit is likely to be smaller
>>> because far fewer folios are expected to contend for the LRU cache.
>>
>> As raised, there is this problem with collect_longterm_unpinnable_folios()
>>
>> (see
>> https://lore.kernel.org/r/20260806-lru_cache_drain_for_folio-v1-1-c6287d295e99@kernel.org
>> )
>>
>> whereby we don't know how many refs we actually hold. Certainly not 1.
>>
>> We might have to wait for Hugh's cleanup to handle that cleanly (and avoid all
>> the other LRU cache draining).
> 
> Hi David,
> Thanks for raising this.
> Yes, I saw your comment and took a closer look at it. I think the
> best approach for now is to leave that part untouched until Hugh's
> patch lands? 

I think we should not consider your patch set until Hugh either resolved it or
we have a better way to handle draining. Building on top of this for large
folios now just gets painful.

> We might end up doing some extra draining in that case,
> but that's safe—any value greater than 1 might not be. 

We'll unconditionally drain all LRU caches, which is precisely *not* what we
want to do in the first place. :)

-- 
Cheers,

David

Re: [RFC PATCH v3 0/4] mm: enable lru cache for smaller large folios
Posted by Barry Song 1 month, 1 week ago
On Fri, Aug 21, 2026 at 10:18 PM David Hildenbrand (Arm)
<david@kernel.org> wrote:
>
> On 8/20/26 22:18, Barry Song wrote:
> > On Fri, Aug 21, 2026 at 2:23 AM David Hildenbrand (Arm)
> > <david@kernel.org> wrote:
> >>
> >> On 8/19/26 00:59, Barry Song (Xiaomi) wrote:
> >>> This patchset enables the per-CPU LRU cache for large folios with fewer
> >>> than `FOLIO_BATCH_SIZE` (31) pages. It also limits each per-CPU LRU cache
> >>> to at most `FOLIO_BATCH_SIZE` pages to avoid negatively affecting
> >>> accounting and memory reclamation pressure.
> >>>
> >>> This is particularly beneficial on systems that use relatively small
> >>> large folios. For larger folios, the benefit is likely to be smaller
> >>> because far fewer folios are expected to contend for the LRU cache.
> >>
> >> As raised, there is this problem with collect_longterm_unpinnable_folios()
> >>
> >> (see
> >> https://lore.kernel.org/r/20260806-lru_cache_drain_for_folio-v1-1-c6287d295e99@kernel.org
> >> )
> >>
> >> whereby we don't know how many refs we actually hold. Certainly not 1.
> >>
> >> We might have to wait for Hugh's cleanup to handle that cleanly (and avoid all
> >> the other LRU cache draining).
> >
> > Hi David,
> > Thanks for raising this.
> > Yes, I saw your comment and took a closer look at it. I think the
> > best approach for now is to leave that part untouched until Hugh's
> > patch lands?
>
> I think we should not consider your patch set until Hugh either resolved it or
> we have a better way to handle draining. Building on top of this for large
> folios now just gets painful.

Sure, I’m perfectly fine with this. I swear I asked Hugh whether I
should suspend this job and wait for him, but he didn’t seem to suggest
suspending it. :-) However, I agree that it makes more sense to
prioritize Hugh’s job before this one.

>
> > We might end up doing some extra draining in that case,
> > but that's safe—any value greater than 1 might not be.
>
> We'll unconditionally drain all LRU caches, which is precisely *not* what we
> want to do in the first place. :)

Sure. Maybe because the test case is a kernel build, the
collect_longterm_unpinnable_folios() path is not really a hot path.
But I agree we can probably find some workloads where it is hot.

I’m perfectly fine with putting this large folio lru_cache work on hold
and working with Hugh on optimizing those drain paths first. :-)

Best Regards
Barry
Re: [RFC PATCH v3 0/4] mm: enable lru cache for smaller large folios
Posted by Barry Song 1 month, 1 week ago
On Fri, Aug 21, 2026 at 4:18 AM Barry Song <baohua@kernel.org> wrote:
>
> On Fri, Aug 21, 2026 at 2:23 AM David Hildenbrand (Arm)
> <david@kernel.org> wrote:
> >
> > On 8/19/26 00:59, Barry Song (Xiaomi) wrote:
> > > This patchset enables the per-CPU LRU cache for large folios with fewer
> > > than `FOLIO_BATCH_SIZE` (31) pages. It also limits each per-CPU LRU cache
> > > to at most `FOLIO_BATCH_SIZE` pages to avoid negatively affecting
> > > accounting and memory reclamation pressure.
> > >
> > > This is particularly beneficial on systems that use relatively small
> > > large folios. For larger folios, the benefit is likely to be smaller
> > > because far fewer folios are expected to contend for the LRU cache.
> >
> > As raised, there is this problem with collect_longterm_unpinnable_folios()
> >
> > (see
> > https://lore.kernel.org/r/20260806-lru_cache_drain_for_folio-v1-1-c6287d295e99@kernel.org
> > )
> >
> > whereby we don't know how many refs we actually hold. Certainly not 1.
> >
> > We might have to wait for Hugh's cleanup to handle that cleanly (and avoid all
> > the other LRU cache draining).
>
> Hi David,
> Thanks for raising this.
> Yes, I saw your comment and took a closer look at it. I think the
> best approach for now is to leave that part untouched until Hugh's
> patch lands? We might end up doing some extra draining in that case,
> but that's safe—any value greater than 1 might not be. We may just
> drain more than necessary?

BTW, David. This is also why I didn't move the below checks in patch 3/4[1]
to lru_cache_drain_for_folio() as we are not safe to do it for that "1" in
collect_longterm_unpinnable_folios():

+ if (folio_ref_count(folio) == folio_expected_ref_count(folio) + 1 +
+    folio_may_be_lru_cached(folio))
+ lru_cache_drain_for_folio(folio, 1, NULL);

[1] https://lore.kernel.org/all/20260818225904.55236-4-baohua@kernel.org/

>
> Best Regards
> Barry
Re: [RFC PATCH v3 0/4] mm: enable lru cache for smaller large folios
Posted by Lance Yang 1 month, 1 week ago
On Wed, Aug 19, 2026 at 06:59:00AM +0800, Barry Song (Xiaomi) wrote:
[...]
>
>-RFC v3:
>  * Rather than hard-coding `< COSTLY_ORDER` to enable the lru_cache,
>    allow the lru_cache for larger orders as long as the folio contains
>    fewer than `FOLIO_BATCH_LRU` pages. Also limit the total number of

Tiny typo ... should FOLIO_BATCH_LRU be FOLIO_BATCH_SIZE here?
Re: [RFC PATCH v3 0/4] mm: enable lru cache for smaller large folios
Posted by Barry Song 1 month, 1 week ago
On Wed, Aug 19, 2026 at 11:03 AM Lance Yang <lance.yang@linux.dev> wrote:
>
>
> On Wed, Aug 19, 2026 at 06:59:00AM +0800, Barry Song (Xiaomi) wrote:
> [...]
> >
> >-RFC v3:
> >  * Rather than hard-coding `< COSTLY_ORDER` to enable the lru_cache,
> >    allow the lru_cache for larger orders as long as the folio contains
> >    fewer than `FOLIO_BATCH_LRU` pages. Also limit the total number of
>
> Tiny typo ... should FOLIO_BATCH_LRU be FOLIO_BATCH_SIZE here?

Lance, thanks for the review. You’re right—it should be
FOLIO_BATCH_SIZE.

Best Regards
Barry