[PATCH 0/8] Optimize anonymous swapbacked large folio unmapping

Dev Jain posted 8 patches 1 day, 14 hours ago
include/linux/rmap.h |  51 +++++++++++++-------
mm/internal.h        |  79 +++++++++++++++++++++++++------
mm/memory.c          |  20 +++-----
mm/mprotect.c        |  17 -------
mm/rmap.c            | 109 +++++++++++++++++++++++++++++++------------
mm/shmem.c           |   8 ++--
mm/swap.h            |  35 ++++++++++++--
mm/swapfile.c        |  31 ++++++------
8 files changed, 235 insertions(+), 115 deletions(-)
[PATCH 0/8] Optimize anonymous swapbacked large folio unmapping
Posted by Dev Jain 1 day, 14 hours ago
Speed up unmapping of anonymous swapbacked large folios by clearing
the ptes, and setting swap ptes, in one go.

The following benchmark (stolen from Barry) is used to measure the
time taken to swapout 256M worth of memory backed by 64K large folios:

 #define _GNU_SOURCE
 #include <stdio.h>
 #include <stdlib.h>
 #include <sys/mman.h>
 #include <string.h>
 #include <time.h>
 #include <unistd.h>
 #include <errno.h>

 #define SIZE_MB 256
 #define SIZE_BYTES (SIZE_MB * 1024 * 1024)

 int main() {
     void *addr = mmap(NULL, SIZE_BYTES, PROT_READ | PROT_WRITE,
                       MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
     if (addr == MAP_FAILED) {
         perror("mmap failed");
         return 1;
     }

     memset(addr, 0, SIZE_BYTES);

     struct timespec start, end;
     clock_gettime(CLOCK_MONOTONIC, &start);

     if (madvise(addr, SIZE_BYTES, MADV_PAGEOUT) != 0) {
         perror("madvise(MADV_PAGEOUT) failed");
         munmap(addr, SIZE_BYTES);
         return 1;
     }

     clock_gettime(CLOCK_MONOTONIC, &end);

     long duration_ns = (end.tv_sec - start.tv_sec) * 1e9 +
                        (end.tv_nsec - start.tv_nsec);
     printf("madvise(MADV_PAGEOUT) took %ld ns (%.3f ms)\n",
            duration_ns, duration_ns / 1e6);

     munmap(addr, SIZE_BYTES);
     return 0;
 }

Performance as measured on a Linux VM on Apple M3 (arm64):

Vanilla - Mean: 37401913 ns, std dev: 12%
Patched - Mean: 17420282 ns, std dev: 11%

resulting in more than 2x speedup.

No regression observed on 4K folios.

Performance as measured on bare metal x86:

Vanilla - mean: 54986286 ns, std dev: 1.5%
Patched - mean: 51930795 ns, std dev: 3%

I tried magnifying the difference on x86 by using 1M large folios, but
can't spot an obvious improvement (looks like my system is too fast to
benefit from batched atomic operations!), hinting that the benefit lies
mainly in the reduction of ptep_get() calls and the reduction of TLB
flushes during contpte-unfolding, on arm64.

No regression is observed on 4K folios on x86 too.

---
Applies on mm-new (3d18f3499c48). mm-selftests pass.

This breakout patchset is the final completion of
https://lore.kernel.org/all/20260526063635.61721-1-dev.jain@arm.com/

The extra addition is the generic set_softleaf_ptes() helper.

Dev Jain (8):
  mm/swapfile: add batched version of folio_dup_swap
  mm/swapfile: add batched version of folio_put_swap
  mm/rmap: mm/rmap: Add batched version of folio_try_share_anon_rmap_pte
  mm/internal: rename swap offset helpers to softleaf offset
  mm/internal: add set_softleaf_ptes
  mm/memory: use set_softleaf_ptes for uffd-wp markers
  mm: move anon-exclusive batch helper to internal.h
  mm/rmap: batch unmap anonymous swap-backed large folios

 include/linux/rmap.h |  51 +++++++++++++-------
 mm/internal.h        |  79 +++++++++++++++++++++++++------
 mm/memory.c          |  20 +++-----
 mm/mprotect.c        |  17 -------
 mm/rmap.c            | 109 +++++++++++++++++++++++++++++++------------
 mm/shmem.c           |   8 ++--
 mm/swap.h            |  35 ++++++++++++--
 mm/swapfile.c        |  31 ++++++------
 8 files changed, 235 insertions(+), 115 deletions(-)

-- 
2.43.0