[PATCH v3 0/2] mm/swap: skip empty clusters in the swapoff scan

Youngjun Park posted 2 patches 1 month, 3 weeks ago
There is a newer version of this series
include/linux/swap.h |  2 +-
mm/swapfile.c        | 43 ++++++++++++++++++++++++++++++-------------
2 files changed, 31 insertions(+), 14 deletions(-)
[PATCH v3 0/2] mm/swap: skip empty clusters in the swapoff scan
Posted by Youngjun Park 1 month, 3 weeks ago
find_next_to_unuse() walks a swap device one offset at a time.  Slot
state now lives in a per cluster swap table, so patch 2 dismisses an
empty cluster with one counter read instead of SWAPFILE_CLUSTER table
reads.

Patch 1 is an unrelated one line comment fix noticed on the way.

A debug test confirmed the skip path runs, and swapoff completed
under load with no DEBUG_VM or lockdep splats.

Changes in v3:
- 2/2: clamp the scan end with min_t() so it stops at si->max, rather
  than running into the masked tail of the last cluster, which drops the
  need to explain why walking that tail was safe (Barry)
- 2/2: compute ci_off only where it is used
- 1/2, 2/2: pick up Barry's Reviewed-by
- Rebased on mm-new
- Link to v2: https://lore.kernel.org/r/20260805141146.127776-1-youngjun.park@lge.com

Changes in v2:
- 1/2: reword the comment to "array, one entry per cluster", dropping
  the redundant "on every device" (Barry)
- 1/2: pick up Kairui's Acked-by
- 2/2: drop the min(), the swap table is always SWAPFILE_CLUSTER entries
  and swapon() masks the tail past si->max as bad (Kairui)
- 2/2: mark the unlocked ci->count read with READ_ONCE() for KCSAN
  instead of cluster_is_empty(), whose other callers hold ci->lock
  (Kairui)
- 2/2: expand the commit message to cover both of the above
- Rebased on mm-new
- Link to v1: https://lore.kernel.org/r/20260728155907.391820-1-youngjun.park@lge.com

Youngjun Park (2):
  mm/swap: fix stale comment on swap_info_struct::cluster_info
  mm/swap: scan by cluster in find_next_to_unuse()

 include/linux/swap.h |  2 +-
 mm/swapfile.c        | 43 ++++++++++++++++++++++++++++++-------------
 2 files changed, 31 insertions(+), 14 deletions(-)


base-commit: 1fb556c523f6c18b43b1f52fb366f61c9963ce06
-- 
2.48.1
Re: [PATCH v3 0/2] mm/swap: skip empty clusters in the swapoff scan
Posted by Andrew Morton 1 month, 3 weeks ago
On Fri,  7 Aug 2026 04:32:26 +0900 Youngjun Park <youngjun.park@lge.com> wrote:

> find_next_to_unuse() walks a swap device one offset at a time.  Slot
> state now lives in a per cluster swap table, so patch 2 dismisses an
> empty cluster with one counter read instead of SWAPFILE_CLUSTER table
> reads.

Thanks.

Can you help us understand how significant this change is for users? 
If "not very" then I'd prefer to defer consideraton of the series until
after 7.3-rc1.
Re: [PATCH v3 0/2] mm/swap: skip empty clusters in the swapoff scan
Posted by Youngjun Park 1 month, 3 weeks ago
On Thu, Aug 06, 2026 at 01:06:55PM -0700, Andrew Morton wrote:
> On Fri,  7 Aug 2026 04:32:26 +0900 Youngjun Park <youngjun.park@lge.com> wrote:
> 
> > find_next_to_unuse() walks a swap device one offset at a time.  Slot
> > state now lives in a per cluster swap table, so patch 2 dismisses an
> > empty cluster with one counter read instead of SWAPFILE_CLUSTER table
> > reads.
> 
> Thanks.
> 
> Can you help us understand how significant this change is for users? 
> If "not very" then I'd prefer to defer consideraton of the series until
> after 7.3-rc1.

Hello Andrew

"Not very" in the common case, though there is a case where the win is clear.
No bug and no user report.

For now I would rather defer to after 7.3-rc1.

And for your reference, here is the details.

Every swapoff does a little less work now, because the scan steps over an
unused area one cluster at a time.
But IMHO most of the swapoff time goes to unuse_mm() and to reading the pages back in.

The gain shows on a large swap device that is almost empty, when the last
pages still in use are near the end of it.  The scan has to walk up to
them, and today it looks at every slot on the way.  Now the empty clusters
in between are skipped in one step.

I have no measured times yet, since that case has to be set up on purpose.
What I did is the arithmetic for the case that skips best,
For example 1T of swap with 256M slots with SWAPFILE_CLUSTER = 512
and everything free but the far end:

  - today:          256M table reads
  - with the skip:  512K counter reads

That should be around half a second of scan saved.

What I have checked is that empty clusters are skipped as intended, so the
scan does less work.  That work is a small part of swapoff, so depending on
the situation it may be too small to see in clock time.

Thanks,
Youngjun
Re: [PATCH v3 0/2] mm/swap: skip empty clusters in the swapoff scan
Posted by Baoquan He 1 month, 3 weeks ago
On 08/07/26 at 03:41pm, Youngjun Park wrote:
> On Thu, Aug 06, 2026 at 01:06:55PM -0700, Andrew Morton wrote:
> > On Fri,  7 Aug 2026 04:32:26 +0900 Youngjun Park <youngjun.park@lge.com> wrote:
> > 
> > > find_next_to_unuse() walks a swap device one offset at a time.  Slot
> > > state now lives in a per cluster swap table, so patch 2 dismisses an
> > > empty cluster with one counter read instead of SWAPFILE_CLUSTER table
> > > reads.
> > 
> > Thanks.
> > 
> > Can you help us understand how significant this change is for users? 
> > If "not very" then I'd prefer to defer consideraton of the series until
> > after 7.3-rc1.
> 
> Hello Andrew
> 
> "Not very" in the common case, though there is a case where the win is clear.
> No bug and no user report.
> 
> For now I would rather defer to after 7.3-rc1.
> 
> And for your reference, here is the details.
> 
> Every swapoff does a little less work now, because the scan steps over an
> unused area one cluster at a time.
> But IMHO most of the swapoff time goes to unuse_mm() and to reading the pages back in.
> 
> The gain shows on a large swap device that is almost empty, when the last
> pages still in use are near the end of it.  The scan has to walk up to
> them, and today it looks at every slot on the way.  Now the empty clusters
> in between are skipped in one step.
> 
> I have no measured times yet, since that case has to be set up on purpose.
> What I did is the arithmetic for the case that skips best,
> For example 1T of swap with 256M slots with SWAPFILE_CLUSTER = 512
> and everything free but the far end:
> 
>   - today:          256M table reads
>   - with the skip:  512K counter reads


Maybe just use time to measure swapoff time consuming, just like below
as I did on a kvm guest, I guess a bare metal machine with larger system
ram could be more obvious?

root@fedora:~# free -h
               total        used        free      shared  buff/cache   available
Mem:           3.8Gi       181Mi       3.6Gi       924Ki        72Mi       3.7Gi
Swap:          2.0Gi        18Mi       2.0Gi
root@fedora:~# time swapoff /dev/vdb

real	0m0.101s
user	0m0.001s
sys	0m0.017s
root@fedora:~# swapon /dev/vdb
root@fedora:~# time swapoff /dev/vdb

real	0m0.014s
user	0m0.003s
sys	0m0.001s

Not sure if Andrew is asking for this.


> 
> That should be around half a second of scan saved.

Yeah, a concrete number is shown.
Re: [PATCH v3 0/2] mm/swap: skip empty clusters in the swapoff scan
Posted by Andrew Morton 1 month, 3 weeks ago
On Fri, 7 Aug 2026 16:42:22 +0800 Baoquan He <baoquan.he@linux.dev> wrote:

> > I have no measured times yet, since that case has to be set up on purpose.
> > What I did is the arithmetic for the case that skips best,
> > For example 1T of swap with 256M slots with SWAPFILE_CLUSTER = 512
> > and everything free but the far end:
> > 
> >   - today:          256M table reads
> >   - with the skip:  512K counter reads
> 
> 
> Maybe just use time to measure swapoff time consuming, just like below
> as I did on a kvm guest, I guess a bare metal machine with larger system
> ram could be more obvious?
> 
> root@fedora:~# free -h
>                total        used        free      shared  buff/cache   available
> Mem:           3.8Gi       181Mi       3.6Gi       924Ki        72Mi       3.7Gi
> Swap:          2.0Gi        18Mi       2.0Gi
> root@fedora:~# time swapoff /dev/vdb
> 
> real	0m0.101s
> user	0m0.001s
> sys	0m0.017s
> root@fedora:~# swapon /dev/vdb
> root@fedora:~# time swapoff /dev/vdb
> 
> real	0m0.014s
> user	0m0.003s
> sys	0m0.001s
> 
> Not sure if Andrew is asking for this.

I think it's helpful to include such info.  The audience for changelogs
is more than swap developers!  It's also an MM maintainer and -stable
maintainers and other people who are all wondering "should I backport
this for my users".  Let's give them the means to determine that.
Re: [PATCH v3 0/2] mm/swap: skip empty clusters in the swapoff scan
Posted by Youngjun Park 1 month, 3 weeks ago
On Fri, Aug 07, 2026 at 02:42:56PM -0700, Andrew Morton wrote:
> On Fri, 7 Aug 2026 16:42:22 +0800 Baoquan He <baoquan.he@linux.dev> wrote:
> 
> > > I have no measured times yet, since that case has to be set up on purpose.
> > > What I did is the arithmetic for the case that skips best,
> > > For example 1T of swap with 256M slots with SWAPFILE_CLUSTER = 512
> > > and everything free but the far end:
> > > 
> > >   - today:          256M table reads
> > >   - with the skip:  512K counter reads
> > 
> > 
> > Maybe just use time to measure swapoff time consuming, just like below
> > as I did on a kvm guest, I guess a bare metal machine with larger system
> > ram could be more obvious?
> > 
> > root@fedora:~# free -h
> >                total        used        free      shared  buff/cache   available
> > Mem:           3.8Gi       181Mi       3.6Gi       924Ki        72Mi       3.7Gi
> > Swap:          2.0Gi        18Mi       2.0Gi
> > root@fedora:~# time swapoff /dev/vdb
> > 
> > real	0m0.101s
> > user	0m0.001s
> > sys	0m0.017s
> > root@fedora:~# swapon /dev/vdb
> > root@fedora:~# time swapoff /dev/vdb
> > 
> > real	0m0.014s
> > user	0m0.003s
> > sys	0m0.001s
> > 
> > Not sure if Andrew is asking for this.
> 
> I think it's helpful to include such info.  The audience for changelogs
> is more than swap developers!  It's also an MM maintainer and -stable
> maintainers and other people who are all wondering "should I backport
> this for my users".  Let's give them the means to determine that.

Hello Andrew Here is the experiment.

7.2-rc5, fill a 1 TiB swap up to some amount, then swapoff.  

What is left sits at the top of what was filled, so every slot below it is free.  
Those entries are in the swap cache and cheap to free, so the search is most of the time.  
In short, best scenario my patch shows improvement.

Medians over 11 pairs at 32 and 128 GiB, 3 pairs at 256 and 512.

      filled          swapoff 
                      old        new     
       32 GiB          92.4ms    66.3ms
      128 GiB         157.6ms    94.4ms
      256 GiB         209.7ms    63.8ms  
      512 GiB         391.8ms    73.4ms   

old grows with how much was filled, new does not.  

In the ordinary case swap still holds real data and swapoff spends its time
reading it back.  I tested 4 GiB on an 8 GiB device, where the scan is 1.4%
of try_to_unuse(), and there is no difference either way. meaning no regression

Thanks!
Youngjun
Re: [PATCH v3 0/2] mm/swap: skip empty clusters in the swapoff scan
Posted by Youngjun Park 1 month, 3 weeks ago
On Fri, Aug 07, 2026 at 04:42:22PM +0800, Baoquan He wrote:
> On 08/07/26 at 03:41pm, Youngjun Park wrote:
> > On Thu, Aug 06, 2026 at 01:06:55PM -0700, Andrew Morton wrote:
> > > On Fri,  7 Aug 2026 04:32:26 +0900 Youngjun Park <youngjun.park@lge.com> wrote:
> > > 
> > > > find_next_to_unuse() walks a swap device one offset at a time.  Slot
> > > > state now lives in a per cluster swap table, so patch 2 dismisses an
> > > > empty cluster with one counter read instead of SWAPFILE_CLUSTER table
> > > > reads.
> > > 
> > > Thanks.
> > > 
> > > Can you help us understand how significant this change is for users? 
> > > If "not very" then I'd prefer to defer consideraton of the series until
> > > after 7.3-rc1.
> > 
> > Hello Andrew
> > 
> > "Not very" in the common case, though there is a case where the win is clear.
> > No bug and no user report.
> > 
> > For now I would rather defer to after 7.3-rc1.
> > 
> > And for your reference, here is the details.
> > 
> > Every swapoff does a little less work now, because the scan steps over an
> > unused area one cluster at a time.
> > But IMHO most of the swapoff time goes to unuse_mm() and to reading the pages back in.
> > 
> > The gain shows on a large swap device that is almost empty, when the last
> > pages still in use are near the end of it.  The scan has to walk up to
> > them, and today it looks at every slot on the way.  Now the empty clusters
> > in between are skipped in one step.
> > 
> > I have no measured times yet, since that case has to be set up on purpose.
> > What I did is the arithmetic for the case that skips best,
> > For example 1T of swap with 256M slots with SWAPFILE_CLUSTER = 512
> > and everything free but the far end:
> > 
> >   - today:          256M table reads
> >   - with the skip:  512K counter reads
> 
> 
> Maybe just use time to measure swapoff time consuming, just like below
> as I did on a kvm guest, I guess a bare metal machine with larger system
> ram could be more obvious?
> 
> root@fedora:~# free -h
>                total        used        free      shared  buff/cache   available
> Mem:           3.8Gi       181Mi       3.6Gi       924Ki        72Mi       3.7Gi
> Swap:          2.0Gi        18Mi       2.0Gi
> root@fedora:~# time swapoff /dev/vdb
> 
> real	0m0.101s
> user	0m0.001s
> sys	0m0.017s
> root@fedora:~# swapon /dev/vdb
> root@fedora:~# time swapoff /dev/vdb
> 
> real	0m0.014s
> user	0m0.003s
> sys	0m0.001s
> 
> Not sure if Andrew is asking for this.

Thank you for looking into it.

Yes, it is possible to show this. However, it cannot provide the saturation
time, as that depends on the swap device size and swap slot distribution.

Since a naive swapoff test cannot provide stable evidence, I addressed the
approximate skip time (and also checked whether a skip occurred) using a logical time
calculation as shown below.

I try to take some time to measure the swapoff time under experiment conditions.
I will follow up this in the next patch iteration or sooner in this thread
(depending on Andrew's decision).

> > 
> > That should be around half a second of scan saved.
> 
> Yeah, a concrete number is shown.

Right, based on time complexity arithmetic.

Youngjun
Re: [PATCH v3 0/2] mm/swap: skip empty clusters in the swapoff scan
Posted by Youngjun Park 1 month, 3 weeks ago
On Fri, Aug 07, 2026 at 03:41:22PM +0900, Youngjun Park wrote:
> On Thu, Aug 06, 2026 at 01:06:55PM -0700, Andrew Morton wrote:
> > On Fri,  7 Aug 2026 04:32:26 +0900 Youngjun Park <youngjun.park@lge.com> wrote:
> > 
> > > find_next_to_unuse() walks a swap device one offset at a time.  Slot
> > > state now lives in a per cluster swap table, so patch 2 dismisses an
> > > empty cluster with one counter read instead of SWAPFILE_CLUSTER table
> > > reads.
> > 
> > Thanks.
> > 
> > Can you help us understand how significant this change is for users? 
> > If "not very" then I'd prefer to defer consideraton of the series until
> > after 7.3-rc1.
> 
> Hello Andrew
> 
> "Not very" in the common case, though there is a case where the win is clear.
> No bug and no user report.

Something I forgot to mention,

This only affects swapoff.
and few users run swapoff often, so the impact is limited either way.

> For now I would rather defer to after 7.3-rc1.
> 
> And for your reference, here is the details.
> 
> Every swapoff does a little less work now, because the scan steps over an
> unused area one cluster at a time.
> But IMHO most of the swapoff time goes to unuse_mm() and to reading the pages back in.
> 
> The gain shows on a large swap device that is almost empty, when the last
> pages still in use are near the end of it.  The scan has to walk up to
> them, and today it looks at every slot on the way.  Now the empty clusters
> in between are skipped in one step.
> 
> I have no measured times yet, since that case has to be set up on purpose.
> What I did is the arithmetic for the case that skips best,
> For example 1T of swap with 256M slots with SWAPFILE_CLUSTER = 512
> and everything free but the far end:
> 
>   - today:          256M table reads
>   - with the skip:  512K counter reads
> 
> That should be around half a second of scan saved.

+ benefit.

Reclaim can put a cached folio back into a page table with the slot it
already had, so try_to_unuse() retries and the scan starts over.  The
saving then applies once per pass.