[PATCH v2 0/4] mm, swap: keep hibernation swap slots out of the swap cache

Youngjun Park posted 4 patches 1 month, 2 weeks ago
There is a newer version of this series
mm/swap_state.c |  9 +++++++--
mm/swap_table.h | 13 +++++++++++++
mm/swapfile.c   | 16 +++++++---------
3 files changed, 27 insertions(+), 11 deletions(-)
[PATCH v2 0/4] mm, swap: keep hibernation swap slots out of the swap cache
Posted by Youngjun Park 1 month, 2 weeks ago
Cluster readahead walks a raw page_cluster sized window of offsets around
the faulting entry.  A hibernation slot looks like an ordinary swapped out
slot, so __swap_cache_add_check() lets it in.  Readahead reads the offset
off the device into a folio and puts that folio in the swap table where the
hibernation entry was.  This has been possible for a long time.  It only
wasted a folio and a read.

That changed with commit 0d6af9bcf383 ("mm, swap: use the swap table to
track the swap count").  A slot with a folio in the swap cache should only
be freed when the folio leaves the cache.  swap_put_entries_cluster() still
does that, but the conversion left swap_free_hibernation_slot() freeing the
slot either way.  Nothing points at the folio after that, and when reclaim
drops it later, it writes to the table entry at the old offset, which
someone else may own by then.

Patch 1 is the fix and the only patch for stable.  It puts the missing
check back, so both free paths behave the same again.

The rest removes the cause.  Readahead should not touch these slots at all,
so patch 2 lets only swapped out slots into the swap cache, which also puts
back a bad slot check the swap cache rework dropped, patch 3 gives
hibernation slots their own swap table entry type so that check covers them
too, and patch 4 drops the guard and the reclaim, since no such folio can
exist any more.

For any of this a task has to hold hibernation slots while the system is
still running.  The in kernel path does not, it allocates, writes and frees
the slots with everything frozen.  Userspace hibernation is different.  The
process writing the image is not frozen, and SNAPSHOT_ALLOC_SWAP_PAGE does
not check that anything is frozen.  The swap device must also not be
SWP_SYNCHRONOUS_IO, or swapin takes the direct path and never reaches
cluster readahead.

Tested with a debug patch generated by AI that counts hibernation slots
through the swap cache paths.  virtio-blk swap, page-cluster 3,
SNAPSHOT_ALLOC_SWAP_PAGE interleaved with MADV_PAGEOUT of a shmem region
so the hibernation slots land in the readahead windows.

                                unpatched   +patch 1   patches 1-4
  hibernation slots allocated       2732       2732         2732
  readahead landed on the slot      2731       2731         2731
  admitted to the swap cache        2731       2731            0
  reclaim found the folio              0       2731            -
  slot left unfreed                    0          0            0
  VM_WARN in the free path          2731          0            0

The VM_WARN is the existing assertion in __swap_cluster_free_entries(),
not something the debug patch adds.

v1: https://lore.kernel.org/linux-mm/20260806190636.446205-1-youngjun.park@lge.com/

Changes since v1:

- Cc stable and Kairui's ack on patch 1
- Reordered, the swap cache check is now patch 2, so the hibernation type
  is turned away from the commit that adds it
- Patch 2 gained a Fixes tag, the bad slot half is a regression.  No
  stable Cc, that commit is only in v7.2-rc
- SWP_TB_HIB keeps its count field at 0 (Kairui)

Youngjun Park (4):
  mm, swap: don't free a hibernation slot that is in the swap cache
  mm, swap: give hibernation swap slots their own swap table entry type
  mm, swap: only allow swapped-out slots into the swap cache
  mm, swap: drop the swap cache guard and reclaim in
    swap_free_hibernation_slot()

 mm/swap_state.c |  9 +++++++--
 mm/swap_table.h | 13 +++++++++++++
 mm/swapfile.c   | 16 +++++++---------
 3 files changed, 27 insertions(+), 11 deletions(-)


base-commit: 0b53bff4fa05ff0d3ffbd3d3bb10fae69dfab498
-- 
2.48.1
Re: [PATCH v2 0/4] mm, swap: keep hibernation swap slots out of the swap cache
Posted by Andrew Morton 1 month, 2 weeks ago
On Sun,  9 Aug 2026 23:45:55 +0900 Youngjun Park <youngjun.park@lge.com> wrote:

> Cluster readahead walks a raw page_cluster sized window of offsets around
> the faulting entry.  A hibernation slot looks like an ordinary swapped out
> slot, so __swap_cache_add_check() lets it in.  Readahead reads the offset
> off the device into a folio and puts that folio in the swap table where the
> hibernation entry was.  This has been possible for a long time.  It only
> wasted a folio and a read.
> 
> That changed with commit 0d6af9bcf383 ("mm, swap: use the swap table to
> track the swap count").  A slot with a folio in the swap cache should only
> be freed when the folio leaves the cache.  swap_put_entries_cluster() still
> does that, but the conversion left swap_free_hibernation_slot() freeing the
> slot either way.  Nothing points at the folio after that, and when reclaim
> drops it later, it writes to the table entry at the old offset, which
> someone else may own by then.
> 
> Patch 1 is the fix and the only patch for stable.  It puts the missing
> check back, so both free paths behave the same again.
> 
> The rest removes the cause.  Readahead should not touch these slots at all,
> so patch 2 lets only swapped out slots into the swap cache, which also puts
> back a bad slot check the swap cache rework dropped, patch 3 gives
> hibernation slots their own swap table entry type so that check covers them
> too, and patch 4 drops the guard and the reclaim, since no such folio can
> exist any more.

Thanks.  It's late and review is only partial.  I'd prefer to wait
until after 7.3-rc1 to process this series.

AI review suggests that more such fixing is needed in
__swap_cache_add_check():
	https://sashiko.dev/#/patchset/20260809144559.2104856-1-youngjun.park@lge.com
Re: [PATCH v2 0/4] mm, swap: keep hibernation swap slots out of the swap cache
Posted by Youngjun Park 1 month, 2 weeks ago
On Mon, Aug 10, 2026 at 03:16:42PM -0700, Andrew Morton wrote:
> On Sun,  9 Aug 2026 23:45:55 +0900 Youngjun Park <youngjun.park@lge.com> wrote:
> 
> > Cluster readahead walks a raw page_cluster sized window of offsets around
> > the faulting entry.  A hibernation slot looks like an ordinary swapped out
> > slot, so __swap_cache_add_check() lets it in.  Readahead reads the offset
> > off the device into a folio and puts that folio in the swap table where the
> > hibernation entry was.  This has been possible for a long time.  It only
> > wasted a folio and a read.
> > 
> > That changed with commit 0d6af9bcf383 ("mm, swap: use the swap table to
> > track the swap count").  A slot with a folio in the swap cache should only
> > be freed when the folio leaves the cache.  swap_put_entries_cluster() still
> > does that, but the conversion left swap_free_hibernation_slot() freeing the
> > slot either way.  Nothing points at the folio after that, and when reclaim
> > drops it later, it writes to the table entry at the old offset, which
> > someone else may own by then.
> > 
> > Patch 1 is the fix and the only patch for stable.  It puts the missing
> > check back, so both free paths behave the same again.
> > 
> > The rest removes the cause.  Readahead should not touch these slots at all,
> > so patch 2 lets only swapped out slots into the swap cache, which also puts
> > back a bad slot check the swap cache rework dropped, patch 3 gives
> > hibernation slots their own swap table entry type so that check covers them
> > too, and patch 4 drops the guard and the reclaim, since no such folio can
> > exist any more.
> 
> Thanks.  It's late and review is only partial.  I'd prefer to wait
> until after 7.3-rc1 to process this series.
> 
> AI review suggests that more such fixing is needed in
> __swap_cache_add_check():
> 	https://sashiko.dev/#/patchset/20260809144559.2104856-1-youngjun.park@lge.com

False positive. But, more fix needed by different reason.

The walk the AI review points at is the large folio swapin path, where
swapin pulls the neighboring slots in together with the faulting one.  A
bad slot cannot show up there.  The range comes from present swap
entries in the page table or the shmem mapping, so every slot in it was
handed out by the allocator, and bad slots never are.

The spot is still worth fixing.  The range is not pinned between the
locked check and the locked recheck, a concurrent free can empty a slot
in it, and hibernation can claim the emptied slot from the allocator.
The raw count read in the walk then trips the countable assertion under
DEBUG_VM, even though the -EBUSY fallback handles the changed range
fine.

I sent v3 with the shadow requirement of patch 2 extended to that routine.

Thanks!
Youngjun