mm/swap_state.c | 11 ++++++++--- mm/swap_table.h | 13 +++++++++++++ mm/swapfile.c | 16 +++++++--------- 3 files changed, 28 insertions(+), 12 deletions(-)
Cluster readahead walks a raw page_cluster sized window of offsets around
the faulting entry. A hibernation slot looks like an ordinary swapped out
slot, so __swap_cache_add_check() lets it in. Readahead reads the offset
off the device into a folio and puts that folio in the swap table where the
hibernation entry was. This has been possible for a long time. It only
wasted a folio and a read.
That changed with commit 0d6af9bcf383 ("mm, swap: use the swap table to
track the swap count"). A slot with a folio in the swap cache should only
be freed when the folio leaves the cache. swap_put_entries_cluster() still
does that, but the conversion left swap_free_hibernation_slot() freeing the
slot either way. Nothing points at the folio after that, and when reclaim
drops it later, it writes to the table entry at the old offset, which
someone else may own by then.
Patch 1 is the fix and the only patch for stable. It puts the missing
check back, so both free paths behave the same again.
The rest removes the cause. Readahead should not touch these slots at all,
so patch 2 lets only swapped out slots into the swap cache, which also puts
back a bad slot check the swap cache rework dropped, patch 3 gives
hibernation slots their own swap table entry type so that check covers them
too, and patch 4 drops the guard and the reclaim, since no such folio can
exist any more.
For any of this a task has to hold hibernation slots while the system is
still running. The in kernel path does not, it allocates, writes and frees
the slots with everything frozen. Userspace hibernation is different. The
process writing the image is not frozen, and SNAPSHOT_ALLOC_SWAP_PAGE does
not check that anything is frozen. The swap device must also not be
SWP_SYNCHRONOUS_IO, or swapin takes the direct path and never reaches
cluster readahead.
Tested with a debug patch generated by AI that counts hibernation slots
through the swap cache paths. virtio-blk swap, page-cluster 3,
SNAPSHOT_ALLOC_SWAP_PAGE interleaved with MADV_PAGEOUT of a shmem region
so the hibernation slots land in the readahead windows.
unpatched +patch 1 patches 1-4
hibernation slots allocated 2732 2732 2732
readahead landed on the slot 2731 2731 2731
admitted to the swap cache 2731 2731 0
reclaim found the folio 0 2731 -
slot left unfreed 0 0 0
VM_WARN in the free path 2731 0 0
The VM_WARN is the existing assertion in __swap_cluster_free_entries(),
not something the debug patch adds.
v2: https://lore.kernel.org/linux-mm/20260809144559.2104856-1-youngjun.park@lge.com/
v1: https://lore.kernel.org/linux-mm/20260806190636.446205-1-youngjun.park@lge.com/
Changes since v2:
- Give the large folio walk in __swap_cache_add_check() the same shadow
test, pointed at by the AI review Andrew linked. No bad slot can be
in that range, but a slot freed and then taken by hibernation could
trip the countable assertion there
- Picked up Kairui's ack on patch 2, given before the walk change
- Rebased onto mm-new
Youngjun Park (4):
mm, swap: don't free a hibernation slot that is in the swap cache
mm, swap: only allow swapped-out slots into the swap cache
mm, swap: give hibernation swap slots their own swap table entry type
mm, swap: drop the swap cache guard and reclaim in
swap_free_hibernation_slot()
mm/swap_state.c | 11 ++++++++---
mm/swap_table.h | 13 +++++++++++++
mm/swapfile.c | 16 +++++++---------
3 files changed, 28 insertions(+), 12 deletions(-)
base-commit: 480a31230b426efb005b6e71a14ef80f405f18b6
--
2.48.1
On Tue, 11 Aug 2026 22:22:05 +0900 Youngjun Park <youngjun.park@lge.com> wrote:
> Cluster readahead walks a raw page_cluster sized window of offsets around
> the faulting entry. A hibernation slot looks like an ordinary swapped out
> slot, so __swap_cache_add_check() lets it in. Readahead reads the offset
> off the device into a folio and puts that folio in the swap table where the
> hibernation entry was. This has been possible for a long time. It only
> wasted a folio and a read.
>
> That changed with commit 0d6af9bcf383 ("mm, swap: use the swap table to
> track the swap count"). A slot with a folio in the swap cache should only
> be freed when the folio leaves the cache. swap_put_entries_cluster() still
> does that, but the conversion left swap_free_hibernation_slot() freeing the
> slot either way. Nothing points at the folio after that, and when reclaim
> drops it later, it writes to the table entry at the old offset, which
> someone else may own by then.
>
> Patch 1 is the fix and the only patch for stable. It puts the missing
> check back, so both free paths behave the same again.
Thanks.
When fixing things, please always take care to describe the
userspace-visible runtime effects of the bug, particularly when
proposing a -stable backport.
For [1/4] Gemini tells me "At a high level, this bug triggers silent
memory corruption, process crashes, or data instability across
completely unrelated userspace applications - typically occurring after a
system resumes from hibernation (suspend-to-disk)." Which is what I
figured too.
Do we have any reports of this? Reported-by/Closes?
I'd like to grab [1/4] only, and defer the other three until 7.3-rc1.
This might be mistaken, but from a quick read, it's not clear what
benefit those three patches offer our users.
And there's value in merging the backportable fix alone, to avoid the
risk that the other three patches accidentally fix misbehavior in
[1/4].
On Tue, Aug 11, 2026 at 11:46:22AM -0700, Andrew Morton wrote: > When fixing things, please always take care to describe the > userspace-visible runtime effects of the bug, particularly when > proposing a -stable backport. Thank you for the advice and sorry for it. I will make sure future patches describe the userspace-visible effects clearly! > For [1/4] Gemini tells me "At a high level, this bug triggers silent > memory corruption, process crashes, or data instability across > completely unrelated userspace applications - typically occurring after a > system resumes from hibernation (suspend-to-disk)." Which is what I > figured too. That is right, with one small correction. this can happen with uswsusp while the hibernation image is being created, not at resume time. Memory corruption, process crashes, or data instability across completely unrelated userspace applications can occur there. > Do we have any reports of this? Reported-by/Closes? I found this while working on giving hibernation slots their own marker in the swap table, which I had discussed with Kairui. (https://lore.kernel.org/linux-mm/abp7aDgYLrxF3Me8@KASONG-MC4/) As far as I know there are no reports, so there is no Reported-by/Closes to add. > I'd like to grab [1/4] only, and defer the other three until 7.3-rc1. > This might be mistaken, but from a quick read, it's not clear what > benefit those three patches offer our users. Simply put, they are close to cleanups that remove unneeded work. (with a some little optimization.) Patch 2 keeps bad slots out of the swap cache, so readahead no longer wastes a folio on them. Patch 3 keeps hibernation slots out of the swap cache, with the same effect. no wasted folio at readahead time. Patch 4 removes dead code. Thanks, Youngjun Park
On Wed, 12 Aug 2026 21:13:37 +0900 Youngjun Park <youngjun.park@lge.com> wrote:
> On Tue, Aug 11, 2026 at 11:46:22AM -0700, Andrew Morton wrote:
> > When fixing things, please always take care to describe the
> > userspace-visible runtime effects of the bug, particularly when
> > proposing a -stable backport.
>
> Thank you for the advice and sorry for it.
> I will make sure future patches describe the userspace-visible effects clearly!
It's common, trust me ;)
> > For [1/4] Gemini tells me "At a high level, this bug triggers silent
> > memory corruption, process crashes, or data instability across
> > completely unrelated userspace applications - typically occurring after a
> > system resumes from hibernation (suspend-to-disk)." Which is what I
> > figured too.
>
> That is right, with one small correction. this can happen with
> uswsusp while the hibernation image is being created, not at resume
> time. Memory corruption, process crashes, or data instability across
> completely unrelated userspace applications can occur there.
OK.
> > Do we have any reports of this? Reported-by/Closes?
>
> I found this while working on giving hibernation slots their own
> marker in the swap table, which I had discussed with Kairui.
> (https://lore.kernel.org/linux-mm/abp7aDgYLrxF3Me8@KASONG-MC4/)
> As far as I know there are no reports, so there is no Reported-by/Closes to
> add.
OK.
> > I'd like to grab [1/4] only, and defer the other three until 7.3-rc1.
> > This might be mistaken, but from a quick read, it's not clear what
> > benefit those three patches offer our users.
>
> Simply put, they are close to cleanups that remove unneeded work.
> (with a some little optimization.)
>
> Patch 2 keeps bad slots out of the swap cache, so readahead no
> longer wastes a folio on them.
>
> Patch 3 keeps hibernation slots out of the swap cache, with the same
> effect. no wasted folio at readahead time.
>
> Patch 4 removes dead code.
So I'll retain [1/4] with the below changelog. Please plan on
sending out the other patches after 7.3-rc1.
From: Youngjun Park <youngjun.park@lge.com>
Subject: mm, swap: don't free a hibernation slot that is in the swap cache
Date: Tue, 11 Aug 2026 22:22:06 +0900
A slot with a folio in the swap cache is freed when the folio leaves the
cache, not when its count drops. swap_put_entries_cluster() follows that
rule. swap_free_hibernation_slot() does not, it calls
__swap_cluster_free_entries() whether or not a folio sits on the slot.
Cluster readahead can put one there. It walks a raw page_cluster sized
window of offsets around the faulting entry, and a hibernation slot passes
__swap_cache_add_check() because it is not a folio and its count is not
zero. Freeing the slot then clears the entry under that folio.
The folio is now unreachable from the swap table, and the offset goes back
to the allocator. The folio is still on the LRU though, so reclaim can
pick it up later. It then takes the old offset out of folio->swap and
overwrites the table entry there, which by then may belong to someone
else.
This bug can trigger silent memory corruption, process crashes, or data
instability across completely unrelated userspace applications - typically
occurring when uswsusp is preparing the hibernation image.
I found this while working on giving hibernation slots their own marker in
the swap table, which I had discussed with Kairui.
(https://lore.kernel.org/linux-mm/abp7aDgYLrxF3Me8@KASONG-MC4/) As far as
I know there are no reports, so there is no Reported-by/Closes to add.
Check for a cached folio before freeing. The slot is then left in the
ordinary state where only the swap cache holds it, and it is freed when
the folio leaves the cache, either through the reclaim below or through
normal reclaim later.
Link: https://lore.kernel.org/20260811132209.2862708-2-youngjun.park@lge.com
Fixes: 0d6af9bcf383 ("mm, swap: use the swap table to track the swap count")
Signed-off-by: Youngjun Park <youngjun.park@lge.com>
Acked-by: Kairui Song <kasong@tencent.com>
Cc: Baoquan He <baoquan.he@linux.dev>
Cc: Barry Song <baohua@kernel.org>
Cc: Chris Li <chrisl@kernel.org>
Cc: Kemeng Shi <shikemeng@huaweicloud.com>
Cc: Nhat Pham <nphamcs@gmail.com>
Cc: <stable@vger.kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
---
mm/swapfile.c | 9 ++++++++-
1 file changed, 8 insertions(+), 1 deletion(-)
--- a/mm/swapfile.c~mm-swap-dont-free-a-hibernation-slot-that-is-in-the-swap-cache
+++ a/mm/swapfile.c
@@ -2196,7 +2196,14 @@ void swap_free_hibernation_slot(swp_entr
ci = swap_cluster_lock(si, offset);
__swap_cluster_put_entry(ci, offset % SWAPFILE_CLUSTER);
- __swap_cluster_free_entries(si, ci, offset % SWAPFILE_CLUSTER, 1);
+ /*
+ * A slot with a folio in the swap cache is freed when the folio
+ * leaves the cache, the same rule swap_put_entries_cluster() follows.
+ * Readahead can put a folio here, and freeing the slot now would
+ * leave that folio with no entry behind it.
+ */
+ if (!swp_tb_is_folio(__swap_table_get(ci, offset % SWAPFILE_CLUSTER)))
+ __swap_cluster_free_entries(si, ci, offset % SWAPFILE_CLUSTER, 1);
swap_cluster_unlock(ci);
/* In theory readahead might add it to the swap cache by accident */
_
© 2016 - 2026 Red Hat, Inc.