[PATCH v2 0/3] mm, swap: don't spin or flood the console on a bad swap entry

Breno Leitao posted 3 patches 1 month, 2 weeks ago
mm/memory.c      |  9 +++++++--
mm/mincore.c     |  2 +-
mm/shmem.c       |  2 +-
mm/swap_state.c  |  4 ++--
mm/swapfile.c    | 20 ++++++++++++--------
mm/userfaultfd.c |  3 ++-
mm/zswap.c       |  2 +-
7 files changed, 26 insertions(+), 16 deletions(-)
[PATCH v2 0/3] mm, swap: don't spin or flood the console on a bad swap entry
Posted by Breno Leitao 1 month, 2 weeks ago
I've seen some machines at Meta fleet that show the following type of
problem:

1) It gets some weird warning:

  BUG: Bad page map in process khugepaged  pte:f000eef300000017 pmd:00000067
  addr:00007f57c0a01000 vm_flags:20200073 anon_vma:ffff88829af7c340 mapping:0000000000000000 index:7f57c0a01

The corruption is most likely the collapse/PT_RECLAIM race fixed by
commit 366a4532d96f ("mm: fix the race between collapse and PT_RECLAIM
under per-vma lock"). But this series is not about this one.

2) Then it floods all the monitoring of the fleet, sending the same
   message in the loop, crashing the our fleet kernel monitoring
   subsystem (which is the part that I am interested in protecting)

  get_swap_device: Bad swap offset entry 3ffffffc043c5

For instance, in a host today it logged 6M in a few hours, and it is still
going forever. Two things go wrong.

1) get_swap_device() prints unconditionally, unlike print_bad_pte() next
   door which suppresses itself with is_bad_page_map_ratelimited().

1) do_swap_page() returns 0 when get_swap_device() fails, so the
   fault is retried, reads the same entry and faults again.
   Nothing in the round trip changes the PTE.

Trying to fix it in a naive way:

Patch 1 is super simple, and rate limits the two prints.

Patch 2 makes get_swap_device() return ERR_PTR(-EINVAL) for an entry
that can never name a slot on any device, keeping NULL for a device
swapoff is taking away, and converts the callers. No functional change
expected.

Patch 3 uses that to return VM_FAULT_SIGBUS instead of retrying.

PS: Sashiko flagged several pre-existing issues, and get_swap_device()
returning an error opens the door to fixing some of them.  For this
series, I am focused in landing the basic cases first and build on top,
if needed.

---
Changes in v2:
- Rate limit swap_dup_entry_direct()'s print too (Andrew)
- Drop "in get_swap_device()" from patch 1's subject, it now covers all
  three prints
- Return ERR_PTR(-EIO) rather than ERR_PTR(-EINVAL) for a malformed
  entry; -EINVAL is too soft for a corrupt page table (David)
- Document the malformed entry case in get_swap_device()'s kerneldoc,
  in patch 2 instead of patch 3 (David)
- Reword patch 2's changelog, "an entry that can never name a slot on
  any device" was unclear (David)
- Link to v1: https://patch.msgid.link/20260810-swap-v1-0-375ef0767206@debian.org

To: Andrew Morton <akpm@linux-foundation.org>
To: Chris Li <chrisl@kernel.org>
To: Kairui Song <kasong@tencent.com>
To: Kemeng Shi <shikemeng@huaweicloud.com>
To: Nhat Pham <nphamcs@gmail.com>
To: Baoquan He <baoquan.he@linux.dev>
To: Barry Song <baohua@kernel.org>
To: Youngjun Park <youngjun.park@lge.com>
To: David Hildenbrand <david@kernel.org>
To: Lorenzo Stoakes <ljs@kernel.org>
To: "Liam R. Howlett" <liam@infradead.org>
To: Vlastimil Babka <vbabka@kernel.org>
To: Mike Rapoport <rppt@kernel.org>
To: Suren Baghdasaryan <surenb@google.com>
To: Michal Hocko <mhocko@suse.com>
To: Jann Horn <jannh@google.com>
To: Pedro Falcato <pfalcato@suse.de>
To: Hugh Dickins <hughd@google.com>
To: Baolin Wang <baolin.wang@linux.alibaba.com>
To: Peter Xu <peterx@redhat.com>
To: Johannes Weiner <hannes@cmpxchg.org>
To: Yosry Ahmed <yosry@kernel.org>
To: Chengming Zhou <chengming.zhou@linux.dev>
Cc: linux-mm@kvack.org
Cc: linux-kernel@vger.kernel.org

---
Breno Leitao (3):
      mm, swap: ratelimit bad swap entry reports
      mm, swap: distinguish a malformed swap entry from a dying device
      mm: fail the fault on a malformed swap entry instead of retrying it

 mm/memory.c      |  9 +++++++--
 mm/mincore.c     |  2 +-
 mm/shmem.c       |  2 +-
 mm/swap_state.c  |  4 ++--
 mm/swapfile.c    | 20 ++++++++++++--------
 mm/userfaultfd.c |  3 ++-
 mm/zswap.c       |  2 +-
 7 files changed, 26 insertions(+), 16 deletions(-)
---
base-commit: 6b8c8af514d739d0335f5579b585e02babe8a727
change-id: 20260810-swap-25420f9c8ba9

Best regards,
--  
Breno Leitao <leitao@debian.org>
Re: [PATCH v2 0/3] mm, swap: don't spin or flood the console on a bad swap entry
Posted by Andrew Morton 1 month, 2 weeks ago
On Thu, 13 Aug 2026 03:02:19 -0700 Breno Leitao <leitao@debian.org> wrote:

> I've seen some machines at Meta fleet that show the following type of
> problem:
> 
> 1) It gets some weird warning:
> 
>   BUG: Bad page map in process khugepaged  pte:f000eef300000017 pmd:00000067
>   addr:00007f57c0a01000 vm_flags:20200073 anon_vma:ffff88829af7c340 mapping:0000000000000000 index:7f57c0a01
> 
> The corruption is most likely the collapse/PT_RECLAIM race fixed by
> commit 366a4532d96f ("mm: fix the race between collapse and PT_RECLAIM
> under per-vma lock"). But this series is not about this one.
> 
> 2) Then it floods all the monitoring of the fleet, sending the same
>    message in the loop, crashing the our fleet kernel monitoring
>    subsystem (which is the part that I am interested in protecting)
> 
>   get_swap_device: Bad swap offset entry 3ffffffc043c5
> 
> For instance, in a host today it logged 6M in a few hours, and it is still
> going forever. Two things go wrong.
> 
> 1) get_swap_device() prints unconditionally, unlike print_bad_pte() next
>    door which suppresses itself with is_bad_page_map_ratelimited().
> 
> 1) do_swap_page() returns 0 when get_swap_device() fails, so the
>    fault is retried, reads the same entry and faults again.
>    Nothing in the round trip changes the PTE.
> 
> Trying to fix it in a naive way:

Cool.

These behaviors sound pretty obnoxious.  And the patches are quite
simple so hopefully the swap maintainers will make quick work of them.

I'm assuming that users of earlier kernels will want these things fixed
so please let's work on identifying suitable Fixes: targets and
deciding which of them should get a cc:stable.



In a spirit of experimentation I asked Gemini to identify suitable Fixes:
targets and it said

[1/3]: Fixes: 122e201211e4 ("mm, swap: get_swap_device() to get reference count of swap_info_struct")

[2/3]: Fixes: 122e201211e4 ("mm, swap: get_swap_device() to get reference count of swap_info_struct")
	(and it complained that this patch doesn't fix anything)

[3/3] Fixes: 122e201211e4 ("mm, swap: get_swap_device() to get reference count of swap_info_struct")

And I cannot find such a commit anywhere, so wtf.

chatgpt didn't give me anything useful.



[2/3] is "no functional change" so ideally it simply wouldn't be
present in the series - we should aim for minimal changes when fixing
bugs, then leave the cleanups for later.


> PS: Sashiko flagged several pre-existing issues, and get_swap_device()
> returning an error opens the door to fixing some of them.  For this
> series, I am focused in landing the basic cases first and build on top,
> if needed.

Yeah. probably these are the same issues:
	https://sashiko.dev/#/patchset/20260813-swap-v2-0-4a625ccabdae@debian.org


As usual, they're all mishandled error-path things.  It's axiomatic,
really - nobody hits error-path bugs, so they never get reported so
they never get fixed.

otoh, now that these bugs are out there and known about, it's possible
that a Black Hat can find a way of exploiting them, which increases the
pressure to get these bugs addressed.
Re: [PATCH v2 0/3] mm, swap: don't spin or flood the console on a bad swap entry
Posted by Breno Leitao 1 month, 1 week ago
Hello Andrew,

On Thu, Aug 13, 2026 at 01:34:55PM -0700, Andrew Morton wrote:
> On Thu, 13 Aug 2026 03:02:19 -0700 Breno Leitao <leitao@debian.org> wrote:
> 
> > I've seen some machines at Meta fleet that show the following type of
> > problem:
> > 
> > 1) It gets some weird warning:
> > 
> >   BUG: Bad page map in process khugepaged  pte:f000eef300000017 pmd:00000067
> >   addr:00007f57c0a01000 vm_flags:20200073 anon_vma:ffff88829af7c340 mapping:0000000000000000 index:7f57c0a01
> > 
> > The corruption is most likely the collapse/PT_RECLAIM race fixed by
> > commit 366a4532d96f ("mm: fix the race between collapse and PT_RECLAIM
> > under per-vma lock"). But this series is not about this one.
> > 
> > 2) Then it floods all the monitoring of the fleet, sending the same
> >    message in the loop, crashing the our fleet kernel monitoring
> >    subsystem (which is the part that I am interested in protecting)
> > 
> >   get_swap_device: Bad swap offset entry 3ffffffc043c5
> > 
> > For instance, in a host today it logged 6M in a few hours, and it is still
> > going forever. Two things go wrong.
> > 
> > 1) get_swap_device() prints unconditionally, unlike print_bad_pte() next
> >    door which suppresses itself with is_bad_page_map_ratelimited().
> > 
> > 1) do_swap_page() returns 0 when get_swap_device() fails, so the
> >    fault is retried, reads the same entry and faults again.
> >    Nothing in the round trip changes the PTE.
> > 
> > Trying to fix it in a naive way:
> 
> Cool.
> 
> These behaviors sound pretty obnoxious.  And the patches are quite
> simple so hopefully the swap maintainers will make quick work of them.
> 
> I'm assuming that users of earlier kernels will want these things fixed
> so please let's work on identifying suitable Fixes: targets and
> deciding which of them should get a cc:stable.
> 
> 
> 
> In a spirit of experimentation I asked Gemini to identify suitable Fixes:
> targets and it said
> 
> [1/3]: Fixes: 122e201211e4 ("mm, swap: get_swap_device() to get reference count of swap_info_struct")
> 
> [2/3]: Fixes: 122e201211e4 ("mm, swap: get_swap_device() to get reference count of swap_info_struct")
> 	(and it complained that this patch doesn't fix anything)
> 
> [3/3] Fixes: 122e201211e4 ("mm, swap: get_swap_device() to get reference count of swap_info_struct")
> 
> And I cannot find such a commit anywhere, so wtf.

I think only 1/3 should be getting a Fixes: in v3. The message I am
drowning in is the Bad_offset one:

  get_swap_device: Bad swap offset entry 3ffffffc043c5

63d8620ecf93b5 ("mm/swapfile: use percpu_ref to serialize against
concurrent swapoff") added the put_out: label with just the
percpu_ref_put(), so that arm was silent. The pr_err() landed in v5.19:

So, if I need to update it, I will include:

Fixes: 23b230ba8ac3 ("mm/swap: print bad swap offset entry in get_swap_device")
Cc: <stable@vger.kernel.org>

> [2/3] is "no functional change" so ideally it simply wouldn't be
> present in the series - we should aim for minimal changes when fixing
> bugs, then leave the cleanups for later.

I need 2/3 to expose the difference in the first place.
get_swap_device() returns NULL both for a malformed entry and for
a device swapoff is taking away, so no caller can tell whether the
failure is worth retrying. 

2/3 adds that distinction and converts the callers, but none of them act
on it yet, so it is no functional change on its own. 

Then 3/3 is the actual fix, now that do_swap_page() can differentiate
a retry from give up.

Do you want me to squash them?
Re: [PATCH v2 0/3] mm, swap: don't spin or flood the console on a bad swap entry
Posted by Andrew Morton 1 month, 1 week ago
On Mon, 17 Aug 2026 05:21:32 -0700 Breno Leitao <leitao@debian.org> wrote:

> I think only 1/3 should be getting a Fixes: in v3. The message I am
> drowning in is the Bad_offset one:
> 
>   get_swap_device: Bad swap offset entry 3ffffffc043c5
> 
> 63d8620ecf93b5 ("mm/swapfile: use percpu_ref to serialize against
> concurrent swapoff") added the put_out: label with just the
> percpu_ref_put(), so that arm was silent. The pr_err() landed in v5.19:
> 
> So, if I need to update it, I will include:
> 
> Fixes: 23b230ba8ac3 ("mm/swap: print bad swap offset entry in get_swap_device")
> Cc: <stable@vger.kernel.org>

OK, so you think that only [1/3] should have cc:stable?

> > [2/3] is "no functional change" so ideally it simply wouldn't be
> > present in the series - we should aim for minimal changes when fixing
> > bugs, then leave the cleanups for later.
> 
> I need 2/3 to expose the difference in the first place.
> get_swap_device() returns NULL both for a malformed entry and for
> a device swapoff is taking away, so no caller can tell whether the
> failure is worth retrying. 
> 
> 2/3 adds that distinction and converts the callers, but none of them act
> on it yet, so it is no functional change on its own. 
> 
> Then 3/3 is the actual fix, now that do_swap_page() can differentiate
> a retry from give up.
> 
> Do you want me to squash them?

If I'm correct above then please send along [1/3] as a separate thing
and I can queue it as a backportable hotfix.  Then [2/3] and [3/3] as a
separate two-patch series for 7.3-rcX.
Re: [PATCH v2 0/3] mm, swap: don't spin or flood the console on a bad swap entry
Posted by Breno Leitao 1 month, 1 week ago
On Mon, Aug 17, 2026 at 02:53:37PM -0700, Andrew Morton wrote:
> On Mon, 17 Aug 2026 05:21:32 -0700 Breno Leitao <leitao@debian.org> wrote:
> 
> > I think only 1/3 should be getting a Fixes: in v3. The message I am
> > drowning in is the Bad_offset one:
> > 
> >   get_swap_device: Bad swap offset entry 3ffffffc043c5
> > 
> > 63d8620ecf93b5 ("mm/swapfile: use percpu_ref to serialize against
> > concurrent swapoff") added the put_out: label with just the
> > percpu_ref_put(), so that arm was silent. The pr_err() landed in v5.19:
> > 
> > So, if I need to update it, I will include:
> > 
> > Fixes: 23b230ba8ac3 ("mm/swap: print bad swap offset entry in get_swap_device")
> > Cc: <stable@vger.kernel.org>
> 
> OK, so you think that only [1/3] should have cc:stable?

Correct, that is my suggestion. The other patches are more improvements
than a proper fix, I would say.

> > > [2/3] is "no functional change" so ideally it simply wouldn't be
> > > present in the series - we should aim for minimal changes when fixing
> > > bugs, then leave the cleanups for later.
> > 
> > I need 2/3 to expose the difference in the first place.
> > get_swap_device() returns NULL both for a malformed entry and for
> > a device swapoff is taking away, so no caller can tell whether the
> > failure is worth retrying. 
> > 
> > 2/3 adds that distinction and converts the callers, but none of them act
> > on it yet, so it is no functional change on its own. 
> > 
> > Then 3/3 is the actual fix, now that do_swap_page() can differentiate
> > a retry from give up.
> > 
> > Do you want me to squash them?
> 
> If I'm correct above then please send along [1/3] as a separate thing
> and I can queue it as a backportable hotfix.  Then [2/3] and [3/3] as a
> separate two-patch series for 7.3-rcX.

ack, I will send [1/3] with the Fixes: tag, and then [2/3] and [3/3] as
a new version of this series.

Thanks,
--breno