[PATCH] mm: swap_cgroup: fix NULL deref in lookup_swap_cgroup_id on swapless host

mambaxin@163.com posted 1 patch 1 month, 2 weeks ago
mm/swap_cgroup.c | 5 +++++
1 file changed, 5 insertions(+)
[PATCH] mm: swap_cgroup: fix NULL deref in lookup_swap_cgroup_id on swapless host
Posted by mambaxin@163.com 1 month, 2 weeks ago
From: "Jose Fernandez (Anthropic)" <jose.fernandez@linux.dev>

[ Upstream commit 63b02a9409cb5180398491b093e48bcb5315f5fb ]

lookup_swap_cgroup_id() passes swap_cgroup_ctrl[type].map to
__swap_cgroup_id_lookup() without checking that the type was ever
registered via swap_cgroup_swapon().  On a swapless host every ctrl->map
is NULL, so __swap_cgroup_id_lookup() dereferences NULL + a scaled
swp_offset().

Since commit bea67dcc5eea ("mm: attempt to batch free swap entries for
zap_pte_range()"), zap_pte_range() -> swap_pte_batch() calls
lookup_swap_cgroup_id() on any non-present, non-none PTE that decodes as a
real swap entry, without first validating it against swap_info[].  A
single PTE corrupted into a type-0 swap entry takes the host down at
process exit.

We hit this in production on a swapless 6.12.58 host: ~1s of
"get_swap_device: Bad swap file entry 3f800204222bb" (do_swap_page() being
correctly defensive about the same entry) followed by

  BUG: unable to handle page fault for address: 000003f800204220
  RIP: 0010:lookup_swap_cgroup_id+0x2b/0x60
  Call Trace:
   swap_pte_batch+0xbf/0x230
   zap_pte_range+0x4c8/0x780
   unmap_page_range+0x190/0x3e0
   exit_mmap+0xd9/0x3c0
   do_exit+0x20c/0x4b0

syzbot has reported the identical stack.

The source of the PTE corruption is a separate bug; this change makes the
teardown path as robust as the fault path already is.  Every other caller
of lookup_swap_cgroup_id() is downstream of a get_swap_device() that has
already validated the entry, so the new branch is cold.

Link: https://lore.kernel.org/20260504-swap-cgroup-fix-7-0-v1-1-f53ff41ee553@linux.dev
Fixes: bea67dcc5eea ("mm: attempt to batch free swap entries for zap_pte_range()")
Signed-off-by: Jose Fernandez (Anthropic) <jose.fernandez@linux.dev>
Reported-by: syzbot+e12bd9ca48157add237a@syzkaller.appspotmail.com
Link: https://lore.kernel.org/r/69859728.050a0220.3b3015.0033.GAE@google.com
Assisted-by: Claude:unspecified
Cc: Barry Song <baohua@kernel.org>
Cc: David Hildenbrand <david@kernel.org>
Cc: Hugh Dickins <hughd@google.com>
Cc: Johannes Weiner <hannes@cmpxchg.org>
Cc: Kairui Song <ryncsn@gmail.com>
Cc: Michal Hocko <mhocko@kernel.org>
Cc: Muchun Song <muchun.song@linux.dev>
Cc: Roman Gushchin <roman.gushchin@linux.dev>
Cc: Shakeel Butt <shakeel.butt@linux.dev>
Cc: <stable@vger.kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Sasha Levin <sashal@kernel.org>
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Signed-off-by: chenxin <chenxinxin@xiaomi.com>
---
 mm/swap_cgroup.c | 5 +++++
 1 file changed, 5 insertions(+)

diff --git a/mm/swap_cgroup.c b/mm/swap_cgroup.c
index db6c4a26cf59..2d0425f4a6b9 100644
--- a/mm/swap_cgroup.c
+++ b/mm/swap_cgroup.c
@@ -161,6 +161,11 @@ unsigned short swap_cgroup_record(swp_entry_t ent, unsigned short id,
  */
 unsigned short lookup_swap_cgroup_id(swp_entry_t ent)
 {
+	struct swap_cgroup_ctrl *ctrl;
+
+	ctrl = &swap_cgroup_ctrl[swp_type(ent)];
+	if (unlikely(!ctrl->map))
+		return 0;
 	return lookup_swap_cgroup(ent, NULL)->id;
 }
 
-- 
2.50.1
Re: [PATCH] mm: swap_cgroup: fix NULL deref in lookup_swap_cgroup_id on swapless host
Posted by Barry Song 1 month, 2 weeks ago
On Fri, Aug 14, 2026 at 3:48 PM <mambaxin@163.com> wrote:
>
> From: "Jose Fernandez (Anthropic)" <jose.fernandez@linux.dev>
>
> [ Upstream commit 63b02a9409cb5180398491b093e48bcb5315f5fb ]
>
> lookup_swap_cgroup_id() passes swap_cgroup_ctrl[type].map to
> __swap_cgroup_id_lookup() without checking that the type was ever
> registered via swap_cgroup_swapon().  On a swapless host every ctrl->map
> is NULL, so __swap_cgroup_id_lookup() dereferences NULL + a scaled
> swp_offset().
>
> Since commit bea67dcc5eea ("mm: attempt to batch free swap entries for
> zap_pte_range()"), zap_pte_range() -> swap_pte_batch() calls
> lookup_swap_cgroup_id() on any non-present, non-none PTE that decodes as a
> real swap entry, without first validating it against swap_info[].  A
> single PTE corrupted into a type-0 swap entry takes the host down at
> process exit.

Thanks for the patch. However, we have a strict check to ensure that
this is only done for valid swap entries:

static inline int swap_pte_batch(pte_t *start_ptep, int max_nr, pte_t pte)
{
        pte_t expected_pte = pte_next_swp_offset(pte);
        const pte_t *end_ptep = start_ptep + max_nr;
        pte_t *ptep = start_ptep + 1;

        VM_WARN_ON(max_nr < 1);
        VM_WARN_ON(!softleaf_is_swap(softleaf_from_pte(pte)));

        while (ptep < end_ptep) {
                pte = ptep_get(ptep);

                if (!pte_same(pte, expected_pte))
                        break;
                expected_pte = pte_next_swp_offset(expected_pte);
                ptep++;
        }

        return ptep - start_ptep;
}

I don't know why this can happen on a swapless system.

>
> We hit this in production on a swapless 6.12.58 host: ~1s of
> "get_swap_device: Bad swap file entry 3f800204222bb" (do_swap_page() being
> correctly defensive about the same entry) followed by
>
>   BUG: unable to handle page fault for address: 000003f800204220
>   RIP: 0010:lookup_swap_cgroup_id+0x2b/0x60
>   Call Trace:
>    swap_pte_batch+0xbf/0x230
>    zap_pte_range+0x4c8/0x780
>    unmap_page_range+0x190/0x3e0
>    exit_mmap+0xd9/0x3c0
>    do_exit+0x20c/0x4b0
>
> syzbot has reported the identical stack.
>
> The source of the PTE corruption is a separate bug; this change makes the
> teardown path as robust as the fault path already is.  Every other caller
> of lookup_swap_cgroup_id() is downstream of a get_swap_device() that has
> already validated the entry, so the new branch is cold.

If the source is PTE corruption, I think we should fix the corruption
itself rather than work around it here.

Best Regards
Barry
Re: [PATCH] mm: swap_cgroup: fix NULL deref in lookup_swap_cgroup_id on swapless host
Posted by Shakeel Butt 1 month, 2 weeks ago
On Fri, Aug 14, 2026 at 07:23:41PM +0800, Barry Song wrote:
> On Fri, Aug 14, 2026 at 3:48 PM <mambaxin@163.com> wrote:
> >
> > From: "Jose Fernandez (Anthropic)" <jose.fernandez@linux.dev>
> >
> > [ Upstream commit 63b02a9409cb5180398491b093e48bcb5315f5fb ]
> >
> > lookup_swap_cgroup_id() passes swap_cgroup_ctrl[type].map to
> > __swap_cgroup_id_lookup() without checking that the type was ever
> > registered via swap_cgroup_swapon().  On a swapless host every ctrl->map
> > is NULL, so __swap_cgroup_id_lookup() dereferences NULL + a scaled
> > swp_offset().
> >
> > Since commit bea67dcc5eea ("mm: attempt to batch free swap entries for
> > zap_pte_range()"), zap_pte_range() -> swap_pte_batch() calls
> > lookup_swap_cgroup_id() on any non-present, non-none PTE that decodes as a
> > real swap entry, without first validating it against swap_info[].  A
> > single PTE corrupted into a type-0 swap entry takes the host down at
> > process exit.
> 
> Thanks for the patch. However, we have a strict check to ensure that
> this is only done for valid swap entries:

This patch is already in the upstream linus tree.

Mambaxin, what do you want to do with this patch? Are you requesting to backport
this to stable tree? Which one? It already had stable CCed, so I would expect
the stable tree maintainers would pick this up automatically.

Re: [PATCH] mm: swap_cgroup: fix NULL deref in lookup_swap_cgroup_id on swapless host
Posted by mambaxin@163.com 1 month, 1 week ago
> On Fri, Aug 14, 2026 at 07:23:41PM +0800, Barry Song wrote:
> > On Fri, Aug 14, 2026 at 3:48 PM <mambaxin@163.com> wrote:
> > >
> > > From: "Jose Fernandez (Anthropic)" <jose.fernandez@linux.dev>
> > >
> > > [ Upstream commit 63b02a9409cb5180398491b093e48bcb5315f5fb ]
> > >
> > > lookup_swap_cgroup_id() passes swap_cgroup_ctrl[type].map to
> > > __swap_cgroup_id_lookup() without checking that the type was ever
> > > registered via swap_cgroup_swapon().  On a swapless host every ctrl->map
> > > is NULL, so __swap_cgroup_id_lookup() dereferences NULL + a scaled
> > > swp_offset().
> > >
> > > Since commit bea67dcc5eea ("mm: attempt to batch free swap entries for
> > > zap_pte_range()"), zap_pte_range() -> swap_pte_batch() calls
> > > lookup_swap_cgroup_id() on any non-present, non-none PTE that decodes as a
> > > real swap entry, without first validating it against swap_info[].  A
> > > single PTE corrupted into a type-0 swap entry takes the host down at
> > > process exit.
> > 
> > Thanks for the patch. However, we have a strict check to ensure that
> > this is only done for valid swap entries:
> 
> This patch is already in the upstream linus tree.
> 
> Mambaxin, what do you want to do with this patch? Are you requesting to backport
> this to stable tree? Which one? It already had stable CCed, so I would expect
> the stable tree maintainers would pick this up automatically.

Hi Shakeel,Yes, we would like to backport this patch to the 6.6.y stable tree. 
We have encountered the same issue on a product running the Android GKI 6.6 kernel. 
We are still investigating the underlying PTE corruption using various debugging approaches, 
but the reproduction rate is low. Given the constraints of our product release schedule, 
we would like to first bring in this fix from upstream so that we can subsequently pull it into the Android GKI kernel.
Re: [PATCH] mm: swap_cgroup: fix NULL deref in lookup_swap_cgroup_id on swapless host
Posted by Sasha Levin 1 month, 1 week ago
> Yes, we would like to backport this patch to the 6.6.y stable tree.
> We have encountered the same issue on a product running the Android GKI 6.6 kernel.

This is not a 6.6.y fix. The NULL deref is only reachable through the
batching added by bea67dcc5eea ("mm: attempt to batch free swap entries
for zap_pte_range()"), which is v6.12 and newer. 6.6.y has no
swap_pte_batch() at all, so the exit_mmap -> zap_pte_range ->
lookup_swap_cgroup_id path you are hitting does not exist there and the
guard would be dead code.

Your oops is on Android GKI 6.6, which carries the batching backport
downstream - that is the tree that should carry this fix, alongside the
commit that makes it reachable.

-- 
Thanks,
Sasha
Re: [PATCH] mm: swap_cgroup: fix NULL deref in lookup_swap_cgroup_id on swapless host
Posted by chenxin 1 month, 1 week ago
> > Yes, we would like to backport this patch to the 6.6.y stable tree.
> > We have encountered the same issue on a product running the Android GKI 6.6 kernel.
>
> This is not a 6.6.y fix. The NULL deref is only reachable through the
> batching added by bea67dcc5eea ("mm: attempt to batch free swap entries
> for zap_pte_range()"), which is v6.12 and newer. 6.6.y has no
> swap_pte_batch() at all, so the exit_mmap -> zap_pte_range ->
> lookup_swap_cgroup_id path you are hitting does not exist there and the
> guard would be dead code.
> 
> Your oops is on Android GKI 6.6, which carries the batching backport
> downstream - that is the tree that should carry this fix, alongside the
> commit that makes it reachable.
> 
> -- 
> Thanks,
> Sasha

Thanks a lot for clarifying, as you noted, this commit was backported
to GKI 6.6 by Google, which is what makes the NULL deref reachable there.
I'll follow up with Google to get the fix backported onto the GKI kernel.

Thanks again for your help.

--
Thanks,
Chenxin