From nobody Sat Jul 25 18:53:41 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9A97239CD00; Tue, 14 Jul 2026 17:24:52 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784049893; cv=none; b=HwIgWdOFgImkpBTXGKEDrcnhpDcTmS02tAL7AuJehCz6eC3K5jzyFpl/3KrF89rhUT/AfU5o2fB2FmlRz34MEAmiJh6Ix3b552Dd9ObcXFHXrGwA0TQk/iLnUInigUdFe8MAGvzYYfDTW+9OrH9QqVF6LHpsvaZjD+940ZCG2v0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784049893; c=relaxed/simple; bh=VgBPfIW0qO8qs1oxwHrwgRLvaSj1B2gVcQMgBYESrY4=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=TvOzqeuVtCBlJiVKBmMQ9TjS9U6Y0JeMxYjdPwXsUk7TwS63RirIGP0+bV4WUm1XEihVy/qWWjN5kFQI/CpGi0+VzZ4YTBWAY22BzEF7xjynpggwChd658E7ulH8wlDEgmQKVg1GTKGwa+Qmyvmtiv3okDye5m69cSnaEpvtMf4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=dkCg7y8A; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="dkCg7y8A" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 59B431F00A3A; Tue, 14 Jul 2026 17:24:45 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784049892; bh=TdyW4Y7bH4m/445xWB+ioB8s7SF8KHb3T29qlCFwons=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=dkCg7y8AUG2qBdSFxdwAORuaXUG5SkssaCeVeMM/vBKGJX/yusOnplHactjMTh5Qu GSsN06jQv15Wb1Kubj8kCqFQfy+EwTr0Cz6y9fwA8eN/BGhkKKFGKyx3Go3fNnKPBb LIIgowYASPGyx9wMJMXryQkK6+tLRC4mB9Z/xVu+PxZNpdHcCejy6vWfVa+YLmHZHq qHDq6rfeBwnQ2zqbo31yN9v4quCfd6yhetwEufkZJMQ8X9CZJtqPGoiwgZIhIT2LMu onoCsIH0gp4nToM9Bk2hDfcw5Vr54I5fjeiCOCdzccp0aKy4GYHhNYm4lh0SWP9CYS o+JObs9IyxcYQ== From: Lorenzo Stoakes Date: Tue, 14 Jul 2026 18:24:23 +0100 Subject: [PATCH mm-hotfixes v3 1/4] mm/vmalloc: acquire init_mm lock on huge vmap to avoid ptdump UAF Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260714-series-vmap-race-fix-v3-1-b812eccfa0f9@kernel.org> References: <20260714-series-vmap-race-fix-v3-0-b812eccfa0f9@kernel.org> In-Reply-To: <20260714-series-vmap-race-fix-v3-0-b812eccfa0f9@kernel.org> To: Andrew Morton , Suren Baghdasaryan , "Liam R. Howlett" , Vlastimil Babka , Shakeel Butt , David Hildenbrand , Mike Rapoport , Michal Hocko , Uladzislau Rezki , Toshi Kani , Dave Hansen , Andy Lutomirski , Peter Zijlstra , Thomas Gleixner , Ingo Molnar , Borislav Petkov , x86@kernel.org, "H. Peter Anvin" , Kiryl Shutsemau , Catalin Marinas , Will Deacon , Dev Jain , Ryan Roberts Cc: David Carlier , ljs@kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, stable@vger.kernel.org, syzbot+fd95a72470f5a44e464c@syzkaller.appspotmail.com X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=8523; i=ljs@kernel.org; h=from:subject:message-id; bh=VgBPfIW0qO8qs1oxwHrwgRLvaSj1B2gVcQMgBYESrY4=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLLCCs7tiha8tWCh8hXmja8VyvTW/hJLtH9gx6y5UzlG5 /zjU886OkpZGMS4GGTFFFmefxHfHyQSNq/zgr8bzBxWJpAhDFycAjCRuZWMDDu8tHbKcesrz70c tvrYQcnMFVnlPdv/mseLVBRf5PjaacDwh2PVpMueovPFp83/Wa9+Ly5GxPK6Unu1SRajg3tN3M6 JDAA= X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 Currently there is a nasty race between ptdump and vmap when attempting to map a huge P4D, PMD or PUD entry: * ptdump walks kernel page table ranges it doesn't own. * When vmap maps ranges it tries to promotes existing ones to huge page tables in vmap_try_huge_[p4d,pud,pmd]() at P4D, PUD and PMD level, freeing the lower page table in [p4d,pud,pmd]_free_[pud,pmd,pte]_page() when it succeeds. Both of these things can happen at the same time and as a result ptdump can access a freed page table, resulting in a use-after-free and memory corruption. This is possible because while ptdump_walk_pgd() holds both the mem hotplug lock and the mmap write lock before invoking walk_page_range_debug(), vmap takes no relevant locks at all. Fix this by holding the mmap read lock in vmap_try_huge_*() when freeing page tables. We also hold the lock while assigning the huge page table entry, which means page table walkers observe only the huge or non-huge page table entry. We use a trylock to prevent ptdump from blocking vmap making forward progress. This is fine because it's an optimisation in any case, and thus the vmap can safely proceed regardless. All other kernel page table walkers that touch vmalloc ranges either exclusively own the memory walked or acquire the mmap lock, so this correctly excludes those walkers. One wrinkle here is commit fa93b45fd397 ("arm64: Enable vmalloc-huge with ptdump"), which addresses the issue for arm64 only by explicitly acquiring the mmap read lock on kernel page table freeing should a concurrent ptdump be in progress. This is problematic as vmap may acquire the mmap read lock prior to ptdump attempting to acquire an mmap write lock, leading to a deadlock when the mmap read lock is slept upon on page table freeing due to rwsem anti-starvation. We work around this by predicating the mmap lock being taken on !CONFIG_ARM64 for the time being. With this patch applied, a follow up will partially revert commit fa93b45fd397 ("arm64: Enable vmalloc-huge with ptdump") and at that stage remove the arm64 ifdeffery. We also update walk_page_range_debug() to assert the mmap write lock unconditionally and update the comment here to reflect this change. The issue has existed as long as ptdump was available and vmap freed page tables when promoting to a huge leaf entry, that is, since commit b6bdb7517c3d ("mm/vmalloc: add interfaces to free unmapped page table") for huge ioremap, and commit 121e6f3258fe ("mm/vmalloc: hugepage vmalloc mappings") for huge vmalloc. Since the former is the earlier of the two we choose that for our Fixes tag. We also define a guard class for mmap_read_trylock() so we can use cleanup.h to make the scope handling cleaner in the implementation. This patch is based on work by David Carlier (linked), with gratitude! Fixes: b6bdb7517c3d ("mm/vmalloc: add interfaces to free unmapped page tabl= e") Cc: stable@vger.kernel.org Reported-by: syzbot+fd95a72470f5a44e464c@syzkaller.appspotmail.com Closes: https://lore.kernel.org/all/6a287988.39669fcc.33b062.00a0.GAE@googl= e.com/T/ Link: https://lore.kernel.org/linux-mm/20260706203128.162335-1-devnexen@gma= il.com/ Reviewed-by: Mike Rapoport (Microsoft) Signed-off-by: Lorenzo Stoakes Reviewed-by: Dev Jain Acked-by: David Hildenbrand (Arm) Reviewed-by: Kiryl Shutsemau --- include/linux/mmap_lock.h | 1 + mm/pagewalk.c | 22 +++++++++++---------- mm/vmalloc.c | 50 ++++++++++++++++++++++++++++++++++++++-----= ---- 3 files changed, 54 insertions(+), 19 deletions(-) diff --git a/include/linux/mmap_lock.h b/include/linux/mmap_lock.h index 04b8f61ece5d..6b5c2390cc30 100644 --- a/include/linux/mmap_lock.h +++ b/include/linux/mmap_lock.h @@ -621,6 +621,7 @@ static inline void mmap_read_unlock(struct mm_struct *m= m) =20 DEFINE_GUARD(mmap_read_lock, struct mm_struct *, mmap_read_lock(_T), mmap_read_unlock(_T)) +DEFINE_GUARD_COND(mmap_read_lock, _try, mmap_read_trylock(_T)) =20 static inline void mmap_read_unlock_non_owner(struct mm_struct *mm) { diff --git a/mm/pagewalk.c b/mm/pagewalk.c index 3ae2586ff45b..bbcfd68d0907 100644 --- a/mm/pagewalk.c +++ b/mm/pagewalk.c @@ -678,6 +678,8 @@ int walk_kernel_page_table_range_lockless(unsigned long= start, unsigned long end * will also not lock the PTEs for the pte_entry() callback. * * This is for debugging purposes ONLY. + * + * The mmap write lock must be held. */ int walk_page_range_debug(struct mm_struct *mm, unsigned long start, unsigned long end, const struct mm_walk_ops *ops, @@ -691,6 +693,16 @@ int walk_page_range_debug(struct mm_struct *mm, unsign= ed long start, .no_vma =3D true }; =20 + /* + * When walking userland page tables, an mmap write lock must be held to + * account for munmap() downgrading to an mmap read lock when tearing + * down page tables. + * + * When walking kernel page tables, an mmap write lock must also be held + * to account for page table freeing on vmap huge page mapping. + */ + mmap_assert_write_locked(mm); + /* For convenience, we allow traversal of kernel mappings. */ if (mm =3D=3D &init_mm) return walk_kernel_page_table_range(start, end, ops, @@ -700,16 +712,6 @@ int walk_page_range_debug(struct mm_struct *mm, unsign= ed long start, if (!check_ops_safe(ops)) return -EINVAL; =20 - /* - * The mmap lock protects the page walker from changes to the page - * tables during the walk. However a read lock is insufficient to - * protect those areas which don't have a VMA as munmap() detaches - * the VMAs before downgrading to a read lock and actually tearing - * down PTEs/page tables. In which case, the mmap write lock should - * be held. - */ - mmap_assert_write_locked(mm); - return walk_pgd_range(start, end, &walk); } =20 diff --git a/mm/vmalloc.c b/mm/vmalloc.c index 1afca3568b9b..1fa9ac6e43d4 100644 --- a/mm/vmalloc.c +++ b/mm/vmalloc.c @@ -43,6 +43,7 @@ #include #include #include +#include =20 #define CREATE_TRACE_POINTS #include @@ -158,10 +159,25 @@ static int vmap_try_huge_pmd(pmd_t *pmd, unsigned lon= g addr, unsigned long end, if (!IS_ALIGNED(phys_addr, PMD_SIZE)) return 0; =20 - if (pmd_present(*pmd) && !pmd_free_pte_page(pmd, addr)) - return 0; + if (!pmd_present(*pmd)) + return pmd_set_huge(pmd, phys_addr, prot); =20 - return pmd_set_huge(pmd, phys_addr, prot); + /* + * Kernel page table walkers either walk ranges they own exclusively or + * hold the mmap write lock on init_mm (ptdump being the motivating + * case). + * + * Therefore, acquire the mmap read lock to prevent use-after-free when + * freeing page tables. + */ +#ifndef CONFIG_ARM64 + scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) +#endif + { + if (!pmd_free_pte_page(pmd, addr)) + return 0; + return pmd_set_huge(pmd, phys_addr, prot); + } } =20 static int vmap_pmd_range(pud_t *pud, unsigned long addr, unsigned long en= d, @@ -210,10 +226,18 @@ static int vmap_try_huge_pud(pud_t *pud, unsigned lon= g addr, unsigned long end, if (!IS_ALIGNED(phys_addr, PUD_SIZE)) return 0; =20 - if (pud_present(*pud) && !pud_free_pmd_page(pud, addr)) - return 0; + if (!pud_present(*pud)) + return pud_set_huge(pud, phys_addr, prot); =20 - return pud_set_huge(pud, phys_addr, prot); + /* See comment in vmap_try_huge_pmd(). */ +#ifndef CONFIG_ARM64 + scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) +#endif + { + if (!pud_free_pmd_page(pud, addr)) + return 0; + return pud_set_huge(pud, phys_addr, prot); + } } =20 static int vmap_pud_range(p4d_t *p4d, unsigned long addr, unsigned long en= d, @@ -262,10 +286,18 @@ static int vmap_try_huge_p4d(p4d_t *p4d, unsigned lon= g addr, unsigned long end, if (!IS_ALIGNED(phys_addr, P4D_SIZE)) return 0; =20 - if (p4d_present(*p4d) && !p4d_free_pud_page(p4d, addr)) - return 0; + if (!p4d_present(*p4d)) + return p4d_set_huge(p4d, phys_addr, prot); =20 - return p4d_set_huge(p4d, phys_addr, prot); + /* See comment in vmap_try_huge_pmd(). */ +#ifndef CONFIG_ARM64 + scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) +#endif + { + if (!p4d_free_pud_page(p4d, addr)) + return 0; + return p4d_set_huge(p4d, phys_addr, prot); + } } =20 static int vmap_p4d_range(pgd_t *pgd, unsigned long addr, unsigned long en= d, --=20 2.55.0 From nobody Sat Jul 25 18:53:41 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BCE522AD37; Tue, 14 Jul 2026 17:24:59 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784049900; cv=none; b=KBdoeUjr/f3/a5zIe0pi+Pd76xgJHfwkz6dekM6FkckLRtrvRds3llDqJW30Ygshvcsogqf6v5uXguucr8BC9Kve5xqqMaNfdTt+wk9bqNzqHPM63QnrGroYoeMg4fALzVjm6NpQGTr4U1W0V9bXbyHcmN2pXtzvYS3PBHjNYyA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784049900; c=relaxed/simple; bh=1IRMkBuBSJn8c0dhLGk5dFenOrN5JSkDTuylzdcXQUQ=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=P7JVg0Y10h8y25IlqNU+tG9ohad3PTKmkCyJ16Lg2EGm/1y/EPCTyra7LbsZfwp62tNIdPDnrp7VJ9QF+eXBfiWNp8/0QulEZidfTaN/wjOzrveJssZ0QGvngZYcUIb2YWJsyiISFtlYJSmuSGKxAV6lIFgYxWdEyNzj2YOXf8Y= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=A+kxTrnX; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="A+kxTrnX" Received: by smtp.kernel.org (Postfix) with ESMTPSA id B0DF91F00A3D; Tue, 14 Jul 2026 17:24:52 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784049899; bh=BYfKvJYN8Ofzx4ZcF2C0yaTq6AqCLt0T09rDQbFZEtU=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=A+kxTrnXg6Io/GxtciuQ8gUMV/UHNsl8xfZDXmio287Uq1eq0AB/sirE9DHnMeL55 spsQknwWtzixnvri9S7+b1ZoAHYYaoFXquGjVStimNk5YfB9wHH/WVtgglXU3m54/D 7HYTFusEKpmBvCPbd4aZRaHdwPEMj/1rc52yjuPKNBmynU58cfrjReDxBCNQAsOKf9 l2V7BFZBr3kQwXUAKTXb9gQvLFdPLqECx0sI83DZqFtL0EoaGwK8xDcM5AGefLIoII BpcW0cq5ehnHH5qa1pzpq2VMVA3bIawmGGCJt8ifvxADVJGyqCLUEF07ptVJjY1LZi IRz8CTEyXAAGg== From: Lorenzo Stoakes Date: Tue, 14 Jul 2026 18:24:24 +0100 Subject: [PATCH mm-hotfixes v3 2/4] x86/mm/pat: acquire mmap lock on page table free to avoid ptdump UAF Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260714-series-vmap-race-fix-v3-2-b812eccfa0f9@kernel.org> References: <20260714-series-vmap-race-fix-v3-0-b812eccfa0f9@kernel.org> In-Reply-To: <20260714-series-vmap-race-fix-v3-0-b812eccfa0f9@kernel.org> To: Andrew Morton , Suren Baghdasaryan , "Liam R. Howlett" , Vlastimil Babka , Shakeel Butt , David Hildenbrand , Mike Rapoport , Michal Hocko , Uladzislau Rezki , Toshi Kani , Dave Hansen , Andy Lutomirski , Peter Zijlstra , Thomas Gleixner , Ingo Molnar , Borislav Petkov , x86@kernel.org, "H. Peter Anvin" , Kiryl Shutsemau , Catalin Marinas , Will Deacon , Dev Jain , Ryan Roberts Cc: David Carlier , ljs@kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, stable@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=2859; i=ljs@kernel.org; h=from:subject:message-id; bh=1IRMkBuBSJn8c0dhLGk5dFenOrN5JSkDTuylzdcXQUQ=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLLCCs6x5etcWPDhb4c5X+Hv5b1Xc349/WTx8ort+4YT0 yXsZVYt7ihlYRDjYpAVU2R5/kV8f5BI2LzOC/5uMHNYmUCGMHBxCsBE9j9m+B8ld9rf4Hmdcmzz 8kB317VWRsGfPI89nX/E8NeTTrmkaE2G/9HXsroSfPSfSCpkf2OZzvbr7yzG+tCfR6s18mKfGs9 OZgQA X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 x86 implements page attribute modification using its Change Page Attributes (CPA) mechanism. This tracks properties of ranges such as cache mode through x86 page attributes, and as part of that logic manipulates kernel page tables. Since commit 41d88484c71c ("x86/mm/pat: restore large ROX pages after fragmentation") ranges of kernel page table entries can be collapsed into huge page table entries as part of this logic. As part of this collapse, it frees the page tables which the collapsed entries previously pointed to, and it does so without any relevant locks being held to preclude concurrent kernel page table walkers. The only way this code can be reached is if CPA_COLLAPSE is specified, and this is only set in set_memory_rox() via: set_memory_rox() -> change_page_attr_set_clr() -> cpa_flush() -> cpa_collapse_large_pages() Notable users of this are execmem and bpf when manipulating executable mappings. However, this is problematic for ptdump as it walks ranges it does not own and thus runs the risk of a use-after-free on page tables freed underneath it. Resolve the issue by acquiring the mmap read lock on init_mm which prevents a concurrent ptdump as it acquires the write lock. It is safe to acquire a sleeping lock as all the callers invoke set_memory_rox() from process context and in any case, change_page_attr_set_clr() calls vm_unmap_alias() which ultimately takes a mutex, disallowing atomic context here. Fixes: 41d88484c71c ("x86/mm/pat: restore large ROX pages after fragmentati= on") Cc: stable@vger.kernel.org Reviewed-by: Mike Rapoport (Microsoft) Reviewed-by: Kiryl Shutsemau (Meta) Signed-off-by: Lorenzo Stoakes Reviewed-by: Dave Hansen Reviewed-by: David Hildenbrand (Arm) --- arch/x86/mm/pat/set_memory.c | 14 +++++++++++--- 1 file changed, 11 insertions(+), 3 deletions(-) diff --git a/arch/x86/mm/pat/set_memory.c b/arch/x86/mm/pat/set_memory.c index d023a40a1e03..4c4b8244502f 100644 --- a/arch/x86/mm/pat/set_memory.c +++ b/arch/x86/mm/pat/set_memory.c @@ -22,6 +22,7 @@ #include #include #include +#include =20 #include #include @@ -436,9 +437,16 @@ static void cpa_collapse_large_pages(struct cpa_data *= cpa) =20 flush_tlb_all(); =20 - list_for_each_entry_safe(ptdesc, tmp, &pgtables, pt_list) { - list_del(&ptdesc->pt_list); - pagetable_free(ptdesc); + /* + * ptdump might read these page tables, so avoid a use-after-free by + * acquiring the mmap read lock on init_mm (ptdump acquires the mmap + * write lock). + */ + scoped_guard(mmap_read_lock, &init_mm) { + list_for_each_entry_safe(ptdesc, tmp, &pgtables, pt_list) { + list_del(&ptdesc->pt_list); + pagetable_free(ptdesc); + } } } =20 --=20 2.55.0 From nobody Sat Jul 25 18:53:41 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0A32847ECF7; Tue, 14 Jul 2026 17:25:06 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784049908; cv=none; b=rCcEXvnNamxNWS2mIG5KxR/cAtg8IVbFztS6k86C9WHkRU6ALD1K592q5we9ECpOFjkCLHTTgRhItu9VV1bC98W3Ldg9yhHo5CLeoiJIJohDA3YifalrqOvSV6d9lO/cfz2BxUOf8gF5qGYRNCakD4rSeGuZNH3MsZqho+7XRg0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784049908; c=relaxed/simple; bh=+da+wQab4iV0i+3djD3hSaZE5QbecH6NvEFAGihVKmE=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=WFnWwPm0/KczZsVHeuRVJ2lpP5peWKpuimf+nHlow+Xel4ZTDbbU/oYakhBHeBTRX2QZWdb5kEDN68DiNd1/BCNI1WPiBz8gvu4fmAIiSUOt3jcRLqrViXP8SRly3+38pGiux2kl74K/fIU9ZNQvoe8EN1rYsj4yte5DJyQknhQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=OJlvXF9S; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="OJlvXF9S" Received: by smtp.kernel.org (Postfix) with ESMTPSA id D3C221F000E9; Tue, 14 Jul 2026 17:24:59 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784049906; bh=7c4JIh+xXXzUf3RuGROVnJmdpEMXk82+j8IbVgKDa6Q=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=OJlvXF9SDCYElXNumMqDXpyahca0HdWzyR4wK5diDEgqmbgJmZxbHSAmgZtsY66mJ bxTgooysHMBrODx4u6eF8rwqriH74jqLi5TwDCa9BpCVWGUCoSreK/wp529S6l3A6W JhsYCDV9C3rAa/ohqxc71Us0gLNQ8pn64J+rWkVgaiqooJHhXOaOJbJKLgxdEqY6t6 GnXC5B5wxHArdwF528qzs2ZmhWbkBl+6NoR29hMloLREUfslmoJi9qiPBpP78dWK9z DQoD+2SamlgwqAqBjDfeB3EAsBSoyHc1Qh0CQKrfheXphhSAYkNR62Rh4avwgei7KB MVtvAV6CyT3Uw== From: Lorenzo Stoakes Date: Tue, 14 Jul 2026 18:24:25 +0100 Subject: [PATCH mm-hotfixes v3 3/4] mm/ptdump: always stabilise against page table freeing using init_mm Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260714-series-vmap-race-fix-v3-3-b812eccfa0f9@kernel.org> References: <20260714-series-vmap-race-fix-v3-0-b812eccfa0f9@kernel.org> In-Reply-To: <20260714-series-vmap-race-fix-v3-0-b812eccfa0f9@kernel.org> To: Andrew Morton , Suren Baghdasaryan , "Liam R. Howlett" , Vlastimil Babka , Shakeel Butt , David Hildenbrand , Mike Rapoport , Michal Hocko , Uladzislau Rezki , Toshi Kani , Dave Hansen , Andy Lutomirski , Peter Zijlstra , Thomas Gleixner , Ingo Molnar , Borislav Petkov , x86@kernel.org, "H. Peter Anvin" , Kiryl Shutsemau , Catalin Marinas , Will Deacon , Dev Jain , Ryan Roberts Cc: David Carlier , ljs@kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, stable@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=3218; i=ljs@kernel.org; h=from:subject:message-id; bh=+da+wQab4iV0i+3djD3hSaZE5QbecH6NvEFAGihVKmE=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLLCCs5J7kqZej7tJJ/tEu1Z3zYd/lkiw+B3/oJ83J5XP brVjMFfO0pZGMS4GGTFFFmefxHfHyQSNq/zgr8bzBxWJpAhDFycAjCRC7MY/pn16AjX77hp7Hdo +83gjqAb2f+WTop+GxJcrPMq0i7PeQEjw+FrCicDTl1s5Kz61Wepp/mPj/uT7MLn01UuO61Z+5/ nADcA X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 Previous commits have established the invariant that kernel page table freeing is performed while an mmap read lock on init_mm is held, which fixes races between ptdump and kernel page table freeing over init_mm. However, x86 and arm64 can perform a ptdump over an mm other than init_mm via ptdump_walk_pgd() and since kernel memory ranges are shared across non-kernel mm's, this means that the race still exists for these cases. Fix this by acquiring a nested mmap write lock for init_mm in ptdump_walk_pgd(). This is safe as we take this after mmap write locking the mm, and nothing acquires the init_mm lock first before locking an arbitrary mm, so no deadlock is possible. Also update walk_page_range_debug() to assert that init_mm is write locked, add a comment explaining why and remove some redundant code, and eliminate the unnecessary and confusing invocation of walk_kernel_page_table_range(). We can safely remove the non-NULL check for walk.mm, as the mmap lock asserts would NULL pointer deref if it was (and of course no callers do this). The first point at which ptdump can race kernel page table freeing is commit b6bdb7517c3d ("mm/vmalloc: add interfaces to free unmapped page table"), so we target this in the Fixes tag. Fixes: b6bdb7517c3d ("mm/vmalloc: add interfaces to free unmapped page tabl= e") Cc: stable@vger.kernel.org Reviewed-by: Mike Rapoport (Microsoft) Signed-off-by: Lorenzo Stoakes Acked-by: David Hildenbrand (Arm) Reviewed-by: Kiryl Shutsemau --- mm/pagewalk.c | 14 +++++++++----- mm/ptdump.c | 7 +++++++ 2 files changed, 16 insertions(+), 5 deletions(-) diff --git a/mm/pagewalk.c b/mm/pagewalk.c index bbcfd68d0907..5d87c632a255 100644 --- a/mm/pagewalk.c +++ b/mm/pagewalk.c @@ -702,12 +702,16 @@ int walk_page_range_debug(struct mm_struct *mm, unsig= ned long start, * to account for page table freeing on vmap huge page mapping. */ mmap_assert_write_locked(mm); + /* + * x86, arm64 ptdump allow walks of efi mm's and x86 ptdump allows walks + * of arbitrary mm's. + * + * However, they both must also hold the init_mm lock to account for + * concurrent kernel page table freeing. + */ + mmap_assert_write_locked(&init_mm); =20 - /* For convenience, we allow traversal of kernel mappings. */ - if (mm =3D=3D &init_mm) - return walk_kernel_page_table_range(start, end, ops, - pgd, private); - if (start >=3D end || !walk.mm) + if (start >=3D end) return -EINVAL; if (!check_ops_safe(ops)) return -EINVAL; diff --git a/mm/ptdump.c b/mm/ptdump.c index 973020000096..5851096e6f65 100644 --- a/mm/ptdump.c +++ b/mm/ptdump.c @@ -178,11 +178,18 @@ void ptdump_walk_pgd(struct ptdump_state *st, struct = mm_struct *mm, pgd_t *pgd) =20 get_online_mems(); mmap_write_lock(mm); + /* To stabilise kernel page tables we must hold the init_mm lock too. */ + if (mm !=3D &init_mm) + mmap_write_lock_nested(&init_mm, SINGLE_DEPTH_NESTING); + while (range->start !=3D range->end) { walk_page_range_debug(mm, range->start, range->end, &ptdump_ops, pgd, st); range++; } + + if (mm !=3D &init_mm) + mmap_write_unlock(&init_mm); mmap_write_unlock(mm); put_online_mems(); =20 --=20 2.55.0 From nobody Sat Jul 25 18:53:41 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 25F6147DD7A; Tue, 14 Jul 2026 17:25:13 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784049915; cv=none; b=GifNEkL8Kg6oZw2/aLZbF6UscyN51fUjHTT46n7lZSQLsCcQ5EQxRmJju8IifJxPdJgVHmCvWTcUiY+pXmEfSndNQdPlT/vf7NTlkPjHkGuRGIyGh9w9pVfHHztmp/Hkw8Wr17K0toKkzYZRloxCbttADwr/k8wKWQqRo44H/gA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784049915; c=relaxed/simple; bh=QfkQiauRr1PTaQoDS6zXfYa+8p/BKdWZL6M6VQqRgS0=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=Hm8ViroIN6Y32b5YUuGVUcv6nmYQEjlVUm2TmoxhlLtx+916daNYH5Y9gA58t8kGzXkm0gBhDVwjn+GnvCU+qaU56kNxyn1eJ0kY9d/RvRhJiI8xPVub95bw5rFaaoSNdHJKCI12kwXP3YZwPh3O16iEfeReujAPk7jotMqTNBQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=PC17pXS2; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="PC17pXS2" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 279171F00A3A; Tue, 14 Jul 2026 17:25:06 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784049913; bh=6X5YMXpvKlN1rtFZuXWP5ezzGqrHft5cjQkbgIqQagg=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=PC17pXS2WHaTYBIL1Slqx4PrZnyY8dMc2smCQYBF/0yu4Wi4jL/xj6gSsjw9azwsH h2PBo0E+0nEYjR6EkiARz92BXhIqmjmW4a6zK34FpYcajUMzqcWO7elmfDzKcaoqr7 dsgozGajBE7ht5bjTWhV9wjAMcphziVzf5Wcuoh1s1jsrCx8EIiwXgedkDdnqhJv/y KyApJKXEtaOZR2x9QFGqMRtDHMBG3iTiLsarzRks95WAQ6NzX5+pCm4pDlv19UfUKt ZtJx4I31XHtkTgIb9cpQS7AVdPrHfrSmd0TmDewmewNFNK/PJKo1UzNVdV3ivRcQKH uERmQC+kxoAVA== From: Lorenzo Stoakes Date: Tue, 14 Jul 2026 18:24:26 +0100 Subject: [PATCH mm-hotfixes v3 4/4] arm64: remove redundant concurrent ptdump UAF mitigation Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260714-series-vmap-race-fix-v3-4-b812eccfa0f9@kernel.org> References: <20260714-series-vmap-race-fix-v3-0-b812eccfa0f9@kernel.org> In-Reply-To: <20260714-series-vmap-race-fix-v3-0-b812eccfa0f9@kernel.org> To: Andrew Morton , Suren Baghdasaryan , "Liam R. Howlett" , Vlastimil Babka , Shakeel Butt , David Hildenbrand , Mike Rapoport , Michal Hocko , Uladzislau Rezki , Toshi Kani , Dave Hansen , Andy Lutomirski , Peter Zijlstra , Thomas Gleixner , Ingo Molnar , Borislav Petkov , x86@kernel.org, "H. Peter Anvin" , Kiryl Shutsemau , Catalin Marinas , Will Deacon , Dev Jain , Ryan Roberts Cc: David Carlier , ljs@kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, stable@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=6970; i=ljs@kernel.org; h=from:subject:message-id; bh=QfkQiauRr1PTaQoDS6zXfYa+8p/BKdWZL6M6VQqRgS0=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLLCCs6xNmyP28KoEf/gWoHNpG1/fXYf2zhlta4S669uy QXX9WxudZSyMIhxMciKKbI8/yK+P0gkbF7nBX83mDmsTCBDGLg4BWAidnEMv5gadq15wVohonZt 46eHq3Ii/Q9Jlk7gEDx3kf/ot3Ui4jcYGToXsTs9EnRJnC5rsyE2cO1k6wOFbXMibap49z40L3u 3lAEA X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 This partially reverts commit fa93b45fd397 ("arm64: Enable vmalloc-huge with ptdump"), retaining vmalloc-huge support but eliminating the now redundant mitigation against a race between huge vmap page table freeing and ptdump, as this issue has now been fixed at core. We also simultaneously remove the arm64 if-deffery when acquiring the mmap read lock upon vmap huge page table promotion as it is no longer required. Note that this patch relies on the preceding vmalloc patch, and should not be backported alone. Fixes: fa93b45fd397 ("arm64: Enable vmalloc-huge with ptdump") Cc: stable@vger.kernel.org Reviewed-by: Dev Jain Acked-by: Mike Rapoport (Microsoft) Signed-off-by: Lorenzo Stoakes Acked-by: Kiryl Shutsemau (Meta) Acked-by: Will Deacon Reviewed-by: David Hildenbrand (Arm) --- arch/arm64/include/asm/ptdump.h | 2 -- arch/arm64/mm/mmu.c | 43 ++++---------------------------------= ---- arch/arm64/mm/ptdump.c | 11 ++--------- mm/vmalloc.c | 15 +++----------- 4 files changed, 9 insertions(+), 62 deletions(-) diff --git a/arch/arm64/include/asm/ptdump.h b/arch/arm64/include/asm/ptdum= p.h index 5b374a6ab34a..50a195eda8ed 100644 --- a/arch/arm64/include/asm/ptdump.h +++ b/arch/arm64/include/asm/ptdump.h @@ -7,8 +7,6 @@ =20 #include =20 -DECLARE_STATIC_KEY_FALSE(arm64_ptdump_lock_key); - #ifdef CONFIG_PTDUMP =20 #include diff --git a/arch/arm64/mm/mmu.c b/arch/arm64/mm/mmu.c index a25d8beacc83..bd52fca6e872 100644 --- a/arch/arm64/mm/mmu.c +++ b/arch/arm64/mm/mmu.c @@ -49,8 +49,6 @@ #define NO_CONT_MAPPINGS BIT(1) #define NO_EXEC_MAPPINGS BIT(2) /* assumes FEAT_HPDS is not used */ =20 -DEFINE_STATIC_KEY_FALSE(arm64_ptdump_lock_key); - u64 kimage_voffset __ro_after_init; EXPORT_SYMBOL(kimage_voffset); =20 @@ -1864,8 +1862,7 @@ int pmd_clear_huge(pmd_t *pmdp) return 1; } =20 -static int __pmd_free_pte_page(pmd_t *pmdp, unsigned long addr, - bool acquire_mmap_lock) +int pmd_free_pte_page(pmd_t *pmdp, unsigned long addr) { pte_t *table; pmd_t pmd; @@ -1877,25 +1874,13 @@ static int __pmd_free_pte_page(pmd_t *pmdp, unsigne= d long addr, return 1; } =20 - /* See comment in pud_free_pmd_page for static key logic */ table =3D pte_offset_kernel(pmdp, addr); pmd_clear(pmdp); __flush_tlb_kernel_pgtable(addr); - if (static_branch_unlikely(&arm64_ptdump_lock_key) && acquire_mmap_lock) { - mmap_read_lock(&init_mm); - mmap_read_unlock(&init_mm); - } - pte_free_kernel(NULL, table); return 1; } =20 -int pmd_free_pte_page(pmd_t *pmdp, unsigned long addr) -{ - /* If ptdump is walking the pagetables, acquire init_mm.mmap_lock */ - return __pmd_free_pte_page(pmdp, addr, /* acquire_mmap_lock =3D */ true); -} - int pud_free_pmd_page(pud_t *pudp, unsigned long addr) { pmd_t *table; @@ -1911,36 +1896,16 @@ int pud_free_pmd_page(pud_t *pudp, unsigned long ad= dr) } =20 table =3D pmd_offset(pudp, addr); - - /* - * Our objective is to prevent ptdump from reading a PMD table which has - * been freed. In this race, if pud_free_pmd_page observes the key on - * (which got flipped by ptdump) then the mmap lock sequence here will, - * as a result of the mmap write lock/unlock sequence in ptdump, give - * us the correct synchronization. If not, this means that ptdump has - * yet not started walking the pagetables - the sequence of barriers - * issued by __flush_tlb_kernel_pgtable() guarantees that ptdump will - * observe an empty PUD. - */ - pud_clear(pudp); - __flush_tlb_kernel_pgtable(addr); - if (static_branch_unlikely(&arm64_ptdump_lock_key)) { - mmap_read_lock(&init_mm); - mmap_read_unlock(&init_mm); - } - pmdp =3D table; next =3D addr; end =3D addr + PUD_SIZE; do { if (pmd_present(pmdp_get(pmdp))) - /* - * PMD has been isolated, so ptdump won't see it. No - * need to acquire init_mm.mmap_lock. - */ - __pmd_free_pte_page(pmdp, next, /* acquire_mmap_lock =3D */ false); + pmd_free_pte_page(pmdp, next); } while (pmdp++, next +=3D PMD_SIZE, next !=3D end); =20 + pud_clear(pudp); + __flush_tlb_kernel_pgtable(addr); pmd_free(NULL, table); return 1; } diff --git a/arch/arm64/mm/ptdump.c b/arch/arm64/mm/ptdump.c index 1c20144700d7..5a76c59b5ada 100644 --- a/arch/arm64/mm/ptdump.c +++ b/arch/arm64/mm/ptdump.c @@ -283,13 +283,6 @@ void note_page_flush(struct ptdump_state *pt_st) note_page(pt_st, 0, -1, pte_val(pte_zero)); } =20 -static void arm64_ptdump_walk_pgd(struct ptdump_state *st, struct mm_struc= t *mm) -{ - static_branch_inc(&arm64_ptdump_lock_key); - ptdump_walk_pgd(st, mm, NULL); - static_branch_dec(&arm64_ptdump_lock_key); -} - void ptdump_walk(struct seq_file *s, struct ptdump_info *info) { unsigned long end =3D ~0UL; @@ -318,7 +311,7 @@ void ptdump_walk(struct seq_file *s, struct ptdump_info= *info) } }; =20 - arm64_ptdump_walk_pgd(&st.ptdump, info->mm); + ptdump_walk_pgd(&st.ptdump, info->mm, NULL); } =20 static void __init ptdump_initialize(void) @@ -360,7 +353,7 @@ bool ptdump_check_wx(void) } }; =20 - arm64_ptdump_walk_pgd(&st.ptdump, &init_mm); + ptdump_walk_pgd(&st.ptdump, &init_mm, NULL); =20 if (st.wx_pages || st.uxn_pages) { pr_warn("Checked W+X mappings: FAILED, %lu W+X pages found, %lu non-UXN = pages found\n", diff --git a/mm/vmalloc.c b/mm/vmalloc.c index 1fa9ac6e43d4..400563ac6d5d 100644 --- a/mm/vmalloc.c +++ b/mm/vmalloc.c @@ -170,10 +170,7 @@ static int vmap_try_huge_pmd(pmd_t *pmd, unsigned long= addr, unsigned long end, * Therefore, acquire the mmap read lock to prevent use-after-free when * freeing page tables. */ -#ifndef CONFIG_ARM64 - scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) -#endif - { + scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) { if (!pmd_free_pte_page(pmd, addr)) return 0; return pmd_set_huge(pmd, phys_addr, prot); @@ -230,10 +227,7 @@ static int vmap_try_huge_pud(pud_t *pud, unsigned long= addr, unsigned long end, return pud_set_huge(pud, phys_addr, prot); =20 /* See comment in vmap_try_huge_pmd(). */ -#ifndef CONFIG_ARM64 - scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) -#endif - { + scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) { if (!pud_free_pmd_page(pud, addr)) return 0; return pud_set_huge(pud, phys_addr, prot); @@ -290,10 +284,7 @@ static int vmap_try_huge_p4d(p4d_t *p4d, unsigned long= addr, unsigned long end, return p4d_set_huge(p4d, phys_addr, prot); =20 /* See comment in vmap_try_huge_pmd(). */ -#ifndef CONFIG_ARM64 - scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) -#endif - { + scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) { if (!p4d_free_pud_page(p4d, addr)) return 0; return p4d_set_huge(p4d, phys_addr, prot); --=20 2.55.0