From nobody Sat Jul 25 04:54:45 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 12EE7175A5; Fri, 17 Jul 2026 17:30:55 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784309457; cv=none; b=WO0w8zty013K6SBkBoU43tp8K9ku+pB2OVzJa0TCUD7lcltPboS1Z8Cp5JzREWg34g4yn003zsCPiW97pbbeesglyq8bl6n/PaXwEL/iWdXVndohsTrAK8RLeax0LC4lGhhfEvtxvVyUL0BoUirz2q9ETLtTMc/dDcsXpXzm4ro= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784309457; c=relaxed/simple; bh=y0jFTQOJ+Dq2wgdO+3QOYkjdotoBMq7hRls4SEbmZAo=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=s/mHqX/Yzn6mG03/s2XtoB6A4hC9nGr8D9WxNr3iCRFh90psp5ND/5nxB2w7/Kr/EBW1sCDOPaWgDpeSD0CET9qbIMbtb9INjbXiBr6o4AEiWhgSpZnTybm3QF9ogSDqqRKcFT+6iVueJ1jqEImZZXzlY5TEr9cif5Yj+q85nRI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=WASkYcCH; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="WASkYcCH" Received: by smtp.kernel.org (Postfix) with ESMTPSA id DD3D11F00A3A; Fri, 17 Jul 2026 17:30:48 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784309455; bh=VUDPRdQphPMrECclGKpOqw0nMGRICTbCUk3WOCCAA1A=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=WASkYcCHBUDKKnfXbCGeSPf5kDOuVpR194ki/SN/lZBzNwpojwTWPoripPyE9IPsJ vOrVxnqDgxjwc6ws0TCpUqLwQ/he31BZlqLBBRUOfP0SdKrV/Iv6iEhbNC9i3oS3g8 C58UuxqQ8feH9Zm1RI0pm8dE3csGJ1T12b6aedALWnsNujf8UUGXRKTotHPUXLZ3Tt d3HtgklZbEN3/x8UI1VnT3eeikE2xFMpWheBa03vTFyzaCws0xf9j64ahg99kuJ+jH dEQtGV2bTJmTV22NROODyaTv9Wd+UfiK3W1Ta373JBqv8/ydetOjC7yr19o8/7FgVh PcIOIQ3vmPxLA== From: "Lorenzo Stoakes (ARM)" Date: Fri, 17 Jul 2026 18:30:07 +0100 Subject: [PATCH mm-hotfixes v5 1/5] mm/vmalloc: acquire init_mm lock on huge vmap to avoid ptdump UAF Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260717-series-vmap-race-fix-v5-1-606a0ac6d3e5@kernel.org> References: <20260717-series-vmap-race-fix-v5-0-606a0ac6d3e5@kernel.org> In-Reply-To: <20260717-series-vmap-race-fix-v5-0-606a0ac6d3e5@kernel.org> To: Andrew Morton , Suren Baghdasaryan , "Liam R. Howlett" , Vlastimil Babka , Shakeel Butt , David Hildenbrand , Mike Rapoport , Michal Hocko , Uladzislau Rezki , Toshi Kani , Dave Hansen , Andy Lutomirski , Peter Zijlstra , Thomas Gleixner , Ingo Molnar , Borislav Petkov , x86@kernel.org, "H. Peter Anvin" , Kiryl Shutsemau , Catalin Marinas , Will Deacon , Dev Jain , Ryan Roberts Cc: David Carlier , ljs@kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, "Denis V. Lunev" , stable@vger.kernel.org, syzbot+fd95a72470f5a44e464c@syzkaller.appspotmail.com X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=8900; i=ljs@kernel.org; h=from:subject:message-id; bh=y0jFTQOJ+Dq2wgdO+3QOYkjdotoBMq7hRls4SEbmZAo=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLKi0rZpNpubpcYWZK9ZZCXWeYRjyTnzwJRjiv3/f31Wv /X2plVyRykLgxgXg6yYIsvzL+L7g0TC5nVe8HeDmcPKBDKEgYtTACbSs4Dhr6TFMv7zJffK2ITj 9vJeYVcwufHql8Cbu33yMyovtfxQVmJkaHgwQTdbzOdzW8jtuo6tP4SfhuswGUbckcy6uKK7vuc ODwA= X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 Currently there is a nasty race between ptdump and vmap when attempting to map a huge P4D, PUD or PMD entry: * ptdump walks kernel page table ranges it doesn't own. * When vmap maps ranges it tries to promotes existing ones to huge page tables in vmap_try_huge_[p4d,pud,pmd]() at P4D, PUD and PMD level, freeing the lower page table in [p4d,pud,pmd]_free_[pud,pmd,pte]_page() when it succeeds. Both of these things can happen at the same time and as a result ptdump can access a freed page table, resulting in a use-after-free and memory corruption. This is possible because while ptdump_walk_pgd() holds both the mem hotplug lock and the mmap write lock before invoking walk_page_range_debug(), vmap takes no relevant locks at all. Fix this by holding the mmap read lock in vmap_try_huge_*() when freeing page tables. The read lock is sufficient: ptdump is the only walker that must be excluded and it holds the mmap write lock. Other holders of the read lock may run concurrently, but each exclusively owns the range it operates on and cannot reach the page tables freed here. We also hold the lock while assigning the huge page table entry, which means page table walkers observe only the huge or non-huge page table entry. We use a trylock to prevent ptdump from blocking vmap making forward progress. This is fine because it's an optimisation in any case, and thus the vmap can safely proceed regardless. All other kernel page table walkers that touch vmalloc ranges either exclusively own the memory walked or acquire the mmap lock, so this correctly excludes those walkers. One wrinkle here is commit fa93b45fd397 ("arm64: Enable vmalloc-huge with ptdump"), which addresses the issue for arm64 only by explicitly acquiring the mmap read lock on kernel page table freeing should a concurrent ptdump be in progress. This is problematic as vmap may acquire the mmap read lock prior to ptdump attempting to acquire an mmap write lock, leading to a deadlock when the mmap read lock is slept upon on page table freeing due to rwsem anti-starvation. We work around this by predicating the mmap lock being taken on !CONFIG_ARM64 for the time being. With this patch applied, a follow up will partially revert commit fa93b45fd397 ("arm64: Enable vmalloc-huge with ptdump") and at that stage remove the arm64 ifdeffery. We also update walk_page_range_debug() to assert the mmap write lock unconditionally and update the comment here to reflect this change. The issue has existed as long as ptdump was available and vmap freed page tables when promoting to a huge leaf entry, that is, since commit b6bdb7517c3d ("mm/vmalloc: add interfaces to free unmapped page table") for huge ioremap, and commit 121e6f3258fe ("mm/vmalloc: hugepage vmalloc mappings") for huge vmalloc. Since the former is the earlier of the two we choose that for our Fixes tag. We also define a guard class for mmap_read_trylock() so we can use cleanup.h to make the scope handling cleaner in the implementation. This patch is based on work by David Carlier (linked), with gratitude! Fixes: b6bdb7517c3d ("mm/vmalloc: add interfaces to free unmapped page tabl= e") Cc: stable@vger.kernel.org Reported-by: syzbot+fd95a72470f5a44e464c@syzkaller.appspotmail.com Closes: https://lore.kernel.org/all/6a287988.39669fcc.33b062.00a0.GAE@googl= e.com/T/ Link: https://lore.kernel.org/linux-mm/20260706203128.162335-1-devnexen@gma= il.com/ Reviewed-by: Mike Rapoport (Microsoft) Reviewed-by: Dev Jain Acked-by: David Hildenbrand (Arm) Reviewed-by: Kiryl Shutsemau Signed-off-by: Lorenzo Stoakes (ARM) --- include/linux/mmap_lock.h | 1 + mm/pagewalk.c | 22 +++++++++++---------- mm/vmalloc.c | 49 ++++++++++++++++++++++++++++++++++++++-----= ---- 3 files changed, 53 insertions(+), 19 deletions(-) diff --git a/include/linux/mmap_lock.h b/include/linux/mmap_lock.h index 04b8f61ece5d..6b5c2390cc30 100644 --- a/include/linux/mmap_lock.h +++ b/include/linux/mmap_lock.h @@ -621,6 +621,7 @@ static inline void mmap_read_unlock(struct mm_struct *m= m) =20 DEFINE_GUARD(mmap_read_lock, struct mm_struct *, mmap_read_lock(_T), mmap_read_unlock(_T)) +DEFINE_GUARD_COND(mmap_read_lock, _try, mmap_read_trylock(_T)) =20 static inline void mmap_read_unlock_non_owner(struct mm_struct *mm) { diff --git a/mm/pagewalk.c b/mm/pagewalk.c index 3ae2586ff45b..bbcfd68d0907 100644 --- a/mm/pagewalk.c +++ b/mm/pagewalk.c @@ -678,6 +678,8 @@ int walk_kernel_page_table_range_lockless(unsigned long= start, unsigned long end * will also not lock the PTEs for the pte_entry() callback. * * This is for debugging purposes ONLY. + * + * The mmap write lock must be held. */ int walk_page_range_debug(struct mm_struct *mm, unsigned long start, unsigned long end, const struct mm_walk_ops *ops, @@ -691,6 +693,16 @@ int walk_page_range_debug(struct mm_struct *mm, unsign= ed long start, .no_vma =3D true }; =20 + /* + * When walking userland page tables, an mmap write lock must be held to + * account for munmap() downgrading to an mmap read lock when tearing + * down page tables. + * + * When walking kernel page tables, an mmap write lock must also be held + * to account for page table freeing on vmap huge page mapping. + */ + mmap_assert_write_locked(mm); + /* For convenience, we allow traversal of kernel mappings. */ if (mm =3D=3D &init_mm) return walk_kernel_page_table_range(start, end, ops, @@ -700,16 +712,6 @@ int walk_page_range_debug(struct mm_struct *mm, unsign= ed long start, if (!check_ops_safe(ops)) return -EINVAL; =20 - /* - * The mmap lock protects the page walker from changes to the page - * tables during the walk. However a read lock is insufficient to - * protect those areas which don't have a VMA as munmap() detaches - * the VMAs before downgrading to a read lock and actually tearing - * down PTEs/page tables. In which case, the mmap write lock should - * be held. - */ - mmap_assert_write_locked(mm); - return walk_pgd_range(start, end, &walk); } =20 diff --git a/mm/vmalloc.c b/mm/vmalloc.c index 1afca3568b9b..d5c4d2bb770b 100644 --- a/mm/vmalloc.c +++ b/mm/vmalloc.c @@ -43,6 +43,7 @@ #include #include #include +#include =20 #define CREATE_TRACE_POINTS #include @@ -158,10 +159,24 @@ static int vmap_try_huge_pmd(pmd_t *pmd, unsigned lon= g addr, unsigned long end, if (!IS_ALIGNED(phys_addr, PMD_SIZE)) return 0; =20 - if (pmd_present(*pmd) && !pmd_free_pte_page(pmd, addr)) - return 0; + if (!pmd_present(*pmd)) + return pmd_set_huge(pmd, phys_addr, prot); =20 - return pmd_set_huge(pmd, phys_addr, prot); + /* + * Acquire the mmap read lock to exclude ptdump, which walks + * kernel page tables it does not own under the mmap write lock. + * + * Concurrent read lock holders are safe: each exclusively owns + * the range it operates on and cannot reach this page table. + */ +#ifndef CONFIG_ARM64 + scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) +#endif + { + if (!pmd_free_pte_page(pmd, addr)) + return 0; + return pmd_set_huge(pmd, phys_addr, prot); + } } =20 static int vmap_pmd_range(pud_t *pud, unsigned long addr, unsigned long en= d, @@ -210,10 +225,18 @@ static int vmap_try_huge_pud(pud_t *pud, unsigned lon= g addr, unsigned long end, if (!IS_ALIGNED(phys_addr, PUD_SIZE)) return 0; =20 - if (pud_present(*pud) && !pud_free_pmd_page(pud, addr)) - return 0; + if (!pud_present(*pud)) + return pud_set_huge(pud, phys_addr, prot); =20 - return pud_set_huge(pud, phys_addr, prot); + /* See comment in vmap_try_huge_pmd(). */ +#ifndef CONFIG_ARM64 + scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) +#endif + { + if (!pud_free_pmd_page(pud, addr)) + return 0; + return pud_set_huge(pud, phys_addr, prot); + } } =20 static int vmap_pud_range(p4d_t *p4d, unsigned long addr, unsigned long en= d, @@ -262,10 +285,18 @@ static int vmap_try_huge_p4d(p4d_t *p4d, unsigned lon= g addr, unsigned long end, if (!IS_ALIGNED(phys_addr, P4D_SIZE)) return 0; =20 - if (p4d_present(*p4d) && !p4d_free_pud_page(p4d, addr)) - return 0; + if (!p4d_present(*p4d)) + return p4d_set_huge(p4d, phys_addr, prot); =20 - return p4d_set_huge(p4d, phys_addr, prot); + /* See comment in vmap_try_huge_pmd(). */ +#ifndef CONFIG_ARM64 + scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) +#endif + { + if (!p4d_free_pud_page(p4d, addr)) + return 0; + return p4d_set_huge(p4d, phys_addr, prot); + } } =20 static int vmap_p4d_range(pgd_t *pgd, unsigned long addr, unsigned long en= d, --=20 2.55.0 From nobody Sat Jul 25 04:54:45 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C651F37268D; Fri, 17 Jul 2026 17:31:02 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784309464; cv=none; b=Ivir/eKIq0Ns+00IuN1PHgepx8Hl3toj+jppbOGoeRHXIjdznOV0bC9jZIzZsENGS/1BA7kxhb3Y64dGqrXVU+qoM2RspfsBenzGvlOjLgXHiueuVjFIfSkAIM9BUCy/S72Ar1YtsiTVH7tna59AsK2gp3Ag42cHcx/z+fFA2No= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784309464; c=relaxed/simple; bh=3isjtK28wb1BiLlMBIpy71VvY1jYYrdz7YUUYuIwCk8=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=qvhvMh2CMuFBK5cijnl8ODRRhFwGKnz1IVf4rGKcBHiWvRKr6bf70Q6VYqp41pHOttRLyotHFyaFr0Ze8IpkJsZ8H3E0yZvd88xbpZ2KOPGneomiTzz/0Xc/TjIH2KPolYzdxz6pMyVgE5WSXcminA89iygWUsSTNvJwMylh8XA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=DYQijovd; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="DYQijovd" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 1CF911F000E9; Fri, 17 Jul 2026 17:30:55 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784309462; bh=jGfvVz5phYfszb4vs/wrQwKbo5YfQ9htQgjlD4h0uzo=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=DYQijovd054GG2wU3z7XMoKsnqNogdfpJ+OCYKtauwMoBZtvbLhprNGmeh+1j3wkt /kqqZmxWLbM2EfHqTuQPWREroE7kObjwev+cigZ3gExShhUSE05tjJ/eLczaGqQ4qK cFj2JveVjO1V9Cc0nFDhAiWYGGHJABDU4EuJpvsG4w7U0cS8/Le4Pgl4xh1hR7DIEw p+wyed3zInNffSwUkz/PC7sGtzuaio9mHkKltJaAb5WkHCqTvRLYGjvjVnZamvd6rx 0q8ZQe5y1FGmXnpocY1jnKhEZrute7gnN8qhFvMfonPTZ618WoLrJQl3KtWeZGP8Az mS3L78Qs9gIpw== From: "Lorenzo Stoakes (ARM)" Date: Fri, 17 Jul 2026 18:30:08 +0100 Subject: [PATCH mm-hotfixes v5 2/5] x86/mm/pat: acquire init_mm write lock on collapse to avoid UAF Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260717-series-vmap-race-fix-v5-2-606a0ac6d3e5@kernel.org> References: <20260717-series-vmap-race-fix-v5-0-606a0ac6d3e5@kernel.org> In-Reply-To: <20260717-series-vmap-race-fix-v5-0-606a0ac6d3e5@kernel.org> To: Andrew Morton , Suren Baghdasaryan , "Liam R. Howlett" , Vlastimil Babka , Shakeel Butt , David Hildenbrand , Mike Rapoport , Michal Hocko , Uladzislau Rezki , Toshi Kani , Dave Hansen , Andy Lutomirski , Peter Zijlstra , Thomas Gleixner , Ingo Molnar , Borislav Petkov , x86@kernel.org, "H. Peter Anvin" , Kiryl Shutsemau , Catalin Marinas , Will Deacon , Dev Jain , Ryan Roberts Cc: David Carlier , ljs@kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, "Denis V. Lunev" , stable@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=4047; i=ljs@kernel.org; h=from:subject:message-id; bh=3isjtK28wb1BiLlMBIpy71VvY1jYYrdz7YUUYuIwCk8=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLKi0rZJZfsl737mbpD55EnArs8aCv9C575vYr5c8sBJ+ E1ubYdpRykLgxgXg6yYIsvzL+L7g0TC5nVe8HeDmcPKBDKEgYtTACZy6irD/+hLhquku2w+zP3V zF6zRaj+pt3Je7ZPdoaVxEr/2t40/SQjw+fK9LjUr3KzDSfMK5f6Oj/vWOWjymcrT+ctUGHd4mP yhhUA X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 x86 implements page attribute modification using its Change Page Attributes (CPA) mechanism. This tracks properties of ranges such as cache mode through x86 page attributes, and as part of that logic manipulates kernel page tables. Since commit 41d88484c71c ("x86/mm/pat: restore large ROX pages after fragmentation") ranges of kernel page table entries can be collapsed into huge page table entries as part of this logic. As part of this collapse, it frees the page tables which the collapsed entries previously pointed to, and it does so without any relevant locks being held to preclude concurrent kernel page table walkers. The only way this code can be reached is if CPA_COLLAPSE is specified, and this is only set in set_memory_rox() via: set_memory_rox() -> change_page_attr_set_clr() -> cpa_flush() -> cpa_collapse_large_pages() Notable users of this are execmem and bpf when manipulating executable mappings. However, this is problematic for ptdump as it walks ranges it does not own and thus runs the risk of a use-after-free on page tables freed underneath it. In addition, concurrent CPA collapse operations are possible which can also cause races. Resolve the issue by acquiring the mmap write lock on init_mm across the whole operation. It is safe to acquire a sleeping lock as all the callers invoke set_memory_rox() from process context and in any case, change_page_attr_set_clr() calls vm_unmap_alias() which ultimately takes a mutex, disallowing atomic context here. Fixes: 41d88484c71c ("x86/mm/pat: restore large ROX pages after fragmentati= on") Cc: stable@vger.kernel.org Reviewed-by: Mike Rapoport (Microsoft) Reviewed-by: Kiryl Shutsemau (Meta) Reviewed-by: David Hildenbrand (Arm) Reviewed-by: Dave Hansen Reviewed-by: Will Deacon Reviewed-by: David Carlier Signed-off-by: Lorenzo Stoakes (ARM) --- arch/x86/mm/pat/set_memory.c | 15 ++++++++++++++- include/linux/mmap_lock.h | 2 ++ 2 files changed, 16 insertions(+), 1 deletion(-) diff --git a/arch/x86/mm/pat/set_memory.c b/arch/x86/mm/pat/set_memory.c index d023a40a1e03..d1e63f7d267f 100644 --- a/arch/x86/mm/pat/set_memory.c +++ b/arch/x86/mm/pat/set_memory.c @@ -22,6 +22,7 @@ #include #include #include +#include =20 #include #include @@ -410,7 +411,7 @@ static void __cpa_flush_tlb(void *data) =20 static int collapse_large_pages(unsigned long addr, struct list_head *pgta= bles); =20 -static void cpa_collapse_large_pages(struct cpa_data *cpa) +static void __cpa_collapse_large_pages(struct cpa_data *cpa) { unsigned long start, addr, end; struct ptdesc *ptdesc, *tmp; @@ -442,6 +443,18 @@ static void cpa_collapse_large_pages(struct cpa_data *= cpa) } } =20 +static void cpa_collapse_large_pages(struct cpa_data *cpa) +{ + /* + * Take the mmap write lock on init_mm to: + * - Avoid a use-after-free if raced by ptdump (which takes its own + * write lock on init_mm). + * - Serialise concurrent CPA walkers. + */ + scoped_guard(mmap_write_lock, &init_mm) + __cpa_collapse_large_pages(cpa); +} + static void cpa_flush(struct cpa_data *cpa, int cache) { unsigned int i; diff --git a/include/linux/mmap_lock.h b/include/linux/mmap_lock.h index 6b5c2390cc30..047f5f5e2c34 100644 --- a/include/linux/mmap_lock.h +++ b/include/linux/mmap_lock.h @@ -621,6 +621,8 @@ static inline void mmap_read_unlock(struct mm_struct *m= m) =20 DEFINE_GUARD(mmap_read_lock, struct mm_struct *, mmap_read_lock(_T), mmap_read_unlock(_T)) +DEFINE_GUARD(mmap_write_lock, struct mm_struct *, + mmap_write_lock(_T), mmap_write_unlock(_T)) DEFINE_GUARD_COND(mmap_read_lock, _try, mmap_read_trylock(_T)) =20 static inline void mmap_read_unlock_non_owner(struct mm_struct *mm) --=20 2.55.0 From nobody Sat Jul 25 04:54:45 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1D645353A8F; Fri, 17 Jul 2026 17:31:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784309471; cv=none; b=BVlNrVLaA7UsBoIBgeHaZnLmKNSDagCzP8NnUNlXjX6eOxR1gkRyYiztyWypt7gObwawh7fxFV1VJRFh5M9m7xtmersqhYIZA5JVMvFq5dzn9XcwcczJaVK43YHJb90EaM8qRwbUHYMgyWwvcgzqB5VW+nxVcKUt57ipqULDULQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784309471; c=relaxed/simple; bh=FncxTVoaDCe8XgBgxmxR3F4pYkqYCxDSgA+rKimYCJI=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=fdjIdTtApgiO8ukvF284iuErTLSdlHozW8f7HKyHINea4b29pWXtUm19TY+0eEm1NmF4UjsdEn5Hs8wieoqbjAPOL7r5IGhtcKzxRqjmNNLf0MQcs/hmuEaHE3iazOT6wCZT0TCLjn89UeWGQm2Xs850gP6JLiCI8D+scxNlCCo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=mW7ymBxJ; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="mW7ymBxJ" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 2FAC61F00A3D; Fri, 17 Jul 2026 17:31:02 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784309469; bh=adqfM1VF52UApKZ4esto4yd6ee7VDcwIxlypGM7JBY0=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=mW7ymBxJXJPEnbbaFA0UsYPLhTxbTPMblfLKS7QeqW568DbUxgaRUY7zLQHxGBoBV 86680uEMKDxD8AyU/EC/HTv00K7gMLCsLcIMgFw/h4+sd1JxGQP12ypOycjVQk4sQ4 XBw5xkgSCJZHStzGysvvQFQc3q+2SLQya0hcPMCY5poXVoUzXJOJktSpfzSwZuLWcP l6v+rJjiTfPQIRFJXhiCRlaPUFoIlfz+2DFQ7bdvXQzDmIezeWGk5RAeuBcpAaCee0 9I+8T6Sqetq+fpVtXU+xXV2uy6o8CWbtwIW34asPTGI6DYAusFhqkhi3WBhe+G+ICz XlFmoeK/VQs2g== From: "Lorenzo Stoakes (ARM)" Date: Fri, 17 Jul 2026 18:30:09 +0100 Subject: [PATCH mm-hotfixes v5 3/5] x86/mm/pat: acquire init_mm read lock on attribute change to avoid UAF Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260717-series-vmap-race-fix-v5-3-606a0ac6d3e5@kernel.org> References: <20260717-series-vmap-race-fix-v5-0-606a0ac6d3e5@kernel.org> In-Reply-To: <20260717-series-vmap-race-fix-v5-0-606a0ac6d3e5@kernel.org> To: Andrew Morton , Suren Baghdasaryan , "Liam R. Howlett" , Vlastimil Babka , Shakeel Butt , David Hildenbrand , Mike Rapoport , Michal Hocko , Uladzislau Rezki , Toshi Kani , Dave Hansen , Andy Lutomirski , Peter Zijlstra , Thomas Gleixner , Ingo Molnar , Borislav Petkov , x86@kernel.org, "H. Peter Anvin" , Kiryl Shutsemau , Catalin Marinas , Will Deacon , Dev Jain , Ryan Roberts Cc: David Carlier , ljs@kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, "Denis V. Lunev" , stable@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=2831; i=ljs@kernel.org; h=from:subject:message-id; bh=FncxTVoaDCe8XgBgxmxR3F4pYkqYCxDSgA+rKimYCJI=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLKi0rbrCT16ub2mRjsq5L1Ck+OmWsWrey2ezPoumvndp //7rluuHaUsDGJcDLJiiizPv4jvDxIJm9d5wd8NZg4rE8gQBi5OAZjI8lSG/463b91xC/KqPrxk pv+ST1+Lb6zYqchaXR3jecN8wv/f+nMZGdbd3WG6dcPVjzoTrIqfHzq179POoAUh3IsL5wj9t9L dvp8HAA== X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 A previous commit protected us against races between ptdump and CPA collapse, however one still exists between attribute changes and collapse as reported by Denis V. Lunev (linked). When an attribute change arises, a lockless page table walker obtains a PTE entry, which is later written to via set_pte_atomic(): ... -> change_page_attr_set_clr() -> __change_page_attr_set_clr() -> __change_page_attr() -> _lookup_address_cpa() -> lookup_address_in_pgd_attr() -> [ lockless page table walker ] -> set_pte_atomic() There is nothing preventing a concurrent CPA collapse which can free the PTE that was retrieved here, resulting in a use-after-free. With the mmap write lock taken on init_mm over CPA collapse, we can now resolve this race by acquiring an mmap read lock on init_mm over __change_page_attr_set_clr(). This locks across the whole operation over which the walk and the PTE entry write occurs, solving the race. It is safe to do this here, as no spinlocks are held upon entry to __change_page_attr_set_clr(). The CPA_COLLAPSE flag is only set by set_memory_rox(), which exclusively operates upon vmalloc ranges, and on x86 only within the module mapping space. This is important, because some callers directly invoke __change_page_attr_set_clr(), bypassing this lock. However, none of these operate within the module mapping space. * cpa_process_alias() - a recursive helper called by __change_page_attr_set_clr(). * __set_memory_enc_pgtable() - operates on the direct mapping and (via __vmbus_establish_gpadl()) the vmalloc mapping space. * __set_pages_[n]p() - called by set_direct_map_[invalid, default, valid]_noflush(), __kernel_map_pages() - operates on the direct map. * kernel_[un]map_pages_in_pgd() - operates on EFI ranges. This work is based upon Denis V. Lunev's excellent analysis of the bug with gratitude. Link: https://lore.kernel.org/all/20260626163213.2284080-1-den@openvz.org/ Fixes: 41d88484c71c ("x86/mm/pat: restore large ROX pages after fragmentati= on") Cc: stable@vger.kernel.org Signed-off-by: Lorenzo Stoakes (ARM) --- arch/x86/mm/pat/set_memory.c | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/arch/x86/mm/pat/set_memory.c b/arch/x86/mm/pat/set_memory.c index d1e63f7d267f..301fb9e77d91 100644 --- a/arch/x86/mm/pat/set_memory.c +++ b/arch/x86/mm/pat/set_memory.c @@ -2122,7 +2122,9 @@ static int change_page_attr_set_clr(unsigned long *ad= dr, int numpages, cpa.curpage =3D 0; cpa.force_split =3D force_split; =20 - ret =3D __change_page_attr_set_clr(&cpa, 1); + /* Avoid race with concurrent CPA collapse. */ + scoped_guard(mmap_read_lock, &init_mm) + ret =3D __change_page_attr_set_clr(&cpa, 1); =20 /* * Check whether we really changed something: --=20 2.55.0 From nobody Sat Jul 25 04:54:45 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C44812E091B; Fri, 17 Jul 2026 17:31:16 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784309477; cv=none; b=iPbekwi7x730GvGdTHjjkpDXdvhR94uVIQEP8IptcEUbm69er8u+hI0Ca+Gfdb/bcYsiq6aHzxh1XYa5nQlMNOwrryHhi428uazoecM1Kar6eK4dEw9QqIBNslHLgqVTjD9P9t2jK2bfNSrm16SHNXJt2tl7TCaUSR0mhFVayAw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784309477; c=relaxed/simple; bh=VITXU838ytmI5wjkeVjJjzJ76fi6Ga2j8Ljay/OkRtw=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=WkFD7HoUEj2DGT9Xy2YEdGJcU4CVj8Pyg6NUl154IJZnmpIElON0pBdxtfCWH2ho6X9UGritfhyMWo/mJtiXF8UIYHkDEDEVr5WrfwKWX//ZDhbos/E0RoW53kzK1vs17n7fxJEps/yHMJ9EZAIdEe61+O8wQE0ahpsfd3eQgf8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=j7PegHdq; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="j7PegHdq" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 30E8F1F000E9; Fri, 17 Jul 2026 17:31:09 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784309476; bh=4TmiXyJvQ92TKxUnkF2EBAbSx4BVAEx8E3qvkE5M06w=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=j7PegHdq1OFesFoUaayKABXnDJpVNSvSPDwYaiVx8/2g+pwbNCJ3ZfJja6eAtQZ8p cIyc+DGp8cE1mY5T3MVxHeoEgCBsIQO/nHC6nOlsBe2FoOyLceFjBJ0vPPfhsUb9SM 5m/5O+dZndx9pKSDw52Ga8CFym1IitqeHcVgrB3OBwnEfcA6QcwjBgRM++V4FQwvYi aOZwa7704A9u7bk3CoW7eomhkOF0zM5mBg5cz6qywFa0Ar92EAzklWTyarUQg4dv7m BSQW/EW4hzjEqX2YoDLXo5uYFzkwDvt0LRfUvjOsP2TYV3cvXr7QKgDnltalcad/JW +Q9P+xVsQPKwA== From: "Lorenzo Stoakes (ARM)" Date: Fri, 17 Jul 2026 18:30:10 +0100 Subject: [PATCH mm-hotfixes v5 4/5] mm/ptdump: always stabilise against page table freeing using init_mm Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260717-series-vmap-race-fix-v5-4-606a0ac6d3e5@kernel.org> References: <20260717-series-vmap-race-fix-v5-0-606a0ac6d3e5@kernel.org> In-Reply-To: <20260717-series-vmap-race-fix-v5-0-606a0ac6d3e5@kernel.org> To: Andrew Morton , Suren Baghdasaryan , "Liam R. Howlett" , Vlastimil Babka , Shakeel Butt , David Hildenbrand , Mike Rapoport , Michal Hocko , Uladzislau Rezki , Toshi Kani , Dave Hansen , Andy Lutomirski , Peter Zijlstra , Thomas Gleixner , Ingo Molnar , Borislav Petkov , x86@kernel.org, "H. Peter Anvin" , Kiryl Shutsemau , Catalin Marinas , Will Deacon , Dev Jain , Ryan Roberts Cc: David Carlier , ljs@kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, "Denis V. Lunev" , stable@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=3325; i=ljs@kernel.org; h=from:subject:message-id; bh=VITXU838ytmI5wjkeVjJjzJ76fi6Ga2j8Ljay/OkRtw=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLKi0rZ/XWP6/n/eThWz4m8aOyLm1Qpq9ujrr9sb/v3mt bR522oudJSyMIhxMciKKbI8/yK+P0gkbF7nBX83mDmsTCBDGLg4BWAih+MYfjHL3hNP3Nvg+Jtp 1eU2vcCDZsmnLsjIzVir+DA4257NqJnhf4l0GpOQ+mypRYdmc1x2ZFwq37jKX7A5avOxL9Nz3lU KMQAA X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 Previous commits have established the invariant that kernel page table freeing is performed while an mmap read lock on init_mm is held, which fixes races between ptdump and kernel page table freeing over init_mm. However, x86 and arm64 can perform a ptdump over an mm other than init_mm via ptdump_walk_pgd() and since kernel memory ranges are shared across non-kernel mm's, this means that the race still exists for these cases. Fix this by acquiring a nested mmap write lock for init_mm in ptdump_walk_pgd(). This is safe as we take this after mmap write locking the mm, and nothing acquires the init_mm lock first before locking an arbitrary mm, so no deadlock is possible. Also update walk_page_range_debug() to assert that init_mm is write locked, add a comment explaining why and remove some redundant code, and eliminate the unnecessary and confusing invocation of walk_kernel_page_table_range(). We can safely remove the non-NULL check for walk.mm, as the mmap lock asserts would NULL pointer deref if it was (and of course no callers do this). The first point at which ptdump can race kernel page table freeing is commit b6bdb7517c3d ("mm/vmalloc: add interfaces to free unmapped page table"), so we target this in the Fixes tag. Fixes: b6bdb7517c3d ("mm/vmalloc: add interfaces to free unmapped page tabl= e") Cc: stable@vger.kernel.org Reviewed-by: Mike Rapoport (Microsoft) Acked-by: David Hildenbrand (Arm) Reviewed-by: Kiryl Shutsemau Signed-off-by: Lorenzo Stoakes (ARM) --- mm/pagewalk.c | 14 +++++++++----- mm/ptdump.c | 7 +++++++ 2 files changed, 16 insertions(+), 5 deletions(-) diff --git a/mm/pagewalk.c b/mm/pagewalk.c index bbcfd68d0907..5d87c632a255 100644 --- a/mm/pagewalk.c +++ b/mm/pagewalk.c @@ -702,12 +702,16 @@ int walk_page_range_debug(struct mm_struct *mm, unsig= ned long start, * to account for page table freeing on vmap huge page mapping. */ mmap_assert_write_locked(mm); + /* + * x86, arm64 ptdump allow walks of efi mm's and x86 ptdump allows walks + * of arbitrary mm's. + * + * However, they both must also hold the init_mm lock to account for + * concurrent kernel page table freeing. + */ + mmap_assert_write_locked(&init_mm); =20 - /* For convenience, we allow traversal of kernel mappings. */ - if (mm =3D=3D &init_mm) - return walk_kernel_page_table_range(start, end, ops, - pgd, private); - if (start >=3D end || !walk.mm) + if (start >=3D end) return -EINVAL; if (!check_ops_safe(ops)) return -EINVAL; diff --git a/mm/ptdump.c b/mm/ptdump.c index 973020000096..5851096e6f65 100644 --- a/mm/ptdump.c +++ b/mm/ptdump.c @@ -178,11 +178,18 @@ void ptdump_walk_pgd(struct ptdump_state *st, struct = mm_struct *mm, pgd_t *pgd) =20 get_online_mems(); mmap_write_lock(mm); + /* To stabilise kernel page tables we must hold the init_mm lock too. */ + if (mm !=3D &init_mm) + mmap_write_lock_nested(&init_mm, SINGLE_DEPTH_NESTING); + while (range->start !=3D range->end) { walk_page_range_debug(mm, range->start, range->end, &ptdump_ops, pgd, st); range++; } + + if (mm !=3D &init_mm) + mmap_write_unlock(&init_mm); mmap_write_unlock(mm); put_online_mems(); =20 --=20 2.55.0 From nobody Sat Jul 25 04:54:45 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 07F842C11D6; Fri, 17 Jul 2026 17:31:23 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784309485; cv=none; b=mKHgMv96PiOqxK1cpSOADmxaDjeoUSrcwYgGJiOZN8DgxzBzboYklnqJcX5k0SxM8TIrsbC1ZSGcS+A1I/t6gsoKhUq/HvD9W3S6VkPAjDVFsCpedzenpEYW1oin5u4rOzwHKZLgBX+FAVmDnAj/Om4Ujn08iYHPMMBDgmAtIaE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784309485; c=relaxed/simple; bh=kE9XtB6iyw7BSFNAuuPgI01clqwgzMyGQ2yqQHsLjJM=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=IT+7aRYBxf7xwFk60eMLdQ+o4xyhve1FTtl56NBPlU++jgTKd3+vWWQFep7NTq0cJdHH+4QAVjRpbH9m87wqjxmHtka9lqgZ7GjoLHRUQUAJ3PR1w4I5Pd33qu6VfelRI0ihwQ/p8n2dnNBJn4aknglK2d2ATN5bWqdD8q1E0nk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=Wxu9zgqI; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="Wxu9zgqI" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 31A7C1F00A3D; Fri, 17 Jul 2026 17:31:16 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784309483; bh=6tB+p5fpJihvWmKYtOqn6XBXbSKHjdM1g/8Wc7lnKAc=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=Wxu9zgqIph5rtq6bTKYeHsaiTgxmWpTOwq98kp7/FgXG55S4zw0mETpYJT0yfbyPq v4V5g+nDLg/ghKYBi5mk8rvjZlAhytiKJcx45c4Brl9l2yYq3Oq6NsJfFC417wfHzQ nEBMYY/F2Hn0ehBJGokTLjGbrdWaaUUhz4wU8dEmGho81WigKlnV7ujaEaUP1nKnn2 9KjKwty8unA4lldmoJS4SwYuTYkCO8/kkwO6cuknvoWVtcbs1W8JJfUWjmO6QUDDLH dBZTx4vjOUiYj9omJD2y5FfAmxxaZXP1/vo2uMdE8m4CbzOsEyj2QKTHzh5//tJKCN LXQOR0tArV9qA== From: "Lorenzo Stoakes (ARM)" Date: Fri, 17 Jul 2026 18:30:11 +0100 Subject: [PATCH mm-hotfixes v5 5/5] arm64: remove redundant concurrent ptdump UAF mitigation Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260717-series-vmap-race-fix-v5-5-606a0ac6d3e5@kernel.org> References: <20260717-series-vmap-race-fix-v5-0-606a0ac6d3e5@kernel.org> In-Reply-To: <20260717-series-vmap-race-fix-v5-0-606a0ac6d3e5@kernel.org> To: Andrew Morton , Suren Baghdasaryan , "Liam R. Howlett" , Vlastimil Babka , Shakeel Butt , David Hildenbrand , Mike Rapoport , Michal Hocko , Uladzislau Rezki , Toshi Kani , Dave Hansen , Andy Lutomirski , Peter Zijlstra , Thomas Gleixner , Ingo Molnar , Borislav Petkov , x86@kernel.org, "H. Peter Anvin" , Kiryl Shutsemau , Catalin Marinas , Will Deacon , Dev Jain , Ryan Roberts Cc: David Carlier , ljs@kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, "Denis V. Lunev" , stable@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=7063; i=ljs@kernel.org; h=from:subject:message-id; bh=kE9XtB6iyw7BSFNAuuPgI01clqwgzMyGQ2yqQHsLjJM=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLKi0rbPmiD/OZTXw4v7yvRv651ei/3cuKD+SZn7s71WQ nr3o9af7ShlYRDjYpAVU2R5/kV8f5BI2LzOC/5uMHNYmUCGMHBxCsBEHn1lZJhx43Hz5ID+kmJe 3r7zmyZwv3D7KhTYJX5kG2fK/jdH7mgw/GTkPXPjdkyIsd0f84dBtVY3N+yRF7h5eeHxcueUi5Y sW1gA X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 This partially reverts commit fa93b45fd397 ("arm64: Enable vmalloc-huge with ptdump"), retaining vmalloc-huge support but eliminating the now redundant mitigation against a race between huge vmap page table freeing and ptdump, as this issue has now been fixed at core. We also simultaneously remove the arm64 if-deffery when acquiring the mmap read lock upon vmap huge page table promotion as it is no longer required. Note that this patch relies on the preceding vmalloc patch, and should not be backported alone. Fixes: fa93b45fd397 ("arm64: Enable vmalloc-huge with ptdump") Cc: stable@vger.kernel.org Reviewed-by: Dev Jain Acked-by: Mike Rapoport (Microsoft) Acked-by: Kiryl Shutsemau (Meta) Acked-by: Will Deacon Reviewed-by: David Hildenbrand (Arm) Signed-off-by: Lorenzo Stoakes (ARM) --- arch/arm64/include/asm/ptdump.h | 2 -- arch/arm64/mm/mmu.c | 43 ++++---------------------------------= ---- arch/arm64/mm/ptdump.c | 11 ++--------- mm/vmalloc.c | 15 +++----------- 4 files changed, 9 insertions(+), 62 deletions(-) diff --git a/arch/arm64/include/asm/ptdump.h b/arch/arm64/include/asm/ptdum= p.h index 5b374a6ab34a..50a195eda8ed 100644 --- a/arch/arm64/include/asm/ptdump.h +++ b/arch/arm64/include/asm/ptdump.h @@ -7,8 +7,6 @@ =20 #include =20 -DECLARE_STATIC_KEY_FALSE(arm64_ptdump_lock_key); - #ifdef CONFIG_PTDUMP =20 #include diff --git a/arch/arm64/mm/mmu.c b/arch/arm64/mm/mmu.c index a25d8beacc83..bd52fca6e872 100644 --- a/arch/arm64/mm/mmu.c +++ b/arch/arm64/mm/mmu.c @@ -49,8 +49,6 @@ #define NO_CONT_MAPPINGS BIT(1) #define NO_EXEC_MAPPINGS BIT(2) /* assumes FEAT_HPDS is not used */ =20 -DEFINE_STATIC_KEY_FALSE(arm64_ptdump_lock_key); - u64 kimage_voffset __ro_after_init; EXPORT_SYMBOL(kimage_voffset); =20 @@ -1864,8 +1862,7 @@ int pmd_clear_huge(pmd_t *pmdp) return 1; } =20 -static int __pmd_free_pte_page(pmd_t *pmdp, unsigned long addr, - bool acquire_mmap_lock) +int pmd_free_pte_page(pmd_t *pmdp, unsigned long addr) { pte_t *table; pmd_t pmd; @@ -1877,25 +1874,13 @@ static int __pmd_free_pte_page(pmd_t *pmdp, unsigne= d long addr, return 1; } =20 - /* See comment in pud_free_pmd_page for static key logic */ table =3D pte_offset_kernel(pmdp, addr); pmd_clear(pmdp); __flush_tlb_kernel_pgtable(addr); - if (static_branch_unlikely(&arm64_ptdump_lock_key) && acquire_mmap_lock) { - mmap_read_lock(&init_mm); - mmap_read_unlock(&init_mm); - } - pte_free_kernel(NULL, table); return 1; } =20 -int pmd_free_pte_page(pmd_t *pmdp, unsigned long addr) -{ - /* If ptdump is walking the pagetables, acquire init_mm.mmap_lock */ - return __pmd_free_pte_page(pmdp, addr, /* acquire_mmap_lock =3D */ true); -} - int pud_free_pmd_page(pud_t *pudp, unsigned long addr) { pmd_t *table; @@ -1911,36 +1896,16 @@ int pud_free_pmd_page(pud_t *pudp, unsigned long ad= dr) } =20 table =3D pmd_offset(pudp, addr); - - /* - * Our objective is to prevent ptdump from reading a PMD table which has - * been freed. In this race, if pud_free_pmd_page observes the key on - * (which got flipped by ptdump) then the mmap lock sequence here will, - * as a result of the mmap write lock/unlock sequence in ptdump, give - * us the correct synchronization. If not, this means that ptdump has - * yet not started walking the pagetables - the sequence of barriers - * issued by __flush_tlb_kernel_pgtable() guarantees that ptdump will - * observe an empty PUD. - */ - pud_clear(pudp); - __flush_tlb_kernel_pgtable(addr); - if (static_branch_unlikely(&arm64_ptdump_lock_key)) { - mmap_read_lock(&init_mm); - mmap_read_unlock(&init_mm); - } - pmdp =3D table; next =3D addr; end =3D addr + PUD_SIZE; do { if (pmd_present(pmdp_get(pmdp))) - /* - * PMD has been isolated, so ptdump won't see it. No - * need to acquire init_mm.mmap_lock. - */ - __pmd_free_pte_page(pmdp, next, /* acquire_mmap_lock =3D */ false); + pmd_free_pte_page(pmdp, next); } while (pmdp++, next +=3D PMD_SIZE, next !=3D end); =20 + pud_clear(pudp); + __flush_tlb_kernel_pgtable(addr); pmd_free(NULL, table); return 1; } diff --git a/arch/arm64/mm/ptdump.c b/arch/arm64/mm/ptdump.c index 1c20144700d7..5a76c59b5ada 100644 --- a/arch/arm64/mm/ptdump.c +++ b/arch/arm64/mm/ptdump.c @@ -283,13 +283,6 @@ void note_page_flush(struct ptdump_state *pt_st) note_page(pt_st, 0, -1, pte_val(pte_zero)); } =20 -static void arm64_ptdump_walk_pgd(struct ptdump_state *st, struct mm_struc= t *mm) -{ - static_branch_inc(&arm64_ptdump_lock_key); - ptdump_walk_pgd(st, mm, NULL); - static_branch_dec(&arm64_ptdump_lock_key); -} - void ptdump_walk(struct seq_file *s, struct ptdump_info *info) { unsigned long end =3D ~0UL; @@ -318,7 +311,7 @@ void ptdump_walk(struct seq_file *s, struct ptdump_info= *info) } }; =20 - arm64_ptdump_walk_pgd(&st.ptdump, info->mm); + ptdump_walk_pgd(&st.ptdump, info->mm, NULL); } =20 static void __init ptdump_initialize(void) @@ -360,7 +353,7 @@ bool ptdump_check_wx(void) } }; =20 - arm64_ptdump_walk_pgd(&st.ptdump, &init_mm); + ptdump_walk_pgd(&st.ptdump, &init_mm, NULL); =20 if (st.wx_pages || st.uxn_pages) { pr_warn("Checked W+X mappings: FAILED, %lu W+X pages found, %lu non-UXN = pages found\n", diff --git a/mm/vmalloc.c b/mm/vmalloc.c index d5c4d2bb770b..f4fa227a8d7f 100644 --- a/mm/vmalloc.c +++ b/mm/vmalloc.c @@ -169,10 +169,7 @@ static int vmap_try_huge_pmd(pmd_t *pmd, unsigned long= addr, unsigned long end, * Concurrent read lock holders are safe: each exclusively owns * the range it operates on and cannot reach this page table. */ -#ifndef CONFIG_ARM64 - scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) -#endif - { + scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) { if (!pmd_free_pte_page(pmd, addr)) return 0; return pmd_set_huge(pmd, phys_addr, prot); @@ -229,10 +226,7 @@ static int vmap_try_huge_pud(pud_t *pud, unsigned long= addr, unsigned long end, return pud_set_huge(pud, phys_addr, prot); =20 /* See comment in vmap_try_huge_pmd(). */ -#ifndef CONFIG_ARM64 - scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) -#endif - { + scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) { if (!pud_free_pmd_page(pud, addr)) return 0; return pud_set_huge(pud, phys_addr, prot); @@ -289,10 +283,7 @@ static int vmap_try_huge_p4d(p4d_t *p4d, unsigned long= addr, unsigned long end, return p4d_set_huge(p4d, phys_addr, prot); =20 /* See comment in vmap_try_huge_pmd(). */ -#ifndef CONFIG_ARM64 - scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) -#endif - { + scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) { if (!p4d_free_pud_page(p4d, addr)) return 0; return p4d_set_huge(p4d, phys_addr, prot); --=20 2.55.0