From nobody Fri Jul 24 05:21:56 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3DFB14756AE; Thu, 23 Jul 2026 15:17:11 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784819835; cv=none; b=jwsRapNOcDhryIChg7bb9ClSkT2liF6NG7qBOEMF/iMePtxMyAbonClYYSlBzF02in6DpYezM096gTmK/s7i6ibyZwaGFG7V1+pYKi3/ZI43gBlEUM7sMDrC5HhOvkBy8RtLGhSKpNMbOTANErGZzbHnyn3yGgOgk2AiTyv4IoA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784819835; c=relaxed/simple; bh=y0jFTQOJ+Dq2wgdO+3QOYkjdotoBMq7hRls4SEbmZAo=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=nxz7RmvOTClSy31+OwF81VoN0hhjH48RJROYs+sFWjJIuuGKUufbkS7VzoIGkOh58UjIVvWWbHi1NCrHM150m+LASefhEUBcVbGyDt0YsFpoqiB1YRDWKuTO/LIC51DJ1DHjwDjyvY67XChsfKT3pHYAX5mWbXKwv0VKIRkr8jk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=SQWOUjeW; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="SQWOUjeW" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 413C71F00A3A; Thu, 23 Jul 2026 15:17:05 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784819831; bh=VUDPRdQphPMrECclGKpOqw0nMGRICTbCUk3WOCCAA1A=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=SQWOUjeW2dB3kjlC4odUADa9Jp2xJqiFX/7ZcgihbP0S64TwuIV2ks6zIvSWFIvvF xxWvfA7I3/PG5inDeTcHwTDlr544Hm1wv2BR1pIiiYuPWWIVEsqSozelC4GqGacO9e yKiakA8eLI22ykmTv99SLVJho+mtp4JueZhUd16wxuYhrkYyU5SVQPHcoQznNe9v6U hrreponpC2Elqr0MF5CfVPAHgv8m5DnGG72PqHDND+yjA1nZpPAz/3OEwAptnZZNVW VqrVj3BAopHlyQ242U24ibg6z6JBRmMpiVCELJPJ2mXRESt8wgun67jUQ5sDtxxmE9 hPvZvyFUhbZDQ== From: "Lorenzo Stoakes (ARM)" Date: Thu, 23 Jul 2026 16:16:31 +0100 Subject: [PATCH mm-hotfixes v6 1/5] mm/vmalloc: acquire init_mm lock on huge vmap to avoid ptdump UAF Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260723-series-vmap-race-fix-v6-1-8cc77dcc0018@kernel.org> References: <20260723-series-vmap-race-fix-v6-0-8cc77dcc0018@kernel.org> In-Reply-To: <20260723-series-vmap-race-fix-v6-0-8cc77dcc0018@kernel.org> To: Andrew Morton , Suren Baghdasaryan , "Liam R. Howlett" , Vlastimil Babka , Shakeel Butt , David Hildenbrand , Mike Rapoport , Michal Hocko , Uladzislau Rezki , Toshi Kani , Dave Hansen , Andy Lutomirski , Peter Zijlstra , Thomas Gleixner , Ingo Molnar , Borislav Petkov , x86@kernel.org, "H. Peter Anvin" , Kiryl Shutsemau , Catalin Marinas , Will Deacon , Dev Jain , Ryan Roberts Cc: David Carlier , ljs@kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, "Denis V. Lunev" , "Borah, Chaitanya Kumar" , stable@vger.kernel.org, syzbot+fd95a72470f5a44e464c@syzkaller.appspotmail.com X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=8900; i=ljs@kernel.org; h=from:subject:message-id; bh=y0jFTQOJ+Dq2wgdO+3QOYkjdotoBMq7hRls4SEbmZAo=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLKSDGItizZ7PHgwo3lue/LTT5IeaudLdv1gqsg49sz98 fonJkXfOkpZGMS4GGTFFFmefxHfHyQSNq/zgr8bzBxWJpAhDFycAjCRu46MDFdNRW6+X9lYLxpT /yeWiUNNKezc08qJL9z74tg1/+5nWMnw3//inBM9L7/uyhQ9Ztr2pJF7w5/s+1v2ee+c7ZmQJbN nOS8A X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 Currently there is a nasty race between ptdump and vmap when attempting to map a huge P4D, PUD or PMD entry: * ptdump walks kernel page table ranges it doesn't own. * When vmap maps ranges it tries to promotes existing ones to huge page tables in vmap_try_huge_[p4d,pud,pmd]() at P4D, PUD and PMD level, freeing the lower page table in [p4d,pud,pmd]_free_[pud,pmd,pte]_page() when it succeeds. Both of these things can happen at the same time and as a result ptdump can access a freed page table, resulting in a use-after-free and memory corruption. This is possible because while ptdump_walk_pgd() holds both the mem hotplug lock and the mmap write lock before invoking walk_page_range_debug(), vmap takes no relevant locks at all. Fix this by holding the mmap read lock in vmap_try_huge_*() when freeing page tables. The read lock is sufficient: ptdump is the only walker that must be excluded and it holds the mmap write lock. Other holders of the read lock may run concurrently, but each exclusively owns the range it operates on and cannot reach the page tables freed here. We also hold the lock while assigning the huge page table entry, which means page table walkers observe only the huge or non-huge page table entry. We use a trylock to prevent ptdump from blocking vmap making forward progress. This is fine because it's an optimisation in any case, and thus the vmap can safely proceed regardless. All other kernel page table walkers that touch vmalloc ranges either exclusively own the memory walked or acquire the mmap lock, so this correctly excludes those walkers. One wrinkle here is commit fa93b45fd397 ("arm64: Enable vmalloc-huge with ptdump"), which addresses the issue for arm64 only by explicitly acquiring the mmap read lock on kernel page table freeing should a concurrent ptdump be in progress. This is problematic as vmap may acquire the mmap read lock prior to ptdump attempting to acquire an mmap write lock, leading to a deadlock when the mmap read lock is slept upon on page table freeing due to rwsem anti-starvation. We work around this by predicating the mmap lock being taken on !CONFIG_ARM64 for the time being. With this patch applied, a follow up will partially revert commit fa93b45fd397 ("arm64: Enable vmalloc-huge with ptdump") and at that stage remove the arm64 ifdeffery. We also update walk_page_range_debug() to assert the mmap write lock unconditionally and update the comment here to reflect this change. The issue has existed as long as ptdump was available and vmap freed page tables when promoting to a huge leaf entry, that is, since commit b6bdb7517c3d ("mm/vmalloc: add interfaces to free unmapped page table") for huge ioremap, and commit 121e6f3258fe ("mm/vmalloc: hugepage vmalloc mappings") for huge vmalloc. Since the former is the earlier of the two we choose that for our Fixes tag. We also define a guard class for mmap_read_trylock() so we can use cleanup.h to make the scope handling cleaner in the implementation. This patch is based on work by David Carlier (linked), with gratitude! Fixes: b6bdb7517c3d ("mm/vmalloc: add interfaces to free unmapped page tabl= e") Cc: stable@vger.kernel.org Reported-by: syzbot+fd95a72470f5a44e464c@syzkaller.appspotmail.com Closes: https://lore.kernel.org/all/6a287988.39669fcc.33b062.00a0.GAE@googl= e.com/T/ Link: https://lore.kernel.org/linux-mm/20260706203128.162335-1-devnexen@gma= il.com/ Reviewed-by: Mike Rapoport (Microsoft) Reviewed-by: Dev Jain Acked-by: David Hildenbrand (Arm) Reviewed-by: Kiryl Shutsemau Signed-off-by: Lorenzo Stoakes (ARM) --- include/linux/mmap_lock.h | 1 + mm/pagewalk.c | 22 +++++++++++---------- mm/vmalloc.c | 49 ++++++++++++++++++++++++++++++++++++++-----= ---- 3 files changed, 53 insertions(+), 19 deletions(-) diff --git a/include/linux/mmap_lock.h b/include/linux/mmap_lock.h index 04b8f61ece5d..6b5c2390cc30 100644 --- a/include/linux/mmap_lock.h +++ b/include/linux/mmap_lock.h @@ -621,6 +621,7 @@ static inline void mmap_read_unlock(struct mm_struct *m= m) =20 DEFINE_GUARD(mmap_read_lock, struct mm_struct *, mmap_read_lock(_T), mmap_read_unlock(_T)) +DEFINE_GUARD_COND(mmap_read_lock, _try, mmap_read_trylock(_T)) =20 static inline void mmap_read_unlock_non_owner(struct mm_struct *mm) { diff --git a/mm/pagewalk.c b/mm/pagewalk.c index 3ae2586ff45b..bbcfd68d0907 100644 --- a/mm/pagewalk.c +++ b/mm/pagewalk.c @@ -678,6 +678,8 @@ int walk_kernel_page_table_range_lockless(unsigned long= start, unsigned long end * will also not lock the PTEs for the pte_entry() callback. * * This is for debugging purposes ONLY. + * + * The mmap write lock must be held. */ int walk_page_range_debug(struct mm_struct *mm, unsigned long start, unsigned long end, const struct mm_walk_ops *ops, @@ -691,6 +693,16 @@ int walk_page_range_debug(struct mm_struct *mm, unsign= ed long start, .no_vma =3D true }; =20 + /* + * When walking userland page tables, an mmap write lock must be held to + * account for munmap() downgrading to an mmap read lock when tearing + * down page tables. + * + * When walking kernel page tables, an mmap write lock must also be held + * to account for page table freeing on vmap huge page mapping. + */ + mmap_assert_write_locked(mm); + /* For convenience, we allow traversal of kernel mappings. */ if (mm =3D=3D &init_mm) return walk_kernel_page_table_range(start, end, ops, @@ -700,16 +712,6 @@ int walk_page_range_debug(struct mm_struct *mm, unsign= ed long start, if (!check_ops_safe(ops)) return -EINVAL; =20 - /* - * The mmap lock protects the page walker from changes to the page - * tables during the walk. However a read lock is insufficient to - * protect those areas which don't have a VMA as munmap() detaches - * the VMAs before downgrading to a read lock and actually tearing - * down PTEs/page tables. In which case, the mmap write lock should - * be held. - */ - mmap_assert_write_locked(mm); - return walk_pgd_range(start, end, &walk); } =20 diff --git a/mm/vmalloc.c b/mm/vmalloc.c index 1afca3568b9b..d5c4d2bb770b 100644 --- a/mm/vmalloc.c +++ b/mm/vmalloc.c @@ -43,6 +43,7 @@ #include #include #include +#include =20 #define CREATE_TRACE_POINTS #include @@ -158,10 +159,24 @@ static int vmap_try_huge_pmd(pmd_t *pmd, unsigned lon= g addr, unsigned long end, if (!IS_ALIGNED(phys_addr, PMD_SIZE)) return 0; =20 - if (pmd_present(*pmd) && !pmd_free_pte_page(pmd, addr)) - return 0; + if (!pmd_present(*pmd)) + return pmd_set_huge(pmd, phys_addr, prot); =20 - return pmd_set_huge(pmd, phys_addr, prot); + /* + * Acquire the mmap read lock to exclude ptdump, which walks + * kernel page tables it does not own under the mmap write lock. + * + * Concurrent read lock holders are safe: each exclusively owns + * the range it operates on and cannot reach this page table. + */ +#ifndef CONFIG_ARM64 + scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) +#endif + { + if (!pmd_free_pte_page(pmd, addr)) + return 0; + return pmd_set_huge(pmd, phys_addr, prot); + } } =20 static int vmap_pmd_range(pud_t *pud, unsigned long addr, unsigned long en= d, @@ -210,10 +225,18 @@ static int vmap_try_huge_pud(pud_t *pud, unsigned lon= g addr, unsigned long end, if (!IS_ALIGNED(phys_addr, PUD_SIZE)) return 0; =20 - if (pud_present(*pud) && !pud_free_pmd_page(pud, addr)) - return 0; + if (!pud_present(*pud)) + return pud_set_huge(pud, phys_addr, prot); =20 - return pud_set_huge(pud, phys_addr, prot); + /* See comment in vmap_try_huge_pmd(). */ +#ifndef CONFIG_ARM64 + scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) +#endif + { + if (!pud_free_pmd_page(pud, addr)) + return 0; + return pud_set_huge(pud, phys_addr, prot); + } } =20 static int vmap_pud_range(p4d_t *p4d, unsigned long addr, unsigned long en= d, @@ -262,10 +285,18 @@ static int vmap_try_huge_p4d(p4d_t *p4d, unsigned lon= g addr, unsigned long end, if (!IS_ALIGNED(phys_addr, P4D_SIZE)) return 0; =20 - if (p4d_present(*p4d) && !p4d_free_pud_page(p4d, addr)) - return 0; + if (!p4d_present(*p4d)) + return p4d_set_huge(p4d, phys_addr, prot); =20 - return p4d_set_huge(p4d, phys_addr, prot); + /* See comment in vmap_try_huge_pmd(). */ +#ifndef CONFIG_ARM64 + scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) +#endif + { + if (!p4d_free_pud_page(p4d, addr)) + return 0; + return p4d_set_huge(p4d, phys_addr, prot); + } } =20 static int vmap_p4d_range(pgd_t *pgd, unsigned long addr, unsigned long en= d, --=20 2.55.0 From nobody Fri Jul 24 05:21:56 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B54E84570C3; Thu, 23 Jul 2026 15:17:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784819845; cv=none; b=XV/zFo3e3qdDgMMYam3Mw4vWLoFCIJhkAdTcdBrGc5Ri3obpzZDJi4uSVCTcSUZ6pjHYARzlCq3ODZANf4BlI+5Ao90jv0InGqBvPr0fIBVLpbIen4Ev5gsZq+V4MOM2KUck+zhyZxtdm1JM0pOmxzCL1Eub3BIXNXw+pjPVbMk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784819845; c=relaxed/simple; bh=3isjtK28wb1BiLlMBIpy71VvY1jYYrdz7YUUYuIwCk8=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=S1mjxFrf2boq70rFjiCtNis4vXvYGGaImiIUXmhEoP3J7cwaxqRtJUbDTDaK4W11m+ujGEMUrkqVdUFZw6i1aBiPhohmHJIwvtnQky6tUS2uxb+UfNom68dVxdzMVe6bk6J14PhbaP93imn7jUaI3IBle702k0dalDBAoio5VNA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=OEW0t8w/; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="OEW0t8w/" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 448C61F000E9; Thu, 23 Jul 2026 15:17:12 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784819838; bh=jGfvVz5phYfszb4vs/wrQwKbo5YfQ9htQgjlD4h0uzo=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=OEW0t8w/ZdlXeERCjc75StArDzDUqbCfBmj/Evab41GuYCBr4ySbLeSLVRFCMGDv+ l/ctMnIh60mbWkTdAktAFmL9qN05WJa8i0aOq5ilbl+d8NsWSz3RRjVbBSouVtRrVM iIkARpos/gy9E6Y6xap1sb5R8kGxxStmgv2YxQD92ZPMdQzDm71GyCKg7SOaoHWpSk msBqA3iq8JznuaUxh2t5IkX+7Si6Kyh0nWI4L6NqqzbExl7eiMUQCzwJtzlqMH+b9D MSEzlmyzhr+pkSk+W/MEA/VPlHCL+klN7I9sn8F3dEM8FlMuu2UxEFYJRfR5FBnSQn jDYMIua2YFiTA== From: "Lorenzo Stoakes (ARM)" Date: Thu, 23 Jul 2026 16:16:32 +0100 Subject: [PATCH mm-hotfixes v6 2/5] x86/mm/pat: acquire init_mm write lock on collapse to avoid UAF Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260723-series-vmap-race-fix-v6-2-8cc77dcc0018@kernel.org> References: <20260723-series-vmap-race-fix-v6-0-8cc77dcc0018@kernel.org> In-Reply-To: <20260723-series-vmap-race-fix-v6-0-8cc77dcc0018@kernel.org> To: Andrew Morton , Suren Baghdasaryan , "Liam R. Howlett" , Vlastimil Babka , Shakeel Butt , David Hildenbrand , Mike Rapoport , Michal Hocko , Uladzislau Rezki , Toshi Kani , Dave Hansen , Andy Lutomirski , Peter Zijlstra , Thomas Gleixner , Ingo Molnar , Borislav Petkov , x86@kernel.org, "H. Peter Anvin" , Kiryl Shutsemau , Catalin Marinas , Will Deacon , Dev Jain , Ryan Roberts Cc: David Carlier , ljs@kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, "Denis V. Lunev" , "Borah, Chaitanya Kumar" , stable@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=4047; i=ljs@kernel.org; h=from:subject:message-id; bh=3isjtK28wb1BiLlMBIpy71VvY1jYYrdz7YUUYuIwCk8=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLKSDGJ/vL0Rc3P9G8usFe9mixwynvfEO80n4NQqqbCXw cfUbMpaO0pZGMS4GGTFFFmefxHfHyQSNq/zgr8bzBxWJpAhDFycAjCRG7sY/ufu5Fgs37pTbkL8 3TldD9vP771waZK/w1Pec69qPkXfyX3LyLB0xzdr+6U3D30QfFtaOetH0DOr+mDn5iN5szhtf4p uecAAAA== X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 x86 implements page attribute modification using its Change Page Attributes (CPA) mechanism. This tracks properties of ranges such as cache mode through x86 page attributes, and as part of that logic manipulates kernel page tables. Since commit 41d88484c71c ("x86/mm/pat: restore large ROX pages after fragmentation") ranges of kernel page table entries can be collapsed into huge page table entries as part of this logic. As part of this collapse, it frees the page tables which the collapsed entries previously pointed to, and it does so without any relevant locks being held to preclude concurrent kernel page table walkers. The only way this code can be reached is if CPA_COLLAPSE is specified, and this is only set in set_memory_rox() via: set_memory_rox() -> change_page_attr_set_clr() -> cpa_flush() -> cpa_collapse_large_pages() Notable users of this are execmem and bpf when manipulating executable mappings. However, this is problematic for ptdump as it walks ranges it does not own and thus runs the risk of a use-after-free on page tables freed underneath it. In addition, concurrent CPA collapse operations are possible which can also cause races. Resolve the issue by acquiring the mmap write lock on init_mm across the whole operation. It is safe to acquire a sleeping lock as all the callers invoke set_memory_rox() from process context and in any case, change_page_attr_set_clr() calls vm_unmap_alias() which ultimately takes a mutex, disallowing atomic context here. Fixes: 41d88484c71c ("x86/mm/pat: restore large ROX pages after fragmentati= on") Cc: stable@vger.kernel.org Reviewed-by: Mike Rapoport (Microsoft) Reviewed-by: Kiryl Shutsemau (Meta) Reviewed-by: David Hildenbrand (Arm) Reviewed-by: Dave Hansen Reviewed-by: Will Deacon Reviewed-by: David Carlier Signed-off-by: Lorenzo Stoakes (ARM) --- arch/x86/mm/pat/set_memory.c | 15 ++++++++++++++- include/linux/mmap_lock.h | 2 ++ 2 files changed, 16 insertions(+), 1 deletion(-) diff --git a/arch/x86/mm/pat/set_memory.c b/arch/x86/mm/pat/set_memory.c index d023a40a1e03..d1e63f7d267f 100644 --- a/arch/x86/mm/pat/set_memory.c +++ b/arch/x86/mm/pat/set_memory.c @@ -22,6 +22,7 @@ #include #include #include +#include =20 #include #include @@ -410,7 +411,7 @@ static void __cpa_flush_tlb(void *data) =20 static int collapse_large_pages(unsigned long addr, struct list_head *pgta= bles); =20 -static void cpa_collapse_large_pages(struct cpa_data *cpa) +static void __cpa_collapse_large_pages(struct cpa_data *cpa) { unsigned long start, addr, end; struct ptdesc *ptdesc, *tmp; @@ -442,6 +443,18 @@ static void cpa_collapse_large_pages(struct cpa_data *= cpa) } } =20 +static void cpa_collapse_large_pages(struct cpa_data *cpa) +{ + /* + * Take the mmap write lock on init_mm to: + * - Avoid a use-after-free if raced by ptdump (which takes its own + * write lock on init_mm). + * - Serialise concurrent CPA walkers. + */ + scoped_guard(mmap_write_lock, &init_mm) + __cpa_collapse_large_pages(cpa); +} + static void cpa_flush(struct cpa_data *cpa, int cache) { unsigned int i; diff --git a/include/linux/mmap_lock.h b/include/linux/mmap_lock.h index 6b5c2390cc30..047f5f5e2c34 100644 --- a/include/linux/mmap_lock.h +++ b/include/linux/mmap_lock.h @@ -621,6 +621,8 @@ static inline void mmap_read_unlock(struct mm_struct *m= m) =20 DEFINE_GUARD(mmap_read_lock, struct mm_struct *, mmap_read_lock(_T), mmap_read_unlock(_T)) +DEFINE_GUARD(mmap_write_lock, struct mm_struct *, + mmap_write_lock(_T), mmap_write_unlock(_T)) DEFINE_GUARD_COND(mmap_read_lock, _try, mmap_read_trylock(_T)) =20 static inline void mmap_read_unlock_non_owner(struct mm_struct *mm) --=20 2.55.0 From nobody Fri Jul 24 05:21:56 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DE04A4756B3; Thu, 23 Jul 2026 15:17:27 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784819851; cv=none; b=pQz1/rDvJKCRY14knz+hzO1GlF+d3d0SiMF/v28VhGtDJlE4Mv1dwZxZmMQfJRdPpVulVnQ87/K5zSXG+ZmEer9Xl1GLSmriVhRL5w1NkocRZq4fmfrR/y3Fak7OD+EimTJzqBkVse2Zti9mP00dhfE/KoTA4YLOA1MQ7Ua1OeM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784819851; c=relaxed/simple; bh=akxrj13lq/jN4z92xXrWGgQjWx+k4pmtig1tmCI1Rgs=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=Gjio6uvtkRdziyWLyZoCKrS4wkJbwNBFiFWs2l29m89R0MssV3BIESGq4oYWGdENe2KWKV/wPTyX3IfB/ga4SPicyENWn9+Fc952CMdAMNdHopcEueLLNtOOkRvn4e9YHJHRjKR8u1xzDOaHT3aYNtQUN7PYPQo9PjjKliofaCY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=NotWd1pq; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="NotWd1pq" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 16B531F00A3A; Thu, 23 Jul 2026 15:17:18 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784819845; bh=Cp2mZdK6fJ1/si4CysQ56zitoOpwB+w35kScbOb2foM=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=NotWd1pqW5cSI0cTR7v9yu46zVCF47GX5WwS13EO559ek9XdS1VKv+VhM7bYjlSjk wfnuNCgkjM3KADUhnUfjJfLN5ae+8XXCPt/u4afIW7HDlXbUJUOxjEYlQhu/hscZz8 x6v3t7J0Vpu2pc/e1Nmwx1p91nv2Ag7knaRRoq976gNfbXqe5eCupHWjc8TYkYPty2 lCBrXMHaDy0HoqayKPYfKfgY5QBb6qCpX2nBPGuyCCKN90/Gpt/blL2mfVIqg1zUEx kDLCmK6XBlKYIfgtLBeKg42eXyRgwoKoscq5AxIdb8J2oYzOB5/geYfhDBpmtTYuPQ 2U0mqevzggA2Q== From: "Lorenzo Stoakes (ARM)" Date: Thu, 23 Jul 2026 16:16:33 +0100 Subject: [PATCH mm-hotfixes v6 3/5] x86/mm/pat: acquire init_mm read lock on attribute change to avoid UAF Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260723-series-vmap-race-fix-v6-3-8cc77dcc0018@kernel.org> References: <20260723-series-vmap-race-fix-v6-0-8cc77dcc0018@kernel.org> In-Reply-To: <20260723-series-vmap-race-fix-v6-0-8cc77dcc0018@kernel.org> To: Andrew Morton , Suren Baghdasaryan , "Liam R. Howlett" , Vlastimil Babka , Shakeel Butt , David Hildenbrand , Mike Rapoport , Michal Hocko , Uladzislau Rezki , Toshi Kani , Dave Hansen , Andy Lutomirski , Peter Zijlstra , Thomas Gleixner , Ingo Molnar , Borislav Petkov , x86@kernel.org, "H. Peter Anvin" , Kiryl Shutsemau , Catalin Marinas , Will Deacon , Dev Jain , Ryan Roberts Cc: David Carlier , ljs@kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, "Denis V. Lunev" , "Borah, Chaitanya Kumar" , stable@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=4639; i=ljs@kernel.org; h=from:subject:message-id; bh=akxrj13lq/jN4z92xXrWGgQjWx+k4pmtig1tmCI1Rgs=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLKSDGJZv/xpV/pzpMI449Xn74qeNttef6kPX1t9YNmRP xOnsp3O6ShlYRDjYpAVU2R5/kV8f5BI2LzOC/5uMHNYmUCGMHBxCsBEOGcyMqyt+rVf2Dh9loKt YdgqI4aJmooduV+fbbvNsUvG5kDfyqWMDNPaH/qzLrXY/599euTFu+mnXu/jTlStY9GPubdItt5 jHi8A X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 A previous commit protected us against races between ptdump and CPA collapse, however one still exists between attribute changes and collapse as reported by Denis V. Lunev (linked). When an attribute change arises, a lockless page table walker obtains a PTE entry, which is later written to via set_pte_atomic(): ... -> change_page_attr_set_clr() -> __change_page_attr_set_clr() -> __change_page_attr() -> _lookup_address_cpa() -> lookup_address_in_pgd_attr() -> [ lockless page table walker ] -> set_pte_atomic() There is nothing preventing a concurrent CPA collapse which can free the PTE that was retrieved here, resulting in a use-after-free. With the mmap write lock taken on init_mm over CPA collapse, we can now resolve this race by acquiring an mmap read lock on init_mm over __change_page_attr_set_clr(). This locks across the whole operation over which the walk and the PTE entry write occurs, solving the race. It is safe to do this here, as no spinlocks are held upon entry to __change_page_attr_set_clr(). However, the lock must not be held over an allocation, as allocation can trigger reclaim and shrinkers may call into CPA recursively, making deadlocks possible (init_mm -> ... -> fs_reclaim -> init_mm). A page table is allocated when a huge page needs to be split: -> change_page_attr_set_clr() -> __change_page_attr_set_clr() -> __change_page_attr() -> split_large_page() [ pagetable_alloc() ] -> __split_large_page() Avoid deadlocks by dropping the mmap lock across pagetable_alloc() in split_large_page() and track whether this is needed by adding a new 'init_mm_read_locked' flag to struct cpa_data. This is safe as __split_large_page() (called with locks re-established) revalidates that the page table entry is the same as it was prior to the locks being dropped and __change_page_attr() repeats the entire page table walk whenever a split occurs, so concurrent split and collapse are accounted for. Concurrent ptdump is also safe as the lock is only dropped over page table allocation during which time the page table has not yet been modified. The CPA_COLLAPSE flag is only set by set_memory_rox(), which exclusively operates upon vmalloc ranges, and on x86 only within the module mapping space. This is important, because some callers directly invoke __change_page_attr_set_clr(), bypassing this lock. However, none of these operate within the module mapping space. * cpa_process_alias() - a recursive helper called by __change_page_attr_set_clr(). * __set_memory_enc_pgtable() - operates on the direct mapping and (via __vmbus_establish_gpadl()) the vmalloc mapping space. * __set_pages_[n]p() - called by set_direct_map_[invalid, default, valid]_noflush(), __kernel_map_pages() - operates on the direct map. * kernel_[un]map_pages_in_pgd() - operates on EFI ranges. This work is based upon Denis V. Lunev's excellent analysis of the bug with gratitude. Link: https://lore.kernel.org/all/20260626163213.2284080-1-den@openvz.org/ Fixes: 41d88484c71c ("x86/mm/pat: restore large ROX pages after fragmentati= on") Cc: stable@vger.kernel.org Signed-off-by: Lorenzo Stoakes (ARM) --- arch/x86/mm/pat/set_memory.c | 13 +++++++++++-- 1 file changed, 11 insertions(+), 2 deletions(-) diff --git a/arch/x86/mm/pat/set_memory.c b/arch/x86/mm/pat/set_memory.c index d1e63f7d267f..618f72096480 100644 --- a/arch/x86/mm/pat/set_memory.c +++ b/arch/x86/mm/pat/set_memory.c @@ -50,7 +50,8 @@ struct cpa_data { unsigned int flags; unsigned int force_split : 1, force_static_prot : 1, - force_flush_all : 1; + force_flush_all : 1, + init_mm_read_locked : 1; struct page **pages; }; =20 @@ -1250,7 +1251,11 @@ static int split_large_page(struct cpa_data *cpa, pt= e_t *kpte, =20 if (!debug_pagealloc_enabled()) spin_unlock(&cpa_lock); + if (cpa->init_mm_read_locked) + mmap_read_unlock(&init_mm); ptdesc =3D pagetable_alloc(GFP_KERNEL, 0); + if (cpa->init_mm_read_locked) + mmap_read_lock(&init_mm); if (!debug_pagealloc_enabled()) spin_lock(&cpa_lock); if (!ptdesc) @@ -2122,7 +2127,11 @@ static int change_page_attr_set_clr(unsigned long *a= ddr, int numpages, cpa.curpage =3D 0; cpa.force_split =3D force_split; =20 - ret =3D __change_page_attr_set_clr(&cpa, 1); + /* Avoid race with concurrent CPA collapse. */ + cpa.init_mm_read_locked =3D true; + scoped_guard(mmap_read_lock, &init_mm) + ret =3D __change_page_attr_set_clr(&cpa, 1); + cpa.init_mm_read_locked =3D false; =20 /* * Check whether we really changed something: --=20 2.55.0 From nobody Fri Jul 24 05:21:56 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D82EF4570FD; Thu, 23 Jul 2026 15:17:34 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784819858; cv=none; b=LXbYhXNwaUd6WYrZb9Av5DFL3wGhWU97xIDi0WMUj9CouytYdHJBDL57Iw8YCD7/Zyo4kz15UWveujjuROHdbQPyIFpO9OewfizsqS5JcasetDsvt4LtufBXzEw218QFMhlZOBlmYd0VJy/plo8NjASfQ47902bpdDMUlh4wZZI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784819858; c=relaxed/simple; bh=VITXU838ytmI5wjkeVjJjzJ76fi6Ga2j8Ljay/OkRtw=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=Tx+7EsFE5Mjim4Evm/GYQRavaWJ/CX4laWFjICRvLEdADzcV/to7bzkFRhVaxFyd6OtjfgoPe5yPTi+0dj5uOqXg1vgtzUfpKHX62QCMXhZAQJs+LvWo9Vy59wEKCCG5TYARDh35bDsqi3tAWQznZA2gl+8/F9fBMjAAGR5Ct/M= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=keE8qT7v; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="keE8qT7v" Received: by smtp.kernel.org (Postfix) with ESMTPSA id DDCFB1F000E9; Thu, 23 Jul 2026 15:17:25 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784819852; bh=4TmiXyJvQ92TKxUnkF2EBAbSx4BVAEx8E3qvkE5M06w=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=keE8qT7vBTXkbMs26asTsL8pW67WGqXgwzV1Ac5f45Fk6q+/p0+JWXxKODKzKG2ED DZgmEz0u7tnsHrqOJVLTUrY6Qc9if4n1X3ZkO2XIPXZyfpazO1aHpvfk4le7g7Kb9U J3ZwgKTuEoAtiz+bAz/60D8uKMNeX+k1NhBe5UfZimuno6656dW8Qyjs7P+2lJh4qj gPCAEud3BZEadvF4O9II4RK6a8GfKCB4CRik21ABhe2MO0v5DtfVsiKGx0Jpqc+zox xUMBc7WDk1jje3vkntaNN9cecPbAVoijItvVKYQn0OQjnUDlxqTfyiUahPWon6/hGP u9MfyDp+58GEQ== From: "Lorenzo Stoakes (ARM)" Date: Thu, 23 Jul 2026 16:16:34 +0100 Subject: [PATCH mm-hotfixes v6 4/5] mm/ptdump: always stabilise against page table freeing using init_mm Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260723-series-vmap-race-fix-v6-4-8cc77dcc0018@kernel.org> References: <20260723-series-vmap-race-fix-v6-0-8cc77dcc0018@kernel.org> In-Reply-To: <20260723-series-vmap-race-fix-v6-0-8cc77dcc0018@kernel.org> To: Andrew Morton , Suren Baghdasaryan , "Liam R. Howlett" , Vlastimil Babka , Shakeel Butt , David Hildenbrand , Mike Rapoport , Michal Hocko , Uladzislau Rezki , Toshi Kani , Dave Hansen , Andy Lutomirski , Peter Zijlstra , Thomas Gleixner , Ingo Molnar , Borislav Petkov , x86@kernel.org, "H. Peter Anvin" , Kiryl Shutsemau , Catalin Marinas , Will Deacon , Dev Jain , Ryan Roberts Cc: David Carlier , ljs@kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, "Denis V. Lunev" , "Borah, Chaitanya Kumar" , stable@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=3325; i=ljs@kernel.org; h=from:subject:message-id; bh=VITXU838ytmI5wjkeVjJjzJ76fi6Ga2j8Ljay/OkRtw=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLKSDGJXXL3g06vGuumGcyFrSqK5zkz7B84HBAO1XojWL UqxWvS8o5SFQYyLQVZMkeX5F/H9QSJh8zov+LvBzGFlAhnCwMUpABPx28/I8LtPcvGmhtd7xKYe yuPN0LVhOPaq5/kH15cfy7JcGZiLwhj+cHu+7FFcyTpljrjO661+H5q0Q74szbw//Uuwtii/t4c iCwA= X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 Previous commits have established the invariant that kernel page table freeing is performed while an mmap read lock on init_mm is held, which fixes races between ptdump and kernel page table freeing over init_mm. However, x86 and arm64 can perform a ptdump over an mm other than init_mm via ptdump_walk_pgd() and since kernel memory ranges are shared across non-kernel mm's, this means that the race still exists for these cases. Fix this by acquiring a nested mmap write lock for init_mm in ptdump_walk_pgd(). This is safe as we take this after mmap write locking the mm, and nothing acquires the init_mm lock first before locking an arbitrary mm, so no deadlock is possible. Also update walk_page_range_debug() to assert that init_mm is write locked, add a comment explaining why and remove some redundant code, and eliminate the unnecessary and confusing invocation of walk_kernel_page_table_range(). We can safely remove the non-NULL check for walk.mm, as the mmap lock asserts would NULL pointer deref if it was (and of course no callers do this). The first point at which ptdump can race kernel page table freeing is commit b6bdb7517c3d ("mm/vmalloc: add interfaces to free unmapped page table"), so we target this in the Fixes tag. Fixes: b6bdb7517c3d ("mm/vmalloc: add interfaces to free unmapped page tabl= e") Cc: stable@vger.kernel.org Reviewed-by: Mike Rapoport (Microsoft) Acked-by: David Hildenbrand (Arm) Reviewed-by: Kiryl Shutsemau Signed-off-by: Lorenzo Stoakes (ARM) --- mm/pagewalk.c | 14 +++++++++----- mm/ptdump.c | 7 +++++++ 2 files changed, 16 insertions(+), 5 deletions(-) diff --git a/mm/pagewalk.c b/mm/pagewalk.c index bbcfd68d0907..5d87c632a255 100644 --- a/mm/pagewalk.c +++ b/mm/pagewalk.c @@ -702,12 +702,16 @@ int walk_page_range_debug(struct mm_struct *mm, unsig= ned long start, * to account for page table freeing on vmap huge page mapping. */ mmap_assert_write_locked(mm); + /* + * x86, arm64 ptdump allow walks of efi mm's and x86 ptdump allows walks + * of arbitrary mm's. + * + * However, they both must also hold the init_mm lock to account for + * concurrent kernel page table freeing. + */ + mmap_assert_write_locked(&init_mm); =20 - /* For convenience, we allow traversal of kernel mappings. */ - if (mm =3D=3D &init_mm) - return walk_kernel_page_table_range(start, end, ops, - pgd, private); - if (start >=3D end || !walk.mm) + if (start >=3D end) return -EINVAL; if (!check_ops_safe(ops)) return -EINVAL; diff --git a/mm/ptdump.c b/mm/ptdump.c index 973020000096..5851096e6f65 100644 --- a/mm/ptdump.c +++ b/mm/ptdump.c @@ -178,11 +178,18 @@ void ptdump_walk_pgd(struct ptdump_state *st, struct = mm_struct *mm, pgd_t *pgd) =20 get_online_mems(); mmap_write_lock(mm); + /* To stabilise kernel page tables we must hold the init_mm lock too. */ + if (mm !=3D &init_mm) + mmap_write_lock_nested(&init_mm, SINGLE_DEPTH_NESTING); + while (range->start !=3D range->end) { walk_page_range_debug(mm, range->start, range->end, &ptdump_ops, pgd, st); range++; } + + if (mm !=3D &init_mm) + mmap_write_unlock(&init_mm); mmap_write_unlock(mm); put_online_mems(); =20 --=20 2.55.0 From nobody Fri Jul 24 05:21:56 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 62B4E4570FC; Thu, 23 Jul 2026 15:17:39 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784819863; cv=none; b=TwBlL05Je/mg/1+EZjUoQRt0V5rwsFNiWuA+8WcYAmw82OsN05gCoKYuZbcEoHekxaJxNHSzB1AsAiC/Ui45LWD9VruGV1TFeolUwH42PQQnjb6o89fEhFphhiZjX+/WBWOaTU/OHKFJR+SHD+c0+IA9ilMl+IU/HE6p47RXuYc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784819863; c=relaxed/simple; bh=kE9XtB6iyw7BSFNAuuPgI01clqwgzMyGQ2yqQHsLjJM=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=Wn4I5brM3pi7CFliX3pCC3whguZAqMMp9+iZWeratF3aGRbxApyilMG4WVy8uH2cSrtlQxR7RpQInViyDbF7I1D95mitWa6O3dhrjaD5KFpR3BDydGGTFE4zQ/ERoYsvXY1/d/F1gvZhfsK8REZfoB5x6p0qJIBUyyC6Gx0628E= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=KRq8PkcA; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="KRq8PkcA" Received: by smtp.kernel.org (Postfix) with ESMTPSA id AFF771F00A3D; Thu, 23 Jul 2026 15:17:32 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784819859; bh=6tB+p5fpJihvWmKYtOqn6XBXbSKHjdM1g/8Wc7lnKAc=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=KRq8PkcAI15VvLDhvamyrojQCEWXOJom5dmcVSRO3GGI0ZVtmBWxnlnE4AfAvyW5R xSyNV65KtAqJ+AvR5Wuxbd8pExLFLRPPQeUuAcx6AV3s33G+KLwTWzaSddN0y/DdoF MWRWqVKt3EnHVokgFuwjSt8Aw/UEvPGcBZm2/gKqTMgr6P949HYeNEnGcMXF2TpnUh dY0jpQKJ62SmMj9IWNm7sEVu9ncYGmQm4sUlXSCzAkl+ZvG4gEsu7AZ96Rvg42W4Y3 NG691MHau9f3QMzSeaejz1V6rUC3jsXt1wPdRFfZqeZp/CaXPrITTT7J6iFG2I+0qG KMvqqHSJXD1wg== From: "Lorenzo Stoakes (ARM)" Date: Thu, 23 Jul 2026 16:16:35 +0100 Subject: [PATCH mm-hotfixes v6 5/5] arm64: remove redundant concurrent ptdump UAF mitigation Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260723-series-vmap-race-fix-v6-5-8cc77dcc0018@kernel.org> References: <20260723-series-vmap-race-fix-v6-0-8cc77dcc0018@kernel.org> In-Reply-To: <20260723-series-vmap-race-fix-v6-0-8cc77dcc0018@kernel.org> To: Andrew Morton , Suren Baghdasaryan , "Liam R. Howlett" , Vlastimil Babka , Shakeel Butt , David Hildenbrand , Mike Rapoport , Michal Hocko , Uladzislau Rezki , Toshi Kani , Dave Hansen , Andy Lutomirski , Peter Zijlstra , Thomas Gleixner , Ingo Molnar , Borislav Petkov , x86@kernel.org, "H. Peter Anvin" , Kiryl Shutsemau , Catalin Marinas , Will Deacon , Dev Jain , Ryan Roberts Cc: David Carlier , ljs@kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, "Denis V. Lunev" , "Borah, Chaitanya Kumar" , stable@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=7063; i=ljs@kernel.org; h=from:subject:message-id; bh=kE9XtB6iyw7BSFNAuuPgI01clqwgzMyGQ2yqQHsLjJM=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLKSDGJX1utw8GwSKzbiF8gXVHh1sqQieU/oivnb0i5vD 2qse5zeUcrCIMbFICumyPL8i/j+IJGweZ0X/N1g5rAygQxh4OIUgIkcYGH4n6NXKbo/bEO/hMzX CQcZJhQf9Fpxozz4pO7M2XPF9wnliTP8ldJbzXqqN+T1UY/L+/2PmS3rVPbftWy1jrJfkvNqgak /WAE= X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 This partially reverts commit fa93b45fd397 ("arm64: Enable vmalloc-huge with ptdump"), retaining vmalloc-huge support but eliminating the now redundant mitigation against a race between huge vmap page table freeing and ptdump, as this issue has now been fixed at core. We also simultaneously remove the arm64 if-deffery when acquiring the mmap read lock upon vmap huge page table promotion as it is no longer required. Note that this patch relies on the preceding vmalloc patch, and should not be backported alone. Fixes: fa93b45fd397 ("arm64: Enable vmalloc-huge with ptdump") Cc: stable@vger.kernel.org Reviewed-by: Dev Jain Acked-by: Mike Rapoport (Microsoft) Acked-by: Kiryl Shutsemau (Meta) Acked-by: Will Deacon Reviewed-by: David Hildenbrand (Arm) Signed-off-by: Lorenzo Stoakes (ARM) --- arch/arm64/include/asm/ptdump.h | 2 -- arch/arm64/mm/mmu.c | 43 ++++---------------------------------= ---- arch/arm64/mm/ptdump.c | 11 ++--------- mm/vmalloc.c | 15 +++----------- 4 files changed, 9 insertions(+), 62 deletions(-) diff --git a/arch/arm64/include/asm/ptdump.h b/arch/arm64/include/asm/ptdum= p.h index 5b374a6ab34a..50a195eda8ed 100644 --- a/arch/arm64/include/asm/ptdump.h +++ b/arch/arm64/include/asm/ptdump.h @@ -7,8 +7,6 @@ =20 #include =20 -DECLARE_STATIC_KEY_FALSE(arm64_ptdump_lock_key); - #ifdef CONFIG_PTDUMP =20 #include diff --git a/arch/arm64/mm/mmu.c b/arch/arm64/mm/mmu.c index a25d8beacc83..bd52fca6e872 100644 --- a/arch/arm64/mm/mmu.c +++ b/arch/arm64/mm/mmu.c @@ -49,8 +49,6 @@ #define NO_CONT_MAPPINGS BIT(1) #define NO_EXEC_MAPPINGS BIT(2) /* assumes FEAT_HPDS is not used */ =20 -DEFINE_STATIC_KEY_FALSE(arm64_ptdump_lock_key); - u64 kimage_voffset __ro_after_init; EXPORT_SYMBOL(kimage_voffset); =20 @@ -1864,8 +1862,7 @@ int pmd_clear_huge(pmd_t *pmdp) return 1; } =20 -static int __pmd_free_pte_page(pmd_t *pmdp, unsigned long addr, - bool acquire_mmap_lock) +int pmd_free_pte_page(pmd_t *pmdp, unsigned long addr) { pte_t *table; pmd_t pmd; @@ -1877,25 +1874,13 @@ static int __pmd_free_pte_page(pmd_t *pmdp, unsigne= d long addr, return 1; } =20 - /* See comment in pud_free_pmd_page for static key logic */ table =3D pte_offset_kernel(pmdp, addr); pmd_clear(pmdp); __flush_tlb_kernel_pgtable(addr); - if (static_branch_unlikely(&arm64_ptdump_lock_key) && acquire_mmap_lock) { - mmap_read_lock(&init_mm); - mmap_read_unlock(&init_mm); - } - pte_free_kernel(NULL, table); return 1; } =20 -int pmd_free_pte_page(pmd_t *pmdp, unsigned long addr) -{ - /* If ptdump is walking the pagetables, acquire init_mm.mmap_lock */ - return __pmd_free_pte_page(pmdp, addr, /* acquire_mmap_lock =3D */ true); -} - int pud_free_pmd_page(pud_t *pudp, unsigned long addr) { pmd_t *table; @@ -1911,36 +1896,16 @@ int pud_free_pmd_page(pud_t *pudp, unsigned long ad= dr) } =20 table =3D pmd_offset(pudp, addr); - - /* - * Our objective is to prevent ptdump from reading a PMD table which has - * been freed. In this race, if pud_free_pmd_page observes the key on - * (which got flipped by ptdump) then the mmap lock sequence here will, - * as a result of the mmap write lock/unlock sequence in ptdump, give - * us the correct synchronization. If not, this means that ptdump has - * yet not started walking the pagetables - the sequence of barriers - * issued by __flush_tlb_kernel_pgtable() guarantees that ptdump will - * observe an empty PUD. - */ - pud_clear(pudp); - __flush_tlb_kernel_pgtable(addr); - if (static_branch_unlikely(&arm64_ptdump_lock_key)) { - mmap_read_lock(&init_mm); - mmap_read_unlock(&init_mm); - } - pmdp =3D table; next =3D addr; end =3D addr + PUD_SIZE; do { if (pmd_present(pmdp_get(pmdp))) - /* - * PMD has been isolated, so ptdump won't see it. No - * need to acquire init_mm.mmap_lock. - */ - __pmd_free_pte_page(pmdp, next, /* acquire_mmap_lock =3D */ false); + pmd_free_pte_page(pmdp, next); } while (pmdp++, next +=3D PMD_SIZE, next !=3D end); =20 + pud_clear(pudp); + __flush_tlb_kernel_pgtable(addr); pmd_free(NULL, table); return 1; } diff --git a/arch/arm64/mm/ptdump.c b/arch/arm64/mm/ptdump.c index 1c20144700d7..5a76c59b5ada 100644 --- a/arch/arm64/mm/ptdump.c +++ b/arch/arm64/mm/ptdump.c @@ -283,13 +283,6 @@ void note_page_flush(struct ptdump_state *pt_st) note_page(pt_st, 0, -1, pte_val(pte_zero)); } =20 -static void arm64_ptdump_walk_pgd(struct ptdump_state *st, struct mm_struc= t *mm) -{ - static_branch_inc(&arm64_ptdump_lock_key); - ptdump_walk_pgd(st, mm, NULL); - static_branch_dec(&arm64_ptdump_lock_key); -} - void ptdump_walk(struct seq_file *s, struct ptdump_info *info) { unsigned long end =3D ~0UL; @@ -318,7 +311,7 @@ void ptdump_walk(struct seq_file *s, struct ptdump_info= *info) } }; =20 - arm64_ptdump_walk_pgd(&st.ptdump, info->mm); + ptdump_walk_pgd(&st.ptdump, info->mm, NULL); } =20 static void __init ptdump_initialize(void) @@ -360,7 +353,7 @@ bool ptdump_check_wx(void) } }; =20 - arm64_ptdump_walk_pgd(&st.ptdump, &init_mm); + ptdump_walk_pgd(&st.ptdump, &init_mm, NULL); =20 if (st.wx_pages || st.uxn_pages) { pr_warn("Checked W+X mappings: FAILED, %lu W+X pages found, %lu non-UXN = pages found\n", diff --git a/mm/vmalloc.c b/mm/vmalloc.c index d5c4d2bb770b..f4fa227a8d7f 100644 --- a/mm/vmalloc.c +++ b/mm/vmalloc.c @@ -169,10 +169,7 @@ static int vmap_try_huge_pmd(pmd_t *pmd, unsigned long= addr, unsigned long end, * Concurrent read lock holders are safe: each exclusively owns * the range it operates on and cannot reach this page table. */ -#ifndef CONFIG_ARM64 - scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) -#endif - { + scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) { if (!pmd_free_pte_page(pmd, addr)) return 0; return pmd_set_huge(pmd, phys_addr, prot); @@ -229,10 +226,7 @@ static int vmap_try_huge_pud(pud_t *pud, unsigned long= addr, unsigned long end, return pud_set_huge(pud, phys_addr, prot); =20 /* See comment in vmap_try_huge_pmd(). */ -#ifndef CONFIG_ARM64 - scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) -#endif - { + scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) { if (!pud_free_pmd_page(pud, addr)) return 0; return pud_set_huge(pud, phys_addr, prot); @@ -289,10 +283,7 @@ static int vmap_try_huge_p4d(p4d_t *p4d, unsigned long= addr, unsigned long end, return p4d_set_huge(p4d, phys_addr, prot); =20 /* See comment in vmap_try_huge_pmd(). */ -#ifndef CONFIG_ARM64 - scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) -#endif - { + scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) { if (!p4d_free_pud_page(p4d, addr)) return 0; return p4d_set_huge(p4d, phys_addr, prot); --=20 2.55.0