From nobody Sat Jul 25 06:09:49 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D5E6A3BFAEB; Thu, 16 Jul 2026 21:31:47 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784237509; cv=none; b=jU57CUOimH3T62Tko9cwtcKtFobtkUwhG1nNMPpZ0g2IEfZhIWtHkzWL9/jrnPBXORGdQguS5I4u/yXiEO94qN+GpSlITJIje6ihp8b1i1pyQ9OYggifJg0vAw+H0u2TO4qAycqO7H1nFqw13p9zNfELAa0k6dxcpdTRD4UWBIQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784237509; c=relaxed/simple; bh=y0jFTQOJ+Dq2wgdO+3QOYkjdotoBMq7hRls4SEbmZAo=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=h6ZNT8zuRT7kOAEXwlRbNXilVWRFI0VPCn1Vr8kmjNWJ3rcazPncuu5q2cMy8xs7X+cUSGPiUPKZYMk7WBdwLiFr9sr/qFvy62VnWfIMjAGp7n9dBajdG6uok/Aaz0ztBokwwvX55DIWpGKFiAN8fQ4vbQMncEmSnFqO3NCfXXM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=D58q0M8+; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="D58q0M8+" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 3019C1F00A3A; Thu, 16 Jul 2026 21:31:41 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784237507; bh=VUDPRdQphPMrECclGKpOqw0nMGRICTbCUk3WOCCAA1A=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=D58q0M8+Hpf8Dkt3GCqnMEE5JOKkP0nBGdiPf2iJVmdnjHAkCgo1+Z7Fs8Jrlo/9D bdx/vBgJX2NV2AoFWh6icjk/R4EXrTCDNjRp+KhS7BbbUfE7sqnKG4MSYIBwGyBCn8 jZV1+Jo43GZdCNQcX8dFJVAp6MHz2agji3U8ZfY9ITXYrfw54Itb9lSU6SwAYZMlbz ujFWDXuMLYH8R4aA4wO5kT6s64CXPZH80un5jn33Z7jslRc8DgZEqHUfCQrm6kXAi6 ZcZoh83wZ174UURkU7+xz/uLZ5rC6bVki9PyH+GHwuMYkBe+0mknZeXFh+bHG4FNS1 DDtMtBXs/jr1w== From: "Lorenzo Stoakes (ARM)" Date: Thu, 16 Jul 2026 22:31:12 +0100 Subject: [PATCH mm-hotfixes v4 1/4] mm/vmalloc: acquire init_mm lock on huge vmap to avoid ptdump UAF Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260716-series-vmap-race-fix-v4-1-8c108c4317df@kernel.org> References: <20260716-series-vmap-race-fix-v4-0-8c108c4317df@kernel.org> In-Reply-To: <20260716-series-vmap-race-fix-v4-0-8c108c4317df@kernel.org> To: Andrew Morton , Suren Baghdasaryan , "Liam R. Howlett" , Vlastimil Babka , Shakeel Butt , David Hildenbrand , Mike Rapoport , Michal Hocko , Uladzislau Rezki , Toshi Kani , Dave Hansen , Andy Lutomirski , Peter Zijlstra , Thomas Gleixner , Ingo Molnar , Borislav Petkov , x86@kernel.org, "H. Peter Anvin" , Kiryl Shutsemau , Catalin Marinas , Will Deacon , Dev Jain , Ryan Roberts Cc: David Carlier , ljs@kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, "Denis V. Lunev" , stable@vger.kernel.org, syzbot+fd95a72470f5a44e464c@syzkaller.appspotmail.com X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=8900; i=ljs@kernel.org; h=from:subject:message-id; bh=y0jFTQOJ+Dq2wgdO+3QOYkjdotoBMq7hRls4SEbmZAo=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLIifdfeP/e/0ik0SPHi9TBVx/edlp7Ghq+TzFZ2VKjzq 997KvW2o5SFQYyLQVZMkeX5F/H9QSJh8zov+LvBzGFlAhnCwMUpABORUGb47yguEWjeZ7UmPWPZ /uMdtz9P8A05OvNdypf7gv5tj9/K6jP8927O3+tWYr3txq/fB1o+OEyT3v1sZsiM+dm+Zi07VK6 0MQIA X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 Currently there is a nasty race between ptdump and vmap when attempting to map a huge P4D, PUD or PMD entry: * ptdump walks kernel page table ranges it doesn't own. * When vmap maps ranges it tries to promotes existing ones to huge page tables in vmap_try_huge_[p4d,pud,pmd]() at P4D, PUD and PMD level, freeing the lower page table in [p4d,pud,pmd]_free_[pud,pmd,pte]_page() when it succeeds. Both of these things can happen at the same time and as a result ptdump can access a freed page table, resulting in a use-after-free and memory corruption. This is possible because while ptdump_walk_pgd() holds both the mem hotplug lock and the mmap write lock before invoking walk_page_range_debug(), vmap takes no relevant locks at all. Fix this by holding the mmap read lock in vmap_try_huge_*() when freeing page tables. The read lock is sufficient: ptdump is the only walker that must be excluded and it holds the mmap write lock. Other holders of the read lock may run concurrently, but each exclusively owns the range it operates on and cannot reach the page tables freed here. We also hold the lock while assigning the huge page table entry, which means page table walkers observe only the huge or non-huge page table entry. We use a trylock to prevent ptdump from blocking vmap making forward progress. This is fine because it's an optimisation in any case, and thus the vmap can safely proceed regardless. All other kernel page table walkers that touch vmalloc ranges either exclusively own the memory walked or acquire the mmap lock, so this correctly excludes those walkers. One wrinkle here is commit fa93b45fd397 ("arm64: Enable vmalloc-huge with ptdump"), which addresses the issue for arm64 only by explicitly acquiring the mmap read lock on kernel page table freeing should a concurrent ptdump be in progress. This is problematic as vmap may acquire the mmap read lock prior to ptdump attempting to acquire an mmap write lock, leading to a deadlock when the mmap read lock is slept upon on page table freeing due to rwsem anti-starvation. We work around this by predicating the mmap lock being taken on !CONFIG_ARM64 for the time being. With this patch applied, a follow up will partially revert commit fa93b45fd397 ("arm64: Enable vmalloc-huge with ptdump") and at that stage remove the arm64 ifdeffery. We also update walk_page_range_debug() to assert the mmap write lock unconditionally and update the comment here to reflect this change. The issue has existed as long as ptdump was available and vmap freed page tables when promoting to a huge leaf entry, that is, since commit b6bdb7517c3d ("mm/vmalloc: add interfaces to free unmapped page table") for huge ioremap, and commit 121e6f3258fe ("mm/vmalloc: hugepage vmalloc mappings") for huge vmalloc. Since the former is the earlier of the two we choose that for our Fixes tag. We also define a guard class for mmap_read_trylock() so we can use cleanup.h to make the scope handling cleaner in the implementation. This patch is based on work by David Carlier (linked), with gratitude! Fixes: b6bdb7517c3d ("mm/vmalloc: add interfaces to free unmapped page tabl= e") Cc: stable@vger.kernel.org Reported-by: syzbot+fd95a72470f5a44e464c@syzkaller.appspotmail.com Closes: https://lore.kernel.org/all/6a287988.39669fcc.33b062.00a0.GAE@googl= e.com/T/ Link: https://lore.kernel.org/linux-mm/20260706203128.162335-1-devnexen@gma= il.com/ Reviewed-by: Mike Rapoport (Microsoft) Reviewed-by: Dev Jain Acked-by: David Hildenbrand (Arm) Reviewed-by: Kiryl Shutsemau Signed-off-by: Lorenzo Stoakes (ARM) --- include/linux/mmap_lock.h | 1 + mm/pagewalk.c | 22 +++++++++++---------- mm/vmalloc.c | 49 ++++++++++++++++++++++++++++++++++++++-----= ---- 3 files changed, 53 insertions(+), 19 deletions(-) diff --git a/include/linux/mmap_lock.h b/include/linux/mmap_lock.h index 04b8f61ece5d..6b5c2390cc30 100644 --- a/include/linux/mmap_lock.h +++ b/include/linux/mmap_lock.h @@ -621,6 +621,7 @@ static inline void mmap_read_unlock(struct mm_struct *m= m) =20 DEFINE_GUARD(mmap_read_lock, struct mm_struct *, mmap_read_lock(_T), mmap_read_unlock(_T)) +DEFINE_GUARD_COND(mmap_read_lock, _try, mmap_read_trylock(_T)) =20 static inline void mmap_read_unlock_non_owner(struct mm_struct *mm) { diff --git a/mm/pagewalk.c b/mm/pagewalk.c index 3ae2586ff45b..bbcfd68d0907 100644 --- a/mm/pagewalk.c +++ b/mm/pagewalk.c @@ -678,6 +678,8 @@ int walk_kernel_page_table_range_lockless(unsigned long= start, unsigned long end * will also not lock the PTEs for the pte_entry() callback. * * This is for debugging purposes ONLY. + * + * The mmap write lock must be held. */ int walk_page_range_debug(struct mm_struct *mm, unsigned long start, unsigned long end, const struct mm_walk_ops *ops, @@ -691,6 +693,16 @@ int walk_page_range_debug(struct mm_struct *mm, unsign= ed long start, .no_vma =3D true }; =20 + /* + * When walking userland page tables, an mmap write lock must be held to + * account for munmap() downgrading to an mmap read lock when tearing + * down page tables. + * + * When walking kernel page tables, an mmap write lock must also be held + * to account for page table freeing on vmap huge page mapping. + */ + mmap_assert_write_locked(mm); + /* For convenience, we allow traversal of kernel mappings. */ if (mm =3D=3D &init_mm) return walk_kernel_page_table_range(start, end, ops, @@ -700,16 +712,6 @@ int walk_page_range_debug(struct mm_struct *mm, unsign= ed long start, if (!check_ops_safe(ops)) return -EINVAL; =20 - /* - * The mmap lock protects the page walker from changes to the page - * tables during the walk. However a read lock is insufficient to - * protect those areas which don't have a VMA as munmap() detaches - * the VMAs before downgrading to a read lock and actually tearing - * down PTEs/page tables. In which case, the mmap write lock should - * be held. - */ - mmap_assert_write_locked(mm); - return walk_pgd_range(start, end, &walk); } =20 diff --git a/mm/vmalloc.c b/mm/vmalloc.c index 1afca3568b9b..d5c4d2bb770b 100644 --- a/mm/vmalloc.c +++ b/mm/vmalloc.c @@ -43,6 +43,7 @@ #include #include #include +#include =20 #define CREATE_TRACE_POINTS #include @@ -158,10 +159,24 @@ static int vmap_try_huge_pmd(pmd_t *pmd, unsigned lon= g addr, unsigned long end, if (!IS_ALIGNED(phys_addr, PMD_SIZE)) return 0; =20 - if (pmd_present(*pmd) && !pmd_free_pte_page(pmd, addr)) - return 0; + if (!pmd_present(*pmd)) + return pmd_set_huge(pmd, phys_addr, prot); =20 - return pmd_set_huge(pmd, phys_addr, prot); + /* + * Acquire the mmap read lock to exclude ptdump, which walks + * kernel page tables it does not own under the mmap write lock. + * + * Concurrent read lock holders are safe: each exclusively owns + * the range it operates on and cannot reach this page table. + */ +#ifndef CONFIG_ARM64 + scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) +#endif + { + if (!pmd_free_pte_page(pmd, addr)) + return 0; + return pmd_set_huge(pmd, phys_addr, prot); + } } =20 static int vmap_pmd_range(pud_t *pud, unsigned long addr, unsigned long en= d, @@ -210,10 +225,18 @@ static int vmap_try_huge_pud(pud_t *pud, unsigned lon= g addr, unsigned long end, if (!IS_ALIGNED(phys_addr, PUD_SIZE)) return 0; =20 - if (pud_present(*pud) && !pud_free_pmd_page(pud, addr)) - return 0; + if (!pud_present(*pud)) + return pud_set_huge(pud, phys_addr, prot); =20 - return pud_set_huge(pud, phys_addr, prot); + /* See comment in vmap_try_huge_pmd(). */ +#ifndef CONFIG_ARM64 + scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) +#endif + { + if (!pud_free_pmd_page(pud, addr)) + return 0; + return pud_set_huge(pud, phys_addr, prot); + } } =20 static int vmap_pud_range(p4d_t *p4d, unsigned long addr, unsigned long en= d, @@ -262,10 +285,18 @@ static int vmap_try_huge_p4d(p4d_t *p4d, unsigned lon= g addr, unsigned long end, if (!IS_ALIGNED(phys_addr, P4D_SIZE)) return 0; =20 - if (p4d_present(*p4d) && !p4d_free_pud_page(p4d, addr)) - return 0; + if (!p4d_present(*p4d)) + return p4d_set_huge(p4d, phys_addr, prot); =20 - return p4d_set_huge(p4d, phys_addr, prot); + /* See comment in vmap_try_huge_pmd(). */ +#ifndef CONFIG_ARM64 + scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) +#endif + { + if (!p4d_free_pud_page(p4d, addr)) + return 0; + return p4d_set_huge(p4d, phys_addr, prot); + } } =20 static int vmap_p4d_range(pgd_t *pgd, unsigned long addr, unsigned long en= d, --=20 2.55.0 From nobody Sat Jul 25 06:09:49 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 842233BFAEB; Thu, 16 Jul 2026 21:31:54 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784237515; cv=none; b=iiE36P/3lNELdCdMyrnyZbuDv8h2Ai/iQvIYvfN0V1z91+Xyi9k0mYSsegOmW+Isi3DSAQOlMbR6z6RjEdG9j7SxY4cwGQIczjlGKYBS7Iq/GO0bW7DC7X/X0AWHeeCFqAFnhtZSYvaabdGOeQBWxVsPwZ6xL8n5vfE4gz4Jeak= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784237515; c=relaxed/simple; bh=Tr/TB2AVwbvv7zzaS9pkOVNBodaCW3bWJxJVOhSXA9M=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=kbO2okDkJ5Dpp7ITD5OIMwd7uS7fnmWI2C8x1JStAYSvxmH6q2G+3QzA0Hc6FAKtxOK+RTOzYh/1IrzzFnkR3I87D2TDeoL9Ayxg3HM86917VkJdXR5c8fwrWu4OwXT6YJUczqmPRjZUs6t0xOX8HFx2Yc1zS6w0ItPm64N+4Bg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=HVcYQQqp; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="HVcYQQqp" Received: by smtp.kernel.org (Postfix) with ESMTPSA id F0E601F000E9; Thu, 16 Jul 2026 21:31:47 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784237514; bh=Op085aHgq7EGWjw704qIT6T5yZRCkSyhbHQ1mSMEI8s=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=HVcYQQqpLoh6Q92sYkqOm0bXJh+/09opzvs/PvUscy8+isf3kIAL8ZxMw8DPBpRUZ sElAaONNpmvFtREwnkiCc9thHVE3vOzVGkgrqTF4A0CPCLUZSGGqQ/JIlBymedlrth Qabki4DyCOSHdyGgt/kNhbk06oxFYVqN5CslunAP6MEv8eyO5mPTDVbCRuRDINEf93 q5tq93tz5Rg6gIrfqN9mCUOqHTDFydtEEEN49/0tCFjZnFebXrQGXpc9PWnem0oe0z P4nn+geMJHZJgprDbxzA+3wgXHpaQi2xvxFL8Ybima+1iG4mArdAkJQ9wRXlPQ3iH5 t2o8ckdjevp0Q== From: "Lorenzo Stoakes (ARM)" Date: Thu, 16 Jul 2026 22:31:13 +0100 Subject: [PATCH mm-hotfixes v4 2/4] x86/mm/pat: acquire init_mm write lock to avoid UAF Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260716-series-vmap-race-fix-v4-2-8c108c4317df@kernel.org> References: <20260716-series-vmap-race-fix-v4-0-8c108c4317df@kernel.org> In-Reply-To: <20260716-series-vmap-race-fix-v4-0-8c108c4317df@kernel.org> To: Andrew Morton , Suren Baghdasaryan , "Liam R. Howlett" , Vlastimil Babka , Shakeel Butt , David Hildenbrand , Mike Rapoport , Michal Hocko , Uladzislau Rezki , Toshi Kani , Dave Hansen , Andy Lutomirski , Peter Zijlstra , Thomas Gleixner , Ingo Molnar , Borislav Petkov , x86@kernel.org, "H. Peter Anvin" , Kiryl Shutsemau , Catalin Marinas , Will Deacon , Dev Jain , Ryan Roberts Cc: David Carlier , ljs@kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, "Denis V. Lunev" , stable@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=3954; i=ljs@kernel.org; h=from:subject:message-id; bh=Tr/TB2AVwbvv7zzaS9pkOVNBodaCW3bWJxJVOhSXA9M=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLIifdeu/9y1Rs3Jfe1W7Rldp5bbKgfO8/reO2mZU871y nuTn5RpdZSyMIhxMciKKbI8/yK+P0gkbF7nBX83mDmsTCBDGLg4BWAiE4sZ/lfu/fMi2iYgTWL3 lP9NRxWPKwf2N4e7PlYOv9JRLV8oFMzwv0Blc0TXr6sXurTN5dUy+k2SD8lnn55fvj+K+YqCfm8 KHwA= X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 x86 implements page attribute modification using its Change Page Attributes (CPA) mechanism. This tracks properties of ranges such as cache mode through x86 page attributes, and as part of that logic manipulates kernel page tables. Since commit 41d88484c71c ("x86/mm/pat: restore large ROX pages after fragmentation") ranges of kernel page table entries can be collapsed into huge page table entries as part of this logic. As part of this collapse, it frees the page tables which the collapsed entries previously pointed to, and it does so without any relevant locks being held to preclude concurrent kernel page table walkers. The only way this code can be reached is if CPA_COLLAPSE is specified, and this is only set in set_memory_rox() via: set_memory_rox() -> change_page_attr_set_clr() -> cpa_flush() -> cpa_collapse_large_pages() Notable users of this are execmem and bpf when manipulating executable mappings. However, this is problematic for ptdump as it walks ranges it does not own and thus runs the risk of a use-after-free on page tables freed underneath it. In addition, concurrent CPA collapse operations are possible which can also cause races. Resolve the issue by acquiring the mmap write lock on init_mm across the whole operation. It is safe to acquire a sleeping lock as all the callers invoke set_memory_rox() from process context and in any case, change_page_attr_set_clr() calls vm_unmap_alias() which ultimately takes a mutex, disallowing atomic context here. Fixes: 41d88484c71c ("x86/mm/pat: restore large ROX pages after fragmentati= on") Cc: stable@vger.kernel.org Reviewed-by: Mike Rapoport (Microsoft) Reviewed-by: Kiryl Shutsemau (Meta) Reviewed-by: David Hildenbrand (Arm) Reviewed-by: Dave Hansen Signed-off-by: Lorenzo Stoakes (ARM) Reviewed-by: David Carlier Reviewed-by: Will Deacon --- arch/x86/mm/pat/set_memory.c | 15 ++++++++++++++- include/linux/mmap_lock.h | 2 ++ 2 files changed, 16 insertions(+), 1 deletion(-) diff --git a/arch/x86/mm/pat/set_memory.c b/arch/x86/mm/pat/set_memory.c index d023a40a1e03..d1e63f7d267f 100644 --- a/arch/x86/mm/pat/set_memory.c +++ b/arch/x86/mm/pat/set_memory.c @@ -22,6 +22,7 @@ #include #include #include +#include =20 #include #include @@ -410,7 +411,7 @@ static void __cpa_flush_tlb(void *data) =20 static int collapse_large_pages(unsigned long addr, struct list_head *pgta= bles); =20 -static void cpa_collapse_large_pages(struct cpa_data *cpa) +static void __cpa_collapse_large_pages(struct cpa_data *cpa) { unsigned long start, addr, end; struct ptdesc *ptdesc, *tmp; @@ -442,6 +443,18 @@ static void cpa_collapse_large_pages(struct cpa_data *= cpa) } } =20 +static void cpa_collapse_large_pages(struct cpa_data *cpa) +{ + /* + * Take the mmap write lock on init_mm to: + * - Avoid a use-after-free if raced by ptdump (which takes its own + * write lock on init_mm). + * - Serialise concurrent CPA walkers. + */ + scoped_guard(mmap_write_lock, &init_mm) + __cpa_collapse_large_pages(cpa); +} + static void cpa_flush(struct cpa_data *cpa, int cache) { unsigned int i; diff --git a/include/linux/mmap_lock.h b/include/linux/mmap_lock.h index 6b5c2390cc30..047f5f5e2c34 100644 --- a/include/linux/mmap_lock.h +++ b/include/linux/mmap_lock.h @@ -621,6 +621,8 @@ static inline void mmap_read_unlock(struct mm_struct *m= m) =20 DEFINE_GUARD(mmap_read_lock, struct mm_struct *, mmap_read_lock(_T), mmap_read_unlock(_T)) +DEFINE_GUARD(mmap_write_lock, struct mm_struct *, + mmap_write_lock(_T), mmap_write_unlock(_T)) DEFINE_GUARD_COND(mmap_read_lock, _try, mmap_read_trylock(_T)) =20 static inline void mmap_read_unlock_non_owner(struct mm_struct *mm) --=20 2.55.0 From nobody Sat Jul 25 06:09:49 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0B8F63C3F69; Thu, 16 Jul 2026 21:32:00 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784237522; cv=none; b=um5Y2TP1bZD8KntBJbOnZt2NFUsdw3GW/6eoaiXeMNJLAD+RmUyMs9mz7pjyjFKUV+8eNACxzeOcyQbbgdzxZGPxQx7UmTWM8vjCabJiItjzlIbbmsWkDX5oaiof1VjvOaeKvksMqdajzbVOuUTOipfacFKVzS+3ipdkgj6PqzM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784237522; c=relaxed/simple; bh=VITXU838ytmI5wjkeVjJjzJ76fi6Ga2j8Ljay/OkRtw=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=DB/PenWhXDhEN8ik8BH+xDnHAmxaSJFSOUvXguzVJ2fe9SB1hEvGb40WMDlfb+xAQVJCCoa2xPPA1DO3uskHPUKAlyT807Rob4CbPxJAYlKqyyU/2gZyH1QY6mgAEQAu2gDt45HCEN8ZIa3SxNVYZaaMzDnhTxHvy3zM7x3UBfw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=GYgZQmvm; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="GYgZQmvm" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 8ECAE1F00A3A; Thu, 16 Jul 2026 21:31:54 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784237520; bh=4TmiXyJvQ92TKxUnkF2EBAbSx4BVAEx8E3qvkE5M06w=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=GYgZQmvmtJSCVwlxmpjDgZpiPmaoPOUEQojBo7h8SSWNVzSGj11N/GUGKOHB/4X7E xjyPpfeJfxld/9aBWN4fwcoHpUX+rynA8+i1bnLjwmsTf1Dn3tQf0ldvNMYAOJ3T+g ZYr19KOkbGbG3bwcoR43r0vZOR9u/ENab3upSYP1tLsLhaU5z9woe4i3Vv9e1ZRora iZW4tyHUfHzthiX+ve4ZGiM+50GiUG9l/3ow+618/ct2YyCbuNXjqRtNyjwymX+5oS /kTxJlA8i/wPaUpCmiRfbpO4XiuFMRJncNBkmpIftb1iLfatGaI/pp2BbHa+qq1gny p6Jb+93YStcmQ== From: "Lorenzo Stoakes (ARM)" Date: Thu, 16 Jul 2026 22:31:14 +0100 Subject: [PATCH mm-hotfixes v4 3/4] mm/ptdump: always stabilise against page table freeing using init_mm Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260716-series-vmap-race-fix-v4-3-8c108c4317df@kernel.org> References: <20260716-series-vmap-race-fix-v4-0-8c108c4317df@kernel.org> In-Reply-To: <20260716-series-vmap-race-fix-v4-0-8c108c4317df@kernel.org> To: Andrew Morton , Suren Baghdasaryan , "Liam R. Howlett" , Vlastimil Babka , Shakeel Butt , David Hildenbrand , Mike Rapoport , Michal Hocko , Uladzislau Rezki , Toshi Kani , Dave Hansen , Andy Lutomirski , Peter Zijlstra , Thomas Gleixner , Ingo Molnar , Borislav Petkov , x86@kernel.org, "H. Peter Anvin" , Kiryl Shutsemau , Catalin Marinas , Will Deacon , Dev Jain , Ryan Roberts Cc: David Carlier , ljs@kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, "Denis V. Lunev" , stable@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=3325; i=ljs@kernel.org; h=from:subject:message-id; bh=VITXU838ytmI5wjkeVjJjzJ76fi6Ga2j8Ljay/OkRtw=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLIifdeK1H72/9qooXMh7OxZvidbk/uXLSj4tneKLMPTC Ya+NpmbOkpZGMS4GGTFFFmefxHfHyQSNq/zgr8bzBxWJpAhDFycAjARTkFGhjPp5b9fLnBdf+/V qrsRtnMETV8qZqkLlsdxxuhvz93xRISR4YmPyWsx52chzOyJO2bXi16zvJtqLaHdqDpniUxxyfd 4bgA= X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 Previous commits have established the invariant that kernel page table freeing is performed while an mmap read lock on init_mm is held, which fixes races between ptdump and kernel page table freeing over init_mm. However, x86 and arm64 can perform a ptdump over an mm other than init_mm via ptdump_walk_pgd() and since kernel memory ranges are shared across non-kernel mm's, this means that the race still exists for these cases. Fix this by acquiring a nested mmap write lock for init_mm in ptdump_walk_pgd(). This is safe as we take this after mmap write locking the mm, and nothing acquires the init_mm lock first before locking an arbitrary mm, so no deadlock is possible. Also update walk_page_range_debug() to assert that init_mm is write locked, add a comment explaining why and remove some redundant code, and eliminate the unnecessary and confusing invocation of walk_kernel_page_table_range(). We can safely remove the non-NULL check for walk.mm, as the mmap lock asserts would NULL pointer deref if it was (and of course no callers do this). The first point at which ptdump can race kernel page table freeing is commit b6bdb7517c3d ("mm/vmalloc: add interfaces to free unmapped page table"), so we target this in the Fixes tag. Fixes: b6bdb7517c3d ("mm/vmalloc: add interfaces to free unmapped page tabl= e") Cc: stable@vger.kernel.org Reviewed-by: Mike Rapoport (Microsoft) Acked-by: David Hildenbrand (Arm) Reviewed-by: Kiryl Shutsemau Signed-off-by: Lorenzo Stoakes (ARM) --- mm/pagewalk.c | 14 +++++++++----- mm/ptdump.c | 7 +++++++ 2 files changed, 16 insertions(+), 5 deletions(-) diff --git a/mm/pagewalk.c b/mm/pagewalk.c index bbcfd68d0907..5d87c632a255 100644 --- a/mm/pagewalk.c +++ b/mm/pagewalk.c @@ -702,12 +702,16 @@ int walk_page_range_debug(struct mm_struct *mm, unsig= ned long start, * to account for page table freeing on vmap huge page mapping. */ mmap_assert_write_locked(mm); + /* + * x86, arm64 ptdump allow walks of efi mm's and x86 ptdump allows walks + * of arbitrary mm's. + * + * However, they both must also hold the init_mm lock to account for + * concurrent kernel page table freeing. + */ + mmap_assert_write_locked(&init_mm); =20 - /* For convenience, we allow traversal of kernel mappings. */ - if (mm =3D=3D &init_mm) - return walk_kernel_page_table_range(start, end, ops, - pgd, private); - if (start >=3D end || !walk.mm) + if (start >=3D end) return -EINVAL; if (!check_ops_safe(ops)) return -EINVAL; diff --git a/mm/ptdump.c b/mm/ptdump.c index 973020000096..5851096e6f65 100644 --- a/mm/ptdump.c +++ b/mm/ptdump.c @@ -178,11 +178,18 @@ void ptdump_walk_pgd(struct ptdump_state *st, struct = mm_struct *mm, pgd_t *pgd) =20 get_online_mems(); mmap_write_lock(mm); + /* To stabilise kernel page tables we must hold the init_mm lock too. */ + if (mm !=3D &init_mm) + mmap_write_lock_nested(&init_mm, SINGLE_DEPTH_NESTING); + while (range->start !=3D range->end) { walk_page_range_debug(mm, range->start, range->end, &ptdump_ops, pgd, st); range++; } + + if (mm !=3D &init_mm) + mmap_write_unlock(&init_mm); mmap_write_unlock(mm); put_online_mems(); =20 --=20 2.55.0 From nobody Sat Jul 25 06:09:49 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A4FCB222590; Thu, 16 Jul 2026 21:32:07 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784237529; cv=none; b=dLb0MbXFL4Bpprh255VVXv+k8W5vmYNtGSNp/wIGfAJlfdmnvxdlgtbUUXpCOEhHTzIDRPuvvPuo7EH/D3j9t7NRsejaKezyG0x0wzNb7sMyJBI6b2PHOoatHTEOMftYx6XDVRhri66xjxjUR3uCSCond4AcA/Mtp56FI+gEDxQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784237529; c=relaxed/simple; bh=kE9XtB6iyw7BSFNAuuPgI01clqwgzMyGQ2yqQHsLjJM=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=L+Nu3b4dD6/1YiX2nP3Wp1ZeNOB9G8qawstbR0WJW0uXw6np0vvzBTDgxYWYVF8MCbOEI+W9i0ILjH8j8ijrFXeSO2BA65dNLL7zut/pOMgxuTGKbWYfUHZzSObQf3/2C6QOfB6frvL4GRc/Xgm0r7JSo9KxQqYyfrXRLd8v0cU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=NmwgOa92; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="NmwgOa92" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 2EEF91F000E9; Thu, 16 Jul 2026 21:32:00 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784237527; bh=6tB+p5fpJihvWmKYtOqn6XBXbSKHjdM1g/8Wc7lnKAc=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=NmwgOa92PeD48+bLHUAnvL16XjRKn0h3kz3vESeNv5Xhqnj3ggmzIWBXdFEPyMcTX fJhBMOQOJQ+2iCO5mPto30hTGE6qL8Oe7ldjnmQ9Re0i2HQiGK2V9lXBhR1CfCV3XB 5q6AyYOfIWL1Rr+Cqd18GE735zTE8lg4QQxygLEX8KLtETwV6squjFXrHuiRYNMgeu cTNJq84ysVaj6Ztf3t+vk3f1YCJkYod87ejGP8R56cJT4K19UAtzIYDZFCRZR65HI9 IQHJF9wIu3f6XvEO+sRaVkGaiehTmKdKS1mEnLxprdPDXi0g1kLIHifWgM14ekZJBC Dt1c6JEj6f3CQ== From: "Lorenzo Stoakes (ARM)" Date: Thu, 16 Jul 2026 22:31:15 +0100 Subject: [PATCH mm-hotfixes v4 4/4] arm64: remove redundant concurrent ptdump UAF mitigation Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260716-series-vmap-race-fix-v4-4-8c108c4317df@kernel.org> References: <20260716-series-vmap-race-fix-v4-0-8c108c4317df@kernel.org> In-Reply-To: <20260716-series-vmap-race-fix-v4-0-8c108c4317df@kernel.org> To: Andrew Morton , Suren Baghdasaryan , "Liam R. Howlett" , Vlastimil Babka , Shakeel Butt , David Hildenbrand , Mike Rapoport , Michal Hocko , Uladzislau Rezki , Toshi Kani , Dave Hansen , Andy Lutomirski , Peter Zijlstra , Thomas Gleixner , Ingo Molnar , Borislav Petkov , x86@kernel.org, "H. Peter Anvin" , Kiryl Shutsemau , Catalin Marinas , Will Deacon , Dev Jain , Ryan Roberts Cc: David Carlier , ljs@kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, bpf@vger.kernel.org, linux-arm-kernel@lists.infradead.org, "Denis V. Lunev" , stable@vger.kernel.org X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=openpgp-sha256; l=7063; i=ljs@kernel.org; h=from:subject:message-id; bh=kE9XtB6iyw7BSFNAuuPgI01clqwgzMyGQ2yqQHsLjJM=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLIifdfqWl5lylpfsPmAvrPrt6Vyx7bu6VxW84JjddT7M 4Xb/+671FHKwiDGxSArpsjy/Iv4/iCRsHmdF/zdYOawMoEMYeDiFICJvLNjZHi6UU7r53+HlR+L P+QsMfBrOaYzf+0HlpYiUeYfKoXi+c8Y/kfnLNGy2d6aEJ72yVTwOdPS/lK523P+Poy9fOGY1Td vERYA X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 This partially reverts commit fa93b45fd397 ("arm64: Enable vmalloc-huge with ptdump"), retaining vmalloc-huge support but eliminating the now redundant mitigation against a race between huge vmap page table freeing and ptdump, as this issue has now been fixed at core. We also simultaneously remove the arm64 if-deffery when acquiring the mmap read lock upon vmap huge page table promotion as it is no longer required. Note that this patch relies on the preceding vmalloc patch, and should not be backported alone. Fixes: fa93b45fd397 ("arm64: Enable vmalloc-huge with ptdump") Cc: stable@vger.kernel.org Reviewed-by: Dev Jain Acked-by: Mike Rapoport (Microsoft) Acked-by: Kiryl Shutsemau (Meta) Acked-by: Will Deacon Reviewed-by: David Hildenbrand (Arm) Signed-off-by: Lorenzo Stoakes (ARM) --- arch/arm64/include/asm/ptdump.h | 2 -- arch/arm64/mm/mmu.c | 43 ++++---------------------------------= ---- arch/arm64/mm/ptdump.c | 11 ++--------- mm/vmalloc.c | 15 +++----------- 4 files changed, 9 insertions(+), 62 deletions(-) diff --git a/arch/arm64/include/asm/ptdump.h b/arch/arm64/include/asm/ptdum= p.h index 5b374a6ab34a..50a195eda8ed 100644 --- a/arch/arm64/include/asm/ptdump.h +++ b/arch/arm64/include/asm/ptdump.h @@ -7,8 +7,6 @@ =20 #include =20 -DECLARE_STATIC_KEY_FALSE(arm64_ptdump_lock_key); - #ifdef CONFIG_PTDUMP =20 #include diff --git a/arch/arm64/mm/mmu.c b/arch/arm64/mm/mmu.c index a25d8beacc83..bd52fca6e872 100644 --- a/arch/arm64/mm/mmu.c +++ b/arch/arm64/mm/mmu.c @@ -49,8 +49,6 @@ #define NO_CONT_MAPPINGS BIT(1) #define NO_EXEC_MAPPINGS BIT(2) /* assumes FEAT_HPDS is not used */ =20 -DEFINE_STATIC_KEY_FALSE(arm64_ptdump_lock_key); - u64 kimage_voffset __ro_after_init; EXPORT_SYMBOL(kimage_voffset); =20 @@ -1864,8 +1862,7 @@ int pmd_clear_huge(pmd_t *pmdp) return 1; } =20 -static int __pmd_free_pte_page(pmd_t *pmdp, unsigned long addr, - bool acquire_mmap_lock) +int pmd_free_pte_page(pmd_t *pmdp, unsigned long addr) { pte_t *table; pmd_t pmd; @@ -1877,25 +1874,13 @@ static int __pmd_free_pte_page(pmd_t *pmdp, unsigne= d long addr, return 1; } =20 - /* See comment in pud_free_pmd_page for static key logic */ table =3D pte_offset_kernel(pmdp, addr); pmd_clear(pmdp); __flush_tlb_kernel_pgtable(addr); - if (static_branch_unlikely(&arm64_ptdump_lock_key) && acquire_mmap_lock) { - mmap_read_lock(&init_mm); - mmap_read_unlock(&init_mm); - } - pte_free_kernel(NULL, table); return 1; } =20 -int pmd_free_pte_page(pmd_t *pmdp, unsigned long addr) -{ - /* If ptdump is walking the pagetables, acquire init_mm.mmap_lock */ - return __pmd_free_pte_page(pmdp, addr, /* acquire_mmap_lock =3D */ true); -} - int pud_free_pmd_page(pud_t *pudp, unsigned long addr) { pmd_t *table; @@ -1911,36 +1896,16 @@ int pud_free_pmd_page(pud_t *pudp, unsigned long ad= dr) } =20 table =3D pmd_offset(pudp, addr); - - /* - * Our objective is to prevent ptdump from reading a PMD table which has - * been freed. In this race, if pud_free_pmd_page observes the key on - * (which got flipped by ptdump) then the mmap lock sequence here will, - * as a result of the mmap write lock/unlock sequence in ptdump, give - * us the correct synchronization. If not, this means that ptdump has - * yet not started walking the pagetables - the sequence of barriers - * issued by __flush_tlb_kernel_pgtable() guarantees that ptdump will - * observe an empty PUD. - */ - pud_clear(pudp); - __flush_tlb_kernel_pgtable(addr); - if (static_branch_unlikely(&arm64_ptdump_lock_key)) { - mmap_read_lock(&init_mm); - mmap_read_unlock(&init_mm); - } - pmdp =3D table; next =3D addr; end =3D addr + PUD_SIZE; do { if (pmd_present(pmdp_get(pmdp))) - /* - * PMD has been isolated, so ptdump won't see it. No - * need to acquire init_mm.mmap_lock. - */ - __pmd_free_pte_page(pmdp, next, /* acquire_mmap_lock =3D */ false); + pmd_free_pte_page(pmdp, next); } while (pmdp++, next +=3D PMD_SIZE, next !=3D end); =20 + pud_clear(pudp); + __flush_tlb_kernel_pgtable(addr); pmd_free(NULL, table); return 1; } diff --git a/arch/arm64/mm/ptdump.c b/arch/arm64/mm/ptdump.c index 1c20144700d7..5a76c59b5ada 100644 --- a/arch/arm64/mm/ptdump.c +++ b/arch/arm64/mm/ptdump.c @@ -283,13 +283,6 @@ void note_page_flush(struct ptdump_state *pt_st) note_page(pt_st, 0, -1, pte_val(pte_zero)); } =20 -static void arm64_ptdump_walk_pgd(struct ptdump_state *st, struct mm_struc= t *mm) -{ - static_branch_inc(&arm64_ptdump_lock_key); - ptdump_walk_pgd(st, mm, NULL); - static_branch_dec(&arm64_ptdump_lock_key); -} - void ptdump_walk(struct seq_file *s, struct ptdump_info *info) { unsigned long end =3D ~0UL; @@ -318,7 +311,7 @@ void ptdump_walk(struct seq_file *s, struct ptdump_info= *info) } }; =20 - arm64_ptdump_walk_pgd(&st.ptdump, info->mm); + ptdump_walk_pgd(&st.ptdump, info->mm, NULL); } =20 static void __init ptdump_initialize(void) @@ -360,7 +353,7 @@ bool ptdump_check_wx(void) } }; =20 - arm64_ptdump_walk_pgd(&st.ptdump, &init_mm); + ptdump_walk_pgd(&st.ptdump, &init_mm, NULL); =20 if (st.wx_pages || st.uxn_pages) { pr_warn("Checked W+X mappings: FAILED, %lu W+X pages found, %lu non-UXN = pages found\n", diff --git a/mm/vmalloc.c b/mm/vmalloc.c index d5c4d2bb770b..f4fa227a8d7f 100644 --- a/mm/vmalloc.c +++ b/mm/vmalloc.c @@ -169,10 +169,7 @@ static int vmap_try_huge_pmd(pmd_t *pmd, unsigned long= addr, unsigned long end, * Concurrent read lock holders are safe: each exclusively owns * the range it operates on and cannot reach this page table. */ -#ifndef CONFIG_ARM64 - scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) -#endif - { + scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) { if (!pmd_free_pte_page(pmd, addr)) return 0; return pmd_set_huge(pmd, phys_addr, prot); @@ -229,10 +226,7 @@ static int vmap_try_huge_pud(pud_t *pud, unsigned long= addr, unsigned long end, return pud_set_huge(pud, phys_addr, prot); =20 /* See comment in vmap_try_huge_pmd(). */ -#ifndef CONFIG_ARM64 - scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) -#endif - { + scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) { if (!pud_free_pmd_page(pud, addr)) return 0; return pud_set_huge(pud, phys_addr, prot); @@ -289,10 +283,7 @@ static int vmap_try_huge_p4d(p4d_t *p4d, unsigned long= addr, unsigned long end, return p4d_set_huge(p4d, phys_addr, prot); =20 /* See comment in vmap_try_huge_pmd(). */ -#ifndef CONFIG_ARM64 - scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) -#endif - { + scoped_cond_guard(mmap_read_lock_try, return 0, &init_mm) { if (!p4d_free_pud_page(p4d, addr)) return 0; return p4d_set_huge(p4d, phys_addr, prot); --=20 2.55.0