From nobody Wed Sep 30 05:00:31 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 80C563B95E3; Thu, 13 Aug 2026 09:01:48 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786611712; cv=none; b=Ctv5hlOs9i+A5M2Zpru3FMtrzF2Yw/+RfopSMRlhJ7oc4l/uw+aldULik9g2dbk5VyANyiCrfQF9c4ShbVrZjVWv6t8YaK8UVKjhbt+Ar4nKdQ9ImrJKMpBrZMVVyzITqxAtquwj0cGZjt+1Xf3Zt9NmScFfUYDGn0cB0x0uN30= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786611712; c=relaxed/simple; bh=Mjrsg1JU/mJdl1NQFbdbtQ8yw6hXZhbIGt1umXBnvD8=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=GGuPDfZjjSsIAywcRCaYqGJeo9N1TZxBkxBtw0y/ucWMXfgllhknew2wxiUCu7sjwv5VkmTLGfEnIAbQQ8n2E139z+3fLaKZk7ZO3Dnq/VAXqUoKwGJ7+5ODoA1Bar8bKXRFWx1jLAtRY2r4++P0vmNyjoEyM6MgBaxxI3f8CWE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=PH2JAqBI; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="PH2JAqBI" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 2C6581F00A3D; Thu, 13 Aug 2026 09:01:40 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786611708; bh=CRSJV6+wHMCp/NANau/roLEK+lDHPucZMiMPxBMKL5U=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=PH2JAqBIaEBKerWzSh1Dt4eMzbUkBQlgGpe8Iwq8jlPTpKtSGwCa7MMcudT4pyKdb m2SP9Zyw4VfKA2fYBdSvAE8xpn1sJ7+yjEOUVNRh+M8FA5uUlX19MD5iPybqbOygZH pHzw9NcRz7HVK43LC/p6b71fHO2zdidVMPtdlpZfvlT5T+26dEYfe3iH1OHBYxC8Ic FIuwIqjq4lWbS3HVvSm+cDUjGKf/1scH974kVr0txKC9FAa8ZL1Nz7WJPHwaTHp1KN eC0V9tjfcJpKdwj6famWSz4jdYfd2EuxwkAcCVF8nhMInbu/gdamdyYhdpBBB1mM2P 3gjX+Vb6NoUeQ== From: Mike Rapoport Date: Thu, 13 Aug 2026 12:01:24 +0300 Subject: [PATCH v2 1/5] x86/mm/pat: acquire init_mm write lock on collapse to avoid UAF Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260813-cpa-fixes-v2-1-39b4ff90f91d@kernel.org> References: <20260813-cpa-fixes-v2-0-39b4ff90f91d@kernel.org> In-Reply-To: <20260813-cpa-fixes-v2-0-39b4ff90f91d@kernel.org> To: Dave Hansen Cc: Andrew Morton , Andy Lutomirski , Borislav Petkov , David CARLIER , David Hildenbrand , Ingo Molnar , Jason Gunthorpe , Jiri Slaby , Juergen Gross , Kevin Tian , Kiryl Shutsemau , "Liam R. Howlett" , Lorenzo Stoakes , Lu Baolu , Mike Rapoport , Nikunj A Dadhania , Pedro Falcato , "H. Peter Anvin" , Peter Zijlstra , Shakeel Butt , Steffen Dirkwinkel , Suren Baghdasaryan , Thomas Gleixner , Toshi Kani , Vishal Moola , Vlastimil Babka , Will Deacon , iommu@lists.linux.dev, linux-kernel@vger.kernel.org, linux-mm@kvack.org, stable@vger.kernel.org, syzbot@syzkaller.appspotmail.com, x86@kernel.org X-Mailer: b4 0.17-dev From: "Lorenzo Stoakes (ARM)" x86 implements page attribute modification using its Change Page Attributes (CPA) mechanism. This tracks properties of ranges such as cache mode through x86 page attributes, and as part of that logic manipulates kernel page tables. Since commit 41d88484c71c ("x86/mm/pat: restore large ROX pages after fragmentation") ranges of kernel page table entries can be collapsed into huge page table entries as part of this logic. As part of this collapse, it frees the page tables which the collapsed entries previously pointed to, and it does so without any relevant locks being held to preclude concurrent kernel page table walkers. The only way this code can be reached is if CPA_COLLAPSE is specified, and this is only set in set_memory_rox() via: set_memory_rox() -> change_page_attr_set_clr() -> cpa_flush() -> cpa_collapse_large_pages() Notable users of this are execmem and bpf when manipulating executable mappings. However, this is problematic for ptdump as it walks ranges it does not own and thus runs the risk of a use-after-free on page tables freed underneath it. In addition, concurrent CPA collapse operations are possible which can also cause races. Resolve the issue by acquiring the mmap write lock on init_mm across the whole operation. It is safe to acquire a sleeping lock as all the callers invoke set_memory_rox() from process context and in any case, change_page_attr_set_clr() calls vm_unmap_alias() which ultimately takes a mutex, disallowing atomic context here. Fixes: 41d88484c71c ("x86/mm/pat: restore large ROX pages after fragmentati= on") Cc: stable@vger.kernel.org Reviewed-by: Mike Rapoport (Microsoft) Reviewed-by: Kiryl Shutsemau (Meta) Reviewed-by: David Hildenbrand (Arm) Reviewed-by: Dave Hansen Reviewed-by: Will Deacon Reviewed-by: David Carlier Signed-off-by: Lorenzo Stoakes (ARM) Signed-off-by: Mike Rapoport (Microsoft) Tested-by: Atish Patra Tested-by: Nikunj A Dadhania --- arch/x86/mm/pat/set_memory.c | 15 ++++++++++++++- include/linux/mmap_lock.h | 2 ++ 2 files changed, 16 insertions(+), 1 deletion(-) diff --git a/arch/x86/mm/pat/set_memory.c b/arch/x86/mm/pat/set_memory.c index d8d057f44417..c18b887ee4c2 100644 --- a/arch/x86/mm/pat/set_memory.c +++ b/arch/x86/mm/pat/set_memory.c @@ -22,6 +22,7 @@ #include #include #include +#include =20 #include #include @@ -409,7 +410,7 @@ static void __cpa_flush_tlb(void *data) =20 static int collapse_large_pages(unsigned long addr, struct list_head *pgta= bles); =20 -static void cpa_collapse_large_pages(struct cpa_data *cpa) +static void __cpa_collapse_large_pages(struct cpa_data *cpa) { unsigned long start, addr, end; struct ptdesc *ptdesc, *tmp; @@ -443,6 +444,18 @@ static void cpa_collapse_large_pages(struct cpa_data *= cpa) } } =20 +static void cpa_collapse_large_pages(struct cpa_data *cpa) +{ + /* + * Take the mmap write lock on init_mm to: + * - Avoid a use-after-free if raced by ptdump (which takes its own + * write lock on init_mm). + * - Serialise concurrent CPA walkers. + */ + scoped_guard(mmap_write_lock, &init_mm) + __cpa_collapse_large_pages(cpa); +} + static void cpa_flush(struct cpa_data *cpa, int cache) { unsigned int i; diff --git a/include/linux/mmap_lock.h b/include/linux/mmap_lock.h index 04b8f61ece5d..f4ceb968aeb3 100644 --- a/include/linux/mmap_lock.h +++ b/include/linux/mmap_lock.h @@ -621,6 +621,8 @@ static inline void mmap_read_unlock(struct mm_struct *m= m) =20 DEFINE_GUARD(mmap_read_lock, struct mm_struct *, mmap_read_lock(_T), mmap_read_unlock(_T)) +DEFINE_GUARD(mmap_write_lock, struct mm_struct *, + mmap_write_lock(_T), mmap_write_unlock(_T)) =20 static inline void mmap_read_unlock_non_owner(struct mm_struct *mm) { --=20 2.53.0 From nobody Wed Sep 30 05:00:31 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7A33E3B893F; Thu, 13 Aug 2026 09:01:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786611722; cv=none; b=KZjs1bsb4PqJ0QxPrTrmB8gR0iMOaiTb0I19pZs1NvmcAXjsKfTFuzRLL47tZsIx/r1MNPxSAXd9D99br0HYX1wM0R4Z76iZDVCzzU+EoSGfGPvvCDXQF4borm/YC2Ls23L87dytLmzmqdP8l7x6Zuzt7WJMJSYRmAXNxSJViGU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786611722; c=relaxed/simple; bh=WSIPSNqUu+ryx5Rg2F8NkgOUmUQiNtKd7tK5+iTJPm4=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=SIvLmdfTq7BREt6nCx+/wBUbkOA3wIztsZIzEn/X1IRgKOFt+rE5aJDu6mZXvmOSi1yB/nZbNsitoVP2keLF1FXuoG5dFsZ5Ym0PSOsyhAZYvKnfTk8TGEKCnkhr2qKvgvHM9i8cv30TrsFqGdlPxd2XR9vsLA8cWtew1DDk5AE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=fpl58WiJ; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="fpl58WiJ" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 337AD1F00A3A; Thu, 13 Aug 2026 09:01:48 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786611716; bh=QbqAHeBmTdNz5nx7DbozdY5ynkW62XANFKaegUYp63U=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=fpl58WiJQbCPqzesROD3VU1p5FQ723FiIYReD8pXtJgCqw0JSoITh9dAMypwqb/lD 3l7vR+MjHy6ki83qrUVjTdoYWSt3ylKYcOZXU/vB56miSBcq0kh9twp2J4UwphiM0N ON7qNylifK3/bx6E4gCUQCVx0lcL6XBXSPh/OcJolZjvX0Z626PIaGEe+GigZnCIA3 gJUvZMuuuk14KNU0trmRx1BkKlpwsGEmt5boJ+VmnB0UDXD2rGHU2gDiDK1DCgzo55 stgpR546uhuIoSEsbC5m3b1h0drzy+VhB8AAri1nBM5LgJLiUcsKAMyX9TzmVE/xAb G8H9lBJc7qXFA== From: Mike Rapoport Date: Thu, 13 Aug 2026 12:01:25 +0300 Subject: [PATCH v2 2/5] x86/mm/pat: acquire init_mm read lock on attribute change to avoid UAF Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260813-cpa-fixes-v2-2-39b4ff90f91d@kernel.org> References: <20260813-cpa-fixes-v2-0-39b4ff90f91d@kernel.org> In-Reply-To: <20260813-cpa-fixes-v2-0-39b4ff90f91d@kernel.org> To: Dave Hansen Cc: Andrew Morton , Andy Lutomirski , Borislav Petkov , David CARLIER , David Hildenbrand , Ingo Molnar , Jason Gunthorpe , Jiri Slaby , Juergen Gross , Kevin Tian , Kiryl Shutsemau , "Liam R. Howlett" , Lorenzo Stoakes , Lu Baolu , Mike Rapoport , Nikunj A Dadhania , Pedro Falcato , "H. Peter Anvin" , Peter Zijlstra , Shakeel Butt , Steffen Dirkwinkel , Suren Baghdasaryan , Thomas Gleixner , Toshi Kani , Vishal Moola , Vlastimil Babka , Will Deacon , iommu@lists.linux.dev, linux-kernel@vger.kernel.org, linux-mm@kvack.org, stable@vger.kernel.org, syzbot@syzkaller.appspotmail.com, x86@kernel.org X-Mailer: b4 0.17-dev From: "Lorenzo Stoakes (ARM)" A previous commit protected us against races between ptdump and CPA collapse, however one still exists between attribute changes and collapse as reported by Denis V. Lunev (linked). When an attribute change arises, a lockless page table walker obtains a PTE entry, which is later written to via set_pte_atomic(): ... -> change_page_attr_set_clr() -> __change_page_attr_set_clr() -> __change_page_attr() -> _lookup_address_cpa() -> lookup_address_in_pgd_attr() -> [ lockless page table walker ] -> set_pte_atomic() There is nothing preventing a concurrent CPA collapse which can free the PTE that was retrieved here, resulting in a use-after-free. With the mmap write lock taken on init_mm over CPA collapse, we can now resolve this race by acquiring an mmap read lock on init_mm over __change_page_attr_set_clr(). This locks across the whole operation over which the walk and the PTE entry write occurs, solving the race. It is safe to do this here, as no spinlocks are held upon entry to __change_page_attr_set_clr(). However, the lock must not be held over an allocation, as allocation can trigger reclaim and shrinkers may call into CPA recursively, making deadlocks possible (init_mm -> ... -> fs_reclaim -> init_mm). A page table is allocated when a huge page needs to be split: -> change_page_attr_set_clr() -> __change_page_attr_set_clr() -> __change_page_attr() -> split_large_page() [ pagetable_alloc() ] -> __split_large_page() Avoid deadlocks by dropping the mmap lock across pagetable_alloc() in split_large_page() and track whether this is needed by adding a new 'init_mm_read_locked' flag to struct cpa_data. This is safe as __split_large_page() (called with locks re-established) revalidates that the page table entry is the same as it was prior to the locks being dropped and __change_page_attr() repeats the entire page table walk whenever a split occurs, so concurrent split and collapse are accounted for. Concurrent ptdump is also safe as the lock is only dropped over page table allocation during which time the page table has not yet been modified. The CPA_COLLAPSE flag is only set by set_memory_rox(), which exclusively operates upon vmalloc ranges, and on x86 only within the module mapping space. This is important, because some callers directly invoke __change_page_attr_set_clr(), bypassing this lock. However, none of these operate within the module mapping space. * cpa_process_alias() - a recursive helper called by __change_page_attr_set_clr(). * __set_memory_enc_pgtable() - operates on the direct mapping and (via __vmbus_establish_gpadl()) the vmalloc mapping space. * __set_pages_[n]p() - called by set_direct_map_[invalid, default, valid]_noflush(), __kernel_map_pages() - operates on the direct map. * kernel_[un]map_pages_in_pgd() - operates on EFI ranges. This work is based upon Denis V. Lunev's excellent analysis of the bug with gratitude. Link: https://lore.kernel.org/all/20260626163213.2284080-1-den@openvz.org/ Fixes: 41d88484c71c ("x86/mm/pat: restore large ROX pages after fragmentati= on") Cc: stable@vger.kernel.org Signed-off-by: Lorenzo Stoakes (ARM) Signed-off-by: Mike Rapoport (Microsoft) Tested-by: Atish Patra Tested-by: Nikunj A Dadhania --- arch/x86/mm/pat/set_memory.c | 13 +++++++++++-- 1 file changed, 11 insertions(+), 2 deletions(-) diff --git a/arch/x86/mm/pat/set_memory.c b/arch/x86/mm/pat/set_memory.c index c18b887ee4c2..2d04a4bf34aa 100644 --- a/arch/x86/mm/pat/set_memory.c +++ b/arch/x86/mm/pat/set_memory.c @@ -50,7 +50,8 @@ struct cpa_data { unsigned int flags; unsigned int force_split : 1, force_static_prot : 1, - force_flush_all : 1; + force_flush_all : 1, + init_mm_read_locked : 1; struct page **pages; }; =20 @@ -1240,7 +1241,11 @@ static int split_large_page(struct cpa_data *cpa, pt= e_t *kpte, struct ptdesc *ptdesc; =20 spin_unlock(&cpa_lock); + if (cpa->init_mm_read_locked) + mmap_read_unlock(&init_mm); ptdesc =3D pagetable_alloc(GFP_KERNEL, 0); + if (cpa->init_mm_read_locked) + mmap_read_lock(&init_mm); spin_lock(&cpa_lock); if (!ptdesc) return -ENOMEM; @@ -2134,7 +2139,11 @@ static int change_page_attr_set_clr(unsigned long *a= ddr, int numpages, cpa.curpage =3D 0; cpa.force_split =3D force_split; =20 - ret =3D __change_page_attr_set_clr(&cpa, 1); + /* Avoid race with concurrent CPA collapse. */ + cpa.init_mm_read_locked =3D true; + scoped_guard(mmap_read_lock, &init_mm) + ret =3D __change_page_attr_set_clr(&cpa, 1); + cpa.init_mm_read_locked =3D false; =20 /* * Check whether we really changed something: --=20 2.53.0 From nobody Wed Sep 30 05:00:31 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 09D943B7B8C; Thu, 13 Aug 2026 09:02:06 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786611730; cv=none; b=vARsp16lxVnOU8imzrtP/GOp0aMyCxS7uUqQar3leBlG0aJ0pS1fqfHZL7UMG+aEI9c3bIg9nQJlZajVjjLgkttyiQL/pXWa1pNhuEezogLESPD38rewz3poQIDAIUwe6H51VUySXaoklJEsDeBRrmg/WQQ8tc3Nh8meDo/henA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786611730; c=relaxed/simple; bh=ZwwMINBHW+RxGwt3/tsNmCqt1HXar34b1qTvwZacLfw=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=nECH6FPUA7VPVioauKlbIOohdAe/GFLoQ5PE5BKYJ/wqpMD/KPiwq7SowZ0fgGxS8HfM1ZRt1qS4ZgihLoHyWuaICYUmuDq4lEX0bsj5G4dgXe+f/6KSbjsBT+jY1wdogKP3rTEi0jVWJxoy6fwr5fAP+ZulTGb3hXRf4YrAZXw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=m/6e4+c/; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="m/6e4+c/" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 397821F00A3E; Thu, 13 Aug 2026 09:01:57 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786611724; bh=FNUYuddkzTS87Bqr3ijSBGzCz9N6C4j7LbHsu2d+TBQ=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=m/6e4+c/zGvPnNjQbgUwONVEiSwH+bp/iQjpj1MA7E41naGzpipAaNubMIxubt2uR KKusNZRHuKIKaP4Zw9vmnxT+scVXx2ALxu59GeP9XGLDwQANmqy+qvGyFlJEAOFjDr iivR3X2njcirCvJ5uWY4/eciRy5BACg2yDT/CP/cDymcfwwjaaaUiZG0xmFM3cOcED kgFf7ohl58bMf1tcfUnsSMxT9hmfAa7q24MDu3HKgQNnsF7UneMnlPU64C+AdAGsmU IUaEjJD+tWaLuAq+Em9ACzrt7m07ZaZQVmkp0lP37UM3E7DZjK8/PyAlGs1JkMB219 dRZcLMlUwuaKA== From: Mike Rapoport Date: Thu, 13 Aug 2026 12:01:26 +0300 Subject: [PATCH v2 3/5] x86/alternative: exclude text poking against change_page_attr() Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260813-cpa-fixes-v2-3-39b4ff90f91d@kernel.org> References: <20260813-cpa-fixes-v2-0-39b4ff90f91d@kernel.org> In-Reply-To: <20260813-cpa-fixes-v2-0-39b4ff90f91d@kernel.org> To: Dave Hansen Cc: Andrew Morton , Andy Lutomirski , Borislav Petkov , David CARLIER , David Hildenbrand , Ingo Molnar , Jason Gunthorpe , Jiri Slaby , Juergen Gross , Kevin Tian , Kiryl Shutsemau , "Liam R. Howlett" , Lorenzo Stoakes , Lu Baolu , Mike Rapoport , Nikunj A Dadhania , Pedro Falcato , "H. Peter Anvin" , Peter Zijlstra , Shakeel Butt , Steffen Dirkwinkel , Suren Baghdasaryan , Thomas Gleixner , Toshi Kani , Vishal Moola , Vlastimil Babka , Will Deacon , iommu@lists.linux.dev, linux-kernel@vger.kernel.org, linux-mm@kvack.org, stable@vger.kernel.org, syzbot@syzkaller.appspotmail.com, x86@kernel.org X-Mailer: b4 0.17-dev From: Pedro Falcato From time to time, the following BUG can be observed[0]: > kernel BUG at arch/x86/kernel/alternative.c:2576! > Oops: invalid opcode: 0000 [#1] SMP NOPTI > CPU: 0 UID: 0 PID: 355 Comm: (udev-worker) Not tainted 7.1.3-1-default #1= PREEMPT(full) openSUSE Tumbleweed 8c1795b03ec64f997e57a8ad38b1161e3b98da64 > Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS unknown 02/02= /2022 > RIP: 0010:__text_poke+0x2aa/0x450 > Call Trace: > > smp_text_poke_batch_finish+0x2a7/0x320 > __static_call_transform+0xb7/0x220 > arch_static_call_transform+0x5b/0xb0 > __static_call_init+0xe9/0x270 > static_call_module_notify+0x11f/0x150 > notifier_call_chain+0x61/0xe0 > blocking_notifier_call_chain_robust+0x63/0xc0 > load_module+0x1c92/0x20c0 > init_module_from_file+0xd8/0x140 > idempotent_init_module+0x100/0x2f0 > __x64_sys_finit_module+0x71/0xe0 > do_syscall_64+0xe1/0x610 > entry_SYSCALL_64_after_hwframe+0x76/0x7e which matches the following BUG_ON in alternative.c: /* * If something went wrong, crash and burn since recovery paths are not * implemented. */ BUG_ON(!pages[0] || (cross_page_boundary && !pages[1])); This can happen if vmalloc_to_page() fails, for any reason. Such can happen if text poking races with CPA, which can possibly result in the collapsing of page tables (or breaking of PMD hugepages). It is not a problem for most users of vmalloc_to_page() (they solely own the vmalloc'd range) but, when CONFIG_ARCH_HAS_EXECMEM_ROX=3Dy, various modules own a single execmem vmall= oc range, and can call set_memory_*() in parallel on it. This can happen to race against __text_poke and cause havoc in vmalloc_to_page(). Fix it by excluding against CPA using the init_mm mmap read lock. Fixes: 64f6a4e10c05 ("x86: re-enable EXECMEM_ROX support") Reported-by: Jiri Slaby Link: https://bugzilla.opensuse.org/show_bug.cgi?id=3D1271202 [0] Reported-by: Steffen Dirkwinkel Link: https://lore.kernel.org/linux-mm/555ea1d43a12c30a8f1eaf10c899b3790d72= 8f33.camel@dirkwinkel.cc/ Cc: stable@vger.kernel.org Co-developed-by: "Lorenzo Stoakes (ARM)" Signed-off-by: "Lorenzo Stoakes (ARM)" Signed-off-by: Pedro Falcato Signed-off-by: Mike Rapoport (Microsoft) Tested-by: Atish Patra Tested-by: Jiri Slaby Tested-by: Nikunj A Dadhania --- arch/x86/kernel/alternative.c | 39 ++++++++++++++++++++++++++++++++++++--- 1 file changed, 36 insertions(+), 3 deletions(-) diff --git a/arch/x86/kernel/alternative.c b/arch/x86/kernel/alternative.c index 62936a3bde19..f81d6bc90a6f 100644 --- a/arch/x86/kernel/alternative.c +++ b/arch/x86/kernel/alternative.c @@ -6,6 +6,9 @@ #include #include #include +#include +#include +#include =20 #include #include @@ -2543,6 +2546,38 @@ static void text_poke_memset(void *dst, const void *= src, size_t len) =20 typedef void text_poke_f(void *dst, const void *src, size_t len); =20 +static void __poke_vmalloc_pages(struct page **pages, void *addr, + bool cross_page_boundary) +{ + pages[0] =3D vmalloc_to_page(addr); + if (cross_page_boundary) + pages[1] =3D vmalloc_to_page(addr + PAGE_SIZE); +} + +static void poke_vmalloc_pages(struct page **pages, void *addr, + bool cross_page_boundary) +{ + if (in_dbg_master()) { + /* + * If called from kgdb cannot sleep, but all other CPUs stopped + * anyway so safe to proceed without locks + */ + __poke_vmalloc_pages(pages, addr, cross_page_boundary); + } else { + /* + * execmem ROX ranges are shared between modules and can be + * collapsed to huge PMD entries, and this collapse can happen + * concurrently with a racing set_memory_rox(). + * + * Prevent vmalloc_to_page() from racing by acquiring an + * init_mm read lock which pairs with the init_mm write lock in + * cpa_collapse_large_pages(). + */ + guard(mmap_read_lock)(&init_mm); + __poke_vmalloc_pages(pages, addr, cross_page_boundary); + } +} + static void *__text_poke(text_poke_f func, void *addr, const void *src, si= ze_t len) { bool cross_page_boundary =3D offset_in_page(addr) + len > PAGE_SIZE; @@ -2560,9 +2595,7 @@ static void *__text_poke(text_poke_f func, void *addr= , const void *src, size_t l BUG_ON(!after_bootmem); =20 if (!core_kernel_text((unsigned long)addr)) { - pages[0] =3D vmalloc_to_page(addr); - if (cross_page_boundary) - pages[1] =3D vmalloc_to_page(addr + PAGE_SIZE); + poke_vmalloc_pages(pages, addr, cross_page_boundary); } else { pages[0] =3D virt_to_page(addr); WARN_ON(!PageReserved(pages[0])); --=20 2.53.0 From nobody Wed Sep 30 05:00:31 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 35B0F3B71B6; Thu, 13 Aug 2026 09:02:14 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786611738; cv=none; b=KO8GQRp41hTGEdHebb9tEPp4KzsJauIO2TBJqHBg8YFOXNuIei5Gbmp9on+UTgL/MhuyXJpYU9ojyojgeqh3j3Gozfx7aq+bg2IybIki3JWGixkl+wZo3MdmoYCEHN8U0ivjLCZ0VyKhhxuoyO9351oDFKP943Dfm2RAUvFxiC0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786611738; c=relaxed/simple; bh=eR90IvSUViA5rAhz+GIOPWKSgTreSS8lW89G4Kl2ZFw=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=bsh53+hlPCtAV66HhHvcc8wypNIZpbOyHIXp8iUdgbrks0vKzCxkge94CkMCTEus05uDdlEsp2MkPliOutK4MgozSHYi/sgP0iooeiN4CIPKX255Tow6S7jStuLCsVC01KmRcnEdU9brNtqStWHu3z/RenMN5LYBle2yuN1NY04= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=cTHjlUqt; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="cTHjlUqt" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 450831F00A3D; Thu, 13 Aug 2026 09:02:05 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786611732; bh=X5zSwpgxBZvb+myrFl6S3DxQdJeWuTT7rdxmIVQZYGg=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=cTHjlUqtkSdbH+1I9do/NPyiH+MMWzDRET3NOH61eOVwd1VJxTrpsVOqpwVdcd3o3 1a+P8AZvdG4dTwAhdsVlXsHIPTE4Wk30hWvVnuTVs0j7UE9FzV4XjCqDWN6Dq1M2W4 YdIw8U7uzuXowdqqbZpsvhuOLRLJcugSJ8Xvi8+QDSzeE4i89uXTdfi8i5gcV3z+SE mkyJvWT6Uvehy3Nzzc4qwCLqdFIT3iHGNXe/SlE/07853x12BnoUN1Tmnss5+9MbS+ BmiNToCMRCiMdxRSmec7BspR3mYd1XzZelKof57VO9NVgp2H32tQfDvp7tfgGDMpvJ 3XNuJzUVXQKyQ== From: Mike Rapoport Date: Thu, 13 Aug 2026 12:01:27 +0300 Subject: [PATCH v2 4/5] x86/mm/pat: allocate split page tables as kernel page tables Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260813-cpa-fixes-v2-4-39b4ff90f91d@kernel.org> References: <20260813-cpa-fixes-v2-0-39b4ff90f91d@kernel.org> In-Reply-To: <20260813-cpa-fixes-v2-0-39b4ff90f91d@kernel.org> To: Dave Hansen Cc: Andrew Morton , Andy Lutomirski , Borislav Petkov , David CARLIER , David Hildenbrand , Ingo Molnar , Jason Gunthorpe , Jiri Slaby , Juergen Gross , Kevin Tian , Kiryl Shutsemau , "Liam R. Howlett" , Lorenzo Stoakes , Lu Baolu , Mike Rapoport , Nikunj A Dadhania , Pedro Falcato , "H. Peter Anvin" , Peter Zijlstra , Shakeel Butt , Steffen Dirkwinkel , Suren Baghdasaryan , Thomas Gleixner , Toshi Kani , Vishal Moola , Vlastimil Babka , Will Deacon , iommu@lists.linux.dev, linux-kernel@vger.kernel.org, linux-mm@kvack.org, stable@vger.kernel.org, syzbot@syzkaller.appspotmail.com, x86@kernel.org X-Mailer: b4 0.17-dev From: "Lorenzo Stoakes (ARM)" When splitting a large page in CPA in __split_large_page() we allocate a PTE directly without going through the standard page table allocation routines such as pte_alloc_one_kernel(). This means the page table constructor is never called nor is the page table marked as a kernel page table. The former results in the folio associated with the page table not being marked as a page table (__pagetable_ctor() is never called thus neither is __folio_set_pgtable()) nor are statistics updated to reflect it (lruvec_stat_add_folio() is never called). The latter issue of failing to mark the page table as a kernel page table (ptdesc_set_kernel() is never called) is far more problematic. Since commit 5ba2f0a15564 ("mm: introduce deferred freeing for kernel page tables") kernel page table freeing has been batched and since the subsequent commit e37d5a2d60a3 ("iommu/sva: invalidate stale IOTLB entries for kernel address space") IOTLB cache entries for kernel page tables have been invalidated upon being freed. Since split page tables are freed without this invalidation, the IOTLB can contain stale entries for them. Resolve the issue by using the ordinary PTE allocation API at split time. This results in these kernel page tables invoking a page table constructor, and thus requires a page table destructor. Since we cannot assume one is always present (early allocated direct map page tables are not marked as such), we conditionally call pagetable_dtor_free() if the PG_table folio flag for the ptdesc is set, otherwise we free the page table via pagetable_free(). Regardless of which path is taken page tables marked as kernel page tables, which now includes split page tables, take the correct route through pagetable_free_kernel(). There is a user-visible side effect in that split page tables will appear in nr_page_table_pages in /proc/vmstat (as do other kernel page tables allocated after early boot), however this is a positive change. This issue started being markedly problematic after commit 5ba2f0a15564 ("mm: introduce deferred freeing for kernel page tables") so choose this as the Fixes target. Fixes: 5ba2f0a15564 ("mm: introduce deferred freeing for kernel page tables= ") Cc: stable@vger.kernel.org Signed-off-by: Lorenzo Stoakes (ARM) Acked-by: Vishal Moola Signed-off-by: Mike Rapoport (Microsoft) Tested-by: Atish Patra Tested-by: Nikunj A Dadhania --- arch/x86/mm/pat/set_memory.c | 25 ++++++++++++++++--------- 1 file changed, 16 insertions(+), 9 deletions(-) diff --git a/arch/x86/mm/pat/set_memory.c b/arch/x86/mm/pat/set_memory.c index 2d04a4bf34aa..fbc418dfc597 100644 --- a/arch/x86/mm/pat/set_memory.c +++ b/arch/x86/mm/pat/set_memory.c @@ -441,7 +441,15 @@ static void __cpa_collapse_large_pages(struct cpa_data= *cpa) =20 list_for_each_entry_safe(ptdesc, tmp, &pgtables, pt_list) { list_del(&ptdesc->pt_list); - pagetable_free(ptdesc); + /* + * Only early alloc'd direct map should not be flagged PG_table + * here and those shouldn't be collapsed. However be abundantly + * cautious and handle the !PG_table case too. + */ + if (PageTable((ptdesc_page(ptdesc)))) + pagetable_dtor_free(ptdesc); + else + pagetable_free(ptdesc); } } =20 @@ -1134,11 +1142,10 @@ static void split_set_pte(struct cpa_data *cpa, pte= _t *pte, unsigned long pfn, =20 static int __split_large_page(struct cpa_data *cpa, pte_t *kpte, unsigned long addres= s, - struct ptdesc *ptdesc) + pte_t *pbase) { unsigned long lpaddr, lpinc, ref_pfn, pfn, pfninc =3D 1; - struct page *base =3D ptdesc_page(ptdesc); - pte_t *pbase =3D (pte_t *)page_address(base); + struct page *base =3D virt_to_page(pbase); unsigned int i, level; pgprot_t ref_prot; bool nx, rw; @@ -1238,20 +1245,20 @@ __split_large_page(struct cpa_data *cpa, pte_t *kpt= e, unsigned long address, static int split_large_page(struct cpa_data *cpa, pte_t *kpte, unsigned long address) { - struct ptdesc *ptdesc; + pte_t *pte; =20 spin_unlock(&cpa_lock); if (cpa->init_mm_read_locked) mmap_read_unlock(&init_mm); - ptdesc =3D pagetable_alloc(GFP_KERNEL, 0); + pte =3D pte_alloc_one_kernel(&init_mm); if (cpa->init_mm_read_locked) mmap_read_lock(&init_mm); spin_lock(&cpa_lock); - if (!ptdesc) + if (!pte) return -ENOMEM; =20 - if (__split_large_page(cpa, kpte, address, ptdesc)) - pagetable_free(ptdesc); + if (__split_large_page(cpa, kpte, address, pte)) + pte_free_kernel(&init_mm, pte); =20 return 0; } --=20 2.53.0 From nobody Wed Sep 30 05:00:31 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5DCA33B813C; Thu, 13 Aug 2026 09:02:22 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786611744; cv=none; b=oMm9GxUAjAnvezYeQV2tylEpM7yOnywHK/1KldJREZoxcM25ChWL3iR3fImgROMUjsVrtvlunriKh1CXdNMPmztIQH42IeVZNXxNYZwVhPDdzXFzY7o7QlxQzaWWW4rqYdQE4DFYpOgw16AWaubUS52MjICMDIjiuGfp8F+GivM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786611744; c=relaxed/simple; bh=10CIl+Hb8vhPX9MjJbLKvl8gjfb6BmZT4g3+HIsWfkM=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=dVkUGEZyrBDrdO9o8se8rywXlYgrTE2V6OEy8gZEAXJL/kGKYiwvU48+1+waNlelVO7JL3aG16pfLkg9U4slSFCRtGOEqwU+i8feFrbYRoVlvVtuzY1LnYjSlr/tG5LXZvAqCh0knARnfHAIsqmGxf0zz4XIZKuib0rNDavQAEc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=Cq9h/biY; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="Cq9h/biY" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 4B3E31F00A3A; Thu, 13 Aug 2026 09:02:13 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1786611740; bh=itFc+ei9wQ2AisPFkcdPDGQXQgvzV1zw/qg30HEZpkg=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=Cq9h/biYez7ZDggx1XDAyzOQtOqE3YYHzs46iGAd2rMJH+kJ6f4CsQzyDIQEKlHLL qqxDLE7ZdCbpGYdokBQR6rehGnNDLU/HD1bmmhBJ17QiPWej9urwhBlftexAsFFhNP PxXyIg49vQ57AlcYJtxGOs8kCJt53ADPmY/AJJpdxs3R2tElKhI2feDZARQdhtIplK AYOxBo7dxrdwdhdm8Ro90htp3Ie1zkYOjil7DSKlZ4RVfOKZOPNFIQPSft63kM5oMz 7kpA6swRbGH4Z7d8HV8uBp67PYWfhkhbp5LPncNJO09xCKi6mIJNbMm6VDrVcuHja6 QWXfeu4J7BtBg== From: "Mike Rapoport (Microsoft)" Date: Thu, 13 Aug 2026 12:01:28 +0300 Subject: [PATCH v2 5/5] x86/mm/pat: fix effective RW computation in lookup_address_in_pgd_attr() Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260813-cpa-fixes-v2-5-39b4ff90f91d@kernel.org> References: <20260813-cpa-fixes-v2-0-39b4ff90f91d@kernel.org> In-Reply-To: <20260813-cpa-fixes-v2-0-39b4ff90f91d@kernel.org> To: Dave Hansen Cc: Andrew Morton , Andy Lutomirski , Borislav Petkov , David CARLIER , David Hildenbrand , Ingo Molnar , Jason Gunthorpe , Jiri Slaby , Juergen Gross , Kevin Tian , Kiryl Shutsemau , "Liam R. Howlett" , Lorenzo Stoakes , Lu Baolu , Mike Rapoport , Nikunj A Dadhania , Pedro Falcato , "H. Peter Anvin" , Peter Zijlstra , Shakeel Butt , Steffen Dirkwinkel , Suren Baghdasaryan , Thomas Gleixner , Toshi Kani , Vishal Moola , Vlastimil Babka , Will Deacon , iommu@lists.linux.dev, linux-kernel@vger.kernel.org, linux-mm@kvack.org, stable@vger.kernel.org, syzbot@syzkaller.appspotmail.com, x86@kernel.org X-Mailer: b4 0.17-dev lookup_address_in_pgd_attr() accumulates the effective NX and RW bits of the walked page table levels so that verify_rwx() can detect mappings that are both writable and executable. The RW bits are folded into a bool with rw &=3D pXd_flags(*pXd) & _PAGE_RW; but _PAGE_RW is 0x2. So consider the accumulation line: rw &=3D pXd_flags(*pXd) & _PAGE_RW; where rw=3D0x1 and the right side evaluates down to 0x2. It'll end up doing: rw =3D 0x1 & 0x2 and rw always ends up 0. This way rw becomes false at the first level walked, regardless of the actual permissions, and verify_rwx() treats every mapping as non-writable and never reports a W^X violation. Add double negation to the right side to normalize the _PAGE_RW flag to 0 or 1. Fixes: ceb647b4b529 ("x86/pat: Introduce lookup_address_in_pgd_attr()") Cc: stable@vger.kernel.org Assisted-by: Copilot:claude-opus-4.8 Reviewed-by: Juergen Gross Tested-by: syzbot@syzkaller.appspotmail.com Signed-off-by: Mike Rapoport (Microsoft) Reviewed-by: Lorenzo Stoakes (ARM) Tested-by: Atish Patra Tested-by: Nikunj A Dadhania --- arch/x86/mm/pat/set_memory.c | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/arch/x86/mm/pat/set_memory.c b/arch/x86/mm/pat/set_memory.c index fbc418dfc597..430d0b448371 100644 --- a/arch/x86/mm/pat/set_memory.c +++ b/arch/x86/mm/pat/set_memory.c @@ -754,7 +754,7 @@ pte_t *lookup_address_in_pgd_attr(pgd_t *pgd, unsigned = long address, =20 *level =3D PG_LEVEL_512G; *nx |=3D pgd_flags(*pgd) & _PAGE_NX; - *rw &=3D pgd_flags(*pgd) & _PAGE_RW; + *rw &=3D !!(pgd_flags(*pgd) & _PAGE_RW); =20 p4d =3D p4d_offset(pgd, address); if (p4d_none(*p4d)) @@ -765,7 +765,7 @@ pte_t *lookup_address_in_pgd_attr(pgd_t *pgd, unsigned = long address, =20 *level =3D PG_LEVEL_1G; *nx |=3D p4d_flags(*p4d) & _PAGE_NX; - *rw &=3D p4d_flags(*p4d) & _PAGE_RW; + *rw &=3D !!(p4d_flags(*p4d) & _PAGE_RW); =20 pud =3D pud_offset(p4d, address); if (pud_none(*pud)) @@ -776,7 +776,7 @@ pte_t *lookup_address_in_pgd_attr(pgd_t *pgd, unsigned = long address, =20 *level =3D PG_LEVEL_2M; *nx |=3D pud_flags(*pud) & _PAGE_NX; - *rw &=3D pud_flags(*pud) & _PAGE_RW; + *rw &=3D !!(pud_flags(*pud) & _PAGE_RW); =20 pmd =3D pmd_offset(pud, address); if (pmd_none(*pmd)) @@ -787,7 +787,7 @@ pte_t *lookup_address_in_pgd_attr(pgd_t *pgd, unsigned = long address, =20 *level =3D PG_LEVEL_4K; *nx |=3D pmd_flags(*pmd) & _PAGE_NX; - *rw &=3D pmd_flags(*pmd) & _PAGE_RW; + *rw &=3D !!(pmd_flags(*pmd) & _PAGE_RW); =20 return pte_offset_kernel(pmd, address); } --=20 2.53.0