[PATCH v2] x86/mm/pat: take cpa_lock around large-page collapse

Denis V. Lunev posted 1 patch 1 week, 2 days ago
arch/x86/mm/pat/set_memory.c | 8 +++++++-
1 file changed, 7 insertions(+), 1 deletion(-)
[PATCH v2] x86/mm/pat: take cpa_lock around large-page collapse
Posted by Denis V. Lunev 1 week, 2 days ago
Loading and unloading modules concurrently on several CPUs on a KASAN
build, with a short delay injected at the CPA page-table lookup to
widen the window, faults within minutes:

  BUG: KASAN: use-after-free in __change_page_attr+0x7cc/0x7e0
  Write of size 8 at addr ffff888181139718 by task modprobe
  ...
  The buggy address belongs to the physical page:
   pfn:0x181139 ... page_type: f2(table)

cpa_collapse_large_pages() rebuilds a leaf PMD from its 4K PTEs and
frees the old PTE-table pages, while __change_page_attr() fetches a
PTE pointer from a lockless lookup_address_in_pgd_attr() and writes
it with set_pte_atomic() only later. When module text is served from
a shared large ROX mapping the two run on the same PMD:

  CPU A (module load)              CPU B (module finalize)
  -------------------              -----------------------
  execmem_make_temp_rw
   set_memory_nx
    __change_page_attr
     split 2M -> 4K table P
     kpte = &P[i]  (lockless)
                                   execmem_restore_rox
                                    set_memory_rox (CPA_COLLAPSE)
                                     cpa_collapse_large_pages
                                      rebuild leaf PMD
                                      flush_tlb_all
                                      pagetable_free(P)
     set_pte_atomic(kpte, ...)
       -> writes into freed P

P is a page-table page (page_type: table), reused at once, so the
write corrupts whatever got the page next: a bad-pte or bad-page
splat, or a fatal fault once P has been turned into read-only text.

The flush_tlb_all() before the free does not close this: its IPI only
serializes against page-table walkers that run with interrupts off
(e.g. GUP-fast); the walk in __change_page_attr() runs with interrupts
on, so nothing stops it from holding a stale pointer into P.

Serialize the collapse - the PMD rebuild, TLB flush and PTE-table
free - under cpa_lock, the same lock __change_page_attr() now takes
unconditionally since commit ("x86/mm/pat: stop gating cpa_lock on
debug_pagealloc_enabled()"), so a concurrent walker can no longer
hold a pointer into a table the collapse is about to free.

Fixes: 41d88484c71c ("x86/mm/pat: restore large ROX pages after fragmentation")
Signed-off-by: Denis V. Lunev <den@openvz.org>
Acked-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
---
v2:
- drop the debug_pagealloc_enabled() skip and its comment: now that
  __change_page_attr() takes cpa_lock unconditionally, the skip is no
  longer needed to avoid locking only one side of the race

 arch/x86/mm/pat/set_memory.c | 8 +++++++-
 1 file changed, 7 insertions(+), 1 deletion(-)

diff --git a/arch/x86/mm/pat/set_memory.c b/arch/x86/mm/pat/set_memory.c
index e8316f5ffa8a..6dda81a629d6 100644
--- a/arch/x86/mm/pat/set_memory.c
+++ b/arch/x86/mm/pat/set_memory.c
@@ -417,6 +417,8 @@ static void cpa_collapse_large_pages(struct cpa_data *cpa)
 	int collapsed = 0;
 	int i;
 
+	spin_lock(&cpa_lock);
+
 	if (cpa->flags & (CPA_PAGES_ARRAY | CPA_ARRAY)) {
 		for (i = 0; i < cpa->numpages; i++)
 			collapsed += collapse_large_pages(__cpa_addr(cpa, i),
@@ -430,8 +432,10 @@ static void cpa_collapse_large_pages(struct cpa_data *cpa)
 			collapsed += collapse_large_pages(addr, &pgtables);
 	}
 
-	if (!collapsed)
+	if (!collapsed) {
+		spin_unlock(&cpa_lock);
 		return;
+	}
 
 	flush_tlb_all();
 
@@ -439,6 +443,8 @@ static void cpa_collapse_large_pages(struct cpa_data *cpa)
 		list_del(&ptdesc->pt_list);
 		pagetable_free(ptdesc);
 	}
+
+	spin_unlock(&cpa_lock);
 }
 
 static void cpa_flush(struct cpa_data *cpa, int cache)

base-commit: 4a0e2d0aa6a44155f9bf0289be3dd94024cee3b4
-- 
2.53.0
Re: [PATCH v2] x86/mm/pat: take cpa_lock around large-page collapse
Posted by Lorenzo Stoakes (ARM) 4 days ago
On Wed, Jul 15, 2026 at 08:34:52PM +0200, Denis V. Lunev wrote:
> Loading and unloading modules concurrently on several CPUs on a KASAN
> build, with a short delay injected at the CPA page-table lookup to
> widen the window, faults within minutes:
>
>   BUG: KASAN: use-after-free in __change_page_attr+0x7cc/0x7e0
>   Write of size 8 at addr ffff888181139718 by task modprobe
>   ...
>   The buggy address belongs to the physical page:
>    pfn:0x181139 ... page_type: f2(table)
>
> cpa_collapse_large_pages() rebuilds a leaf PMD from its 4K PTEs and
> frees the old PTE-table pages, while __change_page_attr() fetches a
> PTE pointer from a lockless lookup_address_in_pgd_attr() and writes
> it with set_pte_atomic() only later. When module text is served from
> a shared large ROX mapping the two run on the same PMD:
>
>   CPU A (module load)              CPU B (module finalize)
>   -------------------              -----------------------
>   execmem_make_temp_rw
>    set_memory_nx
>     __change_page_attr
>      split 2M -> 4K table P
>      kpte = &P[i]  (lockless)
>                                    execmem_restore_rox
>                                     set_memory_rox (CPA_COLLAPSE)
>                                      cpa_collapse_large_pages
>                                       rebuild leaf PMD
>                                       flush_tlb_all
>                                       pagetable_free(P)
>      set_pte_atomic(kpte, ...)
>        -> writes into freed P

So now this is resolved via:

	init_mm read lock		init_mm write lock

>
> P is a page-table page (page_type: table), reused at once, so the
> write corrupts whatever got the page next: a bad-pte or bad-page
> splat, or a fatal fault once P has been turned into read-only text.
>
> The flush_tlb_all() before the free does not close this: its IPI only
> serializes against page-table walkers that run with interrupts off
> (e.g. GUP-fast); the walk in __change_page_attr() runs with interrupts
> on, so nothing stops it from holding a stale pointer into P.
>
> Serialize the collapse - the PMD rebuild, TLB flush and PTE-table
> free - under cpa_lock, the same lock __change_page_attr() now takes
> unconditionally since commit ("x86/mm/pat: stop gating cpa_lock on
> debug_pagealloc_enabled()"), so a concurrent walker can no longer
> hold a pointer into a table the collapse is about to free.

So this isn't quite correct any more; the init_mm locks fix this race in my
series ([0]), annotated above.

I feel terrible to have raced with your patch, and the commit message is
gloriously well-written and the patch is really good, it's just unfortunate
that ptdump ALSO races and can't use cpa_lock.

And my series ended up growing into a monstrous rabbit hole in general :)

However importantly, this patch does not conflict with mine, as the latest
revision holds the init_mm write lock over what was
cpa_collapse_large_pages() and what is now __cpa_collapse_large_pages() so
the spin lock is nested within the rwsem.

(There'll be a small merge conflict there, easily resolved.)

But, I wonder whether it's achieving much at this stage? But I think there
might be other races here potentially with lockless walkers, which are
probably worth looking at.

But certainly perhaps the commit message should be altered to reflect that.

>
> Fixes: 41d88484c71c ("x86/mm/pat: restore large ROX pages after fragmentation")
> Signed-off-by: Denis V. Lunev <den@openvz.org>
> Acked-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
> ---
> v2:
> - drop the debug_pagealloc_enabled() skip and its comment: now that
>   __change_page_attr() takes cpa_lock unconditionally, the skip is no
>   longer needed to avoid locking only one side of the race

So there's also an unfortunate issues as a result of doing this:

As commented over there, Mike really needs to convert the locks in his
patch ([1]) to be irq-safe since debug pagealloc wonderfully calls
__change_page_attr_set_clr() from potentially-atomic context.

(And obviously then your spin locks should also be irq-safe)

But with all of the spin locks converted, you potentially deadlock, since
you're holding cpa_lock across flush_tlb_all() which can end up doing an
IPI:


CPU 0
-----------------------------------|----------------------------------------
< GFP_ATOMIC context >		   |
__free_pages_prepare()             | set_memory_rox()
-> debug_pagealloc_unmap_pages()   | -> ...
-> __kernel_map_pages()            | -> __cpa_collapse_large_pages()
-> __change_page_attr_set_clr()    | -> cpa_lock ACQUIRED IRQs off
-> cpa_lock spins [irqs off]       | -> IPI every CPU...

<DEADLOCK wait forever>

BUT I think you can fix that by just reinstating from v1 the 'give up on
collapse if debug_pagealloc_enabled()' thing, as that's the only way the
locks can be held by anything in irq context.

Then perhaps it's actually valid to use spin_[un]lock() here also and we
need only update those code paths accessible with debug page alloc enabled?

>
>  arch/x86/mm/pat/set_memory.c | 8 +++++++-
>  1 file changed, 7 insertions(+), 1 deletion(-)
>
> diff --git a/arch/x86/mm/pat/set_memory.c b/arch/x86/mm/pat/set_memory.c
> index e8316f5ffa8a..6dda81a629d6 100644
> --- a/arch/x86/mm/pat/set_memory.c
> +++ b/arch/x86/mm/pat/set_memory.c
> @@ -417,6 +417,8 @@ static void cpa_collapse_large_pages(struct cpa_data *cpa)
>  	int collapsed = 0;
>  	int i;
>
> +	spin_lock(&cpa_lock);
> +
>  	if (cpa->flags & (CPA_PAGES_ARRAY | CPA_ARRAY)) {
>  		for (i = 0; i < cpa->numpages; i++)
>  			collapsed += collapse_large_pages(__cpa_addr(cpa, i),
> @@ -430,8 +432,10 @@ static void cpa_collapse_large_pages(struct cpa_data *cpa)
>  			collapsed += collapse_large_pages(addr, &pgtables);
>  	}
>
> -	if (!collapsed)
> +	if (!collapsed) {
> +		spin_unlock(&cpa_lock);
>  		return;
> +	}
>
>  	flush_tlb_all();
>
> @@ -439,6 +443,8 @@ static void cpa_collapse_large_pages(struct cpa_data *cpa)
>  		list_del(&ptdesc->pt_list);
>  		pagetable_free(ptdesc);
>  	}
> +
> +	spin_unlock(&cpa_lock);


>  }
>
>  static void cpa_flush(struct cpa_data *cpa, int cache)
>
> base-commit: 4a0e2d0aa6a44155f9bf0289be3dd94024cee3b4
> --
> 2.53.0
>
>

Thanks, Lorenzo

[0]:https://lore.kernel.org/linux-mm/20260716-series-vmap-race-fix-v4-0-8c108c4317df@kernel.org/
[1]:https://lore.kernel.org/all/20260715144519.934289-1-rppt@kernel.org/
Re: [PATCH v2] x86/mm/pat: take cpa_lock around large-page collapse
Posted by Lorenzo Stoakes (ARM) 3 days, 22 hours ago
On Tue, Jul 21, 2026 at 05:07:05PM +0100, Lorenzo Stoakes (ARM) wrote:
> [0]:https://lore.kernel.org/linux-mm/20260716-series-vmap-race-fix-v4-0-8c108c4317df@kernel.org/

Oops should be
https://lore.kernel.org/linux-mm/20260717-series-vmap-race-fix-v5-0-606a0ac6d3e5@kernel.org

Losing track of my own series :))

Cheers, Lorenzo
Re: [PATCH v2] x86/mm/pat: take cpa_lock around large-page collapse
Posted by Mike Rapoport 1 week, 2 days ago
On Wed, Jul 15, 2026 at 08:34:52PM +0200, Denis V. Lunev wrote:
> Loading and unloading modules concurrently on several CPUs on a KASAN
> build, with a short delay injected at the CPA page-table lookup to
> widen the window, faults within minutes:
> 
>   BUG: KASAN: use-after-free in __change_page_attr+0x7cc/0x7e0
>   Write of size 8 at addr ffff888181139718 by task modprobe
>   ...
>   The buggy address belongs to the physical page:
>    pfn:0x181139 ... page_type: f2(table)

...

> Serialize the collapse - the PMD rebuild, TLB flush and PTE-table
> free - under cpa_lock, the same lock __change_page_attr() now takes
> unconditionally since commit ("x86/mm/pat: stop gating cpa_lock on
> debug_pagealloc_enabled()"), so a concurrent walker can no longer
> hold a pointer into a table the collapse is about to free.
> 
> Fixes: 41d88484c71c ("x86/mm/pat: restore large ROX pages after fragmentation")
> Signed-off-by: Denis V. Lunev <den@openvz.org>
> Acked-by: Kiryl Shutsemau (Meta) <kas@kernel.org>

I though I acked v1, but apparently I didn't :/

Acked-by: Mike Rapoport (Microsoft) <rppt@kernel.org>

-- 
Sincerely yours,
Mike.
[tip: x86/mm] x86/mm/pat: Take cpa_lock around large-page collapse
Posted by tip-bot2 for Denis V. Lunev 1 week, 2 days ago
The following commit has been merged into the x86/mm branch of tip:

Commit-ID:     1aac65f3e651334259ecb2a5f5ddb81c01f02599
Gitweb:        https://git.kernel.org/tip/1aac65f3e651334259ecb2a5f5ddb81c01f02599
Author:        Denis V. Lunev <den@openvz.org>
AuthorDate:    Wed, 15 Jul 2026 20:34:52 +02:00
Committer:     Dave Hansen <dave.hansen@linux.intel.com>
CommitterDate: Wed, 15 Jul 2026 13:06:55 -07:00

x86/mm/pat: Take cpa_lock around large-page collapse

Loading and unloading modules concurrently on several CPUs on a KASAN
build, with a short delay injected at the CPA page-table lookup to
widen the window, faults within minutes:

  BUG: KASAN: use-after-free in __change_page_attr+0x7cc/0x7e0
  Write of size 8 at addr ffff888181139718 by task modprobe
  ...
  The buggy address belongs to the physical page:
   pfn:0x181139 ... page_type: f2(table)

cpa_collapse_large_pages() rebuilds a leaf PMD from its 4K PTEs and
frees the old PTE-table pages, while __change_page_attr() fetches a
PTE pointer from a lockless lookup_address_in_pgd_attr() and writes
it with set_pte_atomic() only later. When module text is served from
a shared large ROX mapping the two run on the same PMD:

  CPU A (module load)              CPU B (module finalize)
  -------------------              -----------------------
  execmem_make_temp_rw
   set_memory_nx
    __change_page_attr
     split 2M -> 4K table P
     kpte = &P[i]  (lockless)
                                   execmem_restore_rox
                                    set_memory_rox (CPA_COLLAPSE)
                                     cpa_collapse_large_pages
                                      rebuild leaf PMD
                                      flush_tlb_all
                                      pagetable_free(P)
     set_pte_atomic(kpte, ...)
       -> writes into freed P

P is a page-table page (page_type: table), reused at once, so the
write corrupts whatever got the page next: a bad-pte or bad-page
splat, or a fatal fault once P has been turned into read-only text.

The flush_tlb_all() before the free does not close this: its IPI only
serializes against page-table walkers that run with interrupts off
(e.g. GUP-fast); the walk in __change_page_attr() runs with interrupts
on, so nothing stops it from holding a stale pointer into P.

Serialize the collapse - the PMD rebuild, TLB flush and PTE-table
free - under cpa_lock, the same lock __change_page_attr() now takes
unconditionally since commit ("x86/mm/pat: stop gating cpa_lock on
debug_pagealloc_enabled()"), so a concurrent walker can no longer
hold a pointer into a table the collapse is about to free.

Fixes: 41d88484c71c ("x86/mm/pat: restore large ROX pages after fragmentation")
Signed-off-by: Denis V. Lunev <den@openvz.org>
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Acked-by: Kiryl Shutsemau (Meta) <kas@kernel.org>
Link: https://patch.msgid.link/20260715183453.2381141-1-den@openvz.org
---
 arch/x86/mm/pat/set_memory.c | 8 +++++++-
 1 file changed, 7 insertions(+), 1 deletion(-)

diff --git a/arch/x86/mm/pat/set_memory.c b/arch/x86/mm/pat/set_memory.c
index e9b4083..b1e780a 100644
--- a/arch/x86/mm/pat/set_memory.c
+++ b/arch/x86/mm/pat/set_memory.c
@@ -417,6 +417,8 @@ static void cpa_collapse_large_pages(struct cpa_data *cpa)
 	int collapsed = 0;
 	int i;
 
+	spin_lock(&cpa_lock);
+
 	if (cpa->flags & (CPA_PAGES_ARRAY | CPA_ARRAY)) {
 		for (i = 0; i < cpa->numpages; i++)
 			collapsed += collapse_large_pages(__cpa_addr(cpa, i),
@@ -430,8 +432,10 @@ static void cpa_collapse_large_pages(struct cpa_data *cpa)
 			collapsed += collapse_large_pages(addr, &pgtables);
 	}
 
-	if (!collapsed)
+	if (!collapsed) {
+		spin_unlock(&cpa_lock);
 		return;
+	}
 
 	flush_tlb_all();
 
@@ -439,6 +443,8 @@ static void cpa_collapse_large_pages(struct cpa_data *cpa)
 		list_del(&ptdesc->pt_list);
 		pagetable_free(ptdesc);
 	}
+
+	spin_unlock(&cpa_lock);
 }
 
 static void cpa_flush(struct cpa_data *cpa, int cache)