From nobody Sat Oct 3 04:50:54 2026 Received: from galois.linutronix.de (Galois.linutronix.de [193.142.43.55]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E9C643E2771; Wed, 5 Aug 2026 12:26:50 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=193.142.43.55 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785932812; cv=none; b=uQpZf5FIS/xZXPmzIlbRMagh+OkDOnVhOMOKNZQk/Z0eYNonhdA+gJyjIHE7FkoMbNe30t8q7niPrKSVOHi5R3sYm/TwzTJJYph4K8aL49Cjvju6ZYSW2M6Sp1vc2rBUqyhi88ZGzFVVN54/XOMCNw+S1NjlESMw7fCUMGPEmFs= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785932812; c=relaxed/simple; bh=ks9Hg1Qv0HtHexOOOXP3WXwqX2peK8zefFOkapPfhP8=; h=Date:From:To:Subject:Cc:In-Reply-To:References:MIME-Version: Message-ID:Content-Type; b=T237rBGMHloxm0QMamvX+slU9V2wWRj9MLfk7emidAmW/wnuogHzkt9hIdpI3xVuxIc12dEtN4yATTY6mGiZ8sx4kCfsLHVxYOCPZdQ62T5M4OerGkLxs6mDmh9ix1y3DQfmLiUkC7I0NCfvXZfCPHAOSf5JfETbxGOlWI35KJo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linutronix.de; spf=pass smtp.mailfrom=linutronix.de; dkim=pass (2048-bit key) header.d=linutronix.de header.i=@linutronix.de header.b=e7CnIZyW; dkim=permerror (0-bit key) header.d=linutronix.de header.i=@linutronix.de header.b=tnM2Luq0; arc=none smtp.client-ip=193.142.43.55 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linutronix.de Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linutronix.de Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=linutronix.de header.i=@linutronix.de header.b="e7CnIZyW"; dkim=permerror (0-bit key) header.d=linutronix.de header.i=@linutronix.de header.b="tnM2Luq0" Date: Wed, 05 Aug 2026 12:26:46 -0000 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linutronix.de; s=2020; t=1785932808; h=from:from:sender:sender:reply-to:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=6JObeyWwWOrFbteKsWTi8xdRGfNRn63ZsxeQAuaIpC4=; b=e7CnIZyW1JCYPbQhQYutBrlfSnYQbGIshnKT63hrdJmsXJpZFt4T9tj94a6KxtnBrKL811 sEjOaeuWuRX+z8w66CCiujwY+8hqefaDKq4FWSWrlfu6PCd2OuoMSGFCjO3evoyykVDsoj GjpkJZ619G1QiEVrYciIgRH6xCn7ufn2xzFfOP6wgTHvsv3lML3ypi4NIB/6nMXIW49y8v ZviJoE8V26mChBVb0o/t/bIa2GnJ4itNatXkKJIpL/YKzF5JfpD6wq2IJYt3LXwGJ00TCZ wsj7/o2iNof6waFPKCbOL6sQhF3J256M5ApwsmObIdvmfzMPTNqn8ziSD80DXA== DKIM-Signature: v=1; a=ed25519-sha256; c=relaxed/relaxed; d=linutronix.de; s=2020e; t=1785932808; h=from:from:sender:sender:reply-to:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=6JObeyWwWOrFbteKsWTi8xdRGfNRn63ZsxeQAuaIpC4=; b=tnM2Luq0df2vsjqm+IEMzhBAirh8r3rQUHfRr62i7hCEvHd6kUsSm35ylKXm8zQwN7N7LQ 15VDTTtygaV+EHAg== From: "tip-bot2 for Peter Zijlstra" Sender: tip-bot2@linutronix.de Reply-to: linux-kernel@vger.kernel.org To: linux-tip-commits@vger.kernel.org Subject: [tip: x86/mm] x86/mm: Fix and document DEBUG_PAGEALLOC Cc: "Peter Zijlstra (Intel)" , x86@kernel.org, linux-kernel@vger.kernel.org In-Reply-To: <20260729111119.604452135@infradead.org> References: <20260729111119.604452135@infradead.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Message-ID: <178593280656.708.5338231791833080639.tip-bot2@tip-bot2> Robot-ID: Robot-Unsubscribe: Contact to get blacklisted from these emails Precedence: bulk Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable The following commit has been merged into the x86/mm branch of tip: Commit-ID: 7da514d819a0afb148634aac92b3d190f34947c3 Gitweb: https://git.kernel.org/tip/7da514d819a0afb148634aac92b3d190f= 34947c3 Author: Peter Zijlstra AuthorDate: Wed, 29 Jul 2026 13:08:10 +02:00 Committer: Peter Zijlstra CommitterDate: Wed, 05 Aug 2026 14:19:35 +02:00 x86/mm: Fix and document DEBUG_PAGEALLOC It turns out that commit 5fce67641a3e ("x86/mm/pat: Don't gate cpa_lock on debug_pagealloc_enabled()") was a little too quick to remove the debug_pagealloc exception for cpa_lock. Notably __kernel_map_pages() is used by the page-allocator from any context the page-allocator itself is used, which violates the cpa_lock rules. Re-instate the exception, except make it specific to the __kernel_map_pages() such that any other cpa() usage is still fully serialized by cpa_lock. Also note that since cpa() should not be used on memory that isn't allocated, the page-allocator locking and cpa are infact mutually exclusive and all cpa usage in fully serialized. Add a comment explaining this and other 'funnies' surrounding DEBUG_PAGEALLOC, including how pgd_lock is not affected and the TLB trickery. Fixes: 5fce67641a3e ("x86/mm/pat: Don't gate cpa_lock on debug_pagealloc_en= abled()") Signed-off-by: Peter Zijlstra (Intel) Link: https://patch.msgid.link/20260729111119.604452135@infradead.org --- arch/x86/mm/pat/set_memory.c | 80 +++++++++++++++++++++++++---------- 1 file changed, 58 insertions(+), 22 deletions(-) diff --git a/arch/x86/mm/pat/set_memory.c b/arch/x86/mm/pat/set_memory.c index ad52c9d..d8d057f 100644 --- a/arch/x86/mm/pat/set_memory.c +++ b/arch/x86/mm/pat/set_memory.c @@ -68,11 +68,12 @@ static const int cpa_warn_level =3D CPA_PROTECT; */ static DEFINE_SPINLOCK(cpa_lock); =20 -#define CPA_FLUSHTLB 1 -#define CPA_ARRAY 2 -#define CPA_PAGES_ARRAY 4 -#define CPA_NO_CHECK_ALIAS 8 /* Do not search for aliases */ -#define CPA_COLLAPSE 16 /* try to collapse large pages */ +#define CPA_FLUSHTLB 0x01 +#define CPA_ARRAY 0x02 +#define CPA_PAGES_ARRAY 0x04 +#define CPA_NO_CHECK_ALIAS 0x08 /* Do not search for aliases */ +#define CPA_COLLAPSE 0x10 /* try to collapse large pages */ +#define CPA_DEBUG_PAGEALLOC 0x20 =20 static inline pgprot_t cachemode2pgprot(enum page_cache_mode pcm) { @@ -1990,6 +1991,7 @@ static int __change_page_attr_set_clr(struct cpa_data= *cpa, int primary) { unsigned long numpages =3D cpa->numpages; unsigned long rempages =3D numpages; + bool lock =3D true; int ret =3D 0; =20 /* @@ -1999,6 +2001,29 @@ static int __change_page_attr_set_clr(struct cpa_dat= a *cpa, int primary) !cpa->force_split) return ret; =20 + /* + * DEBUG_PAGEALLOC is special; it is called from any context the + * page-allocator is, which violates the normal cpa_lock locking + * rules. + * + * However, since it is part of the page-allocator, things are still + * properly serialized by the page-allocator locking and the fact that + * when a page is owned by the page-allocator, it isn't owned by + * anybody else. That is, you *SHOULD NOT* be calling cpa() on memory + * that isn't allocated. + * + * Additionally, DEBUG_PAGEALLOC ensures (per probe_page_size_mask()) + * that the kernel mapping is 4k pages, therefore there are no large + * pages to split/collapse. + * + * Furthermore, the page-allocator strictly manages pages that + * *exist*, avoiding pgd_lock. + * + * Therefore, it is safe to not take cpa_lock. + */ + if (debug_pagealloc_enabled() && (cpa->flags & CPA_DEBUG_PAGEALLOC)) + lock =3D false; + while (rempages) { /* * Store the remaining nr of pages for the large page @@ -2009,9 +2034,12 @@ static int __change_page_attr_set_clr(struct cpa_dat= a *cpa, int primary) if (cpa->flags & (CPA_ARRAY | CPA_PAGES_ARRAY)) cpa->numpages =3D 1; =20 - spin_lock(&cpa_lock); - ret =3D __change_page_attr(cpa, primary); - spin_unlock(&cpa_lock); + if (lock) { + guard(spinlock)(&cpa_lock); + ret =3D __change_page_attr(cpa, primary); + } else { + ret =3D __change_page_attr(cpa, primary); + } if (ret) goto out; =20 @@ -2590,7 +2618,7 @@ int set_pages_rw(struct page *page, int numpages) return set_memory_rw(addr, numpages); } =20 -static int __set_pages_p(struct page *page, int numpages) +static int __set_pages_p(struct page *page, int numpages, unsigned int cpa= _flags) { unsigned long tempaddr =3D (unsigned long) page_address(page); struct cpa_data cpa =3D { .vaddr =3D &tempaddr, @@ -2598,7 +2626,7 @@ static int __set_pages_p(struct page *page, int numpa= ges) .numpages =3D numpages, .mask_set =3D __pgprot(_PAGE_PRESENT | _PAGE_RW), .mask_clr =3D __pgprot(0), - .flags =3D CPA_NO_CHECK_ALIAS }; + .flags =3D CPA_NO_CHECK_ALIAS | cpa_flags }; =20 /* * No alias checking needed for setting present flag. otherwise, @@ -2609,7 +2637,7 @@ static int __set_pages_p(struct page *page, int numpa= ges) return __change_page_attr_set_clr(&cpa, 1); } =20 -static int __set_pages_np(struct page *page, int numpages) +static int __set_pages_np(struct page *page, int numpages, unsigned int cp= a_flags) { unsigned long tempaddr =3D (unsigned long) page_address(page); struct cpa_data cpa =3D { .vaddr =3D &tempaddr, @@ -2617,7 +2645,7 @@ static int __set_pages_np(struct page *page, int nump= ages) .numpages =3D numpages, .mask_set =3D __pgprot(0), .mask_clr =3D __pgprot(_PAGE_PRESENT | _PAGE_RW | _PAGE_DIRTY), - .flags =3D CPA_NO_CHECK_ALIAS }; + .flags =3D CPA_NO_CHECK_ALIAS | cpa_flags }; =20 /* * No alias checking needed for setting not present flag. otherwise, @@ -2630,20 +2658,20 @@ static int __set_pages_np(struct page *page, int nu= mpages) =20 int set_direct_map_invalid_noflush(struct page *page) { - return __set_pages_np(page, 1); + return __set_pages_np(page, 1, 0); } =20 int set_direct_map_default_noflush(struct page *page) { - return __set_pages_p(page, 1); + return __set_pages_p(page, 1, 0); } =20 int set_direct_map_valid_noflush(struct page *page, unsigned nr, bool vali= d) { if (valid) - return __set_pages_p(page, nr); + return __set_pages_p(page, nr, 0); =20 - return __set_pages_np(page, nr); + return __set_pages_np(page, nr, 0); } =20 #ifdef CONFIG_DEBUG_PAGEALLOC @@ -2662,15 +2690,23 @@ void __kernel_map_pages(struct page *page, int nump= ages, int enable) * and hence no memory allocations during large page split. */ if (enable) - __set_pages_p(page, numpages); + __set_pages_p(page, numpages, CPA_DEBUG_PAGEALLOC); else - __set_pages_np(page, numpages); + __set_pages_np(page, numpages, CPA_DEBUG_PAGEALLOC); =20 /* - * We should perform an IPI and flush all tlbs, - * but that can deadlock->flush only current cpu. - * Preemption needs to be disabled around __flush_tlb_all() due to - * CR3 reload in __native_flush_tlb(). + * We should perform an IPI and flush all tlbs, but that can + * deadlock, settle for a local flush. + * + * Not doing a global TLB flush means that remote CPUs will retain + * stale TLB entries. In case of P->NP (on free) this means the remote + * CPUs will not take the faults, making the debug scheme less + * reliable. On the NP->P (on alloc) this means the remote CPUs can + * take a spurious fault. However spurious_kernel_fault() will observe + * *_present() and fix it up. + * + * Preemption needs to be disabled around __flush_tlb_all() due to CR3 + * reload in __native_flush_tlb(). */ preempt_disable(); __flush_tlb_all();