From nobody Mon Sep 28 19:34:00 2026 Received: from mta1.migadu.com (out-141.mta1.migadu.com [95.215.58.141]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D2AF045FFA2 for ; Tue, 18 Aug 2026 13:12:15 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=95.215.58.141 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787058737; cv=none; b=aOkf5rSorIVL6noH2YH2bXJ1bGhLHg9mpi4WULWHnpja7tzzyawBFFuc2nhXkBS5E5vFnG2MRQuRxFYBthQYNzXLLHISZQr+SdWqSDPwEjiVcziQHg+b+i8kQa4IPqYEiuUGQIK3w+BccXNus6gPwe5g9Nwmu/fhj6hkKNDBwFU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787058737; c=relaxed/simple; bh=z3Nuo7rdKfAlilt7Qu0TMklie3/gOO8sgolPIL27pOs=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=pklNnv2jcOc0gofcHXfXUKYZ4GC4X6OEiGe5AFSWApfBt3YnTjTRZa5LaYytktRVugaNSJo8QzNd/RWttHnkrBEXDtMEWA4TFSmgC4XhpxyeQhxsy953XxZfck2qf1xtz5S+Se471BKzeTj6z+D5p6CvEmdnsMI2YNpqYLhGteE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=AJwMRauu; arc=none smtp.client-ip=95.215.58.141 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="AJwMRauu" X-Envelope-To: linux-kernel@vger.kernel.org DKIM-Signature: a=rsa-sha256; bh=z3Nuo7rdKfAlilt7Qu0TMklie3/gOO8sgolPIL27pOs=; c=simple/simple; d=linux.dev; h=from:to:subject:date:message-id:mime-version:content-type; s=key1; t=1787058733; v=1; x=1787663533; b=AJwMRauuRm37dhI27xz8Gcs6xoET4jfC2mUjAq85bfG2PC1u/mcEQJxO9BIP5Q+gWFjkaxEW SvZaPlkmLsuxrZJpAcaF9D3qTAm2NFC1Xzl1Mnt76jkN1h3wpPk8zzvYDp7CFMpjNPoZSpbP75t SiEhzWq/sB6ZNXHlJaowWzV0= X-Envelope-To: linux-kernel@vger.kernel.org Received: from localhost (2a03:2880:10ff:c::) by mta10.migadu.com with ESMTPS id 0ff7917604fd4d10; Tue, 18 Aug 2026 13:12:12 +0000 X-Migadu-Flow: FLOW_OUT From: Usama Arif To: Andrew Morton , david@kernel.org, chrisl@kernel.org, kasong@tencent.com, ljs@kernel.org, ziy@nvidia.com, linux-mm@kvack.org Cc: ying.huang@linux.alibaba.com, Baoquan He , willy@infradead.org, youngjun.park@lge.com, hannes@cmpxchg.org, riel@surriel.com, shakeel.butt@linux.dev, alex@ghiti.fr, kas@kernel.org, baohua@kernel.org, dev.jain@arm.com, baolin.wang@linux.alibaba.com, Nico Pache , Liam R. Howlett , ryan.roberts@arm.com, Vlastimil Babka , lance.yang@linux.dev, linux-kernel@vger.kernel.org, nphamcs@gmail.com, shikemeng@huaweicloud.com, yosry@kernel.org, kernel-team@meta.com, Usama Arif Subject: [PATCH v6 01/12] mm: rename pmd_to_softleaf_folio() to pmd_softleaf_to_folio() Date: Tue, 18 Aug 2026 06:09:42 -0700 Message-ID: <20260818131202.494754-2-usama.arif@linux.dev> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260818131202.494754-1-usama.arif@linux.dev> References: <20260818131202.494754-1-usama.arif@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" pmd_to_softleaf_folio() converts the softleaf entry encoded by a PMD to a folio. Rename it to pmd_softleaf_to_folio() to make the conversion direction explicit and align it with softleaf_to_folio(). No functional change. Suggested-by: Dev Jain Signed-off-by: Usama Arif Acked-by: David Hildenbrand (Arm) Acked-by: Kiryl Shutsemau (Meta) Reviewed-by: Lorenzo Stoakes (ARM) Reviewed-by: Zi Yan --- include/linux/leafops.h | 4 ++-- mm/huge_memory.c | 2 +- 2 files changed, 3 insertions(+), 3 deletions(-) diff --git a/include/linux/leafops.h b/include/linux/leafops.h index 4c1476ae32343..7c13c58a5e218 100644 --- a/include/linux/leafops.h +++ b/include/linux/leafops.h @@ -657,7 +657,7 @@ static inline bool pmd_is_valid_softleaf(pmd_t pmd) } =20 /** - * pmd_to_softleaf_folio() - Convert the PMD entry to a folio. + * pmd_softleaf_to_folio() - Convert the PMD softleaf entry to a folio. * @pmd: PMD entry. * * The PMD entry is expected to be a valid PMD softleaf entry. @@ -665,7 +665,7 @@ static inline bool pmd_is_valid_softleaf(pmd_t pmd) * Returns: the folio the softleaf entry references if this is a valid sof= tleaf * entry, otherwise NULL. */ -static inline struct folio *pmd_to_softleaf_folio(pmd_t pmd) +static inline struct folio *pmd_softleaf_to_folio(pmd_t pmd) { const softleaf_t entry =3D softleaf_from_pmd(pmd); =20 diff --git a/mm/huge_memory.c b/mm/huge_memory.c index ced400f72d43a..1b6b0aa2baa3b 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -2467,7 +2467,7 @@ static struct folio *normal_or_softleaf_folio_pmd(str= uct vm_area_struct *vma, =20 if (!thp_migration_supported()) WARN_ONCE(1, "Non present huge pmd without pmd migration enabled!"); - return pmd_to_softleaf_folio(pmdval); + return pmd_softleaf_to_folio(pmdval); } =20 static bool has_deposited_pgtable(struct vm_area_struct *vma, pmd_t pmdval, --=20 2.53.0-Meta From nobody Mon Sep 28 19:34:00 2026 Received: from mta1.migadu.com (out-146.mta1.migadu.com [95.215.58.146]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1E5094749CF for ; Tue, 18 Aug 2026 13:12:16 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=95.215.58.146 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787058739; cv=none; b=PONsN5ydyZo0HIqJ8FFHpQNIIBHSUjQ+CAC/R8O6cQ8yKap6VqlDQFx4c11p+RdCJSFX4Yc5A4xAd66RM4+lEpXKpsP/dBS40eYkTsrW+sgdCxDuV7h9r1Was8UhV63Ezq8LIwpJ2ine+mmvweLp5cNzZVtHiacVshIzQNxrx2Y= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787058739; c=relaxed/simple; bh=Zq+Ctczt2l05WWIcApO8BE7BxqptNyjFszeHYSoNb5Y=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=HR95fQVsdFogGi7TiyEcBIOWbYOdxQRWy49nTmm9vFS0FvBKO/sdU/g0LDtxsYbhFKG6T2DAESk/12S2WGaI7zxvuYwAByq4JvwXgJrnlbOuKPOxCKE/8WG1K4xKk8LY1DsY2SQkwYCGtTRhwJsn6+HnmN63DxLBiveX8jyweww= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=iETa7Yfq; arc=none smtp.client-ip=95.215.58.146 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="iETa7Yfq" X-Envelope-To: linux-kernel@vger.kernel.org DKIM-Signature: a=rsa-sha256; bh=Zq+Ctczt2l05WWIcApO8BE7BxqptNyjFszeHYSoNb5Y=; c=simple/simple; d=linux.dev; h=from:to:subject:date:message-id:mime-version:content-type; s=key1; t=1787058735; v=1; x=1787663535; b=iETa7YfqCDf16oR4bL68jzlFBT6sD29y3n9ar2+BVMl22DWq+0bLgBlACwhW40WElExtNBSX sa6j2vts48wAzP2pOB4uzgGQEo6e+JnRIwIBsjVs56HKM2zoGOvTnc6WCQSoALypIrApDAfowH5 RV1xGlf+O4VLvg6AumaCzI9A= X-Envelope-To: linux-kernel@vger.kernel.org Received: from localhost (2a03:2880:10ff:42::) by mta10.migadu.com with ESMTPS id 2117379505cae7b9; Tue, 18 Aug 2026 13:12:14 +0000 X-Migadu-Flow: FLOW_OUT From: Usama Arif To: Andrew Morton , david@kernel.org, chrisl@kernel.org, kasong@tencent.com, ljs@kernel.org, ziy@nvidia.com, linux-mm@kvack.org Cc: ying.huang@linux.alibaba.com, Baoquan He , willy@infradead.org, youngjun.park@lge.com, hannes@cmpxchg.org, riel@surriel.com, shakeel.butt@linux.dev, alex@ghiti.fr, kas@kernel.org, baohua@kernel.org, dev.jain@arm.com, baolin.wang@linux.alibaba.com, Nico Pache , Liam R. Howlett , ryan.roberts@arm.com, Vlastimil Babka , lance.yang@linux.dev, linux-kernel@vger.kernel.org, nphamcs@gmail.com, shikemeng@huaweicloud.com, yosry@kernel.org, kernel-team@meta.com, Usama Arif Subject: [PATCH v6 02/12] mm: add PMD swap entry detection support Date: Tue, 18 Aug 2026 06:09:43 -0700 Message-ID: <20260818131202.494754-3-usama.arif@linux.dev> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260818131202.494754-1-usama.arif@linux.dev> References: <20260818131202.494754-1-usama.arif@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Currently when a PMD-mapped THP is swapped out, the PMD is always split into HPAGE_PMD_NR PTE-level swap entries. To preserve huge page information across swap cycles, later patches will install a single PMD-level swap entry instead. Add the infrastructure to detect those entries. Teach the softleaf layer to recognise PMD swap entries: pmd_is_swap_entry() detects them and softleaf_is_valid_pmd_entry() accepts them as a valid non-present type. Because swap entries do not encode a PFN, make pmd_softleaf_to_folio() warn and return NULL for them instead of passing the swap offset to softleaf_to_folio(). Clear the exclusive overlay bit in softleaf_from_pmd() before decoding, matching how soft_dirty and uffd bits are already stripped. Add pmd_swp_mkexclusive(), pmd_swp_exclusive(), and pmd_swp_clear_exclusive() helpers to each architecture that supports PMD softleaf entries (x86, arm64, s390, riscv, loongarch, powerpc), mirroring the existing PTE swap exclusive helpers in each arch's pgtable.h. Provide generic no-op PMD swap exclusive fallbacks for architectures without PMD softleaf support, matching the generic PMD swap soft-dirty fallbacks. Signed-off-by: Usama Arif --- arch/arm64/include/asm/pgtable.h | 6 +++++ arch/loongarch/include/asm/pgtable.h | 19 ++++++++++++++ arch/powerpc/include/asm/book3s/64/pgtable.h | 17 +++++++++++++ arch/riscv/include/asm/pgtable.h | 15 +++++++++++ arch/s390/include/asm/pgtable.h | 17 +++++++++++++ arch/x86/include/asm/pgtable.h | 17 +++++++++++++ include/linux/leafops.h | 26 ++++++++++++++++---- include/linux/pgtable.h | 17 +++++++++++++ 8 files changed, 129 insertions(+), 5 deletions(-) diff --git a/arch/arm64/include/asm/pgtable.h b/arch/arm64/include/asm/pgta= ble.h index a2681d7553584..860f95573d1fb 100644 --- a/arch/arm64/include/asm/pgtable.h +++ b/arch/arm64/include/asm/pgtable.h @@ -598,6 +598,12 @@ static inline int pmd_protnone(pmd_t pmd) #define pmd_swp_clear_uffd(pmd) \ pte_pmd(pte_swp_clear_uffd(pmd_pte(pmd))) #endif /* CONFIG_HAVE_ARCH_USERFAULTFD_WP */ +#ifdef CONFIG_ARCH_HAS_PMD_SOFTLEAVES +#define pmd_swp_exclusive(pmd) pte_swp_exclusive(pmd_pte(pmd)) +#define pmd_swp_mkexclusive(pmd) pte_pmd(pte_swp_mkexclusive(pmd_pte(pmd))) +#define pmd_swp_clear_exclusive(pmd) \ + pte_pmd(pte_swp_clear_exclusive(pmd_pte(pmd))) +#endif =20 #define pmd_write(pmd) pte_write(pmd_pte(pmd)) =20 diff --git a/arch/loongarch/include/asm/pgtable.h b/arch/loongarch/include/= asm/pgtable.h index 1952e34bc8ee0..aa8e1223d3973 100644 --- a/arch/loongarch/include/asm/pgtable.h +++ b/arch/loongarch/include/asm/pgtable.h @@ -357,6 +357,25 @@ static inline pte_t pte_swp_clear_exclusive(pte_t pte) return pte; } =20 +#ifdef CONFIG_ARCH_HAS_PMD_SOFTLEAVES +static inline pmd_t pmd_swp_mkexclusive(pmd_t pmd) +{ + pmd_val(pmd) |=3D _PAGE_SWP_EXCLUSIVE; + return pmd; +} + +static inline bool pmd_swp_exclusive(pmd_t pmd) +{ + return pmd_val(pmd) & _PAGE_SWP_EXCLUSIVE; +} + +static inline pmd_t pmd_swp_clear_exclusive(pmd_t pmd) +{ + pmd_val(pmd) &=3D ~_PAGE_SWP_EXCLUSIVE; + return pmd; +} +#endif + #define pte_none(pte) (!(pte_val(pte) & ~_PAGE_GLOBAL)) #define pte_present(pte) (pte_val(pte) & (_PAGE_PRESENT | _PAGE_PROTNONE)) #define pte_no_exec(pte) (pte_val(pte) & _PAGE_NO_EXEC) diff --git a/arch/powerpc/include/asm/book3s/64/pgtable.h b/arch/powerpc/in= clude/asm/book3s/64/pgtable.h index f4db7d7fbd5c6..6a899d0793b3b 100644 --- a/arch/powerpc/include/asm/book3s/64/pgtable.h +++ b/arch/powerpc/include/asm/book3s/64/pgtable.h @@ -699,6 +699,23 @@ static inline pte_t pte_swp_clear_exclusive(pte_t pte) return __pte_raw(pte_raw(pte) & cpu_to_be64(~_PAGE_SWP_EXCLUSIVE)); } =20 +#ifdef CONFIG_ARCH_HAS_PMD_SOFTLEAVES +static inline pmd_t pmd_swp_mkexclusive(pmd_t pmd) +{ + return __pmd_raw(pmd_raw(pmd) | cpu_to_be64(_PAGE_SWP_EXCLUSIVE)); +} + +static inline bool pmd_swp_exclusive(pmd_t pmd) +{ + return !!(pmd_raw(pmd) & cpu_to_be64(_PAGE_SWP_EXCLUSIVE)); +} + +static inline pmd_t pmd_swp_clear_exclusive(pmd_t pmd) +{ + return __pmd_raw(pmd_raw(pmd) & cpu_to_be64(~_PAGE_SWP_EXCLUSIVE)); +} +#endif + static inline bool check_pte_access(unsigned long access, unsigned long pt= ev) { /* diff --git a/arch/riscv/include/asm/pgtable.h b/arch/riscv/include/asm/pgta= ble.h index 1225cf05696a2..65b4181c62d9a 100644 --- a/arch/riscv/include/asm/pgtable.h +++ b/arch/riscv/include/asm/pgtable.h @@ -1213,6 +1213,21 @@ static inline pte_t pte_swp_clear_exclusive(pte_t pt= e) } =20 #ifdef CONFIG_ARCH_HAS_PMD_SOFTLEAVES +static inline bool pmd_swp_exclusive(pmd_t pmd) +{ + return pte_swp_exclusive(pmd_pte(pmd)); +} + +static inline pmd_t pmd_swp_mkexclusive(pmd_t pmd) +{ + return pte_pmd(pte_swp_mkexclusive(pmd_pte(pmd))); +} + +static inline pmd_t pmd_swp_clear_exclusive(pmd_t pmd) +{ + return pte_pmd(pte_swp_clear_exclusive(pmd_pte(pmd))); +} + #define __pmd_to_swp_entry(pmd) ((swp_entry_t) { pmd_val(pmd) }) #define __swp_entry_to_pmd(swp) __pmd((swp).val) #endif /* CONFIG_ARCH_HAS_PMD_SOFTLEAVES */ diff --git a/arch/s390/include/asm/pgtable.h b/arch/s390/include/asm/pgtabl= e.h index e882663a58e77..490e4a3464b19 100644 --- a/arch/s390/include/asm/pgtable.h +++ b/arch/s390/include/asm/pgtable.h @@ -870,6 +870,23 @@ static inline pte_t pte_swp_clear_exclusive(pte_t pte) return clear_pte_bit(pte, __pgprot(_PAGE_SWP_EXCLUSIVE)); } =20 +#ifdef CONFIG_ARCH_HAS_PMD_SOFTLEAVES +static inline pmd_t pmd_swp_mkexclusive(pmd_t pmd) +{ + return set_pmd_bit(pmd, __pgprot(_PAGE_SWP_EXCLUSIVE)); +} + +static inline bool pmd_swp_exclusive(pmd_t pmd) +{ + return pmd_val(pmd) & _PAGE_SWP_EXCLUSIVE; +} + +static inline pmd_t pmd_swp_clear_exclusive(pmd_t pmd) +{ + return clear_pmd_bit(pmd, __pgprot(_PAGE_SWP_EXCLUSIVE)); +} +#endif + static inline int pte_soft_dirty(pte_t pte) { return pte_val(pte) & _PAGE_SOFT_DIRTY; diff --git a/arch/x86/include/asm/pgtable.h b/arch/x86/include/asm/pgtable.h index 8e0018fadd14e..b5da5447e83d0 100644 --- a/arch/x86/include/asm/pgtable.h +++ b/arch/x86/include/asm/pgtable.h @@ -1525,6 +1525,23 @@ static inline pte_t pte_swp_clear_exclusive(pte_t pt= e) return pte_clear_flags(pte, _PAGE_SWP_EXCLUSIVE); } =20 +#ifdef CONFIG_ARCH_HAS_PMD_SOFTLEAVES +static inline pmd_t pmd_swp_mkexclusive(pmd_t pmd) +{ + return pmd_set_flags(pmd, _PAGE_SWP_EXCLUSIVE); +} + +static inline int pmd_swp_exclusive(pmd_t pmd) +{ + return pmd_flags(pmd) & _PAGE_SWP_EXCLUSIVE; +} + +static inline pmd_t pmd_swp_clear_exclusive(pmd_t pmd) +{ + return pmd_clear_flags(pmd, _PAGE_SWP_EXCLUSIVE); +} +#endif + #ifdef CONFIG_HAVE_ARCH_SOFT_DIRTY static inline pte_t pte_swp_mksoft_dirty(pte_t pte) { diff --git a/include/linux/leafops.h b/include/linux/leafops.h index 7c13c58a5e218..4a6c52974b305 100644 --- a/include/linux/leafops.h +++ b/include/linux/leafops.h @@ -102,6 +102,8 @@ static inline softleaf_t softleaf_from_pmd(pmd_t pmd) pmd =3D pmd_swp_clear_soft_dirty(pmd); if (pmd_swp_uffd(pmd)) pmd =3D pmd_swp_clear_uffd(pmd); + if (pmd_swp_exclusive(pmd)) + pmd =3D pmd_swp_clear_exclusive(pmd); arch_entry =3D __pmd_to_swp_entry(pmd); =20 /* Temporary until swp_entry_t eliminated. */ @@ -634,18 +636,30 @@ static inline bool pmd_is_migration_entry(pmd_t pmd) */ static inline bool softleaf_is_valid_pmd_entry(softleaf_t entry) { - /* Only device private, migration entries valid for PMD. */ + /* Device private, migration, and swap entries valid for PMD. */ return softleaf_is_device_private(entry) || - softleaf_is_migration(entry); + softleaf_is_migration(entry) || + softleaf_is_swap(entry); +} + +/** + * pmd_is_swap_entry() - Does this PMD entry encode an actual swap entry? + * @pmd: PMD entry. + * + * Returns: true if the PMD encodes a swap entry, otherwise false. + */ +static inline bool pmd_is_swap_entry(pmd_t pmd) +{ + return softleaf_is_swap(softleaf_from_pmd(pmd)); } =20 /** * pmd_is_valid_softleaf() - Is this PMD entry a valid softleaf entry? * @pmd: PMD entry. * - * PMD leaf entries are valid only if they are device private or migration - * entries. This function asserts that a PMD leaf entry is valid in this - * respect. + * PMD leaf entries are valid only if they are device private, migration, + * or swap entries. This function asserts that a PMD leaf entry is valid + * in this respect. * * Returns: true if the PMD entry is a valid leaf entry, otherwise false. */ @@ -673,6 +687,8 @@ static inline struct folio *pmd_softleaf_to_folio(pmd_t= pmd) VM_WARN_ON_ONCE(true); return NULL; } + if (WARN_ON_ONCE(!softleaf_has_pfn(entry))) + return NULL; return softleaf_to_folio(entry); } =20 diff --git a/include/linux/pgtable.h b/include/linux/pgtable.h index 8c093c119e5a8..e10a7e91e4260 100644 --- a/include/linux/pgtable.h +++ b/include/linux/pgtable.h @@ -1917,6 +1917,23 @@ static inline pmd_t pmd_swp_clear_soft_dirty(pmd_t p= md) } #endif =20 +#ifndef CONFIG_ARCH_HAS_PMD_SOFTLEAVES +static inline pmd_t pmd_swp_mkexclusive(pmd_t pmd) +{ + return pmd; +} + +static inline bool pmd_swp_exclusive(pmd_t pmd) +{ + return false; +} + +static inline pmd_t pmd_swp_clear_exclusive(pmd_t pmd) +{ + return pmd; +} +#endif + #ifndef __HAVE_PFNMAP_TRACKING /* * Interfaces that can be used by architecture code to keep track of --=20 2.53.0-Meta From nobody Mon Sep 28 19:34:00 2026 Received: from mta0.migadu.com (out-110.mta0.migadu.com [91.218.175.110]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 90D7D47278D for ; Tue, 18 Aug 2026 13:12:23 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.110 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787058745; cv=none; b=aMoeaqy5gKTZVC4NRjrF1+hAEQPKLoiMb33hwtPMj1oB8CTkIXPaphdiYg2PIpmMeZnXqD2xqzeLgI5Yi/iB8X4f702y6KhIf/wF+GMCkBsPZu0EH/Di5mFeFncCg+Q0FEE5/wLWhHEGTcdVUuAHUh9SV4hyqFjvZg+e2DGPxKU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787058745; c=relaxed/simple; bh=1Dz5t2k/qHsfuuzKgM9VwuvNQXNi7pjePtn1+ZkgY2I=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=XHZtZQt2NzDwR+USXVUPg+otQjDABZqD6ZpvRueeR0m0gWybI6MFTIwposBHLDRblveLvYaVNRUeeUyCTyKMbKSP4a7AeLHENleh6fIFgZhdXeU2DJk5+JCi9NQZm5lC2V9ExgInpo8VhB10EpQvw0hnHfzqB+X7eWag2FbWce0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=ePR1CsFS; arc=none smtp.client-ip=91.218.175.110 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="ePR1CsFS" X-Envelope-To: linux-kernel@vger.kernel.org DKIM-Signature: a=rsa-sha256; bh=1Dz5t2k/qHsfuuzKgM9VwuvNQXNi7pjePtn1+ZkgY2I=; c=simple/simple; d=linux.dev; h=from:to:subject:date:message-id:mime-version:content-type; s=key1; t=1787058741; v=1; x=1787663541; b=ePR1CsFSfhtZnZDwdzp278rbXm8Fa/19ikhEl7qhC9YsSoeftTdy+Lf6sIPvRiKMyN1OB208 DTU5CEAEzL+MTIT3nlp1lfEVSmD0gQu61fhKKMUP+r7SWZ4IvXz7XZVKBeflgERBPyvmBtrjFYt u3wJ9i5U9DwEHpNIK9QKyAXs= X-Envelope-To: linux-kernel@vger.kernel.org Received: from localhost (2a03:2880:10ff:48::) by mta11.migadu.com with ESMTPS id 841d0e040e1652fc; Tue, 18 Aug 2026 13:12:21 +0000 X-Migadu-Flow: FLOW_OUT From: Usama Arif To: Andrew Morton , david@kernel.org, chrisl@kernel.org, kasong@tencent.com, ljs@kernel.org, ziy@nvidia.com, linux-mm@kvack.org Cc: ying.huang@linux.alibaba.com, Baoquan He , willy@infradead.org, youngjun.park@lge.com, hannes@cmpxchg.org, riel@surriel.com, shakeel.butt@linux.dev, alex@ghiti.fr, kas@kernel.org, baohua@kernel.org, dev.jain@arm.com, baolin.wang@linux.alibaba.com, Nico Pache , Liam R. Howlett , ryan.roberts@arm.com, Vlastimil Babka , lance.yang@linux.dev, linux-kernel@vger.kernel.org, nphamcs@gmail.com, shikemeng@huaweicloud.com, yosry@kernel.org, kernel-team@meta.com, Usama Arif Subject: [PATCH v6 03/12] mm: add PMD swap entry splitting support Date: Tue, 18 Aug 2026 06:09:44 -0700 Message-ID: <20260818131202.494754-4-usama.arif@linux.dev> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260818131202.494754-1-usama.arif@linux.dev> References: <20260818131202.494754-1-usama.arif@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add a swap branch in __split_huge_pmd_locked() that splits a PMD swap entry into 512 PTE swap entries. No folio reference is needed because swap entries point to swap slots rather than pages. Each PTE inherits the correct sub-slot offset and preserves soft_dirty, uffd_wp, and exclusive flags. The folio_remove_rmap_pmd() gate at the end must inspect old_pmd rather than *pmd: for a present THP split, *pmd has already been cleared by pmdp_invalidate(), and that invalidated bit pattern can decode as a plausible swap entry. This branch is reached from the explicit __split_huge_pmd() callers that hit a non-present PMD: partial-range mprotect / munmap, the wp_huge_pmd() PMD-COW fallback, and the swap-in / swapoff fallbacks added in later patches when the cached folio is no longer PMD-sized. page_vma_mapped_walk() does not iterate PMD swap entries, so try_to_unmap_one() and try_to_migrate_one() do not reach this branch and freeze=3Dtrue cannot occur in this branch today. page and folio are therefore left uninitialized in the swap branch; a VM_WARN_ON_ONCE(freeze) catches any future caller that breaks this invariant before the freeze path dereferences page_to_pfn(page + i) or put_page(page). Signed-off-by: Usama Arif --- mm/huge_memory.c | 29 ++++++++++++++++++++++++++++- 1 file changed, 28 insertions(+), 1 deletion(-) diff --git a/mm/huge_memory.c b/mm/huge_memory.c index 1b6b0aa2baa3b..a473e85d30f51 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -3252,6 +3252,14 @@ static void __split_huge_pmd_locked(struct vm_area_s= truct *vma, pmd_t *pmd, folio_add_anon_rmap_ptes(folio, page, HPAGE_PMD_NR, vma, haddr, rmap_flags); } + } else if (pmd_is_swap_entry(*pmd)) { + VM_WARN_ON_ONCE(freeze); + /* Swap entries have no page for the migration freeze path. */ + freeze =3D false; + old_pmd =3D *pmd; + soft_dirty =3D pmd_swp_soft_dirty(old_pmd); + uffd_wp =3D pmd_swp_uffd(old_pmd); + anon_exclusive =3D pmd_swp_exclusive(old_pmd); } else { /* * Up to this point the pmd is present and huge and userland has @@ -3388,6 +3396,25 @@ static void __split_huge_pmd_locked(struct vm_area_s= truct *vma, pmd_t *pmd, VM_WARN_ON(!pte_none(ptep_get(pte + i))); set_pte_at(mm, addr, pte + i, entry); } + } else if (pmd_is_swap_entry(old_pmd)) { + softleaf_t sl_entry =3D softleaf_from_pmd(old_pmd); + pte_t swp_pte; + swp_entry_t sub_entry; + + for (i =3D 0, addr =3D haddr; i < HPAGE_PMD_NR; + i++, addr +=3D PAGE_SIZE) { + sub_entry =3D swp_entry(swp_type(sl_entry), + swp_offset(sl_entry) + i); + swp_pte =3D swp_entry_to_pte(sub_entry); + if (soft_dirty) + swp_pte =3D pte_swp_mksoft_dirty(swp_pte); + if (uffd_wp) + swp_pte =3D pte_swp_mkuffd(swp_pte); + if (anon_exclusive) + swp_pte =3D pte_swp_mkexclusive(swp_pte); + VM_WARN_ON(!pte_none(ptep_get(pte + i))); + set_pte_at(mm, addr, pte + i, swp_pte); + } } else { pte_t entry; =20 @@ -3415,7 +3442,7 @@ static void __split_huge_pmd_locked(struct vm_area_st= ruct *vma, pmd_t *pmd, } pte_unmap(pte); =20 - if (!pmd_is_migration_entry(*pmd)) + if (!pmd_is_migration_entry(old_pmd) && !pmd_is_swap_entry(old_pmd)) folio_remove_rmap_pmd(folio, page, vma); if (freeze) put_page(page); --=20 2.53.0-Meta From nobody Mon Sep 28 19:34:00 2026 Received: from mta1.migadu.com (out-171.mta1.migadu.com [95.215.58.171]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 37666472520 for ; Tue, 18 Aug 2026 13:12:25 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=95.215.58.171 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787058748; cv=none; b=nT55tD84NIFHQQA1yWcrnkCSrbyMseZ/jasF+VNmS78dXH9UNrbNHFMVZvqMPA3bE6LW9deZO8MskezbXY2iGTRLE0ojJ9vTl+6JXcYxrJRE/KS3IoIwWmGbcdMDvQ8ZDJYldLT91xZy+ae/dUl6228isg74cEGBrj8d8PJD4Zg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787058748; c=relaxed/simple; bh=U1HZxdR4YZz99uGbDOKosgWNZREILwS2Y40gIn/27F0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=uMqf7T0JHGWmsako2t086JGbHFXiipgOk6U1QxbOupBR4h0Tid0qbUufaYkxLzySgdazpA82CapE9Vn7uLM8Wl9QERDr/+jhnAZoDe38bYwN3MIZ73N6wERWOEzcBsyPtt3wDZhmd/0IsFruE5LFWvg31ZBsIp6Yym1loaiIFHs= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=QOoQBRwi; arc=none smtp.client-ip=95.215.58.171 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="QOoQBRwi" X-Envelope-To: linux-kernel@vger.kernel.org DKIM-Signature: a=rsa-sha256; bh=U1HZxdR4YZz99uGbDOKosgWNZREILwS2Y40gIn/27F0=; c=simple/simple; d=linux.dev; h=from:to:subject:date:message-id:mime-version:content-type; s=key1; t=1787058744; v=1; x=1787663544; b=QOoQBRwiW8eEyszUwOzR/3F9C42Gdx3XyvaGpG8KjilO2wn7wFcuV4rmi7jY51rus3pQRoED 0hFMvfrFTxhvrItomsqjve6V0lKWyavjUOPCy/bBF7S95NjPfV4RnScKs2GFkpk6SlFtIY1S+mJ gLF8aUad4WqVAZfDf1/NFs94= X-Envelope-To: linux-kernel@vger.kernel.org Received: from localhost (2a03:2880:10ff:58::) by mta11.migadu.com with ESMTPS id d2282fcef2b2083d; Tue, 18 Aug 2026 13:12:23 +0000 X-Migadu-Flow: FLOW_OUT From: Usama Arif To: Andrew Morton , david@kernel.org, chrisl@kernel.org, kasong@tencent.com, ljs@kernel.org, ziy@nvidia.com, linux-mm@kvack.org Cc: ying.huang@linux.alibaba.com, Baoquan He , willy@infradead.org, youngjun.park@lge.com, hannes@cmpxchg.org, riel@surriel.com, shakeel.butt@linux.dev, alex@ghiti.fr, kas@kernel.org, baohua@kernel.org, dev.jain@arm.com, baolin.wang@linux.alibaba.com, Nico Pache , Liam R. Howlett , ryan.roberts@arm.com, Vlastimil Babka , lance.yang@linux.dev, linux-kernel@vger.kernel.org, nphamcs@gmail.com, shikemeng@huaweicloud.com, yosry@kernel.org, kernel-team@meta.com, Usama Arif Subject: [PATCH v6 04/12] mm: handle PMD swap entries in fork path Date: Tue, 18 Aug 2026 06:09:45 -0700 Message-ID: <20260818131202.494754-5-usama.arif@linux.dev> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260818131202.494754-1-usama.arif@linux.dev> References: <20260818131202.494754-1-usama.arif@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Teach copy_huge_pmd()/copy_huge_non_present_pmd() about swap entries, mirroring copy_nonpresent_pte(). swap_dup_entry_direct() gains a nr parameter (and is renamed to swap_dup_entries_direct()) so it can duplicate a contiguous range of swap slots in one call, matching the existing swap_put_entries_direct(entry, nr) API. Existing callers pass 1. swap_retry_table_alloc() likewise gains a nr parameter so the outer retry knows how many slots the caller was trying to duplicate. The underlying swap_extend_table_alloc() now scans every slot in [ci_off, ci_off + nr) to confirm that at least one still needs the per-cluster extend table before committing an allocation. copy_huge_non_present_pmd() "copies" PMD swap entries during fork instead of splitting, preserving the THP. This mirrors copy_nonpresent_pte() which duplicates the swap slot refcount, clears the exclusive bit on the source, and adds the destination mm to mmlist. If swap_dup_entries_direct() fails (GFP_ATOMIC table alloc), copy_huge_pmd() retries once after swap_retry_table_alloc(entry, HPAGE_PMD_NR, GFP_KERNEL). Signed-off-by: Usama Arif --- include/linux/swap.h | 4 +-- mm/huge_memory.c | 60 ++++++++++++++++++++++++++++++++++++++------ mm/memory.c | 4 +-- mm/swap.h | 5 ++-- mm/swapfile.c | 40 ++++++++++++++++++----------- 5 files changed, 86 insertions(+), 27 deletions(-) diff --git a/include/linux/swap.h b/include/linux/swap.h index 0f953ed9c8630..e465069361733 100644 --- a/include/linux/swap.h +++ b/include/linux/swap.h @@ -395,7 +395,7 @@ sector_t swap_folio_sector(struct folio *folio); * All entries must be allocated by folio_alloc_swap(). And they must have * a swap count > 1. See comments of folio_*_swap helpers for more info. */ -int swap_dup_entry_direct(swp_entry_t entry); +int swap_dup_entries_direct(swp_entry_t entry, int nr); void swap_put_entries_direct(swp_entry_t entry, int nr); =20 /* @@ -439,7 +439,7 @@ static inline void free_swap_cache(struct folio *folio) { } =20 -static inline int swap_dup_entry_direct(swp_entry_t ent) +static inline int swap_dup_entries_direct(swp_entry_t ent, int nr) { return 0; } diff --git a/mm/huge_memory.c b/mm/huge_memory.c index a473e85d30f51..2735c3de7029c 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -1849,7 +1849,7 @@ bool touch_pmd(struct vm_area_struct *vma, unsigned l= ong addr, return false; } =20 -static void copy_huge_non_present_pmd( +static int copy_huge_non_present_pmd( struct mm_struct *dst_mm, struct mm_struct *src_mm, pmd_t *dst_pmd, pmd_t *src_pmd, unsigned long addr, struct vm_area_struct *dst_vma, struct vm_area_struct *src_vma, @@ -1895,14 +1895,35 @@ static void copy_huge_non_present_pmd( */ folio_try_dup_anon_rmap_pmd(src_folio, &src_folio->page, dst_vma, src_vma); + } else if (softleaf_is_swap(entry)) { + int err; + + /* + * PMD swap entry: duplicate swap references and clear + * exclusive on source, matching copy_nonpresent_pte(). + */ + err =3D swap_dup_entries_direct(entry, HPAGE_PMD_NR); + if (err < 0) + return err; + + mm_prepare_for_swap_entries(dst_mm); + + if (pmd_swp_exclusive(pmd)) { + pmd =3D pmd_swp_clear_exclusive(pmd); + set_pmd_at(src_mm, addr, src_pmd, pmd); + } } =20 - add_mm_counter(dst_mm, MM_ANONPAGES, HPAGE_PMD_NR); + if (softleaf_is_swap(entry)) + add_mm_counter(dst_mm, MM_SWAPENTS, HPAGE_PMD_NR); + else + add_mm_counter(dst_mm, MM_ANONPAGES, HPAGE_PMD_NR); mm_inc_nr_ptes(dst_mm); pgtable_trans_huge_deposit(dst_mm, dst_pmd, pgtable); if (!userfaultfd_protected(dst_vma)) pmd =3D pmd_swp_clear_uffd(pmd); set_pmd_at(dst_mm, addr, dst_pmd, pmd); + return 0; } =20 int copy_huge_pmd(struct mm_struct *dst_mm, struct mm_struct *src_mm, @@ -1912,6 +1933,7 @@ int copy_huge_pmd(struct mm_struct *dst_mm, struct mm= _struct *src_mm, spinlock_t *dst_ptl, *src_ptl; struct page *src_page; struct folio *src_folio; + bool retried =3D false; pmd_t pmd; pgtable_t pgtable =3D NULL; int ret =3D -ENOMEM; @@ -1943,6 +1965,7 @@ int copy_huge_pmd(struct mm_struct *dst_mm, struct mm= _struct *src_mm, if (unlikely(!pgtable)) goto out; =20 +retry: dst_ptl =3D pmd_lock(dst_mm, dst_pmd); src_ptl =3D pmd_lockptr(src_mm, src_pmd); spin_lock_nested(src_ptl, SINGLE_DEPTH_NESTING); @@ -1950,11 +1973,34 @@ int copy_huge_pmd(struct mm_struct *dst_mm, struct = mm_struct *src_mm, ret =3D -EAGAIN; pmd =3D *src_pmd; =20 - if (unlikely(thp_migration_supported() && - pmd_is_valid_softleaf(pmd))) { - copy_huge_non_present_pmd(dst_mm, src_mm, dst_pmd, src_pmd, addr, - dst_vma, src_vma, pmd, pgtable); - ret =3D 0; + if (unlikely(pmd_is_valid_softleaf(pmd))) { + ret =3D copy_huge_non_present_pmd(dst_mm, src_mm, dst_pmd, src_pmd, + addr, dst_vma, src_vma, pmd, + pgtable); + if (ret) { + spin_unlock(src_ptl); + spin_unlock(dst_ptl); + /* + * For PMD swap entries -ENOMEM means the per-cluster + * swap-extend table couldn't be GFP_ATOMIC-allocated. + * Try the GFP_KERNEL fallback once before giving up. + * swap_retry_table_alloc() also returns 0 when it + * decides the table is not needed after all, so bound + * this to a single retry rather than looping on it. + */ + if (ret =3D=3D -ENOMEM && !retried) { + softleaf_t entry =3D softleaf_from_pmd(pmd); + + retried =3D true; + if (softleaf_is_swap(entry) && + !swap_retry_table_alloc(entry, HPAGE_PMD_NR, + GFP_KERNEL)) + goto retry; + } + pte_free(dst_mm, pgtable); + ret =3D -ENOMEM; + goto out; + } goto out_unlock; } =20 diff --git a/mm/memory.c b/mm/memory.c index 4134ac607ee0f..36ddca806be2f 100644 --- a/mm/memory.c +++ b/mm/memory.c @@ -1016,7 +1016,7 @@ copy_nonpresent_pte(struct mm_struct *dst_mm, struct = mm_struct *src_mm, struct page *page; =20 if (likely(softleaf_is_swap(entry))) { - if (swap_dup_entry_direct(entry) < 0) + if (swap_dup_entries_direct(entry, 1) < 0) return -EIO; =20 mm_prepare_for_swap_entries(dst_mm); @@ -1431,7 +1431,7 @@ copy_pte_range(struct vm_area_struct *dst_vma, struct= vm_area_struct *src_vma, =20 if (ret =3D=3D -EIO) { VM_WARN_ON_ONCE(!entry.val); - if (swap_retry_table_alloc(entry, GFP_KERNEL) < 0) { + if (swap_retry_table_alloc(entry, 1, GFP_KERNEL) < 0) { ret =3D -ENOMEM; goto out; } diff --git a/mm/swap.h b/mm/swap.h index 90a551a88df63..e225527b2eb7a 100644 --- a/mm/swap.h +++ b/mm/swap.h @@ -222,7 +222,7 @@ static inline void swap_cluster_unlock_irq(struct swap_= cluster_info *ci) spin_unlock_irq(&ci->lock); } =20 -extern int swap_retry_table_alloc(swp_entry_t entry, gfp_t gfp); +int swap_retry_table_alloc(swp_entry_t entry, unsigned int nr, gfp_t gfp); =20 /* * Below are the core routines for doing swap for a folio. @@ -428,7 +428,8 @@ static inline int swap_writeout(struct swap_io_ctx *ctx= , struct folio *folio) return 0; } =20 -static inline int swap_retry_table_alloc(swp_entry_t entry, gfp_t gfp) +static inline int swap_retry_table_alloc(swp_entry_t entry, unsigned int n= r, + gfp_t gfp) { return -EINVAL; } diff --git a/mm/swapfile.c b/mm/swapfile.c index f5dfc7e59191e..44b9ebbe7229f 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -1462,9 +1462,11 @@ static bool swap_sync_discard(void) =20 static int swap_extend_table_alloc(struct swap_info_struct *si, struct swap_cluster_info *ci, - unsigned int ci_off, gfp_t gfp) + unsigned int ci_off, unsigned int nr, + gfp_t gfp) { int count; + unsigned int i; void *table; =20 table =3D kzalloc(sizeof(ci->extend_table[0]) * SWAPFILE_CLUSTER, gfp); @@ -1480,15 +1482,21 @@ static int swap_extend_table_alloc(struct swap_info= _struct *si, */ if (!cluster_table_is_alloced(ci)) goto out_free; - count =3D swp_tb_get_count(__swap_table_get(ci, ci_off)); - if (count < (SWP_TB_COUNT_MAX - 1)) - goto out_free; if (ci->extend_table) goto out_free; - - ci->extend_table =3D table; - spin_unlock(&ci->lock); - return 0; + /* + * The caller may not know which slot in [ci_off, ci_off + nr) hit + * SWP_TB_COUNT_MAX - 1. Confirm at least one slot in the range still + * needs the extend table before committing the allocation. + */ + for (i =3D 0; i < nr; i++) { + count =3D swp_tb_get_count(__swap_table_get(ci, ci_off + i)); + if (count >=3D (SWP_TB_COUNT_MAX - 1)) { + ci->extend_table =3D table; + spin_unlock(&ci->lock); + return 0; + } + } =20 out_free: spin_unlock(&ci->lock); @@ -1496,7 +1504,7 @@ static int swap_extend_table_alloc(struct swap_info_s= truct *si, return 0; } =20 -int swap_retry_table_alloc(swp_entry_t entry, gfp_t gfp) +int swap_retry_table_alloc(swp_entry_t entry, unsigned int nr, gfp_t gfp) { int ret; struct swap_info_struct *si; @@ -1508,7 +1516,8 @@ int swap_retry_table_alloc(swp_entry_t entry, gfp_t g= fp) return 0; =20 ci =3D __swap_offset_to_cluster(si, offset); - ret =3D swap_extend_table_alloc(si, ci, swp_cluster_offset(entry), gfp); + ret =3D swap_extend_table_alloc(si, ci, swp_cluster_offset(entry), nr, + gfp); =20 put_swap_device(si); return ret; @@ -1709,7 +1718,8 @@ static int swap_dup_entries_cluster(struct swap_info_= struct *si, if (unlikely(err)) { if (err =3D=3D -ENOMEM) { spin_unlock(&ci->lock); - err =3D swap_extend_table_alloc(si, ci, ci_off, GFP_ATOMIC); + err =3D swap_extend_table_alloc(si, ci, ci_off, 1, + GFP_ATOMIC); spin_lock(&ci->lock); if (!err) goto restart; @@ -1720,6 +1730,7 @@ static int swap_dup_entries_cluster(struct swap_info_= struct *si, swap_cluster_unlock(ci); return 0; failed: + /* The caller's page-table or swap-cache reference pins every slot. */ while (ci_off-- > ci_start) __swap_cluster_put_entry(ci, ci_off); swap_extend_table_try_free(ci); @@ -3925,8 +3936,9 @@ void si_swapinfo(struct sysinfo *val) } =20 /* - * swap_dup_entry_direct() - Increase reference count of a swap entry by o= ne. + * swap_dup_entries_direct() - Increase reference count of swap entries by= one. * @entry: first swap entry from which we want to increase the refcount. + * @nr: number of contiguous swap entries to duplicate. * * Returns 0 for success, or -ENOMEM if the extend table is required * but could not be atomically allocated. Returns -EINVAL if the swap @@ -3938,7 +3950,7 @@ void si_swapinfo(struct sysinfo *val) * Also the swap entry must have a count >=3D 1. Otherwise folio_dup_swap = should * be used. */ -int swap_dup_entry_direct(swp_entry_t entry) +int swap_dup_entries_direct(swp_entry_t entry, int nr) { struct swap_info_struct *si; =20 @@ -3955,7 +3967,7 @@ int swap_dup_entry_direct(swp_entry_t entry) */ VM_WARN_ON_ONCE(!swap_entry_swapped(si, entry)); =20 - return swap_dup_entries_cluster(si, swp_offset(entry), 1); + return swap_dup_entries_cluster(si, swp_offset(entry), nr); } =20 #if defined(CONFIG_MEMCG) && defined(CONFIG_BLK_CGROUP) --=20 2.53.0-Meta From nobody Mon Sep 28 19:34:00 2026 Received: from mta1.migadu.com (out-180.mta1.migadu.com [95.215.58.180]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0FF2F28850E for ; Tue, 18 Aug 2026 13:12:32 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=95.215.58.180 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787058754; cv=none; b=Z+5LSGMBdPAwUb2XerLGbGOofSuzkn0ojiWcvpNn808yyePQAkykMFMS82LzjmMw28447a5jVGR0ZzT27HBmEcOApWSFzxmVo7zIc5CgfbpBxneTVNmVts0WtZhrFWPu2q3ZWqqwIweOfDUZzlOA/SdhgSGIuHX5JDDVSACK284= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787058754; c=relaxed/simple; bh=FPWpaNmTRU+xlShU64gQ7An1URF32LkC2T1D7RJM8B8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Jh9uGAjH6nK4cVMYfIeiWAK231jtco3Sq2s4lKFo9hgx+wZSTiy3AC/pAjiPnokta0wTpH0kL4P4bzvbXXUzCSvtmZ8oSgp/jEXBvtUkc5C3poZpiOLfKWAQ3F5dNeR38wIbyacQqOZWxkVgKecP4KFfXWM/CeW9PO+r6U/mcWc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=fea0Iqtr; arc=none smtp.client-ip=95.215.58.180 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="fea0Iqtr" X-Envelope-To: linux-kernel@vger.kernel.org DKIM-Signature: a=rsa-sha256; bh=FPWpaNmTRU+xlShU64gQ7An1URF32LkC2T1D7RJM8B8=; c=simple/simple; d=linux.dev; h=from:to:subject:date:message-id:mime-version:content-type; s=key1; t=1787058751; v=1; x=1787663551; b=fea0IqtrG9VfdW11Ia+oI7Ez+E9ZswSNW2EYTFMalLV10nXvOUsWkhe47GUGBpIjpR1bQRac /rz211dnpHetCJsAXDxxbkV89Wo+Vp5FPwMfIcczEO4oDcFktLRA4g8J/b10Ynn6NZb4bg3koob OnWCfJSN9goBD4Dec+hZeV+I= X-Envelope-To: linux-kernel@vger.kernel.org Received: from localhost (2a03:2880:10ff:1a::) by mta12.migadu.com with ESMTPS id 7c31fdcbe05a5cd5; Tue, 18 Aug 2026 13:12:30 +0000 X-Migadu-Flow: FLOW_OUT From: Usama Arif To: Andrew Morton , david@kernel.org, chrisl@kernel.org, kasong@tencent.com, ljs@kernel.org, ziy@nvidia.com, linux-mm@kvack.org Cc: ying.huang@linux.alibaba.com, Baoquan He , willy@infradead.org, youngjun.park@lge.com, hannes@cmpxchg.org, riel@surriel.com, shakeel.butt@linux.dev, alex@ghiti.fr, kas@kernel.org, baohua@kernel.org, dev.jain@arm.com, baolin.wang@linux.alibaba.com, Nico Pache , Liam R. Howlett , ryan.roberts@arm.com, Vlastimil Babka , lance.yang@linux.dev, linux-kernel@vger.kernel.org, nphamcs@gmail.com, shikemeng@huaweicloud.com, yosry@kernel.org, kernel-team@meta.com, Alexandre Ghiti , Usama Arif Subject: [PATCH v6 05/12] mm: zswap: add range lookup for large-folio swapin Date: Tue, 18 Aug 2026 06:09:46 -0700 Message-ID: <20260818131202.494754-6-usama.arif@linux.dev> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260818131202.494754-1-usama.arif@linux.dev> References: <20260818131202.494754-1-usama.arif@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: Alexandre Ghiti A large folio reaches zswap_load() only when the caller expects the whole range to be on disk. Zswap still stores large folios as independent order-0 entries, so reconstructing a large folio from zswap entries would risk returning partially initialized data. Teach zswap_load() to scan the covered range. If no slot is in zswap, return -ENOENT so swap_read_folio() reads the backing device. If any slot is still in zswap, fail the large-folio read so the caller can fall back to per-page swapin. Return -EIO rather than -EINVAL for that conflict. Large-folio loads are now valid requests; the error means zswap cannot safely satisfy the request from partial per-page compressed state, not that the request is unsupported. Existing callers only distinguish -ENOENT, so this is a semantic clarification rather than a behavioral change. Add zswap_is_present() so PMD swap-entry consumers can make the same range decision before attempting PMD-order swapin. Also use it from __swap_cache_add_check() for multi-page insertions while holding the swap cluster lock. That check runs before folio allocation and again immediately before swap-cache insertion, closing the race with zswap writeback and rejecting mixed zswap/disk backing with -EBUSY. Signed-off-by: Alexandre Ghiti Signed-off-by: Usama Arif --- include/linux/zswap.h | 6 ++++++ mm/swap_state.c | 10 ++++++++++ mm/zswap.c | 46 +++++++++++++++++++++++++++++++------------ 3 files changed, 49 insertions(+), 13 deletions(-) diff --git a/include/linux/zswap.h b/include/linux/zswap.h index 30c193a1207e1..cd9efcf9dec94 100644 --- a/include/linux/zswap.h +++ b/include/linux/zswap.h @@ -35,6 +35,7 @@ void zswap_lruvec_state_init(struct lruvec *lruvec); void zswap_folio_swapin(struct folio *folio); bool zswap_is_enabled(void); bool zswap_never_enabled(void); +bool zswap_is_present(swp_entry_t entry, unsigned int nr); #else =20 struct zswap_lruvec_state {}; @@ -69,6 +70,11 @@ static inline bool zswap_never_enabled(void) return true; } =20 +static inline bool zswap_is_present(swp_entry_t entry, unsigned int nr) +{ + return false; +} + #endif =20 #endif /* _LINUX_ZSWAP_H */ diff --git a/mm/swap_state.c b/mm/swap_state.c index b76eb3d876fd7..15e200d6966b9 100644 --- a/mm/swap_state.c +++ b/mm/swap_state.c @@ -12,6 +12,7 @@ #include #include #include +#include #include #include #include @@ -191,6 +192,15 @@ static int __swap_cache_add_check(struct swap_cluster_= info *ci, if (nr =3D=3D 1) return 0; =20 + /* + * The cluster lock serializes swap-cache insertion with zswap + * writeback. Reject mixed zswap/disk backing before allocating a + * large folio and recheck it before adding the folio to swap cache. + */ + if (zswap_is_present(swp_entry(swp_type(targ_entry), + round_down(swp_offset(targ_entry), nr)), nr)) + return -EBUSY; + is_zero =3D __swap_table_test_zero(ci, ci_off); ci_off =3D round_down(ci_off, nr); ci_end =3D ci_off + nr; diff --git a/mm/zswap.c b/mm/zswap.c index 37f34e406c8e3..32671dc2bf84d 100644 --- a/mm/zswap.c +++ b/mm/zswap.c @@ -1571,6 +1571,23 @@ bool zswap_store(struct folio *folio) return ret; } =20 +/** + * zswap_is_present() - is any slot in [entry, entry + nr) in zswap? + * @entry: base swap entry of the range + * @nr: number of contiguous slots to check (pass 1 for a single-slot quer= y) + */ +bool zswap_is_present(swp_entry_t entry, unsigned int nr) +{ + pgoff_t offset =3D swp_offset(entry); + struct xarray *tree =3D swap_zswap_tree(entry); + unsigned long index =3D offset; + + if (!nr || zswap_never_enabled()) + return false; + + return xa_find(tree, &index, offset + nr - 1, XA_PRESENT); +} + /** * zswap_load() - load a folio from zswap * @folio: folio to load @@ -1578,13 +1595,9 @@ bool zswap_store(struct folio *folio) * Return: 0 on success, with the folio unlocked and marked up-to-date, or= one * of the following error codes: * - * -EIO: if the swapped out content was in zswap, but could not be loaded - * into the page due to a decompression failure. The folio is unlocked, b= ut - * NOT marked up-to-date, so that an IO error is emitted (e.g. do_swap_pa= ge() - * will SIGBUS). - * - * -EINVAL: if the swapped out content was in zswap, but the page belongs - * to a large folio, which is not supported by zswap. The folio is unlock= ed, + * -EIO: if the swapped out content was in zswap but could not be handed + * back, either because decompression failed or because a slot in a + * large-folio range is unexpectedly still in zswap. The folio is unlocke= d, * but NOT marked up-to-date, so that an IO error is emitted (e.g. * do_swap_page() will SIGBUS). * @@ -1605,13 +1618,20 @@ int zswap_load(struct folio *folio) return -ENOENT; =20 /* - * Large folios should not be swapped in while zswap is being used, as - * they are not properly handled. Zswap does not properly load large - * folios, and a large folio may only be partially in zswap. + * A large folio reaches zswap_load() only when its whole range is + * expected to be on disk: PMD swap-entry consumers split before + * calling into PMD-order swapin whenever any slot is still in zswap. + * Confirm the range is entirely absent from zswap and return -ENOENT + * so the caller reads it from disk; if a slot is unexpectedly still in + * zswap, fail the read rather than return partially-initialized data. */ - if (WARN_ON_ONCE(folio_test_large(folio))) { - folio_unlock(folio); - return -EINVAL; + if (folio_test_large(folio)) { + if (WARN_ON_ONCE(zswap_is_present(swp, + folio_nr_pages(folio)))) { + folio_unlock(folio); + return -EIO; + } + return -ENOENT; } =20 entry =3D xa_load(tree, offset); --=20 2.53.0-Meta From nobody Mon Sep 28 19:34:00 2026 Received: from mta1.migadu.com (out-187.mta1.migadu.com [95.215.58.187]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CCD2F4749D1 for ; Tue, 18 Aug 2026 13:12:35 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=95.215.58.187 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787058757; cv=none; b=MAYUdSJzVVCLhDif1cjayECqCfJxOlv2AjnLSi23pyoHT3cLSSI8YDwJnI7ucIRIpkCuMx47SQYt3SMlm3NSkesbJh3SbpW88RqC34VpoYZA0KZ0+nXdqWvwNouRbsqIlCucO+g7oqdZGlkn5Qhv148baoEkiGAL/8uCvTYRB9Y= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787058757; c=relaxed/simple; bh=669qLI089dH17Rw7Kr9YX56S160m9VKovF8ZGZBWnlo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=O0yUfkxqmaapimiLnWxhyu4jm2zKavv4zqpuavkCfCYvIN0Vb9xDHVyL9iG4w+apcVPZBeVA7ZVDfYCU5jIswqaNvK8gfClLgzYZZaXz5R14Q7/bGeI5KBnlnXI3i9PRU/MhT3OqsDeGQo2B7/MCTBLa7xzVh6tNNjzVe3jq0BU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=oLdvZbYB; arc=none smtp.client-ip=95.215.58.187 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="oLdvZbYB" X-Envelope-To: linux-kernel@vger.kernel.org DKIM-Signature: a=rsa-sha256; bh=669qLI089dH17Rw7Kr9YX56S160m9VKovF8ZGZBWnlo=; c=simple/simple; d=linux.dev; h=from:to:subject:date:message-id:mime-version:content-type; s=key1; t=1787058753; v=1; x=1787663553; b=oLdvZbYBjeHUUhHBs4AsNXLNoGAVbMHDEFgg3PsPhLzE6sOmYPPyzlKoN57ZFwRbTC5E1b3y 1jmM44ii7akcjig3aFLrYysiknx+ll3AQXVHUILXyRsls4MKwxn+oaJCZ2bAfIzoeZFGI9RaTnj RkYm7cIhBHK/CEcen+rO7LAg= X-Envelope-To: linux-kernel@vger.kernel.org Received: from localhost (2a03:2880:10ff:d::) by mta12.migadu.com with ESMTPS id 0dd70e55678ba656; Tue, 18 Aug 2026 13:12:33 +0000 X-Migadu-Flow: FLOW_OUT From: Usama Arif To: Andrew Morton , david@kernel.org, chrisl@kernel.org, kasong@tencent.com, ljs@kernel.org, ziy@nvidia.com, linux-mm@kvack.org Cc: ying.huang@linux.alibaba.com, Baoquan He , willy@infradead.org, youngjun.park@lge.com, hannes@cmpxchg.org, riel@surriel.com, shakeel.butt@linux.dev, alex@ghiti.fr, kas@kernel.org, baohua@kernel.org, dev.jain@arm.com, baolin.wang@linux.alibaba.com, Nico Pache , Liam R. Howlett , ryan.roberts@arm.com, Vlastimil Babka , lance.yang@linux.dev, linux-kernel@vger.kernel.org, nphamcs@gmail.com, shikemeng@huaweicloud.com, yosry@kernel.org, kernel-team@meta.com, Usama Arif Subject: [PATCH v6 06/12] mm: swap in PMD swap entries as whole THPs during swapoff Date: Tue, 18 Aug 2026 06:09:47 -0700 Message-ID: <20260818131202.494754-7-usama.arif@linux.dev> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260818131202.494754-1-usama.arif@linux.dev> References: <20260818131202.494754-1-usama.arif@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add swap_pmd_cache_lookup() to classify the swap cache behind a PMD swap entry as empty, backed by one PMD-sized folio, or requiring per-page handling because at least one covered slot has a smaller folio in the swap cache. PMD swap entries are handled at PMD granularity only while the covered cache range is empty or backed by a PMD-sized folio; a split cache forces the entry to be split and retried through the PTE path. Add unuse_pmd() and call it from unuse_pmd_range() to swap in PMD-level swap entries as whole THPs during swapoff. This mirrors the existing unuse_pte_range() but operates at PMD granularity. Preserve soft-dirty, exclusive, and UFFD state when installing the present PMD. If an RWP VMA has a UFFD-marked swap entry, also restore PAGE_NONE so the first subsequent access still generates a userfault. If the PMD-order folio cannot be allocated or read, the swap cache already contains per-page folios in the covered range (e.g. split in the swap cache by deferred_split_scan() or memory_failure() while the PMD swap entry was installed), or any subpage is hardware-poisoned, the PMD swap entry is split into PTE-level entries via __split_huge_pmd() and a non-zero error is returned so unuse_pmd_range() falls through to unuse_pte_range(), which handles the individual entries at order-0. Remove a failed !uptodate PMD-sized folio from swap cache before splitting so PTE fallback rereads each slot independently. Keep hwpoisoned folios cached so PTE fallback can isolate bad subpages. Signed-off-by: Usama Arif --- mm/swap.h | 17 +++++ mm/swap_state.c | 44 +++++++++++++ mm/swapfile.c | 169 ++++++++++++++++++++++++++++++++++++++++++++++++ 3 files changed, 230 insertions(+) diff --git a/mm/swap.h b/mm/swap.h index e225527b2eb7a..441e017afc67b 100644 --- a/mm/swap.h +++ b/mm/swap.h @@ -311,6 +311,23 @@ static inline bool folio_matches_swap_entry(const stru= ct folio *folio, bool swap_cache_has_folio(swp_entry_t entry); struct folio *swap_cache_get_folio(swp_entry_t entry); void *swap_cache_get_shadow(swp_entry_t entry); +enum swap_pmd_cache { + SWAP_PMD_CACHE_EMPTY, + SWAP_PMD_CACHE_HUGE, + SWAP_PMD_CACHE_SPLIT, +}; + +#ifdef CONFIG_THP_SWAP +enum swap_pmd_cache swap_pmd_cache_lookup(swp_entry_t entry, + struct folio **foliop); +#else +static inline enum swap_pmd_cache swap_pmd_cache_lookup(swp_entry_t entry, + struct folio **foliop) +{ + *foliop =3D NULL; + return SWAP_PMD_CACHE_EMPTY; +} +#endif void swap_cache_del_folio(struct folio *folio); struct folio *swap_cache_alloc_folio(swp_entry_t target_entry, gfp_t gfp_m= ask, unsigned long orders, struct vm_fault *vmf, diff --git a/mm/swap_state.c b/mm/swap_state.c index 15e200d6966b9..559dc00b28f95 100644 --- a/mm/swap_state.c +++ b/mm/swap_state.c @@ -125,6 +125,50 @@ bool swap_cache_has_folio(swp_entry_t entry) return swp_tb_is_folio(swp_tb); } =20 +#ifdef CONFIG_THP_SWAP +/** + * swap_pmd_cache_lookup - classify the swap cache behind a PMD swap entry + * @entry: first swap slot encoded by the PMD swap entry + * @foliop: returned PMD-sized folio, with a reference, if present + * + * A PMD swap entry is a compact page-table encoding for HPAGE_PMD_NR + * consecutive swap slots. The swap cache behind those slots can be empty, + * one PMD-sized folio, or per-slot folios after the original folio was sp= lit. + * + * Context: Caller must keep @entry valid using the usual swap cache rules. + * Return: SWAP_PMD_CACHE_EMPTY if no slot in the PMD range has a cached f= olio, + * SWAP_PMD_CACHE_HUGE if one PMD-sized folio covers the range, or + * SWAP_PMD_CACHE_SPLIT if the range needs per-page handling. + */ +enum swap_pmd_cache swap_pmd_cache_lookup(swp_entry_t entry, + struct folio **foliop) +{ + unsigned int type =3D swp_type(entry); + pgoff_t offset =3D swp_offset(entry); + struct folio *folio; + int i; + + *foliop =3D NULL; + + folio =3D swap_cache_get_folio(entry); + if (folio) { + if (folio_nr_pages(folio) =3D=3D HPAGE_PMD_NR) { + *foliop =3D folio; + return SWAP_PMD_CACHE_HUGE; + } + folio_put(folio); + return SWAP_PMD_CACHE_SPLIT; + } + + for (i =3D 1; i < HPAGE_PMD_NR; i++) { + if (swap_cache_has_folio(swp_entry(type, offset + i))) + return SWAP_PMD_CACHE_SPLIT; + } + + return SWAP_PMD_CACHE_EMPTY; +} +#endif + /** * swap_cache_get_shadow - Looks up a shadow in the swap cache. * @entry: swap entry used for the lookup. diff --git a/mm/swapfile.c b/mm/swapfile.c index 44b9ebbe7229f..67a7e2053dc12 100644 --- a/mm/swapfile.c +++ b/mm/swapfile.c @@ -42,6 +42,7 @@ #include #include #include +#include =20 #include #include @@ -2670,6 +2671,160 @@ static int unuse_pte_range(struct vm_area_struct *v= ma, pmd_t *pmd, return 0; } =20 +#ifdef CONFIG_THP_SWAP +/* + * unuse_pmd - Map a locked folio at PMD granularity during swapoff. + * + * The caller provides a locked, swapped-in folio. Returns 0 on success + * (PMD was mapped). Returns -EAGAIN if the swap cache folio no longer + * matches the entry or the PMD changed under the lock (try_to_unuse will + * rescan). Returns -EIO if the folio is not uptodate or contains a poison= ed + * subpage; in that case the PMD is split so unuse_pte_range() can handle + * individual pages. + */ +static int unuse_pmd(struct vm_area_struct *vma, pmd_t *pmd, + unsigned long addr, softleaf_t entry, + struct folio *folio) +{ + struct mm_struct *mm =3D vma->vm_mm; + struct page *page; + pmd_t new_pmd, old_pmd; + spinlock_t *ptl; + rmap_t rmap_flags =3D RMAP_NONE; + bool exclusive; + + if (unlikely(!folio_matches_swap_entry(folio, entry))) + return -EAGAIN; + + if (unlikely(!folio_test_uptodate(folio))) { + /* Let PTE fallback reread each slot independently. */ + swap_cache_del_folio(folio); + __split_huge_pmd(vma, pmd, addr, false); + return -EIO; + } + + if (unlikely(folio_contain_hwpoisoned_page(folio))) { + /* Let PTE fallback isolate the poisoned subpages. */ + __split_huge_pmd(vma, pmd, addr, false); + return -EIO; + } + + page =3D folio_page(folio, 0); + + ptl =3D pmd_lock(mm, pmd); + old_pmd =3D pmdp_get(pmd); + + if (!pmd_is_swap_entry(old_pmd) || + softleaf_from_pmd(old_pmd).val !=3D entry.val) { + spin_unlock(ptl); + return -EAGAIN; + } + + exclusive =3D pmd_swp_exclusive(old_pmd); + + /* + * Some architectures may have to restore extra metadata to the folio + * when reading from swap. This metadata may be indexed by swap entry + * so this must be called before folio_put_swap(). + */ + arch_swap_restore(folio_swap(entry, folio), folio); + + add_mm_counter(mm, MM_ANONPAGES, HPAGE_PMD_NR); + add_mm_counter(mm, MM_SWAPENTS, -HPAGE_PMD_NR); + + new_pmd =3D folio_mk_pmd(folio, vma->vm_page_prot); + new_pmd =3D pmd_mkold(new_pmd); + if (pmd_swp_soft_dirty(old_pmd)) + new_pmd =3D pmd_mksoft_dirty(new_pmd); + if (pmd_swp_uffd(old_pmd)) + new_pmd =3D pmd_mkuffd(new_pmd); + if (pmd_swp_uffd(old_pmd) && userfaultfd_rwp(vma)) + new_pmd =3D pmd_modify(new_pmd, PAGE_NONE); + + if (exclusive) + rmap_flags |=3D RMAP_EXCLUSIVE; + + folio_get(folio); + if (!folio_test_anon(folio)) + folio_add_new_anon_rmap(folio, vma, addr, rmap_flags); + else + folio_add_anon_rmap_pmd(folio, page, vma, addr, rmap_flags); + + set_pmd_at(mm, addr, pmd, new_pmd); + folio_put_swap(folio, NULL); + + spin_unlock(ptl); + + folio_free_swap(folio); + return 0; +} + +/* + * Try to swap in a PMD swap entry as a whole THP. Returns 0 on success. + * If the swap cache no longer has one PMD-sized folio, zswap may require + * per-page loading, or a PMD-order allocation/read fails, split the PMD so + * the caller can fall back to unuse_pte_range(). Otherwise propagates the + * error from unuse_pmd(). + */ +static int unuse_pmd_entry(struct vm_area_struct *vma, pmd_t *pmd, + unsigned long addr, softleaf_t entry) +{ + struct folio *folio; + enum swap_pmd_cache cache_state; + int ret; + + cache_state =3D swap_pmd_cache_lookup(entry, &folio); + if (cache_state =3D=3D SWAP_PMD_CACHE_SPLIT) { + ret =3D -EAGAIN; + goto split_fallback; + } + if (!folio) { + struct vm_fault vmf =3D { + .vma =3D vma, + .address =3D addr, + .real_address =3D addr, + .pmd =3D pmd, + }; + + if (zswap_is_present(entry, HPAGE_PMD_NR)) { + ret =3D -EAGAIN; + goto split_fallback; + } + + folio =3D swapin_sync(entry, GFP_HIGHUSER_MOVABLE, + BIT(HPAGE_PMD_ORDER), &vmf, NULL, 0); + if (IS_ERR_OR_NULL(folio)) { + ret =3D folio ? PTR_ERR(folio) : -ENOMEM; + goto split_fallback; + } + } + + folio_lock(folio); + folio_wait_writeback(folio); + /* + * If the cached folio is no longer PMD-sized (e.g. split in the + * swap cache by deferred_split_scan() or memory_failure() while + * the PMD swap entry was installed), the PMD swap entry no longer + * maps a single contiguous folio. Split the PMD swap entry so + * unuse_pte_range() can swap the per-slot folios in individually. + */ + if (folio_nr_pages(folio) !=3D HPAGE_PMD_NR) { + folio_unlock(folio); + folio_put(folio); + ret =3D -EAGAIN; + goto split_fallback; + } + ret =3D unuse_pmd(vma, pmd, addr, entry, folio); + folio_unlock(folio); + folio_put(folio); + return ret; + +split_fallback: + __split_huge_pmd(vma, pmd, addr, false); + return ret; +} +#endif + static inline int unuse_pmd_range(struct vm_area_struct *vma, pud_t *pud, unsigned long addr, unsigned long end, unsigned int type) @@ -2682,6 +2837,20 @@ static inline int unuse_pmd_range(struct vm_area_str= uct *vma, pud_t *pud, do { cond_resched(); next =3D pmd_addr_end(addr, end); + +#ifdef CONFIG_THP_SWAP + pmd_t pmdval =3D pmdp_get(pmd); + + if (pmd_is_swap_entry(pmdval)) { + softleaf_t sl =3D softleaf_from_pmd(pmdval); + + if (swp_type(sl) =3D=3D type) { + if (!unuse_pmd_entry(vma, pmd, addr, sl)) + continue; + } + } +#endif + ret =3D unuse_pte_range(vma, pmd, addr, next, type); if (ret) return ret; --=20 2.53.0-Meta From nobody Mon Sep 28 19:34:00 2026 Received: from mta0.migadu.com (out-160.mta0.migadu.com [91.218.175.160]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 582804749CF for ; Tue, 18 Aug 2026 13:12:37 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.160 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787058760; cv=none; b=Feya8xg+F1uKL8NpwPBEKJetU40pGV3hMF63fa6y/CNAPgK/m77fROrSfF6nwY2QoON32RnXKP6wX7rDmFFWawAjaZM2rAb/5TTIA0kTE4QftIs5DYbPrwVCigFf3q0NYfxK07dlXuRp0PxD/x0E6WGcn+bhdm/xw4mxo/s3L2I= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787058760; c=relaxed/simple; bh=h6V8hB3omjRe0c7aiousJ/TjmZXePRZFl9ICyKKrXLM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=KYYtpL1wgB6f7CrAcgRMEA9arVDOH1QYCcNsmVDCR4d8arrtVM+XjeSYPj9Ml01R+7+nRc5QJ0TyEyG1HT8fA6uCdcRyQLF6RTAuqzbS6m7nS2wCDjsgrn9vmENbde9Vcv7e8S0zEHpSdEfH+uurjChrjpgnl4CYRI4I9DsFEog= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=Ksqyv3AZ; arc=none smtp.client-ip=91.218.175.160 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="Ksqyv3AZ" X-Envelope-To: linux-kernel@vger.kernel.org DKIM-Signature: a=rsa-sha256; bh=h6V8hB3omjRe0c7aiousJ/TjmZXePRZFl9ICyKKrXLM=; c=simple/simple; d=linux.dev; h=from:to:subject:date:message-id:mime-version:content-type; s=key1; t=1787058756; v=1; x=1787663556; b=Ksqyv3AZH6TJB5WMhf7ZJVTfmthuQVvWUjwB8Rv08X80FiBaacLe8U5XszODj7DpGVvGU4lu RyvrihUlmcO+mpPWpzPidD+ETPKPyiu8ZlYhpio+F4v4cd5GJbIzaYNJ1TDmvGqdkhkEnvZLGQe z0o67wxXvoTyi31jOSArywVA= X-Envelope-To: linux-kernel@vger.kernel.org Received: from localhost (2a03:2880:10ff:1d::) by mta12.migadu.com with ESMTPS id 0920b0673411bc84; Tue, 18 Aug 2026 13:12:35 +0000 X-Migadu-Flow: FLOW_OUT From: Usama Arif To: Andrew Morton , david@kernel.org, chrisl@kernel.org, kasong@tencent.com, ljs@kernel.org, ziy@nvidia.com, linux-mm@kvack.org Cc: ying.huang@linux.alibaba.com, Baoquan He , willy@infradead.org, youngjun.park@lge.com, hannes@cmpxchg.org, riel@surriel.com, shakeel.butt@linux.dev, alex@ghiti.fr, kas@kernel.org, baohua@kernel.org, dev.jain@arm.com, baolin.wang@linux.alibaba.com, Nico Pache , Liam R. Howlett , ryan.roberts@arm.com, Vlastimil Babka , lance.yang@linux.dev, linux-kernel@vger.kernel.org, nphamcs@gmail.com, shikemeng@huaweicloud.com, yosry@kernel.org, kernel-team@meta.com, Usama Arif Subject: [PATCH v6 07/12] mm: handle PMD swap entries in non-present PMD walkers Date: Tue, 18 Aug 2026 06:09:48 -0700 Message-ID: <20260818131202.494754-8-usama.arif@linux.dev> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260818131202.494754-1-usama.arif@linux.dev> References: <20260818131202.494754-1-usama.arif@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Teach the remaining non-present PMD walkers about swap entries, mirroring the PTE-level equivalents. smaps_pmd_entry() accounts swap and swap_pss via a new shared smaps_account_swap() helper used by both PTE and PMD paths. For a PMD range, it accounts each covered slot separately because their swap reference counts can differ. move_soft_dirty_pmd(), clear_soft_dirty_pmd(), and make_uffd_wp_pmd(), pagemap_pmd_range_thp() and change_huge_pmd() handle swap entries alongside migration entries. hmm_vma_handle_absent_pmd() records a fault for PMD swap entries via hmm_record_fault() instead of returning -EFAULT, allowing hmm_range_fault() to fault them in. The first per-page handle_mm_fault() call triggers do_huge_pmd_swap_page(), which maps the entire folio; subsequent calls become harmless huge_pmd_set_accessed() and the walker retries with a present PMD. When no fault is requested, it reports the range as non-resident rather than HMM_PFN_ERROR, matching the PTE swap entry path. madvise_free_huge_pmd() handles PMD swap entries directly: for a full-range MADV_FREE it clears the PMD, frees the deposited page table, and releases the swap slots; for a partial range it splits to PTE swap entries. Without this, MADV_FREE silently becomes a no-op on swapped-out THPs and leaves the swap slots allocated, unlike the PTE path which frees them. zap_huge_pmd() frees swap slots via swap_put_entries_direct(), matching zap_nonpresent_ptes(). change_non_present_huge_pmd() skips write-permission changes for swap entries and only updates uffd_wp, matching change_softleaf_pte(). madvise_cold_or_pageout_pte_range() skips PMD swap entries early. MADV_COLD and MADV_PAGEOUT operate on resident folios, so a swapped-out THP has nothing to deactivate or reclaim; skipping also prevents the walker from descending into or splitting the PMD swap entry. The locked THP path also treats a racing PMD swap entry as handled before checking for other non-present PMD types. mincore_pte_range() routes PMD swap entries through a new mincore_pmd_swap() helper, matching how the PTE path already calls mincore_swap() for non-present PTEs. Without this a swapped-out PMD-mapped THP would be reported as resident, because pmd_is_huge() (and therefore pmd_trans_huge_lock()) accepts any non-present non-none PMD and the old branch unconditionally did memset(vec, 1, nr). mincore_pmd_swap() checks the PMD-sized swap-cache folio, or the individual slots if it was split. Migration and device-private PMDs keep their previous behavior of reporting 1. check_pmd_state() in khugepaged returns SCAN_PMD_MAPPED for PMD swap entries, treating a swapped-out THP as still being a THP from khugepaged's perspective and matching the existing migration-entry handling. change_huge_pmd() and pagemap_pmd_range_thp() drop redundant thp_migration_supported() gates: when PMD softleaves are unsupported, softleaf_from_pmd() returns a none entry and pmd_is_valid_softleaf() is false. Signed-off-by: Usama Arif --- fs/proc/task_mmu.c | 46 ++++++++++++++++++++++++++-------------- mm/hmm.c | 11 +++++++++- mm/huge_memory.c | 53 ++++++++++++++++++++++++++++++++++++---------- mm/khugepaged.c | 6 ++++++ mm/madvise.c | 14 +++++++++++- mm/mincore.c | 45 ++++++++++++++++++++++++++++++++++++++- 6 files changed, 145 insertions(+), 30 deletions(-) diff --git a/fs/proc/task_mmu.c b/fs/proc/task_mmu.c index 5c54aebe21182..8926392e33338 100644 --- a/fs/proc/task_mmu.c +++ b/fs/proc/task_mmu.c @@ -1046,6 +1046,27 @@ static void smaps_pte_hole_lookup(unsigned long addr= , struct mm_walk *walk) #endif } =20 +static void smaps_account_swap(struct mem_size_stats *mss, + softleaf_t entry, unsigned long size) +{ + unsigned long nr_pages =3D size >> PAGE_SHIFT; + + mss->swap +=3D size; + do { + int mapcount =3D swp_swapcount(entry); + + if (mapcount >=3D 2) { + u64 pss_delta =3D (u64)PAGE_SIZE << PSS_SHIFT; + + do_div(pss_delta, mapcount); + mss->swap_pss +=3D pss_delta; + } else { + mss->swap_pss +=3D (u64)PAGE_SIZE << PSS_SHIFT; + } + entry.val++; + } while (--nr_pages); +} + static void smaps_pte_entry(pte_t *pte, unsigned long addr, struct mm_walk *walk) { @@ -1067,18 +1088,7 @@ static void smaps_pte_entry(pte_t *pte, unsigned lon= g addr, const softleaf_t entry =3D softleaf_from_pte(ptent); =20 if (softleaf_is_swap(entry)) { - int mapcount; - - mss->swap +=3D PAGE_SIZE; - mapcount =3D swp_swapcount(entry); - if (mapcount >=3D 2) { - u64 pss_delta =3D (u64)PAGE_SIZE << PSS_SHIFT; - - do_div(pss_delta, mapcount); - mss->swap_pss +=3D pss_delta; - } else { - mss->swap_pss +=3D (u64)PAGE_SIZE << PSS_SHIFT; - } + smaps_account_swap(mss, entry, PAGE_SIZE); } else if (softleaf_has_pfn(entry)) { if (softleaf_is_device_private(entry)) present =3D true; @@ -1108,9 +1118,13 @@ static void smaps_pmd_entry(pmd_t *pmd, unsigned lon= g addr, if (pmd_present(*pmd)) { page =3D vm_normal_page_pmd(vma, addr, *pmd); present =3D true; - } else if (unlikely(thp_migration_supported())) { + } else { const softleaf_t entry =3D softleaf_from_pmd(*pmd); =20 + if (softleaf_is_swap(entry)) { + smaps_account_swap(mss, entry, HPAGE_PMD_SIZE); + return; + } if (softleaf_has_pfn(entry)) page =3D softleaf_to_page(entry); } @@ -1755,7 +1769,7 @@ static inline void clear_soft_dirty_pmd(struct vm_are= a_struct *vma, pmd =3D pmd_clear_soft_dirty(pmd); =20 set_pmd_at(vma->vm_mm, addr, pmdp, pmd); - } else if (pmd_is_migration_entry(pmd)) { + } else if (pmd_is_migration_entry(pmd) || pmd_is_swap_entry(pmd)) { pmd =3D pmd_swp_clear_soft_dirty(pmd); set_pmd_at(vma->vm_mm, addr, pmdp, pmd); } @@ -2115,7 +2129,7 @@ static int pagemap_pmd_range_thp(pmd_t *pmdp, unsigne= d long addr, flags |=3D PM_UFFD_WP; if (pm->show_pfn) frame =3D pmd_pfn(pmd) + idx; - } else if (thp_migration_supported()) { + } else if (pmd_is_valid_softleaf(pmd)) { const softleaf_t entry =3D softleaf_from_pmd(pmd); unsigned long offset; =20 @@ -2581,7 +2595,7 @@ static void make_uffd_wp_pmd(struct vm_area_struct *v= ma, old =3D pmdp_invalidate_ad(vma, addr, pmdp); pmd =3D pmd_mkuffd(old); set_pmd_at(vma->vm_mm, addr, pmdp, pmd); - } else if (pmd_is_migration_entry(pmd)) { + } else if (pmd_is_migration_entry(pmd) || pmd_is_swap_entry(pmd)) { pmd =3D pmd_swp_mkuffd(pmd); set_pmd_at(vma->vm_mm, addr, pmdp, pmd); } diff --git a/mm/hmm.c b/mm/hmm.c index 2f1e98c6b6440..95575ac378888 100644 --- a/mm/hmm.c +++ b/mm/hmm.c @@ -377,12 +377,21 @@ static int hmm_vma_handle_absent_pmd(struct mm_walk *= walk, unsigned long start, required_fault =3D hmm_range_need_fault(hmm_vma_walk, hmm_pfns, npages, 0); if (required_fault) { - if (softleaf_is_device_private(entry)) + if (softleaf_is_device_private(entry) || + softleaf_is_swap(entry)) return hmm_record_fault(addr, end, required_fault, walk); else return -EFAULT; } =20 + /* + * A swapped-out THP is not resident. Report it as not-valid, + * matching what hmm_vma_handle_pte() does for a PTE swap entry when + * no fault was requested. + */ + if (softleaf_is_swap(entry)) + return hmm_pfns_fill(start, end, range, 0); + return hmm_pfns_fill(start, end, range, HMM_PFN_ERROR); } #else diff --git a/mm/huge_memory.c b/mm/huge_memory.c index 2735c3de7029c..54aef8394fa25 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -2397,6 +2397,14 @@ vm_fault_t do_huge_pmd_numa_page(struct vm_fault *vm= f) return 0; } =20 +static inline void zap_deposited_table(struct mm_struct *mm, pmd_t *pmd) +{ + pgtable_t pgtable; + + pgtable =3D pgtable_trans_huge_withdraw(mm, pmd); + pte_free(mm, pgtable); + mm_dec_nr_ptes(mm); +} /* * Return true if we do MADV_FREE successfully on entire pmd page. * Otherwise, return false. @@ -2421,6 +2429,21 @@ bool madvise_free_huge_pmd(struct mmu_gather *tlb, s= truct vm_area_struct *vma, goto out; =20 if (unlikely(!pmd_present(orig_pmd))) { + if (pmd_is_swap_entry(orig_pmd)) { + if (next - addr !=3D HPAGE_PMD_SIZE) { + spin_unlock(ptl); + __split_huge_pmd(vma, pmd, addr, false); + goto out_unlocked; + } + softleaf_t sl =3D softleaf_from_pmd(orig_pmd); + + pmdp_huge_get_and_clear(mm, addr, pmd); + zap_deposited_table(mm, pmd); + spin_unlock(ptl); + swap_put_entries_direct(sl, HPAGE_PMD_NR); + add_mm_counter(mm, MM_SWAPENTS, -HPAGE_PMD_NR); + return true; + } VM_WARN_ON_ONCE(!pmd_is_migration_entry(orig_pmd) && !pmd_is_device_private_entry(orig_pmd)); goto out; @@ -2471,15 +2494,6 @@ bool madvise_free_huge_pmd(struct mmu_gather *tlb, s= truct vm_area_struct *vma, return ret; } =20 -static inline void zap_deposited_table(struct mm_struct *mm, pmd_t *pmd) -{ - pgtable_t pgtable; - - pgtable =3D pgtable_trans_huge_withdraw(mm, pmd); - pte_free(mm, pgtable); - mm_dec_nr_ptes(mm); -} - static void zap_huge_pmd_folio(struct mm_struct *mm, struct vm_area_struct= *vma, pmd_t pmdval, struct folio *folio, bool is_present) { @@ -2572,6 +2586,16 @@ bool zap_huge_pmd(struct mmu_gather *tlb, struct vm_= area_struct *vma, arch_check_zapped_pmd(vma, orig_pmd); tlb_remove_pmd_tlb_entry(tlb, pmd, addr); =20 + if (pmd_is_swap_entry(orig_pmd)) { + softleaf_t sl =3D softleaf_from_pmd(orig_pmd); + + zap_deposited_table(mm, pmd); + spin_unlock(ptl); + swap_put_entries_direct(sl, HPAGE_PMD_NR); + add_mm_counter(mm, MM_SWAPENTS, -HPAGE_PMD_NR); + return true; + } + is_present =3D pmd_present(orig_pmd); folio =3D normal_or_softleaf_folio_pmd(vma, addr, orig_pmd, is_present); has_deposit =3D has_deposited_pgtable(vma, orig_pmd, folio); @@ -2604,7 +2628,8 @@ static inline int pmd_move_must_withdraw(spinlock_t *= new_pmd_ptl, static pmd_t move_soft_dirty_pmd(pmd_t pmd) { if (pgtable_supports_soft_dirty()) { - if (unlikely(pmd_is_migration_entry(pmd))) + if (unlikely(pmd_is_migration_entry(pmd) || + pmd_is_swap_entry(pmd))) pmd =3D pmd_swp_mksoft_dirty(pmd); else if (pmd_present(pmd)) pmd =3D pmd_mksoft_dirty(pmd); @@ -2695,6 +2720,12 @@ static void change_non_present_huge_pmd(struct mm_st= ruct *mm, pmd_t newpmd; =20 VM_WARN_ON(!pmd_is_valid_softleaf(*pmd)); + + /* + * Note that a PMD swap entry falls into the default branch below: it + * does not encode write permission in the entry type, so only the + * uffd_wp flag update at the end applies to it. + */ if (softleaf_is_migration_write(entry)) { const struct folio *folio =3D softleaf_to_folio(entry); =20 @@ -2755,7 +2786,7 @@ int change_huge_pmd(struct mmu_gather *tlb, struct vm= _area_struct *vma, if (!ptl) return 0; =20 - if (thp_migration_supported() && pmd_is_valid_softleaf(*pmd)) { + if (pmd_is_valid_softleaf(*pmd)) { change_non_present_huge_pmd(mm, addr, pmd, uffd_prot, uffd_prot_resolve); goto unlock; diff --git a/mm/khugepaged.c b/mm/khugepaged.c index 5a06e3942e889..15e2d0c7384d3 100644 --- a/mm/khugepaged.c +++ b/mm/khugepaged.c @@ -1125,6 +1125,12 @@ static inline enum scan_result check_pmd_state(pmd_t= *pmd) */ if (pmd_is_migration_entry(pmde)) return SCAN_PMD_MAPPED; + /* + * A PMD-mapped THP that has been swapped out is still a THP from + * khugepaged's perspective; treat it like a present huge PMD. + */ + if (pmd_is_swap_entry(pmde)) + return SCAN_PMD_MAPPED; if (!pmd_present(pmde)) return SCAN_NO_PTE_TABLE; if (pmd_trans_huge(pmde)) diff --git a/mm/madvise.c b/mm/madvise.c index c179938097bf0..16b39a06b038f 100644 --- a/mm/madvise.c +++ b/mm/madvise.c @@ -375,6 +375,15 @@ static int madvise_cold_or_pageout_pte_range(pmd_t *pm= d, !can_do_file_pageout(vma); =20 #ifdef CONFIG_TRANSPARENT_HUGEPAGE + /* + * Swapped-out THPs have no resident folio to deactivate or reclaim. + * Avoid descending into or splitting a PMD swap entry. + */ + if (pmd_is_swap_entry(*pmd)) { + walk->action =3D ACTION_CONTINUE; + return 0; + } + if (pmd_trans_huge(*pmd)) { pmd_t orig_pmd; unsigned long next =3D pmd_addr_end(addr, end); @@ -385,6 +394,9 @@ static int madvise_cold_or_pageout_pte_range(pmd_t *pmd, return 0; =20 orig_pmd =3D *pmd; + if (pmd_is_swap_entry(orig_pmd)) + goto huge_unlock; + if (is_huge_zero_pmd(orig_pmd)) goto huge_unlock; =20 @@ -666,7 +678,7 @@ static int madvise_free_pte_range(pmd_t *pmd, unsigned = long addr, int nr, max_nr; =20 next =3D pmd_addr_end(addr, end); - if (pmd_trans_huge(*pmd)) + if (pmd_trans_huge(*pmd) || pmd_is_swap_entry(*pmd)) if (madvise_free_huge_pmd(tlb, vma, pmd, addr, next)) return 0; =20 diff --git a/mm/mincore.c b/mm/mincore.c index ff4ac82817683..3f0fba964c8a3 100644 --- a/mm/mincore.c +++ b/mm/mincore.c @@ -85,6 +85,41 @@ static unsigned char mincore_swap(swp_entry_t entry, boo= l shmem) return present; } =20 +#ifdef CONFIG_THP_SWAP +static void mincore_pmd_swap(swp_entry_t entry, unsigned long addr, + unsigned long end, unsigned char *vec) +{ + unsigned long haddr =3D addr & HPAGE_PMD_MASK; + unsigned long start =3D (addr - haddr) >> PAGE_SHIFT; + unsigned long nr =3D (end - addr) >> PAGE_SHIFT; + struct folio *folio; + enum swap_pmd_cache state; + int i; + + state =3D swap_pmd_cache_lookup(entry, &folio); + if (state =3D=3D SWAP_PMD_CACHE_HUGE) { + memset(vec, folio_test_uptodate(folio), nr); + folio_put(folio); + return; + } + + if (state =3D=3D SWAP_PMD_CACHE_EMPTY) { + memset(vec, 0, nr); + return; + } + + /* + * The PMD swap entry is only a compact encoding for consecutive swap + * slots. If the PMD-sized swapcache folio was split, report residency + * from the individual slots covered by this mincore() range. + */ + for (i =3D 0; i < nr; i++) + vec[i] =3D mincore_swap(swp_entry(swp_type(entry), + swp_offset(entry) + start + i), + false); +} +#endif + /* * Later we can get more picky about what "in core" means precisely. * For now, simply check to see if the page is in the page cache, @@ -171,7 +206,15 @@ static int mincore_pte_range(pmd_t *pmd, unsigned long= addr, unsigned long end, =20 ptl =3D pmd_trans_huge_lock(pmd, vma); if (ptl) { - memset(vec, 1, nr); + if (pmd_is_swap_entry(*pmd)) { +#ifdef CONFIG_THP_SWAP + mincore_pmd_swap(softleaf_from_pmd(*pmd), addr, end, vec); +#else + memset(vec, 0, nr); +#endif + } else { + memset(vec, 1, nr); + } spin_unlock(ptl); goto out; } --=20 2.53.0-Meta From nobody Mon Sep 28 19:34:00 2026 Received: from mta1.migadu.com (out-215.mta1.migadu.com [95.215.58.215]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1457445FFA2 for ; Tue, 18 Aug 2026 13:12:44 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=95.215.58.215 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787058766; cv=none; b=lFcKY32bKSngIit9AWE1sBJP+2k7vimg6wqpA8+F2huhLYU/l/+MVhN0ygoulPDmAGqihHSdgBYzzqCjdAz1eT3Rw4An/QNKM0cg8TdPevbSf4YipowQ+8ODj69DSp8655DKKW5cQ2dBjOUhOpxqc812TAbM0hSEfQc3l4ar8rs= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787058766; c=relaxed/simple; bh=aLAdrmSbqCsLxg33maT0h/TBnMjvDOsyYsDfAadsCew=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=n44kphkAAZDfP/fDyipG9S8Qus/yQEeU5a/+eGrbsS5Yle0NjhEg9nt1X0DlgmpesTwwpQ+yo7zst2h4yBL1sLMVzqZxx/a2ZIwK7aL/3s/3d8yeZ8D+A+Oc2UUNoYktUK7eyUBb9uG6YA23MgtbN0BDwGfeIV0B8cmHY3hkpXM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=qCrOxkVA; arc=none smtp.client-ip=95.215.58.215 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="qCrOxkVA" X-Envelope-To: linux-kernel@vger.kernel.org DKIM-Signature: a=rsa-sha256; bh=aLAdrmSbqCsLxg33maT0h/TBnMjvDOsyYsDfAadsCew=; c=simple/simple; d=linux.dev; h=from:to:subject:date:message-id:mime-version:content-type; s=key1; t=1787058763; v=1; x=1787663563; b=qCrOxkVAt/HY/nGLXJPlGWqdDAfBrskbY8w3O2kCahluie7bg+PPTXJjc5KX36n5I7yK43Q5 q81Ur+GVM0h0Sha1oG/+iQtYgkle6h1gWQ33ehf/rxmZ9gmN13hdfk4LCDqJIXPN9SGTR/EotjN ZGdecgCNrQDF8vMv7yjN1YXc= X-Envelope-To: linux-kernel@vger.kernel.org Received: from localhost (2a03:2880:10ff:44::) by mta11.migadu.com with ESMTPS id 3c6aca1e0754bc79; Tue, 18 Aug 2026 13:12:43 +0000 X-Migadu-Flow: FLOW_OUT From: Usama Arif To: Andrew Morton , david@kernel.org, chrisl@kernel.org, kasong@tencent.com, ljs@kernel.org, ziy@nvidia.com, linux-mm@kvack.org Cc: ying.huang@linux.alibaba.com, Baoquan He , willy@infradead.org, youngjun.park@lge.com, hannes@cmpxchg.org, riel@surriel.com, shakeel.butt@linux.dev, alex@ghiti.fr, kas@kernel.org, baohua@kernel.org, dev.jain@arm.com, baolin.wang@linux.alibaba.com, Nico Pache , Liam R. Howlett , ryan.roberts@arm.com, Vlastimil Babka , lance.yang@linux.dev, linux-kernel@vger.kernel.org, nphamcs@gmail.com, shikemeng@huaweicloud.com, yosry@kernel.org, kernel-team@meta.com, Usama Arif Subject: [PATCH v6 08/12] mm: handle PMD swap entries in MADV_WILLNEED Date: Tue, 18 Aug 2026 06:09:49 -0700 Message-ID: <20260818131202.494754-9-usama.arif@linux.dev> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260818131202.494754-1-usama.arif@linux.dev> References: <20260818131202.494754-1-usama.arif@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" swapin_walk_pmd_entry() walks PTEs and skips non-present PMDs, so MADV_WILLNEED is a no-op on a PMD swap entry. Handle PMD swap entries under pmd_trans_huge_lock(). If the covered swap-cache range already has a PMD-sized folio, there is nothing left to prefetch. If the range has split cache state, or any covered slot currently has a zswap entry, split the PMD swap entry and ask the walker to retry so the PTE path can handle the individual slots. Otherwise pin the swap device and read the folio in at PMD order via swapin_sync(BIT(HPAGE_PMD_ORDER)). This keeps the subsequent fault on the do_huge_pmd_swap_page() path and avoids order-0 readahead needlessly splitting the PMD swap entry. Any failure of the PMD-order swapin splits the entry and retries through the PTE path. That covers losing a race with per-slot swap-cache population (-EBUSY) after dropping the PMD lock, but also the -ENOMEM that a PMD-order allocation can easily hit: leaving the entry alone would make MADV_WILLNEED prefetch nothing at all for the range, while the PTE path can still read the 512 slots at order 0. If per-page zswap state reappears during the read, remove the failed clean PMD-sized folio from swap cache before splitting so the PTE path can load each slot. This uses folio_trylock(): the lock is only free once the read has completed, so MADV_WILLNEED never blocks on in-flight I/O, and unlike testing folio_test_locked() directly it cannot race with an unrelated lock holder. Signed-off-by: Usama Arif --- mm/madvise.c | 106 +++++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 106 insertions(+) diff --git a/mm/madvise.c b/mm/madvise.c index 16b39a06b038f..3ef1af1e76daa 100644 --- a/mm/madvise.c +++ b/mm/madvise.c @@ -33,6 +33,7 @@ #include #include #include +#include =20 #include =20 @@ -185,6 +186,93 @@ static int madvise_update_vma(vm_flags_t new_flags, } =20 #ifdef CONFIG_SWAP +/* + * Prefetch a whole PMD swap entry as one PMD-order folio. + * + * Called with the PMD lock held; always drops it. Returns true when the + * caller should ask the walker to retry so the PTE path can handle the + * covered slots individually. + */ +static bool swapin_pmd_swap_entry(struct vm_area_struct *vma, pmd_t *pmd, + unsigned long addr, softleaf_t entry, + spinlock_t *ptl) +{ + struct vm_fault vmf =3D { + .vma =3D vma, + .address =3D addr, + .real_address =3D addr, + .pmd =3D pmd, + }; + enum swap_pmd_cache cache_state; + struct swap_info_struct *si; + struct folio *folio; + bool split =3D false; + + cache_state =3D swap_pmd_cache_lookup(entry, &folio); + if (cache_state =3D=3D SWAP_PMD_CACHE_HUGE) { + /* Already cached as one PMD-sized folio, nothing to do. */ + folio_put(folio); + spin_unlock(ptl); + return false; + } + if (cache_state =3D=3D SWAP_PMD_CACHE_SPLIT || + zswap_is_present(entry, HPAGE_PMD_NR)) { + spin_unlock(ptl); + return true; + } + + /* + * Pin the swap device under the PMD lock so the PMD-swap-entry + * observation keeps the entry valid for swapin_sync(). + */ + si =3D get_swap_device(entry); + spin_unlock(ptl); + if (!si) + return false; + + folio =3D swapin_sync(entry, GFP_HIGHUSER_MOVABLE, BIT(HPAGE_PMD_ORDER), + &vmf, NULL, 0); + + /* + * Fall back to PTE-order swapin: a PMD-order failure does not mean + * that individual slots cannot be read. + */ + if (IS_ERR_OR_NULL(folio)) { + split =3D true; + goto out; + } + + if (folio_nr_pages(folio) !=3D HPAGE_PMD_NR) { + split =3D true; + goto out_put; + } + + /* + * A trylock only succeeds once the read has completed, so this never + * blocks MADV_WILLNEED on in-flight I/O. A failed PMD-order zswap load + * leaves the folio clean and not uptodate; drop it from the swap cache + * so the PTE retry can load the per-page state. Another thread may + * have removed it already, so revalidate the association first. + */ + if (!folio_trylock(folio)) + goto out_put; + + if (!folio_test_uptodate(folio) && + zswap_is_present(entry, HPAGE_PMD_NR)) { + if (folio_matches_swap_entry(folio, entry)) + swap_cache_del_folio(folio); + split =3D true; + } + folio_unlock(folio); + +out_put: + folio_put(folio); +out: + /* Keep the device pinned until the last use of @entry. */ + put_swap_device(si); + return split; +} + static int swapin_walk_pmd_entry(pmd_t *pmd, unsigned long start, unsigned long end, struct mm_walk *walk) { @@ -194,6 +282,23 @@ static int swapin_walk_pmd_entry(pmd_t *pmd, unsigned = long start, spinlock_t *ptl; unsigned long addr; =20 + ptl =3D pmd_trans_huge_lock(pmd, vma); + if (ptl) { + pmd_t pmdval =3D *pmd; + + if (pmd_is_swap_entry(pmdval)) { + /* swapin_pmd_swap_entry() always drops the PMD lock. */ + if (swapin_pmd_swap_entry(vma, pmd, start, + softleaf_from_pmd(pmdval), + ptl)) { + __split_huge_pmd(vma, pmd, start, false); + walk->action =3D ACTION_AGAIN; + } + goto ret; + } + spin_unlock(ptl); + } + for (addr =3D start; addr < end; addr +=3D PAGE_SIZE) { pte_t pte; softleaf_t entry; @@ -222,6 +327,7 @@ static int swapin_walk_pmd_entry(pmd_t *pmd, unsigned l= ong start, if (ptep) pte_unmap_unlock(ptep, ptl); swap_read_submit(&ctx); +ret: cond_resched(); =20 return 0; --=20 2.53.0-Meta From nobody Mon Sep 28 19:34:00 2026 Received: from mta1.migadu.com (out-230.mta1.migadu.com [95.215.58.230]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 460AB45FFA2 for ; Tue, 18 Aug 2026 13:12:51 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=95.215.58.230 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787058773; cv=none; b=T7yVrnqYwqcYW7BnzVc7leLT3wbJYymueajp2BvZxdJpHAF9VccYKUEdx+LRt9x2qEAbCVP+Iz2wYx7+G0pZIzPlAZN8Bx7NhjBHR+/2SoTRuQ/JRn7u0lYlqUCv2MCKivJOsxw3pZjnU+GzTQPgfNQpiyzoJVFm9HBbSwdU1u4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787058773; c=relaxed/simple; bh=283hSMUTf8gUQjFzslSC5pHDKwFP9cHOoLBS8rXTLtE=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=n5iNj3h5+hA/ZmiwdPH+vNbbENCHJcX7RisNgysBMSqwl5OEeSLjoSsdOrg0tfAO0bErIkhVY1CjXckYzmxWpm8ZFqi5o/yGVud0YpuG3+t5xAGu7dlytt7z/FUKxVsgyGY4KtMJz81Iogz70UUUg2JvvoqAlumC4Uq5FitnlQQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=UJA8fo0N; arc=none smtp.client-ip=95.215.58.230 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="UJA8fo0N" X-Envelope-To: linux-kernel@vger.kernel.org DKIM-Signature: a=rsa-sha256; bh=283hSMUTf8gUQjFzslSC5pHDKwFP9cHOoLBS8rXTLtE=; c=simple/simple; d=linux.dev; h=from:to:subject:date:message-id:mime-version:content-type; s=key1; t=1787058769; v=1; x=1787663569; b=UJA8fo0NMOR69hKHQibRS33CHzOeSWXnP30WedXLCzFu6kzymAsStkn6775HlO+POsYw9NL1 91km0re/pHAnCWeOwOw/OgGfyF/fA6otJuUaaZdIMs/9iUO5nR0yq3Cct8n4fnv1eO7Z3cFn6QE 0ySoVNMk9IIsDNeIyQJFjyPM= X-Envelope-To: linux-kernel@vger.kernel.org Received: from localhost (2a03:2880:10ff:28::) by mta11.migadu.com with ESMTPS id 7b774479dcad1c3c; Tue, 18 Aug 2026 13:12:49 +0000 X-Migadu-Flow: FLOW_OUT From: Usama Arif To: Andrew Morton , david@kernel.org, chrisl@kernel.org, kasong@tencent.com, ljs@kernel.org, ziy@nvidia.com, linux-mm@kvack.org Cc: ying.huang@linux.alibaba.com, Baoquan He , willy@infradead.org, youngjun.park@lge.com, hannes@cmpxchg.org, riel@surriel.com, shakeel.butt@linux.dev, alex@ghiti.fr, kas@kernel.org, baohua@kernel.org, dev.jain@arm.com, baolin.wang@linux.alibaba.com, Nico Pache , Liam R. Howlett , ryan.roberts@arm.com, Vlastimil Babka , lance.yang@linux.dev, linux-kernel@vger.kernel.org, nphamcs@gmail.com, shikemeng@huaweicloud.com, yosry@kernel.org, kernel-team@meta.com, Usama Arif Subject: [PATCH v6 09/12] mm: handle PMD swap entries in UFFDIO_MOVE Date: Tue, 18 Aug 2026 06:09:50 -0700 Message-ID: <20260818131202.494754-10-usama.arif@linux.dev> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260818131202.494754-1-usama.arif@linux.dev> References: <20260818131202.494754-1-usama.arif@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" move_pages_huge_pmd() returned -ENOENT for any non-trans_huge, non-migration PMD, which fails aligned UFFDIO_MOVE on a swapped-out THP -- the PMD swap entry is a perfectly valid mapping that should move whole. Splitting via the move_pages_ptes() fallback isn't a substitute either: __split_huge_pmd_locked() splits a PMD swap entry into HPAGE_PMD_NR PTE swap entries pointing at the same swap-cache folio, but move_pages_pte() refuses any swap-cache folio that is still large and returns -EBUSY. Add move_swap_pmd(), modeled on move_swap_pte(), that moves the swap entry whole-PMD and re-anchors a PMD-sized swap-cache folio's anon rmap to the destination VMA. Reject !pmd_swp_exclusive() entries with -EBUSY to preserve UFFDIO_MOVE's single-owner semantics, propagate soft-dirty, arm the UFFD marker for an RWP-registered destination, and carry the deposited page table across with the entry. The marker handling matches move_swap_pte(): the source marker rides along with the entry and is additionally armed for an RWP destination. The dispatcher in move_pages_huge_pmd() routes PMD swap entries through move_swap_pmd() after pinning the swap device and arming an mmu_notifier range. Both are guarded by CONFIG_THP_SWAP, since swap_pmd_cache_lookup() and friends only exist under CONFIG_SWAP and PMD swap entries cannot exist without CONFIG_THP_SWAP. Before moving, classify the whole PMD swap-cache range with swap_pmd_cache_lookup(). A PMD swap entry can be moved whole only if the covered range is empty or backed by one PMD-sized folio. If the range already has per-slot cache state, split the PMD swap entry and return -EAGAIN so the caller retries through the PTE path. If a PMD-sized folio is cached, lock and revalidate that it still matches the PMD swap entry. If no folio is cached, recheck all HPAGE_PMD_NR slots under both PMD locks before moving the entry; any per-slot folio that appears needs the PTE move path to update its rmap metadata. This avoids moving the PMD while cached folios still point at the old anon_vma/index. Finally, reject a PMD swap entry at the *destination* with -EEXIST in move_pages(). Such a destination is not a hole, and unlike a PMD migration entry it does not resolve on its own: pte_alloc() skips a !pmd_none PMD, pte_offset_map_rw_nolock() then fails on the non-present PMD, and move_pages() would retry the resulting -EAGAIN forever, only escapable with a fatal signal. Signed-off-by: Usama Arif --- mm/huge_memory.c | 140 ++++++++++++++++++++++++++++++++++++++++++++++- mm/userfaultfd.c | 14 +++++ 2 files changed, 153 insertions(+), 1 deletion(-) diff --git a/mm/huge_memory.c b/mm/huge_memory.c index 54aef8394fa25..0076d206c6a8a 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -2913,6 +2913,77 @@ int change_huge_pud(struct mmu_gather *tlb, struct v= m_area_struct *vma, #endif =20 #ifdef CONFIG_USERFAULTFD +#ifdef CONFIG_THP_SWAP +/* + * Move a PMD-level swap entry from src_pmd to dst_pmd. Both PMD locks are + * acquired here; src_folio (if present) must already be locked. The depos= ited + * page table backing the source THP is moved across with the entry. + */ +static int move_swap_pmd(struct mm_struct *mm, struct vm_area_struct *dst_= vma, + unsigned long dst_addr, unsigned long src_addr, + pmd_t *dst_pmd, pmd_t *src_pmd, + pmd_t orig_dst_pmd, pmd_t orig_src_pmd, + spinlock_t *dst_ptl, spinlock_t *src_ptl, + struct folio *src_folio, swp_entry_t entry) +{ + pgtable_t src_pgtable; + pmd_t moved_pmd; + + /* + * The folio may have been freed and reused for a different swap entry + * while it was unlocked. Re-verify the association. + */ + if (src_folio && unlikely(!folio_matches_swap_entry(src_folio, entry) || + folio_nr_pages(src_folio) !=3D HPAGE_PMD_NR)) + return -EAGAIN; + + double_pt_lock(dst_ptl, src_ptl); + + if (!pmd_same(*src_pmd, orig_src_pmd) || + !pmd_same(*dst_pmd, orig_dst_pmd)) { + double_pt_unlock(dst_ptl, src_ptl); + return -EAGAIN; + } + + /* + * If the folio is in the swap cache, re-anchor its anon rmap to the + * destination VMA so a future swap-in fault at dst_addr finds it. + * Otherwise, re-check the whole PMD swap range: a PMD swap entry is + * only a compact encoding for 512 swap slots, and any per-slot cached + * folio would need the PTE move path to update its rmap metadata. + */ + if (src_folio) { + folio_move_anon_rmap(src_folio, dst_vma); + src_folio->index =3D linear_page_index(dst_vma, dst_addr); + } else { + unsigned int type =3D swp_type(entry); + pgoff_t offset =3D swp_offset(entry); + int i; + + for (i =3D 0; i < HPAGE_PMD_NR; i++) { + if (swap_cache_has_folio(swp_entry(type, offset + i))) { + double_pt_unlock(dst_ptl, src_ptl); + return -EAGAIN; + } + } + } + + moved_pmd =3D pmdp_huge_get_and_clear(mm, src_addr, src_pmd); + if (pgtable_supports_soft_dirty()) + moved_pmd =3D pmd_swp_mksoft_dirty(moved_pmd); + /* Re-arm RWP on the moved swap entry if dst_vma is RWP-registered. */ + if (userfaultfd_rwp(dst_vma)) + moved_pmd =3D pmd_swp_mkuffd(moved_pmd); + set_pmd_at(mm, dst_addr, dst_pmd, moved_pmd); + + src_pgtable =3D pgtable_trans_huge_withdraw(mm, src_pmd); + pgtable_trans_huge_deposit(mm, dst_pmd, src_pgtable); + + double_pt_unlock(dst_ptl, src_ptl); + return 0; +} +#endif /* CONFIG_THP_SWAP */ + /* * The PT lock for src_pmd and dst_vma/src_vma (for reading) are locked by * the caller, but it must return after releasing the page_table_lock. @@ -2947,11 +3018,78 @@ int move_pages_huge_pmd(struct mm_struct *mm, pmd_t= *dst_pmd, pmd_t *src_pmd, pm } =20 if (!pmd_trans_huge(src_pmdval)) { - spin_unlock(src_ptl); if (pmd_is_migration_entry(src_pmdval)) { + spin_unlock(src_ptl); pmd_migration_entry_wait(mm, src_pmd); return -EAGAIN; } +#ifdef CONFIG_THP_SWAP + if (pmd_is_swap_entry(src_pmdval)) { + swp_entry_t entry; + struct swap_info_struct *si; + enum swap_pmd_cache cache_state; + + /* + * UFFDIO_MOVE on anon mappings requires single-owner + * semantics; refuse to move a shared swap entry. + */ + if (!pmd_swp_exclusive(src_pmdval)) { + spin_unlock(src_ptl); + return -EBUSY; + } + + entry =3D softleaf_from_pmd(src_pmdval); + spin_unlock(src_ptl); + + /* Pin the swap device against a racing swapoff. */ + si =3D get_swap_device(entry); + if (unlikely(!si)) + return -EAGAIN; + + src_folio =3D NULL; + cache_state =3D swap_pmd_cache_lookup(entry, &src_folio); + if (cache_state =3D=3D SWAP_PMD_CACHE_SPLIT) { + put_swap_device(si); + __split_huge_pmd(src_vma, src_pmd, src_addr, false); + return -EAGAIN; + } + + mmu_notifier_range_init(&range, MMU_NOTIFY_CLEAR, 0, + mm, src_addr, + src_addr + HPAGE_PMD_SIZE); + mmu_notifier_invalidate_range_start(&range); + + if (src_folio) { + folio_lock(src_folio); + if (!folio_matches_swap_entry(src_folio, entry) || + folio_nr_pages(src_folio) !=3D HPAGE_PMD_NR) { + err =3D -EAGAIN; + folio_unlock(src_folio); + folio_put(src_folio); + mmu_notifier_invalidate_range_end(&range); + put_swap_device(si); + __split_huge_pmd(src_vma, src_pmd, + src_addr, false); + return err; + } + } + + dst_ptl =3D pmd_lockptr(mm, dst_pmd); + err =3D move_swap_pmd(mm, dst_vma, dst_addr, src_addr, + dst_pmd, src_pmd, dst_pmdval, + src_pmdval, dst_ptl, src_ptl, + src_folio, entry); + + mmu_notifier_invalidate_range_end(&range); + if (src_folio) { + folio_unlock(src_folio); + folio_put(src_folio); + } + put_swap_device(si); + return err; + } +#endif /* CONFIG_THP_SWAP */ + spin_unlock(src_ptl); return -ENOENT; } =20 diff --git a/mm/userfaultfd.c b/mm/userfaultfd.c index 23fb68fce000e..3692ffb326ffd 100644 --- a/mm/userfaultfd.c +++ b/mm/userfaultfd.c @@ -2106,6 +2106,20 @@ static ssize_t move_pages(struct userfaultfd_ctx *ct= x, unsigned long dst_start, break; } =20 + /* + * A PMD swap entry at dst is a swapped-out THP, not a hole, + * and unlike a PMD migration entry it will not resolve on its + * own. Nothing below faults it back in: pte_alloc() skips a + * !pmd_none PMD, pte_offset_map_rw_nolock() then fails on the + * non-present PMD, and the -EAGAIN that produces would be + * retried forever by the loop below. Be strict, exactly as for + * a present THP. + */ + if (unlikely(pmd_is_swap_entry(dst_pmdval))) { + err =3D -EEXIST; + break; + } + ptl =3D pmd_trans_huge_lock(src_pmd, src_vma); if (ptl) { /* Check if we can move the pmd without splitting it. */ --=20 2.53.0-Meta From nobody Mon Sep 28 19:34:00 2026 Received: from mta1.migadu.com (out-243.mta1.migadu.com [95.215.58.243]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C6E6C476CDD for ; Tue, 18 Aug 2026 13:12:55 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=95.215.58.243 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787058778; cv=none; b=H5iD5ZfCA1U+LOVsmtpO5BFGYMufek46CvABrrjZbA41+uhve6gPeimyxVfAwUtnS0u4aXSuXHonx+8RFs4xa3VNKJKG7+OZYo+SvmWYv3gKHX6Jwlnwwa6EkLlC1mr5jmWUXbFgABYS59ynvM2nuL8zFDYqZr6O1kfah72w/Is= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787058778; c=relaxed/simple; bh=nfAomkfVSHX4HliseOLQoYgtxdVGOCZmkB003762vy8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Ln2bMt6hJT7daOHdwX8xWP1E3bjegxZfwg/nCo7iJnFB7qaL4XOoE/1U9OBXW8CZeQVlWz3oefieiiQkkK39QQhZiwvM/ZDD7iANpTlTm+ESxAIENHoQUwbbbJQOuuHwQfl4o/wiOl5HDLf/KVDGK1PxNMZM19ARq2eywJ5+IwE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=QIK2Qo/n; arc=none smtp.client-ip=95.215.58.243 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="QIK2Qo/n" X-Envelope-To: linux-kernel@vger.kernel.org DKIM-Signature: a=rsa-sha256; bh=nfAomkfVSHX4HliseOLQoYgtxdVGOCZmkB003762vy8=; c=simple/simple; d=linux.dev; h=from:to:subject:date:message-id:mime-version:content-type; s=key1; t=1787058773; v=1; x=1787663573; b=QIK2Qo/nwr8duLqkuIJShi0aL5AJuoGB5PzZlmUnCA+XDwfCJw0Nb1sD+JKCXURl6suJjlB+ KFJH7NiOKCysOKUw4DGpPIYU1KVzQvT2y3tWTYOertMadr6RbMHukkXTohCes7AuqUF7Fli3o/9 1Pv3N5xZAsp9vwxtjJ6jD6ws= X-Envelope-To: linux-kernel@vger.kernel.org Received: from localhost (2a03:2880:10ff:58::) by mta11.migadu.com with ESMTPS id 9c9323b9ba099be2; Tue, 18 Aug 2026 13:12:51 +0000 X-Migadu-Flow: FLOW_OUT From: Usama Arif To: Andrew Morton , david@kernel.org, chrisl@kernel.org, kasong@tencent.com, ljs@kernel.org, ziy@nvidia.com, linux-mm@kvack.org Cc: ying.huang@linux.alibaba.com, Baoquan He , willy@infradead.org, youngjun.park@lge.com, hannes@cmpxchg.org, riel@surriel.com, shakeel.butt@linux.dev, alex@ghiti.fr, kas@kernel.org, baohua@kernel.org, dev.jain@arm.com, baolin.wang@linux.alibaba.com, Nico Pache , Liam R. Howlett , ryan.roberts@arm.com, Vlastimil Babka , lance.yang@linux.dev, linux-kernel@vger.kernel.org, nphamcs@gmail.com, shikemeng@huaweicloud.com, yosry@kernel.org, kernel-team@meta.com, Usama Arif Subject: [PATCH v6 10/12] mm: handle PMD swap entry faults on swap-in Date: Tue, 18 Aug 2026 06:09:51 -0700 Message-ID: <20260818131202.494754-11-usama.arif@linux.dev> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260818131202.494754-1-usama.arif@linux.dev> References: <20260818131202.494754-1-usama.arif@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Add do_huge_pmd_swap_page() and dispatch to it from __handle_mm_fault() when vmf->orig_pmd encodes a swap entry. The handler resolves the entire 2 MB mapping in one shot, mirroring do_swap_page() (PTE path) at PMD granularity: - Look up the folio in the swap cache; on a miss, allocate a PMD-order folio via swapin_sync(BIT(HPAGE_PMD_ORDER)) and read from swap. This deliberately skips the existing order-0 swap readahead paths: the fault already asks for the whole PMD range, while order-0 readahead would populate per-page swap cache state and force the PMD swap entry to split. If the range already has per-page swap-cache or zswap state, split and retry through PTEs. - After locking, re-validate that the folio still corresponds to our entry and is still PMD-sized. Between the unlocked cache lookup and the lock, a racing swap-in on the same entry may have removed it from the cache via folio_free_swap(), or reclaim / memory_failure / deferred-split may have split the folio into smaller folios. - Refuse to map a folio that contains a hardware-poisoned subpage. On a hit, split the PMD swap entry so do_swap_page() can return VM_FAULT_HWPOISON per subpage, matching the PTE swap-in PageHWPoison check. Disable large-folio PTE batching for such a folio so the fallback cannot map the poisoned subpage as part of a batch. - Restore soft_dirty and uffd_wp from the swap PMD. For an RWP VMA, also restore PAGE_NONE so the next access reaches userfaultfd. Map writable only when the entry was exclusive, the VMA permits writes, and uffd-wp is not armed. Drop the exclusive marker when the cached folio is under writeback to an SWP_STABLE_WRITES backend (zram) so the PMD is mapped read-only; a later write COWs into a fresh folio rather than corrupting the in-flight writeback. Mirrors do_swap_page(). - When the resulting PMD is read-only but the fault was a write, update vmf->orig_pmd and call wp_huge_pmd() in the same handler to COW without waiting for a second fault, unless wp_huge_pmd() itself has to fall back to PTE level. Leave an RWP-restored PMD for userfaultfd. Mask VM_FAULT_FALLBACK from the return: a PMD-COW that splits to PTE-level is normal, but the bit is part of VM_FAULT_ERROR and arch fault handlers BUG() on it without SIGBUS/HWPOISON/SIGSEGV. - Free the swap slot via should_try_to_free_swap() (hoisted from mm/memory.c into mm/internal.h so PTE- and PMD-level swap-in share the heuristic). When PMD-order resources are unavailable (folio allocation fails, the cached folio was split, memcg charge fails, or swap-cache insertion races) split the PMD swap entry into 512 PTE swap entries via __split_huge_pmd() and return 0. The fault retries and do_swap_page() takes over per-PTE. This avoids returning VM_FAULT_OOM for transient PMD-order allocation failures. Signed-off-by: Usama Arif --- include/linux/huge_mm.h | 14 +++ mm/huge_memory.c | 242 ++++++++++++++++++++++++++++++++++++++++ mm/internal.h | 42 +++++++ mm/memory.c | 43 ++----- 4 files changed, 306 insertions(+), 35 deletions(-) diff --git a/include/linux/huge_mm.h b/include/linux/huge_mm.h index c745f7ad22987..e7107e0991ad7 100644 --- a/include/linux/huge_mm.h +++ b/include/linux/huge_mm.h @@ -552,6 +552,15 @@ vm_fault_t do_huge_pmd_uffd_rwp(struct vm_fault *vmf); =20 vm_fault_t do_huge_pmd_device_private(struct vm_fault *vmf); =20 +#ifdef CONFIG_THP_SWAP +vm_fault_t do_huge_pmd_swap_page(struct vm_fault *vmf); +#else +static inline vm_fault_t do_huge_pmd_swap_page(struct vm_fault *vmf) +{ + return 0; +} +#endif + extern struct folio *huge_zero_folio; extern unsigned long huge_zero_pfn; =20 @@ -754,6 +763,11 @@ static inline vm_fault_t do_huge_pmd_device_private(st= ruct vm_fault *vmf) return 0; } =20 +static inline vm_fault_t do_huge_pmd_swap_page(struct vm_fault *vmf) +{ + return 0; +} + static inline bool is_huge_zero_folio(const struct folio *folio) { return false; diff --git a/mm/huge_memory.c b/mm/huge_memory.c index 0076d206c6a8a..10d265c7e6331 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -42,6 +42,7 @@ #include #include #include +#include =20 #include #include "internal.h" @@ -2397,6 +2398,247 @@ vm_fault_t do_huge_pmd_numa_page(struct vm_fault *v= mf) return 0; } =20 +#ifdef CONFIG_THP_SWAP +/** + * do_huge_pmd_swap_page() - Handle a fault on a PMD-level swap entry. + * @vmf: Fault context. vmf->orig_pmd contains the swap PMD. + * + * A PMD swap entry is a compact encoding for HPAGE_PMD_NR consecutive swap + * slots. If the swap cache still has one PMD-sized folio covering the ran= ge, + * map it directly at PMD level. If the range has been split into per-page + * cache state, or zswap may have per-page state for it, split the PMD swap + * entry and retry at PTE granularity. + * + * Return: VM_FAULT_* flags. + */ +vm_fault_t do_huge_pmd_swap_page(struct vm_fault *vmf) +{ + struct vm_area_struct *vma =3D vmf->vma; + struct mm_struct *mm =3D vma->vm_mm; + struct folio *folio; + struct page *page; + struct swap_info_struct *si; + unsigned long haddr =3D vmf->address & HPAGE_PMD_MASK; + softleaf_t entry; + swp_entry_t swp_entry; + pmd_t pmd; + vm_fault_t ret =3D 0; + bool exclusive, rwp_restore =3D false; + bool write =3D vmf->flags & FAULT_FLAG_WRITE; + rmap_t rmap_flags =3D RMAP_NONE; + enum swap_pmd_cache cache_state; + + entry =3D softleaf_from_pmd(vmf->orig_pmd); + if (unlikely(!softleaf_is_swap(entry))) + return 0; + + if (!thp_vma_allowable_order(vma, vma->vm_flags, TVA_PAGEFAULT, + HPAGE_PMD_ORDER)) { + __split_huge_pmd(vma, vmf->pmd, haddr, false); + return 0; + } + + swp_entry =3D entry; + + /* Prevent swapoff from happening to us. */ + si =3D get_swap_device(swp_entry); + if (unlikely(!si)) + return 0; + + cache_state =3D swap_pmd_cache_lookup(swp_entry, &folio); + if (cache_state =3D=3D SWAP_PMD_CACHE_SPLIT) + goto split_fallback; + if (!folio) { + /* + * PMD swap entries encode ordinary per-page swap slots. If any + * slot is in zswap, split and let the PTE swap path load the + * range per page. Otherwise the range is all on disk and can be + * read back as one PMD-sized folio. + */ + if (zswap_is_present(swp_entry, HPAGE_PMD_NR)) + goto split_fallback; + + folio =3D swapin_sync(swp_entry, GFP_HIGHUSER_MOVABLE, + BIT(HPAGE_PMD_ORDER), vmf, NULL, 0); + if (IS_ERR_OR_NULL(folio)) + goto split_fallback; + + /* Had to read from swap area: Major fault */ + ret =3D VM_FAULT_MAJOR; + count_vm_event(PGMAJFAULT); + count_memcg_event_mm(mm, PGMAJFAULT); + } + + ret |=3D folio_lock_or_retry(folio, vmf); + if (ret & VM_FAULT_RETRY) + goto out_release; + + /* Verify the folio is still in swap cache and matches our entry */ + if (unlikely(!folio_matches_swap_entry(folio, swp_entry))) + goto out_page; + + /* + * Folio should be PMD-sized; if not (e.g. split in swap cache), + * split the PMD swap entry and retry at PTE level. + */ + if (folio_nr_pages(folio) !=3D HPAGE_PMD_NR) { + folio_unlock(folio); + folio_put(folio); + goto split_fallback; + } + + if (unlikely(!folio_test_uptodate(folio))) { + if (zswap_is_present(swp_entry, HPAGE_PMD_NR)) { + folio_unlock(folio); + folio_put(folio); + goto split_fallback; + } + ret =3D VM_FAULT_SIGBUS; + goto out_page; + } + + /* + * If any subpage is hardware-poisoned, split the PMD swap entry and + * let the PTE swap-in path handle each page individually so + * do_swap_page() can return VM_FAULT_HWPOISON for the poisoned + * subpage rather than mapping the corrupted memory as one THP. + */ + if (unlikely(folio_contain_hwpoisoned_page(folio))) { + folio_unlock(folio); + folio_put(folio); + goto split_fallback; + } + + page =3D folio_page(folio, 0); + arch_swap_restore(folio_swap(swp_entry, folio), folio); + + folio_throttle_swaprate(folio, GFP_KERNEL); + + /* Lock the PMD and verify it hasn't changed */ + vmf->ptl =3D pmd_lock(mm, vmf->pmd); + if (unlikely(!pmd_same(vmf->orig_pmd, pmdp_get(vmf->pmd)))) { + spin_unlock(vmf->ptl); + goto out_page; + } + + exclusive =3D pmd_swp_exclusive(vmf->orig_pmd); + + /* + * Some swap backends (e.g. zram) don't support concurrent page + * modifications while under writeback. If we map exclusive on such + * a backend while the folio is still under writeback, the writeback + * may see partial modifications and corrupt the swap slot. Drop the + * exclusive marker and only map R/O for that case; further GUP + * references can't appear once the page is fully unmapped, so this + * is safe. + */ + if (exclusive && folio_test_writeback(folio) && + data_race(si->flags & SWP_STABLE_WRITES)) + exclusive =3D false; + + /* + * Set up the PMD mapping. Similar to do_swap_page() but at PMD level. + */ + add_mm_counter(mm, MM_ANONPAGES, HPAGE_PMD_NR); + add_mm_counter(mm, MM_SWAPENTS, -HPAGE_PMD_NR); + + pmd =3D folio_mk_pmd(folio, vma->vm_page_prot); + pmd =3D pmd_mkyoung(pmd); + + if (pmd_swp_soft_dirty(vmf->orig_pmd)) + pmd =3D pmd_mksoft_dirty(pmd); + if (pmd_swp_uffd(vmf->orig_pmd)) + pmd =3D pmd_mkuffd(pmd); + if (pmd_swp_uffd(vmf->orig_pmd) && userfaultfd_rwp(vma)) { + pmd =3D pmd_modify(pmd, PAGE_NONE); + rwp_restore =3D true; + } + + /* + * Check exclusivity to determine if we can map writable. + */ + if (exclusive) { + if (!rwp_restore && (vma->vm_flags & VM_WRITE) && + !userfaultfd_huge_pmd_wp(vma, pmd) && + !pmd_needs_soft_dirty_wp(vma, pmd)) { + pmd =3D pmd_mkwrite(pmd, vma); + if (write) + pmd =3D pmd_mkdirty(pmd); + } + rmap_flags |=3D RMAP_EXCLUSIVE; + } + + flush_icache_pages(vma, page, HPAGE_PMD_NR); + + if (!folio_test_anon(folio)) + folio_add_new_anon_rmap(folio, vma, haddr, rmap_flags); + else + folio_add_anon_rmap_pmd(folio, page, vma, haddr, rmap_flags); + + folio_put_swap(folio, NULL); + + set_pmd_at(mm, haddr, vmf->pmd, pmd); + update_mmu_cache_pmd(vma, haddr, vmf->pmd); + + /* Update orig_pmd for any follow-up wp_huge_pmd() below. */ + vmf->orig_pmd =3D pmd; + + /* + * Conditionally try to free up the swap cache. Do it after mapping, + * so raced page faults will likely see the folio in swap cache and + * wait on the folio lock. + */ + if (should_try_to_free_swap(si, folio, vma, exclusive, vmf->flags)) + folio_free_swap(folio); + + spin_unlock(vmf->ptl); + + folio_unlock(folio); + put_swap_device(si); + + /* + * If the write fault wasn't satisfied above (folio is shared without + * exclusivity), call wp_huge_pmd() to handle COW or + * userfaultfd-wp without forcing a second fault. + * + * wp_huge_pmd() may return VM_FAULT_FALLBACK if it had to split the + * PMD; that's a normal outcome, and the natural PTE-level refault will + * complete the COW. Mask it so callers (and the arch fault handler) + * don't see VM_FAULT_FALLBACK as a fatal VM_FAULT_ERROR. + */ + if (write && !pmd_write(pmd) && !rwp_restore) { + vm_fault_t wp_ret =3D wp_huge_pmd(vmf); + + wp_ret &=3D ~VM_FAULT_FALLBACK; + ret |=3D wp_ret; + if (ret & VM_FAULT_ERROR) + ret &=3D VM_FAULT_ERROR; + } + + return ret; + +out_page: + folio_unlock(folio); +out_release: + folio_put(folio); + put_swap_device(si); + return ret; + +split_fallback: + /* + * Only split if the PMD is still the swap entry we were called for. + * All the reasons we get here (allocation failure, zswap state, a + * split or poisoned cached folio) were observed without the PMD lock, + * so a racing thread may already have swapped the range back in as a + * THP -- splitting that would silently demote a perfectly good huge + * mapping. + */ + if (pmd_same(vmf->orig_pmd, pmdp_get_lockless(vmf->pmd))) + __split_huge_pmd(vma, vmf->pmd, haddr, false); + put_swap_device(si); + return 0; +} +#endif /* CONFIG_THP_SWAP */ static inline void zap_deposited_table(struct mm_struct *mm, pmd_t *pmd) { pgtable_t pgtable; diff --git a/mm/internal.h b/mm/internal.h index 38b1165212c94..ae4e18e1b14ee 100644 --- a/mm/internal.h +++ b/mm/internal.h @@ -574,6 +574,48 @@ static inline vm_fault_t vmf_anon_prepare(struct vm_fa= ult *vmf) } =20 vm_fault_t do_swap_page(struct vm_fault *vmf); + +#ifdef CONFIG_TRANSPARENT_HUGEPAGE +vm_fault_t wp_huge_pmd(struct vm_fault *vmf); +#else +static inline vm_fault_t wp_huge_pmd(struct vm_fault *vmf) +{ + return VM_FAULT_FALLBACK; +} +#endif + +/* + * Check if we should call folio_free_swap to free the swap cache. + * folio_free_swap only frees the swap cache to release the slot if swap + * count is zero, so we don't need to check the swap count here. + */ +static inline bool should_try_to_free_swap(struct swap_info_struct *si, + struct folio *folio, + struct vm_area_struct *vma, + bool exclusive, + unsigned int fault_flags) +{ + if (!folio_test_swapcache(folio)) + return false; + /* + * Always try to free swap cache for SWP_SYNCHRONOUS_IO devices. Swap + * cache can help save some IO or memory overhead, but these devices + * are fast, and meanwhile, swap cache pinning the slot deferring the + * release of metadata or fragmentation is a more critical issue. + */ + if (data_race(si->flags & SWP_SYNCHRONOUS_IO)) + return true; + if (mem_cgroup_swap_full(folio) || (vma->vm_flags & VM_LOCKED) || + folio_test_mlocked(folio)) + return true; + + /* + * Free the swapcache only if we are the exclusive user and + * this is a write fault. + */ + return (fault_flags & FAULT_FLAG_WRITE) && exclusive; +} + void folio_rotate_reclaimable(struct folio *folio); bool __folio_end_writeback(struct folio *folio); void deactivate_file_folio(struct folio *folio); diff --git a/mm/memory.c b/mm/memory.c index 36ddca806be2f..6f2cf12d3a45f 100644 --- a/mm/memory.c +++ b/mm/memory.c @@ -4640,38 +4640,6 @@ static vm_fault_t remove_device_exclusive_entry(stru= ct vm_fault *vmf) return 0; } =20 -/* - * Check if we should call folio_free_swap to free the swap cache. - * folio_free_swap only frees the swap cache to release the slot if swap - * count is zero, so we don't need to check the swap count here. - */ -static inline bool should_try_to_free_swap(struct swap_info_struct *si, - struct folio *folio, - struct vm_area_struct *vma, - bool exclusive, - unsigned int fault_flags) -{ - if (!folio_test_swapcache(folio)) - return false; - /* - * Always try to free swap cache for SWP_SYNCHRONOUS_IO devices. Swap - * cache can help save some IO or memory overhead, but these devices - * are fast, and meanwhile, swap cache pinning the slot deferring the - * release of metadata or fragmentation is a more critical issue. - */ - if (data_race(si->flags & SWP_SYNCHRONOUS_IO)) - return true; - if (mem_cgroup_swap_full(folio) || (vma->vm_flags & VM_LOCKED) || - folio_test_mlocked(folio)) - return true; - - /* - * Free the swapcache only if we are the exclusive user and - * this is a write fault. - */ - return (fault_flags & FAULT_FLAG_WRITE) && exclusive; -} - static vm_fault_t pte_marker_clear(struct vm_fault *vmf) { vmf->pte =3D pte_offset_map_lock(vmf->vma->vm_mm, vmf->pmd, @@ -5052,7 +5020,8 @@ vm_fault_t do_swap_page(struct vm_fault *vmf) page_idx =3D 0; address =3D vmf->address; ptep =3D vmf->pte; - if (folio_test_large(folio) && folio_test_swapcache(folio)) { + if (folio_test_large(folio) && folio_test_swapcache(folio) && + !folio_contain_hwpoisoned_page(folio)) { int nr =3D folio_nr_pages(folio); unsigned long idx =3D folio_page_idx(folio, page); unsigned long folio_start =3D address - idx * PAGE_SIZE; @@ -6387,8 +6356,8 @@ static inline vm_fault_t create_huge_pmd(struct vm_fa= ult *vmf) return VM_FAULT_FALLBACK; } =20 -/* `inline' is required to avoid gcc 4.1.2 build error */ -static inline vm_fault_t wp_huge_pmd(struct vm_fault *vmf) +#ifdef CONFIG_TRANSPARENT_HUGEPAGE +vm_fault_t wp_huge_pmd(struct vm_fault *vmf) { struct vm_area_struct *vma =3D vmf->vma; const bool unshare =3D vmf->flags & FAULT_FLAG_UNSHARE; @@ -6418,6 +6387,7 @@ static inline vm_fault_t wp_huge_pmd(struct vm_fault = *vmf) =20 return VM_FAULT_FALLBACK; } +#endif /* CONFIG_TRANSPARENT_HUGEPAGE */ =20 static vm_fault_t create_huge_pud(struct vm_fault *vmf) { @@ -6681,6 +6651,9 @@ static vm_fault_t __handle_mm_fault(struct vm_area_st= ruct *vma, =20 if (pmd_is_migration_entry(vmf.orig_pmd)) pmd_migration_entry_wait(mm, vmf.pmd); + else if (IS_ENABLED(CONFIG_THP_SWAP) && + pmd_is_swap_entry(vmf.orig_pmd)) + return do_huge_pmd_swap_page(&vmf); return 0; } if (pmd_trans_huge(vmf.orig_pmd)) { --=20 2.53.0-Meta From nobody Mon Sep 28 19:34:00 2026 Received: from mta0.migadu.com (out-206.mta0.migadu.com [91.218.175.206]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 82E0E476CF2 for ; Tue, 18 Aug 2026 13:12:56 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.206 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787058778; cv=none; b=IF+IEYIDY9sbKOuSVJu86R4E63pli8DOKiQbAe6wpmt0u/zgFCX/eWCXRpsX882pLnmBsCaH+HOYoYA3fNYVU+cVRbCkcykMiMSVOQmbO6JMZs9k5J2Ee06iN+JKbUbtbL4cVCXOfuOmbY9p7AyaPJr/2Q0zZao2Tb8tSmH+epA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787058778; c=relaxed/simple; bh=znqDwJ9Ce0OIJka+MopdalndNHzzmD2CA+PlSkpLojM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=uqCdKqWRjCFuQ4wZKOPcbUKzkdLZCC6VuFNZhBhg5wL16OoXtVx5v0GKQSYRzPpC9DeMwdQ+XGKcRhqgjI/B2ZYeG+jI698r/a2B5FNvCzwcXZQTLKw8/5gXs3f0GuF1S2sqWWFaoOzCUC0TEI1cTW08E5rQd67yRemRBJEVlP8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=jqvy0oKX; arc=none smtp.client-ip=91.218.175.206 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="jqvy0oKX" X-Envelope-To: linux-kernel@vger.kernel.org DKIM-Signature: a=rsa-sha256; bh=znqDwJ9Ce0OIJka+MopdalndNHzzmD2CA+PlSkpLojM=; c=simple/simple; d=linux.dev; h=from:to:subject:date:message-id:mime-version:content-type; s=key1; t=1787058774; v=1; x=1787663574; b=jqvy0oKX1w+CwCHplnlbH8DqhkhfvreVFJg2REbesx71iFxthEIdcBwh0tmv6KGdE9Pp/MYW lihEchfopdzwt5B88No/0gNdoaWkPGS1uEJnt3uzSQpywAAFn9abgJs5msBAgtgFGT7jcPKC+vD efQ9N8QLvBnLOc8mfSTGNUkg= X-Envelope-To: linux-kernel@vger.kernel.org Received: from localhost (2a03:2880:10ff:52::) by mta11.migadu.com with ESMTPS id 7af44e08486162c3; Tue, 18 Aug 2026 13:12:53 +0000 X-Migadu-Flow: FLOW_OUT From: Usama Arif To: Andrew Morton , david@kernel.org, chrisl@kernel.org, kasong@tencent.com, ljs@kernel.org, ziy@nvidia.com, linux-mm@kvack.org Cc: ying.huang@linux.alibaba.com, Baoquan He , willy@infradead.org, youngjun.park@lge.com, hannes@cmpxchg.org, riel@surriel.com, shakeel.butt@linux.dev, alex@ghiti.fr, kas@kernel.org, baohua@kernel.org, dev.jain@arm.com, baolin.wang@linux.alibaba.com, Nico Pache , Liam R. Howlett , ryan.roberts@arm.com, Vlastimil Babka , lance.yang@linux.dev, linux-kernel@vger.kernel.org, nphamcs@gmail.com, shikemeng@huaweicloud.com, yosry@kernel.org, kernel-team@meta.com, Usama Arif Subject: [PATCH v6 11/12] mm: install PMD swap entries on swap-out Date: Tue, 18 Aug 2026 06:09:52 -0700 Message-ID: <20260818131202.494754-12-usama.arif@linux.dev> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260818131202.494754-1-usama.arif@linux.dev> References: <20260818131202.494754-1-usama.arif@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Reclaim today splits a PMD-mapped anonymous THP into 512 PTE swap entries before unmap, losing the huge mapping across the swap round-trip and forcing khugepaged to rebuild it later. The contiguous swap range was already secured when the folio was added to the swap cache (a non-contiguous allocation would have split the folio earlier), so the PMD can be replaced by a single PMD-level swap entry instead. This patch mirrors the existing PTE swap-out path at PMD granularity: - shrink_folio_list() drops TTU_SPLIT_HUGE_PMD for PMD-mappable swapcache folios. zswap is handled by the PMD swap-in users: if any covered slot currently has a zswap entry, they split the PMD swap entry and fall back to the per-PTE path. - try_to_unmap_one() now has a PMD branch that calls set_pmd_swap_entry() and adjusts MM_ANONPAGES / MM_SWAPENTS by HPAGE_PMD_NR before walk_done. TTU_SPLIT_HUGE_PMD remains the fallback. - set_pmd_swap_entry() is the installer. Mirroring the PTE swap-out sequence at PMD granularity, it clears the present mapping (keeping the original for rollback), bumps the swap_map refcount for the folio's 512 slots, transfers the exclusive state in the swap entry, propagates the dirty bit to the folio so writeback is not lost, and installs a swap PMD that preserves the original soft-dirty / uffd-wp / exclusive bits. Any failing step rolls back the present mapping. The swap entry value matches what 512 PTE swap entries would encode, so swap_map refcounting is unchanged: each of the 512 slots carries a count of 1, released individually on later split or together on swap-in. Add thp_swpout_pmd to count each PMD mapping replaced by a PMD-level swap entry. Unlike the folio-level thp_swpout counter, a fork-shared THP can increment this counter once for each mapping; document that distinction. Signed-off-by: Usama Arif --- Documentation/admin-guide/mm/transhuge.rst | 5 ++ include/linux/huge_mm.h | 2 + include/linux/vm_event_item.h | 1 + mm/huge_memory.c | 81 ++++++++++++++++++++++ mm/rmap.c | 19 +++++ mm/vmscan.c | 9 ++- mm/vmstat.c | 1 + 7 files changed, 117 insertions(+), 1 deletion(-) diff --git a/Documentation/admin-guide/mm/transhuge.rst b/Documentation/adm= in-guide/mm/transhuge.rst index b187d618452f4..64d413d9fd83e 100644 --- a/Documentation/admin-guide/mm/transhuge.rst +++ b/Documentation/admin-guide/mm/transhuge.rst @@ -632,6 +632,11 @@ thp_swpout is incremented every time a huge page is swapout in one piece without splitting. =20 +thp_swpout_pmd + is incremented every time a PMD mapping is replaced by a PMD-level + swap entry. A fork-shared THP can increment this counter once for each + PMD mapping that is swapped out. + thp_swpout_fallback is incremented if a huge page has to be split before swapout. Usually because failed to allocate some continuous swap space diff --git a/include/linux/huge_mm.h b/include/linux/huge_mm.h index e7107e0991ad7..41cf643a3f55f 100644 --- a/include/linux/huge_mm.h +++ b/include/linux/huge_mm.h @@ -554,6 +554,8 @@ vm_fault_t do_huge_pmd_device_private(struct vm_fault *= vmf); =20 #ifdef CONFIG_THP_SWAP vm_fault_t do_huge_pmd_swap_page(struct vm_fault *vmf); +int set_pmd_swap_entry(struct page_vma_mapped_walk *pvmw, + struct folio *folio); #else static inline vm_fault_t do_huge_pmd_swap_page(struct vm_fault *vmf) { diff --git a/include/linux/vm_event_item.h b/include/linux/vm_event_item.h index 2628ccda076a0..f8fd4e13698c3 100644 --- a/include/linux/vm_event_item.h +++ b/include/linux/vm_event_item.h @@ -108,6 +108,7 @@ enum vm_event_item { PGPGIN, PGPGOUT, PSWPIN, PSWPOUT, THP_ZERO_PAGE_ALLOC_FAILED, THP_SWPOUT, THP_SWPOUT_FALLBACK, + THP_SWPOUT_PMD, #endif #ifdef CONFIG_BALLOON BALLOON_INFLATE, diff --git a/mm/huge_memory.c b/mm/huge_memory.c index 10d265c7e6331..d2f7a22aae3ad 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -5640,3 +5640,84 @@ void remove_migration_pmd(struct page_vma_mapped_wal= k *pvmw, struct folio *folio trace_remove_migration_pmd(address, pmd_val(pmde)); } #endif + +#ifdef CONFIG_THP_SWAP +/** + * set_pmd_swap_entry() - Replace a PMD mapping with a PMD-level swap entr= y. + * @pvmw: Page vma mapped walk context, must have pvmw->pmd set and + * pvmw->pte NULL (i.e. PMD-mapped). + * @folio: The folio being swapped out. Must be in the swap cache. + * + * This installs a PMD-level swap entry in place of a present PMD mapping, + * avoiding the need to split the PMD into PTE-level swap entries. + * + * Return: 0 on success, negative error code on failure. + */ +int set_pmd_swap_entry(struct page_vma_mapped_walk *pvmw, + struct folio *folio) +{ + struct vm_area_struct *vma =3D pvmw->vma; + struct mm_struct *mm =3D vma->vm_mm; + unsigned long address =3D pvmw->address; + unsigned long haddr =3D address & HPAGE_PMD_MASK; + struct page *page =3D folio_page(folio, 0); + bool anon_exclusive; + pmd_t pmdval; + swp_entry_t entry; + pmd_t pmdswp; + + if (WARN_ON_ONCE(!pvmw->pmd || pvmw->pte)) + return -EINVAL; + + VM_BUG_ON_FOLIO(!folio_test_swapcache(folio), folio); + VM_BUG_ON_FOLIO(!folio_test_anon(folio), folio); + VM_BUG_ON_FOLIO(folio_nr_pages(folio) !=3D HPAGE_PMD_NR, folio); + + if (unlikely(folio_test_swapbacked(folio) !=3D + folio_test_swapcache(folio))) { + WARN_ON_ONCE(1); + return -EBUSY; + } + + flush_cache_range(vma, haddr, haddr + HPAGE_PMD_SIZE); + + pmdval =3D pmdp_invalidate(vma, haddr, pvmw->pmd); + + /* Update high watermark before we lower rss */ + update_hiwater_rss(mm); + + if (folio_dup_swap(folio, NULL) < 0) { + set_pmd_at(mm, haddr, pvmw->pmd, pmdval); + return -ENOMEM; + } + + /* See folio_try_share_anon_rmap_pmd(): invalidate PMD first. */ + anon_exclusive =3D PageAnonExclusive(page); + if (anon_exclusive && folio_try_share_anon_rmap_pmd(folio, page)) { + folio_put_swap(folio, NULL); + set_pmd_at(mm, haddr, pvmw->pmd, pmdval); + return -EBUSY; + } + + mm_prepare_for_swap_entries(mm); + + if (pmd_dirty(pmdval)) + folio_mark_dirty(folio); + + entry =3D folio->swap; + pmdswp =3D softleaf_to_pmd(entry); + if (pmd_soft_dirty(pmdval)) + pmdswp =3D pmd_swp_mksoft_dirty(pmdswp); + if (pmd_uffd(pmdval)) + pmdswp =3D pmd_swp_mkuffd(pmdswp); + if (anon_exclusive) + pmdswp =3D pmd_swp_mkexclusive(pmdswp); + set_pmd_at(mm, haddr, pvmw->pmd, pmdswp); + + folio_remove_rmap_pmd(folio, page, vma); + folio_put(folio); + + count_vm_event(THP_SWPOUT_PMD); + return 0; +} +#endif /* CONFIG_THP_SWAP */ diff --git a/mm/rmap.c b/mm/rmap.c index d1819fd699380..4997724c2dd85 100644 --- a/mm/rmap.c +++ b/mm/rmap.c @@ -2282,6 +2282,25 @@ static bool try_to_unmap_one(struct folio *folio, st= ruct vm_area_struct *vma, goto walk_abort; } =20 +#ifdef CONFIG_THP_SWAP + /* + * If the folio is in the swap cache and we're not + * asked to split, install a PMD-level swap entry. + */ + if (!(flags & TTU_SPLIT_HUGE_PMD) && + folio_test_anon(folio) && + folio_test_swapcache(folio)) { + if (set_pmd_swap_entry(&pvmw, folio)) + goto walk_abort; + + add_mm_counter(mm, MM_ANONPAGES, + -HPAGE_PMD_NR); + add_mm_counter(mm, MM_SWAPENTS, + HPAGE_PMD_NR); + goto walk_done; + } +#endif + if (flags & TTU_SPLIT_HUGE_PMD) { /* * We temporarily have to drop the PTL and diff --git a/mm/vmscan.c b/mm/vmscan.c index c1404a59523d6..94038e8cc64d0 100644 --- a/mm/vmscan.c +++ b/mm/vmscan.c @@ -1329,7 +1329,14 @@ static unsigned int shrink_folio_list(struct list_he= ad *folio_list, enum ttu_flags flags =3D TTU_BATCH_FLUSH; bool was_swapbacked =3D folio_test_swapbacked(folio); =20 - if (folio_test_pmd_mappable(folio)) + /* + * With THP_SWAP, PMD-mappable folios already in the + * swap cache can be unmapped with a PMD-level swap + * entry, avoiding the cost of splitting the PMD. + */ + if (folio_test_pmd_mappable(folio) && + !(IS_ENABLED(CONFIG_THP_SWAP) && + folio_test_swapcache(folio))) flags |=3D TTU_SPLIT_HUGE_PMD; /* * Without TTU_SYNC, try_to_unmap will only begin to diff --git a/mm/vmstat.c b/mm/vmstat.c index cb57714539fb5..f40bca6aa45a0 100644 --- a/mm/vmstat.c +++ b/mm/vmstat.c @@ -1435,6 +1435,7 @@ const char * const vmstat_text[] =3D { [I(THP_ZERO_PAGE_ALLOC_FAILED)] =3D "thp_zero_page_alloc_failed", [I(THP_SWPOUT)] =3D "thp_swpout", [I(THP_SWPOUT_FALLBACK)] =3D "thp_swpout_fallback", + [I(THP_SWPOUT_PMD)] =3D "thp_swpout_pmd", #endif #ifdef CONFIG_BALLOON [I(BALLOON_INFLATE)] =3D "balloon_inflate", --=20 2.53.0-Meta From nobody Mon Sep 28 19:34:00 2026 Received: from mta1.migadu.com (out-12.mta1.migadu.com [95.215.58.12]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D0554477289 for ; Tue, 18 Aug 2026 13:13:01 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=95.215.58.12 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787058784; cv=none; b=tFotYfEtcZkF4hwgJBtlp4UJci/aihq1IWGJAC1/QOoviFLPjlKGsEjL9Q0OghXubPtk4oTA/2QY4u3uommbWojR94sz2T8b7K00i7yb6k1qq++TBMHWcYr65yUQnZ0Szd+RL4wi9fel7Z8RaK/tVcGgxDF/09swSVMmREZCYag= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787058784; c=relaxed/simple; bh=uDZbM4DLAiIeCyYKYiSCL2eLy15fVFg3+DOrPDm/rRA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=mXpJwZx/VN0m8V7waf6rJgAPz4e1E4AKlrcos/oz3ZGPtcpI7aVyfAqis3J1LQkLqRglkq3JruFMvR6d2/368OQpPiZTbNYRuW7MFdN7ZTF+mAozn9bAIpQVIOQe5jQwYP9B8XXsSRRZCucJHFofh1MvsD9g0AKkGA1L8pnfEEg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=Yx3sMq+A; arc=none smtp.client-ip=95.215.58.12 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="Yx3sMq+A" X-Envelope-To: linux-kernel@vger.kernel.org DKIM-Signature: a=rsa-sha256; bh=uDZbM4DLAiIeCyYKYiSCL2eLy15fVFg3+DOrPDm/rRA=; c=simple/simple; d=linux.dev; h=from:to:subject:date:message-id:mime-version:content-type; s=key1; t=1787058779; v=1; x=1787663579; b=Yx3sMq+A3CkrRWo6Jr1S36nhGOFlHEk50dErjDYt244xqxQEQhZ8FQl3RbSUzWIwFPDBB+xE fZ1uqiZ5gl5LbagmUS9zEFLX9tTXKVq1TR2SLgXB4Lqu3ZSde5mMdqtIU8crbDvmSYVB4d1l1xu f1BOpMpJDc1fXTDZlPB9DIl4= X-Envelope-To: linux-kernel@vger.kernel.org Received: from localhost (2a03:2880:10ff:14::) by mta11.migadu.com with ESMTPS id d54d858bd5989204; Tue, 18 Aug 2026 13:12:59 +0000 X-Migadu-Flow: FLOW_OUT From: Usama Arif To: Andrew Morton , david@kernel.org, chrisl@kernel.org, kasong@tencent.com, ljs@kernel.org, ziy@nvidia.com, linux-mm@kvack.org Cc: ying.huang@linux.alibaba.com, Baoquan He , willy@infradead.org, youngjun.park@lge.com, hannes@cmpxchg.org, riel@surriel.com, shakeel.butt@linux.dev, alex@ghiti.fr, kas@kernel.org, baohua@kernel.org, dev.jain@arm.com, baolin.wang@linux.alibaba.com, Nico Pache , Liam R. Howlett , ryan.roberts@arm.com, Vlastimil Babka , lance.yang@linux.dev, linux-kernel@vger.kernel.org, nphamcs@gmail.com, shikemeng@huaweicloud.com, yosry@kernel.org, kernel-team@meta.com, Usama Arif Subject: [PATCH v6 12/12] selftests/mm: add PMD swap entry tests Date: Tue, 18 Aug 2026 06:09:53 -0700 Message-ID: <20260818131202.494754-13-usama.arif@linux.dev> X-Mailer: git-send-email 2.53.0 In-Reply-To: <20260818131202.494754-1-usama.arif@linux.dev> References: <20260818131202.494754-1-usama.arif@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" Exercise the PMD swap entry paths. Each test gets a fresh PMD-mapped THP from fixture setup, fills it with a page-distinct pattern, swaps it out with MADV_PAGEOUT, and verifies that thp_swpout_pmd increased. The tests are: - basic: fault in a swapped PMD and verify its contents. - fork: verify parent and child can fault in the shared swap entry. - fork_cow: verify parent and child writes remain isolated. - write: fault in by writing one byte and preserve the rest of the THP. - rwp_swapin: verify userfaultfd RWP survives PMD-order swap-in. - munmap: unmap the full entry and check that VmSwap drops. - mprotect: change full-range protections without faulting the entry in. - split_mprotect: change half-range protections and verify the data. - split_munmap: unmap half, drop its accounting, and preserve the rest. - uffdio_move: move the entry and RWP state, then fault it in at dst. - mremap: force the entry to a new aligned address and verify the data. - pagemap: verify swapped bits and consecutive swap-slot offsets. - mincore: walk the entry without faulting it in. - madvise_free: release the slots, clear the entry, and verify zeroes. - madvise_willneed: prefetch the entry and verify subsequent swap-in. - swapoff: unuse the entry and preserve data and PMD/RWP state. Fixture teardown owns the mappings and file descriptors and restores swap after assertion failures. PMD_SWAP_DEVICE remains optional for swapoff. Distinguish an environment that cannot allocate a PMD THP from a failure to install a PMD swap entry, so the former skips while the latter fails. Also check VmSwap accounting, pagemap slot offsets, swapped state after non-faulting operations, and PMD restoration when zswap does not require PTE fallback. Register the test with run_vmtests.sh and the default kselftest runner. Signed-off-by: Usama Arif --- tools/testing/selftests/mm/Makefile | 2 + tools/testing/selftests/mm/ksft_pmd_swap.sh | 4 + tools/testing/selftests/mm/pmd_swap.c | 742 ++++++++++++++++++++ tools/testing/selftests/mm/run_vmtests.sh | 4 + 4 files changed, 752 insertions(+) create mode 100755 tools/testing/selftests/mm/ksft_pmd_swap.sh create mode 100644 tools/testing/selftests/mm/pmd_swap.c diff --git a/tools/testing/selftests/mm/Makefile b/tools/testing/selftests/= mm/Makefile index 2d5366196e309..dafa3a482451d 100644 --- a/tools/testing/selftests/mm/Makefile +++ b/tools/testing/selftests/mm/Makefile @@ -104,6 +104,7 @@ TEST_GEN_FILES +=3D guard-regions TEST_GEN_FILES +=3D merge TEST_GEN_FILES +=3D rmap TEST_GEN_FILES +=3D folio_split_race_test +TEST_GEN_FILES +=3D pmd_swap =20 ifneq ($(ARCH),arm64) TEST_GEN_FILES +=3D soft-dirty @@ -165,6 +166,7 @@ TEST_PROGS +=3D ksft_mremap.sh TEST_PROGS +=3D ksft_pagemap.sh TEST_PROGS +=3D ksft_pfnmap.sh TEST_PROGS +=3D ksft_pkey.sh +TEST_PROGS +=3D ksft_pmd_swap.sh TEST_PROGS +=3D ksft_process_madv.sh TEST_PROGS +=3D ksft_process_mrelease.sh TEST_PROGS +=3D ksft_rmap.sh diff --git a/tools/testing/selftests/mm/ksft_pmd_swap.sh b/tools/testing/se= lftests/mm/ksft_pmd_swap.sh new file mode 100755 index 0000000000000..0f070b4729a89 --- /dev/null +++ b/tools/testing/selftests/mm/ksft_pmd_swap.sh @@ -0,0 +1,4 @@ +#!/bin/sh -e +# SPDX-License-Identifier: GPL-2.0 + +./run_vmtests.sh -t pmd_swap diff --git a/tools/testing/selftests/mm/pmd_swap.c b/tools/testing/selftest= s/mm/pmd_swap.c new file mode 100644 index 0000000000000..30911ef6480f3 --- /dev/null +++ b/tools/testing/selftests/mm/pmd_swap.c @@ -0,0 +1,742 @@ +// SPDX-License-Identifier: GPL-2.0 +/* Test PMD-level swap entries and their users. */ +#define _GNU_SOURCE +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "kselftest_harness.h" +#include "vm_util.h" + +#define ZSWAP_ENABLED_PATH "/sys/module/zswap/parameters/enabled" + +/* pagemap: bits 0-54 hold the PFN, or type|offset for a swap entry. */ +#define PM_PFRAME_MASK ((1ULL << 55) - 1) +/* Must match MAX_SWAPFILES_SHIFT in include/linux/swap.h. */ +#define MAX_SWAPFILES_SHIFT 5 + +static bool check_swapped(int pagemap_fd, char *addr, unsigned long size) +{ + unsigned long off; + + for (off =3D 0; off < size; off +=3D getpagesize()) + if (!pagemap_is_swapped(pagemap_fd, addr + off)) + return false; + return true; +} + +static bool zswap_enabled(void) +{ + char enabled =3D 0; + FILE *f; + + f =3D fopen(ZSWAP_ENABLED_PATH, "r"); + if (!f) + return false; + + if (fscanf(f, " %c", &enabled) !=3D 1) + enabled =3D 0; + fclose(f); + + return enabled =3D=3D 'Y' || enabled =3D=3D 'y' || enabled =3D=3D '1'; +} + +static bool swap_available(unsigned long required_bytes) +{ + unsigned long required_kb =3D (required_bytes + 1023) / 1024; + unsigned long size_kb, used_kb; + char line[256]; + bool ret =3D false; + FILE *f; + + f =3D fopen("/proc/swaps", "r"); + if (!f) + return false; + + /* Skip the header. */ + if (!fgets(line, sizeof(line), f)) + goto out; + + while (fgets(line, sizeof(line), f)) { + if (sscanf(line, "%*s %*s %lu %lu", &size_kb, &used_kb) =3D=3D 2 && + size_kb >=3D used_kb && size_kb - used_kb >=3D required_kb) { + ret =3D true; + break; + } + } + +out: + fclose(f); + return ret; +} + +static unsigned long read_vm_event(const char *name) +{ + char line[256]; + size_t name_len =3D strlen(name); + unsigned long val =3D 0; + FILE *f; + + f =3D fopen("/proc/vmstat", "r"); + if (!f) + return 0; + while (fgets(line, sizeof(line), f)) { + if (!strncmp(line, name, name_len) && line[name_len] =3D=3D ' ') { + val =3D strtoul(line + name_len + 1, NULL, 10); + break; + } + } + fclose(f); + return val; +} + +static unsigned int random_seed(void) +{ + unsigned int seed; + + if (getrandom(&seed, sizeof(seed), 0) !=3D sizeof(seed)) + seed =3D (unsigned int)time(NULL); + return seed; +} + +static unsigned char pattern_byte(unsigned int seed, unsigned long off) +{ + return (unsigned char)(seed + off + (off >> 8) + (off >> 16)); +} + +static void fill_pattern(char *buf, unsigned long size, unsigned int seed) +{ + unsigned long i; + + for (i =3D 0; i < size; i++) + buf[i] =3D (char)pattern_byte(seed, i); +} + +static bool verify_pattern_range(char *buf, unsigned long size, + unsigned int seed, unsigned long offset) +{ + unsigned long i; + + for (i =3D 0; i < size; i++) + if ((unsigned char)buf[i] !=3D pattern_byte(seed, offset + i)) + return false; + return true; +} + +static bool verify_pattern(char *buf, unsigned long size, unsigned int see= d) +{ + return verify_pattern_range(buf, size, seed, 0); +} + +static bool verify_zero(char *buf, unsigned long size) +{ + unsigned long i; + + for (i =3D 0; i < size; i++) + if (buf[i]) + return false; + return true; +} + +/* + * mmap an anonymous PMD-aligned region of pmd_size bytes. Over-allocates + * by one PMD and trims the unaligned head/tail so the returned address is + * PMD-aligned (required for whole-PMD UFFDIO_MOVE). + */ +static char *mmap_pmd_aligned(unsigned long pmd_size) +{ + unsigned long pad =3D pmd_size; + char *raw, *aligned; + + raw =3D mmap(NULL, pmd_size + pad, PROT_READ | PROT_WRITE, + MAP_PRIVATE | MAP_ANONYMOUS, -1, 0); + if (raw =3D=3D MAP_FAILED) + return MAP_FAILED; + + aligned =3D (char *)(((uintptr_t)raw + pmd_size - 1) & ~(pmd_size - 1)); + if (aligned !=3D raw) + munmap(raw, aligned - raw); + if (aligned + pmd_size !=3D raw + pmd_size + pad) + munmap(aligned + pmd_size, + (raw + pmd_size + pad) - (aligned + pmd_size)); + return aligned; +} + +enum swap_thp_result { + SWAP_THP_OK, + SWAP_THP_UNAVAILABLE, + SWAP_THP_FAILED, +}; + +/* Per-process swapped size in bytes, from /proc/self/status VmSwap. */ +static unsigned long read_vmswap(void) +{ + char line[256]; + unsigned long kb =3D 0; + FILE *f; + + f =3D fopen("/proc/self/status", "r"); + if (!f) + return 0; + while (fgets(line, sizeof(line), f)) { + if (!strncmp(line, "VmSwap:", 7)) { + kb =3D strtoul(line + 7, NULL, 10); + break; + } + } + fclose(f); + return kb * 1024; +} + +static bool swap_out_pmd(char *mem, unsigned long pmd_size, int pagemap_fd) +{ + unsigned long before =3D read_vm_event("thp_swpout_pmd"); + unsigned long after; + + if (madvise(mem, pmd_size, MADV_PAGEOUT)) { + ksft_print_msg("MADV_PAGEOUT failed: %s\n", strerror(errno)); + return false; + } + if (!check_swapped(pagemap_fd, mem, pmd_size)) { + ksft_print_msg("MADV_PAGEOUT did not swap the whole PMD range\n"); + return false; + } + + after =3D read_vm_event("thp_swpout_pmd"); + ksft_print_msg("thp_swpout_pmd: %lu -> %lu\n", before, after); + return after > before; +} + +static char *alloc_fill_swap_thp(unsigned long pmd_size, int pagemap_fd, + unsigned int seed, enum swap_thp_result *res) +{ + char *mem; + + *res =3D SWAP_THP_UNAVAILABLE; + + mem =3D mmap_pmd_aligned(pmd_size); + if (mem =3D=3D MAP_FAILED) + return MAP_FAILED; + + if (madvise(mem, pmd_size, MADV_HUGEPAGE)) { + ksft_print_msg("MADV_HUGEPAGE failed: %s\n", strerror(errno)); + munmap(mem, pmd_size); + return MAP_FAILED; + } + fill_pattern(mem, pmd_size, seed); + + if (!check_huge_anon(mem, pmd_size, 1, pmd_size)) { + munmap(mem, pmd_size); + return MAP_FAILED; + } + *res =3D SWAP_THP_FAILED; + + if (!swap_out_pmd(mem, pmd_size, pagemap_fd)) { + munmap(mem, pmd_size); + return MAP_FAILED; + } + + *res =3D SWAP_THP_OK; + return mem; +} + +struct rwp_access_args { + unsigned char *addr; + unsigned char expected; + bool write; + bool ok; +}; + +static void *rwp_access_thread(void *data) +{ + struct rwp_access_args *args =3D data; + + if (args->write) + *args->addr =3D args->expected; + args->ok =3D *args->addr =3D=3D args->expected; + return NULL; +} + +static int register_rwp(char *addr, unsigned long size, bool protect) +{ + struct uffdio_register reg =3D {}; + struct uffdio_rwprotect rwp =3D {}; + struct uffdio_api api =3D {}; + int uffd; + + uffd =3D syscall(__NR_userfaultfd, O_CLOEXEC | O_NONBLOCK); + if (uffd < 0) + return -1; + + api.api =3D UFFD_API; + api.features =3D UFFD_FEATURE_RWP; + if (ioctl(uffd, UFFDIO_API, &api) || + !(api.features & UFFD_FEATURE_RWP)) + goto error; + + reg.range.start =3D (unsigned long)addr; + reg.range.len =3D size; + reg.mode =3D UFFDIO_REGISTER_MODE_RWP; + if (ioctl(uffd, UFFDIO_REGISTER, ®)) + goto error; + + if (!protect) + return uffd; + + rwp.range.start =3D (unsigned long)addr; + rwp.range.len =3D size; + rwp.mode =3D UFFDIO_RWPROTECT_MODE_RWP; + if (!ioctl(uffd, UFFDIO_RWPROTECT, &rwp)) + return uffd; + +error: + close(uffd); + return -1; +} + +static bool expect_rwp_fault(int uffd, char *addr, unsigned long size, + unsigned char expected, bool write) +{ + struct rwp_access_args args =3D { + .addr =3D (unsigned char *)addr, + .expected =3D expected, + .write =3D write, + }; + struct uffdio_rwprotect rwp =3D { + .range =3D { + .start =3D (unsigned long)addr, + .len =3D size, + }, + }; + struct pollfd pollfd =3D { + .fd =3D uffd, + .events =3D POLLIN, + }; + struct uffd_msg msg =3D {}; + pthread_t thread; + bool saw_rwp =3D false; + int ret; + + if (pthread_create(&thread, NULL, rwp_access_thread, &args)) + return false; + + ret =3D poll(&pollfd, 1, 5000); + if (ret =3D=3D 1 && (pollfd.revents & POLLIN) && + read(uffd, &msg, sizeof(msg)) =3D=3D (ssize_t)sizeof(msg)) { + saw_rwp =3D msg.event =3D=3D UFFD_EVENT_PAGEFAULT && + (msg.arg.pagefault.flags & UFFD_PAGEFAULT_FLAG_RWP); + } + + /* Resolve the access even on failure so the worker cannot remain blocked= . */ + ioctl(uffd, UFFDIO_RWPROTECT, &rwp); + if (pthread_join(thread, NULL)) + return false; + return saw_rwp && args.ok; +} + +FIXTURE(pmd_swap) +{ + unsigned long pmd_size; + unsigned long mem_len; + int pagemap_fd; + int uffd; + unsigned int seed; + bool zswap_enabled; + bool swap_disabled; + const char *swap_dev; + char *mem; + char *aux; +}; + +FIXTURE_SETUP(pmd_swap) +{ + enum swap_thp_result res; + + self->pagemap_fd =3D -1; + self->uffd =3D -1; + self->mem =3D MAP_FAILED; + self->aux =3D MAP_FAILED; + self->mem_len =3D 0; + self->swap_disabled =3D false; + self->swap_dev =3D getenv("PMD_SWAP_DEVICE"); + if (!strcmp(_metadata->name, "swapoff") && !self->swap_dev) + SKIP(return, "PMD_SWAP_DEVICE env var not set\n"); + + self->pmd_size =3D read_pmd_pagesize(); + if (!self->pmd_size) + SKIP(return, "Cannot determine PMD size\n"); + + self->pagemap_fd =3D open("/proc/self/pagemap", O_RDONLY); + if (self->pagemap_fd < 0) + SKIP(return, "Cannot open /proc/self/pagemap\n"); + + if (!swap_available(self->pmd_size)) + SKIP(return, "No active swap device has enough free space\n"); + + self->seed =3D random_seed(); + self->zswap_enabled =3D zswap_enabled(); + self->mem =3D alloc_fill_swap_thp(self->pmd_size, self->pagemap_fd, + self->seed, &res); + if (self->mem =3D=3D MAP_FAILED) { + ASSERT_NE(res, SWAP_THP_FAILED); + SKIP(return, "Could not create swapped THP\n"); + } + self->mem_len =3D self->pmd_size; +} + +FIXTURE_TEARDOWN(pmd_swap) +{ + int swap_err =3D 0; + int swap_ret =3D 0; + + if (self->swap_disabled) { + swap_ret =3D swapon(self->swap_dev, 0); + swap_err =3D errno; + } + if (self->uffd >=3D 0) + close(self->uffd); + if (self->aux !=3D MAP_FAILED) + munmap(self->aux, self->pmd_size); + if (self->mem !=3D MAP_FAILED) + munmap(self->mem, self->mem_len); + if (self->pagemap_fd >=3D 0) + close(self->pagemap_fd); + + EXPECT_EQ(swap_ret, 0) { + TH_LOG("swapon(%s) failed: %s", self->swap_dev, + strerror(swap_err)); + } +} + +TEST_F(pmd_swap, basic) +{ + ASSERT_TRUE(verify_pattern(self->mem, self->pmd_size, self->seed)); +} + +TEST_F(pmd_swap, fork) +{ + pid_t pid; + int status; + + pid =3D fork(); + ASSERT_GE(pid, 0); + + if (pid =3D=3D 0) + _exit(verify_pattern(self->mem, self->pmd_size, + self->seed) ? 0 : 1); + + ASSERT_TRUE(verify_pattern(self->mem, self->pmd_size, self->seed)); + + ASSERT_EQ(waitpid(pid, &status, 0), pid); + ASSERT_TRUE(WIFEXITED(status)); + ASSERT_EQ(WEXITSTATUS(status), 0); +} + +TEST_F(pmd_swap, fork_cow) +{ + unsigned int parent_seed =3D self->seed; + unsigned int child_seed =3D ~self->seed; + unsigned int new_seed =3D self->seed ^ 0xa5a5a5a5; + int release_child[2]; + bool parent_ok; + char c =3D 0; + pid_t pid; + int status, ret; + + ASSERT_EQ(pipe(release_child), 0); + + pid =3D fork(); + ASSERT_GE(pid, 0); + + if (pid =3D=3D 0) { + close(release_child[1]); + if (read(release_child[0], &c, 1) !=3D 1) + _exit(1); + if (!verify_pattern(self->mem, self->pmd_size, parent_seed)) + _exit(2); + fill_pattern(self->mem, self->pmd_size, child_seed); + if (!verify_pattern(self->mem, self->pmd_size, child_seed)) + _exit(3); + _exit(0); + } + + close(release_child[0]); + fill_pattern(self->mem, self->pmd_size, new_seed); + parent_ok =3D verify_pattern(self->mem, self->pmd_size, new_seed); + ret =3D write(release_child[1], &c, 1); + close(release_child[1]); + ASSERT_EQ(waitpid(pid, &status, 0), pid); + ASSERT_EQ(ret, 1); + ASSERT_TRUE(parent_ok); + ASSERT_TRUE(WIFEXITED(status)); + ASSERT_EQ(WEXITSTATUS(status), 0); + ASSERT_TRUE(verify_pattern(self->mem, self->pmd_size, new_seed)); +} + +TEST_F(pmd_swap, write) +{ + self->mem[0] =3D 0xbb; + ASSERT_EQ(self->mem[0], (char)0xbb); + ASSERT_TRUE(verify_pattern_range(self->mem + 1, self->pmd_size - 1, + self->seed, 1)); + if (!self->zswap_enabled) + ASSERT_TRUE(check_huge_anon(self->mem, self->pmd_size, 1, + self->pmd_size)); +} + +TEST_F(pmd_swap, rwp_swapin) +{ + self->uffd =3D register_rwp(self->mem, self->pmd_size, true); + if (self->uffd < 0) + SKIP(return, "Userfaultfd RWP unsupported\n"); + + ASSERT_TRUE(expect_rwp_fault(self->uffd, self->mem, self->pmd_size, + pattern_byte(self->seed, 0), false)); + ASSERT_TRUE(verify_pattern(self->mem, self->pmd_size, self->seed)); +} + +TEST_F(pmd_swap, munmap) +{ + unsigned long swap_before, swap_after; + int ret; + + swap_before =3D read_vmswap(); + ASSERT_GE(swap_before, self->pmd_size); + + ret =3D munmap(self->mem, self->pmd_size); + if (!ret) { + self->mem =3D MAP_FAILED; + self->mem_len =3D 0; + } + ASSERT_EQ(ret, 0); + + swap_after =3D read_vmswap(); + ASSERT_LE(swap_after, swap_before - self->pmd_size); +} + +TEST_F(pmd_swap, mprotect) +{ + ASSERT_EQ(mprotect(self->mem, self->pmd_size, PROT_READ), 0); + ASSERT_TRUE(check_swapped(self->pagemap_fd, self->mem, + self->pmd_size)); + ASSERT_EQ(mprotect(self->mem, self->pmd_size, + PROT_READ | PROT_WRITE), 0); + ASSERT_TRUE(check_swapped(self->pagemap_fd, self->mem, + self->pmd_size)); + ASSERT_TRUE(verify_pattern(self->mem, self->pmd_size, self->seed)); +} + +TEST_F(pmd_swap, split_mprotect) +{ + unsigned long half =3D self->pmd_size / 2; + + ASSERT_EQ(mprotect(self->mem, half, PROT_READ), 0); + ASSERT_TRUE(check_swapped(self->pagemap_fd, self->mem, + self->pmd_size)); + ASSERT_EQ(mprotect(self->mem, half, PROT_READ | PROT_WRITE), 0); + ASSERT_TRUE(verify_pattern(self->mem, self->pmd_size, self->seed)); +} + +TEST_F(pmd_swap, split_munmap) +{ + unsigned long half =3D self->pmd_size / 2; + unsigned long swap_before =3D read_vmswap(); + unsigned long i; + char *base =3D self->mem; + int ret; + + ASSERT_GE(swap_before, half); + ret =3D munmap(base, half); + if (!ret) { + self->mem =3D base + half; + self->mem_len =3D half; + } + ASSERT_EQ(ret, 0); + ASSERT_LE(read_vmswap(), swap_before - half); + + for (i =3D 0; i < half; i +=3D getpagesize()) + ASSERT_TRUE(pagemap_is_swapped(self->pagemap_fd, + self->mem + i)); + ASSERT_TRUE(verify_pattern_range(self->mem, half, self->seed, half)); +} + +TEST_F(pmd_swap, uffdio_move) +{ + struct uffdio_register reg =3D {}; + struct uffdio_move move =3D {}; + struct uffdio_api api =3D {}; + bool rwp; + + self->aux =3D mmap_pmd_aligned(self->pmd_size); + if (self->aux =3D=3D MAP_FAILED) + SKIP(return, "Could not mmap aligned dst\n"); + ASSERT_EQ(madvise(self->aux, self->pmd_size, MADV_HUGEPAGE), 0); + + self->uffd =3D syscall(__NR_userfaultfd, O_CLOEXEC | O_NONBLOCK); + if (self->uffd < 0) + SKIP(return, "userfaultfd unavailable\n"); + + api.api =3D UFFD_API; + api.features =3D UFFD_FEATURE_MOVE | UFFD_FEATURE_RWP; + if (ioctl(self->uffd, UFFDIO_API, &api) || + !(api.features & UFFD_FEATURE_MOVE)) + SKIP(return, "UFFD_FEATURE_MOVE unsupported\n"); + rwp =3D api.features & UFFD_FEATURE_RWP; + + reg.range.start =3D (unsigned long)self->aux; + reg.range.len =3D self->pmd_size; + reg.mode =3D UFFDIO_REGISTER_MODE_MISSING | + (rwp ? UFFDIO_REGISTER_MODE_RWP : 0); + ASSERT_EQ(ioctl(self->uffd, UFFDIO_REGISTER, ®), 0); + + move.dst =3D (unsigned long)self->aux; + move.src =3D (unsigned long)self->mem; + move.len =3D self->pmd_size; + ASSERT_EQ(ioctl(self->uffd, UFFDIO_MOVE, &move), 0); + ASSERT_EQ(move.move, self->pmd_size); + + ASSERT_TRUE(check_swapped(self->pagemap_fd, self->aux, + self->pmd_size)); + if (rwp) + ASSERT_TRUE(expect_rwp_fault(self->uffd, self->aux, + self->pmd_size, + pattern_byte(self->seed, 0), false)); + ASSERT_TRUE(verify_pattern(self->aux, self->pmd_size, self->seed)); + if (!self->zswap_enabled) + ASSERT_TRUE(check_huge_anon(self->aux, self->pmd_size, 1, + self->pmd_size)); +} + +TEST_F(pmd_swap, mremap) +{ + char *new_mem, *dst; + + self->aux =3D mmap_pmd_aligned(self->pmd_size); + if (self->aux =3D=3D MAP_FAILED) + SKIP(return, "Could not mmap aligned dst\n"); + dst =3D self->aux; + + new_mem =3D mremap(self->mem, self->pmd_size, self->pmd_size, + MREMAP_MAYMOVE | MREMAP_FIXED, dst); + if (new_mem !=3D MAP_FAILED) { + self->mem =3D new_mem; + self->aux =3D MAP_FAILED; + } + ASSERT_NE(new_mem, MAP_FAILED); + ASSERT_EQ(new_mem, dst); + + ASSERT_TRUE(check_swapped(self->pagemap_fd, new_mem, self->pmd_size)); + ASSERT_TRUE(verify_pattern(new_mem, self->pmd_size, self->seed)); +} + +TEST_F(pmd_swap, pagemap) +{ + uint64_t entry, first =3D 0; + unsigned long off; + + for (off =3D 0; off < self->pmd_size; off +=3D getpagesize()) { + entry =3D pagemap_get_entry(self->pagemap_fd, self->mem + off); + ASSERT_TRUE(entry & (1ULL << 62)); + ASSERT_FALSE(entry & (1ULL << 63)); + + if (entry & PM_PFRAME_MASK) { + uint64_t idx =3D off / getpagesize(); + + if (!off) + first =3D entry & PM_PFRAME_MASK; + ASSERT_EQ(entry & PM_PFRAME_MASK, + first + (idx << MAX_SWAPFILES_SHIFT)); + } + } +} + +TEST_F(pmd_swap, mincore) +{ + unsigned long pages =3D self->pmd_size / getpagesize(); + unsigned char vec[pages]; + + ASSERT_EQ(mincore(self->mem, self->pmd_size, vec), 0); + ASSERT_TRUE(check_swapped(self->pagemap_fd, self->mem, + self->pmd_size)); +} + +TEST_F(pmd_swap, madvise_free) +{ + unsigned long swap_before =3D read_vmswap(); + unsigned long i; + + ASSERT_TRUE(check_swapped(self->pagemap_fd, self->mem, + self->pmd_size)); + ASSERT_GE(swap_before, self->pmd_size); + ASSERT_EQ(madvise(self->mem, self->pmd_size, MADV_FREE), 0); + for (i =3D 0; i < self->pmd_size; i +=3D getpagesize()) + ASSERT_FALSE(pagemap_is_swapped(self->pagemap_fd, + self->mem + i)); + ASSERT_LE(read_vmswap(), swap_before - self->pmd_size); + ASSERT_TRUE(verify_zero(self->mem, self->pmd_size)); +} + +TEST_F(pmd_swap, madvise_willneed) +{ + ASSERT_EQ(madvise(self->mem, self->pmd_size, MADV_WILLNEED), 0); + ASSERT_TRUE(check_swapped(self->pagemap_fd, self->mem, + self->pmd_size)); + ASSERT_TRUE(verify_pattern(self->mem, self->pmd_size, self->seed)); + if (!self->zswap_enabled) + ASSERT_TRUE(check_huge_anon(self->mem, self->pmd_size, 1, + self->pmd_size)); +} + +TEST_F(pmd_swap, swapoff) +{ + int ret, err; + + self->uffd =3D register_rwp(self->mem, self->pmd_size, true); + + ret =3D swapoff(self->swap_dev); + err =3D errno; + if (!ret) + self->swap_disabled =3D true; + ASSERT_EQ(ret, 0) { + TH_LOG("swapoff(%s) failed: %s", self->swap_dev, strerror(err)); + } + + /* + * Check residency before touching the memory. If we read + * first, a bug that left a PMD swap entry in place after swapoff + * would silently trigger do_huge_pmd_swap_page() and reinstall a + * PMD mapping, masking the regression. + */ + if (!self->zswap_enabled) + ASSERT_TRUE(check_huge_anon(self->mem, self->pmd_size, 1, + self->pmd_size)); + if (self->uffd >=3D 0) + ASSERT_TRUE(expect_rwp_fault(self->uffd, self->mem, + self->pmd_size, + pattern_byte(self->seed, 0), false)); + ASSERT_TRUE(verify_pattern(self->mem, self->pmd_size, self->seed)); + + ret =3D swapon(self->swap_dev, 0); + err =3D errno; + if (!ret) + self->swap_disabled =3D false; + ASSERT_EQ(ret, 0) { + TH_LOG("swapon(%s) failed: %s", self->swap_dev, strerror(err)); + } +} + +TEST_HARNESS_MAIN diff --git a/tools/testing/selftests/mm/run_vmtests.sh b/tools/testing/self= tests/mm/run_vmtests.sh index d09f9f6a384ee..ff53ff28c0042 100755 --- a/tools/testing/selftests/mm/run_vmtests.sh +++ b/tools/testing/selftests/mm/run_vmtests.sh @@ -69,6 +69,8 @@ separated by spaces: test pagemap_scan IOCTL - pfnmap tests for VM_PFNMAP handling +- pmd_swap + tests for PMD-level swap entries - process_madv test for process_madv - cow @@ -399,6 +401,8 @@ CATEGORY=3D"pagemap" run_test ./pagemap_ioctl =20 CATEGORY=3D"pfnmap" run_test ./pfnmap =20 +CATEGORY=3D"pmd_swap" run_test ./pmd_swap + # COW tests CATEGORY=3D"cow" run_test ./cow =20 --=20 2.53.0-Meta