From nobody Tue Sep 29 11:56:00 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0C5C1379EFC for ; Fri, 7 Aug 2026 21:17:49 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786137470; cv=none; b=kjXAm+u4DzasaLVaOgwRZgMpoQ2J6EzwJYswJLpgek8zLBJZ2WK3GO4Q97xGfon3jAvZkxFiqJQaz2B9aYd78nghx/CmAs4znjpE+LLYE/vVkZ6CKeMULFgAWPD0uUUAXIBsq8YABsMzBOu0edVyTSupGR3tPYlxGgMDUk0AH9o= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786137470; c=relaxed/simple; bh=qiYRZIqnSPHS9cwgme1xDdC4tc2BYZzhqH5XAMzhg60=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=ouWOSaCbSFOwVz9TOLSkgwu8ChRd2RLrickPSNgildezXSvUGQitcSEYwL3aYLoxo6Vv3QlZBKGg4i3UH9Zu4lpiTFtDm96eJ1vwffwjuS51oVao7bHwYbYZgWRWAie/E4RR5r7AVgBblgHAbz24ALKXdPviTKUu1HgLbBURE2M= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=TPG83SXJ; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="TPG83SXJ" Received: by smtp.kernel.org (Postfix) with ESMTPS id A1854C2BCC7; Fri, 7 Aug 2026 21:17:49 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1786137469; bh=qiYRZIqnSPHS9cwgme1xDdC4tc2BYZzhqH5XAMzhg60=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=TPG83SXJyY62J7+O3ulzYdAYqN1o7/SgLj7PKS05WPeGwfSMjeBxPdHzOeWKGgOjM Ur1Opvd9Of4mFW5onfsTUdLSfexttD2AiCUFWGKgUHlKSaiyUZBWeL1twUnOjLBOTC AIqZpLrE5rp0eAjIYJu8Tw/TmwdePmWbpgI6jIF6Jryu1IARn8UlX052CI4iIHTxoj 1ZBdsSMqBUXQp9TGp3G19PqawHUOfFgIvQYrwetc+oEN2U1s8KXNWusHdNzbLhp2Ny O2SdHsR51XH6Td3z7lE4nEIjp11V+yKCoUBmOftvM0Z/C9FcxV0F2uxOr/mHle+bhH NFxKpGiVbSUYA== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 8196BC5AC7A; Fri, 7 Aug 2026 21:17:49 +0000 (UTC) From: Kairui Song via B4 Relay Date: Sat, 08 Aug 2026 05:17:02 +0800 Subject: [PATCH RFC 01/13] mm/swap: fix off-by-one in swap cache replace sanity check Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260808-swap-thp-cleanup-v1-1-689939a7ccc3@tencent.com> References: <20260808-swap-thp-cleanup-v1-0-689939a7ccc3@tencent.com> In-Reply-To: <20260808-swap-thp-cleanup-v1-0-689939a7ccc3@tencent.com> To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, Andrew Morton , David Hildenbrand , Lorenzo Stoakes , Zi Yan , Baolin Wang , "Liam R. Howlett" , Nico Pache , Ryan Roberts , Dev Jain , Lance Yang , Usama Arif , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Chris Li , Kemeng Shi , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , Kairui Song X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1786137467; l=1310; i=kasong@tencent.com; s=kasong-sign-tencent; h=from:subject:message-id; bh=UvgIby0zG1FXH2X6ElUOn6PPzkRRWckpaZUTxQTZz5M=; b=Wh+7tf7JuTm0ygd9gFN7krWlLeYMCMlu6zdCeV4nbLXG9ujVY1uzJTuI7Ek7dUydurWjMJF0f lHMeliNanqoDYX3fO/JGiBouhbrGtkSbJlVhkvVTm3ESFyRcKI0lCgh X-Developer-Key: i=kasong@tencent.com; a=ed25519; pk=kCdoBuwrYph+KrkJnrr7Sm1pwwhGDdZKcKrqiK8Y1mI= X-Endpoint-Received: by B4 Relay for kasong@tencent.com/kasong-sign-tencent with auth_id=562 X-Original-From: Kairui Song Reply-To: kasong@tencent.com From: Kairui Song The DEBUG_VM sanity check in __swap_cache_replace_folio() iterates the old folio's range with "while (ci_off++ < ci_end)", so the loop body runs on the already-incremented offset: the first entry is skipped and one entry past the range is read. For a folio split that entry belongs to the first after-split folio and was just repointed by the replacement loop above, so the check would warn spuriously whenever sub-folio orders differ from the head folio's, as non-uniform swapcache splits now do. Use the same do-while pattern as the replacement loop. Fixes: 8578e0c00dcf ("mm, swap: use the swap table for the swap cache and s= witch API") Signed-off-by: Kairui Song Acked-by: Zi Yan --- mm/swap_state.c | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/mm/swap_state.c b/mm/swap_state.c index 5be825911e64..f1405e5b813e 100644 --- a/mm/swap_state.c +++ b/mm/swap_state.c @@ -388,8 +388,9 @@ void __swap_cache_replace_folio(struct swap_cluster_inf= o *ci, folio_order(old) !=3D folio_order(new)) { ci_off =3D swp_cluster_offset(old->swap); ci_end =3D ci_off + folio_nr_pages(old); - while (ci_off++ < ci_end) + do { WARN_ON_ONCE(swp_tb_to_folio(__swap_table_get(ci, ci_off)) !=3D old); + } while (++ci_off < ci_end); } } =20 --=20 2.55.0 From nobody Tue Sep 29 11:56:00 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1C2233CB2FD for ; Fri, 7 Aug 2026 21:17:50 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786137470; cv=none; b=NusdepgYxap0pN2lnRcQNgqXi7WoPP0S9QlCSa7GlyG/7zDbqSU6RdcArE+rXwOPMF1DOPHif8qws1hx0ma8FDjIzebjJD6SNjFbLUqCtsJScaaKnAwI/AmomUdMaAZGJQKpSIVauEEvmlaUAVci1ruF19US3m+2nwAm9UqptvI= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786137470; c=relaxed/simple; bh=undGCia7Jer3tja6gd6S285qvFmBFMFl3F3plDrS6m4=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=c1+AFIQ/U1JWtMIafUugamoOhOQjcW4/y1T44wEGJXwZpH4JyFCB7mye3iI4upVTDd1W9XFFwxr/OXpBv39tGb+CKOmUQf2ZXHW/0o2+W2kmH875NF26Xx3/dkNq9RjSVVAIc4xFhbt4SyHANPIg41PSTgsdus8viBh7fajx/jU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=rgVTuLkT; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="rgVTuLkT" Received: by smtp.kernel.org (Postfix) with ESMTPS id B6B9EC2BCF7; Fri, 7 Aug 2026 21:17:49 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1786137469; bh=undGCia7Jer3tja6gd6S285qvFmBFMFl3F3plDrS6m4=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=rgVTuLkTOyhcFcADjCFH+C6byyBBhX5VCA0gmbW3+UZKMVp4qV/u1yH6nVUHPc+g0 oiS/I7KgGpcThf/gxy4EGjzSYObbdc+EswN5idanBquF6vIw2Ji2ulquf0wb34oW+H nwuyaBShZ/vw996i6kmPzxRXYy97ZAIkt+jCOvJOsmGFr/M8wX00nf+YndR5nm/4Is 7eP7G/2zgnjqqgQb9QcqRIBu3yXHOPyqaWsP6qB0kz6qKFPsPvDaHTnR8wktgS5KBl P8Tzk6uQAKTqD1DiEqS/Gm9396kSK5knTIQKbdKVB3dI1e2vDXYHwKJ0zySrWRd8cD hKwWys3u/Dt5g== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 95E7BC2A09B; Fri, 7 Aug 2026 21:17:49 +0000 (UTC) From: Kairui Song via B4 Relay Date: Sat, 08 Aug 2026 05:17:03 +0800 Subject: [PATCH RFC 02/13] mm/huge_memory: fix rejection of swap cache folios with a mapping Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260808-swap-thp-cleanup-v1-2-689939a7ccc3@tencent.com> References: <20260808-swap-thp-cleanup-v1-0-689939a7ccc3@tencent.com> In-Reply-To: <20260808-swap-thp-cleanup-v1-0-689939a7ccc3@tencent.com> To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, Andrew Morton , David Hildenbrand , Lorenzo Stoakes , Zi Yan , Baolin Wang , "Liam R. Howlett" , Nico Pache , Ryan Roberts , Dev Jain , Lance Yang , Usama Arif , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Chris Li , Kemeng Shi , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , Kairui Song X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1786137467; l=3532; i=kasong@tencent.com; s=kasong-sign-tencent; h=from:subject:message-id; bh=CYIc1grENaJdW23IuxAR2mIbHqlN5p0Qzr5lTwIiQZc=; b=vj4vHBRaXAz7etIULkvV5Z9u1GthgAv8ww+CrE7FRqFWnH9VEyTBmFrzx4yZFf+ogRTl2aZwi kvK0TKpok5OCsScuovbsRcozYfPLcz2WQHypWq0BmnDk24bUVXuYs+I X-Developer-Key: i=kasong@tencent.com; a=ed25519; pk=kCdoBuwrYph+KrkJnrr7Sm1pwwhGDdZKcKrqiK8Y1mI= X-Endpoint-Received: by B4 Relay for kasong@tencent.com/kasong-sign-tencent with auth_id=562 X-Original-From: Kairui Song Reply-To: kasong@tencent.com From: Kairui Song A folio in the swap cache cannot be split if it has a mapping (shmem). The split code only checks for this in __folio_freeze_and_split_unmapped, after the folio ref has been frozen and the NR_SHMEM_THPS/NR_FILE_THPS counters have been decremented, and returns -EINVAL without unfreezing the folio or restoring the counters. That error path is fragile: if it is ever taken, the folio is left frozen and stuck, the counters are skewed, and the VM_WARN_ON_ONCE_FOLIO would fire for a state that is actually legitimate. Check for this case up front in folio_check_splittable and return -EINVAL before any state is modified. Under DEBUG_VM, the existing "Tried to split an unsplittable folio" warning in __folio_split reports the rejection. Fixes: 00527733d0dc ("mm/huge_memory: add two new (not yet used) functions = for folio_split()") Fixes: 714b056c8321 ("mm/huge_memory: convert VM_BUG* to VM_WARN* in __foli= o_split") Signed-off-by: Kairui Song --- mm/huge_memory.c | 26 ++++++++++++++++---------- 1 file changed, 16 insertions(+), 10 deletions(-) diff --git a/mm/huge_memory.c b/mm/huge_memory.c index ced400f72d43..2fa72158e063 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -3878,6 +3878,9 @@ static int __split_unmapped_folio(struct folio *folio= , int new_order, int folio_check_splittable(struct folio *folio, unsigned int new_order, enum split_type split_type) { + bool is_anon =3D folio_test_anon(folio); + bool is_swapcache =3D folio_test_swapcache(folio); + VM_WARN_ON_FOLIO(!folio_test_locked(folio), folio); /* * Folios that just got truncated cannot get split. Signal to the @@ -3886,11 +3889,11 @@ int folio_check_splittable(struct folio *folio, uns= igned int new_order, * TODO: this will also currently refuse folios without a mapping in the * swapcache (shmem or to-be-anon folios). */ - if (!folio->mapping && !folio_test_anon(folio)) + if (!folio->mapping && !is_anon) return -EBUSY; =20 /* order-1 is not supported for anonymous THP. */ - if (folio_test_anon(folio) && new_order =3D=3D 1) + if (is_anon && new_order =3D=3D 1) return -EINVAL; =20 /* @@ -3901,7 +3904,7 @@ int folio_check_splittable(struct folio *folio, unsig= ned int new_order, * swapcache folio split. Only uniform split to order-0 can be used * here. */ - if ((split_type =3D=3D SPLIT_TYPE_NON_UNIFORM || new_order) && folio_test= _swapcache(folio)) { + if ((split_type =3D=3D SPLIT_TYPE_NON_UNIFORM || new_order) && is_swapcac= he) { return -EINVAL; } =20 @@ -3911,6 +3914,15 @@ int folio_check_splittable(struct folio *folio, unsi= gned int new_order, if (folio_test_writeback(folio)) return -EBUSY; =20 + /* + * A non-anon swapcache folio that still has a mapping should only + * be a shmem folio under IO, there is little benefit in splitting + * them hence not supported. Reject it here up front: the split + * routine cannot back out cleanly once the folio ref is frozen. + */ + if (!is_anon && is_swapcache && folio->mapping) + return -EINVAL; + return 0; } =20 @@ -3983,14 +3995,8 @@ static int __folio_freeze_and_split_unmapped(struct = folio *folio, unsigned int n } } =20 - if (folio_test_swapcache(folio)) { - if (mapping) { - VM_WARN_ON_ONCE_FOLIO(mapping, folio); - return -EINVAL; - } - + if (folio_test_swapcache(folio)) ci =3D swap_cluster_get_and_lock(folio); - } =20 /* lock lru list/PageCompound, ref frozen by page_ref_freeze */ if (do_lru) --=20 2.55.0 From nobody Tue Sep 29 11:56:00 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1C188384CD5 for ; Fri, 7 Aug 2026 21:17:50 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786137470; cv=none; b=PxUym62uoHqROPs1ipW0vC+1PSpM69CMhRnrCkbYwd//14Uwis/xd8y1r9bO3CdVD5Nw6RfOhnoUV2FHmMLyvZB7V+aoLJAVt9eImnVVdQU3pl0rGMzFXNVQgqs1tMarNCGrVR2OtEZ933FND9vbpkKMJpQqZvRGnRuOzch6uI0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786137470; c=relaxed/simple; bh=Xjvbc6dxTdwAtdxCb/oqs8fxkZQn50o3lLgMYMga3lU=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=HJ+VcL5VBOq0UeSLgt4If45q7cbELYSK3ksv/aj+7VUfhPIMIT+rQG4Gm74fBnK8iLWIbJKYmXBfZ0ANWoeeJ1pMEbBwyAQBPZad2BSnCplCkh9ETXj5MM47BB5lkfjmu5qU6vnh8IDf/W5tcSO4d9FiNflOd04jmTxqlPbCaDo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=VBeEbGFV; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="VBeEbGFV" Received: by smtp.kernel.org (Postfix) with ESMTPS id CC9CEC2BD04; Fri, 7 Aug 2026 21:17:49 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1786137469; bh=Xjvbc6dxTdwAtdxCb/oqs8fxkZQn50o3lLgMYMga3lU=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=VBeEbGFVP4aY1rkUHKVeKdXQZDIzOAQjS/oNrzFbQAZlvIIQvjQIN/WkPpLYhscjZ Hb5q77v1lMZ/zt73cV13eU2nQtlNDAu+boxTn9N4qcTQXt22tBD1CKVNgeVmdcmp6Z hBADLhOB7n6xxMUd4KzM+UTpfJbWBligvCtvurOpAuSC2ICP49gv5T9FhsfdNx9yyi BOpIcpv6vwYQlN2+QZDjtLKMrNUEtNXJME+kCyd1eEQTPTz7T78xhqmcEnO2annbAU B86+2RHRdPjaB4RZfzsU/Nruv6BjQEIryMJr3VCCzRjeQ/jaknWPkXaW2cfZhwaJI1 0pJKqJcE2jD1A== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id AAB53C5ACD1; Fri, 7 Aug 2026 21:17:49 +0000 (UTC) From: Kairui Song via B4 Relay Date: Sat, 08 Aug 2026 05:17:04 +0800 Subject: [PATCH RFC 03/13] mm/huge_memory: invert folio_ref_freeze() check to reduce indentation Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260808-swap-thp-cleanup-v1-3-689939a7ccc3@tencent.com> References: <20260808-swap-thp-cleanup-v1-0-689939a7ccc3@tencent.com> In-Reply-To: <20260808-swap-thp-cleanup-v1-0-689939a7ccc3@tencent.com> To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, Andrew Morton , David Hildenbrand , Lorenzo Stoakes , Zi Yan , Baolin Wang , "Liam R. Howlett" , Nico Pache , Ryan Roberts , Dev Jain , Lance Yang , Usama Arif , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Chris Li , Kemeng Shi , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , Kairui Song X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1786137467; l=7825; i=kasong@tencent.com; s=kasong-sign-tencent; h=from:subject:message-id; bh=vEiSOg48pfKlibIxJXaAjoXOcM/bHDnJ1mwc1VCXIgI=; b=LMT7kjdS6jvyvicYZiiyyzzNsAhx/DUVFJfuN7IqgCQboDJ6ieR81y1ldneSjPPnU5tDqMOWr FLPV+jm6USLBz6V6nYiNojSqezgdSVklrvLkCovfIfJL0AY5zWvx8Yy X-Developer-Key: i=kasong@tencent.com; a=ed25519; pk=kCdoBuwrYph+KrkJnrr7Sm1pwwhGDdZKcKrqiK8Y1mI= X-Endpoint-Received: by B4 Relay for kasong@tencent.com/kasong-sign-tencent with auth_id=562 X-Original-From: Kairui Song Reply-To: kasong@tencent.com From: Kairui Song Invert the folio_ref_freeze() success check in __folio_freeze_and_split_unmapped() to return early on failure, which removes one level of indentation from the entire success path. This is a pure refactoring with no functional change. It prepares the function to be split into separate helpers for anonymous and file-backed folios in a later patch. Signed-off-by: Kairui Song Reviewed-by: Zi Yan --- mm/huge_memory.c | 181 +++++++++++++++++++++++++++------------------------= ---- 1 file changed, 90 insertions(+), 91 deletions(-) diff --git a/mm/huge_memory.c b/mm/huge_memory.c index 2fa72158e063..cf8f90b94e42 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -3941,9 +3941,11 @@ static int __folio_freeze_and_split_unmapped(struct = folio *folio, unsigned int n pgoff_t end, int *nr_shmem_dropped) { struct folio *end_folio =3D folio_next(folio); + struct swap_cluster_info *ci =3D NULL; struct folio *new_folio, *next; int old_order =3D folio_order(folio); struct list_lru_one *lru; + struct lruvec *lruvec; bool dequeue_deferred; int ret =3D 0; =20 @@ -3964,122 +3966,119 @@ static int __folio_freeze_and_split_unmapped(stru= ct folio *folio, unsigned int n lru =3D list_lru_lock(&deferred_split_lru, folio_nid(folio), &memcg); } - if (folio_ref_freeze(folio, folio_cache_ref_count(folio) + 1)) { - struct swap_cluster_info *ci =3D NULL; - struct lruvec *lruvec; =20 + if (!folio_ref_freeze(folio, folio_cache_ref_count(folio) + 1)) { if (dequeue_deferred) { - __list_lru_del(&deferred_split_lru, lru, - &folio->_deferred_list, folio_nid(folio)); - if (folio_test_partially_mapped(folio)) { - folio_clear_partially_mapped(folio); - mod_mthp_stat(old_order, - MTHP_STAT_NR_ANON_PARTIALLY_MAPPED, -1); - } list_lru_unlock(lru); rcu_read_unlock(); } + return -EAGAIN; + } =20 - if (mapping) { - int nr =3D folio_nr_pages(folio); - - if (folio_test_pmd_mappable(folio) && - new_order < HPAGE_PMD_ORDER) { - if (folio_test_swapbacked(folio)) { - lruvec_stat_mod_folio(folio, - NR_SHMEM_THPS, -nr); - } else { - lruvec_stat_mod_folio(folio, - NR_FILE_THPS, -nr); - } - } + if (dequeue_deferred) { + __list_lru_del(&deferred_split_lru, lru, + &folio->_deferred_list, folio_nid(folio)); + if (folio_test_partially_mapped(folio)) { + folio_clear_partially_mapped(folio); + mod_mthp_stat(old_order, + MTHP_STAT_NR_ANON_PARTIALLY_MAPPED, -1); } + list_lru_unlock(lru); + rcu_read_unlock(); + } =20 - if (folio_test_swapcache(folio)) - ci =3D swap_cluster_get_and_lock(folio); - - /* lock lru list/PageCompound, ref frozen by page_ref_freeze */ - if (do_lru) - lruvec =3D folio_lruvec_lock(folio); - - ret =3D __split_unmapped_folio(folio, new_order, split_at, xas, - mapping, split_type); + if (mapping) { + int nr =3D folio_nr_pages(folio); =20 - /* - * Unfreeze after-split folios and put them back to the right - * list. @folio should be kept frozon until page cache - * entries are updated with all the other after-split folios - * to prevent others seeing stale page cache entries. - * As a result, new_folio starts from the next folio of - * @folio. - */ - for (new_folio =3D folio_next(folio); new_folio !=3D end_folio; - new_folio =3D next) { - unsigned long nr_pages =3D folio_nr_pages(new_folio); + if (folio_test_pmd_mappable(folio) && + new_order < HPAGE_PMD_ORDER) { + if (folio_test_swapbacked(folio)) { + lruvec_stat_mod_folio(folio, + NR_SHMEM_THPS, -nr); + } else { + lruvec_stat_mod_folio(folio, + NR_FILE_THPS, -nr); + } + } + } =20 - next =3D folio_next(new_folio); + if (folio_test_swapcache(folio)) + ci =3D swap_cluster_get_and_lock(folio); =20 - zone_device_private_split_cb(folio, new_folio); + /* lock lru list/PageCompound, ref frozen by page_ref_freeze */ + if (do_lru) + lruvec =3D folio_lruvec_lock(folio); =20 - folio_ref_unfreeze(new_folio, - folio_cache_ref_count(new_folio) + 1); + ret =3D __split_unmapped_folio(folio, new_order, split_at, xas, + mapping, split_type); =20 - if (do_lru) - lru_add_split_folio(folio, new_folio, lruvec, list); + /* + * Unfreeze after-split folios and put them back to the right + * list. @folio should be kept frozon until page cache + * entries are updated with all the other after-split folios + * to prevent others seeing stale page cache entries. + * As a result, new_folio starts from the next folio of + * @folio. + */ + for (new_folio =3D folio_next(folio); new_folio !=3D end_folio; + new_folio =3D next) { + unsigned long nr_pages =3D folio_nr_pages(new_folio); =20 - /* - * Anonymous folio with swap cache. - * NOTE: shmem in swap cache is not supported yet. - */ - if (ci) { - __swap_cache_replace_folio(ci, folio, new_folio); - continue; - } + next =3D folio_next(new_folio); =20 - /* Anonymous folio without swap cache */ - if (!mapping) - continue; + zone_device_private_split_cb(folio, new_folio); =20 - /* Add the new folio to the page cache. */ - if (new_folio->index < end) { - __xa_store(&mapping->i_pages, new_folio->index, - new_folio, 0); - continue; - } + folio_ref_unfreeze(new_folio, + folio_cache_ref_count(new_folio) + 1); =20 - VM_WARN_ON_ONCE(!nr_shmem_dropped); - /* Drop folio beyond EOF: ->index >=3D end */ - if (shmem_mapping(mapping) && nr_shmem_dropped) - *nr_shmem_dropped +=3D nr_pages; - else if (folio_test_clear_dirty(new_folio)) - folio_account_cleaned( - new_folio, inode_to_wb(mapping->host)); - __filemap_remove_folio(new_folio, NULL); - folio_put_refs(new_folio, nr_pages); - } + if (do_lru) + lru_add_split_folio(folio, new_folio, lruvec, list); =20 - zone_device_private_split_cb(folio, NULL); /* - * Unfreeze @folio only after all page cache entries, which - * used to point to it, have been updated with new folios. - * Otherwise, a parallel folio_try_get() can grab @folio - * and its caller can see stale page cache entries. + * Anonymous folio with swap cache. + * NOTE: shmem in swap cache is not supported yet. */ - folio_ref_unfreeze(folio, folio_cache_ref_count(folio) + 1); + if (ci) { + __swap_cache_replace_folio(ci, folio, new_folio); + continue; + } =20 - if (do_lru) - lruvec_unlock(lruvec); + /* Anonymous folio without swap cache */ + if (!mapping) + continue; =20 - if (ci) - swap_cluster_unlock(ci); - } else { - if (dequeue_deferred) { - list_lru_unlock(lru); - rcu_read_unlock(); + /* Add the new folio to the page cache. */ + if (new_folio->index < end) { + __xa_store(&mapping->i_pages, new_folio->index, + new_folio, 0); + continue; } - return -EAGAIN; + + VM_WARN_ON_ONCE(!nr_shmem_dropped); + /* Drop folio beyond EOF: ->index >=3D end */ + if (shmem_mapping(mapping) && nr_shmem_dropped) + *nr_shmem_dropped +=3D nr_pages; + else if (folio_test_clear_dirty(new_folio)) + folio_account_cleaned(new_folio, + inode_to_wb(mapping->host)); + __filemap_remove_folio(new_folio, NULL); + folio_put_refs(new_folio, nr_pages); } =20 + zone_device_private_split_cb(folio, NULL); + /* + * Unfreeze @folio only after all page cache entries, which + * used to point to it, have been updated with new folios. + * Otherwise, a parallel folio_try_get() can grab @folio + * and its caller can see stale page cache entries. + */ + folio_ref_unfreeze(folio, folio_cache_ref_count(folio) + 1); + + if (do_lru) + lruvec_unlock(lruvec); + if (ci) + swap_cluster_unlock(ci); + return ret; } =20 --=20 2.55.0 From nobody Tue Sep 29 11:56:00 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 33AD73E63A8 for ; Fri, 7 Aug 2026 21:17:50 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786137470; cv=none; b=IrHazgd5ceVSGeFh6vRGWMR7sShJKodcso0W5Pc7qJ5Y+AOfy4Bag6yEm7s1KA6Zg3i7lUQH1U+gRi0L9bTgvJ8xvfgD6eNSMAnZ57TXrZ4ZqwUZ3PVXEJBkE8q1rZVI5K69funSfZKoM5eTNu3LaPGf5WvZr7W989YstWec/U4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786137470; c=relaxed/simple; bh=h2zIakGlP4riVOSKtW+5kfZ5EhRMYipRYteot4VyaNs=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=kUj2Cl+Pnz61ye1bjSsofiARbdAWxAV+H6e4nrxFdCPdGAtv5aqhTmJCB6h61/SuF2UhFpUOFlQyEQCtbh7XQDy75HfGw/jk+KSzrJzm3nJ173Q2wN5audrDtfrPcjVexFaFsgE95t5pUq1dquU6raYXD2b9+wl0rluVNESiCGw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=TQ5D4qlB; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="TQ5D4qlB" Received: by smtp.kernel.org (Postfix) with ESMTPS id D9460C32781; Fri, 7 Aug 2026 21:17:49 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1786137469; bh=h2zIakGlP4riVOSKtW+5kfZ5EhRMYipRYteot4VyaNs=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=TQ5D4qlBUX9fhTvwUi8xbqIViGMx3eW7fBZVi5c72iBxOytMY91mqKfv8cm8+Y2qC +5G3XXajkOKwn+u2/dFEmSog4/ZAWUvZUxCFtDo18sO+vxGi2MwcORrKdM4S4b/XtI 3/DrHCJTvrQyIGKPH89Y1BpVdFO2czmy9y05rd5oEsv325ikdhPtvDz7jcs16Hno8S fxdMeIb/BTJK5eTQypeZqIaj9mpdemb5C8xv8xccnUNcg9kggDiHWx2K4DcatCCKQ3 pzrbvS1lGB9qX/gPeZUXXzyau86fXKYJlMu1igytz90useIs3CYRskDBuS+6Fyczhz ji4U2L9164HfA== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id C39CEC5AC7A; Fri, 7 Aug 2026 21:17:49 +0000 (UTC) From: Kairui Song via B4 Relay Date: Sat, 08 Aug 2026 05:17:05 +0800 Subject: [PATCH RFC 04/13] mm/huge_memory: split the routine for splitting anon and file folio Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260808-swap-thp-cleanup-v1-4-689939a7ccc3@tencent.com> References: <20260808-swap-thp-cleanup-v1-0-689939a7ccc3@tencent.com> In-Reply-To: <20260808-swap-thp-cleanup-v1-0-689939a7ccc3@tencent.com> To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, Andrew Morton , David Hildenbrand , Lorenzo Stoakes , Zi Yan , Baolin Wang , "Liam R. Howlett" , Nico Pache , Ryan Roberts , Dev Jain , Lance Yang , Usama Arif , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Chris Li , Kemeng Shi , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , Kairui Song X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1786137467; l=7993; i=kasong@tencent.com; s=kasong-sign-tencent; h=from:subject:message-id; bh=lE4agxpn3+FjF8GNafybHiUXu6DuwUcHCxx6OFIyyko=; b=jXXfNFzNWB6VtTyyD5fFqI1uatlfOQTIT/j+NeGk17u1RduV+pdazf0ClMyT0jhyVcRUiflDM AjI2pnjs8n2Cu+5vYJxZ9KJ6O8TjLytin2u13dIH9nuR48TVn8DzPu7 X-Developer-Key: i=kasong@tencent.com; a=ed25519; pk=kCdoBuwrYph+KrkJnrr7Sm1pwwhGDdZKcKrqiK8Y1mI= X-Endpoint-Received: by B4 Relay for kasong@tencent.com/kasong-sign-tencent with auth_id=562 X-Original-From: Kairui Song Reply-To: kasong@tencent.com From: Kairui Song No functional change intended. Before adding more logic, split __folio_freeze_and_split_unmapped() into an anon and a file variant so each path can evolve independently. The two paths shared little beyond the folio freeze call, the LRU locking, and the unfreeze skeleton, but differed in all other per-folio bookkeeping and routines. While splitting, some cleanups become easy to apply, and helped dropping a few now redundant checks. Signed-off-by: Kairui Song --- mm/huge_memory.c | 121 ++++++++++++++++++++++++++++++++++-----------------= ---- 1 file changed, 76 insertions(+), 45 deletions(-) diff --git a/mm/huge_memory.c b/mm/huge_memory.c index cf8f90b94e42..56a356c30f30 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -3934,22 +3934,19 @@ static unsigned int folio_cache_ref_count(const str= uct folio *folio) return folio_nr_pages(folio); } =20 -static int __folio_freeze_and_split_unmapped(struct folio *folio, unsigned= int new_order, - struct page *split_at, struct xa_state *xas, - struct address_space *mapping, bool do_lru, - struct list_head *list, enum split_type split_type, - pgoff_t end, int *nr_shmem_dropped) +static int __folio_freeze_split_unmapped_anon(struct folio *folio, unsigne= d int new_order, + struct page *split_at, bool do_lru, + struct list_head *list, enum split_type split_type) { struct folio *end_folio =3D folio_next(folio); struct swap_cluster_info *ci =3D NULL; - struct folio *new_folio, *next; + struct folio *new_folio; int old_order =3D folio_order(folio); struct list_lru_one *lru; struct lruvec *lruvec; bool dequeue_deferred; int ret =3D 0; =20 - VM_WARN_ON_ONCE(!mapping && end); /* * If this folio can be on the deferred split queue, lock out * the shrinker before freezing the ref. If the shrinker sees @@ -3957,7 +3954,7 @@ static int __folio_freeze_and_split_unmapped(struct f= olio *folio, unsigned int n * lock and must clean up the LRU state - the same dequeue we * will do below as part of the split. */ - dequeue_deferred =3D folio_test_anon(folio) && old_order > 1; + dequeue_deferred =3D old_order > 1; if (dequeue_deferred) { struct mem_cgroup *memcg; =20 @@ -3987,24 +3984,73 @@ static int __folio_freeze_and_split_unmapped(struct= folio *folio, unsigned int n rcu_read_unlock(); } =20 - if (mapping) { + if (folio_test_swapcache(folio)) + ci =3D swap_cluster_get_and_lock(folio); + + if (do_lru) + lruvec =3D folio_lruvec_lock(folio); + + ret =3D __split_unmapped_folio(folio, new_order, split_at, NULL, + NULL, split_type); + + /* + * Unfreeze the after-split folios and put them back to the right + * place, keeping the head @folio frozen until the end. While the + * folio is in the swap cache, the sub entries must be updated with + * their after-split folios before the head is unfrozen, so a + * concurrent swap_cache_get_folio() cannot return the head folio + * for a sub entry. Keeping the head frozen throughout also stops a + * parallel folio_try_get() from observing a partially split folio. + */ + for (new_folio =3D folio_next(folio); new_folio !=3D end_folio; + new_folio =3D folio_next(new_folio)) { + zone_device_private_split_cb(folio, new_folio); + folio_ref_unfreeze(new_folio, + folio_cache_ref_count(new_folio) + 1); + if (do_lru) + lru_add_split_folio(folio, new_folio, lruvec, list); + if (ci) + __swap_cache_replace_folio(ci, folio, new_folio); + } + + zone_device_private_split_cb(folio, NULL); + folio_ref_unfreeze(folio, folio_cache_ref_count(folio) + 1); + + if (do_lru) + lruvec_unlock(lruvec); + if (ci) + swap_cluster_unlock(ci); + + return ret; +} + +static int __folio_freeze_split_unmapped_file(struct folio *folio, unsigne= d int new_order, + struct page *split_at, struct xa_state *xas, + struct address_space *mapping, bool do_lru, + struct list_head *list, enum split_type split_type, + pgoff_t end, int *nr_shmem_dropped) +{ + struct folio *end_folio =3D folio_next(folio); + struct folio *new_folio, *next; + struct lruvec *lruvec; + int ret; + + if (!folio_ref_freeze(folio, folio_cache_ref_count(folio) + 1)) + return -EAGAIN; + + if (folio_test_pmd_mappable(folio) && + new_order < HPAGE_PMD_ORDER) { int nr =3D folio_nr_pages(folio); =20 - if (folio_test_pmd_mappable(folio) && - new_order < HPAGE_PMD_ORDER) { - if (folio_test_swapbacked(folio)) { - lruvec_stat_mod_folio(folio, - NR_SHMEM_THPS, -nr); - } else { - lruvec_stat_mod_folio(folio, - NR_FILE_THPS, -nr); - } + if (folio_test_swapbacked(folio)) { + lruvec_stat_mod_folio(folio, + NR_SHMEM_THPS, -nr); + } else { + lruvec_stat_mod_folio(folio, + NR_FILE_THPS, -nr); } } =20 - if (folio_test_swapcache(folio)) - ci =3D swap_cluster_get_and_lock(folio); - /* lock lru list/PageCompound, ref frozen by page_ref_freeze */ if (do_lru) lruvec =3D folio_lruvec_lock(folio); @@ -4014,7 +4060,7 @@ static int __folio_freeze_and_split_unmapped(struct f= olio *folio, unsigned int n =20 /* * Unfreeze after-split folios and put them back to the right - * list. @folio should be kept frozon until page cache + * list. @folio should be kept frozen until page cache * entries are updated with all the other after-split folios * to prevent others seeing stale page cache entries. * As a result, new_folio starts from the next folio of @@ -4026,27 +4072,12 @@ static int __folio_freeze_and_split_unmapped(struct= folio *folio, unsigned int n =20 next =3D folio_next(new_folio); =20 - zone_device_private_split_cb(folio, new_folio); - folio_ref_unfreeze(new_folio, folio_cache_ref_count(new_folio) + 1); =20 if (do_lru) lru_add_split_folio(folio, new_folio, lruvec, list); =20 - /* - * Anonymous folio with swap cache. - * NOTE: shmem in swap cache is not supported yet. - */ - if (ci) { - __swap_cache_replace_folio(ci, folio, new_folio); - continue; - } - - /* Anonymous folio without swap cache */ - if (!mapping) - continue; - /* Add the new folio to the page cache. */ if (new_folio->index < end) { __xa_store(&mapping->i_pages, new_folio->index, @@ -4065,7 +4096,6 @@ static int __folio_freeze_and_split_unmapped(struct f= olio *folio, unsigned int n folio_put_refs(new_folio, nr_pages); } =20 - zone_device_private_split_cb(folio, NULL); /* * Unfreeze @folio only after all page cache entries, which * used to point to it, have been updated with new folios. @@ -4076,8 +4106,6 @@ static int __folio_freeze_and_split_unmapped(struct f= olio *folio, unsigned int n =20 if (do_lru) lruvec_unlock(lruvec); - if (ci) - swap_cluster_unlock(ci); =20 return ret; } @@ -4231,10 +4259,14 @@ static int __folio_split(struct folio *folio, unsig= ned int new_order, ret =3D -EAGAIN; goto fail; } + ret =3D __folio_freeze_split_unmapped_file(folio, new_order, split_at, &= xas, mapping, + true, list, split_type, end, + &nr_shmem_dropped); + } else { + ret =3D __folio_freeze_split_unmapped_anon(folio, new_order, split_at, t= rue, + list, split_type); } =20 - ret =3D __folio_freeze_and_split_unmapped(folio, new_order, split_at, &xa= s, mapping, - true, list, split_type, end, &nr_shmem_dropped); fail: if (mapping) xas_unlock(&xas); @@ -4334,9 +4366,8 @@ int folio_split_unmapped(struct folio *folio, unsigne= d int new_order) return -EAGAIN; =20 local_irq_disable(); - ret =3D __folio_freeze_and_split_unmapped(folio, new_order, &folio->page,= NULL, - NULL, false, NULL, SPLIT_TYPE_UNIFORM, - 0, NULL); + ret =3D __folio_freeze_split_unmapped_anon(folio, new_order, &folio->page, + false, NULL, SPLIT_TYPE_UNIFORM); local_irq_enable(); return ret; } --=20 2.55.0 From nobody Tue Sep 29 11:56:00 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3FA2C3F4103 for ; Fri, 7 Aug 2026 21:17:50 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786137470; cv=none; b=EWHOlzQWDkVbr80zwjimeyTwtUOEpBsCXSBGjRd16MDBFVhwTQrQIsQ0gUrllBOc6QW3WhY4tsJ7ex7KgLAV/KBR0rSxKKx3fsXJgO87tsba15F/dGT6Cr6oXnr9ftBqHjOYgnJWHDnjT8dG0+6HSSF+4jOxgmrEXuLcWKz9gGs= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786137470; c=relaxed/simple; bh=qg2voHcSa5lIFMPir4HqweiwxwR7SabRJqMF5KNCN3E=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=QKeoKKl4WrVFuqbvVf04ZYQyr/4DjgJbrt5+QpqQrS6OGaTl5SfnrjAQNFnnDO6rMeA5VJbc9CkO4+R/yynZvlJ7C52MJTNYa267J03MYIajoyt7robcoBt1z/BZa6WGVd/iMRxFigyqwAV+A8Noliuz3awmV1hFDv9v988h/V0= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=A7Ha1ubr; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="A7Ha1ubr" Received: by smtp.kernel.org (Postfix) with ESMTPS id E7956C4AF0E; Fri, 7 Aug 2026 21:17:49 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1786137470; bh=qg2voHcSa5lIFMPir4HqweiwxwR7SabRJqMF5KNCN3E=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=A7Ha1ubrL6o1rXqEntRyJnY2Ksy7+SmDCTKoqOiTs2GfJn8c8OGbhFUtsPGaGyXTT QUyVp13MoH6DqCIqN7owlQMlAHjEIwWvjdhKCVWkvve/rXAgf621BlaY10vhiRYmcZ DzwW0aLKDorNo13j6KIN1GO3Ce2zSRcEtctTHGiXBk9cNMxIKxMTAB71onUNMmmRNG lYODRGQmOL/iV1O6dZuZb/YmVxXDMSN0QJ+ohkDYOlMxHq74YgfHiJ8FAXFRB4fDtQ OpogQBFiOYPSIRsStU9dbZqZmn9sNJbJu5V+4UXCL6ibupavE6xXZs9vnT4NE91yoc XNVha+PeNS7Iw== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id D56FEC5ACD3; Fri, 7 Aug 2026 21:17:49 +0000 (UTC) From: Kairui Song via B4 Relay Date: Sat, 08 Aug 2026 05:17:06 +0800 Subject: [PATCH RFC 05/13] mm/huge_memory: consolidate irq and locking for folio split Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260808-swap-thp-cleanup-v1-5-689939a7ccc3@tencent.com> References: <20260808-swap-thp-cleanup-v1-0-689939a7ccc3@tencent.com> In-Reply-To: <20260808-swap-thp-cleanup-v1-0-689939a7ccc3@tencent.com> To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, Andrew Morton , David Hildenbrand , Lorenzo Stoakes , Zi Yan , Baolin Wang , "Liam R. Howlett" , Nico Pache , Ryan Roberts , Dev Jain , Lance Yang , Usama Arif , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Chris Li , Kemeng Shi , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , Kairui Song X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1786137467; l=3951; i=kasong@tencent.com; s=kasong-sign-tencent; h=from:subject:message-id; bh=yZ53/oBOT/Q/to9ofK8tbkn2yDjvslEEnFYWLgM/+Qk=; b=t+/Q4iPxdj5Y5U1lyou6qmjzoFRjzxv2yON8YWRa/nW5Se1WbrQN+JkcG3tuKXkJthj2quc5f SN7VLiot+jECC4BI2fikxBSaQEF9iDysr49czZcbLdLvX1sEWJGsXmU X-Developer-Key: i=kasong@tencent.com; a=ed25519; pk=kCdoBuwrYph+KrkJnrr7Sm1pwwhGDdZKcKrqiK8Y1mI= X-Endpoint-Received: by B4 Relay for kasong@tencent.com/kasong-sign-tencent with auth_id=562 X-Original-From: Kairui Song Reply-To: kasong@tencent.com From: Kairui Song Let each split helper handle its own locking instead of relying on the caller, so both paths follow the same convention and __folio_split() can drop its local irq handling and fail label, preparing for further cleanup. Signed-off-by: Kairui Song --- mm/huge_memory.c | 52 ++++++++++++++++++++++++---------------------------- 1 file changed, 24 insertions(+), 28 deletions(-) diff --git a/mm/huge_memory.c b/mm/huge_memory.c index 56a356c30f30..ca9430a6f8d1 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -3947,6 +3947,8 @@ static int __folio_freeze_split_unmapped_anon(struct = folio *folio, unsigned int bool dequeue_deferred; int ret =3D 0; =20 + local_irq_disable(); + /* * If this folio can be on the deferred split queue, lock out * the shrinker before freezing the ref. If the shrinker sees @@ -3969,6 +3971,7 @@ static int __folio_freeze_split_unmapped_anon(struct = folio *folio, unsigned int list_lru_unlock(lru); rcu_read_unlock(); } + local_irq_enable(); return -EAGAIN; } =20 @@ -4020,6 +4023,7 @@ static int __folio_freeze_split_unmapped_anon(struct = folio *folio, unsigned int lruvec_unlock(lruvec); if (ci) swap_cluster_unlock(ci); + local_irq_enable(); =20 return ret; } @@ -4035,8 +4039,21 @@ static int __folio_freeze_split_unmapped_file(struct= folio *folio, unsigned int struct lruvec *lruvec; int ret; =20 - if (!folio_ref_freeze(folio, folio_cache_ref_count(folio) + 1)) - return -EAGAIN; + xas_lock_irq(xas); + + /* + * Check if the folio is present in page cache. + * We assume all tail are present too, if folio is there. + */ + if (xas_load(xas) !=3D folio) { + ret =3D -EAGAIN; + goto fail; + } + + if (!folio_ref_freeze(folio, folio_cache_ref_count(folio) + 1)) { + ret =3D -EAGAIN; + goto fail; + } =20 if (folio_test_pmd_mappable(folio) && new_order < HPAGE_PMD_ORDER) { @@ -4107,6 +4124,8 @@ static int __folio_freeze_split_unmapped_file(struct = folio *folio, unsigned int if (do_lru) lruvec_unlock(lruvec); =20 +fail: + xas_unlock_irq(xas); return ret; } =20 @@ -4246,19 +4265,7 @@ static int __folio_split(struct folio *folio, unsign= ed int new_order, =20 unmap_folio(folio); =20 - /* block interrupt reentry in xa_lock and spinlock */ - local_irq_disable(); - if (mapping) { - /* - * Check if the folio is present in page cache. - * We assume all tail are present too, if folio is there. - */ - xas_lock(&xas); - xas_reset(&xas); - if (xas_load(&xas) !=3D folio) { - ret =3D -EAGAIN; - goto fail; - } + if (!is_anon) { ret =3D __folio_freeze_split_unmapped_file(folio, new_order, split_at, &= xas, mapping, true, list, split_type, end, &nr_shmem_dropped); @@ -4267,12 +4274,6 @@ static int __folio_split(struct folio *folio, unsign= ed int new_order, list, split_type); } =20 -fail: - if (mapping) - xas_unlock(&xas); - - local_irq_enable(); - if (nr_shmem_dropped) shmem_uncharge(mapping->host, nr_shmem_dropped); =20 @@ -4355,8 +4356,6 @@ static int __folio_split(struct folio *folio, unsigne= d int new_order, */ int folio_split_unmapped(struct folio *folio, unsigned int new_order) { - int ret =3D 0; - VM_WARN_ON_ONCE_FOLIO(folio_mapped(folio), folio); VM_WARN_ON_ONCE_FOLIO(!folio_test_locked(folio), folio); VM_WARN_ON_ONCE_FOLIO(!folio_test_large(folio), folio); @@ -4365,11 +4364,8 @@ int folio_split_unmapped(struct folio *folio, unsign= ed int new_order) if (folio_expected_ref_count(folio) !=3D folio_ref_count(folio) - 1) return -EAGAIN; =20 - local_irq_disable(); - ret =3D __folio_freeze_split_unmapped_anon(folio, new_order, &folio->page, - false, NULL, SPLIT_TYPE_UNIFORM); - local_irq_enable(); - return ret; + return __folio_freeze_split_unmapped_anon(folio, new_order, &folio->page, + false, NULL, SPLIT_TYPE_UNIFORM); } =20 /* --=20 2.55.0 From nobody Tue Sep 29 11:56:00 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3F8703EEACF for ; Fri, 7 Aug 2026 21:17:50 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786137470; cv=none; b=sWBke2sNQQZ8/3P9dzXzdYjfddR4XQnYJCCAZymz5X1NmJva5O5lc9uAM6+/CMXywMTE/RhueyhXCnAUQ08JBPaigCGL5AZmlhgSIQqA3yWZBobDdyvEM46a4UBlDhdET6JkozcOd4kEs2Iy3IL5eIl31fJKcewliio+6/dzL4Y= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786137470; c=relaxed/simple; bh=RMQj/ADtXLid5GNTm45wXQFW8eJpGFU1MLjh5peezVc=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=a1jOCff2XSsuZ4/BibIfclxDGjJrO2JypbXxMZBhYSorlUn5nSJss6m3/HmqDoEMK6yS5Xj1hF0rLz3wuHpU8gIRoHVu38zUiPk3dG6szRFC6MPrXQIdvgrW28ORxAdqiYDCJ+CH1xJJ8AmzOdt5f9Bw/tDoU18KhvdijxL/EaE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=ub/1Y7Op; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="ub/1Y7Op" Received: by smtp.kernel.org (Postfix) with ESMTPS id 082BCC2BD05; Fri, 7 Aug 2026 21:17:50 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1786137470; bh=RMQj/ADtXLid5GNTm45wXQFW8eJpGFU1MLjh5peezVc=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=ub/1Y7Opi+xUA26V40Lass7adw6P90cGHSNhWBp0a+//G/CG9rgSnv9Zx7UFRB1go FEl03QH69nKJqB3RwVw7bE8OawVEwegfOZoVteR5xqZFm6Z+VW0K16izqWA9dVVn/J aWquVbYdgm3DAUQU/zj8dHn9e6aXKdByw3bputSzv0VxZG07A7boJcqsAJvBiIDyeU mHBmKbPWg0eUFsy4JqCaSwU3qAqJGl2eDqTgFNAx7kkcNVl2oGSg6VyLLC23zzx9A6 W9LGQpQcOr0MqtACrVMcCis8wWuKKUqbjz7w0nbQmbCek971WQ0YrhNzyCiGDfo+Mx ctj6PQ4jmN+hQ== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id E8325C5AC9F; Fri, 7 Aug 2026 21:17:49 +0000 (UTC) From: Kairui Song via B4 Relay Date: Sat, 08 Aug 2026 05:17:07 +0800 Subject: [PATCH RFC 06/13] mm/huge_memory: move EOF trimming into the file split helper Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260808-swap-thp-cleanup-v1-6-689939a7ccc3@tencent.com> References: <20260808-swap-thp-cleanup-v1-0-689939a7ccc3@tencent.com> In-Reply-To: <20260808-swap-thp-cleanup-v1-0-689939a7ccc3@tencent.com> To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, Andrew Morton , David Hildenbrand , Lorenzo Stoakes , Zi Yan , Baolin Wang , "Liam R. Howlett" , Nico Pache , Ryan Roberts , Dev Jain , Lance Yang , Usama Arif , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Chris Li , Kemeng Shi , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , Kairui Song X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1786137467; l=4129; i=kasong@tencent.com; s=kasong-sign-tencent; h=from:subject:message-id; bh=ZL+oPmr0JHpxplESJ7EerMTHHdu34aMC9es8n1iXisQ=; b=TUuL3slW4TdvxV30rEc7OV1KPB5pDFSkgui5vPimNk4tOvZccc4VWNgqV0EZxXrHdc9bAkfeW 7m5+33bUPJyDUPCcJYrU+saVmc62A+8kjy3ie6RTbvOMfdDqPsmOo1l X-Developer-Key: i=kasong@tencent.com; a=ed25519; pk=kCdoBuwrYph+KrkJnrr7Sm1pwwhGDdZKcKrqiK8Y1mI= X-Endpoint-Received: by B4 Relay for kasong@tencent.com/kasong-sign-tencent with auth_id=562 X-Original-From: Kairui Song Reply-To: kasong@tencent.com From: Kairui Song Instead of receiving @end and @nr_shmem_dropped from the caller, the file split helper now computes the EOF boundary and trims pages beyond it itself, as this is only needed for file split. This drops the redundant parameter passing and sanity check. Signed-off-by: Kairui Song Reviewed-by: Zi Yan --- mm/huge_memory.c | 42 +++++++++++++++++++----------------------- 1 file changed, 19 insertions(+), 23 deletions(-) diff --git a/mm/huge_memory.c b/mm/huge_memory.c index ca9430a6f8d1..72f5d0d24127 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -4031,14 +4031,26 @@ static int __folio_freeze_split_unmapped_anon(struc= t folio *folio, unsigned int static int __folio_freeze_split_unmapped_file(struct folio *folio, unsigne= d int new_order, struct page *split_at, struct xa_state *xas, struct address_space *mapping, bool do_lru, - struct list_head *list, enum split_type split_type, - pgoff_t end, int *nr_shmem_dropped) + struct list_head *list, enum split_type split_type) { struct folio *end_folio =3D folio_next(folio); struct folio *new_folio, *next; + int nr_shmem_dropped =3D 0; struct lruvec *lruvec; + pgoff_t end =3D 0; int ret; =20 + /* + *__split_unmapped_folio() may need to trim off pages beyond + * EOF: but on 32-bit, i_size_read() takes an irq-unsafe + * seqlock, which cannot be nested inside the page tree lock. + * So note end now: i_size itself may be changed at any moment, + * but folio lock is good enough to serialize the trimming. + */ + end =3D DIV_ROUND_UP(i_size_read(mapping->host), PAGE_SIZE); + if (shmem_mapping(mapping)) + end =3D shmem_fallocend(mapping->host, end); + xas_lock_irq(xas); =20 /* @@ -4102,10 +4114,9 @@ static int __folio_freeze_split_unmapped_file(struct= folio *folio, unsigned int continue; } =20 - VM_WARN_ON_ONCE(!nr_shmem_dropped); /* Drop folio beyond EOF: ->index >=3D end */ - if (shmem_mapping(mapping) && nr_shmem_dropped) - *nr_shmem_dropped +=3D nr_pages; + if (shmem_mapping(mapping)) + nr_shmem_dropped +=3D nr_pages; else if (folio_test_clear_dirty(new_folio)) folio_account_cleaned(new_folio, inode_to_wb(mapping->host)); @@ -4126,6 +4137,8 @@ static int __folio_freeze_split_unmapped_file(struct = folio *folio, unsigned int =20 fail: xas_unlock_irq(xas); + if (nr_shmem_dropped) + shmem_uncharge(mapping->host, nr_shmem_dropped); return ret; } =20 @@ -4162,9 +4175,7 @@ static int __folio_split(struct folio *folio, unsigne= d int new_order, struct anon_vma *anon_vma =3D NULL; int old_order =3D folio_order(folio); struct folio *new_folio, *next; - int nr_shmem_dropped =3D 0; enum ttu_flags ttu_flags =3D 0; - pgoff_t end =3D 0; int ret; =20 VM_WARN_ON_ONCE_FOLIO(!folio_test_locked(folio), folio); @@ -4241,17 +4252,6 @@ static int __folio_split(struct folio *folio, unsign= ed int new_order, =20 anon_vma =3D NULL; i_mmap_lock_read(mapping); - - /* - *__split_unmapped_folio() may need to trim off pages beyond - * EOF: but on 32-bit, i_size_read() takes an irq-unsafe - * seqlock, which cannot be nested inside the page tree lock. - * So note end now: i_size itself may be changed at any moment, - * but folio lock is good enough to serialize the trimming. - */ - end =3D DIV_ROUND_UP(i_size_read(mapping->host), PAGE_SIZE); - if (shmem_mapping(mapping)) - end =3D shmem_fallocend(mapping->host, end); } =20 /* @@ -4267,16 +4267,12 @@ static int __folio_split(struct folio *folio, unsig= ned int new_order, =20 if (!is_anon) { ret =3D __folio_freeze_split_unmapped_file(folio, new_order, split_at, &= xas, mapping, - true, list, split_type, end, - &nr_shmem_dropped); + true, list, split_type); } else { ret =3D __folio_freeze_split_unmapped_anon(folio, new_order, split_at, t= rue, list, split_type); } =20 - if (nr_shmem_dropped) - shmem_uncharge(mapping->host, nr_shmem_dropped); - if (!ret && is_anon && !folio_is_device_private(folio)) ttu_flags =3D TTU_USE_SHARED_ZEROPAGE; =20 --=20 2.55.0 From nobody Tue Sep 29 11:56:00 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 47EF53F6C3E for ; Fri, 7 Aug 2026 21:17:50 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786137470; cv=none; b=U2zJCQFYsi+wz/0910OcMt/GQq1ZnkxNzhgJBIkJ6g+aJFbaU+cU9dQYqLGp9SeUf0BN55mVjXEJmeUFjoGrHaAPiq1UQlwYQEOLUVtq1kvGixN3SBcKbAinT8dlbDUY3g0e3Z+SKyo+BEIPJ78hAtr23l/0pS4HDnHUiJonnRw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786137470; c=relaxed/simple; bh=BwmezivhFVZFocDKfY/PmLnAxu2FtilBaiVrvwrRpYc=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=bKNejKEDuNhP2QqYwULV0FKe3HScFKgKYbqUn+nU5Qw45JGQ9e/dKQbKO0cggoJ60ASIAVwBiRip0dQ/luE9Z/clBjgdYaE0Ffk3kQmTjheB+SgaFnbrEzAG58pOzq+SQfTvha6/1XUKyESSnmkjehH70tNdZnRlmYBcvijTsJo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=rnw4koTF; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="rnw4koTF" Received: by smtp.kernel.org (Postfix) with ESMTPS id 1AE5AC2BCFC; Fri, 7 Aug 2026 21:17:50 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1786137470; bh=BwmezivhFVZFocDKfY/PmLnAxu2FtilBaiVrvwrRpYc=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=rnw4koTF8epW9Q1mSzbuZ8ArLfRr2Q6HICHPDg3ZLhwNlM+qYKNU1dTYKhoVaDrKi 8IbtgyCSRbN4h38FYNNIgAgRXPgRRD4S3sSdyZ7rer2mWxVF8Dnx6fXsS2a5nwHHpP +7blL5SrSJWmG6hTQi7Y2kgz3/MN20D95pCmf5PV/tSoYKYk8G+BJgmJ4O7csPY0zY jEXJ8vBiG3WqSmwlH0wsXmvG/kDYtjknvhGaVTxH3MDP78yyf6N8T4dnUQTYIdM2Cg Ab8b8br20QrjvDiX+pOQX5w6B5DWuQ0GboDJVZCtUnoaVEciRmZu21r6LDlsKMkhvV URttBuRfqFD6g== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 06D41C5AC7A; Fri, 7 Aug 2026 21:17:50 +0000 (UTC) From: Kairui Song via B4 Relay Date: Sat, 08 Aug 2026 05:17:08 +0800 Subject: [PATCH RFC 07/13] mm/huge_memory: move unmap and remap into the split helpers Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260808-swap-thp-cleanup-v1-7-689939a7ccc3@tencent.com> References: <20260808-swap-thp-cleanup-v1-0-689939a7ccc3@tencent.com> In-Reply-To: <20260808-swap-thp-cleanup-v1-0-689939a7ccc3@tencent.com> To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, Andrew Morton , David Hildenbrand , Lorenzo Stoakes , Zi Yan , Baolin Wang , "Liam R. Howlett" , Nico Pache , Ryan Roberts , Dev Jain , Lance Yang , Usama Arif , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Chris Li , Kemeng Shi , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , Kairui Song X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1786137467; l=5093; i=kasong@tencent.com; s=kasong-sign-tencent; h=from:subject:message-id; bh=XdpcGLla7dWDxeaeoUQbxCu4Vf8PJ1zdPwmIqbBmjvw=; b=m8CSiewgW07eR4+Qa5gk6AQA4FGk9moAjQP7hGwJYLKKjrcBJIRoRtpKFRhbyMR95L58hhp4c ZxfbxvL3ITQCN2kJFkZqySU9ENT7pO4q4BKg/XlN1LYjBS8fDoFbxLI X-Developer-Key: i=kasong@tencent.com; a=ed25519; pk=kCdoBuwrYph+KrkJnrr7Sm1pwwhGDdZKcKrqiK8Y1mI= X-Endpoint-Received: by B4 Relay for kasong@tencent.com/kasong-sign-tencent with auth_id=562 X-Original-From: Kairui Song Reply-To: kasong@tencent.com From: Kairui Song To prepare for further cleanup, move the unmap/remap handling from __folio_split() into the split helpers. Only anon folios need to be remapped, so remap_page() is now only called for anon splits and the anon check in remap_page() is redundant and can be removed. Signed-off-by: Kairui Song Reviewed-by: Zi Yan --- mm/huge_memory.c | 51 ++++++++++++++++++++++++++------------------------- 1 file changed, 26 insertions(+), 25 deletions(-) diff --git a/mm/huge_memory.c b/mm/huge_memory.c index 72f5d0d24127..c0115841d1a0 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -3589,9 +3589,6 @@ static void remap_page(struct folio *folio, unsigned = long nr, int flags) { int i =3D 0; =20 - /* If unmap_folio() uses try_to_migrate() on file, remove this check */ - if (!folio_test_anon(folio)) - return; for (;;) { remove_migration_ptes(folio, folio, TTU_RMAP_LOCKED | flags); i +=3D folio_nr_pages(folio); @@ -3934,19 +3931,23 @@ static unsigned int folio_cache_ref_count(const str= uct folio *folio) return folio_nr_pages(folio); } =20 -static int __folio_freeze_split_unmapped_anon(struct folio *folio, unsigne= d int new_order, - struct page *split_at, bool do_lru, - struct list_head *list, enum split_type split_type) +static int __folio_freeze_split_unmap_anon(struct folio *folio, unsigned i= nt new_order, + struct page *split_at, bool do_lru, bool unmap, + struct list_head *list, enum split_type split_type) { struct folio *end_folio =3D folio_next(folio); struct swap_cluster_info *ci =3D NULL; struct folio *new_folio; int old_order =3D folio_order(folio); + enum ttu_flags ttu_flags =3D 0; struct list_lru_one *lru; struct lruvec *lruvec; bool dequeue_deferred; int ret =3D 0; =20 + if (unmap) + unmap_folio(folio); + local_irq_disable(); =20 /* @@ -3971,8 +3972,8 @@ static int __folio_freeze_split_unmapped_anon(struct = folio *folio, unsigned int list_lru_unlock(lru); rcu_read_unlock(); } - local_irq_enable(); - return -EAGAIN; + ret =3D -EAGAIN; + goto out_no_split; } =20 if (dequeue_deferred) { @@ -4023,15 +4024,21 @@ static int __folio_freeze_split_unmapped_anon(struc= t folio *folio, unsigned int lruvec_unlock(lruvec); if (ci) swap_cluster_unlock(ci); +out_no_split: local_irq_enable(); + if (unmap) { + if (!ret && !folio_is_device_private(folio)) + ttu_flags =3D TTU_USE_SHARED_ZEROPAGE; + remap_page(folio, 1 << old_order, ttu_flags); + } =20 return ret; } =20 -static int __folio_freeze_split_unmapped_file(struct folio *folio, unsigne= d int new_order, - struct page *split_at, struct xa_state *xas, - struct address_space *mapping, bool do_lru, - struct list_head *list, enum split_type split_type) +static int __folio_freeze_split_unmap_file(struct folio *folio, unsigned i= nt new_order, + struct page *split_at, struct xa_state *xas, + struct address_space *mapping, bool do_lru, + struct list_head *list, enum split_type split_type) { struct folio *end_folio =3D folio_next(folio); struct folio *new_folio, *next; @@ -4051,6 +4058,8 @@ static int __folio_freeze_split_unmapped_file(struct = folio *folio, unsigned int if (shmem_mapping(mapping)) end =3D shmem_fallocend(mapping->host, end); =20 + unmap_folio(folio); + xas_lock_irq(xas); =20 /* @@ -4175,7 +4184,6 @@ static int __folio_split(struct folio *folio, unsigne= d int new_order, struct anon_vma *anon_vma =3D NULL; int old_order =3D folio_order(folio); struct folio *new_folio, *next; - enum ttu_flags ttu_flags =3D 0; int ret; =20 VM_WARN_ON_ONCE_FOLIO(!folio_test_locked(folio), folio); @@ -4263,21 +4271,14 @@ static int __folio_split(struct folio *folio, unsig= ned int new_order, goto out_unlock; } =20 - unmap_folio(folio); - if (!is_anon) { - ret =3D __folio_freeze_split_unmapped_file(folio, new_order, split_at, &= xas, mapping, + ret =3D __folio_freeze_split_unmap_file(folio, new_order, split_at, &xas= , mapping, true, list, split_type); } else { - ret =3D __folio_freeze_split_unmapped_anon(folio, new_order, split_at, t= rue, - list, split_type); + ret =3D __folio_freeze_split_unmap_anon(folio, new_order, split_at, true, + true, list, split_type); } =20 - if (!ret && is_anon && !folio_is_device_private(folio)) - ttu_flags =3D TTU_USE_SHARED_ZEROPAGE; - - remap_page(folio, 1 << old_order, ttu_flags); - /* * Drop the mapping while the inode is still pinned. @folio stays * locked and present in the page cache until the loop below, so @@ -4360,8 +4361,8 @@ int folio_split_unmapped(struct folio *folio, unsigne= d int new_order) if (folio_expected_ref_count(folio) !=3D folio_ref_count(folio) - 1) return -EAGAIN; =20 - return __folio_freeze_split_unmapped_anon(folio, new_order, &folio->page, - false, NULL, SPLIT_TYPE_UNIFORM); + return __folio_freeze_split_unmap_anon(folio, new_order, &folio->page, fa= lse, + false, NULL, SPLIT_TYPE_UNIFORM); } =20 /* --=20 2.55.0 From nobody Tue Sep 29 11:56:00 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5937D40DB38 for ; Fri, 7 Aug 2026 21:17:50 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786137470; cv=none; b=S2+nr1u/o2Bo/SIEARatIuqva2XWGMb0dhio+JsgTQ6s9srX+YyhF//97RuLUrytZ9MXXHvpnJJZRkznwp1HRAv0vCuMxQ3FKgh7zmjVsp8PuYnqPuMghEvPx/Y0Rvwbq7pHXAEMNKSUc/MqMGtFl3Yqr3LNY3mBzj2dm/vh3cA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786137470; c=relaxed/simple; bh=5VM1gvICKoR0rxwdh5lvSYtecezYifXxD2IyBpdRBbg=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=JvfoKOpjykduyZcd3Nd4c6JjXm5kq9jAKobmDv+edHtx54Iu6fSSMK/d3ffD2QE90kjprMolybZ63G0mbkEtNYhp8cTHvLAQT+G5TASlSFhRRnoUh4swUo834WJxov8bwys50fXpHrkCBcPlWuR+uM/yX8y+K7q1B1tAqiexWDA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=LqSswVzZ; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="LqSswVzZ" Received: by smtp.kernel.org (Postfix) with ESMTPS id 2F7A1C19425; Fri, 7 Aug 2026 21:17:50 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1786137470; bh=5VM1gvICKoR0rxwdh5lvSYtecezYifXxD2IyBpdRBbg=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=LqSswVzZfBoVoduboQvvWrOG8LS3ET0Nryk2d2lkF5wh+LTXC2AJm6Xgh/mz2Ob7S 2S6g+2sjU6WQ5wuY3ttsjfF9IEFKg0PgU+/8IUil7FgZU6Js+qpcGJbJ4aDZQim0mD 2AefkCQ1iVEquwDaBepkNpycn+pQaF6c40ufUl1E+ALiGOTE1QT4UPYI3dwedWzzVC v8g1S0MowL5ih0egrIq2nKha0Ex2Uu1xDUp4PcpZ36fBf+sDHUJgf9S8yX1O9PpUqu Nq2UaIWxvOELlEcYt8u+QQjrV7OlADlqOeZR7Un9X3+4AdMJkxs4Q0YqLc9mNs4LKf WFz/JDeUe4tLA== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 1A0E6C2A09B; Fri, 7 Aug 2026 21:17:50 +0000 (UTC) From: Kairui Song via B4 Relay Date: Sat, 08 Aug 2026 05:17:09 +0800 Subject: [PATCH RFC 08/13] mm/huge_memory: move anon_vma and filemap management into split helpers Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260808-swap-thp-cleanup-v1-8-689939a7ccc3@tencent.com> References: <20260808-swap-thp-cleanup-v1-0-689939a7ccc3@tencent.com> In-Reply-To: <20260808-swap-thp-cleanup-v1-0-689939a7ccc3@tencent.com> To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, Andrew Morton , David Hildenbrand , Lorenzo Stoakes , Zi Yan , Baolin Wang , "Liam R. Howlett" , Nico Pache , Ryan Roberts , Dev Jain , Lance Yang , Usama Arif , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Chris Li , Kemeng Shi , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , Kairui Song X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1786137467; l=9016; i=kasong@tencent.com; s=kasong-sign-tencent; h=from:subject:message-id; bh=Tdpu3yM2jJD/dLTIquK8QVzk8LI0pYKr2L3DEgrEOZY=; b=vE5S16rk8H6tQL7nMbKAIEEiUR03ZTW7o4PEm5gfVq7qO35uzHPGQa0PaRtrA3KeCbh9yKL8e qEexL8JeNZ1CjPd10TUfXx4AfXATeKc3jN/GArLUOjRiAlTTylK34wK X-Developer-Key: i=kasong@tencent.com; a=ed25519; pk=kCdoBuwrYph+KrkJnrr7Sm1pwwhGDdZKcKrqiK8Y1mI= X-Endpoint-Received: by B4 Relay for kasong@tencent.com/kasong-sign-tencent with auth_id=562 X-Original-From: Kairui Song Reply-To: kasong@tencent.com From: Kairui Song Only anon split needs vma info, and only file split needs the filemap handling. Move the related code into separate helpers so they are genuinely more self-contained. Signed-off-by: Kairui Song Reviewed-by: Zi Yan --- mm/huge_memory.c | 177 +++++++++++++++++++++++++--------------------------= ---- 1 file changed, 79 insertions(+), 98 deletions(-) diff --git a/mm/huge_memory.c b/mm/huge_memory.c index c0115841d1a0..f0ea6e5f53f7 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -3939,12 +3939,33 @@ static int __folio_freeze_split_unmap_anon(struct f= olio *folio, unsigned int new struct swap_cluster_info *ci =3D NULL; struct folio *new_folio; int old_order =3D folio_order(folio); + struct anon_vma *anon_vma =3D NULL; enum ttu_flags ttu_flags =3D 0; struct list_lru_one *lru; struct lruvec *lruvec; bool dequeue_deferred; int ret =3D 0; =20 + /* + * Unmap/remap needs the anon_vma. The caller does not necessarily + * hold an mmap_lock that would prevent the anon_vma from + * disappearing, so we first take a reference and lock it. This is + * similar to folio_lock_anon_vma_read() except the write lock is + * taken to serialize against parallel split or collapse. + */ + if (unmap) { + anon_vma =3D folio_get_anon_vma(folio); + if (!anon_vma) + return -EBUSY; + anon_vma_lock_write(anon_vma); + } + + /* Racy check if we can split the page, before the optional unmap. */ + if (folio_expected_ref_count(folio) !=3D folio_ref_count(folio) - 1) { + ret =3D -EAGAIN; + goto out_unlock; + } + if (unmap) unmap_folio(folio); =20 @@ -4031,21 +4052,58 @@ static int __folio_freeze_split_unmap_anon(struct f= olio *folio, unsigned int new ttu_flags =3D TTU_USE_SHARED_ZEROPAGE; remap_page(folio, 1 << old_order, ttu_flags); } +out_unlock: + if (anon_vma) { + anon_vma_unlock_write(anon_vma); + put_anon_vma(anon_vma); + } =20 return ret; } =20 static int __folio_freeze_split_unmap_file(struct folio *folio, unsigned i= nt new_order, - struct page *split_at, struct xa_state *xas, - struct address_space *mapping, bool do_lru, + struct page *split_at, bool do_lru, struct list_head *list, enum split_type split_type) { + struct address_space *mapping =3D folio->mapping; + XA_STATE(xas, &mapping->i_pages, folio->index); struct folio *end_folio =3D folio_next(folio); struct folio *new_folio, *next; int nr_shmem_dropped =3D 0; + unsigned int min_order; struct lruvec *lruvec; pgoff_t end =3D 0; - int ret; + gfp_t gfp; + int ret =3D 0; + + min_order =3D mapping_min_folio_order(mapping); + if (new_order < min_order) + return -EINVAL; + + gfp =3D current_gfp_context(mapping_gfp_mask(mapping) & GFP_RECLAIM_MASK); + if (!filemap_release_folio(folio, gfp)) + return -EBUSY; + + mapping_set_update(&xas, mapping); + + if (split_type =3D=3D SPLIT_TYPE_UNIFORM) { + int old_order =3D folio_order(folio); + + xas_set_order(&xas, folio->index, new_order); + xas_split_alloc(&xas, folio, old_order, gfp); + if (xas_error(&xas)) { + ret =3D xas_error(&xas); + goto fail_free; + } + } + + i_mmap_lock_read(mapping); + + /* Racy check if we can split the page, before unmap_folio() */ + if (folio_expected_ref_count(folio) !=3D folio_ref_count(folio) - 1) { + ret =3D -EAGAIN; + goto fail_mmap_unlock; + } =20 /* *__split_unmapped_folio() may need to trim off pages beyond @@ -4060,13 +4118,13 @@ static int __folio_freeze_split_unmap_file(struct f= olio *folio, unsigned int new =20 unmap_folio(folio); =20 - xas_lock_irq(xas); + xas_lock_irq(&xas); =20 /* * Check if the folio is present in page cache. * We assume all tail are present too, if folio is there. */ - if (xas_load(xas) !=3D folio) { + if (xas_load(&xas) !=3D folio) { ret =3D -EAGAIN; goto fail; } @@ -4093,7 +4151,7 @@ static int __folio_freeze_split_unmap_file(struct fol= io *folio, unsigned int new if (do_lru) lruvec =3D folio_lruvec_lock(folio); =20 - ret =3D __split_unmapped_folio(folio, new_order, split_at, xas, + ret =3D __split_unmapped_folio(folio, new_order, split_at, &xas, mapping, split_type); =20 /* @@ -4145,9 +4203,19 @@ static int __folio_freeze_split_unmap_file(struct fo= lio *folio, unsigned int new lruvec_unlock(lruvec); =20 fail: - xas_unlock_irq(xas); + xas_unlock_irq(&xas); +fail_mmap_unlock: if (nr_shmem_dropped) shmem_uncharge(mapping->host, nr_shmem_dropped); + /* + * Drop the mapping while the inode is still pinned. @folio stays + * locked and present in the page cache, so eviction cannot free + * the inode yet, nothing past this point may touch the inode or + * the mapping. + */ + i_mmap_unlock_read(mapping); +fail_free: + xas_destroy(&xas); return ret; } =20 @@ -4176,12 +4244,9 @@ static int __folio_split(struct folio *folio, unsign= ed int new_order, struct page *split_at, struct page *lock_at, struct list_head *list, enum split_type split_type) { - XA_STATE(xas, &folio->mapping->i_pages, folio->index); struct folio *end_folio =3D folio_next(folio); bool is_anon =3D folio_test_anon(folio); struct mem_cgroup *memcg, *old_memcg; - struct address_space *mapping =3D NULL; - struct anon_vma *anon_vma =3D NULL; int old_order =3D folio_order(folio); struct folio *new_folio, *next; int ret; @@ -4212,84 +4277,12 @@ static int __folio_split(struct folio *folio, unsig= ned int new_order, memcg =3D get_mem_cgroup_from_folio(folio); old_memcg =3D set_active_memcg(memcg); =20 - if (is_anon) { - /* - * The caller does not necessarily hold an mmap_lock that would - * prevent the anon_vma disappearing so we first we take a - * reference to it and then lock the anon_vma for write. This - * is similar to folio_lock_anon_vma_read except the write lock - * is taken to serialise against parallel split or collapse - * operations. - */ - anon_vma =3D folio_get_anon_vma(folio); - if (!anon_vma) { - ret =3D -EBUSY; - goto out; - } - anon_vma_lock_write(anon_vma); - mapping =3D NULL; - } else { - unsigned int min_order; - gfp_t gfp; - - mapping =3D folio->mapping; - min_order =3D mapping_min_folio_order(mapping); - if (new_order < min_order) { - ret =3D -EINVAL; - goto out; - } - - gfp =3D current_gfp_context(mapping_gfp_mask(mapping) & - GFP_RECLAIM_MASK); - - if (!filemap_release_folio(folio, gfp)) { - ret =3D -EBUSY; - goto out; - } - - mapping_set_update(&xas, mapping); - - if (split_type =3D=3D SPLIT_TYPE_UNIFORM) { - xas_set_order(&xas, folio->index, new_order); - xas_split_alloc(&xas, folio, old_order, gfp); - if (xas_error(&xas)) { - ret =3D xas_error(&xas); - goto out; - } - } - - anon_vma =3D NULL; - i_mmap_lock_read(mapping); - } - - /* - * Racy check if we can split the page, before unmap_folio() will - * split PMDs - */ - if (folio_expected_ref_count(folio) !=3D folio_ref_count(folio) - 1) { - ret =3D -EAGAIN; - goto out_unlock; - } - - if (!is_anon) { - ret =3D __folio_freeze_split_unmap_file(folio, new_order, split_at, &xas= , mapping, - true, list, split_type); - } else { + if (is_anon) ret =3D __folio_freeze_split_unmap_anon(folio, new_order, split_at, true, true, list, split_type); - } - - /* - * Drop the mapping while the inode is still pinned. @folio stays - * locked and present in the page cache until the loop below, so - * eviction cannot free the inode yet; @lock_at is not enough, it may - * be a tail beyond EOF that the split already dropped from the page - * cache. Nothing past this point may touch the inode or the mapping. - */ - if (mapping) { - i_mmap_unlock_read(mapping); - mapping =3D NULL; - } + else + ret =3D __folio_freeze_split_unmap_file(folio, new_order, split_at, + true, list, split_type); =20 /* * Unlock all after-split folios except the one containing @@ -4310,19 +4303,10 @@ static int __folio_split(struct folio *folio, unsig= ned int new_order, free_folio_and_swap_cache(new_folio); } =20 -out_unlock: - if (anon_vma) { - anon_vma_unlock_write(anon_vma); - put_anon_vma(anon_vma); - } - if (mapping) - i_mmap_unlock_read(mapping); -out: /* restore to caller's old_memcg */ set_active_memcg(old_memcg); mem_cgroup_put(memcg); out_no_memcg: - xas_destroy(&xas); if (is_pmd_order(old_order)) count_vm_event(!ret ? THP_SPLIT_PAGE : THP_SPLIT_PAGE_FAILED); count_mthp_stat(old_order, !ret ? MTHP_STAT_SPLIT : MTHP_STAT_SPLIT_FAILE= D); @@ -4358,9 +4342,6 @@ int folio_split_unmapped(struct folio *folio, unsigne= d int new_order) VM_WARN_ON_ONCE_FOLIO(!folio_test_large(folio), folio); VM_WARN_ON_ONCE_FOLIO(!folio_test_anon(folio), folio); =20 - if (folio_expected_ref_count(folio) !=3D folio_ref_count(folio) - 1) - return -EAGAIN; - return __folio_freeze_split_unmap_anon(folio, new_order, &folio->page, fa= lse, false, NULL, SPLIT_TYPE_UNIFORM); } --=20 2.55.0 From nobody Tue Sep 29 11:56:00 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5FE20412BF1 for ; Fri, 7 Aug 2026 21:17:50 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786137470; cv=none; b=alYYayEbmMLUtedqjABrLsqxyyhBzWRbME5progF3fW83G4jG4zOe3VEPBddI4VyQwzjHCHlN2a1s3Ap9rUt1xQ/lkKocgRSaWz48y/UKFzoVP8+63N4xS1xsUlNDZ1kzSrpXS9iYf66Y5WCmpRoT4gmc0JnwwLNNg5wd8jY+pY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786137470; c=relaxed/simple; bh=UBCaRIj71EFnMJbiAnbHRp0dByToFm+xvtA3P1554AM=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=s84Knv1+oVOL03EqYEcs7ry1HCrTsYBN3areh3xj8SpwRCY7rATfq5EdIlsECTGMJ7qxhtDqucfif2DBCvFh9JzS0INjvCl5maQQvNXc/3D766G2MEHH/5rX8DHBsOH5ExfDqy0Pzgwm0bO/0JcxVUDgMeLFXRTBg7O8nwr2hPY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=mGr1d5x3; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="mGr1d5x3" Received: by smtp.kernel.org (Postfix) with ESMTPS id 43203C2BCB8; Fri, 7 Aug 2026 21:17:50 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1786137470; bh=UBCaRIj71EFnMJbiAnbHRp0dByToFm+xvtA3P1554AM=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=mGr1d5x3h1etfnIFtuyaOzHL0IF6eO7NQdcMYJRX0FRpTXgyb479Yl+VQ0Pf4Qgpv qcqFvrTaL2bWFV9dk66CZk8Dy8cQ9MyowCrqJBxXIKz5SaAur2SPOBmlH7gvzoX/GH w+zgkBZLd8HQvs4xp2XY5AAGvM+P0CRXbP3zcemYwqQAfhLtq6P8Aybn4H8sb6Nz85 HU3ZGzOfo6hHLwCo72QYTQJcXG4YPZXq2mDdHYLQFJG4MoT6YYHbLJyxdsn/vGUvGl vT3VZOuZmcXMg2sFYmB/DjPHltOlGfLgJ7lXbTNfmODEKfeKhNDwWIqrrsOOXb9Nm3 5eBFu/PWIk1ZA== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 2CDCCC5ACAB; Fri, 7 Aug 2026 21:17:50 +0000 (UTC) From: Kairui Song via B4 Relay Date: Sat, 08 Aug 2026 05:17:10 +0800 Subject: [PATCH RFC 09/13] mm/huge_memory: move memcg switch into the file split helper Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260808-swap-thp-cleanup-v1-9-689939a7ccc3@tencent.com> References: <20260808-swap-thp-cleanup-v1-0-689939a7ccc3@tencent.com> In-Reply-To: <20260808-swap-thp-cleanup-v1-0-689939a7ccc3@tencent.com> To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, Andrew Morton , David Hildenbrand , Lorenzo Stoakes , Zi Yan , Baolin Wang , "Liam R. Howlett" , Nico Pache , Ryan Roberts , Dev Jain , Lance Yang , Usama Arif , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Chris Li , Kemeng Shi , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , Kairui Song X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1786137467; l=3611; i=kasong@tencent.com; s=kasong-sign-tencent; h=from:subject:message-id; bh=suVTq55eCBPCmmmCFw2hiAz9weomIDRQxX0WoFXbtvE=; b=xvJl0GoXEACDF9WKcNSCH+92kCsa0Ivw5DsKzRhSxuaiPMjVQolHTZ1i8jjrAajeC55ZBwAgB E5Icsv9rd4cCHrKQrHdZzukhuf4/YQWkz+r4D1+hRJjqf+t5at4buGR X-Developer-Key: i=kasong@tencent.com; a=ed25519; pk=kCdoBuwrYph+KrkJnrr7Sm1pwwhGDdZKcKrqiK8Y1mI= X-Endpoint-Received: by B4 Relay for kasong@tencent.com/kasong-sign-tencent with auth_id=562 X-Original-From: Kairui Song Reply-To: kasong@tencent.com From: Kairui Song The xarray node allocations in __folio_freeze_split_unmap_file() need to be charged to the folio's memcg, so move the memcg switch from __folio_split() into the helper. The anon split helper and the after-split folio freeing perform no chargeable allocations, so no memcg handling is left in __folio_split(). Rename its out_no_memcg label to out. Signed-off-by: Kairui Song Acked-by: Zi Yan --- mm/huge_memory.c | 36 +++++++++++++++++++----------------- 1 file changed, 19 insertions(+), 17 deletions(-) diff --git a/mm/huge_memory.c b/mm/huge_memory.c index f0ea6e5f53f7..ab2bb29748d3 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -4068,6 +4068,7 @@ static int __folio_freeze_split_unmap_file(struct fol= io *folio, unsigned int new struct address_space *mapping =3D folio->mapping; XA_STATE(xas, &mapping->i_pages, folio->index); struct folio *end_folio =3D folio_next(folio); + struct mem_cgroup *memcg, *old_memcg; struct folio *new_folio, *next; int nr_shmem_dropped =3D 0; unsigned int min_order; @@ -4080,9 +4081,18 @@ static int __folio_freeze_split_unmap_file(struct fo= lio *folio, unsigned int new if (new_order < min_order) return -EINVAL; =20 + /* + * Switch to folio's memcg as xarray node allocation can happen and + * needs to charge to it. + */ + memcg =3D get_mem_cgroup_from_folio(folio); + old_memcg =3D set_active_memcg(memcg); + gfp =3D current_gfp_context(mapping_gfp_mask(mapping) & GFP_RECLAIM_MASK); - if (!filemap_release_folio(folio, gfp)) - return -EBUSY; + if (!filemap_release_folio(folio, gfp)) { + ret =3D -EBUSY; + goto fail_free; + } =20 mapping_set_update(&xas, mapping); =20 @@ -4215,6 +4225,9 @@ static int __folio_freeze_split_unmap_file(struct fol= io *folio, unsigned int new */ i_mmap_unlock_read(mapping); fail_free: + /* Restore the previously active memcg */ + set_active_memcg(old_memcg); + mem_cgroup_put(memcg); xas_destroy(&xas); return ret; } @@ -4246,7 +4259,6 @@ static int __folio_split(struct folio *folio, unsigne= d int new_order, { struct folio *end_folio =3D folio_next(folio); bool is_anon =3D folio_test_anon(folio); - struct mem_cgroup *memcg, *old_memcg; int old_order =3D folio_order(folio); struct folio *new_folio, *next; int ret; @@ -4256,27 +4268,20 @@ static int __folio_split(struct folio *folio, unsig= ned int new_order, =20 if (folio !=3D page_folio(split_at) || folio !=3D page_folio(lock_at)) { ret =3D -EINVAL; - goto out_no_memcg; + goto out; } =20 if (new_order >=3D old_order) { ret =3D -EINVAL; - goto out_no_memcg; + goto out; } =20 ret =3D folio_check_splittable(folio, new_order, split_type); if (ret) { VM_WARN_ONCE(ret =3D=3D -EINVAL, "Tried to split an unsplittable folio"); - goto out_no_memcg; + goto out; } =20 - /* - * switch to folio's memcg as xarray node allocation can happen and - * needs to charge to it. - */ - memcg =3D get_mem_cgroup_from_folio(folio); - old_memcg =3D set_active_memcg(memcg); - if (is_anon) ret =3D __folio_freeze_split_unmap_anon(folio, new_order, split_at, true, true, list, split_type); @@ -4303,10 +4308,7 @@ static int __folio_split(struct folio *folio, unsign= ed int new_order, free_folio_and_swap_cache(new_folio); } =20 - /* restore to caller's old_memcg */ - set_active_memcg(old_memcg); - mem_cgroup_put(memcg); -out_no_memcg: +out: if (is_pmd_order(old_order)) count_vm_event(!ret ? THP_SPLIT_PAGE : THP_SPLIT_PAGE_FAILED); count_mthp_stat(old_order, !ret ? MTHP_STAT_SPLIT : MTHP_STAT_SPLIT_FAILE= D); --=20 2.55.0 From nobody Tue Sep 29 11:56:00 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 74C3E420472 for ; Fri, 7 Aug 2026 21:17:50 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786137470; cv=none; b=BlowLrGQd/9kGAh9mp+Ph0i9hama8+hFVaUJWjOhHMRpIeYQAnLuHOnb8FdmePnlKPu+1VDgrgCyt+PG1QyRK3o5kwPoO7hKrDH5e7bClNqAnzoA6NOtbHWfctRjDZRks5yIvfEjwwxkvV2OYbjEM9+pgBJUPDOVpIrLg4TNxPQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786137470; c=relaxed/simple; bh=I9ZwUvZv1ol36H9/Rq5s+TZN0X2Qruo9ioQ8LVL316Q=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=KNoto34X5Yvms9oenoodDv4BwVLW7jW4p4c8vs+DjCLLBYVjSJJoB8fHw9LmBLR9M8TJlxrVxCo3MFkN0b4p0E6lEfLmNfFmCCqDCWLQUMfwbD08Nj0oUPkeFowmDk6MVovogD/plcfm0UNsYhCF2dqotECZ2fQCTu59d1KmxeQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=eUnSm5ET; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="eUnSm5ET" Received: by smtp.kernel.org (Postfix) with ESMTPS id 567F0C2BCFA; Fri, 7 Aug 2026 21:17:50 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1786137470; bh=I9ZwUvZv1ol36H9/Rq5s+TZN0X2Qruo9ioQ8LVL316Q=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=eUnSm5ETfyN6i6crs9GnEMS4ry6XhaR5jKbidnB6YiwOoRH4+gLpS6YRTeezgwdLh iu+PoVVJclxJ3QAOHTvKN70SIgGkQkY64utKvN74fo67yCYIp1PiZHAAiaOWCuhKux hkXnnKATGvlYRjHHSoXIG5lWH4CKhs8hglsBYA/qPLQsgEy2eu3IiPAtFIU4j8Yu/w Z9YzYJM8mPYLRw9EpzoudZiU1/9z1MY5fhUKBk1SqQpAOdJa3Elp26F3n9TwSipxRk Qe7JtFpCfZ7apN88TSEwHIkUeeKCn24GLVGkIldQK4uOlJDlaODVYr3THnQgPM2ivk Fv6JI7VbdxmqQ== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 42F23C5ACD1; Fri, 7 Aug 2026 21:17:50 +0000 (UTC) From: Kairui Song via B4 Relay Date: Sat, 08 Aug 2026 05:17:11 +0800 Subject: [PATCH RFC 10/13] mm/huge_memory: allow splitting mappingless swap cache folios Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260808-swap-thp-cleanup-v1-10-689939a7ccc3@tencent.com> References: <20260808-swap-thp-cleanup-v1-0-689939a7ccc3@tencent.com> In-Reply-To: <20260808-swap-thp-cleanup-v1-0-689939a7ccc3@tencent.com> To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, Andrew Morton , David Hildenbrand , Lorenzo Stoakes , Zi Yan , Baolin Wang , "Liam R. Howlett" , Nico Pache , Ryan Roberts , Dev Jain , Lance Yang , Usama Arif , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Chris Li , Kemeng Shi , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , Kairui Song X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1786137467; l=5534; i=kasong@tencent.com; s=kasong-sign-tencent; h=from:subject:message-id; bh=RlM8mpHFdOY4ZKngaHLh0+gwh65R/G4errquU7C7PiU=; b=MDqPUHQlHc8YqRsgbFOiw/gs+clVxgCEnHoReVoVTyTkxhaAB4NXX3KxsTtkkEa6Yz3qns2nk DzI/Sd7a3E6DfST1h2ASXmwqXdemTCpnhxl4nCsz0a3lS1sJqa06qma X-Developer-Key: i=kasong@tencent.com; a=ed25519; pk=kCdoBuwrYph+KrkJnrr7Sm1pwwhGDdZKcKrqiK8Y1mI= X-Endpoint-Received: by B4 Relay for kasong@tencent.com/kasong-sign-tencent with auth_id=562 X-Original-From: Kairui Song Reply-To: kasong@tencent.com From: Kairui Song Lift the restriction that kept swap cache folios without a mapping from being split. All the underlying infrastructure is sound against that with a few more tweaks, no reason to block it anymore. Also rename the split helper, which now handles mappingless swap cache folios that are yet to be anon, or may actually belong to shmem. In either case there is not much difference in how they would be split. A non-anon swap cache folio that still has a mapping (e.g. a shmem swap cache folio) remains rejected up front: it would need both its page cache and swap cache entries updated on split, which the split helpers do not do, and there would be little benefit in doing so. Signed-off-by: Kairui Song --- mm/huge_memory.c | 40 +++++++++++++++++++++++----------------- 1 file changed, 23 insertions(+), 17 deletions(-) diff --git a/mm/huge_memory.c b/mm/huge_memory.c index ab2bb29748d3..b80d0db63225 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -3881,12 +3881,10 @@ int folio_check_splittable(struct folio *folio, uns= igned int new_order, VM_WARN_ON_FOLIO(!folio_test_locked(folio), folio); /* * Folios that just got truncated cannot get split. Signal to the - * caller that there was a race. - * - * TODO: this will also currently refuse folios without a mapping in the - * swapcache (shmem or to-be-anon folios). + * caller that there was a race. A mappingless swap cache folio + * has no page cache entries to update, so it is fine to split. */ - if (!folio->mapping && !is_anon) + if (!folio->mapping && !is_swapcache) return -EBUSY; =20 /* order-1 is not supported for anonymous THP. */ @@ -3931,11 +3929,12 @@ static unsigned int folio_cache_ref_count(const str= uct folio *folio) return folio_nr_pages(folio); } =20 -static int __folio_freeze_split_unmap_anon(struct folio *folio, unsigned i= nt new_order, - struct page *split_at, bool do_lru, bool unmap, - struct list_head *list, enum split_type split_type) +static int __folio_freeze_split_unmap(struct folio *folio, unsigned int ne= w_order, + struct page *split_at, bool do_lru, bool anon_unmap, + struct list_head *list, enum split_type split_type) { struct folio *end_folio =3D folio_next(folio); + bool is_anon =3D folio_test_anon(folio); struct swap_cluster_info *ci =3D NULL; struct folio *new_folio; int old_order =3D folio_order(folio); @@ -3953,7 +3952,7 @@ static int __folio_freeze_split_unmap_anon(struct fol= io *folio, unsigned int new * similar to folio_lock_anon_vma_read() except the write lock is * taken to serialize against parallel split or collapse. */ - if (unmap) { + if (anon_unmap) { anon_vma =3D folio_get_anon_vma(folio); if (!anon_vma) return -EBUSY; @@ -3966,7 +3965,7 @@ static int __folio_freeze_split_unmap_anon(struct fol= io *folio, unsigned int new goto out_unlock; } =20 - if (unmap) + if (anon_unmap) unmap_folio(folio); =20 local_irq_disable(); @@ -3977,8 +3976,11 @@ static int __folio_freeze_split_unmap_anon(struct fo= lio *folio, unsigned int new * a 0-ref folio, it assumes it beat folio_put() to the list * lock and must clean up the LRU state - the same dequeue we * will do below as part of the split. + * + * Only anon folios are ever queued on the deferred split list, + * so non-anon folios (mappingless swapcache) never need dequeuing. */ - dequeue_deferred =3D old_order > 1; + dequeue_deferred =3D old_order > 1 && is_anon; if (dequeue_deferred) { struct mem_cgroup *memcg; =20 @@ -4047,13 +4049,13 @@ static int __folio_freeze_split_unmap_anon(struct f= olio *folio, unsigned int new swap_cluster_unlock(ci); out_no_split: local_irq_enable(); - if (unmap) { + if (anon_unmap) { if (!ret && !folio_is_device_private(folio)) ttu_flags =3D TTU_USE_SHARED_ZEROPAGE; remap_page(folio, 1 << old_order, ttu_flags); } out_unlock: - if (anon_vma) { + if (anon_unmap) { anon_vma_unlock_write(anon_vma); put_anon_vma(anon_vma); } @@ -4257,6 +4259,7 @@ static int __folio_split(struct folio *folio, unsigne= d int new_order, struct page *split_at, struct page *lock_at, struct list_head *list, enum split_type split_type) { + bool is_swapcache =3D folio_test_swapcache(folio); struct folio *end_folio =3D folio_next(folio); bool is_anon =3D folio_test_anon(folio); int old_order =3D folio_order(folio); @@ -4283,8 +4286,11 @@ static int __folio_split(struct folio *folio, unsign= ed int new_order, } =20 if (is_anon) - ret =3D __folio_freeze_split_unmap_anon(folio, new_order, split_at, true, - true, list, split_type); + ret =3D __folio_freeze_split_unmap(folio, new_order, split_at, true, + true, list, split_type); + else if (is_swapcache) + ret =3D __folio_freeze_split_unmap(folio, new_order, split_at, true, + false, list, split_type); else ret =3D __folio_freeze_split_unmap_file(folio, new_order, split_at, true, list, split_type); @@ -4344,8 +4350,8 @@ int folio_split_unmapped(struct folio *folio, unsigne= d int new_order) VM_WARN_ON_ONCE_FOLIO(!folio_test_large(folio), folio); VM_WARN_ON_ONCE_FOLIO(!folio_test_anon(folio), folio); =20 - return __folio_freeze_split_unmap_anon(folio, new_order, &folio->page, fa= lse, - false, NULL, SPLIT_TYPE_UNIFORM); + return __folio_freeze_split_unmap(folio, new_order, &folio->page, false, + false, NULL, SPLIT_TYPE_UNIFORM); } =20 /* --=20 2.55.0 From nobody Tue Sep 29 11:56:00 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8F929443E3A for ; Fri, 7 Aug 2026 21:17:50 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786137470; cv=none; b=Vplp4tfoMPjH60D0VPUPf54GOaj2kpm/oQuDOG5VUyFLWucx3JLYnDO12j2mQPrJwkYeOVkYqZfm8uii6Xhbt88RykNbF3cPs6BmtgR8kiWF3ZMFHn/6ySYBsmQfgMi2SUbhdyDvv7F3q1TWgZPprjwMxtvazBTPXxVedXpRQK0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786137470; c=relaxed/simple; bh=99s4a0IEhVPV4G78O2RX0b+40o7pPLKSxH5+kDAOrxQ=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=Fm97VPJsVetFDC5bZhB40qmXbZakqmCxkUxGS/Scvqy7KNBI78aR8O2cu0vd4UJuaOaLJK7rSozwCRTrACsZkRk0vlMGcI3EX0eknUbJUdmadfKX/ynLhZ23G+L+xyz/ZjLrnVYxzxCkQbPwjFCzONW9mgEfQODDCZlJdxMicjI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=tNh5Xzoa; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="tNh5Xzoa" Received: by smtp.kernel.org (Postfix) with ESMTPS id 70AD8C2BD01; Fri, 7 Aug 2026 21:17:50 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1786137470; bh=99s4a0IEhVPV4G78O2RX0b+40o7pPLKSxH5+kDAOrxQ=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=tNh5XzoawqV6AZ+Fpaw+p9eFndVOCNn0dmfjss0RKx0RN/lHi6gIVflk4gxRjpM0S NiPNtTAFmmLmwAHdyD1U1fNVG0qeTmVFHIYI//px1H9hBW3F5r/K1nbovsTvV2QD3s qcQdJ/79+Mo/SY2Wk36S5Z/ldOoqZgcQ7mgS169W9WI9jPO67gFmhkwMGLZwK9rxer pkOa3JlXF/yOr1Wsfv+C1bCgL10U9dA786S/URpyIrb66sQWex43RxrqfIBcvnVBin uS9l/hqnT/uCWrbU7wpHuUtXcYqRCt1k8J61Whh7II0MQUanlnEBzqt/jdLp8chtqd VDrPx22kHVp2g== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 5AA9FC5ACD3; Fri, 7 Aug 2026 21:17:50 +0000 (UTC) From: Kairui Song via B4 Relay Date: Sat, 08 Aug 2026 05:17:12 +0800 Subject: [PATCH RFC 11/13] mm/huge_memory: clean up after-split folio freeing in __folio_split Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260808-swap-thp-cleanup-v1-11-689939a7ccc3@tencent.com> References: <20260808-swap-thp-cleanup-v1-0-689939a7ccc3@tencent.com> In-Reply-To: <20260808-swap-thp-cleanup-v1-0-689939a7ccc3@tencent.com> To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, Andrew Morton , David Hildenbrand , Lorenzo Stoakes , Zi Yan , Baolin Wang , "Liam R. Howlett" , Nico Pache , Ryan Roberts , Dev Jain , Lance Yang , Usama Arif , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Chris Li , Kemeng Shi , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , Kairui Song X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1786137467; l=1476; i=kasong@tencent.com; s=kasong-sign-tencent; h=from:subject:message-id; bh=eR9RHfz1xKpkHU93Mx9UnV8mmj6dwPjTBi2fADo8ZCE=; b=MRnjAlkZDWKZcwh9ZZOzHGfILV2xObDl270yF4y4lR624UzRVeTzjlIe0yhnhMF8Tdg/4sun/ 4PbZdRJJClJBW6FZBZO7PeCkCAslk3cUW4f6CZXwkeolC1ahBi+9oZi X-Developer-Key: i=kasong@tencent.com; a=ed25519; pk=kCdoBuwrYph+KrkJnrr7Sm1pwwhGDdZKcKrqiK8Y1mI= X-Endpoint-Received: by B4 Relay for kasong@tencent.com/kasong-sign-tencent with auth_id=562 X-Original-From: Kairui Song Reply-To: kasong@tencent.com From: Kairui Song Replace free_folio_and_swap_cache() with an explicit folio_free_swap() and folio_put() in the after-split loop. free_folio_and_swap_cache() unlocks the folio, then free_swap_cache() must trylock it again and re-check folio_mapped() before freeing the swap cache entries; if the trylock loses a race, the entries are left behind even though the folio reference is dropped. The sub folios are still locked and unmapped here, so just directly call folio_free_swap() directly under the lock, unlock and drop the reference. This makes the swap cache freeing deterministic and the reference drop explicit. Signed-off-by: Kairui Song Reviewed-by: Zi Yan --- mm/huge_memory.c | 6 ++++-- 1 file changed, 4 insertions(+), 2 deletions(-) diff --git a/mm/huge_memory.c b/mm/huge_memory.c index b80d0db63225..39c91c8e5bc8 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -4304,14 +4304,16 @@ static int __folio_split(struct folio *folio, unsig= ned int new_order, if (new_folio =3D=3D page_folio(lock_at)) continue; =20 - folio_unlock(new_folio); /* * Subpages whose mapping has been zapped may be freed * earlier, but freeing them requires taking the * lru_lock, so we defer put_page() on tail pages until * after the split completes. */ - free_folio_and_swap_cache(new_folio); + if (is_swapcache) + folio_free_swap(new_folio); + folio_unlock(new_folio); + folio_put(new_folio); } =20 out: --=20 2.55.0 From nobody Tue Sep 29 11:56:00 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A659444683B for ; Fri, 7 Aug 2026 21:17:50 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786137470; cv=none; b=oqiTSPstYJCY9CV7Vhj0E8ib/iy++QQDvYqBm7UbhKRKbcB/f416WXmGWs34QY15VB4pPGmYTb76PqhSwmmRxxWJgVeFPeKpQRdzz9FMrPdsll7q84hviJ0A5N4V/uyKtH34WlrfH65QU3nm5RkQorSV70tH6JakqMv/hXLuE3g= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786137470; c=relaxed/simple; bh=yvf2N/efdIKEj5X5qDj9wcYMNlN3bo6H9UHt3Cv+FOM=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=QUm9rUrhoiQDF6EPpdKRj6GKzt0b6iRDkw/FAt3aFhPyrmFQKI1ofW2RzyjgY0UhN6nCTkRCrdWpjINom/9nWbsm85shdAG3WCVERJ6GVx/w/ZxSPwpFZZI+UxPplF/M4v7mffhfptswVQA+3CmUpQcX62m0eNs88w2v+Q7oz2E= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=YHr0xwRl; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="YHr0xwRl" Received: by smtp.kernel.org (Postfix) with ESMTPS id 892FBC2BCFB; Fri, 7 Aug 2026 21:17:50 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1786137470; bh=yvf2N/efdIKEj5X5qDj9wcYMNlN3bo6H9UHt3Cv+FOM=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=YHr0xwRlJAM2p4ts4Fur2Lgd6hpmWZC7vvWlHSnsJ4NNro1NR7zusHJ1BKTLwVvoB Zn0yJXpc3arFNNff4GPvRLGl3jJ35NXQ+8vaEW8POA4xMQ8Po3ShLtGpJ5ZajNknIr 8NK7dTrAw86F2wCpgpXx8YSOtxcKidIBo/yzF88is1Qj5enLdobNGDu1BqzhlJE/z8 devmNl6/SYYoyZuX0wtJLhYasRCpoFxvBkvnFfgwfhwyHOXzg7v/qr8B58yWqnW++y t3S2cYhH1ORxquiUf5bcMDlFpAOo/Kw6+nLzQx6Le9omhcXTX+EKZP0oDRzRKBgDh/ uIo/GgswsB8Lg== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 753FEC5ACAB; Fri, 7 Aug 2026 21:17:50 +0000 (UTC) From: Kairui Song via B4 Relay Date: Sat, 08 Aug 2026 05:17:13 +0800 Subject: [PATCH RFC 12/13] mm/huge_memory: lift order-0 restriction for swapcache split Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260808-swap-thp-cleanup-v1-12-689939a7ccc3@tencent.com> References: <20260808-swap-thp-cleanup-v1-0-689939a7ccc3@tencent.com> In-Reply-To: <20260808-swap-thp-cleanup-v1-0-689939a7ccc3@tencent.com> To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, Andrew Morton , David Hildenbrand , Lorenzo Stoakes , Zi Yan , Baolin Wang , "Liam R. Howlett" , Nico Pache , Ryan Roberts , Dev Jain , Lance Yang , Usama Arif , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Chris Li , Kemeng Shi , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , Kairui Song X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1786137467; l=4236; i=kasong@tencent.com; s=kasong-sign-tencent; h=from:subject:message-id; bh=7SjtfqfZedSJz4xGQv/+c/6mWUxdnqqN70J6iR/8Q6g=; b=OM9J+C2BnomsrB8R3lY++FWH2mGqLYGpQGghKyKGjk11dm3+MVvTipZ8Nt/+n8N1vq2htd2d0 XDBmdKu+FeQAWv/dgW9asdmXhB8LBOKbhdDqSlFssXLP1nG0ZS3AyQI X-Developer-Key: i=kasong@tencent.com; a=ed25519; pk=kCdoBuwrYph+KrkJnrr7Sm1pwwhGDdZKcKrqiK8Y1mI= X-Endpoint-Received: by B4 Relay for kasong@tencent.com/kasong-sign-tencent with auth_id=562 X-Original-From: Kairui Song Reply-To: kasong@tencent.com From: Kairui Song The restriction that swapcache folios can only be uniformly split to order 0 dates back to when the swap cache was managed via address_space mapping (swap_address_space). The old split loop only created order-0 sub-folios with a fixed stride, so non-uniform split and non-zero order were rightfully blocked. After the swap cache switched to swap table under a cluster lock, __swap_cache_replace_folio already gained the ability to replace any number of entries for any sub-folio size in one cluster, and the old swap_address_space locking and limit was removed. The restriction became obsolete but persisted through multiple refactorings. Drop it now: swapcache folios can be split to any supported order with either uniform or non-uniform split, except order-1 which is not supported for anon folios. Mappingless swap cache folios could be either anon or shmem, so for now we just simply forbid order-1 for all swapcache. Signed-off-by: Kairui Song Acked-by: Zi Yan --- mm/huge_memory.c | 32 +++++++++++++------------------- 1 file changed, 13 insertions(+), 19 deletions(-) diff --git a/mm/huge_memory.c b/mm/huge_memory.c index 39c91c8e5bc8..dba53fbb93a8 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -3797,6 +3797,7 @@ static int __split_unmapped_folio(struct folio *folio= , int new_order, struct address_space *mapping, enum split_type split_type) { const bool is_anon =3D folio_test_anon(folio); + const bool is_swapcache =3D folio_test_swapcache(folio); int old_order =3D folio_order(folio); int start_order =3D split_type =3D=3D SPLIT_TYPE_UNIFORM ? new_order : ol= d_order - 1; struct folio *old_folio =3D folio; @@ -3811,8 +3812,8 @@ static int __split_unmapped_folio(struct folio *folio= , int new_order, split_order--) { int nr_new_folios =3D 1UL << (old_order - split_order); =20 - /* order-1 anonymous folio is not supported */ - if (is_anon && split_order =3D=3D 1) + /* order-1 anonymous or swapcache folio is not supported */ + if ((is_anon || is_swapcache) && split_order =3D=3D 1) continue; =20 if (mapping) { @@ -3887,21 +3888,14 @@ int folio_check_splittable(struct folio *folio, uns= igned int new_order, if (!folio->mapping && !is_swapcache) return -EBUSY; =20 - /* order-1 is not supported for anonymous THP. */ - if (is_anon && new_order =3D=3D 1) - return -EINVAL; - /* - * swapcache folio could only be split to order 0 - * - * non-uniform split creates after-split folios with orders from - * folio_order(folio) - 1 to new_order, making it not suitable for any - * swapcache folio split. Only uniform split to order-0 can be used - * here. + * Order-1 is unsupported: anon folios need subpage 2 for the + * deferred split list, hybrid shmem & swap cache folios are not + * splittable, and a splittable mappingless swap cache folio could + * be either anon or shmem, which we cannot tell apart. */ - if ((split_type =3D=3D SPLIT_TYPE_NON_UNIFORM || new_order) && is_swapcac= he) { + if ((is_anon || is_swapcache) && new_order =3D=3D 1) return -EINVAL; - } =20 if (is_huge_zero_folio(folio)) return -EINVAL; @@ -4372,11 +4366,11 @@ int folio_split_unmapped(struct folio *folio, unsig= ned int new_order) * GUP pins, will result in the folio not getting split; instead, the c= aller * will receive an -EAGAIN. * - * 4) @new_order > 1, usually. Splitting to order-1 anonymous folios is not - * supported for non-file-backed folios, because folio->_deferred_list,= which - * is used by partially mapped folios, is stored in subpage 2, but an o= rder-1 - * folio only has subpages 0 and 1. File-backed order-1 folios are supp= orted, - * since they do not use _deferred_list. + * 4) @new_order > 1, usually. Order-1 is not supported for anon or swapca= che + * folios: anon folios need subpage 2 for _deferred_list, which order-1 + * folios lack, and a swapcache folio may become anon once faulted in. + * File-backed order-1 folios are supported, since they do not use + * _deferred_list. * * After splitting, the caller's folio reference will be transferred to @p= age, * resulting in a raised refcount of @page after this call. The other page= s may --=20 2.55.0 From nobody Tue Sep 29 11:56:00 2026 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BBA42448B8F for ; Fri, 7 Aug 2026 21:17:50 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786137470; cv=none; b=X3lDo2wtz1HYiT7bcB11fFKKU7hCETVWw/FSpAMbatvMPqd/d4131SFLA3OiZ3PW0qRJ4oZO7S71iJLIzp8WbC6U56++dvS53uLPkDZeIR1lyU/hUWbg+fAFSjxsSEDBphoL7i3GTLXAYWcWOOFp/pB1/RM5ROCeU2s8T3FkUoY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786137470; c=relaxed/simple; bh=wdQnGoGIB9o4VEZAXvUmfNak2doCNqndMD3WezFuI88=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=eBasdYJBhRHJ0I1ssn2jMRfF2u8LMhZZC6/8ir9eyANEEs3LNZRr+4kyxnCsbuYaPFIbmdN8II1a+VFafHbVkgthiOVCp0S0swncMDZIdKOt+v3/SliZr9oBaTrDRhua42lMb7K/cxXs3jgIKjVB1kFGh44zD0PoZSy9iNilzLw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=eUKih9pt; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="eUKih9pt" Received: by smtp.kernel.org (Postfix) with ESMTPS id 9EAF3C19425; Fri, 7 Aug 2026 21:17:50 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1786137470; bh=wdQnGoGIB9o4VEZAXvUmfNak2doCNqndMD3WezFuI88=; h=From:Date:Subject:References:In-Reply-To:To:Cc:Reply-To:From; b=eUKih9ptdiudblkV8V+3RVkR3OewHsd5hg1v2hdCthshCJ5LsR21pApOTFbGqeXFK phaOC0Mfm4bilR+bSRMqbImDlYNYVYVae1PrZkCIP25F2rl/cK2WIoLsw0bV3Ys8W1 u4iVjMXlwDKP1BcQlz3oKLzLTc82aVwBZnj3xYMxtho1HYhqpdrR5IAnM01a/Wsrd+ XwfPSmfxhIQnvxxzdK5Z0iKeuH4gRKt4P2iHNuyB15J4+7K8P+05Tw51dsnnk3tz66 BJ6BzWy++phWTutZ7LbgkF6C6IKqIi3qyEo+iSsckJU5nggZSJovugjkURc7raTWre LGFeVT5vpgjcQ== Received: from aws-us-west-2-korg-lkml-1.web.codeaurora.org (localhost.localdomain [127.0.0.1]) by smtp.lore.kernel.org (Postfix) with ESMTP id 8C21CC5ACD1; Fri, 7 Aug 2026 21:17:50 +0000 (UTC) From: Kairui Song via B4 Relay Date: Sat, 08 Aug 2026 05:17:14 +0800 Subject: [PATCH RFC 13/13] mm/huge_memory: count only swap cache refs in anon folio split Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable Message-Id: <20260808-swap-thp-cleanup-v1-13-689939a7ccc3@tencent.com> References: <20260808-swap-thp-cleanup-v1-0-689939a7ccc3@tencent.com> In-Reply-To: <20260808-swap-thp-cleanup-v1-0-689939a7ccc3@tencent.com> To: linux-mm@kvack.org Cc: linux-kernel@vger.kernel.org, Andrew Morton , David Hildenbrand , Lorenzo Stoakes , Zi Yan , Baolin Wang , "Liam R. Howlett" , Nico Pache , Ryan Roberts , Dev Jain , Lance Yang , Usama Arif , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Chris Li , Kemeng Shi , Nhat Pham , Baoquan He , Barry Song , Youngjun Park , Kairui Song X-Mailer: b4 0.15.2 X-Developer-Signature: v=1; a=ed25519-sha256; t=1786137467; l=4434; i=kasong@tencent.com; s=kasong-sign-tencent; h=from:subject:message-id; bh=6txV4z9147MCY3v3Y112JpwmGarss54hNdYDLN2lhpw=; b=XXdvVVkJnsuSqgAphHl1AuGTvhydbgaE1nQ6pBzl6XSvLM9/V7veYYRMKGOANAqlaGxgJnV1b LerXbuHpQtTB/9FMV3d50V87DM7K2D5+CgItBEfmSefWnMWMgXfF5CV X-Developer-Key: i=kasong@tencent.com; a=ed25519; pk=kCdoBuwrYph+KrkJnrr7Sm1pwwhGDdZKcKrqiK8Y1mI= X-Endpoint-Received: by B4 Relay for kasong@tencent.com/kasong-sign-tencent with auth_id=562 X-Original-From: Kairui Song Reply-To: kasong@tencent.com From: Kairui Song Only __folio_freeze_split_unmap() sees anon folios and swap cache folios now. The file split helper only handles page cache folios, which hold exactly folio_nr_pages() references. Rename folio_cache_ref_count() to folio_swapcache_ref_count() and drop the anon check so the helper counts what its name says. The file split helper now uses folio_nr_pages() directly. Signed-off-by: Kairui Song Reviewed-by: Zi Yan --- mm/huge_memory.c | 35 +++++++++++++++-------------------- 1 file changed, 15 insertions(+), 20 deletions(-) diff --git a/mm/huge_memory.c b/mm/huge_memory.c index dba53fbb93a8..0d70ee3017c4 100644 --- a/mm/huge_memory.c +++ b/mm/huge_memory.c @@ -3915,10 +3915,10 @@ int folio_check_splittable(struct folio *folio, uns= igned int new_order, return 0; } =20 -/* Number of folio references from the pagecache or the swapcache. */ -static unsigned int folio_cache_ref_count(const struct folio *folio) +/* Number of folio references from the swapcache. */ +static unsigned int folio_swapcache_ref_count(const struct folio *folio) { - if (folio_test_anon(folio) && !folio_test_swapcache(folio)) + if (!folio_test_swapcache(folio)) return 0; return folio_nr_pages(folio); } @@ -3984,7 +3984,7 @@ static int __folio_freeze_split_unmap(struct folio *f= olio, unsigned int new_orde folio_nid(folio), &memcg); } =20 - if (!folio_ref_freeze(folio, folio_cache_ref_count(folio) + 1)) { + if (!folio_ref_freeze(folio, folio_swapcache_ref_count(folio) + 1)) { if (dequeue_deferred) { list_lru_unlock(lru); rcu_read_unlock(); @@ -4027,7 +4027,7 @@ static int __folio_freeze_split_unmap(struct folio *f= olio, unsigned int new_orde new_folio =3D folio_next(new_folio)) { zone_device_private_split_cb(folio, new_folio); folio_ref_unfreeze(new_folio, - folio_cache_ref_count(new_folio) + 1); + folio_swapcache_ref_count(new_folio) + 1); if (do_lru) lru_add_split_folio(folio, new_folio, lruvec, list); if (ci) @@ -4035,7 +4035,7 @@ static int __folio_freeze_split_unmap(struct folio *f= olio, unsigned int new_orde } =20 zone_device_private_split_cb(folio, NULL); - folio_ref_unfreeze(folio, folio_cache_ref_count(folio) + 1); + folio_ref_unfreeze(folio, folio_swapcache_ref_count(folio) + 1); =20 if (do_lru) lruvec_unlock(lruvec); @@ -4064,6 +4064,7 @@ static int __folio_freeze_split_unmap_file(struct fol= io *folio, unsigned int new struct address_space *mapping =3D folio->mapping; XA_STATE(xas, &mapping->i_pages, folio->index); struct folio *end_folio =3D folio_next(folio); + long old_nr_pages =3D folio_nr_pages(folio); struct mem_cgroup *memcg, *old_memcg; struct folio *new_folio, *next; int nr_shmem_dropped =3D 0; @@ -4135,22 +4136,16 @@ static int __folio_freeze_split_unmap_file(struct f= olio *folio, unsigned int new goto fail; } =20 - if (!folio_ref_freeze(folio, folio_cache_ref_count(folio) + 1)) { + if (!folio_ref_freeze(folio, old_nr_pages + 1)) { ret =3D -EAGAIN; goto fail; } =20 - if (folio_test_pmd_mappable(folio) && - new_order < HPAGE_PMD_ORDER) { - int nr =3D folio_nr_pages(folio); - - if (folio_test_swapbacked(folio)) { - lruvec_stat_mod_folio(folio, - NR_SHMEM_THPS, -nr); - } else { - lruvec_stat_mod_folio(folio, - NR_FILE_THPS, -nr); - } + if (folio_test_pmd_mappable(folio) && new_order < HPAGE_PMD_ORDER) { + if (folio_test_swapbacked(folio)) + lruvec_stat_mod_folio(folio, NR_SHMEM_THPS, -old_nr_pages); + else + lruvec_stat_mod_folio(folio, NR_FILE_THPS, -old_nr_pages); } =20 /* lock lru list/PageCompound, ref frozen by page_ref_freeze */ @@ -4175,7 +4170,7 @@ static int __folio_freeze_split_unmap_file(struct fol= io *folio, unsigned int new next =3D folio_next(new_folio); =20 folio_ref_unfreeze(new_folio, - folio_cache_ref_count(new_folio) + 1); + folio_nr_pages(new_folio) + 1); =20 if (do_lru) lru_add_split_folio(folio, new_folio, lruvec, list); @@ -4203,7 +4198,7 @@ static int __folio_freeze_split_unmap_file(struct fol= io *folio, unsigned int new * Otherwise, a parallel folio_try_get() can grab @folio * and its caller can see stale page cache entries. */ - folio_ref_unfreeze(folio, folio_cache_ref_count(folio) + 1); + folio_ref_unfreeze(folio, folio_nr_pages(folio) + 1); =20 if (do_lru) lruvec_unlock(lruvec); --=20 2.55.0