From nobody Mon Sep 28 23:56:19 2026 Received: from fhigh-a7-smtp.messagingengine.com (fhigh-a7-smtp.messagingengine.com [103.168.172.158]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 922221C84BB; Sat, 15 Aug 2026 01:59:10 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.158 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759152; cv=none; b=Lmtw1R1Hdw5JVQfN89T+ull3T4NKaLfJc5VwNXX0X3MYgu1wqJMU35Lacn0pSfaZvttUzXYJ+XDB0Lk3ixcPHj3rgcDYoAecakQs6bUpfZQOisdJvT8JMeGZivAMAudHBI6/jLYh58oE8O726f3S66xaksKpODJN//9HBmRj68A= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759152; c=relaxed/simple; bh=yZPs9aFWHZjbXOkRvX67/dIHGI0NCldpZ6X6TIYwKhM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=sbiFHsNHK0SiwBcmCqZUJxs2zuxDQvBIIbmgWYGbjRDyG5pR4h5WUbGjoUVJ9/saljMNqHoq0UwE95viWgw4l1mJA776SWKjwRru9LCD6yuNOaBWmO97V/cR/ZiO6e7iAOXBryA7M3SzhketwNCEx8lAuFzxOssSOMK2P76/Mbw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=ZgG47TsU; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=EOH9+raa; arc=none smtp.client-ip=103.168.172.158 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="ZgG47TsU"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="EOH9+raa" Received: from phl-compute-10.internal (phl-compute-10.internal [10.202.2.50]) by mailfhigh.phl.internal (Postfix) with ESMTP id 9617F14000DB; Fri, 14 Aug 2026 21:59:09 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-10.internal (MEProxy); Fri, 14 Aug 2026 21:59:09 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786759149; x= 1786845549; bh=zex4bQI/K2jisTXcoQfqqVox+onfSYGfs1VVhpIHxHY=; b=Z gG47TsUXDc74J69BdsHiaAZ20D9iw7bshmUv1PeRXCFvYTNzdErF8Ps8KU+XK1ct qLY/zKY70LAQCU8YnvxCZzMaCO7wI4DyqiZYNfzm6D2yOQ/+tF1Wm9NPoyRoVZWc TERtXHpx/6wnE/S6A84QIerMPR4g47FL1SKYl8BE65Zn7cvlRWyOLlLaNsf9AC3K 0pno4FyKRKyrHLjfexndPrF7HKYHUSfTkRcXu6Pfn753ecWpiOsrgbmPZaBly/mF 03sXxGMhoy2Jlirnm7TS3y4hEtcGbDSOo3lt7ikRaEKj213yNAXEQMRx4FmdrX6F VnsCJUp1SyGp7x6mJRpLw== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786759149; x=1786845549; bh=z ex4bQI/K2jisTXcoQfqqVox+onfSYGfs1VVhpIHxHY=; b=EOH9+raail+Gpppe3 MTz5no349gPpJZHFYC02c+VJXQQlye/7oZHMjWtxwrTgyVJnhc4sQ5bAc3ESM5e0 QxHWeT7xjmLnXk0dxtWCKGQb+TgGirMOT7Pih/UXQleMRLYAdyKMwDUegknPtJeb QP0tRMARucpNNOysGSRswuVN5zDxqgo6l4v6HiYP7rxqKOkJhixPKgg29Ao7Hrce B6bog9SZNzkLlJnrw7Vjtznr1Q35tA3AGyyAONl7Kvedb4FQzbtz4hkv+sIBhbwf cnpkT2pASLFbOcs6F9a3TJ3E+IGZ7TjLNF0Fpn8AnfSVFpE76QmkSqOCbCVP2bMb twauA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEYpEQiBdLKlCaLp317TOYqYwO99jnuX4YrBujo3RuW+IIGXdYj4qE1ezcGRvbtOZ hc7RfnfkdlTehb5Dkvk4VZTcyRAqm2CcEjaldkIQshIttrB2HSeWyAKQfsR+ykmymFv8qO aC/+wh3joJiQccaWn5NjVrqLVVMYdoVO37A5ZL6SnR75seJlCquk1NTOiA0EvXHIVSJ25Q iFONSccd/wEzNfQZlOvvQQf5bCKiGIALXd0Rd37Vcnrh56q8DfWwmvEu2ZgBGSMa525iaP HkMF5MKO3LDiNiHXJVHjo5f/L4XJttiAOfdXhljkaHgZI649okv5M81/eeF8HfhCZr1Ucq SUpY4/UrUyKSVyvxdWdkzRcAxLXEajvkyHarmVKmBKmPh2QSdP0hn98URAdd5ulxcsE/oG HExvPs8wbd4/c1OD0GCTIQk37WW10K4nncE94L1KpyMGnmEpXO+avOn6/IhVBaxW9RvQIk /IEopIoH4a7oK/nS/pMUtNzIIK2N6EUUROvpKfw2YEamJAq1mzV1j85eudJuXLlCj1ymCl DfwOdwHrJDFbVLGOJX9Pbb/kjXnvRCmFZ0bK8vCEtFQTadr9Mdf2cu+dpzUjrEqRyPJWN5 9kzVnnSKnSddmpbT6pvKRYBGeu90PJNZYVPIn+RtVt6DI1aCGz7rLsIxrSwA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 14 Aug 2026 21:59:07 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v4 01/19] selftests/mm: raise the khugepaged test-case cap Date: Sat, 15 Aug 2026 02:58:43 +0100 Message-ID: <20260815015901.1236937-2-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260815015901.1236937-1-kirill@shutemov.name> References: <20260815015901.1236937-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" TEST() ends the run with "MAX_TEST_CASES is too small" when the table fills, and the table holds 64. A full invocation already registers 63, so the next case added anywhere aborts the whole suite before a single test runs. Raise the cap to 256. The table is a static array of small structs, so the room costs nothing worth counting. Assisted-by: Claude-Code:claude-opus-5 Acked-by: Usama Arif Reviewed-by: Mike Rapoport (Microsoft) Signed-off-by: Kiryl Shutsemau (Meta) Acked-by: Lorenzo Stoakes (ARM) --- tools/testing/selftests/mm/khugepaged.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index d3a53673e1f9..6cfec2b940ac 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -1290,7 +1290,7 @@ struct test_case { test_fn fn; }; =20 -#define MAX_TEST_CASES 64 +#define MAX_TEST_CASES 256 static struct test_case test_cases[MAX_TEST_CASES]; static int nr_test_cases; =20 --=20 2.54.0 From nobody Mon Sep 28 23:56:19 2026 Received: from fout-a6-smtp.messagingengine.com (fout-a6-smtp.messagingengine.com [103.168.172.149]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 37D8D346E44; Sat, 15 Aug 2026 01:59:12 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.149 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759153; cv=none; b=FEMA7UNTDzKy4WzZI5h/mCAwsHnERWw/m7TFwbCjao1iRLwOFijBmz0Q+pzsgHH+GdyFQtggatVRle/W5085S6pjv6F6ftbm5JwcsPR82+rprtvPlu37/iRw4r2MVkMoPQkHhPiSgvOf+1acmJlYQk4emwYcY7l09R41JHGixPs= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759153; c=relaxed/simple; bh=OXxNY3EuFFVkH6euR5PM0ctHXZ1HbJ+ey54oO6xSKng=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=c0LCgqQcqE4IVdO2if+iNQtP8kgd4BzEbb+peH9ssvIZ3MMO9XK+lq97z4E+vdCxlNVakaQjfWejg+uBJ8vlyDm6hRBX7Od8BH8wvyfCjlqzMXviZNTrYuVtWtWbcWxvOrt6ZCbEzEKHReA2a+0y60JyJrlgerCIdN5QLyZOYEk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=JLIYuLmt; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=N6vI9+9M; arc=none smtp.client-ip=103.168.172.149 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="JLIYuLmt"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="N6vI9+9M" Received: from phl-compute-03.internal (phl-compute-03.internal [10.202.2.43]) by mailfout.phl.internal (Postfix) with ESMTP id 53490EC021C; Fri, 14 Aug 2026 21:59:11 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-03.internal (MEProxy); Fri, 14 Aug 2026 21:59:11 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786759151; x= 1786845551; bh=l0UVYTdnMDN95QCW3MAbNL5JjNdjlhisIs8DxT1bDN8=; b=J LIYuLmtVX/3gOR7fmDqT1M0k1GthK/F/Fix3+ncbRus2ugay0vTmnRW+rCbwkraB pU/Ah7OCaoIqpbdMIIwH7zkh1h5RwG6hB9LIr/pv5HQXPEqEB9ueL0lcGiawuvQk l6zW15OxEHGdhdlu0H4VyJS6F0LsD0P25e07OkZA+x1yUqODzzkshRYPfwjnaqod RO601aJg8/a73W+vmpeZIYVBJ73op1mO3QyK2tPwDBzudBCsdKDtJ110Jo4ikna6 T5j5eI52U8Dtupe9XMgZT5U2YXaD5+VYGk2xeyPhoMKwQhfLN6em4nQPGdPYCvVE Pv6G9gDVlfmuSd1ADICrQ== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786759151; x=1786845551; bh=l 0UVYTdnMDN95QCW3MAbNL5JjNdjlhisIs8DxT1bDN8=; b=N6vI9+9MbYO6PRqKI 5txNJZ4S7SFvk39vJFz2bBzr3S50NXDCtnaHHVyIQhrYH0tDeVvucUwpTjl6YCbp 1LuvLhwMK4OePf053V0E03QMD5z6wqu5PXa9Ji40p3UUc2425Z56uqdnywIpYhE8 Exv3W3wn2PdC9Rfh64ZgF5/ja8KF5zsAr7DtG8p3SfcIHvaC9//5OGhHitpDbPnK 4gHecnzKYNCvN6UwdlKU7MVoi/6a7BS6ai0k5WVp33VZt99N85q9L4v9xw/ms1P4 s/uFfUIxBnI4KRGsNAqYlO8HKrzT0C0gQayv4Jnp2vTrOyeIzrZBEV12bVYw3GSA jUCDQ== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEYpEQiBdLKlCaLp317TOYqYwO99jnuX4YrBujo3RuW+IIGXdYj4qE1ezcGRvbtOZ hc7RfnfkdlTehb5Dkvk4VZTcyRAqm2CcEjaldkIQshIttrB2HSeWyAKQfsR+ykmymFv8qO aC/+wh3joJiQccaWn5NjVrqLVVMYdoVO37A5ZL6SnR75seJlCquk1NTOiA0EvXHIVSJ25Q iFONSccd/wEzNfQZlOvvQQf5bCKiGIALXd0Rd37Vcnrh56q8DfWwmvEu2ZgBGSMa525iaP HkMF5MKO3LDiNiHXJVHjo5f/L4XJttiAOfdXhljkaHgZI649okv5M81/eeF8HfhCZr1UAV t1ky7+JLBWzuoA0+1UuhthaJDKx/AJ5gC5tRaHE4ASCIDqAo/+HMX4IhequuAa5CReNL6l FT6eWS0bgjZ/1y3XwbN5wvPtBBvj9u4Dat0nWRZ+88AiEQvZ0/CGMji3azaLAD5AkwGcsE PTycNDqzNLkIUshyvcVXUQlB9T4rQcf6mNQ5fDbD078evTC0lHYGW2tTZzGYM/PJK6M5yT 68VA2q4x0WRNvSONYwIxcRfoeCilJAQT4wbLtdhaTwUZ9/c7I6vSuCLJwFvgcUaSuy6EMa 15VpSFvVzkjxYHqqIP+jKjpfyh5Ltfrd2M2n3X7quhkTKwbGW+8liDZabGpg X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 14 Aug 2026 21:59:10 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v4 02/19] selftests/mm: skip collapse_compound_extreme() where the PMD is too large Date: Sat, 15 Aug 2026 02:58:44 +0100 Message-ID: <20260815015901.1236937-3-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260815015901.1236937-1-kirill@shutemov.name> References: <20260815015901.1236937-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" collapse_compound_extreme() builds a PTE table full of distinct PTE-mapped compound pages by cycling hpage_pmd_nr fault-time THPs through mremap. It therefore needs hpage_pmd_nr PMD-order allocations in a row. That is fine at a 2M PMD (4K base pages) or a 32M one (16K). A 512M PMD -- arm64 with 64K base pages -- makes each of those an order-13 allocation, which the allocator cannot reliably hand out even once, let alone 8192 times. The failure is not a quiet one: the case calls ksft_exit_fail_msg(), so the whole binary stops and every case after it is lost. Skip the case where the PMD is larger than 32M. The MADV_COLLAPSE cases still cover PMD-order collapse on those configurations, and 4K and 16K PMDs are unaffected. Assisted-by: Claude-Code:claude-opus-5 Reviewed-by: Mike Rapoport (Microsoft) Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) Acked-by: Lorenzo Stoakes (ARM) --- tools/testing/selftests/mm/khugepaged.c | 10 ++++++++++ 1 file changed, 10 insertions(+) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index 6cfec2b940ac..dd924edd8557 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -940,6 +940,16 @@ static void collapse_compound_extreme(struct collapse_= context *c, struct mem_ops void *p; int i; =20 + /* + * The test needs hpage_pmd_nr PMD-order allocations, which is likely to + * fail for large PMD sizes. Skip if the PMD size is over 32M. + */ + if (hpage_pmd_size > (32UL << 20)) { + ksft_test_result_skip("%s: PMD too large for fault-time THP construction= \n", + __func__); + return; + } + p =3D ops->setup_area(1); ksft_print_msg("Construct PTE page table full of different PTE-mapped com= pound pages\n"); for (i =3D 0; i < hpage_pmd_nr; i++) { --=20 2.54.0 From nobody Mon Sep 28 23:56:19 2026 Received: from fhigh-a7-smtp.messagingengine.com (fhigh-a7-smtp.messagingengine.com [103.168.172.158]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 987323264D4; Sat, 15 Aug 2026 01:59:14 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.158 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759156; cv=none; b=lvFYrTZqDJLgpJ9DWzACNthRQvE9xru0zJOPYZEUnMUolU/pdDegZYkAbYWEYrcpHBxEsDNo8Wqw2pCxvO9ato5DuMclGa6HjwUichG0Gf8+69N8xBNCo7RLD+VN8afqzBWvqH45aE411Qd0mnRHn6NyCBW3NydyaW4/p4VoSKE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759156; c=relaxed/simple; bh=PXDXjjZJq8F3+mlJDJ7CyW23HiLJFxhx5NF61eOZxCA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=S687aMMtZPVTbgpgbIB8FKWOnTwHjKPiW3PuaVW037QaBMYqxCrPtnd3sy8iZlMZCS4kELTKYpdRPxpg372Suwzf5QwuHVRvO+PhOe0W95dVSAaXL596Melld1moGTxrSE+wa0DK5afmr75ZJv7gX/2kayoRm4RTwWktWv5Sv+4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=ycL0gDqt; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=GGCr6WBx; arc=none smtp.client-ip=103.168.172.158 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="ycL0gDqt"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="GGCr6WBx" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfhigh.phl.internal (Postfix) with ESMTP id 5EA4F140011B; Fri, 14 Aug 2026 21:59:13 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-06.internal (MEProxy); Fri, 14 Aug 2026 21:59:13 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786759153; x= 1786845553; bh=7okFjne0ixAjtjgoZ2s//LBwIwAIBw11vdB2BGCtpCE=; b=y cL0gDqtgPktLYY5aC2FRRb+f7yuCgU8MdWuA0SE7klljNCwNCLsPFn/la/FNvbZF rZGR3cojZqugyPWWRAIuMwKMc2OBYqdPdg/0S8VflgFXzewbEtaQOtDvyZedDI8v SqKRM01cojAM85GkYGtyAmqokDw/8MotmKDo+07Sjp32zyPE/HfCiBVjQaDreRpA YX2wTYK4uvhBwd6D8Gmn/czLgVC/Rc6bTTK2BuNJml818+4L4vA4lLlVcUwFHtuV dC8CKDBXtMfsWA3vCwcb0t0Q1iVWXh0642XG/EhgURvpkQ0bbMy/JU+GO6DV4Abm otKTwGnQ0CmAnhHRpJa5g== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786759153; x=1786845553; bh=7 okFjne0ixAjtjgoZ2s//LBwIwAIBw11vdB2BGCtpCE=; b=GGCr6WBxBJoc7zMXf eLylOaM5YjDd508p2ct27iG2BHTO9mGezIvqpKAZExGHnPwx7Z7ADNCr/y9Lwg++ DNCXcOM5oTF2w62Y10x7HWwXskLsLwrdCXriR5lp2jajbSj/aynyfm4yLXaDA5oj dYLG9UwRmHhyelgR91H7IzBR0AXLkSGDQwIGtIEPstUWMcYbDZMTht/jMgo4jsxl wDBA9hO3Tg9EL6mOU6fZTV+YuGKYymN25mxAim2gJENirnf2EgSGxxJVXoiDDbTT 9/GHkAzgzR/j7V5Ij0ZjUc/BKufY5hV6cyIA8DfLw92AvpjEXXeaX10qfriMm36X 1En5g== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEYpEQiBdLKlCaLp317TOYqYwO99jnuX4YrBujo3RuW+IIGXdYj4qE1ezcGRvbtOZ hc7RfnfkdlTehb5Dkvk4VZTcyRAqm2CcEjaldkIQshIttrB2HSeWyAKQfsR+ykmymFv8qO aC/+wh3joJiQccaWn5NjVrqLVVMYdoVO37A5ZL6SnR75seJlCquk1NTOiA0EvXHIVSJ25Q iFONSccd/wEzNfQZlOvvQQf5bCKiGIALXd0Rd37Vcnrh56q8DfWwmvEu2ZgBGSMa525iaP HkMF5MKO3LDiNiHXJVHjo5f/L4XJttiAOfdXhljkaHgZI649okv5M81/eeF8HfhCZr1Umx 2uYvkBKxn8OCitH4tiS7XbwgtTqmllHcukvryTHSfl6vlnOrj8XmN2F/Pp+RhmKDmPC4+L 3yFq30/qobTXJk46PT6w9GUjK37luykqi8lbz8fN5OsJvmGQ8lwD3kxlTDZ9j3RpCuP3WL 3/N2LainzeMLlxA33wTciNriKMZA5r/iKGejBBnu11DrzDFNEivHfhPNRNe0J6yc7L+3IJ zdLsn9uVZbARUSz2715eFtGk52TKGnDW9r5S9t4RtI6DejdZMEaXeqjQsRojU+ix4bgIuv 1P3KMEL0fom/+GhJUW4Obyvc4whTmI7i/V7wwRCeSIj7Z+wk1QfaxKD9FXCw X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 14 Aug 2026 21:59:12 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v4 03/19] selftests/mm: scale khugepaged's collapse wait with the PMD size Date: Sat, 15 Aug 2026 02:58:45 +0100 Message-ID: <20260815015901.1236937-4-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260815015901.1236937-1-kirill@shutemov.name> References: <20260815015901.1236937-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" wait_for_scan() gives every case the same three seconds, whatever the huge page costs to build. collapse_full() asks for four of them: 8M at a 2M PMD, but 2G at a 512M PMD -- arm64 with 64K base pages. Three seconds is thin at that size rather than generous. Across 80 runs of collapse_full() on arm64 with 64K pages the wait was half a second in 73 of them, with a tail to two seconds. The case has also timed out in a full matrix run, reporting a failure for a collapse that was still going. Keep three seconds as the floor and add a second per 128M collapsed. A 2M PMD is unchanged. A 512M PMD gets 19 seconds. arm64/64K: khugepaged all:anon 21 pass/1 fail -> 22 pass/0 fail. x86-64 is unchanged. Assisted-by: Claude-Code:claude-opus-5 Reviewed-by: Mike Rapoport (Microsoft) Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) Acked-by: Lorenzo Stoakes (ARM) --- tools/testing/selftests/mm/khugepaged.c | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index dd924edd8557..c499804a0ec4 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -561,8 +561,10 @@ static bool wait_for_scan(const char *msg, char *p, si= ze_t len, int nr_hpages, int collap_order, struct mem_ops *ops) { unsigned long hpage_size =3D page_size << collap_order; + /* Three seconds as a floor, plus a second per 128M to collapse */ + const unsigned long bytes =3D (unsigned long)nr_hpages * hpage_size; + int timeout =3D 6 + 2 * (bytes / (128UL << 20)); int full_scans; - int timeout =3D 6; /* 3 seconds */ =20 /* Sanity check */ if (!ops->check_huge(p, len, 0, hpage_size)) --=20 2.54.0 From nobody Mon Sep 28 23:56:19 2026 Received: from fout-a6-smtp.messagingengine.com (fout-a6-smtp.messagingengine.com [103.168.172.149]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E811C3446CE; Sat, 15 Aug 2026 01:59:15 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.149 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759157; cv=none; b=Zm2aTfjvlJiyh+HQINOtahe9m88yqf7Y9cPSkRNRqJzI9e/8/L+Zs5Bw554r6buWRpQTnZbC0VkYuTY/9N4kTOh0UgvOWLom5ESsELz1wqPSyp064R/9STqVxjO2unU2Gksya4Of8uI+VlFeDodCpkH61Kq8kuGLcXPjZOARoAU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759157; c=relaxed/simple; bh=efwAuGkc+QfDFtsrIl0ZRjfy4OsmTsx9SEjf76BeyiI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=dn/R2IcIGY9L0g7sxPPnGgiTeHADEx/TU5JGJzS5/kL/jpVfg4I1t1Mvidyh/wd92wcEnkrv5zLfNw7+w2TgbWJAgSkGKfWeQrJCss9GuO6fyDITVfIQhKi9mtjTVYXzLoZ3pnDPdPmlRqc2TtdKg595XuqG9VV8ImrY0GO3gHY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=qgG71gWC; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=C6EwAi+P; arc=none smtp.client-ip=103.168.172.149 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="qgG71gWC"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="C6EwAi+P" Received: from phl-compute-03.internal (phl-compute-03.internal [10.202.2.43]) by mailfout.phl.internal (Postfix) with ESMTP id 21AA3EC0225; Fri, 14 Aug 2026 21:59:15 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-03.internal (MEProxy); Fri, 14 Aug 2026 21:59:15 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786759155; x= 1786845555; bh=OagyuSQwqQVmdL1ycZwWlmABBXtstTQziQig8mkv9yQ=; b=q gG71gWCMufKBaweTLjyrV/2B+GqdIdc9XNhNl4EDaq8wdvzOjlZp+CCWTQYYN2SG ZuJK9NJvsnOUXPMWgTxGaUPM4MOdNigu9za7ZuHNG9irZxc7v8MJjpH3i3ZfHwO/ H/IKMpuzhDLUmUcbg++btiig/qJGvAC+WqtxsY5hloQJb7jZjAdPHIotB9KBQ805 uWcasdNt2iGu7ZIxzdGYPwNmgVbruDFB0CKmbtJMRSkgp1fITixmjIzqf2+jGm40 LFmwvS91qgptyogFF6qf7a/OQR0eQw1rBve0Kkhhx2zpJRim/JZr3Yyo8SQ5giMc KfGiEgJX5gqiBj5w9VWnQ== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786759155; x=1786845555; bh=O agyuSQwqQVmdL1ycZwWlmABBXtstTQziQig8mkv9yQ=; b=C6EwAi+PHIyxBEzN/ AswsushfQm78lDT+1wG9nTDM++fwGfXMe1BmmgZJ8C3yfoJalZ+SWEKv13LeGj8+ wQf5ktoxc4eoovmIvYHts0O95I11ARSu4C/X706yBhk722QunGFrkAtRIVttgBF1 8rrLOvPiDi4B2SsyXCz8nU2ioJW0PdUq9YTHKyVybxXTivzMZupyDKOIM5VUiqVo seZggNcKAylcqGviHCZoLsBQytcwIaX70S+Nyrl3RrUte0rQAirsrderXB6BQN7j ZtdWDaDAf9qWBJsUfVIC8pZsB9unYY0vmfLqyqnckzIELIreHf4KFDd2B/GQZCu6 iaI2A== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGF7teVIHNqoa5p2A8iL+trOmU4pExV0BsGEGBxavUqSy3vvyW/t2uVaC82ybAhbC FZe2jS418ApiE9B8lCjDVAvs2O/1hVxxIyd2QmVeomYKnxIQlXrdOkpVI2/A6aay1R5wM6 752pE1oD84VvsYvCM8wHufwfHzQzVg5WF64m0thPV1o6CCL0N8QD0816LOjK9LaF9KdRdp ShDgDyNTb/bU59JqJj/g3oSWXYBnQXQHEYh7kafl/PHtgQjoSit3F73LPXI/q1yj04b//y CUybFuAxuhlNqjceRELXNaCNxBn+qS4c4JK+cmo8T/HSoyzm/Kml9mcKob58Fnvz9v9LJG 902yBfSqobVpU5cPlkQIChl1GO+Rmll+tN/Ll/HfeQFiznCNQFe9p3Bw2O/4k1VVLkgnp/ fCNJ9XNlNHoWrBoslggvY1/O0wVbkEVo72nb108ebE4dHY77E43Q/D93CzrVpJI8lpUeSv S7ewJu+RGF/Xk9Qmbq2DrWS1sPrmAz00HaDwyBDmVKiPA7VPtKH+xk7t6VDMzkQ5JWyuWd mUDndNMaYibXgmLXP9wwtQwVc/wLMzkZevCOD+5Nv+emmXcwO5XmKGH1VXQVx3J9ELqhZn 53V+Fj/FwHyfPLhOipV4y47OWc7xGS+LvBp0c/5eGKo3LUBE7uLpEboq9Lgg X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 14 Aug 2026 21:59:14 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v4 04/19] selftests/mm: skip khugepaged page cache cases without a PMD folio Date: Sat, 15 Aug 2026 02:58:46 +0100 Message-ID: <20260815015901.1236937-5-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260815015901.1236937-1-kirill@shutemov.name> References: <20260815015901.1236937-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The page cache caps folio order at MAX_PAGECACHE_ORDER, which is smaller than the PMD order on arm64 with 64K pages, where a PMD is 512M. A PMD-sized page cache folio is impossible there, so MADV_COLLAPSE answers -EINVAL and khugepaged passes over the range. The shmem cases ask for a PMD-sized folio anyway, so four of them fail and the run bails out in the middle. MAX_PAGECACHE_ORDER is not shmem-specific: it caps every file folio. Skip both mem types where the cap is below the PMD order. Anonymous collapse is unaffected: its orders are not capped this way. Assisted-by: Claude-Code:claude-opus-5 Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) Reviewed-by: Mike Rapoport (Microsoft) --- tools/testing/selftests/mm/khugepaged.c | 24 ++++++++++++++++++++++++ 1 file changed, 24 insertions(+) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index c499804a0ec4..ec5c36a19d92 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -1357,6 +1357,30 @@ int main(int argc, char **argv) =20 setbuf(stdout, NULL); =20 + /* + * The page cache caps folio order at MAX_PAGECACHE_ORDER, which is + * below the PMD order on arm64 with 64K pages. A PMD-sized page cache + * folio is impossible there, so the kernel refuses these collapses by + * design and there is nothing to test. The cap is not shmem-specific: + * it rules out regular files too, and the per-order shmem_enabled + * controls exist for exactly the orders it allows, which is what makes + * them readable here. + */ + if (!(thp_shmem_supported_orders() & (1UL << hpage_pmd_order))) { + if (shmem_ops) { + ksft_print_msg("no PMD-order page cache folio: skipping shmem\n"); + shmem_ops =3D NULL; + } + if (read_only_file_ops) { + ksft_print_msg("no PMD-order page cache folio: skipping file\n"); + read_only_file_ops =3D NULL; + read_write_file_read_ops =3D NULL; + read_write_file_write_ops =3D NULL; + } + if (!anon_ops && !shmem_ops && !read_only_file_ops) + ksft_exit_skip("Nothing left to collapse into\n"); + } + default_settings.khugepaged.max_ptes_none =3D hpage_pmd_nr - 1; default_settings.khugepaged.max_ptes_swap =3D hpage_pmd_nr / 8; default_settings.khugepaged.max_ptes_shared =3D hpage_pmd_nr / 2; --=20 2.54.0 From nobody Mon Sep 28 23:56:19 2026 Received: from fout-a6-smtp.messagingengine.com (fout-a6-smtp.messagingengine.com [103.168.172.149]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CB43C34A791; Sat, 15 Aug 2026 01:59:17 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.149 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759161; cv=none; b=sJi2ZPkUuJ5avLf6j8J8zCJW5fYgYVk/yiCpvOvnF0KFBcxvrmapG9/LmG+YsVl5x7FCNtqRWwLRDNGg1uGivKIhXeGuM14Ihu8gk3538Q+nzjmIkNEDguwfW8oidgY0VeE76EYCY4hZpuTKshEd0cw+WGLGBm0NBHBQbsg2hlU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759161; c=relaxed/simple; bh=+QBmOchJZd0Ibl2TQFvBSlNpZPNzygSRb6mv49lsJxY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=fg4GmKwq+XQhYH3hBpE/5wl3rD4D4Vrsko9th7+iy1xYSgAdp0rXjDda7zHi5acJT9cFYxGUTlHAT5j4/T5lKdTW6nyd0lSMn5hCsWBrj3VJzOIBBtnwljkEs9S+KePc9OCC0/G4656SvRsEANdDD8hWmdbbaDV4CHYLCcIA4nM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=v8K+fLOC; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=cA0akQKf; arc=none smtp.client-ip=103.168.172.149 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="v8K+fLOC"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="cA0akQKf" Received: from phl-compute-11.internal (phl-compute-11.internal [10.202.2.51]) by mailfout.phl.internal (Postfix) with ESMTP id 030B8EC0222; Fri, 14 Aug 2026 21:59:17 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-11.internal (MEProxy); Fri, 14 Aug 2026 21:59:17 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786759156; x= 1786845556; bh=bljFEbJ6nOpcyGugcFKzpM9ttTNPFcr/ZQQGmFBGNGQ=; b=v 8K+fLOCozrKSLZYkHw5wvKkubbHtkGy4qBB01a/I2nfGUKJmpYOKaISYKiPm+AF5 GveDWj8JXLDbSfkHMYVogaB/ZIJQBHFexFkr1HikMZ+nvKFF7TrX8LLrTr+zkdEC HcpOE5paeofRUkNPUG5ZZqrRjQPZYJ17vhbq2ljFrWFxt/ZjL+oVB50OeErSUmGc PFC6vEjlcz5RUkAn/T2/AMkRF0ZqFbO9R4WGdQQUDy3eR2HdnIT2IoCzvuie+UDJ JV58kBmZCKDI0W1a4PSdZ6y4sRdkRiuZFnGX97+ERsRmT0Mx65XSz+qjJlkE0Vd1 8vdYphfoc1J1xvAIG41sA== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786759156; x=1786845556; bh=b ljFEbJ6nOpcyGugcFKzpM9ttTNPFcr/ZQQGmFBGNGQ=; b=cA0akQKfu38IBZl4w gV9Ov1ersw8ZlDu4zmZ4150TVSoLkM7I9Qss1DGin13bmPtPcgX0K4P3SkH32+iQ 67hhkc4pjiXqWzT3UU9OwyYUb5K2TlKc9k3bwuylTpd9HLgHGRWbFT9rW0Gv5+MO 0Ohk8UGeuOxvXFzQmBQpkTfv7X4grVIVkABWcDoV4mOJnL5r4iA+7HxD/sxT4pMD KZKJ2UfJDJkZa+bZMdOoW2xHX/GIAWkTVTywyanPRmZTHeyEBsA3efcM+noWn6ZL vZ6Df63xwrM9u+UW4XKOpnw55mvDTZH6QRolpfoaZAyks+ulRZPqQk+shhFElEMk n5LJQ== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEYpEQiBdLKlCaLp317TOYqYwO99jnuX4YrBujo3RuW+IIGXdYj4qE1ezcGRvbtOZ hc7RfnfkdlTehb5Dkvk4VZTcyRAqm2CcEjaldkIQshIttrB2HSeWyAKQfsR+ykmymFv8qO aC/+wh3joJiQccaWn5NjVrqLVVMYdoVO37A5ZL6SnR75seJlCquk1NTOiA0EvXHIVSJ25Q iFONSccd/wEzNfQZlOvvQQf5bCKiGIALXd0Rd37Vcnrh56q8DfWwmvEu2ZgBGSMa525iaP HkMF5MKO3LDiNiHXJVHjo5f/L4XJttiAOfdXhljkaHgZI649okv5M81/eeF8HfhCZr1USB aFDwzaDkfOm7MheaHSu0LnQkdQIwPXvaU0CQJ6NRlHYca+Szys49+C+KYF4TupoV0LnHl9 4oZHEL3wpxshS9kAwizxlp8iJnvvqoORWaf4t/+w2sr70KqtTeuQNBxMGtA6+n0hV1Bbs4 kUiKpye1DmC2Z3IY+5BCqeqlSxHA96oGURes8rVJevUEfVEdY2gZ84ack6+02xHVk3wkKg MCfp5WjQhdx+11zRGO7zRb+mvvFoOEu4JGm+agYo4SaH6Uz0HkHI9pRvk3F6fCs6OR2lTy TGwbWLIoANh0cEv1dt6xM/N3WfoyEI1ftOiK+Z4robHef9T3VSKUsgOcnjsw X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 14 Aug 2026 21:59:16 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v4 05/19] selftests/mm: make the swap cases' swapout reliable Date: Sat, 15 Aug 2026 02:58:47 +0100 Message-ID: <20260815015901.1236937-6-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260815015901.1236937-1-kirill@shutemov.name> References: <20260815015901.1236937-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" collapse_swapin_single_pte() and collapse_max_ptes_swap() swap a range out and then require smaps to report exactly the count they asked for. Two things keep that count from arriving. MADV_PAGEOUT is best effort, so the count often turns up a moment late. And wait_for_scan() leaves MADV_HUGEPAGE behind, so khugepaged is still working on the range. Collapsing a range with up to max_ptes_swap pages swapped out means reading them back in, so the daemon empties the swap as fast as the case fills it. On arm64 with 64K pages max_ptes_swap is 1024 pages, which is 64M a step, and the case loses: # Swapout 1024 of 8192 pages... Fail not ok 10 collapse_max_ptes_swap Ask again for up to two seconds, with the range held out of the daemon's reach while asking. The collapse each case runs next puts MADV_HUGEPAGE back, so only the setup is affected. If the pages still will not go, skip. A machine with no swap, or swap too small, full, capped by a memcg or busy with writeback, is not the kernel under test refusing. An error from madvise() itself still ends the run. Assisted-by: Claude-Code:claude-opus-5 Reviewed-by: Muhammad Usama Anjum Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/khugepaged.c | 53 +++++++++++++++++++------ 1 file changed, 41 insertions(+), 12 deletions(-) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index ec5c36a19d92..7eb9db0005a0 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -219,6 +219,41 @@ static bool check_swap(void *addr, unsigned long size) return swap; } =20 +/* + * Page the range out and wait for the swap count to say so. + * + * Two things get in the way. MADV_PAGEOUT is best effort: + * shrink_folio_list() leaves a folio alone when it cannot reclaim it right + * away, and one still under writeback from an earlier pageout is the comm= on + * case, so the count the caller asks for arrives a moment later. And a r= ange + * an earlier collapse left MADV_HUGEPAGE is one khugepaged is still worki= ng + * on: collapsing a range with up to max_ptes_swap pages swapped out means + * reading those pages back in, so the daemon undoes the pageout as fast a= s it + * is asked for. Keep the range out of its reach; the collapse the caller= runs + * next puts MADV_HUGEPAGE back. + * + * Failing to get the pages out is the machine's answer, not the kernel's = -- + * swap too small, swap full, a memcg cap, a folio still under writeback -= - so + * callers skip rather than fail. An error from madvise() is different, a= nd + * ends the run here. + */ +static bool swapout_range(void *p, unsigned long size) +{ + int i; + + if (madvise(p, size, MADV_NOHUGEPAGE)) + ksft_exit_fail_perror("madvise(MADV_NOHUGEPAGE)"); + + for (i =3D 0; i < 40; i++) { + if (madvise(p, size, MADV_PAGEOUT)) + ksft_exit_fail_perror("madvise(MADV_PAGEOUT)"); + if (check_swap(p, size)) + return true; + usleep(50 * 1000); + } + return false; +} + static void *alloc_mapping(int nr) { void *p; @@ -827,12 +862,10 @@ static void collapse_swapin_single_pte(struct collaps= e_context *c, struct mem_op ops->fault(p, 0, hpage_pmd_size); =20 ksft_print_msg("Swapout one page..."); - if (madvise(p, page_size, MADV_PAGEOUT)) - ksft_exit_fail_perror("madvise(MADV_PAGEOUT)"); - if (check_swap(p, page_size)) { + if (swapout_range(p, page_size)) { success("OK"); } else { - fail("Fail"); + skip("Could not swap out"); goto out; } =20 @@ -853,12 +886,10 @@ static void collapse_max_ptes_swap(struct collapse_co= ntext *c, struct mem_ops *o ops->fault(p, 0, hpage_pmd_size); =20 ksft_print_msg("Swapout %d of %d pages...", max_ptes_swap + 1, hpage_pmd_= nr); - if (madvise(p, (max_ptes_swap + 1) * page_size, MADV_PAGEOUT)) - ksft_exit_fail_perror("madvise(MADV_PAGEOUT)"); - if (check_swap(p, (max_ptes_swap + 1) * page_size)) { + if (swapout_range(p, (max_ptes_swap + 1) * page_size)) { success("OK"); } else { - fail("Fail"); + skip("Could not swap out"); goto out; } =20 @@ -870,12 +901,10 @@ static void collapse_max_ptes_swap(struct collapse_co= ntext *c, struct mem_ops *o ops->fault(p, 0, hpage_pmd_size); ksft_print_msg("Swapout %d of %d pages...", max_ptes_swap, hpage_pmd_nr); - if (madvise(p, max_ptes_swap * page_size, MADV_PAGEOUT)) - ksft_exit_fail_perror("madvise(MADV_PAGEOUT)"); - if (check_swap(p, max_ptes_swap * page_size)) { + if (swapout_range(p, max_ptes_swap * page_size)) { success("OK"); } else { - fail("Fail"); + skip("Could not swap out"); goto out; } =20 --=20 2.54.0 From nobody Mon Sep 28 23:56:19 2026 Received: from fout-a6-smtp.messagingengine.com (fout-a6-smtp.messagingengine.com [103.168.172.149]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 152F2349CD6; Sat, 15 Aug 2026 01:59:19 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.149 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759161; cv=none; b=OQfRw1n0kbt/uW6AyaVqokuIx5tqOfgvh+/77wX2JPqOrifeKi/zGlw+d8MPPDJFQ6fvr2RivTvLN6UeU/n4qxfs1htlNTLkOHXWRH7i9u7ZRCOp0BnzZsKFXESZ8LHZVvN0uBIL8cgSKCjFHGbLW0TPcJzQ91kqpNFzMFksEfc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759161; c=relaxed/simple; bh=q0bgg9uaN4yIMca/bqK4goaQL7KGJcwOOLAjn1Yf5AM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Gz6x3NdOz7q/0eaBeVoZBfNRW1IECdWzicskRuy/150PjqBkWu+4sM/W3ecWHGEdR2O9LU6x/yvNCcl3R4IhEbnweswExV5jCl0dCjJBkfy7S10gFy3CuDyrytFK1TUWheLMyjuC/9XFt5/h9oo1rcCiioj9IGv+RO/zrZ4hSuM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=CgdQW+Jc; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=Ku1e1ZdD; arc=none smtp.client-ip=103.168.172.149 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="CgdQW+Jc"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="Ku1e1ZdD" Received: from phl-compute-09.internal (phl-compute-09.internal [10.202.2.49]) by mailfout.phl.internal (Postfix) with ESMTP id 452B4EC0223; Fri, 14 Aug 2026 21:59:19 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-09.internal (MEProxy); Fri, 14 Aug 2026 21:59:19 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786759159; x= 1786845559; bh=LgG13P0WQQyCBdBYOJL0T47tEW0jH9jShvZjUp74c7I=; b=C gdQW+JcxJmpuInkt1ol9tNRCxUZmZA25NEmIzkRIS8KKN4Yz342KBwmMj2w7GVO8 jomEzYytrPXUhNhKQMVvuJ3NeKxvFK13xpt/kgYIE29Skm3NaxRzNwuO16CWPjPu S7UJ0E1joYJwL0Aj4AbekRoES+ZfNke4zfULB7Ghs3NcW4e6UJ2z3vcNKh6eZks6 08Qgt017uRZI+sRhiF8wSfLTrX6t2FZWJHvIyhhJcPhcslHKhoX9SE+a1nAWwd4w 8botglqSk/iDwjmIVRUWqxPJhFuLMQQGIIAy230mCI5wa8rWLA/+FIgsiJe1IfwG Gg3s1SlXbi8kUe+m7wxIg== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786759159; x=1786845559; bh=L gG13P0WQQyCBdBYOJL0T47tEW0jH9jShvZjUp74c7I=; b=Ku1e1ZdDPyHtg2rcI r2U/q7Hg+TiBVWuAW9EEfj9efA+oj1zwHGR+it846TTNXWG9mQE9PjHvQP0jC39r zSoq7GotYkA+FzqNmrQqcuDAqzFdDhQvSaZ8G9ObFX8h2sDoHDefVQIxKm5H8Y1d rtG0dGQrJUNygXM1X81xmtUsTsxM1h0iLqQApO/iZHvkBhtx1kmSy7hViFuwY5Ui lSGb9nOweCEyOb6/vCnzB2ov+96yTqORv7Vmmk7viIziuc5vc6zCUEB/A2shzmCL BwPFwbX5Udxs2ISIckSJMwLtbi3DRBPXrTBphyp3sfg6fuAQvZcEdpw0z95uMjBw 4r6vg== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFkzjFsAsDS7jWvo8tVKWpCqxoPadMTBhe5gqM9L8x58RJwkgUd6ef20LiiZWkVJ2 5XQFhBj/83jOdOO/Ql8fjehU8OJYV5xIu11NhtpuSlnR0oerfSUAVFPgI11Kw5ramX9VWq n08yQGhgo5srNyqhq1wQoX26nhsBi0/kG/kW3aJyoHPyWIE8LEdKOQSb0eQSt/HvnaOXE5 UcIyVyC42WtFkuUGaM3UI8X0We8hi8iW/Yy96suQXEqLNayQSzJTP+/kt+gGyLyNBuaXQP ClJKWCE1EuRQ5KlZV/VIT/qKrfIiWHiKuiXCdDyNa4K7qcq8xZMDqhI/wNpN/ytfSmuiC9 ASDF64uDYYbJcKpkthtCvX+lL6hbpkE8rcBGht8D6sa5z3Wt/7XEw1SzEiJdE326cyseW0 C4UxUzzt/UaE7luBGDD7sTDLHv5DOkHTzgs1a1G23kMi0QipbeAo+uojtkKGEyV1ttU1E9 2GwrULNZCWkoidglMm4joCnVsrGWrN5jEtdvnnC7T+RL8Kbk0g5h1XcbA/vJJ6kuYy8+Z5 xtj/YfjsvJRrWkzRquVToRNe7Mj19b1w1lZ+y9a7DsIoR7l6OOPPWlQJWwhPyg2WbL+EyR U4MTWsj3igdF1mDl5oVUQW12b2JHGbnAWyVl/npAQUdDcFRMPj/i1qntecmQ X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 14 Aug 2026 21:59:18 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v4 06/19] selftests/mm: stop khugepaged during the MADV_COLLAPSE cases Date: Sat, 15 Aug 2026 02:58:48 +0100 Message-ID: <20260815015901.1236937-7-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260815015901.1236937-1-kirill@shutemov.name> References: <20260815015901.1236937-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" __madvise_collapse() turns THP off before each MADV_COLLAPSE, both to keep khugepaged out of the range and to prove MADV_COLLAPSE ignores the setting. It clears the global controls only, which is no longer enough. A per-order control overrides them, and -s leaves the source order at "always", so khugepaged collapses the very range the case is working on. The case then fails on a collapse that was interfered with rather than refused. Clear the per-order controls too. Set them to "inherit", not "never". khugepaged honours the global never and stays out. A forced shmem collapse takes the order it builds from these very controls, and still finds one. Fixes: 9f0704eae8a4 ("selftests/mm/khugepaged: enlighten for multi-size THP= ") Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/khugepaged.c | 13 ++++++++++++- 1 file changed, 12 insertions(+), 1 deletion(-) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index 7eb9db0005a0..0008862e7cbc 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -550,8 +550,8 @@ static bool is_anon(struct mem_ops *ops) static void __madvise_collapse(const char *msg, char *p, int nr_hpages, struct mem_ops *ops, bool expect) { - int ret; struct thp_settings settings =3D *thp_current_settings(); + int ret, i; =20 ksft_print_msg("%s...", msg); =20 @@ -564,9 +564,20 @@ static void __madvise_collapse(const char *msg, char *= p, int nr_hpages, /* * Prevent khugepaged interference and tests that MADV_COLLAPSE * ignores /sys/kernel/mm/transparent_hugepage/enabled + * + * The per-order controls have to go too, not just the global one: a + * source order left at "always" -- which -s does -- lets khugepaged + * collapse the very range the case is working on. Set them to + * "inherit", not "never". khugepaged honours the global never and + * stays out. A forced shmem collapse takes the order it builds from + * these very controls, and still finds one. */ settings.thp_enabled =3D THP_NEVER; settings.shmem_enabled =3D SHMEM_NEVER; + for (i =3D 0; i < NR_ORDERS; i++) { + settings.hugepages[i].enabled =3D THP_INHERIT; + settings.shmem_hugepages[i].enabled =3D SHMEM_INHERIT; + } thp_push_settings(&settings); =20 /* Clear VM_NOHUGEPAGE */ --=20 2.54.0 From nobody Mon Sep 28 23:56:19 2026 Received: from fhigh-a7-smtp.messagingengine.com (fhigh-a7-smtp.messagingengine.com [103.168.172.158]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 10264287246; Sat, 15 Aug 2026 01:59:21 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.158 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759165; cv=none; b=Q6kETuC/NcKmGw7dvGG97W7FxUx6luRB7hrYdxh9W830UIQlsLJguGyKLceD+qa73m/uZcilDxRPqNb5I0cLEYB+/Sg9QNZmOReFsWUCf+E9oSuhgM/0eyZarixSTgQdRgCJbjeQLKXbB9tBGcq7KBBf92RoPpJADkOVReUpdpw= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759165; c=relaxed/simple; bh=trbNubyEscebeI0IJT2waWmyfGYJrlanwkyVzEhahBU=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=XNClGAWIAjaLWb2LhDKEG0rD28EtVwEpTUShJoR9HEY65DKXHPAusoEUQCerdP833WB025D5Gy145XkUBG+zaks1elLFAaZE3R3x6bh0v0+pQGOHld7oU7s/Cvnp4scguA20S4Qkxl1X4zGk7CGiwdVN4JJGUI4oYLjGycnGtAQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=T7xz6jDy; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=PtseMFA1; arc=none smtp.client-ip=103.168.172.158 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="T7xz6jDy"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="PtseMFA1" Received: from phl-compute-02.internal (phl-compute-02.internal [10.202.2.42]) by mailfhigh.phl.internal (Postfix) with ESMTP id 2E01B140011E; Fri, 14 Aug 2026 21:59:21 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-02.internal (MEProxy); Fri, 14 Aug 2026 21:59:21 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786759161; x= 1786845561; bh=aGZmf/6wBq65F2JoipHZ/CYoHcQpvbh6TEl9rP7Pel0=; b=T 7xz6jDyNnlB0Ez2Ct771Gn2xfcopt+eC2/a8yF8l9BLM8jbCSo56u++JpCCcuuUR AMVPYE8m3uAS3KmUA48P0RyoNRcPHU2nSm/TuwRwKHoBCKvNDXU13OP/v1BTbZK+ 7bXh6+7gqXxJy1MTHuTuRf3MZ/ZkkrapVAdagEd/roTp5F+FDMk0zi9xh0P6DBhF rqsryH4uT3V+UGv5g1PtEBrYTyv++Cuf2WKYFH5FED6MrRtYdJKndkeFcmZ0ux5U EayLhRAKws6A7wdQ6i4eMXh9w56rMhm90mBB25M1tv6eWSm06iWfalRnqVu0YUFn E8TibUC4l1xJeta2irQcg== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786759161; x=1786845561; bh=a GZmf/6wBq65F2JoipHZ/CYoHcQpvbh6TEl9rP7Pel0=; b=PtseMFA1D5viN17wS Nev87fgjogNlQEXUuCElSDKtZmfmV9cdgpc0pPbfJPhx5Paz4vTB6rJiSDrcj4nd MpNLbTN1v5yxSof5BlrtlIhhiHrTjytHFOTTvos3iWt13p0NJFgGrVyig8FnEjBd QupT7J4gax0ZQfoDEld8+/RqLiZUObcsaLwH9043Ge+x1SRSaIqoWfzvlddyzEyJ uA/WBqw0srslWUNzhXFAWvBHjv/Yqlw+kcE1+EiZO2q5IUadXDvxotQovGYKm8yn E3XFn3jV0xavgrmyiXR3feQ9KxUSJvjMe8QMyTjceAoPYFBuSiMS3/jy+GlvrNnv MFgTA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEYpEQiBdLKlCaLp317TOYqYwO99jnuX4YrBujo3RuW+IIGXdYj4qE1ezcGRvbtOZ hc7RfnfkdlTehb5Dkvk4VZTcyRAqm2CcEjaldkIQshIttrB2HSeWyAKQfsR+ykmymFv8qO aC/+wh3joJiQccaWn5NjVrqLVVMYdoVO37A5ZL6SnR75seJlCquk1NTOiA0EvXHIVSJ25Q iFONSccd/wEzNfQZlOvvQQf5bCKiGIALXd0Rd37Vcnrh56q8DfWwmvEu2ZgBGSMa525iaP HkMF5MKO3LDiNiHXJVHjo5f/L4XJttiAOfdXhljkaHgZI649okv5M81/eeF8HfhCZr1Ug5 fH812civBTQ4/FE9yW9xtOi8Y0F9yKDrp1R20YTK463G/7P8yKRYnTTXddn5qzFWbXndys aiCpPcx5TNMkbM0QbjoCI2x1ABR670ocmpPmFHiQMaquLHhMxNhgE5Llu5/gibaYC1zkm9 Ar2yt7xVrGUFjmYY0GTfXvXuILsLD2Q4rlHwJiD5kf1kO9y44H7QZ3a6kgmb9kzlEMLv8E 8gx72OWly2IFG0Ahpn8cYPbn31wFFwmhj+8iYcXwP+ksUUlj7vUgpT7c1PW32uSnUDuyyd JLi9C9d343aoTgtRo0HxjLFqQmY6rUFKunIWpzNVxqmpjuM5j+WPCS7xbVxA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 14 Aug 2026 21:59:20 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v4 07/19] selftests/mm: move is_backed_by_folio() into vm_util Date: Sat, 15 Aug 2026 02:58:49 +0100 Message-ID: <20260815015901.1236937-8-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260815015901.1236937-1-kirill@shutemov.name> References: <20260815015901.1236937-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Checking that an address range is backed by a folio of a given order is useful to any test that builds or collapses large folios. mTHP collapse coverage in the khugepaged selftest needs exactly that. split_huge_page_test.c already has the building block: is_backed_by_folio() reads the compound head and tail flags from /proc/kpageflags to classify the folio behind a page. Move it into vm_util so other tests can use it. No functional change. Assisted-by: Claude-Code:claude-opus-5 Acked-by: Mike Rapoport (Microsoft) Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) Acked-by: Lorenzo Stoakes (ARM) --- .../selftests/mm/split_huge_page_test.c | 62 ------------------- tools/testing/selftests/mm/vm_util.c | 62 +++++++++++++++++++ tools/testing/selftests/mm/vm_util.h | 2 + 3 files changed, 64 insertions(+), 62 deletions(-) diff --git a/tools/testing/selftests/mm/split_huge_page_test.c b/tools/test= ing/selftests/mm/split_huge_page_test.c index 86a603692826..0adfe7dde7e5 100644 --- a/tools/testing/selftests/mm/split_huge_page_test.c +++ b/tools/testing/selftests/mm/split_huge_page_test.c @@ -42,68 +42,6 @@ const char *kpageflags_proc =3D "/proc/kpageflags"; int pagemap_fd; int kpageflags_fd; =20 -static bool is_backed_by_folio(char *vaddr, int order, int pagemap_fd, - int kpageflags_fd) -{ - const uint64_t folio_head_flags =3D KPF_THP | KPF_COMPOUND_HEAD; - const uint64_t folio_tail_flags =3D KPF_THP | KPF_COMPOUND_TAIL; - const unsigned long nr_pages =3D 1UL << order; - unsigned long pfn_head; - uint64_t pfn_flags; - unsigned long pfn; - unsigned long i; - - pfn =3D pagemap_get_pfn(pagemap_fd, vaddr); - - /* non present page */ - if (pfn =3D=3D -1UL) - return false; - - if (pageflags_get(pfn, kpageflags_fd, &pfn_flags)) - goto fail; - - /* check for order-0 pages */ - if (!order) { - if (pfn_flags & (folio_head_flags | folio_tail_flags)) - return false; - return true; - } - - /* non THP folio */ - if (!(pfn_flags & KPF_THP)) - return false; - - pfn_head =3D pfn & ~(nr_pages - 1); - - if (pageflags_get(pfn_head, kpageflags_fd, &pfn_flags)) - goto fail; - - /* head PFN has no compound_head flag set */ - if ((pfn_flags & folio_head_flags) !=3D folio_head_flags) - return false; - - /* check all tail PFN flags */ - for (i =3D 1; i < nr_pages; i++) { - if (pageflags_get(pfn_head + i, kpageflags_fd, &pfn_flags)) - goto fail; - if ((pfn_flags & folio_tail_flags) !=3D folio_tail_flags) - return false; - } - - /* - * check the PFN after this folio, but if its flags cannot be obtained, - * assume this folio has the expected order - */ - if (pageflags_get(pfn_head + nr_pages, kpageflags_fd, &pfn_flags)) - return true; - - /* If we find another tail page, then the folio is larger. */ - return (pfn_flags & folio_tail_flags) !=3D folio_tail_flags; -fail: - ksft_exit_fail_msg("Failed to get folio info\n"); - return false; -} - static int check_after_split_folio_orders(char *vaddr_start, size_t len, int pagemap_fd, int kpageflags_fd, int orders[], int nr_orders) { diff --git a/tools/testing/selftests/mm/vm_util.c b/tools/testing/selftests= /mm/vm_util.c index 80bc9f597b52..5db1a7774f49 100644 --- a/tools/testing/selftests/mm/vm_util.c +++ b/tools/testing/selftests/mm/vm_util.c @@ -494,6 +494,68 @@ int pageflags_get(unsigned long pfn, int kpageflags_fd= , uint64_t *flags) return 0; } =20 +bool is_backed_by_folio(char *vaddr, int order, int pagemap_fd, + int kpageflags_fd) +{ + const uint64_t folio_head_flags =3D KPF_THP | KPF_COMPOUND_HEAD; + const uint64_t folio_tail_flags =3D KPF_THP | KPF_COMPOUND_TAIL; + const unsigned long nr_pages =3D 1UL << order; + unsigned long pfn_head; + uint64_t pfn_flags; + unsigned long pfn; + unsigned long i; + + pfn =3D pagemap_get_pfn(pagemap_fd, vaddr); + + /* non present page */ + if (pfn =3D=3D -1UL) + return false; + + if (pageflags_get(pfn, kpageflags_fd, &pfn_flags)) + goto fail; + + /* check for order-0 pages */ + if (!order) { + if (pfn_flags & (folio_head_flags | folio_tail_flags)) + return false; + return true; + } + + /* non THP folio */ + if (!(pfn_flags & KPF_THP)) + return false; + + pfn_head =3D pfn & ~(nr_pages - 1); + + if (pageflags_get(pfn_head, kpageflags_fd, &pfn_flags)) + goto fail; + + /* head PFN has no compound_head flag set */ + if ((pfn_flags & folio_head_flags) !=3D folio_head_flags) + return false; + + /* check all tail PFN flags */ + for (i =3D 1; i < nr_pages; i++) { + if (pageflags_get(pfn_head + i, kpageflags_fd, &pfn_flags)) + goto fail; + if ((pfn_flags & folio_tail_flags) !=3D folio_tail_flags) + return false; + } + + /* + * check the PFN after this folio, but if its flags cannot be obtained, + * assume this folio has the expected order + */ + if (pageflags_get(pfn_head + nr_pages, kpageflags_fd, &pfn_flags)) + return true; + + /* If we find another tail page, then the folio is larger. */ + return (pfn_flags & folio_tail_flags) !=3D folio_tail_flags; +fail: + ksft_exit_fail_msg("Failed to get folio info\n"); + return false; +} + /* If `ioctls' non-NULL, the allowed ioctls will be returned into the var = */ int uffd_register_with_ioctls(int uffd, void *addr, uint64_t len, bool miss, bool wp, bool minor, uint64_t *ioctls) diff --git a/tools/testing/selftests/mm/vm_util.h b/tools/testing/selftests= /mm/vm_util.h index 9a49af88702e..56a28ce7d029 100644 --- a/tools/testing/selftests/mm/vm_util.h +++ b/tools/testing/selftests/mm/vm_util.h @@ -97,6 +97,8 @@ int64_t allocate_transhuge(void *ptr, int pagemap_fd); int pageflags_get(unsigned long pfn, int kpageflags_fd, uint64_t *flags); int gather_folio_orders(char *vaddr_start, size_t len, int pagemap_fd, int kpageflags_fd, int orders[], int nr_orders); +bool is_backed_by_folio(char *vaddr, int order, int pagemap_fd, + int kpageflags_fd); =20 int uffd_register(int uffd, void *addr, uint64_t len, bool miss, bool wp, bool minor); --=20 2.54.0 From nobody Mon Sep 28 23:56:19 2026 Received: from fout-a6-smtp.messagingengine.com (fout-a6-smtp.messagingengine.com [103.168.172.149]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7126D34C815; Sat, 15 Aug 2026 01:59:25 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.149 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759167; cv=none; b=TZHTvHg2pdiop6RLzKh9tTrJFSdxh5nBr39lBLKfN867vZA4RiI+Yqnx3cOn7w5824VHF1/nCNue/k7RStHshvIAHSJX5trcXD7SpTjt9K2GkqajvWV+kjWWdzdmAsv8Ou2+fMU+KNyg5gyJOAgeTIE9ZFZJWWtuYAozUDIJO8w= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759167; c=relaxed/simple; bh=m5zGJkZM/CUqaRGMqJcW7Lths1PS2gsh5NzoaGJme4U=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=KuUtDcWsIxtRU9BmIIR0XXUbhLF3dEJV8HqSyXoXHZXl1/gealaCoNHu4Ne/pEqhTqgtpivczBVe10M6TD8WLgmsohd6O6GVXzb3N9NmNhPhjNzlt+BfYrVq7tJN3zlwHOwzvGyCDmAN+Bw2W7VhXttj6gczhttXIwYP0E0dc94= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=s5oC2QIz; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=a2z10bm5; arc=none smtp.client-ip=103.168.172.149 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="s5oC2QIz"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="a2z10bm5" Received: from phl-compute-05.internal (phl-compute-05.internal [10.202.2.45]) by mailfout.phl.internal (Postfix) with ESMTP id 01CF1EC022E; Fri, 14 Aug 2026 21:59:23 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-05.internal (MEProxy); Fri, 14 Aug 2026 21:59:23 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786759162; x= 1786845562; bh=3BluNDLCZhWfJh2T22U9KUR4xuqo32ndiTxV1AaNaIU=; b=s 5oC2QIzgW1HHEplcurKvHDjlWHAlUd6YZD2TjDoHobEFka0+vVyYjFAf2VBePPcE LXsOxgFGJ421hTgcFCISjyzgJYU4VMKG6C0PGWeizZj3bpJ+nAcTVXqQljRZbr7i aJDsgLUMkIFVnnjgneyfUbGAbDBBpGwjmU1xiwLLbf1JOQUMIklLQT5vbHwXOyzn IK3cWHI0tJC9EYagt0INL1epWpclZCY2z6LccMWY7OPxicvB1DyurqqphVGxHRex wUKaFAghicwjbymzSL1TFUylF4bvLQKgJ8d6tf8O8cE2sSwIJho8rD/B/nO3/Ohi BcpArZpYAuzpcZz9l7K1w== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786759162; x=1786845562; bh=3 BluNDLCZhWfJh2T22U9KUR4xuqo32ndiTxV1AaNaIU=; b=a2z10bm5YOWcOXb4s dd61rq5lq1mP35OFj/TUe9CmxkmYsLqbcisjbnmVImeks10TRr1zsnsZTBWZW38b 5TjSOKsczJCkvFp2aCzl3kUoSe0acFihBbTYXDT1GM1yTT4gDjX5KzvqnoaUX+yN xOI7ZGYEz37iwaUvGWWTVcgUDsK+zbAOXijNsdxV/Q7WJduUexodyyHSnaYVof02 cLPdEFoX4MRw3ZUVDWUd5IJawdQu4yQ02CI7dyrBorbXqDo/B7mDoeP3KPrlnUyO SZi13+PWti/tNL05Jfncqth0pJ8ikpXUZc/dsrwOL8MYmvrt6lt1gvzz7VrWXDL+ XY1PA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEYpEQiBdLKlCaLp317TOYqYwO99jnuX4YrBujo3RuW+IIGXdYj4qE1ezcGRvbtOZ hc7RfnfkdlTehb5Dkvk4VZTcyRAqm2CcEjaldkIQshIttrB2HSeWyAKQfsR+ykmymFv8qO aC/+wh3joJiQccaWn5NjVrqLVVMYdoVO37A5ZL6SnR75seJlCquk1NTOiA0EvXHIVSJ25Q iFONSccd/wEzNfQZlOvvQQf5bCKiGIALXd0Rd37Vcnrh56q8DfWwmvEu2ZgBGSMa525iaP HkMF5MKO3LDiNiHXJVHjo5f/L4XJttiAOfdXhljkaHgZI649okv5M81/eeF8HfhCZr1UAm rgJ76CjNAg/U/EKabJ+37G6BvMM6OBu2S56U3YWsJffoiZIE7o7yRmzI1OaELgILaiNqOo 1Uq53GRRyG0wmSwHCaXQhUmxPHpjgx3LBvDZvJI2g8hLT3dlzx/tzjmukS8sdOlYlQxFrf pF33J8KwidLs+W9kyG792FlRJ9R+6AJij5y8/PXNhargdOkNqPDsBM/c5bDG3fYqAw/ov9 v6QBN4oqUXjbKiyjShwtcAMIWpXKpukDz46zSH5kDTrXYEI9fAj5YLBZW2zQ19ZOhLXATj WaTc5JbID9LNOP/VpoCk3VliNaVr5JQe7/9AwoupdqfOs0kCJnidsvyZ21hA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 14 Aug 2026 21:59:22 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v4 08/19] selftests/mm: add folio-order check for address ranges Date: Sat, 15 Aug 2026 02:58:50 +0100 Message-ID: <20260815015901.1236937-9-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260815015901.1236937-1-kirill@shutemov.name> References: <20260815015901.1236937-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" An mTHP collapse test needs to know that a range is backed by folios of the target order, and that they sit where a collapse would put them. Nothing answers that today: is_backed_by_folio() classifies the folio behind a single page, and check_huge_anon() reads smaps AnonHugePages, which only accounts PMD mappings. Add is_range_backed_by_folio_orders(). For every order-aligned window of the range it requires a present head PFN at its natural alignment and a contiguous PFN run across the window. A window backed by two smaller folios fails the contiguity check, and a folio mapped off the window's alignment fails the head check. The mTHP cases need both to tell a collapsed window from the one beside it. Assisted-by: Claude-Code:claude-opus-5 Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) Reviewed-by: Mike Rapoport (Microsoft) --- tools/testing/selftests/mm/vm_util.c | 42 ++++++++++++++++++++++++++++ tools/testing/selftests/mm/vm_util.h | 2 ++ 2 files changed, 44 insertions(+) diff --git a/tools/testing/selftests/mm/vm_util.c b/tools/testing/selftests= /mm/vm_util.c index 5db1a7774f49..c9bd6c92fa41 100644 --- a/tools/testing/selftests/mm/vm_util.c +++ b/tools/testing/selftests/mm/vm_util.c @@ -556,6 +556,48 @@ bool is_backed_by_folio(char *vaddr, int order, int pa= gemap_fd, return false; } =20 +/* + * Check whether every order-@order window of [start, len) maps exactly one + * folio of that order, head to tail. The address range must be naturally + * aligned, each window's PFN run must be contiguous, and a window's first + * PFN must be the folio head. + * + * This is the check "did this range collapse into order-@order folios": a + * window assembled from parts of several folios, or mapping a folio shift= ed + * from its natural position, fails. + */ +bool is_range_backed_by_folio_orders(char *start, size_t len, int order, + int pagemap_fd, int kpageflags_fd) +{ + const unsigned long nr_pages =3D 1UL << order; + const size_t window =3D nr_pages * psize(); + char *vaddr; + + if ((uintptr_t)start % window || len % window) + return false; + + for (vaddr =3D start; vaddr < start + len; vaddr +=3D window) { + unsigned long pfn =3D pagemap_get_pfn(pagemap_fd, vaddr); + unsigned long i; + + /* Not present, or not mapping the folio head. */ + if (pfn =3D=3D -1UL || pfn % nr_pages) + return false; + + for (i =3D 1; i < nr_pages; i++) { + if (pagemap_get_pfn(pagemap_fd, vaddr + i * psize()) !=3D + pfn + i) + return false; + } + + if (!is_backed_by_folio(vaddr, order, pagemap_fd, + kpageflags_fd)) + return false; + } + + return true; +} + /* If `ioctls' non-NULL, the allowed ioctls will be returned into the var = */ int uffd_register_with_ioctls(int uffd, void *addr, uint64_t len, bool miss, bool wp, bool minor, uint64_t *ioctls) diff --git a/tools/testing/selftests/mm/vm_util.h b/tools/testing/selftests= /mm/vm_util.h index 56a28ce7d029..39dfb18dc10c 100644 --- a/tools/testing/selftests/mm/vm_util.h +++ b/tools/testing/selftests/mm/vm_util.h @@ -99,6 +99,8 @@ int gather_folio_orders(char *vaddr_start, size_t len, int pagemap_fd, int kpageflags_fd, int orders[], int nr_orders); bool is_backed_by_folio(char *vaddr, int order, int pagemap_fd, int kpageflags_fd); +bool is_range_backed_by_folio_orders(char *start, size_t len, int order, + int pagemap_fd, int kpageflags_fd); =20 int uffd_register(int uffd, void *addr, uint64_t len, bool miss, bool wp, bool minor); --=20 2.54.0 From nobody Mon Sep 28 23:56:19 2026 Received: from fout-a6-smtp.messagingengine.com (fout-a6-smtp.messagingengine.com [103.168.172.149]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7114D3446CE; Sat, 15 Aug 2026 01:59:25 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.149 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759167; cv=none; b=O61yFlXfCpCxFRpyFeYvlwvVSJZ9amJD4Kl81Mk4NRm0tzpaRJ/O1sZ8r+z0JBjl+cYep4NUaw4/qfBt/1F2CoccE4An4Le5SsW1dWD2rOTWz2ZItbAISxR8guzK8Sh7ZRX8x2zZHImRE0gMfy96aCedQx1yBT5+lRmfpv2KNSc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759167; c=relaxed/simple; bh=ZNSZWh8G3mSyfxg12hj716BokikUEVVgwZWBnG7u7n0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=YSbtK5mbRESNXTTseifCl9wvtaLaYXvZqpMDzb86sWNq7DxhUjVCye+mDcK9z6tGf1uDqLsEaAPM2xAaQmVqRe8GuJHBXtnRT1Y/qrbn1AUwmsCYWxujpEQFyX9rUwK2FkqYpSsL2MZqJVYoAUHp6vt1C3YVjE2pLYohP6MyFWk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=XGnhpE2L; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=Stk1zm+1; arc=none smtp.client-ip=103.168.172.149 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="XGnhpE2L"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="Stk1zm+1" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfout.phl.internal (Postfix) with ESMTP id 9BA54EC0230; Fri, 14 Aug 2026 21:59:24 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-06.internal (MEProxy); Fri, 14 Aug 2026 21:59:24 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786759164; x= 1786845564; bh=Qk5GeQLVap1nnpt4hhwsoYmaxkuYR/evO/AVLGk8sYw=; b=X GnhpE2LagaVsynrfBAbtHRm0TW1fCoTOVNgE9FSTahYktmCHyH+9OL9j+Tx9ucMy mTxiB4tvfdPRiM+Gc5my6lrikPRjsoASesBLAi2OqjblNsPBeAH9OkpdQ+J6W3aa I9dNCDl/SP1exx1JkOuNd5u0efx1YNrbecbQl+FecoxJFfqy/X3rj/A1NRFALB49 KUegj4s8CiS7yrc3CfVtt3ypIVHIDLMAiUPaxnBDiGX7sDigRfN3Uo2QtN2ZeRRy SO1ZPEsYdjMXEkvjFYnk0go5n/i2DigJ5paOqw3qPFEJBbxPvUcx1e/O1pqgwyOL MVbOzm65jnFzTDC3YBAEQ== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786759164; x=1786845564; bh=Q k5GeQLVap1nnpt4hhwsoYmaxkuYR/evO/AVLGk8sYw=; b=Stk1zm+1mYZX4Q5qu Wzxe5WnhuNvccHbn5+Gbutl2SnS2kLYvbbL3MQ9Ah8rH+bIi3wqnak66GIwmUEWx r53APhcc5t/mpF/U1ngrEdJEv3KBWdwuhpIHstaUt1+qJ8tg78PNnYk32r/SJS73 alt0gdpOr/mE66QDZ4yrSeVIxoGFs+eLKT+miFh8BvCvkcKK+AfisL/nkg+25joe oAsNdwP/t3v4zNLjJ+g3Qy+hY3+TYw1tq13N4ajCg2ZykD8IIhBzxnlQMLyn2C/d LDtgOMQAXjaAfBw1MC4393Ky/SoMkxyJuvbMvEAUND8fLdmYf9RWMG4lqcDBzERj e2gJg== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGF7teVIHNqoa5p2A8iL+trOmU4pExV0BsGEGBxavUqSy3vvyW/t2uVaC82ybAhbC FZe2jS418ApiE9B8lCjDVAvs2O/1hVxxIyd2QmVeomYKnxIQlXrdOkpVI2/A6aay1R5wM6 752pE1oD84VvsYvCM8wHufwfHzQzVg5WF64m0thPV1o6CCL0N8QD0816LOjK9LaF9KdRdp ShDgDyNTb/bU59JqJj/g3oSWXYBnQXQHEYh7kafl/PHtgQjoSit3F73LPXI/q1yj04b//y CUybFuAxuhlNqjceRELXNaCNxBn+qS4c4JK+cmo8T/HSoyzm/Kml9mcKob58Fnvz9v9LXy eLiUW+PjLxDuAmmBOpd0D1WPVsF0nMCdXCvLj/x/oQhXTYJ09daqpkuW9FM7BIIiu0ytWT DziKyW/IY4QISCcstmaRJfwMSDthdrHiyvBbFqmZ06pJcEtaWZVbAcHlqcm777xHFQ8YXE 4fujkbtuim/zsXlMEYlLDsn3+bznqzHFoK3j+7P4SgpfDdmkn3HAXpOSAXkv/QTSEvCZYM zkqJIQjHY0q119hNY27PpCplxnE01dE3k8L+VKWNJ8RwQjwBtHelaQTyO+lRtWbxngkARy abOG8+C8q57JzMojbeEUCducL49/SSE1TO2N0kHX78KfCiQpFrwJKTFh8Bug X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 14 Aug 2026 21:59:24 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v4 09/19] selftests/mm: add folio-order detection self-check Date: Sat, 15 Aug 2026 02:58:51 +0100 Message-ID: <20260815015901.1236937-10-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260815015901.1236937-1-kirill@shutemov.name> References: <20260815015901.1236937-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The khugepaged mTHP tests detect collapse results with the vm_util folio-order helpers rather than smaps AnonHugePages, which only sees PMD mappings. If those helpers are wrong, every case built on them is wrong the same way, and nothing says so. Check them directly. For every anon THP order the kernel supports, fault memory in with only that order enabled. Require the helpers to classify the backing as exactly that order: not a neighbouring order, and 4K-backed memory as order 0. Run it in the thp category, ahead of ./khugepaged, so a broken helper is reported as itself rather than as a collapse failure. Verified on x86-64 4K (orders 0, 2-9) and arm64 64K (orders 0, 2-13). The test needs ALIGN(), which hmm-tests.c and migration.c each defined privately. Move it to vm_util.h and drop both copies. Assisted-by: Claude-Code:claude-opus-5 Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/Makefile | 1 + .../testing/selftests/mm/folio_order_check.c | 137 ++++++++++++++++++ tools/testing/selftests/mm/hmm-tests.c | 1 - tools/testing/selftests/mm/migration.c | 1 - tools/testing/selftests/mm/run_vmtests.sh | 2 + tools/testing/selftests/mm/vm_util.h | 2 + 6 files changed, 142 insertions(+), 2 deletions(-) create mode 100644 tools/testing/selftests/mm/folio_order_check.c diff --git a/tools/testing/selftests/mm/Makefile b/tools/testing/selftests/= mm/Makefile index 2d5366196e30..2093fcf6e915 100644 --- a/tools/testing/selftests/mm/Makefile +++ b/tools/testing/selftests/mm/Makefile @@ -104,6 +104,7 @@ TEST_GEN_FILES +=3D guard-regions TEST_GEN_FILES +=3D merge TEST_GEN_FILES +=3D rmap TEST_GEN_FILES +=3D folio_split_race_test +TEST_GEN_FILES +=3D folio_order_check =20 ifneq ($(ARCH),arm64) TEST_GEN_FILES +=3D soft-dirty diff --git a/tools/testing/selftests/mm/folio_order_check.c b/tools/testing= /selftests/mm/folio_order_check.c new file mode 100644 index 000000000000..93030a42c3cc --- /dev/null +++ b/tools/testing/selftests/mm/folio_order_check.c @@ -0,0 +1,137 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Self-check for the vm_util folio-order detection helpers, + * is_backed_by_folio() and is_range_backed_by_folio_orders(). + * + * For every anon THP order the kernel supports, fault memory in with only + * that order enabled and verify the helpers report exactly that order: + * not a neighbouring order, and plain 4K memory as order 0. The helpers + * are what the khugepaged mTHP tests use to detect collapse results, so + * they must agree with the kernel's own idea of the backing before any + * collapse test relies on them. + */ +#define _GNU_SOURCE +#include +#include +#include +#include +#include + +#include "kselftest.h" +#include "vm_util.h" +#include "hugepage_settings.h" + +static int pagemap_fd; +static int kpageflags_fd; + +/* mmap an anon VMA of exactly @size bytes at a @size-aligned address. */ +static char *alloc_aligned(size_t size) +{ + size_t len =3D size * 2; + uintptr_t aligned; + char *p; + + p =3D mmap(NULL, len, PROT_READ | PROT_WRITE, + MAP_ANONYMOUS | MAP_PRIVATE, -1, 0); + if (p =3D=3D MAP_FAILED) + ksft_exit_fail_perror("mmap()"); + + aligned =3D ALIGN((uintptr_t)p, size); + if (aligned !=3D (uintptr_t)p) + munmap(p, aligned - (uintptr_t)p); + if (aligned + size !=3D (uintptr_t)p + len) + munmap((char *)aligned + size, + (uintptr_t)p + len - aligned - size); + + return (char *)aligned; +} + +/* + * Enable only @order (order 0: nothing), fault one aligned window in and + * check the helpers see exactly @order. + */ +static void check_order(int order) +{ + struct thp_settings settings =3D *thp_current_settings(); + size_t size =3D psize() << order; + bool ok =3D true; + char *p; + int i; + + for (i =3D 0; i < NR_ORDERS; i++) + settings.hugepages[i].enabled =3D THP_NEVER; + if (order) + settings.hugepages[order].enabled =3D THP_ALWAYS; + thp_push_settings(&settings); + + p =3D alloc_aligned(size); + *p =3D 1; + + if (!is_range_backed_by_folio_orders(p, size, order, + pagemap_fd, kpageflags_fd)) { + ksft_print_msg("order %d not detected after fault\n", order); + ok =3D false; + } + + /* A lower order must be rejected: the folio is larger. */ + if (order && is_range_backed_by_folio_orders(p, size, order - 1, + pagemap_fd, + kpageflags_fd)) { + ksft_print_msg("order %d also reported as order %d\n", + order, order - 1); + ok =3D false; + } + + /* Order 0 pages must not look like any large folio, and vice versa. */ + if (order && is_range_backed_by_folio_orders(p, size, 0, + pagemap_fd, + kpageflags_fd)) { + ksft_print_msg("order %d also reported as order 0\n", order); + ok =3D false; + } + + munmap(p, size); + thp_pop_settings(); + + ksft_test_result(ok, "order %d classified\n", order); +} + +int main(void) +{ + struct thp_settings settings; + unsigned long orders; + int order; + + ksft_print_header(); + + if (!thp_available()) + ksft_exit_skip("Transparent Hugepages not available\n"); + + pagemap_fd =3D open("/proc/self/pagemap", O_RDONLY); + if (pagemap_fd < 0) + ksft_exit_fail_perror("open(/proc/self/pagemap)"); + kpageflags_fd =3D open("/proc/kpageflags", O_RDONLY); + if (kpageflags_fd < 0) + ksft_exit_skip("open(\"/proc/kpageflags\") requires root\n"); + + orders =3D thp_supported_orders(); + if (!orders) + ksft_exit_skip("No supported THP orders\n"); + + ksft_set_plan(__builtin_popcountl(orders) + 1); + + thp_save_settings(); + thp_read_settings(&settings); + /* Base of the settings stack; the bottom entry is never popped. */ + thp_push_settings(&settings); + + check_order(0); + for (order =3D 1; order < NR_ORDERS; order++) { + if (!(orders & (1UL << order))) + continue; + check_order(order); + } + + + ksft_finished(); +} diff --git a/tools/testing/selftests/mm/hmm-tests.c b/tools/testing/selftes= ts/mm/hmm-tests.c index e2642eca0d02..df426f9218e7 100644 --- a/tools/testing/selftests/mm/hmm-tests.c +++ b/tools/testing/selftests/mm/hmm-tests.c @@ -65,7 +65,6 @@ enum { #define HMM_PATH_MAX 64 #define NTIMES 10 =20 -#define ALIGN(x, a) (((x) + (a - 1)) & (~((a) - 1))) /* Just the flags we need, copied from mm.h: */ =20 #ifndef FOLL_WRITE diff --git a/tools/testing/selftests/mm/migration.c b/tools/testing/selftes= ts/mm/migration.c index f19d53c69576..fd35f8a7b5b8 100644 --- a/tools/testing/selftests/mm/migration.c +++ b/tools/testing/selftests/mm/migration.c @@ -20,7 +20,6 @@ =20 #define TWOMEG (2<<20) #define RUNTIME (20) -#define ALIGN(x, a) (((x) + (a - 1)) & (~((a) - 1))) =20 HUGETLB_SETUP_DEFAULT_PAGES(1) =20 diff --git a/tools/testing/selftests/mm/run_vmtests.sh b/tools/testing/self= tests/mm/run_vmtests.sh index d09f9f6a384e..2652a7920b80 100755 --- a/tools/testing/selftests/mm/run_vmtests.sh +++ b/tools/testing/selftests/mm/run_vmtests.sh @@ -402,6 +402,8 @@ CATEGORY=3D"pfnmap" run_test ./pfnmap # COW tests CATEGORY=3D"cow" run_test ./cow =20 +CATEGORY=3D"thp" run_test ./folio_order_check + CATEGORY=3D"thp" run_test ./khugepaged =20 CATEGORY=3D"thp" run_test ./khugepaged -s 2 diff --git a/tools/testing/selftests/mm/vm_util.h b/tools/testing/selftests= /mm/vm_util.h index 39dfb18dc10c..ce05bce4670d 100644 --- a/tools/testing/selftests/mm/vm_util.h +++ b/tools/testing/selftests/mm/vm_util.h @@ -10,6 +10,8 @@ #include =20 #define BIT_ULL(nr) (1ULL << (nr)) +#define ALIGN(x, a) (((x) + (a) - 1) & ~((a) - 1)) + #define PM_SOFT_DIRTY BIT_ULL(55) #define PM_MMAP_EXCLUSIVE BIT_ULL(56) #define PM_UFFD_WP BIT_ULL(57) --=20 2.54.0 From nobody Mon Sep 28 23:56:19 2026 Received: from fhigh-a7-smtp.messagingengine.com (fhigh-a7-smtp.messagingengine.com [103.168.172.158]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3D2E334D38B; Sat, 15 Aug 2026 01:59:27 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.158 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759169; cv=none; b=ZTEMPxGWAEJ8BSSwuaMdDPInSEsMS31zny6fI6/6JxXh6foxFIqFzFUkQet2RSnIkb9tbJYELjLHXq++sclrmXT68kRto+glB5m/XONfYDpbXll1NyP8LPR/1kmjb7Qz7LYs3nVB/ciIjOU2Z72DGjdaa8QIPZrjrqmi1LXd75c= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759169; c=relaxed/simple; bh=GAAP1fJySb/pKKMl9ZT51tN8hzshdxLLNhN49eZapkY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=lRbbNPWeG6XpzNZE725YiY7t6oyk4Z9qMcqMNVy23GkBllVQEIX9advRx9FJHlI2X4BmcFNKcxTP1e827EzfYk3zR/kXqtIRIcPeEY0yud6wiubpZ/tjoEubIPUXAA/OcXcATEXT0h6YEKBByOmLtQGD333aMhXvjxgKl0oH4aE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=rU198r7r; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=I34fodJa; arc=none smtp.client-ip=103.168.172.158 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="rU198r7r"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="I34fodJa" Received: from phl-compute-05.internal (phl-compute-05.internal [10.202.2.45]) by mailfhigh.phl.internal (Postfix) with ESMTP id 51902140011C; Fri, 14 Aug 2026 21:59:26 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-05.internal (MEProxy); Fri, 14 Aug 2026 21:59:26 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786759166; x= 1786845566; bh=CZeRnpjSx97tLsIjgu0PpuEyM1aG2ncyWgDrDsC6Hoo=; b=r U198r7rP/Vc4FNI1a1rsYY0AuCB2FeQfyW9Ri/4ejuwfJMpi3AYfcdrJB0ZXtqBu ZeGi2R+yMTFrrsnIo/co7ZOuqesSYpNeyLVvxg64gFVeItTPa8rKFYs1xe8k2OB+ HO+OPSHVMnwTl3z/yyk/bT2uc0KfGjI/RB0i4KKrg43yXI8f1Hholgg5kPw8ZhjH 0cHpcADSCIBupi7ToBr7GBqPA8az13G46PpE6FY0dgjit/5D7tt+OhdroKhxGGrD LSKIAa81stZFgFwUGlBskXv5thasHWRnU68Bczjp8ehsE2I6//C5z7LjM8m7402q tPAZftuJJuP4hJzB0Uwjw== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786759166; x=1786845566; bh=C ZeRnpjSx97tLsIjgu0PpuEyM1aG2ncyWgDrDsC6Hoo=; b=I34fodJaX6Ch2ggTS A+vh4cgl2R5Ef3LjfuLTe/PiY2DiSng9Hz3KZaT7RjpP26/Lf51K3RUo3AChNCVJ oUS/U+Us3urjBDykrmAaEAVMlTHd7Tt1N0lVBk8tE/EcQVI3PapHSSq0/QS6BXBP vZ8sAzRXefCs3yUvUXQWFLlxrgji7DmsmXEEH8H4LpJ04b5tLOKVYy+gYIBEnhiJ dVgNys0iAa1gPVP0cp40n2QxaaAl3rHUWRVjlOAWaURqJnZ60Io4WNd5JTPhgKoM llSzhplwvz50PR32F7hBqKMT5InxPRrf25Dutq7vUtsJQ/PnqErUK2o262OxA0ti RAijA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGF7teVIHNqoa5p2A8iL+trOmU4pExV0BsGEGBxavUqSy3vvyW/t2uVaC82ybAhbC FZe2jS418ApiE9B8lCjDVAvs2O/1hVxxIyd2QmVeomYKnxIQlXrdOkpVI2/A6aay1R5wM6 752pE1oD84VvsYvCM8wHufwfHzQzVg5WF64m0thPV1o6CCL0N8QD0816LOjK9LaF9KdRdp ShDgDyNTb/bU59JqJj/g3oSWXYBnQXQHEYh7kafl/PHtgQjoSit3F73LPXI/q1yj04b//y CUybFuAxuhlNqjceRELXNaCNxBn+qS4c4JK+cmo8T/HSoyzm/Kml9mcKob58Fnvz9v9Li7 CoAEwNoLKJKn1qX+J1+9JwWz88Soz0MhWY28+Pguh3s45O9cQLypQR3KYAcDBhW6yE7XiM IyLtH1lmqZlpPb/jzlLd6/i0Jk5epL1pZtK304Xa71zdyS0MDI21QIx3u3gG0UcrR26/D9 SkQYYRyecImQhU4ldepeg+YHoG4G0/PcBdHE016KA0hy6E33QauH0LyoPx9juZnXgLF7/4 z9Q7Evr1kSl9XtCYmVfPWcO7x6pujAbU8SvmOSfEfvTyj8vgAPZSxLsx3PFEbuY8tiMQhz w/GIfiAugxuu4ruNTl34Ng8To0SmNlHuT1ksGvmpb2iecpc9aigE6lCO6c1Q X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 14 Aug 2026 21:59:25 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v4 10/19] selftests/mm: add khugepaged completion barrier helper Date: Sat, 15 Aug 2026 02:58:52 +0100 Message-ID: <20260815015901.1236937-11-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260815015901.1236937-1-kirill@shutemov.name> References: <20260815015901.1236937-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Race and functional tests need to drive khugepaged in step: set up a layout, let one full scan pass over it, check the result. The khugepaged selftest already waits for full_scans to advance by two, but only makes progress if scan_sleep_millisecs happens to be short. Lift it into khugepaged_full_pass() and drive it through sysfs: any store to scan_sleep_millisecs wakes the daemon, so the barrier completes whatever the scan cadence. A store can be lost when the daemon is between scans, so it keeps storing until the pass lands; a store to an awake daemon costs nothing and queues no extra pass. One wake completes one pass only if the whole mm list fits in a scan batch, so callers need a large pages_to_scan. Settings pushes must not start passes either. A store to either sleep knob wakes the daemon, so thp_write_settings() now writes a khugepaged knob only when its value changes. The other knobs do not wake, but writing them uniformly costs nothing. thp_update_num() is exported for tests that want the same restraint. Assisted-by: Claude-Code:claude-opus-5 Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) --- .../testing/selftests/mm/hugepage_settings.c | 74 ++++++++++++++++--- .../testing/selftests/mm/hugepage_settings.h | 3 + 2 files changed, 68 insertions(+), 9 deletions(-) diff --git a/tools/testing/selftests/mm/hugepage_settings.c b/tools/testing= /selftests/mm/hugepage_settings.c index d7917dce3aba..992efee17b71 100644 --- a/tools/testing/selftests/mm/hugepage_settings.c +++ b/tools/testing/selftests/mm/hugepage_settings.c @@ -183,6 +183,19 @@ void thp_read_settings(struct thp_settings *settings) } } =20 +/* + * Write only on change: a store to either sleep knob wakes khugepaged -- + * __sleep_millisecs_store() clears khugepaged_sleep_expire and wakes the + * queue -- and settings pushes/pops must not start scan passes nobody + * asked for; khugepaged_full_pass() is the only sanctioned wake. The + * other knobs do not wake, but writing them the same way costs nothing. + */ +void thp_update_num(const char *name, unsigned long num) +{ + if (thp_read_num(name) !=3D num) + thp_write_num(name, num); +} + void thp_write_settings(struct thp_settings *settings) { struct khugepaged_settings *khugepaged =3D &settings->khugepaged; @@ -198,15 +211,15 @@ void thp_write_settings(struct thp_settings *settings) shmem_enabled_strings[settings->shmem_enabled]); thp_write_num("use_zero_page", settings->use_zero_page); =20 - thp_write_num("khugepaged/defrag", khugepaged->defrag); - thp_write_num("khugepaged/alloc_sleep_millisecs", - khugepaged->alloc_sleep_millisecs); - thp_write_num("khugepaged/scan_sleep_millisecs", - khugepaged->scan_sleep_millisecs); - thp_write_num("khugepaged/max_ptes_none", khugepaged->max_ptes_none); - thp_write_num("khugepaged/max_ptes_swap", khugepaged->max_ptes_swap); - thp_write_num("khugepaged/max_ptes_shared", khugepaged->max_ptes_shared); - thp_write_num("khugepaged/pages_to_scan", khugepaged->pages_to_scan); + thp_update_num("khugepaged/defrag", khugepaged->defrag); + thp_update_num("khugepaged/alloc_sleep_millisecs", + khugepaged->alloc_sleep_millisecs); + thp_update_num("khugepaged/scan_sleep_millisecs", + khugepaged->scan_sleep_millisecs); + thp_update_num("khugepaged/max_ptes_none", khugepaged->max_ptes_none); + thp_update_num("khugepaged/max_ptes_swap", khugepaged->max_ptes_swap); + thp_update_num("khugepaged/max_ptes_shared", khugepaged->max_ptes_shared); + thp_update_num("khugepaged/pages_to_scan", khugepaged->pages_to_scan); =20 if (dev_queue_read_ahead_path[0]) write_num(dev_queue_read_ahead_path, settings->read_ahead_kb); @@ -230,6 +243,49 @@ void thp_write_settings(struct thp_settings *settings) } } =20 +/* + * Completion barrier for khugepaged: wait until a full scan pass that + * started after this call has finished. full_scans must advance by two; + * a +1 step may complete a pass that examined this mm before the + * caller's setup was in place. + * + * Any store to scan_sleep_millisecs wakes the daemon, so the barrier works + * whatever the configured scan cadence -- but a store can be lost. + * __sleep_millisecs_store() clears khugepaged_sleep_expire and wakes the + * queue; if the daemon is between scans rather than sleeping, it sets + * khugepaged_sleep_expire itself on the way into khugepaged_wait_work() a= nd + * then sleeps for the full interval, having never seen the store. So keep + * storing until the pass lands; a store while the daemon is awake costs + * nothing and does not queue an extra pass. + * + * One wake completes one full pass only if the whole mm list fits in + * one scan batch, so callers must pair this with a large + * pages_to_scan. + */ +bool khugepaged_full_pass(unsigned int timeout_s) +{ + unsigned long deadline_ms =3D timeout_s * 1000UL; + unsigned long sleep_ms =3D + thp_read_num("khugepaged/scan_sleep_millisecs"); + unsigned long elapsed_ms =3D 0; + int pass; + + for (pass =3D 0; pass < 2; pass++) { + unsigned long target =3D + thp_read_num("khugepaged/full_scans") + 1; + + while (thp_read_num("khugepaged/full_scans") < target) { + if (elapsed_ms >=3D deadline_ms) + return false; + thp_write_num("khugepaged/scan_sleep_millisecs", + sleep_ms); + usleep(10 * 1000); + elapsed_ms +=3D 10; + } + } + return true; +} + struct thp_settings *thp_current_settings(void) { if (!settings_index) { diff --git a/tools/testing/selftests/mm/hugepage_settings.h b/tools/testing= /selftests/mm/hugepage_settings.h index 726c73c43c05..ba7d38370d43 100644 --- a/tools/testing/selftests/mm/hugepage_settings.h +++ b/tools/testing/selftests/mm/hugepage_settings.h @@ -70,6 +70,7 @@ int thp_read_string(const char *name, const char * const = strings[]); void thp_write_string(const char *name, const char *val); unsigned long thp_read_num(const char *name); void thp_write_num(const char *name, unsigned long num); +void thp_update_num(const char *name, unsigned long num); =20 void thp_write_settings(struct thp_settings *settings); void thp_read_settings(struct thp_settings *settings); @@ -83,6 +84,8 @@ static inline void thp_save_settings(void) hugepage_save_settings(/* thp =3D */ true, /* hugetlb =3D */ false); } =20 +bool khugepaged_full_pass(unsigned int timeout_s); + void thp_set_read_ahead_path(char *path); unsigned long thp_supported_orders(void); unsigned long thp_shmem_supported_orders(void); --=20 2.54.0 From nobody Mon Sep 28 23:56:19 2026 Received: from fout-a6-smtp.messagingengine.com (fout-a6-smtp.messagingengine.com [103.168.172.149]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CB93C3515CB; Sat, 15 Aug 2026 01:59:28 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.149 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759171; cv=none; b=R+kRkFxXdCv3SdXjcQ8t2bBXBAA4JbeHWKumb8mL8iixeSP55AHUqt/hJbmSDpmn9gOYYRIqQ/9RBVWQN1Ln9djC33nrpQYWJ7FDJKOTpg+PMk6UGawjzYc17WWrjq0SDmPv5LMQ87VVtVyjihYrOikmEgw+LueK9+/DkzZp4IA= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759171; c=relaxed/simple; bh=FfIoOfVH7GdM5rQ+AWToxqhkuMUUKkDSExEaSJkzVvI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=DfEtZ4SC3W6Dn/vI3ZNK47etp+10VT1eajPBB8Kp9oP8Gzwd/uFiswxmerjOd1yHa/K1HGHjSbb5/e0sf57uJ3OucFIvzujTEd92cn/+eoodX2NCMgGaDrzidldYv1lwq5rJM7OH8rdTg3fYjz2Qf6YXhz6IDbTxrAgivIVFGlg= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=b/g+kqZm; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=RhKRP1ko; arc=none smtp.client-ip=103.168.172.149 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="b/g+kqZm"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="RhKRP1ko" Received: from phl-compute-04.internal (phl-compute-04.internal [10.202.2.44]) by mailfout.phl.internal (Postfix) with ESMTP id 053DAEC0222; Fri, 14 Aug 2026 21:59:28 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-04.internal (MEProxy); Fri, 14 Aug 2026 21:59:28 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786759168; x= 1786845568; bh=TndxdCltkGuvzZg0b5nL1NfZtt1dBi/ccjiEs/49rKo=; b=b /g+kqZmdyT/Fnm90OChIX7THylkcZdL5RzQwijJKaZewSVJuIUdkxYulJNaw5kFj Zwc/lC0J59tY/+pMlMD4DobTXP64Nb5MJB+veuzukFtSp1UGpKB/tpwOT1EI9GJu Jh2YyWFJ/jdRGlLaxSeSoFOBUm9yyalSOsPEr94i8Pmu27Wb2g36N5Za2/sn1Jc7 sgvqrpqL1IY++6wfYetxVhdV/8pYMKWwChmhzaucgj88jzH1d4fYm27H1EiEkeuL 5Tj50eh8rMrG+7RF1yK4yXjgZpvVxlPj429Sq9Glj8tzIZ+mDXmzBhHsVy7nQm7t aHR3oGJhgik+tp6ObZ4Eg== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786759168; x=1786845568; bh=T ndxdCltkGuvzZg0b5nL1NfZtt1dBi/ccjiEs/49rKo=; b=RhKRP1koDI+ptIxl3 HGLgxJxt9j3omxH9cRDrfouD5bJwupqLsPOl83WrtPBwtCFXZy58Tz7ibrLlgcKY /9CRyrwf3nOKQxxXsY0zWzzoE0cF/ZG+pxGr/+nk4/NNRCT+oc1RlEr63EWoNVZ7 ndD4VnzTj9Q1/8EGF6chW0Qr0/s/oXRXaDKRscwVAUBUJHIu36COh/z3Nf/ovIuu pqfXrhF/b1UVOdJwGjuINmyBp9GNF7g/OyiGjK6ziB6F+GQZPXP1+sZ2Cyks8HO2 DDl26DeNPyHMMqpAI2k1iE/AsrkGlDD8w9aZexMT6GqgZ+t49libXTx0mTitVJh3 JbZkg== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEYpEQiBdLKlCaLp317TOYqYwO99jnuX4YrBujo3RuW+IIGXdYj4qE1ezcGRvbtOZ hc7RfnfkdlTehb5Dkvk4VZTcyRAqm2CcEjaldkIQshIttrB2HSeWyAKQfsR+ykmymFv8qO aC/+wh3joJiQccaWn5NjVrqLVVMYdoVO37A5ZL6SnR75seJlCquk1NTOiA0EvXHIVSJ25Q iFONSccd/wEzNfQZlOvvQQf5bCKiGIALXd0Rd37Vcnrh56q8DfWwmvEu2ZgBGSMa525iaP HkMF5MKO3LDiNiHXJVHjo5f/L4XJttiAOfdXhljkaHgZI649okv5M81/eeF8HfhCZr1UQh NZgwgKpraJQ4GfA8YlvrMbphan98J2x248uA7+GQJDmU5vuQuFfP9TqLssry0XavFL6pSe CSbzkCGUFIj/OQN8jsfYXJAUYAFsWIpI2kcVV/xHNSdtuebZ+h4xBojLYlydWLCcjiyZci nHPuSmfGs/mzRGIfDM/mxAj55kilgC8yhXvrJHezys3fagfeVEqBfLZdTM15LcTq09UZz5 qu+H3Ov+ZvdmgtJ7488deEr22TLks+7nZ6eYeFdgP868jPPfsXQmzQ5WLYsRKN51XmETcJ 1gMxzmWboYjDOFCJtnCipj1RCiFNNbyKWTOyNN1WFB+zaT8yhIzXuN7inBrQ X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 14 Aug 2026 21:59:27 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v4 11/19] selftests/mm: add order-parameterized khugepaged collapse cases Date: Sat, 15 Aug 2026 02:58:53 +0100 Message-ID: <20260815015901.1236937-12-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260815015901.1236937-1-kirill@shutemov.name> References: <20260815015901.1236937-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The mthp_khugepaged context runs the generic cases at a sub-PMD order, which answers how many folios of that order a range ends up with. It cannot say which window they are in, so "the populated window collapsed and its neighbour did not" and "one window collapsed twice" look alike. Add four cases that check each aligned window on its own, with the folio-order helpers in vm_util: - collapse_order_single_window(): only the populated window collapses; - collapse_order_partial_window(): the default max_ptes_none lets a window with one present PTE collapse; - collapse_order_max_ptes_none(): with max_ptes_none=3D0 a full window collapses and one missing a page does not; - collapse_order_mixed_sources(): sources that are already large folios of a smaller order collapse to the target. Each case faults its region before MADV_HUGEPAGE with only the target order enabled, so the sources are order 0 and the result can only come from khugepaged. They wait for a full pass rather than for the result to appear: without a completed pass, "not collapsed" and "not scanned yet" are the same thing. Assisted-by: Claude-Code:claude-opus-5 Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/khugepaged.c | 230 ++++++++++++++++++++++++ 1 file changed, 230 insertions(+) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index 0008862e7cbc..0489967d6ee0 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -31,6 +31,8 @@ static unsigned long page_size; static int hpage_pmd_nr; static int anon_order; static int collapse_order; +static int pagemap_fd =3D -1; +static int kpageflags_fd =3D -1; =20 #define PID_SMAPS "/proc/self/smaps" #define TEST_FILE "collapse_test_file" @@ -1227,6 +1229,216 @@ static void madvise_retracted_page_tables(struct co= llapse_context *c, ksft_test_result_report(exit_status, "%s\n", __func__); } =20 +/* Smallest order khugepaged will consider for mTHP collapse. */ +#define MIN_MTHP_ORDER 2 + +/* + * Order-parameterized collapse cases for the mthp_khugepaged context. Wh= at + * they add over the generic cases run under that context is per-window + * detection: which aligned window collapsed, and which of its neighbours = did + * not. check_huge() answers how many folios of the order the range holds, + * which cannot tell one window from another. + * + * The region is faulted before MADV_HUGEPAGE, and the target order is only + * enabled for madvise, so the sources are always order 0 and the collapse + * product can only have come from khugepaged. + */ +static size_t mthp_window_size(void) +{ + return page_size << collapse_order; +} + +static void mthp_push_target_order(void) +{ + struct thp_settings settings =3D *thp_current_settings(); + int i; + + /* + * The target order, for madvise only, and nothing else enabled: the + * cases fault their region before MADV_HUGEPAGE, so the sources are + * order 0 whatever -s asked the fault path for. That matters for the + * cases built around a hole -- a large source folio would fill it in + * and the window would collapse after all. + * collapse_order_mixed_sources enables the source order it wants on + * top of this. + */ + settings.thp_enabled =3D THP_NEVER; + for (i =3D 0; i < NR_ORDERS; i++) + settings.hugepages[i].enabled =3D THP_NEVER; + settings.hugepages[collapse_order].enabled =3D THP_MADVISE; + thp_push_settings(&settings); +} + +static bool window_collapsed(void *p, size_t len) +{ + return is_range_backed_by_folio_orders(p, len, collapse_order, + pagemap_fd, kpageflags_fd); +} + +/* No aligned window in [p, p + len) is backed at the target order. */ +static bool window_not_collapsed(void *p, size_t len) +{ + size_t window =3D mthp_window_size(); + char *addr =3D p; + + for (; len >=3D window; addr +=3D window, len -=3D window) { + if (window_collapsed(addr, window)) + return false; + } + return true; +} + +static bool khugepaged_wait_full_pass(void) +{ + /* Wait up to 30 seconds for the pass to complete. */ + return khugepaged_full_pass(30); +} + +static void collapse_order_single_window(struct collapse_context *c, + struct mem_ops *ops) +{ + size_t window =3D mthp_window_size(); + void *p; + + mthp_push_target_order(); + + p =3D ops->setup_area(1); + ops->fault(p, window, 2 * window); + if (!window_not_collapsed(p, hpage_pmd_size)) + ksft_exit_fail_msg("Unexpected large folio after fault\n"); + + madvise(p, hpage_pmd_size, MADV_HUGEPAGE); + ksft_print_msg("Collapse one fully populated window..."); + if (!khugepaged_wait_full_pass()) + fail("Timeout"); + else if (window_collapsed(p + window, window) && + window_not_collapsed(p, window) && + window_not_collapsed(p + 2 * window, + hpage_pmd_size - 2 * window)) + success("OK"); + else + fail("Fail"); + + validate_memory(p, window, 2 * window); + ops->cleanup_area(p, hpage_pmd_size); + thp_pop_settings(); + ksft_test_result_report(exit_status, "%s\n", __func__); +} + +static void collapse_order_partial_window(struct collapse_context *c, + struct mem_ops *ops) +{ + void *p; + + mthp_push_target_order(); + + p =3D ops->setup_area(1); + ops->fault(p, 0, page_size); + if (!window_not_collapsed(p, hpage_pmd_size)) + ksft_exit_fail_msg("Unexpected large folio after fault\n"); + + madvise(p, hpage_pmd_size, MADV_HUGEPAGE); + ksft_print_msg("Collapse window with single PTE entry present..."); + if (!khugepaged_wait_full_pass()) + fail("Timeout"); + else if (window_collapsed(p, mthp_window_size())) + success("OK"); + else + fail("Fail"); + + validate_memory(p, 0, page_size); + ops->cleanup_area(p, hpage_pmd_size); + thp_pop_settings(); + ksft_test_result_report(exit_status, "%s\n", __func__); +} + +static void collapse_order_max_ptes_none(struct collapse_context *c, + struct mem_ops *ops) +{ + struct thp_settings settings; + size_t window =3D mthp_window_size(); + void *p; + + mthp_push_target_order(); + settings =3D *thp_current_settings(); + settings.khugepaged.max_ptes_none =3D 0; + thp_push_settings(&settings); + + p =3D ops->setup_area(1); + ops->fault(p, 0, 2 * window - page_size); + if (!window_not_collapsed(p, hpage_pmd_size)) + ksft_exit_fail_msg("Unexpected large folio after fault\n"); + + madvise(p, hpage_pmd_size, MADV_HUGEPAGE); + ksft_print_msg("Collapse full window, not the one missing a page..."); + if (!khugepaged_wait_full_pass()) + fail("Timeout"); + else if (window_collapsed(p, window) && + window_not_collapsed(p + window, window)) + success("OK"); + else + fail("Fail"); + + validate_memory(p, 0, 2 * window - page_size); + ops->cleanup_area(p, hpage_pmd_size); + thp_pop_settings(); + thp_pop_settings(); + ksft_test_result_report(exit_status, "%s\n", __func__); +} + +static void collapse_order_mixed_sources(struct collapse_context *c, + struct mem_ops *ops) +{ + struct thp_settings settings; + void *p; + + if (collapse_order <=3D MIN_MTHP_ORDER) { + ksft_test_result_skip("%s: no source order below target\n", + __func__); + return; + } + + mthp_push_target_order(); + + /* Fault the whole region as order-MIN_MTHP_ORDER folios. */ + settings =3D *thp_current_settings(); + settings.hugepages[MIN_MTHP_ORDER].enabled =3D THP_ALWAYS; + thp_push_settings(&settings); + p =3D ops->setup_area(1); + ops->fault(p, 0, hpage_pmd_size); + thp_pop_settings(); + + /* + * The order is enabled, but the allocator can still fall back under + * fragmentation. That leaves nothing to collapse from, which is the + * machine's answer rather than a reason to end the run. + */ + if (!is_range_backed_by_folio_orders(p, hpage_pmd_size, MIN_MTHP_ORDER, + pagemap_fd, kpageflags_fd)) { + ksft_print_msg("No order-%d sources to collapse...", + MIN_MTHP_ORDER); + skip("Skip"); + ops->cleanup_area(p, hpage_pmd_size); + thp_pop_settings(); + ksft_test_result_report(exit_status, "%s\n", __func__); + return; + } + + madvise(p, hpage_pmd_size, MADV_HUGEPAGE); + ksft_print_msg("Collapse region backed by smaller large folios..."); + if (!khugepaged_wait_full_pass()) + fail("Timeout"); + else if (window_collapsed(p, hpage_pmd_size)) + success("OK"); + else + fail("Fail"); + + validate_memory(p, 0, hpage_pmd_size); + ops->cleanup_area(p, hpage_pmd_size); + thp_pop_settings(); + ksft_test_result_report(exit_status, "%s\n", __func__); +} + static void usage(void) { fprintf(stderr, "\nUsage: ./khugepaged [OPTIONS] [dir]\n\n"); @@ -1395,6 +1607,20 @@ int main(int argc, char **argv) =20 parse_test_type(argc, argv); =20 + if (mthp_khugepaged_context && + !(thp_supported_orders() & (1UL << collapse_order))) + ksft_exit_skip("Order %d is not a supported anon THP order\n", + collapse_order); + + if (mthp_khugepaged_context) { + pagemap_fd =3D open("/proc/self/pagemap", O_RDONLY); + if (pagemap_fd < 0) + ksft_exit_fail_perror("open(/proc/self/pagemap)"); + kpageflags_fd =3D open("/proc/kpageflags", O_RDONLY); + if (kpageflags_fd < 0) + ksft_exit_fail_perror("open(/proc/kpageflags)"); + } + setbuf(stdout, NULL); =20 /* @@ -1450,6 +1676,10 @@ int main(int argc, char **argv) TEST(collapse_empty, madvise_context, anon_ops); =20 TEST(collapse_single_mthp, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_single_window, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_partial_window, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_max_ptes_none, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_mixed_sources, mthp_khugepaged_context, anon_ops); =20 TEST(collapse_single_pte_entry, khugepaged_context, anon_ops); TEST(collapse_single_pte_entry, khugepaged_context, read_only_file_ops); --=20 2.54.0 From nobody Mon Sep 28 23:56:19 2026 Received: from fhigh-a7-smtp.messagingengine.com (fhigh-a7-smtp.messagingengine.com [103.168.172.158]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 71DA7351C1E; Sat, 15 Aug 2026 01:59:30 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.158 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759172; cv=none; b=baFSs5htUZii3jOdTaoybVzVltEKfKNnOncI5kKJ15qs34eKlx7IT7MV58UCEaju236GgDsrWjReXYfr9A/gLn4QVdViG7izwewbZKvuvMAUOm2PT7rieIC3NPK5keYYW1ocDw0fsRy/FtEb8dsE73gOqFhn4qKvCGbob+mGMec= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759172; c=relaxed/simple; bh=lyolELMNwifT3vqJ7y6+hOWJqyxzGsPTLeCTX/BOPZA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=B0tJsrEtXXQfBLt25CQEmvq6n629+Q11nJxAOmvoBnecqf4z94FDtlIC5gke46Cv1BCiaclnkw4QGRzeM8NFvBssajnq8ZysDxHTnNSw1FlVOtvQqbklSzfEXXW695Wp2tGEZSo2/sNa/jBoXagrp7Uzfow9+qnffjJ0Dfh7r9A= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=hzITPJN9; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=G9G/E9HX; arc=none smtp.client-ip=103.168.172.158 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="hzITPJN9"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="G9G/E9HX" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfhigh.phl.internal (Postfix) with ESMTP id 884B8140011E; Fri, 14 Aug 2026 21:59:29 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-06.internal (MEProxy); Fri, 14 Aug 2026 21:59:29 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786759169; x= 1786845569; bh=u/N4uH7e9xwqe95L8EVPoUAk/HrCwQSNsTbS9xlOgGE=; b=h zITPJN91H83Wz/W0BY1/dG0cXsoMVzbyhrLEHHbizETwXdKFDBHCaZRBV1Lf81iS CW/XQasGuFG0/3dTp4Fl9qn/4d6OBa5oL/ZikbhFDPUpWrk18AFYKgv3bF5Vhqq3 myxFdJ62NhhN8IGT9gsgsro8VoUDSCpnDkmcm2N27vClzb8kGF4WBpYnNxlrbefk Jp5GkDavEcNt4L3Q729r4sw68L2PKNImRe5mCLlQJJfXxEGGfRnseeo2tDVFPJBO cTqrvTzbOUZhlO19cmOtvEl6b1dHHGw+TsFYtVJPW2tDOv3Tp94Idx/G8vYE0J9y 8VQTa5a9gLVC/NUBRGMgw== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786759169; x=1786845569; bh=u /N4uH7e9xwqe95L8EVPoUAk/HrCwQSNsTbS9xlOgGE=; b=G9G/E9HXiMTOh5B2i mbf1cWioF7YzSd4XH0gEwRzeOCfxSR4RRIAfyWGPywd1clhGRU3Bn+FArilclWJQ vNqAtYbxViHytHAv0RMiqCe17gnE+t0ViL1GWBQVr1+bX4tdv8kcGxNjFuvfK/Aa UgIO6TFLoDMo1uMIFK8xpf+TQsw5geaoDZOsTg/xiaFa1uMkDCiN62x759IJZrpl CIQo9moYgdPGUoyNxCevMBW44oh4Gx6ZuOq8Hn2jPeT1ekxDceiZUJMuiC33B2fT 8DmaZHL2tYuAtair80tP6RHY2cxsHzM73cpN6HtBh8PoksQ/K43dq88XZ402VOTN iqlMg== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGF7teVIHNqoa5p2A8iL+trOmU4pExV0BsGEGBxavUqSy3vvyW/t2uVaC82ybAhbC FZe2jS418ApiE9B8lCjDVAvs2O/1hVxxIyd2QmVeomYKnxIQlXrdOkpVI2/A6aay1R5wM6 752pE1oD84VvsYvCM8wHufwfHzQzVg5WF64m0thPV1o6CCL0N8QD0816LOjK9LaF9KdRdp ShDgDyNTb/bU59JqJj/g3oSWXYBnQXQHEYh7kafl/PHtgQjoSit3F73LPXI/q1yj04b//y CUybFuAxuhlNqjceRELXNaCNxBn+qS4c4JK+cmo8T/HSoyzm/Kml9mcKob58Fnvz9v9LWE nNmRLDy6osiysgx2iCEJ3pwXFBkRpHwTdMWlzFr0VzVRxFjUFgk/QX1VAGpcrxNr0+Cez4 Te957jer6VGOf6T93BBz/Bi+W9G3xUIm82OhU9tP/yJkl6U86+aQidXFudg9c7KcHzc35U aRCO/dV5njQzP1yXynmLgGAtWKPTsLHZX2pa7BFJmmvUrtdwK0AqpmzVDHP2IOXhZ4LZu8 /AHFEeuMAY4IG2J+eVWCE/1SH9hJYpCnI8mpJTva/PcuMNiwTVwiAuC3i28H/nn9npdOe1 nLq2CauCln0aHynZYpMNoY2xUzWncmq04dwcHpKPNLb1dC57B/mxF1p889kQ X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 14 Aug 2026 21:59:28 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v4 12/19] selftests/mm: parameterize the mixed-source collapse case by source order Date: Sat, 15 Aug 2026 02:58:54 +0100 Message-ID: <20260815015901.1236937-13-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260815015901.1236937-1-kirill@shutemov.name> References: <20260815015901.1236937-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" collapse_order_mixed_sources() faults its region as order-2 folios and collapses them to the -c target. Order 2 sits below the contpte threshold on both arm64 page-size configurations, so nothing in this suite unfolds a contpte source on purpose. Let -s name the source order alongside -c. The case then faults at that order, keeping order 2 when -s is absent, and the source order has to be a supported mTHP order below the target. The other mTHP cases are unaffected: mthp_push_target_order() enables only the target order. "-s 5 -c 7" on arm64/64K then collapses contpte-mapped sources into a larger mTHP. Assisted-by: Claude-Code:claude-opus-5 Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) Acked-by: Lorenzo Stoakes (ARM) --- tools/testing/selftests/mm/khugepaged.c | 25 +++++++++++++++---------- 1 file changed, 15 insertions(+), 10 deletions(-) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index 0489967d6ee0..1844ddd77b59 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -1389,10 +1389,13 @@ static void collapse_order_max_ptes_none(struct col= lapse_context *c, static void collapse_order_mixed_sources(struct collapse_context *c, struct mem_ops *ops) { + int source_order =3D anon_order ? anon_order : MIN_MTHP_ORDER; struct thp_settings settings; void *p; =20 - if (collapse_order <=3D MIN_MTHP_ORDER) { + /* Sources must be a supported mTHP order strictly below the target. */ + if (source_order >=3D collapse_order || + !(thp_supported_orders() & (1UL << source_order))) { ksft_test_result_skip("%s: no source order below target\n", __func__); return; @@ -1400,23 +1403,22 @@ static void collapse_order_mixed_sources(struct col= lapse_context *c, =20 mthp_push_target_order(); =20 - /* Fault the whole region as order-MIN_MTHP_ORDER folios. */ + /* Fault the whole region as order-@source_order folios. */ settings =3D *thp_current_settings(); - settings.hugepages[MIN_MTHP_ORDER].enabled =3D THP_ALWAYS; + settings.hugepages[source_order].enabled =3D THP_ALWAYS; thp_push_settings(&settings); p =3D ops->setup_area(1); ops->fault(p, 0, hpage_pmd_size); thp_pop_settings(); =20 /* - * The order is enabled, but the allocator can still fall back under - * fragmentation. That leaves nothing to collapse from, which is the - * machine's answer rather than a reason to end the run. + * The order is enabled and supported, but the allocator can still fall + * back under fragmentation. That leaves nothing to collapse from, + * which is the machine's answer rather than a reason to end the run. */ - if (!is_range_backed_by_folio_orders(p, hpage_pmd_size, MIN_MTHP_ORDER, + if (!is_range_backed_by_folio_orders(p, hpage_pmd_size, source_order, pagemap_fd, kpageflags_fd)) { - ksft_print_msg("No order-%d sources to collapse...", - MIN_MTHP_ORDER); + ksft_print_msg("No order-%d sources to collapse...", source_order); skip("Skip"); ops->cleanup_area(p, hpage_pmd_size); thp_pop_settings(); @@ -1425,7 +1427,8 @@ static void collapse_order_mixed_sources(struct colla= pse_context *c, } =20 madvise(p, hpage_pmd_size, MADV_HUGEPAGE); - ksft_print_msg("Collapse region backed by smaller large folios..."); + ksft_print_msg("Collapse region backed by order-%d sources...", + source_order); if (!khugepaged_wait_full_pass()) fail("Timeout"); else if (window_collapsed(p, hpage_pmd_size)) @@ -1456,6 +1459,8 @@ static void usage(void) fprintf(stderr, "\t\t-s: mTHP size, expressed as page order.\n"); fprintf(stderr, "\t\t Defaults to 0. Use this size for anon or shmem a= llocations.\n"); fprintf(stderr, "\t\t-c: collapse order for mTHP collapse, expressed as p= age order.\n"); + fprintf(stderr, "\t\t With -s, -s names the mTHP source order for the\= n"); + fprintf(stderr, "\t\t mixed-source case (source order below the target= ).\n"); exit(1); } =20 --=20 2.54.0 From nobody Mon Sep 28 23:56:19 2026 Received: from fout-a6-smtp.messagingengine.com (fout-a6-smtp.messagingengine.com [103.168.172.149]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1E9DB349CD1; Sat, 15 Aug 2026 01:59:32 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.149 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759174; cv=none; b=mpAeNlkTsfDAhNCLJVE4Ei/cvm48MRBoJe+eL9dDezRy8wagiAqGJxn2ANeCY8iclwWiDeczmT+FEe2vN4YL+rRrq/ndWtx6OgfM3bWokElyh1jUZ86RcHS/UKbmt0yxGw3vrHGhrbf/XLxTjOME15Ceusc/sFcTtu7i27V0w30= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759174; c=relaxed/simple; bh=aNTW6PiAmAGvcLEdj+NQQiEV6JDXgFJPgYLpmH8IAsg=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=rLdx7MN/u/Uc7sP2QxxqMjuJiGLfASiMUjx9p8di04tFFaxJms50K0sqzakpzP17EiYl8+xQLjvZgQ/ap4zQKeHNNLOE3E3SiUXgHy0gVve9a703T/deEHut4wFU8LxB8kAopTz1+ctcX3uwrONXdEypdcr0jbyBbtRz6YKhcyk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=BPZi3u6h; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=RjjQOBQw; arc=none smtp.client-ip=103.168.172.149 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="BPZi3u6h"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="RjjQOBQw" Received: from phl-compute-04.internal (phl-compute-04.internal [10.202.2.44]) by mailfout.phl.internal (Postfix) with ESMTP id 331D3EC022A; Fri, 14 Aug 2026 21:59:31 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-04.internal (MEProxy); Fri, 14 Aug 2026 21:59:31 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786759171; x= 1786845571; bh=WtJcwZE0X9O8WD49joaw69JwdIiKCMaehT7QHdQfXo8=; b=B PZi3u6hPnlGjeHkeUxXMLh1ILCR3852OQPB4P9FORawFv74kvZbxt8LkjAWYfpUG 9QUZ4tJfuxTk8vmCAD9/CaJN2XR+/JPUz/VlLKSQrzP8g5wObb2t1ip38zI4N1Fb hp92CB6DSvsdR5dQEEtjEuK27ZY780S5hysO5lNdtmGFgXTQDxSq/ah8ZWjIDYL+ W0u4rafOM7yAQmrifjgjf3Wx3Nvr0yLDD2gmiu2k3gWE00a3v/ahohk5fpx/CrKK RMJVjkzuuSkY6bS1410KeVNbl7z08tzR1/qUIMylgN87p9kveOr+3J1rHaYSmWck /b+na5PTkDkhHYM8BL/yA== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786759171; x=1786845571; bh=W tJcwZE0X9O8WD49joaw69JwdIiKCMaehT7QHdQfXo8=; b=RjjQOBQw5Kjl4fMPQ boZ71lSlEYIBG9CB/lZ241QEkTHAjnN0Fq6w9NM9nFO8BTZHh12Tmg3fGj9YbSGq H3+68wmjDQe1/QyxPgCeDzAagoDRukYKg40mj1gDsF99NETMMYXWJcmM0Ieed9Pd x4hu3LoIREIWa3nVRzxIE2kkyg+3vNpH0MDQaa9t/LIHhtE+jyCIb5X8okTLNEFa xQpUwwgQ8JbTUj4seUkoRdFgEMJyOfpaUK+zmpWwIf7iPLKVtr1TsBTFc8ndb9m2 tq08yVyntT0I3aaJ8rgcyfm1/WbFMvKl1TyysDxUkSrXWR21sR4sJv05TALPOsrS +3A4g== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEYpEQiBdLKlCaLp317TOYqYwO99jnuX4YrBujo3RuW+IIGXdYj4qE1ezcGRvbtOZ hc7RfnfkdlTehb5Dkvk4VZTcyRAqm2CcEjaldkIQshIttrB2HSeWyAKQfsR+ykmymFv8qO aC/+wh3joJiQccaWn5NjVrqLVVMYdoVO37A5ZL6SnR75seJlCquk1NTOiA0EvXHIVSJ25Q iFONSccd/wEzNfQZlOvvQQf5bCKiGIALXd0Rd37Vcnrh56q8DfWwmvEu2ZgBGSMa525iaP HkMF5MKO3LDiNiHXJVHjo5f/L4XJttiAOfdXhljkaHgZI649okv5M81/eeF8HfhCZr1UpK vQ86E+2J6nnLLby5GHhBzZLfixCEp9w2OHJJWRdv/Abd6ZARiVy2pfDbYLsLc21fZaluFI VWPrJmrbL4cz0QUuPWs3Sh4v67FQmhOCoMqPzKJm4/OPeMlO1wyIo9JNLAdSBiqsOjtLoc 7iCucaHc1r8lOvx9cxj2Wd7U0ARH2gSOtUrPRaXUkHlOV/zOUYbAo8JOgpo8dxUz2Tj9Bx o2wW1MHe7dAhwOjCi5NH0cPmVImhTA7MM7GdNY1p6RyMhXtY/GJJyNHhSic91vhIuxu7GJ /jQl/uyxuMzGLurX3ckbIzwnPrnT6WJRm+8Km7wkIMOEr1H18KZ2RHWEHw9g X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 14 Aug 2026 21:59:30 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v4 13/19] selftests/mm: cover a shared-source collapse write race Date: Sat, 15 Aug 2026 02:58:55 +0100 Message-ID: <20260815015901.1236937-14-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260815015901.1236937-1-kirill@shutemov.name> References: <20260815015901.1236937-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" collapse_fork() checks that a fork-shared range collapses in the process that asks for it while the co-sharer keeps its own page, but the co-sharer sits still while that happens. Add a case where the co-sharer writes to the shared range throughout the collapse. CoW has to keep the two sides apart under those writes: the collapsing child must see the content from before the fork, and the writing parent must see only its own writes. The co-sharer unshares one page every 10ms, and only once the collapsing side says it is about to start. Writing the range in a burst breaks CoW on all of it before the collapse begins, leaving the child to collapse pages that are already exclusive to it. Nothing here is new behaviour: the case passes on mainline, and locks in isolation that collapse already provides. Assisted-by: Claude-Code:claude-opus-5 Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/khugepaged.c | 94 +++++++++++++++++++++++++ 1 file changed, 94 insertions(+) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index 1844ddd77b59..856decd2950a 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -1184,6 +1184,97 @@ static void collapse_max_ptes_shared(struct collapse= _context *c, struct mem_ops ksft_test_result_report(exit_status, "%s\n", __func__); } =20 +/* + * Content stays isolated while a co-sharer writes concurrently. A shared + * source is copied live (not frozen), relying on it being CoW - immutable + * for the duration of the copy; a co-sharer's write goes to a CoW copy. T= he + * collapsing child must see the pre-fork content, the writing parent only + * its own writes. + */ +static void collapse_fork_cow_race(struct collapse_context *c, struct mem_= ops *ops) +{ + const unsigned long shared =3D 64 * page_size; + const int stride =3D page_size / sizeof(int); + int wstatus, child_status, i, n =3D shared / page_size; + /* volatile: the loop below must really store, on every iteration */ + volatile int *ip; + pid_t child; + int sync[2]; + char go =3D 1; + void *p; + + p =3D ops->setup_area(1); + ip =3D p; + ops->fault(p, 0, shared); /* shared prefix, pre-fork pattern */ + if (pipe(sync)) + ksft_exit_fail_perror("pipe()"); + + ksft_print_msg("Fork, collapse in the child while the parent rewrites..."= ); + child =3D fork(); + if (!child) { + int collapse_status; + + close(sync[0]); + ops->fault(p, shared, hpage_pmd_size); /* private remainder */ + /* Start the parent unsharing, and give it a head start. */ + if (write(sync[1], &go, 1) !=3D 1) + _exit(KSFT_FAIL); + usleep(5000); + c->collapse("Collapse a range shared with a writing co-sharer", + p, 1, ops, true); + collapse_status =3D exit_status; + for (i =3D 0; i < n; i++) + if (ip[i * stride] !=3D i + 0xdead0000) + break; + if (i =3D=3D n) + success("OK"); + else + fail("Fail: child content"); + /* The content check must not bury a failed collapse. */ + if (exit_status !=3D KSFT_FAIL) + exit_status =3D collapse_status; + ops->cleanup_area(p, hpage_pmd_size); + _exit(exit_status); + } + + close(sync[1]); + if (read(sync[0], &go, 1) !=3D 1) + ksft_exit_fail_msg("child never reached the collapse\n"); + + /* + * Unshare one page at a time. A burst would break CoW on all of them + * in microseconds -- wait_for_scan() does not even poll for TICK -- + * and the child would collapse pages already exclusive to it. + */ + i =3D 0; + do { + if (i < n) + ip[i * stride] =3D i + 0xbeef0000; + i++; + usleep(10 * 1000); + } while (waitpid(child, &wstatus, WNOHANG) =3D=3D 0); + + /* Whatever the paced sweep did not reach, so the check below is exact. */ + for (; i < n; i++) + ip[i * stride] =3D i + 0xbeef0000; + /* A child that died reading the racing pages is a failure, not a zero. */ + child_status =3D WIFEXITED(wstatus) ? WEXITSTATUS(wstatus) : KSFT_FAIL; + + ksft_print_msg("Check the parent sees only its own writes..."); + for (i =3D 0; i < n; i++) + if (ip[i * stride] !=3D i + 0xbeef0000) + break; + if (i =3D=3D n) + success("OK"); + else + fail("Fail: parent content"); + ops->cleanup_area(p, hpage_pmd_size); + /* Same again: our own check must not bury the child's verdict. */ + if (exit_status !=3D KSFT_FAIL) + exit_status =3D child_status; + ksft_test_result_report(exit_status, "%s\n", __func__); +} + static void madvise_collapse_existing_thps(struct collapse_context *c, struct mem_ops *ops) { @@ -1740,6 +1831,9 @@ int main(int argc, char **argv) TEST(collapse_max_ptes_shared, khugepaged_context, anon_ops); TEST(collapse_max_ptes_shared, madvise_context, anon_ops); =20 + TEST(collapse_fork_cow_race, khugepaged_context, anon_ops); + TEST(collapse_fork_cow_race, madvise_context, anon_ops); + TEST(madvise_collapse_existing_thps, madvise_context, anon_ops); TEST(madvise_collapse_existing_thps, madvise_context, read_only_file_ops); TEST(madvise_collapse_existing_thps, madvise_context, read_write_file_rea= d_ops); --=20 2.54.0 From nobody Mon Sep 28 23:56:19 2026 Received: from fout-a6-smtp.messagingengine.com (fout-a6-smtp.messagingengine.com [103.168.172.149]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CA5B3352025; Sat, 15 Aug 2026 01:59:33 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.149 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759176; cv=none; b=DhPlgSYgm3Ex+oeolOrvRU6YgLt+BcRHdLNmEa3IjY7Xs9zsWhKMTFYkby63dwgbBxd6TopNMggMSnRUOOGSVw7bJHDjJif9uxrlDH4cqYiFhUgZnLFbSPmfGug8LTyFrJ0YiHF69hevvavYVwqBoLVGi3s8l30zI24vOhS5C28= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759176; c=relaxed/simple; bh=Z0Be8cjYUQX86SNs6Mf05Hfq+gYbxhcEviirMIwIP30=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=iOk6+4kFHDvdGfxrV0y/BWe+AJr3OOMtxg4Hd62EorVjUzGE6kWY7Rpl3YCQcbXCidJmq4xiSfIzQT/Mz9Mk9cCGUC+0G3F5ZIMaUuP/lhbhRT3AkKkztp7I7WevTxCdl87/hL7pB+rAkBcgNQUzl1hSQfyNXl/cK0HLhOYAhWo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=HZzCPkDE; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=Sa/gT1TL; arc=none smtp.client-ip=103.168.172.149 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="HZzCPkDE"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="Sa/gT1TL" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfout.phl.internal (Postfix) with ESMTP id BF4A1EC0222; Fri, 14 Aug 2026 21:59:32 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-06.internal (MEProxy); Fri, 14 Aug 2026 21:59:32 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786759172; x= 1786845572; bh=xu5XlzhsRnyO513iPV0tSkUMOqSCDM46ogGqM3hnkno=; b=H ZzCPkDELIGrsL/SPZkH13po6ok3Xzu82j3gEVLFd4py9K0hQqhsH6tbb3v4s9uUt USKjCa0InwtttB36LBy71X7RGJ/6iWdwahgP8rtEt2KvG9XCXfxbuoUdeNaesCsv 65WytTlI1SzNc0LC9j50V6LTwMAIaekKK4Y4ymiN4WDGI2Uk7ZfGIwQj2DdGvdqR zzz6HE5GFV5ag1fk5kwW5t0OEMgKsT9VdyvZIej4/B/tICPoG2K4cdzEv8g8uQh+ HiNYcTF112LeOCpqidFJL0wX7nshUl2mV/GHtKheDjTVMTHrLsUurx++zV9tGEnh X4i2nB6/7ZBMtpOM1I1+w== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786759172; x=1786845572; bh=x u5XlzhsRnyO513iPV0tSkUMOqSCDM46ogGqM3hnkno=; b=Sa/gT1TLJH2prrjz6 jViLTzOr++s0aYqElJ4N7NepM5yMRkSpi6XX9qQCKXxitcOd9q07RjiGeynTXG9u QgxGWcRAlpvROLZgl3QtuE8dHdP1VWcPxlS5KNWn8qOI4EZsrqNa7VHiNvsEfgQ4 stLgu8jrmAhfBiFklZKfhig3FzZfvFI92r5fHYilkDghoO2SISPqxXQhYZ+kL3Md gF16sd1fEYug5UskwBAmbtEBhIqlRL7RMCAt5I/1Tt2SPVoEj1nAzN3FIVN+evOq 7iDCTcNoGcqfwzV+UwqvZC6DT/UgltfbfISeXM8+NoOTHwEZFQ4dhyEW6CW4tm+H XtOuA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTE4LsllL5mDr4jgINBzNWNyncOxykfAHbUnd5YNeufniPijDKZ2Q74nct+xKCifs1 H2Djkg7H/6yFeHnJG3ZaS3y0lvG0agappzf948+VnP0i9yBGlC75HydJdxE+VaLcaAo1iB npgWdh437yabl0tcqFtMCkfLO+e8hKTIM1JoCtbRkD60i2JQs08ndCsfMPOXT8PgyeyXeh uq46vXE/qDxxY4OZQldbm4iR2S+h71iOzjsUxq79Ie/EJboDCHFVxBfBy1NevQy7IuvBPJ BK4YMBoSmlYE8wkBs5S3mRyvR8BaiOuvXRijp0iKeXpkOL2deb0peaX6b452jxbs1L3IHc r6EiA8aJhaBPU/Y029rs+sLysD1ij4of89NgQURm/ilFRab3dDIN42+Xtstkr7g5Gok8lW yOHfx0QBtVuzJJ/rRXeXzo4Q634fT0KPqQ08vCkS19rKppm+Sa7KTKZKI9TnlH7V6ZIPiv 1MYDx+dy4fSKd4WLXV0SFEtxM2wqzZ+UDXLR3CW8cb3ZiTkkqwZag+Z8850Xgw8JMQTvVE FDWj7+OXg5E26baiBESDB5UmHEERvYLbIdWytfZa1VFSzsinRjay+Bh7TTtu1nK0hnAeki /ebZk994hv1njGYI3bVVndIMsVEYk4Ys3MmNYuCuUtOfEnzrxswPkrSqE/NA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 14 Aug 2026 21:59:32 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v4 14/19] selftests/mm: run every supported collapse order by default Date: Sat, 15 Aug 2026 02:58:56 +0100 Message-ID: <20260815015901.1236937-15-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260815015901.1236937-1-kirill@shutemov.name> References: <20260815015901.1236937-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The mTHP collapse cases only run when the caller names both the context and an order, so a plain ./khugepaged covers the PMD contexts on anon and nothing else. run_vmtests.sh pinned order 4 and covered no other. Run the mTHP cases once per supported anon THP order below the PMD when -c is absent, and pull that context into both the no-argument invocation and "all". Around that: - -c still pins one order, and now says what is wrong instead of printing the usage text. An order at or below the -s source order is skipped: the sources would already be the size being asked for. - Both orders end up as array indices and shift counts, so -s and -c are range-checked before they get there. - The mTHP context has only anon cases, so a run that names a different mem_type -- "all:shmem", say -- drops it again rather than refusing to start. Naming both explicitly still refuses. - A case carries the order it was registered at, so a result names it: # Run test: collapse_single_mthp (mthp_khugepaged:anon, order 6) On x86-64 with 4K pages that is orders 2 through 8, and ./khugepaged goes from 28 results to 77 in 21 seconds, so run_vmtests.sh can drop its pinned order-4 line. Assisted-by: Claude-Code:claude-opus-5 Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/khugepaged.c | 105 +++++++++++++++++----- tools/testing/selftests/mm/run_vmtests.sh | 2 - 2 files changed, 85 insertions(+), 22 deletions(-) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index 856decd2950a..172e7307eeee 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -31,6 +31,9 @@ static unsigned long page_size; static int hpage_pmd_nr; static int anon_order; static int collapse_order; +static bool collapse_order_given; +static int collapse_orders[NR_ORDERS]; +static int nr_collapse_orders; static int pagemap_fd =3D -1; static int kpageflags_fd =3D -1; =20 @@ -1550,6 +1553,7 @@ static void usage(void) fprintf(stderr, "\t\t-s: mTHP size, expressed as page order.\n"); fprintf(stderr, "\t\t Defaults to 0. Use this size for anon or shmem a= llocations.\n"); fprintf(stderr, "\t\t-c: collapse order for mTHP collapse, expressed as p= age order.\n"); + fprintf(stderr, "\t\t Defaults to every supported order below the PMD.= \n"); fprintf(stderr, "\t\t With -s, -s names the mTHP source order for the\= n"); fprintf(stderr, "\t\t mixed-source case (source order below the target= ).\n"); exit(1); @@ -1557,6 +1561,7 @@ static void usage(void) =20 static void parse_test_type(int argc, char **argv) { + bool mthp_context_implied =3D false; int opt; char *buf; const char *token; @@ -1568,6 +1573,7 @@ static void parse_test_type(int argc, char **argv) break; case 'c': collapse_order =3D atoi(optarg); + collapse_order_given =3D true; break; case 'h': default: @@ -1575,12 +1581,25 @@ static void parse_test_type(int argc, char **argv) } } =20 + /* + * Both orders end up as array indices and shift counts, so neither + * can be negative, and a zero collapse order asks for base pages. + */ + if (anon_order < 0 || anon_order > hpage_pmd_order) + ksft_exit_fail_msg("-s takes an order in 0..%d, not %d\n", + hpage_pmd_order, anon_order); + if (collapse_order_given && + (collapse_order <=3D 0 || collapse_order >=3D hpage_pmd_order)) + ksft_exit_fail_msg("-c takes an order in 1..%d, not %d\n", + hpage_pmd_order - 1, collapse_order); + argv +=3D optind; argc -=3D optind; =20 if (argc =3D=3D 0) { - /* Backwards compatibility */ + /* Everything that needs no argument of its own: anon, every context */ khugepaged_context =3D &__khugepaged_context; + mthp_khugepaged_context =3D &__mthp_khugepaged_context; madvise_context =3D &__madvise_context; anon_ops =3D &__anon_ops; return; @@ -1591,13 +1610,19 @@ static void parse_test_type(int argc, char **argv) =20 if (!strcmp(token, "all")) { khugepaged_context =3D &__khugepaged_context; + mthp_khugepaged_context =3D &__mthp_khugepaged_context; madvise_context =3D &__madvise_context; + + /* + * "all" sweeps the mTHP context in, but it only has anon + * cases: step it aside for the other mem_types rather than + * refusing the whole run. + */ + mthp_context_implied =3D true; } else if (!strcmp(token, "khugepaged")) { khugepaged_context =3D &__khugepaged_context; } else if (!strcmp(token, "mthp_khugepaged")) { mthp_khugepaged_context =3D &__mthp_khugepaged_context; - if (collapse_order <=3D 0 || collapse_order >=3D hpage_pmd_order) - usage(); } else if (!strcmp(token, "madvise")) { madvise_context =3D &__madvise_context; } else { @@ -1613,20 +1638,20 @@ static void parse_test_type(int argc, char **argv) read_write_file_write_ops =3D &__read_write_file_write_ops; anon_ops =3D &__anon_ops; shmem_ops =3D &__shmem_ops; - if (mthp_khugepaged_context) - usage(); } else if (!strcmp(buf, "anon")) { anon_ops =3D &__anon_ops; } else if (!strcmp(buf, "file")) { read_only_file_ops =3D &__read_only_file_ops; read_write_file_read_ops =3D &__read_write_file_read_ops; read_write_file_write_ops =3D &__read_write_file_write_ops; - if (mthp_khugepaged_context) + if (mthp_khugepaged_context && !mthp_context_implied) usage(); + mthp_khugepaged_context =3D NULL; } else if (!strcmp(buf, "shmem")) { shmem_ops =3D &__shmem_ops; - if (mthp_khugepaged_context) + if (mthp_khugepaged_context && !mthp_context_implied) usage(); + mthp_khugepaged_context =3D NULL; } else { usage(); } @@ -1648,6 +1673,7 @@ struct test_case { struct mem_ops *ops; const char *desc; test_fn fn; + int order; /* mTHP contexts: the collapse order */ }; =20 #define MAX_TEST_CASES 256 @@ -1663,6 +1689,7 @@ static int nr_test_cases; .ops =3D o, \ .desc =3D #t, \ .fn =3D t, \ + .order =3D collapse_order, \ }; \ } \ } while (0) @@ -1703,10 +1730,37 @@ int main(int argc, char **argv) =20 parse_test_type(argc, argv); =20 - if (mthp_khugepaged_context && - !(thp_supported_orders() & (1UL << collapse_order))) - ksft_exit_skip("Order %d is not a supported anon THP order\n", - collapse_order); + if (mthp_khugepaged_context) { + unsigned long orders =3D thp_supported_orders(); + + if (collapse_order_given) { + /* -c pins one order; it has to be one we can build */ + if (!(orders & (1UL << collapse_order))) + ksft_exit_skip("Order %d is not a supported anon THP order\n", + collapse_order); + if (collapse_order <=3D anon_order) + ksft_exit_skip("-c %d needs a source order below it, -s says %d\n", + collapse_order, anon_order); + collapse_orders[nr_collapse_orders++] =3D collapse_order; + } else { + /* + * Otherwise every order a collapse could produce. -s + * makes the fault path hand out folios of that order, + * so a target at or below it has nothing to collapse: + * the sources are already the size being asked for. + */ + int first =3D anon_order + 1; + + if (first < MIN_MTHP_ORDER) + first =3D MIN_MTHP_ORDER; + for (int i =3D first; i < hpage_pmd_order; i++) { + if (orders & (1UL << i)) + collapse_orders[nr_collapse_orders++] =3D i; + } + if (!nr_collapse_orders) + ksft_print_msg("mTHP cases skipped: no order above the source\n"); + } + } =20 if (mthp_khugepaged_context) { pagemap_fd =3D open("/proc/self/pagemap", O_RDONLY); @@ -1760,7 +1814,17 @@ int main(int argc, char **argv) TEST(collapse_full, khugepaged_context, read_write_file_read_ops); TEST(collapse_full, khugepaged_context, read_write_file_write_ops); TEST(collapse_full, khugepaged_context, shmem_ops); - TEST(collapse_full, mthp_khugepaged_context, anon_ops); + for (int i =3D 0; i < nr_collapse_orders; i++) { + collapse_order =3D collapse_orders[i]; + TEST(collapse_full, mthp_khugepaged_context, anon_ops); + TEST(collapse_empty, mthp_khugepaged_context, anon_ops); + TEST(collapse_single_mthp, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_single_window, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_partial_window, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_max_ptes_none, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_mixed_sources, mthp_khugepaged_context, anon_ops); + } + TEST(collapse_full, madvise_context, anon_ops); TEST(collapse_full, madvise_context, read_only_file_ops); TEST(collapse_full, madvise_context, read_write_file_read_ops); @@ -1768,15 +1832,8 @@ int main(int argc, char **argv) TEST(collapse_full, madvise_context, shmem_ops); =20 TEST(collapse_empty, khugepaged_context, anon_ops); - TEST(collapse_empty, mthp_khugepaged_context, anon_ops); TEST(collapse_empty, madvise_context, anon_ops); =20 - TEST(collapse_single_mthp, mthp_khugepaged_context, anon_ops); - TEST(collapse_order_single_window, mthp_khugepaged_context, anon_ops); - TEST(collapse_order_partial_window, mthp_khugepaged_context, anon_ops); - TEST(collapse_order_max_ptes_none, mthp_khugepaged_context, anon_ops); - TEST(collapse_order_mixed_sources, mthp_khugepaged_context, anon_ops); - TEST(collapse_single_pte_entry, khugepaged_context, anon_ops); TEST(collapse_single_pte_entry, khugepaged_context, read_only_file_ops); TEST(collapse_single_pte_entry, khugepaged_context, read_write_file_read_= ops); @@ -1849,7 +1906,15 @@ int main(int argc, char **argv) for (int i =3D 0; i < nr_test_cases; i++) { struct test_case *t =3D &test_cases[i]; =20 - ksft_print_msg("\n# Run test: %s (%s:%s)\n", t->desc, t->ctx->name, t->o= ps->name); + if (t->ctx =3D=3D &__mthp_khugepaged_context) { + collapse_order =3D t->order; + ksft_print_msg("\n# Run test: %s (%s:%s, order %d)\n", + t->desc, t->ctx->name, t->ops->name, + t->order); + } else { + ksft_print_msg("\n# Run test: %s (%s:%s)\n", t->desc, + t->ctx->name, t->ops->name); + } t->fn(t->ctx, t->ops); } =20 diff --git a/tools/testing/selftests/mm/run_vmtests.sh b/tools/testing/self= tests/mm/run_vmtests.sh index 2652a7920b80..8bf898b71350 100755 --- a/tools/testing/selftests/mm/run_vmtests.sh +++ b/tools/testing/selftests/mm/run_vmtests.sh @@ -412,8 +412,6 @@ CATEGORY=3D"thp" run_test ./khugepaged all:shmem =20 CATEGORY=3D"thp" run_test ./khugepaged -s 4 all:shmem =20 -CATEGORY=3D"thp" run_test ./khugepaged -c 4 mthp_khugepaged:anon - # Try to create XFS if not provided if [ -z "${SPLIT_HUGE_PAGE_TEST_XFS_PATH}" ]; then if test_selected "thp"; then --=20 2.54.0 From nobody Mon Sep 28 23:56:19 2026 Received: from fout-a6-smtp.messagingengine.com (fout-a6-smtp.messagingengine.com [103.168.172.149]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1613B348C76; Sat, 15 Aug 2026 01:59:35 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.149 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759177; cv=none; b=mKQKifHI8bu2WjQ40mfodY3ivKWBDXLnn+AYI4bjun0TKm12JZQdNCsmyzUY+91VdSXp+EE7UJ1Ju5+lnFhzLbH1xMnOdc/75D8qhgqPTTFLAhmKvzAW1In6OJ0hPv2fR/Xuo9v7EHpB6IRf4u2t9xl5A3uCVeGFnAiEm8Kbk/g= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759177; c=relaxed/simple; bh=yyq3kmZ3rgY5a6X2MJJXYtvMV0TvBXhbBzmQioozSkk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=nlkuJ26CCSXTSknrNo9R8fr8BNRikMqpW17VLfCIlgOZmLif4+e1a1EZ+QlDzV3hwksnlM3jHZdckFz2WvQfx4Vq3YjrilpC/dPVva+I/Cm69DKRO0yjw3lv6PRRePDZZhPaNzij9KGk3ctw5F3dBo2btHFtKV8tTVB/sUWlowY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=wUEtkP4L; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=AyEkDJyG; arc=none smtp.client-ip=103.168.172.149 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="wUEtkP4L"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="AyEkDJyG" Received: from phl-compute-04.internal (phl-compute-04.internal [10.202.2.44]) by mailfout.phl.internal (Postfix) with ESMTP id 6EAF3EC0223; Fri, 14 Aug 2026 21:59:34 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-04.internal (MEProxy); Fri, 14 Aug 2026 21:59:34 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786759174; x= 1786845574; bh=OE+z5Z9t7XAE9Lh7e4xto9pNkr14mTVu0K4BOrQwVeA=; b=w UEtkP4LaohjkY96dUrM2/Gg8v1SXvE4oyJieVcO/3ihCC/2FVBAXlb9IK9Bv192m 8uzuj/jYsJxulejGiLNr0DoMsCs8KGQMBRB8g40buy4hLq6NEmRIzw1SFmniHJrS bOXaXgn0hEvlZ84bqBZhAjOEh4kCHn7qcAABhBjLT4kW2hWJ06N26RBpsd/o/PW/ t69YZ+E3WYCgItp+vu6oer/AnnZTMH+LTkFXgc7r/uMV1dX4Ubt9QHkDwys3Nfs4 P0bMi8/u64OEMAStvvpzXk0LGbuR+Sy43ntYzKxUjFPRPa5hJ74XUlUxmBXZIwM6 O+/f3oet+Y6QB+2UzfAWg== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786759174; x=1786845574; bh=O E+z5Z9t7XAE9Lh7e4xto9pNkr14mTVu0K4BOrQwVeA=; b=AyEkDJyGf1lb0vT8i u26VGfBvF2PeRMRwqY7FKd0mcxNsZVGFP11AIdffinftozZWOHOJE4F8iIV14zqv iUvGyPjD0EJiPsEtrvP2NzINUIaBIfLZB+HmP1VD/Py3hGOhFJCr5lPO7SByagyW mlR7bib4aJX5QibPO3vK0/24Xsek/zsX7BU/iQCvbcVHTN7K4nCT6QxtNUqxGk2p 9KxTM8gfBVLHVYPEKzf9OHJawb8fSzcq+KbiIoLKO/mpT4ursaEOPgENXDebj40e MQ0C7QLRkKfHMnafK3TjFoJqbp0TKzMgWp2UuT7BecKeujq+1HsSBjho3xmaIiDc K1x9A== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEHOB8klcfpaewZ3k5dhFGq+0NT52m0JgfrBhlsrZZ6/wIFkiWFy9BKVmBRzT/M60 afg3IBZlSt5ygR3bwXNIcW5ZM0yyzN/pf9UTr2aIVxu9PRgCvu2oUUNVujSBw7FeHwFBsc cXVjrQP9UQmJU5OZxw2U7sN+BvQMOJwzXPSq8ovjEuIpT5m3q0hBt7II5b5sCGi7gkG3/G fAkJjKG33ZxocRSbFLEDKzhP8eL3QUGAVXaSl2VxURMVZW4QVvv1Ki5M31XaRzn4+ok0bg hGdz4SdlVbJh+GxMS7w8oKVQIiajy/o8Z9eSxaHEz+PVJrdtO9FMEbFO2si+6c9liGB/Ff wTXUAXs77/qsjTCBHYxKGokLOK18ZP4RunQBJXnW9+ILE0Q8RXfNQjiO2b7MS7IStm3xGm 9ZTMP3wgI0RcearR/x3pfcMGh1sX409I5pLYM4jEiOqo50fkI/rrfVeQsURZyXI4FPNhfv 6HlOxn/FFx+/3Onfi4EM1QN2/WuBLFP908hXPCTSbsDyuE9vtDgUkEJaOMOcu7tO/57tBx 5VjfcRA6iqSAiym7JeLBedzNgZgvBks8ft3zse8ViQIoSeY82Brpm0EMkCQgHaAfpr+Kga ar1+NqX/iZktB0Uc07Ogq4Na5lrgEbypKolD52EtT0mVHSZl2kffM2vXdC6g X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 14 Aug 2026 21:59:33 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v4 15/19] selftests/mm: check that one khugepaged pass collapses one window Date: Sat, 15 Aug 2026 02:58:57 +0100 Message-ID: <20260815015901.1236937-16-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260815015901.1236937-1-kirill@shutemov.name> References: <20260815015901.1236937-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The khugepaged cases drive the daemon through sysfs: a store to scan_sleep_millisecs wakes it, and full_scans advancing by two marks one pass that started after setup. Every khugepaged result in the suite rests on that pair, and nothing checks it. Add khugepaged_sync_check. Each step: - prepare one aligned window - record its source PFNs from pagemap - run one khugepaged_full_pass() barrier - require the window came out collapsed, with exactly one collapse attempt attributed to it The anon events carry no virtual address, so an attempt is matched by the source folio PFN and order that mm_collapse_huge_page_isolate() reports. Reading the trace buffer takes four small helpers in vm_util: open an event subsystem's enable file, flip it, clear the buffer, and open it for reading. scan_sleep_millisecs is set to a minute, so a step that took a sleep instead of a wake would blow the budget. Passes 5/5 on x86-64 4K and arm64 64K. Assisted-by: Claude-Code:claude-opus-5 Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/Makefile | 1 + .../selftests/mm/khugepaged_sync_check.c | 196 ++++++++++++++++++ tools/testing/selftests/mm/run_vmtests.sh | 2 + tools/testing/selftests/mm/vm_util.c | 41 ++++ tools/testing/selftests/mm/vm_util.h | 4 + 5 files changed, 244 insertions(+) create mode 100644 tools/testing/selftests/mm/khugepaged_sync_check.c diff --git a/tools/testing/selftests/mm/Makefile b/tools/testing/selftests/= mm/Makefile index 2093fcf6e915..b2d6e5c12934 100644 --- a/tools/testing/selftests/mm/Makefile +++ b/tools/testing/selftests/mm/Makefile @@ -105,6 +105,7 @@ TEST_GEN_FILES +=3D merge TEST_GEN_FILES +=3D rmap TEST_GEN_FILES +=3D folio_split_race_test TEST_GEN_FILES +=3D folio_order_check +TEST_GEN_FILES +=3D khugepaged_sync_check =20 ifneq ($(ARCH),arm64) TEST_GEN_FILES +=3D soft-dirty diff --git a/tools/testing/selftests/mm/khugepaged_sync_check.c b/tools/tes= ting/selftests/mm/khugepaged_sync_check.c new file mode 100644 index 000000000000..45001996b57a --- /dev/null +++ b/tools/testing/selftests/mm/khugepaged_sync_check.c @@ -0,0 +1,196 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Synchronous khugepaged driving check. + * + * Race tests drive khugepaged through the existing sysfs controls: a + * store to scan_sleep_millisecs wakes the daemon, and full_scans + * advancing by two is a completion barrier for one full pass that + * started after setup (khugepaged_full_pass()). Verify the pair gives + * deterministic, attributable results: one barrier step over one + * prepared window produces exactly one collapse attempt on that + * window's source pages (mm_collapse_huge_page_isolate events filtered + * by source PFN and order) and the window is collapsed + * afterwards, repeatably. + * + * scan_sleep_millisecs is set to 60s to prove the wake path: without + * the wake, one barrier step would sleep multiples of that and blow + * the timeout. It also keeps the daemon from free-running between + * steps, per the khugepaged_full_pass() discipline. + */ +#define _GNU_SOURCE +#include +#include +#include +#include +#include +#include + +#include "kselftest.h" +#include "vm_util.h" +#include "hugepage_settings.h" + +#define BASE_ADDR ((void *)(1UL << 30)) +#define TARGET_ORDER 2 /* smallest order khugepaged considers */ +#define NR_ITERATIONS 5 + +static int pagemap_fd; +static int kpageflags_fd; +static int trace_events_fd =3D -1; +static unsigned long hpage_pmd_size; + +/* + * Each step switches the events off again, but a helper can still give up + * on us in between (a failing sysfs write ends the test from inside + * thp_write_num()), and huge_memory events left on are the whole machine's + * problem, not this test's. + */ +static void trace_events_off(void) +{ + if (trace_events_fd >=3D 0) + tracing_events_enable(trace_events_fd, false); +} + +/* + * Count collapse attempts attributable to our window: isolate events whose + * scan_pfn is one of the window's source PFNs, reported once per attempt. + */ +static int count_attributed(unsigned long *pfns, int nr_pfns, + unsigned int order) +{ + char line[1024]; + int count =3D 0; + FILE *fp; + + fp =3D tracing_open_trace(); + if (!fp) + ksft_exit_fail_msg("Cannot open trace buffer\n"); + + while (fgets(line, sizeof(line), fp)) { + unsigned long val; + unsigned int ord; + char *s, *o; + int i; + + s =3D strstr(line, "mm_collapse_huge_page_isolate:"); + if (!s) + continue; + if (sscanf(s, "mm_collapse_huge_page_isolate: scan_pfn=3D0x%lx", + &val) !=3D 1) + continue; + o =3D strstr(s, "order=3D"); + if (!o || sscanf(o, "order=3D%u", &ord) !=3D 1 || ord !=3D order) + continue; + for (i =3D 0; i < nr_pfns; i++) { + if (val =3D=3D pfns[i]) { + count++; + break; + } + } + } + fclose(fp); + return count; +} + +static void one_step(int iteration) +{ + const size_t window =3D getpagesize() << TARGET_ORDER; + const int nr_pages =3D 1 << TARGET_ORDER; + unsigned long pfns[1 << TARGET_ORDER]; + bool collapsed, passed; + int attributed; + char *p; + int i; + + p =3D mmap(BASE_ADDR, hpage_pmd_size, PROT_READ | PROT_WRITE, + MAP_ANONYMOUS | MAP_PRIVATE | MAP_FIXED_NOREPLACE, -1, 0); + if (p !=3D BASE_ADDR) + ksft_exit_fail_perror("mmap() window"); + + /* Prepare one window; record its source PFNs. */ + for (i =3D 0; i < nr_pages; i++) { + p[i * getpagesize()] =3D i + 1; + pfns[i] =3D pagemap_get_pfn(pagemap_fd, p + i * getpagesize()); + if (pfns[i] =3D=3D -1UL) + ksft_exit_fail_msg("Source page not present\n"); + } + + /* Clear first: with the events still off there is nothing to undo. */ + if (tracing_clear_trace()) + ksft_exit_fail_msg("Cannot clear the trace buffer\n"); + if (tracing_events_enable(trace_events_fd, true)) + ksft_exit_fail_msg("Cannot enable huge_memory events\n"); + + if (madvise(p, hpage_pmd_size, MADV_HUGEPAGE)) + ksft_exit_fail_perror("madvise(MADV_HUGEPAGE)"); + /* Wait up to 120 seconds for the pass to complete. */ + passed =3D khugepaged_full_pass(120); + + /* Off before anything that can give up: the events are system-wide. */ + if (tracing_events_enable(trace_events_fd, false)) + ksft_exit_fail_msg("Cannot disable huge_memory events\n"); + if (!passed) + ksft_exit_fail_msg("khugepaged did not complete a full pass\n"); + + collapsed =3D is_range_backed_by_folio_orders(p, window, TARGET_ORDER, + pagemap_fd, kpageflags_fd); + attributed =3D count_attributed(pfns, nr_pages, TARGET_ORDER); + + ksft_test_result(collapsed && attributed =3D=3D 1, + "step %d: window collapsed, %d attributed result(s)\n", + iteration, attributed); + + munmap(p, hpage_pmd_size); +} + +int main(void) +{ + struct thp_settings settings; + int i; + + ksft_print_header(); + + if (!thp_available()) + ksft_exit_skip("Transparent Hugepages not available\n"); + if (!(thp_supported_orders() & (1UL << TARGET_ORDER))) + ksft_exit_skip("Order %d is not a supported anon THP order\n", + TARGET_ORDER); + + hpage_pmd_size =3D read_pmd_pagesize(); + if (!hpage_pmd_size) + ksft_exit_fail_msg("Reading PMD pagesize failed\n"); + pagemap_fd =3D open("/proc/self/pagemap", O_RDONLY); + if (pagemap_fd < 0) + ksft_exit_fail_perror("open(/proc/self/pagemap)"); + kpageflags_fd =3D open("/proc/kpageflags", O_RDONLY); + if (kpageflags_fd < 0) + ksft_exit_skip("open(\"/proc/kpageflags\") requires root\n"); + trace_events_fd =3D tracing_events_open("huge_memory"); + if (trace_events_fd < 0) + ksft_exit_skip("huge_memory events require tracefs and root\n"); + atexit(trace_events_off); + + ksft_set_plan(NR_ITERATIONS); + + thp_save_settings(); + thp_read_settings(&settings); + settings.thp_enabled =3D THP_MADVISE; + settings.thp_defrag =3D THP_DEFRAG_ALWAYS; + settings.khugepaged.defrag =3D 1; + settings.khugepaged.scan_sleep_millisecs =3D 60000; + settings.khugepaged.alloc_sleep_millisecs =3D 60000; + settings.khugepaged.max_ptes_none =3D (hpage_pmd_size / getpagesize()) - = 1; + /* One wake must complete one full pass; see khugepaged_full_pass(). */ + settings.khugepaged.pages_to_scan =3D 1UL << 24; + for (i =3D 0; i < NR_ORDERS; i++) + settings.hugepages[i].enabled =3D THP_NEVER; + settings.hugepages[TARGET_ORDER].enabled =3D THP_INHERIT; + /* Base of the settings stack; the bottom entry is never popped. */ + thp_push_settings(&settings); + + for (i =3D 0; i < NR_ITERATIONS; i++) + one_step(i); + + thp_restore_settings(); + + ksft_finished(); +} diff --git a/tools/testing/selftests/mm/run_vmtests.sh b/tools/testing/self= tests/mm/run_vmtests.sh index 8bf898b71350..c0f69da3fd3b 100755 --- a/tools/testing/selftests/mm/run_vmtests.sh +++ b/tools/testing/selftests/mm/run_vmtests.sh @@ -404,6 +404,8 @@ CATEGORY=3D"cow" run_test ./cow =20 CATEGORY=3D"thp" run_test ./folio_order_check =20 +CATEGORY=3D"thp" run_test ./khugepaged_sync_check + CATEGORY=3D"thp" run_test ./khugepaged =20 CATEGORY=3D"thp" run_test ./khugepaged -s 2 diff --git a/tools/testing/selftests/mm/vm_util.c b/tools/testing/selftests= /mm/vm_util.c index c9bd6c92fa41..ee1334778391 100644 --- a/tools/testing/selftests/mm/vm_util.c +++ b/tools/testing/selftests/mm/vm_util.c @@ -598,6 +598,47 @@ bool is_range_backed_by_folio_orders(char *start, size= _t len, int order, return true; } =20 +#define TRACEFS_ROOT "/sys/kernel/tracing" + +/* + * Open the enable file of one ftrace event subsystem (e.g. "huge_memory"). + * Returns a descriptor for tracing_events_enable(), or -1 if tracefs or t= he + * subsystem is not there. The events are system-wide state: whoever + * switches them on owns them until it switches them off, including on the + * paths where the test gives up. + */ +int tracing_events_open(const char *subsys) +{ + char path[256]; + + snprintf(path, sizeof(path), TRACEFS_ROOT "/events/%s/enable", + subsys); + return open(path, O_WRONLY); +} + +int tracing_events_enable(int fd, bool enable) +{ + if (pwrite(fd, enable ? "1" : "0", 1, 0) !=3D 1) + return -1; + return 0; +} + +/* Drop what the trace buffer holds so far. */ +int tracing_clear_trace(void) +{ + int fd =3D open(TRACEFS_ROOT "/trace", O_WRONLY | O_TRUNC); + + if (fd < 0) + return -1; + close(fd); + return 0; +} + +FILE *tracing_open_trace(void) +{ + return fopen(TRACEFS_ROOT "/trace", "r"); +} + /* If `ioctls' non-NULL, the allowed ioctls will be returned into the var = */ int uffd_register_with_ioctls(int uffd, void *addr, uint64_t len, bool miss, bool wp, bool minor, uint64_t *ioctls) diff --git a/tools/testing/selftests/mm/vm_util.h b/tools/testing/selftests= /mm/vm_util.h index ce05bce4670d..10c7be46e44c 100644 --- a/tools/testing/selftests/mm/vm_util.h +++ b/tools/testing/selftests/mm/vm_util.h @@ -119,6 +119,10 @@ int close_procmap(struct procmap_fd *procmap); int write_sysfs(const char *file_path, unsigned long val); int read_sysfs(const char *file_path, unsigned long *val); bool softdirty_supported(void); +int tracing_events_open(const char *subsys); +int tracing_events_enable(int fd, bool enable); +int tracing_clear_trace(void); +FILE *tracing_open_trace(void); =20 static inline int open_self_procmap(struct procmap_fd *procmap_out) { --=20 2.54.0 From nobody Mon Sep 28 23:56:19 2026 Received: from fout-a6-smtp.messagingengine.com (fout-a6-smtp.messagingengine.com [103.168.172.149]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1A46C3644C7; Sat, 15 Aug 2026 01:59:36 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.149 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759179; cv=none; b=CwOmT6zTkYiNkBPSnTcHmqy8Ai5h2dOj2L8SYhLIhSqUc7bQu7mmNlPRizaKLMt2jCTefoG9VIzIw/3UZZrDqVhZ5aaBTbcCs+cwpNG+cy/swU1nKWw325Qkzy+ZOElnUt946BbniH8Vyo1RqkMlMQ9bvXG7QtfTdXGdFwEp1FY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759179; c=relaxed/simple; bh=+TpIFElUgF05+/H2zBLqsxLs4H2qFtRl4ykSwFDTMh0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=rX6/uRFuMwBLk/xC+te63NaFXvcz1d43aShqA8xDCjhqhuw1CqIw7ub39jphus5Q8UKQN94MCHsek7XwYesfBWM+TBHjy9F/pftGgTn9sVzVwoVFmljft9SW/KELAGEizMXsE3MkJqGk86b8DtRbpIX4r2wmWzHFWUKD55Q8+ng= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=BJysKvRs; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=D+SXhRAt; arc=none smtp.client-ip=103.168.172.149 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="BJysKvRs"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="D+SXhRAt" Received: from phl-compute-05.internal (phl-compute-05.internal [10.202.2.45]) by mailfout.phl.internal (Postfix) with ESMTP id 4004CEC0222; Fri, 14 Aug 2026 21:59:36 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-05.internal (MEProxy); Fri, 14 Aug 2026 21:59:36 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786759176; x= 1786845576; bh=AO/oUzqhOXSGRBCY8YhrbuAzNWboTvwfnRq6ibT3Z0I=; b=B JysKvRs0w0dgewcptHTLipsM9Igk808LxEzX3Mw+bdC4IaRgdUCKqn4jxy3LehOS jQgwh/N3prJRrIi3DnddulryRX6pwlFMK055muuDUkHZa6yTRVeClP31cqb0cQO6 BvvkaS2JveX/sVlkFswC+TpglmiAdLrmtc6wE1CDtk2xdCOkbc2xSlzDEuweHPek miTAO+ESM98O2e2bcWPxUvXaiXmzvilm9QdAhI++CHbeSY3LSo4lhwCCPBxJhumn hg8Fwxv2JQD77j/n+f1IUtO1s1lDPAoGcNQwqhHCHtJ794bPv1tbBe78oWmUdcHD FDYKG59fQ9yCmxYM6EZdA== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786759176; x=1786845576; bh=A O/oUzqhOXSGRBCY8YhrbuAzNWboTvwfnRq6ibT3Z0I=; b=D+SXhRAto+UkByM6P JkXZgS2uLdbYaQysPTEeGB8CfgKUiAVzol0jcI6Xffwd8l2LMXK3n5DvbhKvVype VQlr9lgbhAxknZJMpJwJhuz3Dv32kpGEJOkF9FSRx4M+abcpDvqW79t4uz0RyyaM O/NhExdboVOmRZNS/ZoModXqwrFfpdUGwezS1EUbld3/39otTfvqm9Y1fUdbNm3/ VZg+ZoN49Nz0krxkc2juYW02gj9PBV0RTEdxx5Dy/Mbn4pFxlZg7fkDWNUqXPo6h XDpJ1AxyuS7pCCFDoL25RILDfm5xpVuOD3AXjMC1EB0iprbpZ4mVH00JkUMeJ437 6Eh1g== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEHOB8klcfpaewZ3k5dhFGq+0NT52m0JgfrBhlsrZZ6/wIFkiWFy9BKVmBRzT/M60 afg3IBZlSt5ygR3bwXNIcW5ZM0yyzN/pf9UTr2aIVxu9PRgCvu2oUUNVujSBw7FeHwFBsc cXVjrQP9UQmJU5OZxw2U7sN+BvQMOJwzXPSq8ovjEuIpT5m3q0hBt7II5b5sCGi7gkG3/G fAkJjKG33ZxocRSbFLEDKzhP8eL3QUGAVXaSl2VxURMVZW4QVvv1Ki5M31XaRzn4+ok0bg hGdz4SdlVbJh+GxMS7w8oKVQIiajy/o8Z9eSxaHEz+PVJrdtO9FMEbFO2si+6c9liGB/kA mdDnw2Qcq4Dz9ph/PSDADa+0M0gnI/qIYttfy8mfVpS7Bx3F6HWYF1THwbGAI8uraMQmXt lznRPNv+kN2bXaiuCMZFG9ezL4A9dHjVOOnrV36kp2B5apv2ECy2pHuk2iOUz8wofrKSVk cUW4VPaOcgzg7yWwV1Evn72D0dr1cklBFc29DIDa3N4SVHaqCy9a3MkPgotfWam/WmT7Vc SgnpF5IsNI+007S29LNfYEToPBZBYAeAcJPGQmQyEmXSn16ENIDsrTM97ivmRGY9q3omLM Gr/T6NjAIA7QGNsgeAuJVh22TzxTOiRHSUh4hY1w9gyJ7ZYfrmhmKzbcPvbQ X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 14 Aug 2026 21:59:35 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v4 16/19] selftests/mm: add khugepaged race harness Date: Sat, 15 Aug 2026 02:58:58 +0100 Message-ID: <20260815015901.1236937-17-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260815015901.1236937-1-kirill@shutemov.name> References: <20260815015901.1236937-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Collapse serialises against faults, GUP, fork, mremap and zapping through a protocol of locks, TLB flushes and refcount checks. No khugepaged selftest exercises any of it under contention. Add khugepaged_race. Six racing threads work the same address space: - two faulters - an MADV_DONTNEED thread - a transient FOLL_PIN thread (gup_test) - a forker - an mremap thread One of three drivers collapses under them: stepped khugepaged, one full pass at a time via khugepaged_full_pass(), so each step covers a known extent; free khugepaged left to run (scan_sleep_millisecs=3D0), for soak; madvise an MADV_COLLAPSE and MADV_DONTNEED loop. Every mode runs in turn unless -m names one, five seconds each. Every supported anon THP order is set to inherit and max_ptes_none is 0, so a window collapses only once fully populated and the racing MADV_DONTNEED steers selection across orders. The rule is that a racing page reads as its pattern or as zero, never anything else. The faulters and fork children check it throughout, and a final sweep checks it again. The other half of the check is the kernel's own assertions -- DEBUG_VM, page_table_check, KASAN, lockdep -- so read dmesg too. The pin thread goes through gup_test, so the harness skips without CONFIG_GUP_TEST or root. The threads share three PMD-sized areas plus one for the mremap thread; -a sets the count, since at a 512M PMD that is over two gigabytes. -d sets how long each mode runs. Assisted-by: Claude-Code:claude-opus-5 Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/Makefile | 1 + tools/testing/selftests/mm/khugepaged_race.c | 442 +++++++++++++++++++ tools/testing/selftests/mm/run_vmtests.sh | 2 + 3 files changed, 445 insertions(+) create mode 100644 tools/testing/selftests/mm/khugepaged_race.c diff --git a/tools/testing/selftests/mm/Makefile b/tools/testing/selftests/= mm/Makefile index b2d6e5c12934..308bbad73c11 100644 --- a/tools/testing/selftests/mm/Makefile +++ b/tools/testing/selftests/mm/Makefile @@ -106,6 +106,7 @@ TEST_GEN_FILES +=3D rmap TEST_GEN_FILES +=3D folio_split_race_test TEST_GEN_FILES +=3D folio_order_check TEST_GEN_FILES +=3D khugepaged_sync_check +TEST_GEN_FILES +=3D khugepaged_race =20 ifneq ($(ARCH),arm64) TEST_GEN_FILES +=3D soft-dirty diff --git a/tools/testing/selftests/mm/khugepaged_race.c b/tools/testing/s= elftests/mm/khugepaged_race.c new file mode 100644 index 000000000000..3128031e28bc --- /dev/null +++ b/tools/testing/selftests/mm/khugepaged_race.c @@ -0,0 +1,442 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * khugepaged race harness. + * + * Runs collapse against concurrent faults, transient GUP pins + * (gup_test), fork, mremap and MADV_DONTNEED over the same ranges, in + * one of three driver modes: + * + * stepped khugepaged, one full pass at a time through + * khugepaged_full_pass(), so a step covers a known extent; + * free khugepaged left to run (scan_sleep_millisecs=3D0), for soak; + * madvise MADV_COLLAPSE and MADV_DONTNEED in a loop. + * + * All anon THP orders are enabled (inherit) and max_ptes_none is 0, so a + * window has to be fully populated before khugepaged will collapse it, and + * the racing MADV_DONTNEED decides which orders it can still use. + * + * Correctness signals: every racing page must read as its pattern or + * zero (MADV_DONTNEED), never anything else. The faulters and the fork + * children check that continuously, a final sweep checks it once more, pl= us + * whatever DEBUG_VM / page_table_check / KASAN / lockdep report in + * dmesg, which the caller is expected to inspect. + */ +#define _GNU_SOURCE +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "kselftest.h" +#include "vm_util.h" +#include "hugepage_settings.h" +#include "../../../../mm/gup_test.h" + +#define BASE_ADDR ((void *)(1UL << 30)) + +/* + * Shared playground for faults/pins/fork/dontneed: several PMD-sized + * areas the racing threads spread across, plus one area owned by the + * mremap thread. More areas means more independent regions collapsing + * at once; the default suits a normal machine. On a memory-constrained + * host -- or under emulation, where a 512M PMD (arm64/64K) makes the + * default playground multi-gigabyte -- pass -a to shrink it. + */ +#define DEFAULT_SHARED_AREAS 3 +static int nr_shared_areas; +static int nr_areas; + +static unsigned long hpage_pmd_size; +static unsigned long page_size; +static char *region; /* NR_AREAS * hpage_pmd_size */ +static char *mremap_area; /* region + NR_SHARED_AREAS areas */ +static char *mremap_scratch; /* well above the region */ +static int gup_fd =3D -1; +static volatile int stop; +static volatile int corrupted; + +static unsigned int pattern(unsigned long page_idx) +{ + unsigned int val =3D (unsigned int)page_idx * 2654435761U; + + return val ? val : 1; /* never collides with the zero-fill */ +} + +/* Zero means never written; anything else must be this page's pattern */ +static bool page_is_corrupt(unsigned long page_idx, unsigned int *val) +{ + *val =3D *(unsigned int *)(region + page_idx * page_size); + + return *val && *val !=3D pattern(page_idx); +} + +static void check_page(unsigned long page_idx) +{ + unsigned int val; + + if (page_is_corrupt(page_idx, &val)) { + corrupted =3D 1; + ksft_print_msg("Corruption at page %lu: %#x !=3D %#x\n", + page_idx, val, pattern(page_idx)); + } +} + +static unsigned long shared_pages(void) +{ + return nr_shared_areas * hpage_pmd_size / page_size; +} + +static unsigned long rand_page(unsigned int *seed) +{ + return (unsigned long)rand_r(seed) % shared_pages(); +} + +/* Pages left from @page_idx, so a range never reaches the mremap thread's= area */ +static unsigned long room_from(unsigned long page_idx, unsigned long want) +{ + unsigned long left =3D shared_pages() - page_idx; + + return want < left ? want : left; +} + +static void *faulter_fn(void *arg) +{ + unsigned int seed =3D (unsigned long)arg; + + while (!stop) { + unsigned long page_idx =3D rand_page(&seed); + + if (rand_r(&seed) & 1) + *(unsigned int *)(region + page_idx * page_size) =3D + pattern(page_idx); + else + check_page(page_idx); + } + return NULL; +} + +static void *dontneed_fn(void *arg) +{ + unsigned int seed =3D (unsigned long)arg; + + while (!stop) { + unsigned long page_idx =3D rand_page(&seed); + unsigned long nr =3D 1UL << (rand_r(&seed) % 6); /* 1..32 pages */ + + madvise(region + page_idx * page_size, + room_from(page_idx, nr) * page_size, MADV_DONTNEED); + usleep(rand_r(&seed) % 500); + } + return NULL; +} + +static void *pinner_fn(void *arg) +{ + unsigned int seed =3D (unsigned long)arg; + + while (!stop) { + struct gup_test gup =3D {}; + unsigned long page_idx =3D rand_page(&seed); + + unsigned long nr =3D room_from(page_idx, 16); + + gup.addr =3D (unsigned long)(region + page_idx * page_size); + gup.size =3D nr * page_size; + gup.nr_pages_per_call =3D nr; + gup.gup_flags =3D 1; /* FOLL_WRITE */ + /* Racing MADV_DONTNEED makes transient failures expected. */ + ioctl(gup_fd, PIN_FAST_BENCHMARK, &gup); + usleep(rand_r(&seed) % 200); + } + return NULL; +} + +static void *forker_fn(void *arg) +{ + unsigned int seed =3D (unsigned long)arg; + + while (!stop) { + pid_t pid =3D fork(); + + if (pid =3D=3D 0) { + unsigned int val; + int bad =3D 0; + + /* + * No stdio in the child: a thread may have held + * stdout's lock when we forked, and printing under an + * inherited lock hangs. The parent reports what the + * exit status says. + */ + for (int i =3D 0; i < 16; i++) + bad |=3D page_is_corrupt(rand_page(&seed), &val); + _exit(bad); + } + if (pid > 0) { + int wstatus; + + if (waitpid(pid, &wstatus, 0) < 0) + ksft_exit_fail_perror("waitpid()"); + /* A child killed on the read counts too, not just its exit code. */ + if (!WIFEXITED(wstatus) || WEXITSTATUS(wstatus)) + corrupted =3D 1; + } + usleep(rand_r(&seed) % 2000); + } + return NULL; +} + +static void *mremapper_fn(void *arg) +{ + unsigned int seed =3D (unsigned long)arg; + + while (!stop) { + void *p; + + p =3D mremap(mremap_area, hpage_pmd_size, hpage_pmd_size, + MREMAP_MAYMOVE | MREMAP_FIXED, mremap_scratch); + if (p =3D=3D MAP_FAILED) + ksft_exit_fail_perror("mremap() away"); + for (int i =3D 0; i < 8; i++) + mremap_scratch[(rand_r(&seed) % + (hpage_pmd_size / page_size)) * page_size] =3D 1; + p =3D mremap(mremap_scratch, hpage_pmd_size, hpage_pmd_size, + MREMAP_MAYMOVE | MREMAP_FIXED, mremap_area); + if (p =3D=3D MAP_FAILED) + ksft_exit_fail_perror("mremap() back"); + usleep(rand_r(&seed) % 2000); + } + return NULL; +} + +static unsigned long now_ms(void) +{ + struct timeval tv; + + gettimeofday(&tv, NULL); + return tv.tv_sec * 1000UL + tv.tv_usec / 1000; +} + +static void usage(void) +{ + fprintf(stderr, + "Usage: khugepaged_race [-d seconds] [-m stepped|free|madvise] [-a areas= ]\n" + "\tWithout -m, every mode runs in turn.\n" + "\t-d: seconds per mode (default 5)\n" + "\t-a: number of shared PMD-sized playground areas (default 3)\n"); + exit(1); +} + +int main(int argc, char **argv) +{ + static const char * const thread_names[] =3D { + "faulter", "faulter2", "dontneed", "pinner", "forker", + "mremapper", + }; + void *(*const thread_fns[])(void *) =3D { + faulter_fn, faulter_fn, dontneed_fn, pinner_fn, forker_fn, + mremapper_fn, + }; + const int nr_threads =3D ARRAY_SIZE(thread_names); + pthread_t threads[ARRAY_SIZE(thread_names)]; + static const char * const all_modes[] =3D { "stepped", "free", "madvise" = }; + const char *one_mode[1]; + const char * const *modes =3D all_modes; + int nr_modes =3D ARRAY_SIZE(all_modes); + const char *mode_arg =3D NULL; + struct thp_settings settings; + unsigned long end_ms; + int duration_s =3D 5; + unsigned long thread_mask =3D ~0UL; + int nr_areas_arg =3D 0; + unsigned long i; + int steps =3D 0; + int opt; + + while ((opt =3D getopt(argc, argv, "a:d:m:t:h")) !=3D -1) { + switch (opt) { + case 'a': + nr_areas_arg =3D atoi(optarg); + break; + case 'd': + duration_s =3D atoi(optarg); + break; + case 'm': + mode_arg =3D optarg; + break; + case 't': + /* debug: bitmask of racing threads to start */ + thread_mask =3D strtoul(optarg, NULL, 0); + break; + default: + usage(); + } + } + if (mode_arg) { + if (strcmp(mode_arg, "stepped") && strcmp(mode_arg, "free") && + strcmp(mode_arg, "madvise")) + usage(); + one_mode[0] =3D mode_arg; + modes =3D one_mode; + nr_modes =3D 1; + } + + ksft_print_header(); + if (!thp_available()) + ksft_exit_skip("Transparent Hugepages not available\n"); + + page_size =3D getpagesize(); + hpage_pmd_size =3D read_pmd_pagesize(); + if (!hpage_pmd_size) + ksft_exit_fail_msg("Reading PMD pagesize failed\n"); + + gup_fd =3D open("/sys/kernel/debug/gup_test", O_RDWR); + if (gup_fd < 0) + ksft_exit_skip("/sys/kernel/debug/gup_test requires CONFIG_GUP_TEST and = root\n"); + + nr_shared_areas =3D nr_areas_arg > 0 ? nr_areas_arg : DEFAULT_SHARED_AREA= S; + nr_areas =3D nr_shared_areas + 1; + + /* + * The mremap thread moves its area to this address and back, and + * MREMAP_FIXED unmaps whatever is in the way without saying so. Claim + * the address here, so a layout that does not match this assumption + * fails now instead of losing a mapping later. Nothing else in the + * process maps this low: thread stacks and malloc arenas come from the + * top-down mmap area, well above. + */ + mremap_scratch =3D (char *)BASE_ADDR + 2 * nr_areas * hpage_pmd_size; + if (mmap(mremap_scratch, hpage_pmd_size, PROT_NONE, + MAP_ANONYMOUS | MAP_PRIVATE | MAP_FIXED_NOREPLACE, + -1, 0) !=3D (void *)mremap_scratch) + ksft_exit_fail_perror("mmap() mremap scratch"); + + ksft_set_plan(nr_modes); + + thp_save_settings(); + thp_read_settings(&settings); + + /* + * A base entry for the stack, so that the pop at the end of a mode + * always has something to write back: thp_pop_settings() on an empty + * stack has no settings to apply and gives up. + */ + thp_push_settings(&settings); + + for (int m =3D 0; m < nr_modes; m++) { + const char *mode =3D modes[m]; + + thp_read_settings(&settings); + settings.thp_enabled =3D THP_MADVISE; + settings.thp_defrag =3D THP_DEFRAG_ALWAYS; + settings.shmem_enabled =3D SHMEM_NEVER; + settings.khugepaged.defrag =3D 1; + settings.khugepaged.scan_sleep_millisecs =3D + strcmp(mode, "free") ? 1000 : 0; + settings.khugepaged.alloc_sleep_millisecs =3D 10; + /* + * Strict occupancy: mTHP collapse only supports 0 or + * HPAGE_PMD_NR - 1 and coerces anything else to 0 anyway, and 0 + * also keeps khugepaged from burning the whole step in doomed + * PMD-sized allocations on 512M-PMD configs: under racing + * MADV_DONTNEED a fully populated PMD area is rare. + */ + settings.khugepaged.max_ptes_none =3D 0; + settings.khugepaged.pages_to_scan =3D + nr_areas * (hpage_pmd_size / page_size) * 8; + for (i =3D 0; i < NR_ORDERS; i++) { + if (thp_supported_orders() & (1UL << i)) + settings.hugepages[i].enabled =3D THP_INHERIT; + } + /* Popped at the end of this mode, before the next one. */ + thp_push_settings(&settings); + + region =3D mmap(BASE_ADDR, nr_areas * hpage_pmd_size, + PROT_READ | PROT_WRITE, MAP_ANONYMOUS | + MAP_PRIVATE | MAP_FIXED_NOREPLACE, -1, 0); + if (region !=3D BASE_ADDR) + ksft_exit_fail_perror("mmap() playground"); + mremap_area =3D region + nr_shared_areas * hpage_pmd_size; + + /* Populate so the first pass has something to collapse. */ + for (i =3D 0; i < nr_shared_areas * hpage_pmd_size / page_size; i++) + *(unsigned int *)(region + i * page_size) =3D pattern(i); + memset(mremap_area, 1, hpage_pmd_size); + if (madvise(region, nr_areas * hpage_pmd_size, MADV_HUGEPAGE)) + ksft_exit_fail_perror("madvise(MADV_HUGEPAGE)"); + + for (i =3D 0; i < nr_threads; i++) { + if (!(thread_mask & (1UL << i))) { + threads[i] =3D 0; + continue; + } + if (pthread_create(&threads[i], NULL, thread_fns[i], + (void *)(i + 1))) + ksft_exit_fail_perror(thread_names[i]); + } + + end_ms =3D now_ms() + duration_s * 1000UL; + if (!strcmp(mode, "stepped")) { + while (now_ms() < end_ms && !corrupted) { + if (!khugepaged_full_pass(600)) + ksft_exit_fail_msg("khugepaged pass timed out\n"); + steps++; + } + } else if (!strcmp(mode, "free")) { + while (now_ms() < end_ms && !corrupted) + usleep(100 * 1000); + } else { /* madvise */ + while (now_ms() < end_ms && !corrupted) { + for (i =3D 0; i < nr_shared_areas; i++) { + madvise(region + i * hpage_pmd_size, + hpage_pmd_size, MADV_COLLAPSE); + } + madvise(region, nr_shared_areas * hpage_pmd_size, + MADV_DONTNEED); + steps++; + } + } + + stop =3D 1; + for (i =3D 0; i < nr_threads; i++) { + if (threads[i]) + pthread_join(threads[i], NULL); + } + + /* Final integrity sweep. */ + for (i =3D 0; i < nr_shared_areas * hpage_pmd_size / page_size; i++) + check_page(i); + + ksft_test_result(!corrupted, + "%s: %ds, %d steps, no corruption\n", + mode, duration_s, steps); + + /* + * Hand the address space and the settings back before the + * next mode: it maps the region at the same fixed address, + * and its scan cadence differs. + */ + munmap(region, nr_areas * hpage_pmd_size); + thp_pop_settings(); + stop =3D 0; + steps =3D 0; + + if (corrupted) { + /* Memory is suspect; the rest would prove nothing. */ + while (++m < nr_modes) + ksft_test_result_skip("%s: skipped after corruption\n", + modes[m]); + break; + } + } + + thp_restore_settings(); + ksft_finished(); +} diff --git a/tools/testing/selftests/mm/run_vmtests.sh b/tools/testing/self= tests/mm/run_vmtests.sh index c0f69da3fd3b..fc61907aa3b2 100755 --- a/tools/testing/selftests/mm/run_vmtests.sh +++ b/tools/testing/selftests/mm/run_vmtests.sh @@ -406,6 +406,8 @@ CATEGORY=3D"thp" run_test ./folio_order_check =20 CATEGORY=3D"thp" run_test ./khugepaged_sync_check =20 +CATEGORY=3D"thp" run_test ./khugepaged_race + CATEGORY=3D"thp" run_test ./khugepaged =20 CATEGORY=3D"thp" run_test ./khugepaged -s 2 --=20 2.54.0 From nobody Mon Sep 28 23:56:19 2026 Received: from fout-a6-smtp.messagingengine.com (fout-a6-smtp.messagingengine.com [103.168.172.149]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E9CA334D384; Sat, 15 Aug 2026 01:59:38 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.149 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759180; cv=none; b=NNIXKo0zxoSYB61LuwCY+asP+Sk8IS8nFFadIitLbYjbz62IH9A6DYzTEdDwsQYMfrjWAFUVklDWFVWAndA9lXfUKmxZhUDbdI5lfB12Dl1fUZhK8stZmFew0WuRSIqkKbNmiE8VEvVSSTpu00RtQTTf8bUvMM+iH/X6GOAjYt4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759180; c=relaxed/simple; bh=6l8VBCYedhCpNLCqH4lpa4kGPS9nRSFObj0Hr7I8990=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=MzH6c3lYloaW3W2ADU1mj/UNx66+s9o+be4Lsj3FlcYyCqOL2mzm8T1klst44RUTpIjRDT2wnPLMg4cWDBH5kpox4pxlCb2Tnh/yHVKzHmn4+54HAxoIzYyBh2wvkCzCVx3txtI8q2Cl0DklU7PnGMEbNe5Nh8nnuq3D7g1lG1Q= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=rsXM3TNO; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=d2mRcc5i; arc=none smtp.client-ip=103.168.172.149 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="rsXM3TNO"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="d2mRcc5i" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfout.phl.internal (Postfix) with ESMTP id ED405EC022A; Fri, 14 Aug 2026 21:59:37 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-06.internal (MEProxy); Fri, 14 Aug 2026 21:59:37 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786759177; x= 1786845577; bh=4qPwRkvpOUtfuWl8Vd1ZbS3CmUow1OnKorGDRJPSt44=; b=r sXM3TNOpkQ0bDeikayo57HaHNhnZFPxRt75G1/GaU9VjnOJd6Dx28jcDjB9MgJ1m Bi4i7urR9KcvIF0FoY0VY2VPT2HWUOC6f6rwhjNVwyxjhz3BJtAZFFvoXYlJKN0e 56tyL1CldILvlz/uoJ8g3fUSdJxYyTKIhjCT7CzbQcnDsNXAMlD4Qb7lvF/ddzUi yMKRx2KapfEyjrHR6GoJWGvOZ+NobGhglYb2b7COYNG0RBYoj+s7Z/wpZMrHNIpi shDxeiq0kXbNtJ5ckkIWcAnmnG762dE95fF2+WaX29szf2PDmhCLY6ckoZDkRkYG GOEfygKz5PdL1CBcS8/eA== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786759177; x=1786845577; bh=4 qPwRkvpOUtfuWl8Vd1ZbS3CmUow1OnKorGDRJPSt44=; b=d2mRcc5iYTpWw90KA +ukgAtBxmM/BXvSMYtVfjRq7FFNjPCyJMaDVXB/cpO+cUsqB81g4J1KPx0bgwKHU jEZ5rg1aSEV5Bac+0wBeA7EQCDrML2ukAl/3wQr+m8Fg2ykzivLoPQS2WJj01gNs aVD3VPG3hoHvMafkysQr7nV8QO3acQUgoxrkTmBO85icJ3VKIOGATz21kxBhGgK/ vQ6x5urYBBGMg8w/Dxd3egYlfcheggsYeliXz6pkjX704bLAu5NYu3P0M3tlmqJC g7rEvEM5vlmq8NbUpOvwAQGGWnGLlugyeWGvEACUNGYnC2c2cBsfxqJxaMaZ2r/4 MDbjA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFzGpGFm5W4vTTPFWblEwVteikp8ZQmNhgJQ39PW7T7XQcfx50sk9SsorwQ5kiHE1 GkGyp58aXi1XV7w9FWqCEkAAXyywQWDaCYDP4jk79Tj+fneRM/f70WhpHyk68ivSHa8CNO EHCv/WuNeVKTxC+1Am5pTEbXGz94MYKYsyM7kpGX5dHsQKSWoIOFoaOZhong5pPq1D5dt9 kZJFmHpp6Lnk5BJwnqgflVky13nVISEGlBe9NK3zcAVq5kZlGc2PqnLHdjdexqJ0GH5I+D CyEe7NNeMVRnhwRmAQGYSBBJ6RBvmXxfczeNb/t4Esr0tR7Igh8W/aIlPiSt8FyWTQd+Pj ToEO9Wkaygl0tzRAxiOsUatANGv0MavqvzeRLvpV66n2CVAAFVetansC73Poslu2NI22gA LuHvDL56Nnf4nUNnsl4rmbfcyV9TlDqb+JYRr3GQDIpUzFXIeaG+3jlY1/8ryaOqdzRLTK VtHSqOmREIUGrw8+Iys5g1m2eidfl+Bv3ONAEXMe4a64tisQZfxfMPAUAo/gKuVWCkyEEG SUz0UTj52k6v+IGP3gFtAK34688BHEl6p0pyK0rKrQrIGqVfA1wGz39ATLzIvU+mkOx07L 1otNCfkXJ4SeWOft+a/+wzrLVnGem4uE1OwH4u0ngNjV7ERIrg7+vp0XONPQ X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 14 Aug 2026 21:59:37 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v4 17/19] selftests/mm: race collapse of windows with holes Date: Sat, 15 Aug 2026 02:58:59 +0100 Message-ID: <20260815015901.1236937-18-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260815015901.1236937-1-kirill@shutemov.name> References: <20260815015901.1236937-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The harness pins max_ptes_none to 0, so khugepaged only collapses a window once every PTE in it is present. A window with holes takes a different route, and never gets raced. A hole is zero-filled into the new folio by clear_user_highpage() rather than copied from anywhere. Which slots count as holes keeps moving under the racing MADV_DONTNEED, right up to the moment the PMD is detached. Run both ends of the occupancy scale for every driver mode, one after the other. mTHP collapse supports only those two, 0 and HPAGE_PMD_NR - 1, and coerces anything between them to 0. Each result says which end it ran: ok 1 stepped/strict: 5s, 231 steps, no corruption ok 2 stepped/holes: 5s, 194 steps, no corruption Assisted-by: Claude-Code:claude-opus-5 Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/khugepaged_race.c | 45 ++++++++++++-------- 1 file changed, 28 insertions(+), 17 deletions(-) diff --git a/tools/testing/selftests/mm/khugepaged_race.c b/tools/testing/s= elftests/mm/khugepaged_race.c index 3128031e28bc..3b369046edf5 100644 --- a/tools/testing/selftests/mm/khugepaged_race.c +++ b/tools/testing/selftests/mm/khugepaged_race.c @@ -11,9 +11,11 @@ * free khugepaged left to run (scan_sleep_millisecs=3D0), for soak; * madvise MADV_COLLAPSE and MADV_DONTNEED in a loop. * - * All anon THP orders are enabled (inherit) and max_ptes_none is 0, so a - * window has to be fully populated before khugepaged will collapse it, and - * the racing MADV_DONTNEED decides which orders it can still use. + * All anon THP orders are enabled (inherit). Occupancy runs at both ends + * of what mTHP collapse supports: max_ptes_none 0, where a window must be + * fully populated, and HPAGE_PMD_NR - 1, where a window full of holes + * collapses too. A hole is zero-filled into the new folio rather than + * copied, and the racing MADV_DONTNEED keeps moving which slots are holes. * * Correctness signals: every racing page must read as its pattern or * zero (MADV_DONTNEED), never anything else. The faulters and the fork @@ -247,6 +249,8 @@ int main(int argc, char **argv) const int nr_threads =3D ARRAY_SIZE(thread_names); pthread_t threads[ARRAY_SIZE(thread_names)]; static const char * const all_modes[] =3D { "stepped", "free", "madvise" = }; + static const bool occupancies[] =3D { false, true }; /* strict, holes */ + const int nr_occupancies =3D ARRAY_SIZE(occupancies); const char *one_mode[1]; const char * const *modes =3D all_modes; int nr_modes =3D ARRAY_SIZE(all_modes); @@ -279,6 +283,7 @@ int main(int argc, char **argv) usage(); } } + if (mode_arg) { if (strcmp(mode_arg, "stepped") && strcmp(mode_arg, "free") && strcmp(mode_arg, "madvise")) @@ -318,7 +323,7 @@ int main(int argc, char **argv) -1, 0) !=3D (void *)mremap_scratch) ksft_exit_fail_perror("mmap() mremap scratch"); =20 - ksft_set_plan(nr_modes); + ksft_set_plan(nr_modes * nr_occupancies); =20 thp_save_settings(); thp_read_settings(&settings); @@ -330,8 +335,9 @@ int main(int argc, char **argv) */ thp_push_settings(&settings); =20 - for (int m =3D 0; m < nr_modes; m++) { - const char *mode =3D modes[m]; + for (int mn =3D 0; mn < nr_modes * nr_occupancies; mn++) { + const char *mode =3D modes[mn / nr_occupancies]; + bool holes =3D occupancies[mn % nr_occupancies]; =20 thp_read_settings(&settings); settings.thp_enabled =3D THP_MADVISE; @@ -341,14 +347,16 @@ int main(int argc, char **argv) settings.khugepaged.scan_sleep_millisecs =3D strcmp(mode, "free") ? 1000 : 0; settings.khugepaged.alloc_sleep_millisecs =3D 10; + /* - * Strict occupancy: mTHP collapse only supports 0 or - * HPAGE_PMD_NR - 1 and coerces anything else to 0 anyway, and 0 - * also keeps khugepaged from burning the whole step in doomed - * PMD-sized allocations on 512M-PMD configs: under racing - * MADV_DONTNEED a fully populated PMD area is rare. + * mTHP collapse only supports the two ends of the occupancy + * scale: 0 or HPAGE_PMD_NR - 1 (anything else coerces to 0). + * Strict needs a fully populated window, which is rare under + * racing MADV_DONTNEED; hole-heavy windows collapse instead, + * so the two ends race different paths. */ - settings.khugepaged.max_ptes_none =3D 0; + settings.khugepaged.max_ptes_none =3D holes ? + (hpage_pmd_size / page_size) - 1 : 0; settings.khugepaged.pages_to_scan =3D nr_areas * (hpage_pmd_size / page_size) * 8; for (i =3D 0; i < NR_ORDERS; i++) { @@ -415,8 +423,9 @@ int main(int argc, char **argv) check_page(i); =20 ksft_test_result(!corrupted, - "%s: %ds, %d steps, no corruption\n", - mode, duration_s, steps); + "%s/%s: %ds, %d steps, no corruption\n", + mode, holes ? "holes" : "strict", + duration_s, steps); =20 /* * Hand the address space and the settings back before the @@ -430,9 +439,11 @@ int main(int argc, char **argv) =20 if (corrupted) { /* Memory is suspect; the rest would prove nothing. */ - while (++m < nr_modes) - ksft_test_result_skip("%s: skipped after corruption\n", - modes[m]); + while (++mn < nr_modes * nr_occupancies) + ksft_test_result_skip("%s/%s: skipped after corruption\n", + modes[mn / nr_occupancies], + occupancies[mn % nr_occupancies] ? + "holes" : "strict"); break; } } --=20 2.54.0 From nobody Mon Sep 28 23:56:19 2026 Received: from fout-a6-smtp.messagingengine.com (fout-a6-smtp.messagingengine.com [103.168.172.149]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D2AB3367B8E; Sat, 15 Aug 2026 01:59:40 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.149 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759182; cv=none; b=jvUqWYOqg1rgcGh4Hgxg51g7iOpPEymWJylaC1+Ny7D4CxArrIzj2mdlklNO7M+ulzwCWURlyRvAHglCsvtog/20OKZ/vLFX0PrSwe/t5K3gHCylnt7OOB4JCA3jR5fTKoiImR5zdom4f6zD21j5gJPXSdXUEpz1ixWksJeJUOU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759182; c=relaxed/simple; bh=ooI4yMykjGTIPAa7034+pYpwGfXX6+IcoJrym4helMs=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=ppVitnB5IuQDTrzMbmKX1J24MtxW398ixM2AtZQY2nzNAWBc3WrvOJmqdBVONDc69oMi5I0tHFMeh+28ZiPgILO8kT19GkCwpgPtiYwzGrI2W/Kg9E1hGAbn0lpKehr2aaloYGYLd59/Q+cabFcu77j0aqRI3pLq41ZHyglp2mw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=D3mtboiF; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=NsrmYP7W; arc=none smtp.client-ip=103.168.172.149 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="D3mtboiF"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="NsrmYP7W" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfout.phl.internal (Postfix) with ESMTP id A420CEC0238; Fri, 14 Aug 2026 21:59:39 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-06.internal (MEProxy); Fri, 14 Aug 2026 21:59:39 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786759179; x= 1786845579; bh=JK5GVI6pI2lxoa84VhaQKI3w4xKylzP6fAx1cS2WfAU=; b=D 3mtboiFR8kAyvDL6O4AAcA0NezFYPVsOBnyzr59jVjz0lHrw/n7/MXZAkjs4b620 +036gY92utRhm5FoKr3ImhrTHrMvMa0uDnoe/umnyg1yirU+b7UuqD0xYXIVzRZ8 Y8k20gVLeor+svnZLALZeyznV9gJCUo2ztgxE+TnYPRAZg0xzEroHJz/4wTBDmEZ F32gFOJbzh5RUAijasgnvSBy475Wysl/HVM99AkmJnHev3rra1H5YOS8W+sclsuk ER0opUBTumB5T0E9epZUNA/ZYNx9eQjRAqbelVEjljcWyJW41Nx7lW5qbUI1xasV MidifAiUitaU+Q3KHc65w== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786759179; x=1786845579; bh=J K5GVI6pI2lxoa84VhaQKI3w4xKylzP6fAx1cS2WfAU=; b=NsrmYP7WnbCsDiaKE +BtqmC4eB9oen/ucBplSll7oiH2kHU40wlRIPQoFg94/hKjh2lyiSzRTOp8JvSCq JF5337sawFgeLeUC2kRsX2kCG/LeJznHCf2C3k7q/g1fgzy17StdLNjewE5Qrc87 Yqi5h3mmE+Z4Q0lca9w4fFyHkx/s8k4QoSRDLeTa29D7c/TDRajBIelQ1ILe1TQC IIKzxmX8pBMfYjU/Lx1eY32JZ6zCEXr0lEHZfV/n4YkROXIFeNdv41lLGcLavduJ ++bJ+8HVPZg014EKnk8pcmgQFh2mi2DYEOzGa2FZnpYQgk6Rr1XrytpDg4BzxlDg b9OWw== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFzGpGFm5W4vTTPFWblEwVteikp8ZQmNhgJQ39PW7T7XQcfx50sk9SsorwQ5kiHE1 GkGyp58aXi1XV7w9FWqCEkAAXyywQWDaCYDP4jk79Tj+fneRM/f70WhpHyk68ivSHa8CNO EHCv/WuNeVKTxC+1Am5pTEbXGz94MYKYsyM7kpGX5dHsQKSWoIOFoaOZhong5pPq1D5dt9 kZJFmHpp6Lnk5BJwnqgflVky13nVISEGlBe9NK3zcAVq5kZlGc2PqnLHdjdexqJ0GH5I+D CyEe7NNeMVRnhwRmAQGYSBBJ6RBvmXxfczeNb/t4Esr0tR7Igh8W/aIlPiSt8FyWTQd+TG GnssYIczN2LI/Xz94YDefGLowoNoiM6pqpAUodaYbR6hi2S4OnC2iEHjxeNfrDS9HJXd2f 6IXWtuE6XJrzJTWNEAJgdYW2s+C51dwE/YnPLYE+HrvL5uB09u1QbSTWUynuv92A4cmqT1 uMUJOkGjlyorBDYinuQkYUiV5ws9EBM9YxluzUjsCJN7p1yORebiDQdoM8urJIm7L2qpIl +k9nX7OtbQij9IAgZW1EW6gDbm9nndabfK8vO5jsdbZ82g1CKkd8Wj8huKmCfXIRb32UZV aE5H1qmmSf9V/qWZ3bw3wrXcKyUhCHDA11D0RzeVcbH74bO9xK+ZkznWeKog X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 14 Aug 2026 21:59:39 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v4 18/19] selftests/mm: add memory-pressure threads to the khugepaged race harness Date: Sat, 15 Aug 2026 02:59:00 +0100 Message-ID: <20260815015901.1236937-19-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260815015901.1236937-1-kirill@shutemov.name> References: <20260815015901.1236937-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The harness races collapse against faults, pins, fork, mremap and MADV_DONTNEED, but nothing in it elevates a source folio's refcount from the reclaim or compaction side. Add two more threads, and run every mode and occupancy limit both with and without them: - pageout: cycles MADV_PAGEOUT over a dedicated neighbour region, faults it back in and checks the content each round, since a page's pattern must survive the trip through swap. Left out when the host has no swap, because then there is no anon reclaim to drive. - compactor: writes /proc/sys/vm/compact_memory in a loop. Compaction isolates and migrates folios, so it competes with a collapse for the pages it is gathering, with refcount elevations and migration entries of its own. Each result says whether it ran under pressure: ok 2 stepped/strict/pressure: 5s, 88 steps, no corruption A full run is now twelve combinations; -m and -d narrow it. Assisted-by: Claude-Code:claude-opus-5 Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/khugepaged_race.c | 144 +++++++++++++++++-- 1 file changed, 132 insertions(+), 12 deletions(-) diff --git a/tools/testing/selftests/mm/khugepaged_race.c b/tools/testing/s= elftests/mm/khugepaged_race.c index 3b369046edf5..a23e5bfe78af 100644 --- a/tools/testing/selftests/mm/khugepaged_race.c +++ b/tools/testing/selftests/mm/khugepaged_race.c @@ -17,6 +17,12 @@ * collapses too. A hole is zero-filled into the new folio rather than * copied, and the racing MADV_DONTNEED keeps moving which slots are holes. * + * Every combination also runs under memory pressure: MADV_PAGEOUT cycling + * on a dedicated neighbour region (swap traffic and LRU churn; left out + * with a note when the host has no swap) and a compact_memory trigger + * loop (compaction migrates source folios, racing collapse's freeze + * with refcount elevation and migration entries of its own). + * * Correctness signals: every racing page must read as its pattern or * zero (MADV_DONTNEED), never anything else. The faulters and the fork * children check that continuously, a final sweep checks it once more, pl= us @@ -60,6 +66,8 @@ static unsigned long page_size; static char *region; /* NR_AREAS * hpage_pmd_size */ static char *mremap_area; /* region + NR_SHARED_AREAS areas */ static char *mremap_scratch; /* well above the region */ +static char *pageout_area; /* dedicated pressure region */ +static size_t pageout_size; static int gup_fd =3D -1; static volatile int stop; static volatile int corrupted; @@ -218,6 +226,70 @@ static void *mremapper_fn(void *arg) return NULL; } =20 +/* + * Swap traffic and LRU churn on a region of our own. The content + * check is exact: a page out and back through swap must preserve the + * pattern, and nothing else ever writes here. + */ +static void *pageout_fn(void *arg) +{ + unsigned int seed =3D (unsigned long)arg; + unsigned long nr =3D pageout_size / page_size; + unsigned long i; + + for (i =3D 0; i < nr; i++) + *(unsigned int *)(pageout_area + i * page_size) =3D pattern(i); + + while (!stop) { + madvise(pageout_area, pageout_size, MADV_PAGEOUT); + for (i =3D 0; i < nr && !stop; i++) { + unsigned int val =3D *(unsigned int *)(pageout_area + + i * page_size); + + if (val !=3D pattern(i)) { + corrupted =3D 1; + ksft_print_msg("Pageout corruption at page %lu: %#x !=3D %#x\n", + i, val, pattern(i)); + } + } + usleep(rand_r(&seed) % 2000); + } + return NULL; +} + +/* Compaction migrates the collapse sources out from under us. */ +static void *compactor_fn(void *arg) +{ + unsigned int seed =3D (unsigned long)arg; + int fd =3D open("/proc/sys/vm/compact_memory", O_WRONLY); + + if (fd < 0) { + ksft_print_msg("No compact_memory; compactor idle\n"); + return NULL; + } + while (!stop) { + if (write(fd, "1", 1) < 0) + break; + usleep(10000 + rand_r(&seed) % 100000); + } + close(fd); + return NULL; +} + +static bool swap_available(void) +{ + char line[256]; + int lines =3D 0; + FILE *fp =3D fopen("/proc/swaps", "r"); + + if (!fp) + return false; + while (fgets(line, sizeof(line), fp)) + lines++; + fclose(fp); + return lines > 1; +} + static unsigned long now_ms(void) { struct timeval tv; @@ -240,17 +312,23 @@ int main(int argc, char **argv) { static const char * const thread_names[] =3D { "faulter", "faulter2", "dontneed", "pinner", "forker", - "mremapper", + "mremapper", "pageout", "compactor", }; void *(*const thread_fns[])(void *) =3D { faulter_fn, faulter_fn, dontneed_fn, pinner_fn, forker_fn, - mremapper_fn, + mremapper_fn, pageout_fn, compactor_fn, }; + enum { T_FAULTER, T_FAULTER2, T_DONTNEED, T_PINNER, T_FORKER, + T_MREMAPPER, T_PAGEOUT, T_COMPACTOR }; + const unsigned long pageout_bit =3D 1UL << T_PAGEOUT; + const unsigned long compactor_bit =3D 1UL << T_COMPACTOR; const int nr_threads =3D ARRAY_SIZE(thread_names); pthread_t threads[ARRAY_SIZE(thread_names)]; static const char * const all_modes[] =3D { "stepped", "free", "madvise" = }; static const bool occupancies[] =3D { false, true }; /* strict, holes */ + static const bool pressures[] =3D { false, true }; /* quiet, under pressu= re */ const int nr_occupancies =3D ARRAY_SIZE(occupancies); + const int nr_pressures =3D ARRAY_SIZE(pressures); const char *one_mode[1]; const char * const *modes =3D all_modes; int nr_modes =3D ARRAY_SIZE(all_modes); @@ -259,6 +337,7 @@ int main(int argc, char **argv) unsigned long end_ms; int duration_s =3D 5; unsigned long thread_mask =3D ~0UL; + unsigned long base_mask; int nr_areas_arg =3D 0; unsigned long i; int steps =3D 0; @@ -323,7 +402,12 @@ int main(int argc, char **argv) -1, 0) !=3D (void *)mremap_scratch) ksft_exit_fail_perror("mmap() mremap scratch"); =20 - ksft_set_plan(nr_modes * nr_occupancies); + base_mask =3D thread_mask; + if (!swap_available()) + /* No swap, no anon reclaim: compaction-only pressure. */ + ksft_print_msg("no swap: the pageout thread stays idle\n"); + + ksft_set_plan(nr_modes * nr_occupancies * nr_pressures); =20 thp_save_settings(); thp_read_settings(&settings); @@ -335,9 +419,17 @@ int main(int argc, char **argv) */ thp_push_settings(&settings); =20 - for (int mn =3D 0; mn < nr_modes * nr_occupancies; mn++) { - const char *mode =3D modes[mn / nr_occupancies]; - bool holes =3D occupancies[mn % nr_occupancies]; + for (int run =3D 0; run < nr_modes * nr_occupancies * nr_pressures; run++= ) { + int rem =3D run % (nr_occupancies * nr_pressures); + const char *mode =3D modes[run / (nr_occupancies * nr_pressures)]; + bool holes =3D occupancies[rem / nr_pressures]; + bool pressure =3D pressures[rem % nr_pressures]; + + thread_mask =3D base_mask; + if (!pressure) + thread_mask &=3D ~(pageout_bit | compactor_bit); + else if (!swap_available()) + thread_mask &=3D ~pageout_bit; =20 thp_read_settings(&settings); settings.thp_enabled =3D THP_MADVISE; @@ -373,6 +465,24 @@ int main(int argc, char **argv) ksft_exit_fail_perror("mmap() playground"); mremap_area =3D region + nr_shared_areas * hpage_pmd_size; =20 + if (thread_mask & pageout_bit) { + /* + * Big enough to cycle real reclaim, small enough not + * to dominate a TCG guest: 4 PMD areas, clamped to + * [16M, 64M]. + */ + pageout_size =3D 4 * hpage_pmd_size; + if (pageout_size < 16UL << 20) + pageout_size =3D 16UL << 20; + if (pageout_size > 64UL << 20) + pageout_size =3D 64UL << 20; + pageout_area =3D mmap(NULL, pageout_size, + PROT_READ | PROT_WRITE, + MAP_ANONYMOUS | MAP_PRIVATE, -1, 0); + if (pageout_area =3D=3D MAP_FAILED) + ksft_exit_fail_perror("mmap() pageout area"); + } + /* Populate so the first pass has something to collapse. */ for (i =3D 0; i < nr_shared_areas * hpage_pmd_size / page_size; i++) *(unsigned int *)(region + i * page_size) =3D pattern(i); @@ -423,8 +533,9 @@ int main(int argc, char **argv) check_page(i); =20 ksft_test_result(!corrupted, - "%s/%s: %ds, %d steps, no corruption\n", + "%s/%s%s: %ds, %d steps, no corruption\n", mode, holes ? "holes" : "strict", + pressure ? "/pressure" : "", duration_s, steps); =20 /* @@ -433,17 +544,26 @@ int main(int argc, char **argv) * and its scan cadence differs. */ munmap(region, nr_areas * hpage_pmd_size); + if (pageout_area) { + munmap(pageout_area, pageout_size); + pageout_area =3D NULL; + } thp_pop_settings(); stop =3D 0; steps =3D 0; =20 if (corrupted) { /* Memory is suspect; the rest would prove nothing. */ - while (++mn < nr_modes * nr_occupancies) - ksft_test_result_skip("%s/%s: skipped after corruption\n", - modes[mn / nr_occupancies], - occupancies[mn % nr_occupancies] ? - "holes" : "strict"); + while (++run < nr_modes * nr_occupancies * nr_pressures) { + rem =3D run % (nr_occupancies * nr_pressures); + + ksft_test_result_skip("%s/%s%s: skipped after corruption\n", + modes[run / (nr_occupancies * nr_pressures)], + occupancies[rem / nr_pressures] ? + "holes" : "strict", + pressures[rem % nr_pressures] ? + "/pressure" : ""); + } break; } } --=20 2.54.0 From nobody Mon Sep 28 23:56:19 2026 Received: from fhigh-a7-smtp.messagingengine.com (fhigh-a7-smtp.messagingengine.com [103.168.172.158]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 537D336AB7C; Sat, 15 Aug 2026 01:59:42 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.158 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759184; cv=none; b=ARCZ0mQ1UCxh6Gg7JxZup/ps7ZPB2O+YAvQ37QL8rlW8brGX4HPAsMN/7ZuQziPvmiVOfaCc+ly9L4ROqLVem1vp0K3C44xC2rAGQ9bfpJIyr6gZ1BgZQW5R5aih1cIPFOj5hwXt41GSed7RgnMVuSPOhtAMe5sbcaexb1y7auk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786759184; c=relaxed/simple; bh=X9BD6U9xfu679WoD7+9b9ZF4FLOqM9dKkuaiTKZIalM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=rfGgmELNq4fbrnrPpDk5iU16TujPkH9Yqqx7BY66Izacmdd7D3s8ra1cDtiN6amGcOirxyade9JNPFJe9wwZAVDGmzrTnI/0a/OHk1MnPxzBs0etrYkY1D12hiLMBM7WVavoFxf7fL9a7Pzw5fygJmJJzOX6vhy8UiaN4lOQ9Ew= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=Ut/lCJEw; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=JhwiSSGG; arc=none smtp.client-ip=103.168.172.158 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="Ut/lCJEw"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="JhwiSSGG" Received: from phl-compute-01.internal (phl-compute-01.internal [10.202.2.41]) by mailfhigh.phl.internal (Postfix) with ESMTP id 72AD5140011E; Fri, 14 Aug 2026 21:59:41 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-01.internal (MEProxy); Fri, 14 Aug 2026 21:59:41 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786759181; x= 1786845581; bh=S0c1eOgbbGAyxsvdncO2s21v5pX/QpSQtyc3bthgrHY=; b=U t/lCJEwc5aQKPhgw7EwTAdTo+PRmF1IqeIOoZHi9eWxt471XuI22OZtQK6rvfCzg PyqtP2xJXjnttcwAHHPTXq6RWBbl1TILSyEPlOKqqlfbqhlbW4LBQyziuNP3aQYc aKIrMrBrGOdfYk+Nl2ekV8oJNdw5MbAAQtCiOWjRg4zX2kJa6hiMe1Tes3mvOw/7 VAqfAQbUZ96Ltjk+aZWT3EGpyc2NFU5ZdBHi82avmgjDy3KaYWLSsE3q/2QbI0fz Iq4d54umBUGqaK0sAfBzpsVxVlL78nvkRCAHeXcQiiFkOPz2M8A6avRgIRFwQq4T oso0GHW6fyv9rYa+r7WYQ== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786759181; x=1786845581; bh=S 0c1eOgbbGAyxsvdncO2s21v5pX/QpSQtyc3bthgrHY=; b=JhwiSSGG9GcvcRpt2 StUvvTqekt2dNbT/MEipdTfJWgSKaGnX8aavOol8DZSxj0N/b+OQqhO+bR6/wcRQ U2EyqxQZdIiQdPGZ2ivmhTmS5ZlWs7xBqRJG/EFbMN/UtZBZxGMkemYuNcfP5spt 2BHM0nsRpAn3Xy3iJVVTlkpk2Izn4qhFrjZQIfUjWRR2Z/SKQodDpB7KKrwySSPj KEGLALAreBOPdw3tIUmhQfzDHFf1zsuAs3LpMjGQUMaMV5nmlBJgxIsDqh1ClrFi YCw3g5OzfuT148DEXOx0g1SLwA7dMzbw+MUgrS9rTmJKJWI/ZRx4nxaZtco2mErn sNtpA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEYpEQiBdLKlCaLp317TOYqYwO99jnuX4YrBujo3RuW+IIGXdYj4qE1ezcGRvbtOZ hc7RfnfkdlTehb5Dkvk4VZTcyRAqm2CcEjaldkIQshIttrB2HSeWyAKQfsR+ykmymFv8qO aC/+wh3joJiQccaWn5NjVrqLVVMYdoVO37A5ZL6SnR75seJlCquk1NTOiA0EvXHIVSJ25Q iFONSccd/wEzNfQZlOvvQQf5bCKiGIALXd0Rd37Vcnrh56q8DfWwmvEu2ZgBGSMa525iaP HkMF5MKO3LDiNiHXJVHjo5f/L4XJttiAOfdXhljkaHgZI649okv5M81/eeF8HfhCZr1UOF hEaMyy8SNp+WA12XXrvIqlcMW2sZE0kgpE/SOwwrkjJJR1inJnCtzQ+93UM2eF1fnod6oT 9r9RUE9IaQWym0pbVGY3R1/tSevMLn2Wm1fDSecSv3O3MsiiHeeD71o9x7zq70jW1XGMxt 6PxXclH4l5ckE/VqaZuoVIPwRETu0ui1oWDy0ZOTdrEEgVG8tzZNDoMGrwKfGHkXxWDatX F7XN4ZYzRnP0ghZSWu5SVUOAinq0yjUcgWhC9vY4+NzBxMhAdOUGFaUimv51f4+MtooxKw LlXk8IkRDFUfG5OOf3LKOP1EhgQRur3r/DB1mhiHKKozaXnvIDhtuZIUDTng X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 14 Aug 2026 21:59:40 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v4 19/19] selftests/mm: zap whole PTE tables in the khugepaged race harness Date: Sat, 15 Aug 2026 02:59:01 +0100 Message-ID: <20260815015901.1236937-20-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260815015901.1236937-1-kirill@shutemov.name> References: <20260815015901.1236937-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The harness's MADV_DONTNEED thread zaps 1 to 32 pages at a time, never a whole PMD-aligned area, and only a zap that covers a full table frees the table itself (CONFIG_PT_RECLAIM). Make the thread zap a whole PMD-aligned area once every 64 iterations, and keep the fine-grained zaps as the common case. The new case frees page tables, racing that against a collapse walking the same table. Assisted-by: Claude-Code:claude-opus-5 Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/khugepaged_race.c | 20 ++++++++++++++++++-- 1 file changed, 18 insertions(+), 2 deletions(-) diff --git a/tools/testing/selftests/mm/khugepaged_race.c b/tools/testing/s= elftests/mm/khugepaged_race.c index a23e5bfe78af..6682bbae0a8f 100644 --- a/tools/testing/selftests/mm/khugepaged_race.c +++ b/tools/testing/selftests/mm/khugepaged_race.c @@ -140,8 +140,24 @@ static void *dontneed_fn(void *arg) unsigned long page_idx =3D rand_page(&seed); unsigned long nr =3D 1UL << (rand_r(&seed) % 6); /* 1..32 pages */ =20 - madvise(region + page_idx * page_size, - room_from(page_idx, nr) * page_size, MADV_DONTNEED); + /* + * Once in a while zap a whole PMD-aligned area: only a zap + * spanning the full table triggers the empty-table reclaim + * (CONFIG_PT_RECLAIM), which can free the table under a + * collapse that is midway through it. Sub-table zaps never + * reach that path. + */ + if (!(rand_r(&seed) % 64)) { + unsigned long area =3D page_idx / + (hpage_pmd_size / page_size); + + madvise(region + area * hpage_pmd_size, + hpage_pmd_size, MADV_DONTNEED); + } else { + madvise(region + page_idx * page_size, + room_from(page_idx, nr) * page_size, + MADV_DONTNEED); + } usleep(rand_r(&seed) % 500); } return NULL; --=20 2.54.0