From nobody Tue Sep 29 13:20:38 2026 Received: from fout-b5-smtp.messagingengine.com (fout-b5-smtp.messagingengine.com [202.12.124.148]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 81CB4357CEC; Fri, 7 Aug 2026 11:36:55 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.148 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102618; cv=none; b=OvGFzwkNHIoS90V8dUt6zMKel2l8Wwo4UCg43pHUzzNnmfBu7oj53FuQEQH/ZD14Rvoyctq0rAMYIh9RHLprEEDaZv5k3EKZEXcyhiSzCEv4z+tN0FgS2RJcRMh7sEjyTu8j1UO+OH/PdmHR5CR0xe6ObzxablQthkqgkjPcppg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102618; c=relaxed/simple; bh=4sgHzMgtwbejiXl7UU6IlQx6IBQ8o/Y1jLmdT+jHz0Y=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=imKmkvbfvNbdPja1lG5KzYSQYoJxr1h2RiAW1I+ZnevnuuWMRkVG3ywGl8nv+2jUbaER2uHllhjdSY/KQGoLgijvxnkzn2IexGK9Klq7KDwWJjSH0pgNvj5w2ROeDVVCtQacjXLpCLL0Bbk7nV5Z9ipUzFQUuKDypH/DCRx+fnI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=oHsHlbcP; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=FK1/zSah; arc=none smtp.client-ip=202.12.124.148 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="oHsHlbcP"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="FK1/zSah" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfout.stl.internal (Postfix) with ESMTP id A4FE21D0008B; Fri, 7 Aug 2026 07:36:53 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-06.internal (MEProxy); Fri, 07 Aug 2026 07:36:54 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786102613; x= 1786189013; bh=d+ra4YCar9dQSq6xe2qyml8d0zrOh7Pky3vMj7voUQ8=; b=o HsHlbcPm4TPT5V2t2qeaIXKgumuPdLRSljW5ZyuFUeNMRRNiJ5TRNC+51iDmTJZv myUE7guHJp5bakjQZ0gkjsmJukt5wBAPA5DD5rc4WRjXP+s1320hrkhkjf37FZSV o3TIFp2CMr9A/0F7rALfI4KaP9qqgVoLCHDttTN/8BmFQ8nI/jXr1q06O4NMeYEl BAAk2vC6iI+Dv4y8bGcw2tUytncyjzAjEOvMGVLLw+hjIQKjE6pkQvK0WamFhy4w znfg9A3E3khUQLKrWgV7qcnK+izqGIwb4oP9kiftchu9bi9e3W8rHLAmP3bHh+Ci Gylxcjd9HGhHdWiuwZxeQ== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786102613; x=1786189013; bh=d +ra4YCar9dQSq6xe2qyml8d0zrOh7Pky3vMj7voUQ8=; b=FK1/zSahaOB/v1MOw 9AdAbCYB1I9xNiJmn25A9Je6ahiyNkuDaXpOLta9nl6sidp+t2eQblEwZX8YymV5 fRciC8/FLsCQ0nlD5njxttrcFATJZh5p40DuigFktcRJNkBS04c7mjNzk8fxjrMf EGUDJ7DW8xNftDl3z9ejZbI0wxThaTGMfSIiyYNPWEs0i216b6/C9iD2+aP9siSv PzdfhNy5dCrbPttPVDNHcUeH3KYKsVry6wQXwskN/8bZslo3+Zu9YXwY4F7liIww dXfL4Cu9Gy4i3yFjSlnUuCk1bpRyJifg0iegWF6YLjuH+cEjUOeDlyPCa8PqDxql AIQ2Q== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEfW3x2Qrq7ELqxlr/x69vk0aYaa/E/2AWrFAl7yu0Vg+jIGlS0mjOCk5J5e5/NJH T3J+l8M+cQ15gxg/u6FQhwxKudsrf7tN5cWcEThveGDS0KcOU0uMrzGx0YBmylEV4PfT2X SbTqClZY3SINVohR/2NzYRxeOu5tU+SgKqY1sS5oM8U15+ky277S+G1Ww7YAUM5Z6EK46p TQK1COWwyEQl/dhwe51H5lWCLFzTj8K9hkk/JBO/DvOgEgoe5sMihzFP04eKTIrZCF8dte tQwhoNU+zszgTlPncqprVsvTwAgpjyPciGBDZXKQScTz4u1orAuTZg8Fi2VomZsbF2qQad Lkl9l51BZdRN70Hk63ufcfU6wqwALAhmqOWouzSC7gOij5TvqfnKy6ADG69lfLUgh0lpiZ kAUMRLbDr7XMEEcGo44022nvTc8uJNyhrilK5mNWJO3DsLD2SyRs7E7w7dGrX0w6ovYG6B wSdpqjibCk8XQNAn03/7eeAntRw7Jwdq3BxVqzuXY/1OJugBvTCFsmhWssqxoMw/1yXSUs dGRfegmLUWJRDkapVnaicCMpKihTv6gkRRtJStRI6Lxcfxfvo6ULDeXqPAprLZku+PYE6A ojgmJhnV9GF9r7CxueRcqT2zEmqwX7MIi7lOCAc0ZCKaDOxqgp5Uizip9y1g X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 7 Aug 2026 07:36:53 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v2 01/17] selftests/mm: skip collapse_compound_extreme where the PMD is too large Date: Fri, 7 Aug 2026 12:36:31 +0100 Message-ID: <20260807113647.3744609-2-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260807113647.3744609-1-kirill@shutemov.name> References: <20260807113647.3744609-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" collapse_compound_extreme builds a PTE table full of distinct PTE-mapped compound pages by cycling hpage_pmd_nr fault-time THPs through mremap. It therefore needs hpage_pmd_nr PMD-order allocations in a row. That is fine at a 2M PMD (4K base pages) or a 32M one (16K), but a 512M PMD -- arm64 with 64K base pages -- makes each of those an order-13 allocation, which the allocator cannot reliably hand out even once, let alone 8192 times. The failure is not a quiet one: the case calls ksft_exit_fail_msg(), so the whole binary stops and every case after it is lost. Skip the case where the PMD is larger than 32M. The MADV_COLLAPSE cases still cover PMD-order collapse on those configurations, and 4K and 16K PMDs are unaffected. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) Tested-by: Muhammad Usama Anjum --- tools/testing/selftests/mm/khugepaged.c | 10 ++++++++++ 1 file changed, 10 insertions(+) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index 9a90ff5f484e..51fc1f04d0a0 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -1007,6 +1007,16 @@ static void collapse_compound_extreme(struct collaps= e_context *c, struct mem_ops void *p; int i; =20 + /* + * The test needs hpage_pmd_nr PMD-order allocations, which is likely to + * fail for large PMD sizes. Skip if the PMD size is over 32M. + */ + if (hpage_pmd_size > (32UL << 20)) { + ksft_test_result_skip("%s: PMD too large for fault-time THP construction= \n", + __func__); + return; + } + p =3D ops->setup_area(1); ksft_print_msg("Construct PTE page table full of different PTE-mapped com= pound pages\n"); for (i =3D 0; i < hpage_pmd_nr; i++) { --=20 2.54.0 From nobody Tue Sep 29 13:20:38 2026 Received: from fhigh-b6-smtp.messagingengine.com (fhigh-b6-smtp.messagingengine.com [202.12.124.157]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 04D9139EF25; Fri, 7 Aug 2026 11:36:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.157 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102621; cv=none; b=V/6lgc4ZrQ/kJD4yLwvXIfvwbVfw5wJQPlWDFg71Ld6nAec7xynVNI607GY95/OqrbjLJQNtWOabVenh+/IxK4+H3VjaXbm0vnne453odUGlB/zj/gHySoOilHlEFPBUun6zp8/uwykLnDNx/ajMP4/V6Hibu6yZjEhj7mW8Uwg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102621; c=relaxed/simple; bh=ZKss5qEILIcku0QFzZjkA/LCUL/jkVDVD1M174XIDfQ=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=sFqhSamMPIx1T/GrffRr4h//+ymeMSdjL2QSi7GJ0f1xV4zVIdep811AzNvvJpzHwMZgmro7UkW0PSVOiMOyie+r7cLvPsCltto+EKE4+RUj8puagBdSR2Ez6mEQ/CEdq6n6IJTfdObEBbXoFx/LfN64X885J5rBEAqp1JXT0nY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=wRFTL0rL; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=G9yXMvl1; arc=none smtp.client-ip=202.12.124.157 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="wRFTL0rL"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="G9yXMvl1" Received: from phl-compute-09.internal (phl-compute-09.internal [10.202.2.49]) by mailfhigh.stl.internal (Postfix) with ESMTP id 638D47A004F; Fri, 7 Aug 2026 07:36:56 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-09.internal (MEProxy); Fri, 07 Aug 2026 07:36:57 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786102616; x= 1786189016; bh=cIkdUEAwCGtCQWq2Mp8sW0xlTg8F6TBRnofJ/Fu60Hg=; b=w RFTL0rLMg4rHvQy8CWUQTbL2oYxVyCUHhwNRnLoht0C1AO7XQk4knl4UYUOTs895 NCaUt70pAyVx+AO9jT5WzXPHmBKxdZb2sx7CXeQSi7GqWcI+YnnkPymz3gpbQDLT OazpJFPIGINitOehdRW1p3xiXQd/mS7FZy3qvn20cPF0HHPNqA6s6aA6xgIPgU9m 7IxncNPZlHIsQyImxY12xKFEXo7zWbuOMUJYl+2T2brGgubvIz4v1EzI7P/kWezA eHkKpVBOeiMl8+6TC3dPxxA++4egmUWSIaWSev9uQJGjL+pvjE0wZ+ONNNNNIRTK TbVVCiK8ZI/l5RKEhH5zw== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786102616; x=1786189016; bh=c IkdUEAwCGtCQWq2Mp8sW0xlTg8F6TBRnofJ/Fu60Hg=; b=G9yXMvl15P/EqLODP h29CXXo6hMefBq7R0Vhm+RJpa0l0UNQVFna/dDXhn5vjylMS83MJYk5he0uAeZkg 4XAq97piv/yjBwsIpUUHJItcCi6CSuJBilQzlqmjAM4M5uLLSdzAfxsLoqy+yEsQ y8inpySpQnxI5aH4/eZjPxDrXyc5W/tFlxIlWWQcrt9cDpWiEL1DQBi8zQCGVanb NjjlKhVlpyqy0M8+CBBzWMWos1l3uAq/KnmDkIgRPRtmWkvyvxEPNaBk7i5EpaZO mFtF0q7A5z9vDMXvQeEDNsPoLLbmhFbDo/0NtsTrMV9kBQfflflRWQQoFSDMTUjQ 9BaXA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEfW3x2Qrq7ELqxlr/x69vk0aYaa/E/2AWrFAl7yu0Vg+jIGlS0mjOCk5J5e5/NJH T3J+l8M+cQ15gxg/u6FQhwxKudsrf7tN5cWcEThveGDS0KcOU0uMrzGx0YBmylEV4PfT2X SbTqClZY3SINVohR/2NzYRxeOu5tU+SgKqY1sS5oM8U15+ky277S+G1Ww7YAUM5Z6EK46p TQK1COWwyEQl/dhwe51H5lWCLFzTj8K9hkk/JBO/DvOgEgoe5sMihzFP04eKTIrZCF8dte tQwhoNU+zszgTlPncqprVsvTwAgpjyPciGBDZXKQScTz4u1orAuTZg8Fi2VomZsbF2qQAb k5JRNRlEKLgml6X7jbsTp0jsLf6g6Zd60MuoeVfy+Ysl9o2kZSA454ZGpdH/QaN/g9nDo7 Lp0E+Yj9TpHDuTjBVlhsxP8CZiKBTsCPvulUuJTHqivMavkWvjquz4yvRe86vfQpe23O4r 1YFtj5pDmVh+VEGw3bu3Yog3PVDXHZzc3fvQQ6UPbSlg3BvA2YgbQnlcHDtOWt4B3F4Fla HuGj+NlZ8rQdbbE1TBNrzJDV2TzID8bIeJ1Pbxpd+JvxADuRnrA77ZhfcLIvYV2PfbMZyC QoVrNtpUIdiF3f1AWD8fIQLDdou5dNdL02HdxcWEF/EHF6S93onH74yVe8xQ X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 7 Aug 2026 07:36:55 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v2 02/17] selftests/mm: scale khugepaged's collapse wait with the PMD size Date: Fri, 7 Aug 2026 12:36:32 +0100 Message-ID: <20260807113647.3744609-3-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260807113647.3744609-1-kirill@shutemov.name> References: <20260807113647.3744609-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" wait_for_scan() gives every case the same three seconds, whatever the huge page costs to build. collapse_full asks for four of them: 8M at a 2M PMD, but 2G at a 512M PMD -- arm64 with 64K base pages. Three seconds is thin at that size rather than generous. Across 80 runs of collapse_full on arm64 with 64K pages the wait was half a second in 73 of them, with a tail to two seconds, and the case has timed out in a full matrix run, reporting a failure for a collapse that was still going. Keep three seconds as the floor and add a second per 128M to collapse. A 2M PMD is unchanged. A 512M PMD gets 19 seconds, which is headroom over the observed tail rather than a measured requirement. The budget bounds how long a real failure takes to report, not how long a passing case waits: wait_for_scan() returns as soon as the collapse turns up. arm64/64K: khugepaged all:anon 21 pass/1 fail -> 22 pass/0 fail. x86-64 is unchanged. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) Tested-by: Muhammad Usama Anjum --- tools/testing/selftests/mm/khugepaged.c | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index 51fc1f04d0a0..0f828bfee31f 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -602,8 +602,10 @@ static bool wait_for_scan(const char *msg, char *p, si= ze_t len, int nr_hpages, int collap_order, struct mem_ops *ops) { unsigned long hpage_size =3D page_size << collap_order; + /* Three seconds as a floor, plus a second per 128M to collapse */ + const unsigned long bytes =3D (unsigned long)nr_hpages * hpage_size; + int timeout =3D 6 + 2 * (bytes / (128UL << 20)); int full_scans; - int timeout =3D 6; /* 3 seconds */ =20 /* Sanity check */ if (!ops->check_huge(p, len, 0, hpage_size)) --=20 2.54.0 From nobody Tue Sep 29 13:20:38 2026 Received: from fhigh-b6-smtp.messagingengine.com (fhigh-b6-smtp.messagingengine.com [202.12.124.157]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CE8E6471250; Fri, 7 Aug 2026 11:37:01 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.157 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102625; cv=none; b=njGz04mkUMiAfrL8i2S0qKOLHmbkFlw+lIcfSfLuC211iKJNMv2IFu95AIDQYgs//FyIOjUtrDaXURXTTSdNbarCaJhsQAjHD+asRzKGMezJYS3ZRf0ia4phMqZELN0HQo0/SVLd7qc6epv7EZwvbjbC1ZI1KTu5Wah19PoRGA4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102625; c=relaxed/simple; bh=QByATvCRuQCNL2ZnbBirX2salSuSQ+wbWK8XCWZ5bbw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=W7iGQqjWClImtx4g+1ikHmOmvDIKUM3iFuzfkzuHMmT5MYwak4nfhZAM9/GOQ4FfAc5Ba0wVhFWAmg05ZqJJSABRo8CaWvgeed1qr7PpHOk3rsoFx8sixLFyIA5/WimVLHNgNOrbYV+iR9sBwK1FL5dyVC5ysUTvgjVIs4R4D4k= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=h5zdipvp; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=gssPUSgF; arc=none smtp.client-ip=202.12.124.157 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="h5zdipvp"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="gssPUSgF" Received: from phl-compute-10.internal (phl-compute-10.internal [10.202.2.50]) by mailfhigh.stl.internal (Postfix) with ESMTP id 0D1587A00C4; Fri, 7 Aug 2026 07:36:59 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-10.internal (MEProxy); Fri, 07 Aug 2026 07:36:59 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786102618; x= 1786189018; bh=zfbh3GHFEDkfsexwK0f+yG1zta9ZtoxopU3TBc0ckuA=; b=h 5zdipvpvC02jD1Mf8nIqPvjVbTcuzSYiXHGXFxb+7el/Svcri/ADkztxbLz5GWL3 KqfGfeF7oWtK3HIa2iebgzMTmX8hVd2q1d1vWi52xdzrdFR45O4HcAAQh05g/c9o 2qOukZfHTa4OiaUzl86d/xnEBJ7QE4/gaQYEeUj+hoKYH4KKQoLm13HkE49dT540 g7zq+5u+63wrlJttq/T0ie8rKfBGbfg0JcKHnsw7B/x7j/4peSbAQd2iJ8knd+Xy u8erj9JSdUMy2gl5r71rsM72CNrwN8lF1YP47VYsW/CTl8qVeieJDWKwMpR+U93/ 4evmsh1zumRZw9vjclrpA== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786102618; x=1786189018; bh=z fbh3GHFEDkfsexwK0f+yG1zta9ZtoxopU3TBc0ckuA=; b=gssPUSgFT+Q7INiV2 tQ1/YcEp833EIqd+keLsa6xOPmgW0gQlR4bZPJUnvaoOtQm8kJHA/v4/3YL1Jm6L YwZvYMF9sMlMok7tgJT8XrMUP+3oDe+1veAwjdAiR/PN+RKIgtHAr8t9IlUhJcQl c46NfKWL3XuCCdbbYf16kmkSV3u2g4ZTtFYt4bRd6wonYMfLs9QrNWYjhtwTDDiO dS7+Q3arOWlkGrjB30NABNDSFtB9BH+uAm4i/+XZ3TN99dOOfsaxCrM//85El/XB L8/SiYsVD5lFT6xDPZ47cgl1vu3mh4T02YGFvEJnKEvro8Rgm4/7i3Daa8y9ifJE 2z86w== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEfW3x2Qrq7ELqxlr/x69vk0aYaa/E/2AWrFAl7yu0Vg+jIGlS0mjOCk5J5e5/NJH T3J+l8M+cQ15gxg/u6FQhwxKudsrf7tN5cWcEThveGDS0KcOU0uMrzGx0YBmylEV4PfT2X SbTqClZY3SINVohR/2NzYRxeOu5tU+SgKqY1sS5oM8U15+ky277S+G1Ww7YAUM5Z6EK46p TQK1COWwyEQl/dhwe51H5lWCLFzTj8K9hkk/JBO/DvOgEgoe5sMihzFP04eKTIrZCF8dte tQwhoNU+zszgTlPncqprVsvTwAgpjyPciGBDZXKQScTz4u1orAuTZg8Fi2VomZsbF2qQaQ WX9kdIRFJUsNCeKzkNq4ymbXB/KQlMUzSHaY6cbAlDEBIXeOB8GtX5Jz9IxAs98YfmqVys qUOg8y01tLyLpPI4SZiwG0rtop3WKpCDrjgtoKCBxJBbeU2IEMYlxdlFMk+ayLhL+ZdCO/ Y0kh9rPsL4eHcdlVUWt+xNYtPyuayojfxejvL8/ENVfvKva7UlexDsqTLfQFmefHNhwFLC NeoOEpJKUI6VIxeoPeAmv8h5L4IhZEzFBGCicLsjv4uRMpUGftkrAuJrOk0BYmnd2IBnr5 AeBmqb+s6xoyvl+j5St0J+JNuAsOzEMoing/A1ja3RJXzjvRKr2LdXQOvUpw X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 7 Aug 2026 07:36:58 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v2 03/17] selftests/mm: skip khugepaged page cache cases without a PMD folio Date: Fri, 7 Aug 2026 12:36:33 +0100 Message-ID: <20260807113647.3744609-4-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260807113647.3744609-1-kirill@shutemov.name> References: <20260807113647.3744609-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The page cache caps folio order at MAX_PAGECACHE_ORDER, which sits below the PMD order where a PMD is 512M -- arm64 with 64K base pages. A PMD-sized page cache folio is then impossible, so MADV_COLLAPSE answers -EINVAL and khugepaged passes over the range. The shmem cases ask for one anyway, so four of them fail and the run bails out in the middle. The cap is one global, so it rules out every file mapping, not just shmem: shmem_huge_global_enabled() drops the PMD order from what it allows, and file_thp_enabled() refuses a regular file whose mapping cannot hold a PMD folio. Skip both mem types when the PMD order has no per-order shmem_enabled control, which is the readable form of the cap: that control is created for the orders in THP_ORDERS_ALL_FILE_DEFAULT. A run left with nothing to collapse into skips outright. Anonymous collapse is unaffected: its orders are not capped this way. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) Tested-by: Muhammad Usama Anjum --- tools/testing/selftests/mm/khugepaged.c | 31 +++++++++++++++++++++++++ 1 file changed, 31 insertions(+) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index 0f828bfee31f..9b1da0dd0f8e 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -1424,6 +1424,37 @@ int main(int argc, char **argv) =20 setbuf(stdout, NULL); =20 + /* + * The page cache caps folio order at MAX_PAGECACHE_ORDER, which + * xas_split_alloc() puts below the PMD order on arm64 with 64K pages. + * A PMD-sized page cache folio is then impossible, so the kernel + * refuses these collapses by design and there is nothing to test. + * + * The cap is one global, so it rules out every file mapping, not just + * shmem: shmem_huge_global_enabled() drops the PMD order from what it + * allows, and file_thp_enabled() refuses a regular file whose mapping + * cannot hold a PMD folio. + * + * The per-order shmem_enabled control below is what makes the cap + * readable: it is created for the orders in THP_ORDERS_ALL_FILE_DEFAULT, + * which is the cap and nothing else, so whether the PMD order has one + * answers for a regular file as much as for shmem. + */ + if (!(thp_shmem_supported_orders() & (1UL << hpage_pmd_order))) { + if (shmem_ops) { + ksft_print_msg("no PMD-order page cache folio: skipping shmem\n"); + shmem_ops =3D NULL; + } + if (read_only_file_ops) { + ksft_print_msg("no PMD-order page cache folio: skipping file\n"); + read_only_file_ops =3D NULL; + read_write_file_read_ops =3D NULL; + read_write_file_write_ops =3D NULL; + } + if (!anon_ops && !shmem_ops && !read_only_file_ops) + ksft_exit_skip("Nothing left to collapse into\n"); + } + default_settings.khugepaged.max_ptes_none =3D hpage_pmd_nr - 1; default_settings.khugepaged.max_ptes_swap =3D hpage_pmd_nr / 8; default_settings.khugepaged.max_ptes_shared =3D hpage_pmd_nr / 2; --=20 2.54.0 From nobody Tue Sep 29 13:20:38 2026 Received: from fhigh-b6-smtp.messagingengine.com (fhigh-b6-smtp.messagingengine.com [202.12.124.157]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B9F3C470440; Fri, 7 Aug 2026 11:37:03 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.157 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102626; cv=none; b=aPwwo9PQFvSaj2wk4z0sr+3cf5z4GGCcFtU9xcYgvmWlMjdunuNL7iiJ/naQKgS3bOomO1J1Et5GXdRWAh170DPa1JWCmTGVbEZ63CtNr6AIZ9KTMrCEVDDaqlt5grJEkQo2JIaHoZ4byUIPlLCKRtFrvJIsijkBUO3hiG7VkBM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102626; c=relaxed/simple; bh=VTtq9viQF+otnO5yiU7JAUFQCN2Xtr9xJfUI7Jc1JIc=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=bFU/AaelOmsooDASSwsI9k+fNUpvhbf2MSTPBaN04zgWegHy6lORBkEo71u7XNlgBJqAI0YmfEMrt94+ynSReLsoisfPKm3t+OhqsN+unk+yaW+WuP/J1hg6iNBY2hhEwnMtnBYDe0Zl8FblIWU00MRaisOj2llZCDwg3mSwDIw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=uPOnI7Rk; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=bPeb5TJt; arc=none smtp.client-ip=202.12.124.157 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="uPOnI7Rk"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="bPeb5TJt" Received: from phl-compute-04.internal (phl-compute-04.internal [10.202.2.44]) by mailfhigh.stl.internal (Postfix) with ESMTP id AB9D07A0120; Fri, 7 Aug 2026 07:37:01 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-04.internal (MEProxy); Fri, 07 Aug 2026 07:37:02 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786102621; x= 1786189021; bh=JUc0tI3FwUfzqFn1GnkruP1SC6LnWC+xRZZtn+0X3kI=; b=u POnI7Rkap0jtzxCvqzP2YGcxkmZK+a/4axyilNMdTDvIcESS1Y2fPCsL+Prjfqb5 NrcEy0Ufy2kVDlhavodIHVvz2Mn5DP0jkowDH04+3Spnxeol4U3HpEu6UrwDD2b3 FRCi0ILFLGYVCvBSVP8AvKY0KWcrJjifkEfgRhKe8q1MmjmYwjnhx60guPiOHsHH 0qLBmy3DWvImXH8Z1x4CmlAyHDn8Mds36rP4b2ikudSphU1pky8BR4kTVeDLnvA3 AW/XR4tElu2w8mrvyPODUbu6XP9BA9j7Pwd2kM6OWOYsBuKG9b+/ZV/+tKZcWS3u IA11TA5j1QFC9eBeqsZ7w== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786102621; x=1786189021; bh=J Uc0tI3FwUfzqFn1GnkruP1SC6LnWC+xRZZtn+0X3kI=; b=bPeb5TJt0Zq+0BPTM R7G/KrPUR1OO54THEGlZAFS49jMnYOV/zYURrI/PVdi/nQEfE5PpS+6x/RUsHxin aqC+ex4XEnWGP4yiocvfK8BwU4tTCbJlTp9vqIj/cD9kY0JNj0KGMTtQ8PgfR18s srjMzQOAsuPY7yS9ZJjS14hOf2JhTNs73024TUsZjhhiK34By2DrSxqn2tohY3kH OYCOJ5iUt9gO0ARZ/OXVy2e1GMft4ASSjG03tf9GbWOCHpu+wf126okY/vBSoLpq rQWI6HCCvV0YSXFQHHmIbrAl1wqrEPwpK7vrW0kof55QHF/K1DSOTcmzlqSaF3EV vKhFQ== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEfW3x2Qrq7ELqxlr/x69vk0aYaa/E/2AWrFAl7yu0Vg+jIGlS0mjOCk5J5e5/NJH T3J+l8M+cQ15gxg/u6FQhwxKudsrf7tN5cWcEThveGDS0KcOU0uMrzGx0YBmylEV4PfT2X SbTqClZY3SINVohR/2NzYRxeOu5tU+SgKqY1sS5oM8U15+ky277S+G1Ww7YAUM5Z6EK46p TQK1COWwyEQl/dhwe51H5lWCLFzTj8K9hkk/JBO/DvOgEgoe5sMihzFP04eKTIrZCF8dte tQwhoNU+zszgTlPncqprVsvTwAgpjyPciGBDZXKQScTz4u1orAuTZg8Fi2VomZsbF2qQsR eL1kaiHBpp80M4RzDUsAZRJygVAIesBLxby3W54tcHXPJqAPjdtHEFumnKAwb74kc0+J3n KlS5iazk/Ooe6XFvlIqXdqyrpLbfIPetzbxMD3rM8ypA4CDMjbgME1dXJFHpFXaauetX3L dh8aaJt0hh5AXjzk1uAO2icErMYCxtpqxWilkvL508lIKOkNHC29LP6kDYhGogrbbnZu/K rjMBlZ7WOIsjIAt+8knDO1/FrYhkE2E37Z0jou3o9BEH9zLe9kGb6x90gjYFmd/JtJtTbu eT2eET2AvAWQVqUvwVJG9CwXzvl7U10CCbLigNt/K/4Q64lGhESar78BxLJQ X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 7 Aug 2026 07:37:00 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v2 04/17] selftests/mm: retry the swapout the khugepaged swap cases rely on Date: Fri, 7 Aug 2026 12:36:34 +0100 Message-ID: <20260807113647.3744609-5-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260807113647.3744609-1-kirill@shutemov.name> References: <20260807113647.3744609-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" collapse_max_ptes_swap() and collapse_swapin_single_pte() set their precondition with MADV_PAGEOUT and then require smaps to report exactly the range they asked for. MADV_PAGEOUT is best effort: shrink_folio_list() leaves a folio alone when it cannot reclaim it right away, and a folio still under writeback from an earlier pageout is the common case. The count comes up short by a page or two, and the case fails on its precondition without testing anything. It shows on arm64 with 64K pages, where max_ptes_swap is 1024 pages: 64M of swap per step, and the second step pages out folios the first step has only just written back. About one run in four ends in # Swapout 1024 of 8192 pages... Fail not ok 95 collapse_max_ptes_swap A probe around the same two steps hits it on roughly one attempt in twenty, and asking again 50ms later closes the gap within four tries. Ask again for up to two seconds before reporting failure. The wait costs nothing when the first attempt is enough, which is the usual case. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) Tested-by: Muhammad Usama Anjum --- tools/testing/selftests/mm/khugepaged.c | 32 ++++++++++++++++++------- 1 file changed, 23 insertions(+), 9 deletions(-) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index 9b1da0dd0f8e..69e0fc4b12fb 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -241,6 +241,26 @@ static bool check_swap(void *addr, unsigned long size) return swap; } =20 +/* + * MADV_PAGEOUT is best effort: shrink_folio_list() leaves a folio alone + * when it cannot reclaim it right away, and one still under writeback from + * an earlier pageout is the common case. The swap count the caller asks + * for then arrives a moment later, so ask again before giving up. + */ +static bool swapout_range(void *p, unsigned long size) +{ + int i; + + for (i =3D 0; i < 40; i++) { + if (madvise(p, size, MADV_PAGEOUT)) + ksft_exit_fail_perror("madvise(MADV_PAGEOUT)"); + if (check_swap(p, size)) + return true; + usleep(50 * 1000); + } + return false; +} + static bool is_swap_available(unsigned long size) { unsigned long swap_total =3D 0; @@ -881,9 +901,7 @@ static void collapse_swapin_single_pte(struct collapse_= context *c, struct mem_op p =3D ops->setup_area(1); ops->fault(p, 0, hpage_pmd_size); =20 - if (madvise(p, page_size, MADV_PAGEOUT)) - ksft_exit_fail_perror("madvise(MADV_PAGEOUT)"); - if (check_swap(p, page_size)) { + if (swapout_range(p, page_size)) { success("OK"); } else { fail("Fail"); @@ -920,9 +938,7 @@ static void collapse_max_ptes_swap(struct collapse_cont= ext *c, struct mem_ops *o p =3D ops->setup_area(1); ops->fault(p, 0, hpage_pmd_size); =20 - if (madvise(p, (max_ptes_swap + 1) * page_size, MADV_PAGEOUT)) - ksft_exit_fail_perror("madvise(MADV_PAGEOUT)"); - if (check_swap(p, (max_ptes_swap + 1) * page_size)) { + if (swapout_range(p, (max_ptes_swap + 1) * page_size)) { success("OK"); } else { fail("Fail"); @@ -937,9 +953,7 @@ static void collapse_max_ptes_swap(struct collapse_cont= ext *c, struct mem_ops *o ops->fault(p, 0, hpage_pmd_size); ksft_print_msg("Swapout %d of %d pages...", max_ptes_swap, hpage_pmd_nr); - if (madvise(p, max_ptes_swap * page_size, MADV_PAGEOUT)) - ksft_exit_fail_perror("madvise(MADV_PAGEOUT)"); - if (check_swap(p, max_ptes_swap * page_size)) { + if (swapout_range(p, max_ptes_swap * page_size)) { success("OK"); } else { fail("Fail"); --=20 2.54.0 From nobody Tue Sep 29 13:20:38 2026 Received: from fout-b5-smtp.messagingengine.com (fout-b5-smtp.messagingengine.com [202.12.124.148]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4B90347126A; Fri, 7 Aug 2026 11:37:06 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.148 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102630; cv=none; b=hj9U8sy+lYgLORinoZhPlVJWgV+0xHWqhIFBOQbDbmGulCawqaQugEofMxKrskNf00vbO5V9gVtDIBV6kK+RAc5mfflFg945A0qM292A3jc3nVPkJDnJROGHSy8wCUoMWBcOVywAa9/lSSzAk3T0iLOEYRSe4uBeTV5m9M98vsU= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102630; c=relaxed/simple; bh=EtVHDvfcfzGDmNjthWwAJslh8TM8zxxJJqnURhgKwRY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Q8N+tUKh3mkm0WZKrn8c0TvqD32Vj47acCsLqNBixDrdIDGWmx0B3jXGRp5Pr90fjIQS0QqZBT3LR0/yyhKB+gV5dq2nrVvX7myV86WjLZ6HaaRm9T2iv+BXumDOEO2qfRQQ8s8t/0lLvQTM8BEEtQnjFhGIkemXrqK7ifpr63k= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=m2UMoDij; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=DITeJelQ; arc=none smtp.client-ip=202.12.124.148 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="m2UMoDij"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="DITeJelQ" Received: from phl-compute-02.internal (phl-compute-02.internal [10.202.2.42]) by mailfout.stl.internal (Postfix) with ESMTP id 2DF841D00082; Fri, 7 Aug 2026 07:37:04 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-02.internal (MEProxy); Fri, 07 Aug 2026 07:37:04 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786102624; x= 1786189024; bh=A0HHfxeX5RuaRDWyznuRMDrKtp+1CgdzuTwwJX2D5b0=; b=m 2UMoDijXexCVrGBOds9UUyroJkfOqIZhMvFfFDoPFjj5ahoY35kTY/Pv30RU53fV TTr6sAemI4NKo0rrn9qfby+0AoBcOPvpr11lOyb/uABlw1JLHodEtqcXXUqg659k 5rkapFsSj5NcWKdKrp1z0l/qZ/kWG1k2vEIUVryFh7uNG6nSP362NXheGHzjc2ds Ty4ydTPHnCojmOzOhy8G7YxfJ/fzEeDYAV1fHgFT4Z8c8rJdSV0greQOz0MKLXiD 7LqJZOppb3ai8dSFlnWWx8cNMPIAWvDiRH6CMkqwakwPmuea+yNH4aki3mdZuzzZ hMv9AUqlvTygqC/oq+JgA== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786102624; x=1786189024; bh=A 0HHfxeX5RuaRDWyznuRMDrKtp+1CgdzuTwwJX2D5b0=; b=DITeJelQOMZvsRliO L1AswV/l0UkmK8H9jaIiWQwLHU0xJjox9kwCdwMi1mTJvxPGAmgHtmQuQrvEA+Is nQ4ew0Y83uATbbvoobhX3xPqRSrLtidRltxiSezlrHrQGqFL+aORPRwRZJIwB4Z3 3YWupHEHyECwl26ACtVV2XwbRW8LLN8vO1Iq9Pd1ho9TUSRUPDFfL0UwmhPzrRHh +iRD0kKcxbBz4BJ3gesfAP0VFAfLdCdaHJrmxoMaeOQvCvJNp0cdEcYl/e2zhu43 o6O+KG6Tz/kTs7DJAYdYYaQ8rYVew+GRkdkR3NC6wRN9SxMTH479HZwm1bsFt0hV ec4fA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEfW3x2Qrq7ELqxlr/x69vk0aYaa/E/2AWrFAl7yu0Vg+jIGlS0mjOCk5J5e5/NJH T3J+l8M+cQ15gxg/u6FQhwxKudsrf7tN5cWcEThveGDS0KcOU0uMrzGx0YBmylEV4PfT2X SbTqClZY3SINVohR/2NzYRxeOu5tU+SgKqY1sS5oM8U15+ky277S+G1Ww7YAUM5Z6EK46p TQK1COWwyEQl/dhwe51H5lWCLFzTj8K9hkk/JBO/DvOgEgoe5sMihzFP04eKTIrZCF8dte tQwhoNU+zszgTlPncqprVsvTwAgpjyPciGBDZXKQScTz4u1orAuTZg8Fi2VomZsbF2qQo6 JT6Sdos6SVoZTTW3+wS1zU1oDV+fbohsxvhGj0aeg73Xml+3CBIK14V5C8qv4Asbm9bXKZ Eh7U1twb1bxCOrCTVx7DPnpB29er5uVSfoP+NWHNHQEya3MER92j1+HR/PyTKk+R5j3bjf KwuIkMP1Ne0I34rb8/q7DifDC/AnEY0jLOJhKfPfqUvCoYYMQ5L/gOgu1YhwvO/8KVuSCC 5W0edYKA3/t1jWwxDfYFy6oXwGzxzfi38ItQrh4PocwCS75TZdwAPkjtE2s7t4nUL8BiPG 1k6qk00KQU1fwiF0KCDdk7UOhnJRzHRmWaHLty6LpaM3ChWMft+QdOZpcW5w X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 7 Aug 2026 07:37:03 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v2 05/17] selftests/mm: move is_backed_by_folio() into vm_util Date: Fri, 7 Aug 2026 12:36:35 +0100 Message-ID: <20260807113647.3744609-6-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260807113647.3744609-1-kirill@shutemov.name> References: <20260807113647.3744609-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The khugepaged selftest is about to gain mTHP collapse coverage, which needs to check that an address range is backed by a folio of a given order. split_huge_page_test.c already has the building block: is_backed_by_folio() reads the compound head and tail flags from /proc/kpageflags to classify the folio behind a page. Move it into vm_util so other tests can use it. No functional change. Assisted-by: Claude-Code:claude-opus-5 Acked-by: Mike Rapoport (Microsoft) Signed-off-by: Kiryl Shutsemau (Meta) Tested-by: Muhammad Usama Anjum --- .../selftests/mm/split_huge_page_test.c | 62 ------------------- tools/testing/selftests/mm/vm_util.c | 62 +++++++++++++++++++ tools/testing/selftests/mm/vm_util.h | 2 + 3 files changed, 64 insertions(+), 62 deletions(-) diff --git a/tools/testing/selftests/mm/split_huge_page_test.c b/tools/test= ing/selftests/mm/split_huge_page_test.c index 86a603692826..0adfe7dde7e5 100644 --- a/tools/testing/selftests/mm/split_huge_page_test.c +++ b/tools/testing/selftests/mm/split_huge_page_test.c @@ -42,68 +42,6 @@ const char *kpageflags_proc =3D "/proc/kpageflags"; int pagemap_fd; int kpageflags_fd; =20 -static bool is_backed_by_folio(char *vaddr, int order, int pagemap_fd, - int kpageflags_fd) -{ - const uint64_t folio_head_flags =3D KPF_THP | KPF_COMPOUND_HEAD; - const uint64_t folio_tail_flags =3D KPF_THP | KPF_COMPOUND_TAIL; - const unsigned long nr_pages =3D 1UL << order; - unsigned long pfn_head; - uint64_t pfn_flags; - unsigned long pfn; - unsigned long i; - - pfn =3D pagemap_get_pfn(pagemap_fd, vaddr); - - /* non present page */ - if (pfn =3D=3D -1UL) - return false; - - if (pageflags_get(pfn, kpageflags_fd, &pfn_flags)) - goto fail; - - /* check for order-0 pages */ - if (!order) { - if (pfn_flags & (folio_head_flags | folio_tail_flags)) - return false; - return true; - } - - /* non THP folio */ - if (!(pfn_flags & KPF_THP)) - return false; - - pfn_head =3D pfn & ~(nr_pages - 1); - - if (pageflags_get(pfn_head, kpageflags_fd, &pfn_flags)) - goto fail; - - /* head PFN has no compound_head flag set */ - if ((pfn_flags & folio_head_flags) !=3D folio_head_flags) - return false; - - /* check all tail PFN flags */ - for (i =3D 1; i < nr_pages; i++) { - if (pageflags_get(pfn_head + i, kpageflags_fd, &pfn_flags)) - goto fail; - if ((pfn_flags & folio_tail_flags) !=3D folio_tail_flags) - return false; - } - - /* - * check the PFN after this folio, but if its flags cannot be obtained, - * assume this folio has the expected order - */ - if (pageflags_get(pfn_head + nr_pages, kpageflags_fd, &pfn_flags)) - return true; - - /* If we find another tail page, then the folio is larger. */ - return (pfn_flags & folio_tail_flags) !=3D folio_tail_flags; -fail: - ksft_exit_fail_msg("Failed to get folio info\n"); - return false; -} - static int check_after_split_folio_orders(char *vaddr_start, size_t len, int pagemap_fd, int kpageflags_fd, int orders[], int nr_orders) { diff --git a/tools/testing/selftests/mm/vm_util.c b/tools/testing/selftests= /mm/vm_util.c index 360c9ec6702b..3c276865a778 100644 --- a/tools/testing/selftests/mm/vm_util.c +++ b/tools/testing/selftests/mm/vm_util.c @@ -494,6 +494,68 @@ int pageflags_get(unsigned long pfn, int kpageflags_fd= , uint64_t *flags) return 0; } =20 +bool is_backed_by_folio(char *vaddr, int order, int pagemap_fd, + int kpageflags_fd) +{ + const uint64_t folio_head_flags =3D KPF_THP | KPF_COMPOUND_HEAD; + const uint64_t folio_tail_flags =3D KPF_THP | KPF_COMPOUND_TAIL; + const unsigned long nr_pages =3D 1UL << order; + unsigned long pfn_head; + uint64_t pfn_flags; + unsigned long pfn; + unsigned long i; + + pfn =3D pagemap_get_pfn(pagemap_fd, vaddr); + + /* non present page */ + if (pfn =3D=3D -1UL) + return false; + + if (pageflags_get(pfn, kpageflags_fd, &pfn_flags)) + goto fail; + + /* check for order-0 pages */ + if (!order) { + if (pfn_flags & (folio_head_flags | folio_tail_flags)) + return false; + return true; + } + + /* non THP folio */ + if (!(pfn_flags & KPF_THP)) + return false; + + pfn_head =3D pfn & ~(nr_pages - 1); + + if (pageflags_get(pfn_head, kpageflags_fd, &pfn_flags)) + goto fail; + + /* head PFN has no compound_head flag set */ + if ((pfn_flags & folio_head_flags) !=3D folio_head_flags) + return false; + + /* check all tail PFN flags */ + for (i =3D 1; i < nr_pages; i++) { + if (pageflags_get(pfn_head + i, kpageflags_fd, &pfn_flags)) + goto fail; + if ((pfn_flags & folio_tail_flags) !=3D folio_tail_flags) + return false; + } + + /* + * check the PFN after this folio, but if its flags cannot be obtained, + * assume this folio has the expected order + */ + if (pageflags_get(pfn_head + nr_pages, kpageflags_fd, &pfn_flags)) + return true; + + /* If we find another tail page, then the folio is larger. */ + return (pfn_flags & folio_tail_flags) !=3D folio_tail_flags; +fail: + ksft_exit_fail_msg("Failed to get folio info\n"); + return false; +} + /* If `ioctls' non-NULL, the allowed ioctls will be returned into the var = */ int uffd_register_with_ioctls(int uffd, void *addr, uint64_t len, bool miss, bool wp, bool minor, uint64_t *ioctls) diff --git a/tools/testing/selftests/mm/vm_util.h b/tools/testing/selftests= /mm/vm_util.h index 9a49af88702e..56a28ce7d029 100644 --- a/tools/testing/selftests/mm/vm_util.h +++ b/tools/testing/selftests/mm/vm_util.h @@ -97,6 +97,8 @@ int64_t allocate_transhuge(void *ptr, int pagemap_fd); int pageflags_get(unsigned long pfn, int kpageflags_fd, uint64_t *flags); int gather_folio_orders(char *vaddr_start, size_t len, int pagemap_fd, int kpageflags_fd, int orders[], int nr_orders); +bool is_backed_by_folio(char *vaddr, int order, int pagemap_fd, + int kpageflags_fd); =20 int uffd_register(int uffd, void *addr, uint64_t len, bool miss, bool wp, bool minor); --=20 2.54.0 From nobody Tue Sep 29 13:20:38 2026 Received: from fhigh-b6-smtp.messagingengine.com (fhigh-b6-smtp.messagingengine.com [202.12.124.157]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E9B63471408; Fri, 7 Aug 2026 11:37:08 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.157 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102632; cv=none; b=cj4exypBumEHdr7x+DCvyp4AYyU61yaWi9Ps9lH467LeZeHP96h1D5XtzGy4WCCtEHaOY9i/nH2mjS7TisuMjqYQj0YCyuWfFoby540tjEN/OLvTWygVNuHm46yRKt5wG2Cf/HD7Mvd+QUHTFTDvyag5uJFVhgjqEs/0gtDm6ao= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102632; c=relaxed/simple; bh=rIKV1oM8EUwKl691Y4aowMbqUCWENCAnNomPHJ0L0kc=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=lWT9r4+prppZYMtSw7zULN7JcPNsCHs2sE9btaH8jf9XnEknAKDXBOVX4/O2mECvpizORJothcSQdw6EzY+MrSIpTPNgSZh3hQPRWRcF9Kwnb9wS99BZbMaIqRiuSvRBrB1u2f6TBclRlANKgQKgr0+Jx7LfIh9xMFPUncE4BYw= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=DmaVQLPw; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=ik8DZv6U; arc=none smtp.client-ip=202.12.124.157 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="DmaVQLPw"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="ik8DZv6U" Received: from phl-compute-04.internal (phl-compute-04.internal [10.202.2.44]) by mailfhigh.stl.internal (Postfix) with ESMTP id 87F197A00C4; Fri, 7 Aug 2026 07:37:06 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-04.internal (MEProxy); Fri, 07 Aug 2026 07:37:07 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786102626; x= 1786189026; bh=l5gMlhY734ld4sS1EqunE+HNrHgaVjUphYaBo5q1iDU=; b=D maVQLPwpHskr8ZYltXRDXvTDO8EUBoI4blSsCsLSlmOcSIZ5ccYVXjE+i79PZEqg z/sKey8i+8JWvTFRLTNGewLfay5YbjPW/wgYN/Z315ZrVTAzODHNQafq25GNFIZ9 XDwhx6gKAxkNEup6fSy33LYr3OVwCddMUd4TS9hRsdcadROm2yobmhPy8N8+RJB0 IfeFfwU0kTIL28WWDS1hYuJsnDPKesoddmWNIoMTiyeVs1uveNgvs1P61ozuucpe 8PFZJII0akGLU7mCMND5Yw28Lss0WxeEEepFDwc2mjZAa00Pappe0TrAabPM2LeJ 9NiNoZDtHj3958r/lFCxg== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786102626; x=1786189026; bh=l 5gMlhY734ld4sS1EqunE+HNrHgaVjUphYaBo5q1iDU=; b=ik8DZv6UZHXaMv2hq XY6dHfmZtlMNX1O288ZJ2ib/FFhKWLJGAxLrrsxsBbgzEI1aXVNketVyDOmFa7Xh 5a2Ji+qnQyLfpGZ+/ChVJA+ZuDNWkS9PwYqS5nGrx85AHneCY46NTbCeFpteqvhQ BU4z8DVKX3R52DOn1IZ3t9Qr6uBldXiF6Af5KDFFWVID8Vt0mIglXtOVhGLPIzmw SxPJuUGgUE9XxbBtI02WifuQFxNliGlaewXV7tS03dDDMNPxtOL8/jvda4EtzwJ6 4Z8tB0aBqTcVK/h/VKlJ/4cQlYKbcE5em4EXtCl4TJNpe+1NjMkw6T2sgh/4KAgz Ag4Qw== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEcBCtHwgFqhRPkncf0k+1ylINMWkiJotqLKgP2XUmyCnaVLH6jJFUJYUtwVpCetV bzy51nCrHgVQnKUbL6xUc35a2wACIZl+NbR3y/w0ModeWlKP4SKCXfcn2Mi8esZBMBdaWx faB4jzBJOawj2LBnFTgyCBiywrPp7FpNgfXB5cgYmB0wgvQ/HhhWAfKjMf/weJgLhMF/VV yXrMoja9EZiHjuXu8bb94hiDoU2fqI8TuwJJmzzN4TUL6MxSST+GnkpA2ZnKWVewXCphJB Lec89iAB+bDvFyV3zsxjbJ8cA1IEQ3aF+gxi5HDHtN75oTRJtbf/Sd54tDaiXy5fa0cQqS TiLXeHKeb+yhvBs42rwl+ePNPwrcrCpwjXd2vzbNyTKLhIUtrOJAuN15gqmTF5MDKkfuGE uX9cm0I6yOBZ1hS9pgoUq5/SSInRuWjFQciDGDhe5m6S8U21ILwj8Ehydnx3O86lYuTp6L fSAMba/DSXMK6QG7KbZBNIOGbS4ZIfQGxHMWj9JOmhpF3QRHO18vsM+2xfgxQWleRPyUKu 2P9eT8X6NWOWfhoEeN/z+SQeX35x7EzXlyyXFcTy/vMzJnPTvm3GDLShNMCBPaN3cTdM9z iBcgXeu6d//2ZKvvtP6tze5ZH0X0hTji9BezWMplN8bOwaSsxF2RjWKc6Q2Q X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 7 Aug 2026 07:37:05 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v2 06/17] selftests/mm: add folio-order check for address ranges Date: Fri, 7 Aug 2026 12:36:36 +0100 Message-ID: <20260807113647.3744609-7-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260807113647.3744609-1-kirill@shutemov.name> References: <20260807113647.3744609-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" An mTHP collapse test has to ask whether a range came out as folios of one particular order, placed where a collapse would place them. Nothing answers that today: is_backed_by_folio() classifies the folio behind a single page, and check_huge_anon() reads smaps AnonHugePages, which only accounts PMD mappings. Add is_range_backed_by_folio_orders(). For every order-aligned window of the range it requires a present head PFN at its natural alignment and a contiguous PFN run across the window. A window stitched together from several folios, or one holding a folio shifted off its natural position, fails -- which is what makes the helper usable for "this window collapsed and that one did not". Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) Tested-by: Muhammad Usama Anjum --- tools/testing/selftests/mm/vm_util.c | 42 ++++++++++++++++++++++++++++ tools/testing/selftests/mm/vm_util.h | 2 ++ 2 files changed, 44 insertions(+) diff --git a/tools/testing/selftests/mm/vm_util.c b/tools/testing/selftests= /mm/vm_util.c index 3c276865a778..3f586f2c3d33 100644 --- a/tools/testing/selftests/mm/vm_util.c +++ b/tools/testing/selftests/mm/vm_util.c @@ -556,6 +556,48 @@ bool is_backed_by_folio(char *vaddr, int order, int pa= gemap_fd, return false; } =20 +/* + * Check whether every order-@order window of [start, len) maps exactly one + * folio of that order, head to tail. The address range must be naturally + * aligned, each window's PFN run must be contiguous, and a window's first + * PFN must be the folio head. + * + * This is the check "did this range collapse into order-@order folios": a + * window assembled from parts of several folios, or mapping a folio shift= ed + * from its natural position, fails. + */ +bool is_range_backed_by_folio_orders(char *start, size_t len, int order, + int pagemap_fd, int kpageflags_fd) +{ + const unsigned long nr_pages =3D 1UL << order; + const size_t window =3D nr_pages * psize(); + char *vaddr; + + if ((uintptr_t)start % window || len % window) + return false; + + for (vaddr =3D start; vaddr < start + len; vaddr +=3D window) { + unsigned long pfn =3D pagemap_get_pfn(pagemap_fd, vaddr); + unsigned long i; + + /* Not present, or not mapping the folio head. */ + if (pfn =3D=3D -1UL || pfn % nr_pages) + return false; + + for (i =3D 1; i < nr_pages; i++) { + if (pagemap_get_pfn(pagemap_fd, vaddr + i * psize()) !=3D + pfn + i) + return false; + } + + if (!is_backed_by_folio(vaddr, order, pagemap_fd, + kpageflags_fd)) + return false; + } + + return true; +} + /* If `ioctls' non-NULL, the allowed ioctls will be returned into the var = */ int uffd_register_with_ioctls(int uffd, void *addr, uint64_t len, bool miss, bool wp, bool minor, uint64_t *ioctls) diff --git a/tools/testing/selftests/mm/vm_util.h b/tools/testing/selftests= /mm/vm_util.h index 56a28ce7d029..39dfb18dc10c 100644 --- a/tools/testing/selftests/mm/vm_util.h +++ b/tools/testing/selftests/mm/vm_util.h @@ -99,6 +99,8 @@ int gather_folio_orders(char *vaddr_start, size_t len, int pagemap_fd, int kpageflags_fd, int orders[], int nr_orders); bool is_backed_by_folio(char *vaddr, int order, int pagemap_fd, int kpageflags_fd); +bool is_range_backed_by_folio_orders(char *start, size_t len, int order, + int pagemap_fd, int kpageflags_fd); =20 int uffd_register(int uffd, void *addr, uint64_t len, bool miss, bool wp, bool minor); --=20 2.54.0 From nobody Tue Sep 29 13:20:38 2026 Received: from fhigh-b6-smtp.messagingengine.com (fhigh-b6-smtp.messagingengine.com [202.12.124.157]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AA4B23BE633; Fri, 7 Aug 2026 11:37:11 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.157 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102635; cv=none; b=JpdaB/1EcTZfM+38TWBsjQc4mBPJ19SkiZmj/yYvD78HIaSWRDxxmNqyFNpCyVhulZ40LPXdDHgYxA2HuBIrs9WyEBV87DGNwToHpTlEeaAj6k7RX1mFj7qxa56ZgGSZ8guEy8eCuybiy5iiPKkna8qBg2srcLY9LMpTzGf89v0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102635; c=relaxed/simple; bh=GXD55gzsxraPTO9fqQQ6dmtLVrWrm36CCw1CQfcJZ8o=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=YO9O5nJgDJYFzyLnol4zKZblung+CumfIdEVcBXD96cH2z9VFhj60pmFGGgV88GL1f0alTTOOygr7tIy32CXbEfnn6vdQ1tGBCcKp7e72VYXHW5Sl69FsWayX2VWsC/NNvrUFT5/nHcvWkld0AiZ+JDYJWymoo38cc7f64qe/uc= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=oM2T9FT6; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=M8dqq961; arc=none smtp.client-ip=202.12.124.157 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="oM2T9FT6"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="M8dqq961" Received: from phl-compute-09.internal (phl-compute-09.internal [10.202.2.49]) by mailfhigh.stl.internal (Postfix) with ESMTP id 260717A004F; Fri, 7 Aug 2026 07:37:09 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-09.internal (MEProxy); Fri, 07 Aug 2026 07:37:09 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786102628; x= 1786189028; bh=3Q59Aumd9RegCdWlgfhV9psGQQvAucyMQvh9lBiHEw4=; b=o M2T9FT6LfiuXNEslmr2/X0kTm2MS+KkZx1jphaTgf+sKPW5zz+/KUYjMS8F/KvFT toa7Xi7LVW5Va0P+Laa8BIJIT1KudZvodLb7v0Mu6GoijuSWuwFCKX8h0U7Ei5xG Z7ULhQO/pv86M58JEbWOk/DCPfle0ovhgGSNyfrTJsaJ0XjZJqqYYGYQcKiaoMlm E3zwNPJZnn78tIjeLuox93vG4eFJ5EiTUlL4qxSKDeivMJ0mYo/8gyy6UkMaYZgQ 8yTBOfK2BzB5aexsbYXADmElQd+58EK+co+gZAhTL2Bl/78vjDznzMfJYlob6nyo A5JL+9V1o82ThDmIbLZRQ== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786102628; x=1786189028; bh=3 Q59Aumd9RegCdWlgfhV9psGQQvAucyMQvh9lBiHEw4=; b=M8dqq961k0sKrZbvI Y49AP/ju2jYimNQtIwq2jMg2jSX02dmFnOPbRnVs7PuTmwX9O8Bmy5YDCc9dHedz 2HukqFisyIJhhDWBcl7xDxGIEVNaDBJZ1jzOLM+bAfHtT37mDZl+2DBjn0mCXfs2 2z1ychGNFKh4ZTaLO3Zj0aTjRKA5VO5TWRazezqaSTcJyJOANOyL+7uheR0uXUnI WW2ZEgG+bCtsGIL4kCObHqPeCnUJbbM892OWaYgVCI8ovf5SX+3mi11E/CL7yZB/ 6zdqo+/vlceTPPNbE/sdWGa7dMyTkQWk3zCMGJtDbMSVZYH12U7yeTIV1XQVTmGy wWQqQ== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEcBCtHwgFqhRPkncf0k+1ylINMWkiJotqLKgP2XUmyCnaVLH6jJFUJYUtwVpCetV bzy51nCrHgVQnKUbL6xUc35a2wACIZl+NbR3y/w0ModeWlKP4SKCXfcn2Mi8esZBMBdaWx faB4jzBJOawj2LBnFTgyCBiywrPp7FpNgfXB5cgYmB0wgvQ/HhhWAfKjMf/weJgLhMF/VV yXrMoja9EZiHjuXu8bb94hiDoU2fqI8TuwJJmzzN4TUL6MxSST+GnkpA2ZnKWVewXCphJB Lec89iAB+bDvFyV3zsxjbJ8cA1IEQ3aF+gxi5HDHtN75oTRJtbf/Sd54tDaiXy5fa0cQtq r+NnXnL5OKG4MXGIn/QCfDhFT91GGYEPW5Dk96gROsGOBHkEe5+DPqNBeNl7epTFHK7f8p Uw3PVVm+827Wwy42lRc2goJ99LU+sSl2tIXGuLCHea6m6MmZ5cdrvcFZCfwPd3vYBRWNLj YANYl1XnXhaOX5VI3ShqvILc6LQm0UCwIr+TAILVga2RYCcZkuEM8MfUhmy4IA++RCLv4n O6VzE92iHn4+1EAhaOwdUVqHT+rzbqyGDwEHCmbGpBa4oLb4BNWyXPIFPAsAcGgy7fVEEl sbnR+UlqPmFLRMkMOCmGey1wr3UP+cD32tj/ZdQQ2APwiEekon4EBFwLDp+g X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 7 Aug 2026 07:37:08 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v2 07/17] selftests/mm: add folio-order detection self-check Date: Fri, 7 Aug 2026 12:36:37 +0100 Message-ID: <20260807113647.3744609-8-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260807113647.3744609-1-kirill@shutemov.name> References: <20260807113647.3744609-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The upcoming khugepaged mTHP tests detect collapse results with the vm_util folio-order helpers instead of smaps AnonHugePages, which only sees PMD mappings. Before any collapse test trusts those helpers, make sure they agree with the kernel about what backs a mapping. For every anon THP order the kernel supports, fault memory in with only that order enabled and require the helpers to classify the backing as exactly that order: not a neighbouring order, and 4K-backed memory as order 0. Runs in the thp category of run_vmtests.sh. Verified on x86-64 4K (orders 0, 2-9) and arm64 64K (orders 0, 2-13). While here: ALIGN() moves into vm_util.h, where the next test that needs a round-up can find it, and hmm-tests.c drops its own copy. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) Tested-by: Muhammad Usama Anjum --- tools/testing/selftests/mm/Makefile | 1 + .../testing/selftests/mm/folio_order_check.c | 137 ++++++++++++++++++ tools/testing/selftests/mm/hmm-tests.c | 1 - tools/testing/selftests/mm/run_vmtests.sh | 2 + tools/testing/selftests/mm/vm_util.h | 2 + 5 files changed, 142 insertions(+), 1 deletion(-) create mode 100644 tools/testing/selftests/mm/folio_order_check.c diff --git a/tools/testing/selftests/mm/Makefile b/tools/testing/selftests/= mm/Makefile index 2d5366196e30..2093fcf6e915 100644 --- a/tools/testing/selftests/mm/Makefile +++ b/tools/testing/selftests/mm/Makefile @@ -104,6 +104,7 @@ TEST_GEN_FILES +=3D guard-regions TEST_GEN_FILES +=3D merge TEST_GEN_FILES +=3D rmap TEST_GEN_FILES +=3D folio_split_race_test +TEST_GEN_FILES +=3D folio_order_check =20 ifneq ($(ARCH),arm64) TEST_GEN_FILES +=3D soft-dirty diff --git a/tools/testing/selftests/mm/folio_order_check.c b/tools/testing= /selftests/mm/folio_order_check.c new file mode 100644 index 000000000000..93030a42c3cc --- /dev/null +++ b/tools/testing/selftests/mm/folio_order_check.c @@ -0,0 +1,137 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Self-check for the vm_util folio-order detection helpers, + * is_backed_by_folio() and is_range_backed_by_folio_orders(). + * + * For every anon THP order the kernel supports, fault memory in with only + * that order enabled and verify the helpers report exactly that order: + * not a neighbouring order, and plain 4K memory as order 0. The helpers + * are what the khugepaged mTHP tests use to detect collapse results, so + * they must agree with the kernel's own idea of the backing before any + * collapse test relies on them. + */ +#define _GNU_SOURCE +#include +#include +#include +#include +#include + +#include "kselftest.h" +#include "vm_util.h" +#include "hugepage_settings.h" + +static int pagemap_fd; +static int kpageflags_fd; + +/* mmap an anon VMA of exactly @size bytes at a @size-aligned address. */ +static char *alloc_aligned(size_t size) +{ + size_t len =3D size * 2; + uintptr_t aligned; + char *p; + + p =3D mmap(NULL, len, PROT_READ | PROT_WRITE, + MAP_ANONYMOUS | MAP_PRIVATE, -1, 0); + if (p =3D=3D MAP_FAILED) + ksft_exit_fail_perror("mmap()"); + + aligned =3D ALIGN((uintptr_t)p, size); + if (aligned !=3D (uintptr_t)p) + munmap(p, aligned - (uintptr_t)p); + if (aligned + size !=3D (uintptr_t)p + len) + munmap((char *)aligned + size, + (uintptr_t)p + len - aligned - size); + + return (char *)aligned; +} + +/* + * Enable only @order (order 0: nothing), fault one aligned window in and + * check the helpers see exactly @order. + */ +static void check_order(int order) +{ + struct thp_settings settings =3D *thp_current_settings(); + size_t size =3D psize() << order; + bool ok =3D true; + char *p; + int i; + + for (i =3D 0; i < NR_ORDERS; i++) + settings.hugepages[i].enabled =3D THP_NEVER; + if (order) + settings.hugepages[order].enabled =3D THP_ALWAYS; + thp_push_settings(&settings); + + p =3D alloc_aligned(size); + *p =3D 1; + + if (!is_range_backed_by_folio_orders(p, size, order, + pagemap_fd, kpageflags_fd)) { + ksft_print_msg("order %d not detected after fault\n", order); + ok =3D false; + } + + /* A lower order must be rejected: the folio is larger. */ + if (order && is_range_backed_by_folio_orders(p, size, order - 1, + pagemap_fd, + kpageflags_fd)) { + ksft_print_msg("order %d also reported as order %d\n", + order, order - 1); + ok =3D false; + } + + /* Order 0 pages must not look like any large folio, and vice versa. */ + if (order && is_range_backed_by_folio_orders(p, size, 0, + pagemap_fd, + kpageflags_fd)) { + ksft_print_msg("order %d also reported as order 0\n", order); + ok =3D false; + } + + munmap(p, size); + thp_pop_settings(); + + ksft_test_result(ok, "order %d classified\n", order); +} + +int main(void) +{ + struct thp_settings settings; + unsigned long orders; + int order; + + ksft_print_header(); + + if (!thp_available()) + ksft_exit_skip("Transparent Hugepages not available\n"); + + pagemap_fd =3D open("/proc/self/pagemap", O_RDONLY); + if (pagemap_fd < 0) + ksft_exit_fail_perror("open(/proc/self/pagemap)"); + kpageflags_fd =3D open("/proc/kpageflags", O_RDONLY); + if (kpageflags_fd < 0) + ksft_exit_skip("open(\"/proc/kpageflags\") requires root\n"); + + orders =3D thp_supported_orders(); + if (!orders) + ksft_exit_skip("No supported THP orders\n"); + + ksft_set_plan(__builtin_popcountl(orders) + 1); + + thp_save_settings(); + thp_read_settings(&settings); + /* Base of the settings stack; the bottom entry is never popped. */ + thp_push_settings(&settings); + + check_order(0); + for (order =3D 1; order < NR_ORDERS; order++) { + if (!(orders & (1UL << order))) + continue; + check_order(order); + } + + + ksft_finished(); +} diff --git a/tools/testing/selftests/mm/hmm-tests.c b/tools/testing/selftes= ts/mm/hmm-tests.c index e2642eca0d02..df426f9218e7 100644 --- a/tools/testing/selftests/mm/hmm-tests.c +++ b/tools/testing/selftests/mm/hmm-tests.c @@ -65,7 +65,6 @@ enum { #define HMM_PATH_MAX 64 #define NTIMES 10 =20 -#define ALIGN(x, a) (((x) + (a - 1)) & (~((a) - 1))) /* Just the flags we need, copied from mm.h: */ =20 #ifndef FOLL_WRITE diff --git a/tools/testing/selftests/mm/run_vmtests.sh b/tools/testing/self= tests/mm/run_vmtests.sh index d09f9f6a384e..2652a7920b80 100755 --- a/tools/testing/selftests/mm/run_vmtests.sh +++ b/tools/testing/selftests/mm/run_vmtests.sh @@ -402,6 +402,8 @@ CATEGORY=3D"pfnmap" run_test ./pfnmap # COW tests CATEGORY=3D"cow" run_test ./cow =20 +CATEGORY=3D"thp" run_test ./folio_order_check + CATEGORY=3D"thp" run_test ./khugepaged =20 CATEGORY=3D"thp" run_test ./khugepaged -s 2 diff --git a/tools/testing/selftests/mm/vm_util.h b/tools/testing/selftests= /mm/vm_util.h index 39dfb18dc10c..ce05bce4670d 100644 --- a/tools/testing/selftests/mm/vm_util.h +++ b/tools/testing/selftests/mm/vm_util.h @@ -10,6 +10,8 @@ #include =20 #define BIT_ULL(nr) (1ULL << (nr)) +#define ALIGN(x, a) (((x) + (a) - 1) & ~((a) - 1)) + #define PM_SOFT_DIRTY BIT_ULL(55) #define PM_MMAP_EXCLUSIVE BIT_ULL(56) #define PM_UFFD_WP BIT_ULL(57) --=20 2.54.0 From nobody Tue Sep 29 13:20:38 2026 Received: from fhigh-b6-smtp.messagingengine.com (fhigh-b6-smtp.messagingengine.com [202.12.124.157]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7D819470EAA; Fri, 7 Aug 2026 11:37:14 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.157 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102637; cv=none; b=n31FZXRI6XB57X552O3aueWp3XqUG8qqXmvwRI1HQ4eDVId0Tdl+Z6zuOAbUmT/ynfGhHIsiMZJN6tZSAcp2FXPxQkwT7asCLQ9m4feLbjkg8RextjZzOZs0Q81Ufc+Pxglb5Dwqqa3huw8Da8krmeXrvDTK9zVAzSpi21gGqpY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102637; c=relaxed/simple; bh=m38HzX/AsEPAN9/io+d10nxHX/8L6lZcjeaioJ5yNHk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=PrKlGjSmvbcxfVOX9gmamn6W9Jfw8I0TtICTu9qE5TNMmxEFhHEVJN7be70Eeti8UxEJRcpOflPxbw/JQdxxDKdqYWm5rrSIijhScqgDlXRyeOOlS0kINwEWXLbbZ4RqK0/0AE4NFoPRKnm5SK5hKuVhyJg+dEEYM5PH9X1dYBI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=NqdTKVrU; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=DxM7IqhF; arc=none smtp.client-ip=202.12.124.157 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="NqdTKVrU"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="DxM7IqhF" Received: from phl-compute-01.internal (phl-compute-01.internal [10.202.2.41]) by mailfhigh.stl.internal (Postfix) with ESMTP id AAA667A00AC; Fri, 7 Aug 2026 07:37:11 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-01.internal (MEProxy); Fri, 07 Aug 2026 07:37:12 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786102631; x= 1786189031; bh=Mkj1cRzazH/pmulbL8egXXcAWCNFw5TBX1ao39pm6/Y=; b=N qdTKVrU6cJA+055lbDIi8VTIWU5S2Cqn31BcUjc3ydVqCjAWKSL7yCHYMUYjm8eE sRRhxhLsLDgxpSIC8uV+d4WZzukvaa1nMzvOuhADzIwTLnf8GnkpmFI+6vBoXEqY wwDuV3yw/i42DKp97Eo/lerGGChERO1qFNmK1lGZX3WMKEsKnJ+m4TCDKO3eH1OL MrOwKO+5LbA4+uO+2HV4lx45EmDNlPTya0NnchIYEsU+dAq6azKBq6dx9J+Xt7d/ tHfWQ2+Gq38m1TgLkxk9ga7H1melRX2NIJIw+GjESktskiixYUuR6iAbqoPEJCYy 3wgx5g452/tDuWJphDCOg== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786102631; x=1786189031; bh=M kj1cRzazH/pmulbL8egXXcAWCNFw5TBX1ao39pm6/Y=; b=DxM7IqhFYchchtByp rG7R3C5VR7g5dWmducfDxa5JYuJwitQTsVw/cVNZsENZtwBtPx6Lb0mM6TFY4MSH ZcWkPUibhqirCdvKKLce9hpH3gr0Q7OH4rnDSpnIB3V4w9fDPgrD872Unb4HdS8V Ctt3advfVw/VTzgfqIftrACrHWtvl3GgxYmWY28Lp79+KNe1W579W2s/jPQW+N/P lPoDxnukcSTGk/KJthssSHDxA8rlQ/pDLO3HAGkidGuSB+V5feuhjQCsWsXUWUPl 94e7y6/gjG1g7RS1mlNaPzFAKdQVx6Gl9BPqkphq3MNPx5mAx+j1ifauYgxx5MXO AuyBQ== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEfW3x2Qrq7ELqxlr/x69vk0aYaa/E/2AWrFAl7yu0Vg+jIGlS0mjOCk5J5e5/NJH T3J+l8M+cQ15gxg/u6FQhwxKudsrf7tN5cWcEThveGDS0KcOU0uMrzGx0YBmylEV4PfT2X SbTqClZY3SINVohR/2NzYRxeOu5tU+SgKqY1sS5oM8U15+ky277S+G1Ww7YAUM5Z6EK46p TQK1COWwyEQl/dhwe51H5lWCLFzTj8K9hkk/JBO/DvOgEgoe5sMihzFP04eKTIrZCF8dte tQwhoNU+zszgTlPncqprVsvTwAgpjyPciGBDZXKQScTz4u1orAuTZg8Fi2VomZsbF2qQmF OdzRUWGH1aHPNSXFNkQ/9PAn+XNG9LuR9hRuKwh3bHLmpo8AjE5RNKZxmxAHfnSXBa+AwU NZVpbkZwjhB4n015qbMJHOsCvHvddJ8UVQ2fYa3VOWJxfK/oqZN2ML932kSF8d8J99W3x3 o+SPkkDRD8K7lYB2uUqdFwVWHQX86G69MIrzYZbXCVpDUU7hRm63rOyGVst5i4y0M3mQ+2 iMuCZB/JTKCqnPiPhC62N4N/GNt1VqqDJHecDFh9/INnRIgs6I5G7U23FS1xtjQDfC66Ok zmxZB47lUcPyXWWa+lpTOoZImAqY15f+xjqwc1ioj6I7dtteISjLxjlltKbA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 7 Aug 2026 07:37:10 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v2 08/17] selftests/mm: add khugepaged completion barrier helper Date: Fri, 7 Aug 2026 12:36:38 +0100 Message-ID: <20260807113647.3744609-9-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260807113647.3744609-1-kirill@shutemov.name> References: <20260807113647.3744609-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Race and functional tests need to drive khugepaged in step: set up a layout, let one full scan pass over it, check the result. The khugepaged selftest already waits for full_scans to advance by two, but only makes progress if scan_sleep_millisecs happens to be short. Lift it into khugepaged_full_pass() and drive it through sysfs: any store to scan_sleep_millisecs wakes the daemon, so the barrier completes whatever the scan cadence. It wakes once per missing pass -- over-waking queues a straggler pass that overlaps what the caller sets up next. One wake completes one pass only if the whole mm list fits in a scan batch, so callers need a large pages_to_scan. Settings pushes must not start passes either, so thp_write_settings() now writes a khugepaged knob only when its value changes. That helper, thp_update_num(), is exported for tests wanting the same restraint. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) Tested-by: Muhammad Usama Anjum --- .../testing/selftests/mm/hugepage_settings.c | 72 ++++++++++++++++--- .../testing/selftests/mm/hugepage_settings.h | 3 + 2 files changed, 66 insertions(+), 9 deletions(-) diff --git a/tools/testing/selftests/mm/hugepage_settings.c b/tools/testing= /selftests/mm/hugepage_settings.c index d7917dce3aba..8afcdf9793bb 100644 --- a/tools/testing/selftests/mm/hugepage_settings.c +++ b/tools/testing/selftests/mm/hugepage_settings.c @@ -183,6 +183,17 @@ void thp_read_settings(struct thp_settings *settings) } } =20 +/* + * Write only on change: any store to a khugepaged sysfs knob wakes the + * daemon, and settings pushes/pops must not start scan passes nobody + * asked for -- khugepaged_full_pass() is the only sanctioned wake. + */ +void thp_update_num(const char *name, unsigned long num) +{ + if (thp_read_num(name) !=3D num) + thp_write_num(name, num); +} + void thp_write_settings(struct thp_settings *settings) { struct khugepaged_settings *khugepaged =3D &settings->khugepaged; @@ -198,15 +209,15 @@ void thp_write_settings(struct thp_settings *settings) shmem_enabled_strings[settings->shmem_enabled]); thp_write_num("use_zero_page", settings->use_zero_page); =20 - thp_write_num("khugepaged/defrag", khugepaged->defrag); - thp_write_num("khugepaged/alloc_sleep_millisecs", - khugepaged->alloc_sleep_millisecs); - thp_write_num("khugepaged/scan_sleep_millisecs", - khugepaged->scan_sleep_millisecs); - thp_write_num("khugepaged/max_ptes_none", khugepaged->max_ptes_none); - thp_write_num("khugepaged/max_ptes_swap", khugepaged->max_ptes_swap); - thp_write_num("khugepaged/max_ptes_shared", khugepaged->max_ptes_shared); - thp_write_num("khugepaged/pages_to_scan", khugepaged->pages_to_scan); + thp_update_num("khugepaged/defrag", khugepaged->defrag); + thp_update_num("khugepaged/alloc_sleep_millisecs", + khugepaged->alloc_sleep_millisecs); + thp_update_num("khugepaged/scan_sleep_millisecs", + khugepaged->scan_sleep_millisecs); + thp_update_num("khugepaged/max_ptes_none", khugepaged->max_ptes_none); + thp_update_num("khugepaged/max_ptes_swap", khugepaged->max_ptes_swap); + thp_update_num("khugepaged/max_ptes_shared", khugepaged->max_ptes_shared); + thp_update_num("khugepaged/pages_to_scan", khugepaged->pages_to_scan); =20 if (dev_queue_read_ahead_path[0]) write_num(dev_queue_read_ahead_path, settings->read_ahead_kb); @@ -230,6 +241,49 @@ void thp_write_settings(struct thp_settings *settings) } } =20 +/* + * Completion barrier for khugepaged: wait until a full scan pass that + * started after this call has finished. full_scans must advance by two; + * a +1 step may complete a pass that examined this mm before the + * caller's setup was in place. + * + * Any store to scan_sleep_millisecs wakes the daemon, so the barrier works + * whatever the configured scan cadence -- but a store can be lost. + * __sleep_millisecs_store() clears khugepaged_sleep_expire and wakes the + * queue; if the daemon is between scans rather than sleeping, it sets + * khugepaged_sleep_expire itself on the way into khugepaged_wait_work() a= nd + * then sleeps for the full interval, having never seen the store. So keep + * storing until the pass lands; a store while the daemon is awake costs + * nothing and does not queue an extra pass. + * + * One wake completes one full pass only if the whole mm list fits in + * one scan batch, so callers must pair this with a large + * pages_to_scan. + */ +bool khugepaged_full_pass(unsigned int timeout_s) +{ + unsigned long deadline_ms =3D timeout_s * 1000UL; + unsigned long sleep_ms =3D + thp_read_num("khugepaged/scan_sleep_millisecs"); + unsigned long elapsed_ms =3D 0; + int pass; + + for (pass =3D 0; pass < 2; pass++) { + unsigned long target =3D + thp_read_num("khugepaged/full_scans") + 1; + + while (thp_read_num("khugepaged/full_scans") < target) { + if (elapsed_ms >=3D deadline_ms) + return false; + thp_write_num("khugepaged/scan_sleep_millisecs", + sleep_ms); + usleep(10 * 1000); + elapsed_ms +=3D 10; + } + } + return true; +} + struct thp_settings *thp_current_settings(void) { if (!settings_index) { diff --git a/tools/testing/selftests/mm/hugepage_settings.h b/tools/testing= /selftests/mm/hugepage_settings.h index 726c73c43c05..ba7d38370d43 100644 --- a/tools/testing/selftests/mm/hugepage_settings.h +++ b/tools/testing/selftests/mm/hugepage_settings.h @@ -70,6 +70,7 @@ int thp_read_string(const char *name, const char * const = strings[]); void thp_write_string(const char *name, const char *val); unsigned long thp_read_num(const char *name); void thp_write_num(const char *name, unsigned long num); +void thp_update_num(const char *name, unsigned long num); =20 void thp_write_settings(struct thp_settings *settings); void thp_read_settings(struct thp_settings *settings); @@ -83,6 +84,8 @@ static inline void thp_save_settings(void) hugepage_save_settings(/* thp =3D */ true, /* hugetlb =3D */ false); } =20 +bool khugepaged_full_pass(unsigned int timeout_s); + void thp_set_read_ahead_path(char *path); unsigned long thp_supported_orders(void); unsigned long thp_shmem_supported_orders(void); --=20 2.54.0 From nobody Tue Sep 29 13:20:38 2026 Received: from fhigh-b6-smtp.messagingengine.com (fhigh-b6-smtp.messagingengine.com [202.12.124.157]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7FC83471400; Fri, 7 Aug 2026 11:37:17 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.157 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102641; cv=none; b=ijcAS5I8QGRO38euTLx3BLUqLVEpR8YGzAKWRwAZZhUC3sFeyyU99ArZv9mgR7vDDaSLhy5eKl1UVeIxHKt6L9xS7gb1gkAZvtX22Xe6DDUzEdBEm8s3vxcl04HLOtWjiTqe/9vmTZNLUrkUpWohyQ3qiYtTgksnnqezpuhx5A0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102641; c=relaxed/simple; bh=2A41GZsQSZMc2sSLxZwfS0VGxSw1NIh8eKqFevmdOUo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=oTupuALGvl11at69KiVovmbAyzDUPITrK5WYt3qfJ2+4KVCJarVpiq1PNvqz7f1QsuD7Q2P+z1KM9gjiGylw4ozxpaBPlrf9flTLGMfXKYXz84mPu9F1XtlqGBwlGyCpZ2kjwlLVC9h03soVDqP+VqlNb7E7k9GtqlL0pMR3LfQ= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=CTXs/xTv; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=X+ZAYHm6; arc=none smtp.client-ip=202.12.124.157 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="CTXs/xTv"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="X+ZAYHm6" Received: from phl-compute-04.internal (phl-compute-04.internal [10.202.2.44]) by mailfhigh.stl.internal (Postfix) with ESMTP id 142BB7A0120; Fri, 7 Aug 2026 07:37:14 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-04.internal (MEProxy); Fri, 07 Aug 2026 07:37:14 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786102633; x= 1786189033; bh=7RLj+GhKJGvuXr+2Q6khalcVsCGbMqebY7tWpZ0bpic=; b=C TXs/xTvQ8lEPg0Lnd0k4++oouWZs14VRweVFW2r/Qr2HqhNvxOrcYNvRVvpc7uiy 6GYkWA62zEW2Epx3xrcXJgRrGVW+NGC+MrqBlggGDAcQx3cBiXcQ9d5po7XAOpGT AcTh5YyBXV3JOKVQ1G5WqCq/0Deli7fd5GOel5ceKvTW5e7PP7DIN7h5zoVp5evG cAV4Z2V+9md3KCFur7DFfQFfV79s92QEbv62LfyKA8Hk5YFf3pb2qp1434kB35mQ 6Wea0YNl3MF3WxdzOQeBstmlF/pFoHeEV3qCpyq0eLVD0bbtyeP5fxA2XYTnniWD y7c1beBmGd0KVN7hJmWiQ== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786102633; x=1786189033; bh=7 RLj+GhKJGvuXr+2Q6khalcVsCGbMqebY7tWpZ0bpic=; b=X+ZAYHm6MK4dUKukK 0Q2W8JJu0l3879eIhUbw9TxYKSlfA/nu/H10ZLzJqL7qTy/PW9fTEPhUmo46ssYJ KczlrG3Zysc7rDplptPb4wIsSjWaoRboYQWeP95XboREtvmT7tSDjCxIVNoNSjwG Qc2SAWGW/fBMtivhSY3ICwTemus1g6k8FMAlUMyCLUuqmi7mMxai5f/IfCdReBIW CglBf08UdNZ+3kLhCx3OaczQg6fl9jaRYn0WuOw1QILEWkI93vBGn9L6yZYuRFlx yjuwyNwI3StCVSqGyNbCybAZtmcPaQXZxWxC8zTC0Bu62vL+fl26+FQ8uTGnDqv7 BRhTQ== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFHgVqOo5hmjJljb3IAZGkXyyxzQaD6i5RNmNk424gVmZkPzr1NfNWN42YMHBLBJ7 ouwgSkV7uP+O116AnyQ77X3VN5vAdqq1zvzV33EYJRHUa/cQ0OT9UF9tJbf4RJE1rh86et dgK/XsSmEjRvIZ4/DwnHkulf5IWkeizHGDAErUJF9mnMadsJqDG0SAHuOjgfipI02yAqvF a9Gp7D1P4aM7ojieUjy73reTq6pHjAlslJzhWF2Bayc7rzuN1loFnb3tH1A/9l8/u/BKz2 kEn4hJCLyjrltgNf/ADveISAZdc/G8K4aY+0IiJqAeENTN6qSce54NhN3knOmRl67eS+r0 WbB8FDRD8Au+Ul36Ui63//knhvcdpy/2ObtGyXjV1wk+1Ra0l5w3HpclF91ujjSDW50IZn 0JJVkQexVMzgM8el3u286H6VD5m376WwXUFv3p3seAqRjpJN6EZRk7TBof/oTj4rXd7i+0 maLtZ1m8876dtBEWKZQ+82r/HY58GCkcxnfQA9b1j0r884ELtpH6tjrVKq7CCi7YwD1fKb +a6OVn8CxN/lj4pSaWq2EjyF7gRAZaFNyl4pdfyXs88CEvTReS24GkQlClr56kIN5IeM6o jkS54vxV0WnazKVtF6TTM5g+vefrF05VUgLOYNW2R3HdAVdVImRUR2x36BCw X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 7 Aug 2026 07:37:13 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v2 09/17] selftests/mm: add order-parameterized khugepaged collapse cases Date: Fri, 7 Aug 2026 12:36:39 +0100 Message-ID: <20260807113647.3744609-10-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260807113647.3744609-1-kirill@shutemov.name> References: <20260807113647.3744609-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The mthp_khugepaged context runs the generic cases at a sub-PMD order, which answers how many folios of that order a range ends up with. It cannot say which window they are in, so "the populated window collapsed and its neighbour did not" and "one window collapsed twice" look alike. Add four cases that check each aligned window on its own, with the folio-order helpers in vm_util: - collapse_order_single_window: only the populated window collapses; - collapse_order_partial_window: the default max_ptes_none lets a window with one present PTE collapse; - collapse_order_max_ptes_none: with max_ptes_none=3D0 a full window collapses and one missing a page does not; - collapse_order_mixed_sources: sources that are already large folios of a smaller order collapse to the target. Each case faults its region before MADV_HUGEPAGE with only the target order enabled, so the sources are order 0 and the result can only come from khugepaged. They wait for a full pass rather than for the result to appear: without a completed pass, "not collapsed" and "not scanned yet" are the same thing. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) Tested-by: Muhammad Usama Anjum --- tools/testing/selftests/mm/khugepaged.c | 219 ++++++++++++++++++++++++ 1 file changed, 219 insertions(+) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index 69e0fc4b12fb..bcaef17e430d 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -31,6 +31,8 @@ static unsigned long page_size; static int hpage_pmd_nr; static int anon_order; static int collapse_order; +static int pagemap_fd =3D -1; +static int kpageflags_fd =3D -1; =20 #define PID_SMAPS "/proc/self/smaps" #define TEST_FILE "collapse_test_file" @@ -1268,6 +1270,205 @@ static void madvise_retracted_page_tables(struct co= llapse_context *c, ksft_test_result_report(exit_status, "%s\n", __func__); } =20 +/* Smallest order khugepaged will consider for mTHP collapse. */ +#define MIN_MTHP_ORDER 2 + +/* + * Order-parameterized collapse cases for the mthp_khugepaged context. Wh= at + * they add over the generic cases run under that context is per-window + * detection: which aligned window collapsed, and which of its neighbours = did + * not. check_huge() answers how many folios of the order the range holds, + * which cannot tell one window from another. + * + * The region is faulted before MADV_HUGEPAGE, and the target order is only + * enabled for madvise, so the sources are always order 0 and the collapse + * product can only have come from khugepaged. + */ +static size_t mthp_window_size(void) +{ + return page_size << collapse_order; +} + +static void mthp_push_target_order(void) +{ + struct thp_settings settings =3D *thp_current_settings(); + int i; + + /* + * The target order, for madvise only, and nothing else enabled: the + * cases fault their region before MADV_HUGEPAGE, so the sources are + * order 0 whatever -s asked the fault path for. That matters for the + * cases built around a hole -- a large source folio would fill it in + * and the window would collapse after all. + * collapse_order_mixed_sources enables the source order it wants on + * top of this. + */ + settings.thp_enabled =3D THP_NEVER; + for (i =3D 0; i < NR_ORDERS; i++) + settings.hugepages[i].enabled =3D THP_NEVER; + settings.hugepages[collapse_order].enabled =3D THP_MADVISE; + thp_push_settings(&settings); +} + +static bool window_collapsed(void *p, size_t len) +{ + return is_range_backed_by_folio_orders(p, len, collapse_order, + pagemap_fd, kpageflags_fd); +} + +/* No aligned window in [p, p + len) is backed at the target order. */ +static bool window_not_collapsed(void *p, size_t len) +{ + size_t window =3D mthp_window_size(); + char *addr =3D p; + + for (; len >=3D window; addr +=3D window, len -=3D window) { + if (window_collapsed(addr, window)) + return false; + } + return true; +} + +static bool khugepaged_wait_full_pass(void) +{ + /* Wait up to 30 seconds for the pass to complete. */ + return khugepaged_full_pass(30); +} + +static void collapse_order_single_window(struct collapse_context *c, + struct mem_ops *ops) +{ + size_t window =3D mthp_window_size(); + void *p; + + mthp_push_target_order(); + + p =3D ops->setup_area(1); + ops->fault(p, window, 2 * window); + if (!window_not_collapsed(p, hpage_pmd_size)) + ksft_exit_fail_msg("Unexpected large folio after fault\n"); + + madvise(p, hpage_pmd_size, MADV_HUGEPAGE); + ksft_print_msg("Collapse one fully populated window..."); + if (!khugepaged_wait_full_pass()) + fail("Timeout"); + else if (window_collapsed(p + window, window) && + window_not_collapsed(p, window) && + window_not_collapsed(p + 2 * window, + hpage_pmd_size - 2 * window)) + success("OK"); + else + fail("Fail"); + + validate_memory(p, window, 2 * window); + ops->cleanup_area(p, hpage_pmd_size); + thp_pop_settings(); + ksft_test_result_report(exit_status, "%s\n", __func__); +} + +static void collapse_order_partial_window(struct collapse_context *c, + struct mem_ops *ops) +{ + void *p; + + mthp_push_target_order(); + + p =3D ops->setup_area(1); + ops->fault(p, 0, page_size); + if (!window_not_collapsed(p, hpage_pmd_size)) + ksft_exit_fail_msg("Unexpected large folio after fault\n"); + + madvise(p, hpage_pmd_size, MADV_HUGEPAGE); + ksft_print_msg("Collapse window with single PTE entry present..."); + if (!khugepaged_wait_full_pass()) + fail("Timeout"); + else if (window_collapsed(p, mthp_window_size())) + success("OK"); + else + fail("Fail"); + + validate_memory(p, 0, page_size); + ops->cleanup_area(p, hpage_pmd_size); + thp_pop_settings(); + ksft_test_result_report(exit_status, "%s\n", __func__); +} + +static void collapse_order_max_ptes_none(struct collapse_context *c, + struct mem_ops *ops) +{ + struct thp_settings settings; + size_t window =3D mthp_window_size(); + void *p; + + mthp_push_target_order(); + settings =3D *thp_current_settings(); + settings.khugepaged.max_ptes_none =3D 0; + thp_push_settings(&settings); + + p =3D ops->setup_area(1); + ops->fault(p, 0, 2 * window - page_size); + if (!window_not_collapsed(p, hpage_pmd_size)) + ksft_exit_fail_msg("Unexpected large folio after fault\n"); + + madvise(p, hpage_pmd_size, MADV_HUGEPAGE); + ksft_print_msg("Collapse full window, not the one missing a page..."); + if (!khugepaged_wait_full_pass()) + fail("Timeout"); + else if (window_collapsed(p, window) && + window_not_collapsed(p + window, window)) + success("OK"); + else + fail("Fail"); + + validate_memory(p, 0, 2 * window - page_size); + ops->cleanup_area(p, hpage_pmd_size); + thp_pop_settings(); + thp_pop_settings(); + ksft_test_result_report(exit_status, "%s\n", __func__); +} + +static void collapse_order_mixed_sources(struct collapse_context *c, + struct mem_ops *ops) +{ + struct thp_settings settings; + void *p; + + if (collapse_order <=3D MIN_MTHP_ORDER) { + ksft_test_result_skip("%s: no source order below target\n", + __func__); + return; + } + + mthp_push_target_order(); + + /* Fault the whole region as order-MIN_MTHP_ORDER folios. */ + settings =3D *thp_current_settings(); + settings.hugepages[MIN_MTHP_ORDER].enabled =3D THP_ALWAYS; + thp_push_settings(&settings); + p =3D ops->setup_area(1); + ops->fault(p, 0, hpage_pmd_size); + thp_pop_settings(); + + if (!is_range_backed_by_folio_orders(p, hpage_pmd_size, MIN_MTHP_ORDER, + pagemap_fd, kpageflags_fd)) + ksft_exit_fail_msg("Region not backed by order-%d folios after fault\n", + MIN_MTHP_ORDER); + + madvise(p, hpage_pmd_size, MADV_HUGEPAGE); + ksft_print_msg("Collapse region backed by smaller large folios..."); + if (!khugepaged_wait_full_pass()) + fail("Timeout"); + else if (window_collapsed(p, hpage_pmd_size)) + success("OK"); + else + fail("Fail"); + + validate_memory(p, 0, hpage_pmd_size); + ops->cleanup_area(p, hpage_pmd_size); + thp_pop_settings(); + ksft_test_result_report(exit_status, "%s\n", __func__); +} + static void usage(void) { fprintf(stderr, "\nUsage: ./khugepaged [OPTIONS] [dir]\n\n"); @@ -1436,6 +1637,20 @@ int main(int argc, char **argv) =20 parse_test_type(argc, argv); =20 + if (mthp_khugepaged_context && + !(thp_supported_orders() & (1UL << collapse_order))) + ksft_exit_skip("Order %d is not a supported anon THP order\n", + collapse_order); + + if (mthp_khugepaged_context) { + pagemap_fd =3D open("/proc/self/pagemap", O_RDONLY); + if (pagemap_fd < 0) + ksft_exit_fail_perror("open(/proc/self/pagemap)"); + kpageflags_fd =3D open("/proc/kpageflags", O_RDONLY); + if (kpageflags_fd < 0) + ksft_exit_fail_perror("open(/proc/kpageflags)"); + } + setbuf(stdout, NULL); =20 /* @@ -1498,6 +1713,10 @@ int main(int argc, char **argv) TEST(collapse_empty, madvise_context, anon_ops); =20 TEST(collapse_single_mthp, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_single_window, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_partial_window, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_max_ptes_none, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_mixed_sources, mthp_khugepaged_context, anon_ops); =20 TEST(collapse_single_pte_entry, khugepaged_context, anon_ops); TEST(collapse_single_pte_entry, khugepaged_context, read_only_file_ops); --=20 2.54.0 From nobody Tue Sep 29 13:20:38 2026 Received: from fhigh-b6-smtp.messagingengine.com (fhigh-b6-smtp.messagingengine.com [202.12.124.157]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 26B3D472520; Fri, 7 Aug 2026 11:37:18 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.157 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102642; cv=none; b=hv2XNPc2tdpVH73nntwynt8StbQuL30boJA4255UyqTf/4jha4d6pzF0d2RbaPigmGIe4PZBwbdfyX8OHhb9vmFOaTvRjK6Mw2J1UzVz9GlAT0C8C0OMqjWkgG381sadxe+43Xi0gedGFkDNBuI2fJMkzqZiLaMXKyePaWPVwKM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102642; c=relaxed/simple; bh=qp+IRtLlULna7LZoS83zodAAVuFN2LVhsRhdjjr+MWg=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=uVvL5zwde3yDTgwHm8g21G2o8Wx/TlMZhDBj7uZqy1vLAc3nIaWrpRXUQ1Z5GGwcmTrfhustcvbwMoCcnbZtvQmXrPDjNm8hq8egX+d3ndpRzFwNZCSHkMATKpkzbcATX1cCJYSkG1zg39U8zUGuK8g0EUkj10p545AbCNCBK0k= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=QOJdnMDE; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=HcUNMZRs; arc=none smtp.client-ip=202.12.124.157 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="QOJdnMDE"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="HcUNMZRs" Received: from phl-compute-02.internal (phl-compute-02.internal [10.202.2.42]) by mailfhigh.stl.internal (Postfix) with ESMTP id 651537A0125; Fri, 7 Aug 2026 07:37:16 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-02.internal (MEProxy); Fri, 07 Aug 2026 07:37:17 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786102636; x= 1786189036; bh=zDwmOcIYA0DElOIzeueyNhLx2hbSZ552mVI7ZkHow+M=; b=Q OJdnMDEKaEpsGhKkFkXswda3CyYDZquipaju8fR91AVa54DsxnTagFQiTjs+NSbW fXr4brLzf8xHbJkynxgc7STv8PLEvMcjCRKSwWTntt3XztwtxzxWPoj0V6m98C2B OmSe5Ww3PK36yeQjpAQmvt6NF0bwW6b2KW4PbHWPH2d4S+Z+doaIKuMwMT7ONsNj SCro0sbv/IQJLyg1iuqeHgfC1KdzFwtgwLwls4VcRGUYaQMsDV0TYmPcx9Bbnn7h IOXyMWaOSP9aPEtvQnWEWFKjh7J3K430e+muJ+gDmHn/cF3hsZq4FJqELnCeNsFP f5nLbOLmeQ3NbDCAkAfrw== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786102636; x=1786189036; bh=z DwmOcIYA0DElOIzeueyNhLx2hbSZ552mVI7ZkHow+M=; b=HcUNMZRskVTum2QsU He0NQ48mK65A2jGlUCKYJmm24mArfJXKF+DYfd7Rjy31jQ+2snNRIUqJKkec5EES MWmWV4thUnScKCS9Wsh1PmLpdhFFSVOsMx9U0Ccbb0fr2Q3qxdl7dkU4t30SoPWl ufOUCM3rjnfi3GUjsG3rz8EuwrXqQggZA/olmMBV+9vSoqvZDpL7OszQKe52Er8H jDWu6oKS90Wf9OEAnTKL+hhstygmtGJDU7evkDKftsc0fkt+AAPNUwir7lIJcyIK UDRDBQGDhavSL2SUpffpDj4PzgzFsYZLsB2/0ysGQYChs/xQXah3P6EYySvGOA2z T9JSg== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEcBCtHwgFqhRPkncf0k+1ylINMWkiJotqLKgP2XUmyCnaVLH6jJFUJYUtwVpCetV bzy51nCrHgVQnKUbL6xUc35a2wACIZl+NbR3y/w0ModeWlKP4SKCXfcn2Mi8esZBMBdaWx faB4jzBJOawj2LBnFTgyCBiywrPp7FpNgfXB5cgYmB0wgvQ/HhhWAfKjMf/weJgLhMF/VV yXrMoja9EZiHjuXu8bb94hiDoU2fqI8TuwJJmzzN4TUL6MxSST+GnkpA2ZnKWVewXCphJB Lec89iAB+bDvFyV3zsxjbJ8cA1IEQ3aF+gxi5HDHtN75oTRJtbf/Sd54tDaiXy5fa0cQPj U2TbJkcb+9iijbIPlEaK0B6F5bw3UDfJ1MYirL/BHTzp/YpPZRiz1Cym40T7pwZTZFigNO n5dWOQG0yWOfNnsOVtwP6tYvaWKY1vWY9/4oBQpWn3s15AWbMlP1fHZ6MANq18DJSWjWaC RtK+QxWIn/52Vmxhvv0CuwdFUt20EIQvpO6uJltK8t1kkJEFmJIngj5+Q/zb9SjzYQGBIt G67NUe2kLEnHEDNAPWT357EKQWHoJMwAwplVSaoOz7AnxqBriAILhGNlYFtd/5P/7xPw32 rZmePKHNzMbAGsyTYY/V2WtvPfIVkqh1VVxyFxsP7hUrlVk26uJHCY/4rKQA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 7 Aug 2026 07:37:15 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v2 10/17] selftests/mm: parameterize the mixed-source collapse case by source order Date: Fri, 7 Aug 2026 12:36:40 +0100 Message-ID: <20260807113647.3744609-11-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260807113647.3744609-1-kirill@shutemov.name> References: <20260807113647.3744609-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" collapse_order_mixed_sources faults its region as order-2 folios and collapses them to the -c target. Order 2 sits below the contpte threshold on both arm64 page-size configurations, so nothing in this suite unfolds a contpte source on purpose. Let -s name the source order alongside -c, which was rejected before. The case then faults at that order, keeping order 2 when -s is absent, and the source order has to be a supported mTHP order below the target. Other cases are unaffected: they enable their source order locally. "-s 5 -c 7" on arm64/64K then collapses contpte-mapped sources into a larger mTHP. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) Tested-by: Muhammad Usama Anjum --- tools/testing/selftests/mm/khugepaged.c | 18 ++++++++++++------ 1 file changed, 12 insertions(+), 6 deletions(-) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index bcaef17e430d..d3ff1b196e09 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -1430,10 +1430,13 @@ static void collapse_order_max_ptes_none(struct col= lapse_context *c, static void collapse_order_mixed_sources(struct collapse_context *c, struct mem_ops *ops) { + int source_order =3D anon_order ? anon_order : MIN_MTHP_ORDER; struct thp_settings settings; void *p; =20 - if (collapse_order <=3D MIN_MTHP_ORDER) { + /* Sources must be a supported mTHP order strictly below the target. */ + if (source_order >=3D collapse_order || + !(thp_supported_orders() & (1UL << source_order))) { ksft_test_result_skip("%s: no source order below target\n", __func__); return; @@ -1441,21 +1444,22 @@ static void collapse_order_mixed_sources(struct col= lapse_context *c, =20 mthp_push_target_order(); =20 - /* Fault the whole region as order-MIN_MTHP_ORDER folios. */ + /* Fault the whole region as order-@source_order folios. */ settings =3D *thp_current_settings(); - settings.hugepages[MIN_MTHP_ORDER].enabled =3D THP_ALWAYS; + settings.hugepages[source_order].enabled =3D THP_ALWAYS; thp_push_settings(&settings); p =3D ops->setup_area(1); ops->fault(p, 0, hpage_pmd_size); thp_pop_settings(); =20 - if (!is_range_backed_by_folio_orders(p, hpage_pmd_size, MIN_MTHP_ORDER, + if (!is_range_backed_by_folio_orders(p, hpage_pmd_size, source_order, pagemap_fd, kpageflags_fd)) ksft_exit_fail_msg("Region not backed by order-%d folios after fault\n", - MIN_MTHP_ORDER); + source_order); =20 madvise(p, hpage_pmd_size, MADV_HUGEPAGE); - ksft_print_msg("Collapse region backed by smaller large folios..."); + ksft_print_msg("Collapse region backed by order-%d sources...", + source_order); if (!khugepaged_wait_full_pass()) fail("Timeout"); else if (window_collapsed(p, hpage_pmd_size)) @@ -1486,6 +1490,8 @@ static void usage(void) fprintf(stderr, "\t\t-s: mTHP size, expressed as page order.\n"); fprintf(stderr, "\t\t Defaults to 0. Use this size for anon or shmem a= llocations.\n"); fprintf(stderr, "\t\t-c: collapse order for mTHP collapse, expressed as p= age order.\n"); + fprintf(stderr, "\t\t With -s, -s names the mTHP source order for the\= n"); + fprintf(stderr, "\t\t mixed-source case (source order below the target= ).\n"); exit(1); } =20 --=20 2.54.0 From nobody Tue Sep 29 13:20:38 2026 Received: from fout-b5-smtp.messagingengine.com (fout-b5-smtp.messagingengine.com [202.12.124.148]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D0EC047277C; Fri, 7 Aug 2026 11:37:21 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.148 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102645; cv=none; b=dI/xOIC20d255VX7yQ2qygxB93KTNu3lI4uopy5nhh+n8EPHkFFEb+h1neIt0SCdm560Yts7aQkmmUYx4QLPF6RNNzAq7zruPItMwoCfNl3N0qw64wbb2PG/VD8Pr9aRXEwh1ZVcZA7G039mrvQ403ekB4zxVEu7zO6vdxF8seY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102645; c=relaxed/simple; bh=fQMu5pxtI/bJpziVuI0v74HoxYFXAO5iWBBgu6Q5DIo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=aOei/m4fbxr8H50JN2C9Z2voODDcucDnhBuSPcjV5zIv3DSmeh8aJE5kAFNJlYneSIcvkeJqIiEUKBhKEdp1qhTG+ieMP22N12PZT/EfwubGPeTf00p/hxZqDUrfQ9J9xq/MKgh1f+/GSA9DsgHjze4jq/TPsMHfkZByv098TYY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=u0KIY31K; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=JmtyROkD; arc=none smtp.client-ip=202.12.124.148 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="u0KIY31K"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="JmtyROkD" Received: from phl-compute-05.internal (phl-compute-05.internal [10.202.2.45]) by mailfout.stl.internal (Postfix) with ESMTP id 1031D1D00082; Fri, 7 Aug 2026 07:37:19 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-05.internal (MEProxy); Fri, 07 Aug 2026 07:37:19 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786102638; x= 1786189038; bh=AJEIXCOdsK1VdJ+vTB/+b5PTVsVX+FRzDaAsHSTP7IA=; b=u 0KIY31KUUaa8liCCu9KTgWbh9RLdB7yGDqhdRba/oqA6EqTFBxivp+GvevkcPGp0 R7V96Ts6E2FSICBk9yItN3plLUxEtbnEYAszGtb28aIN3a4KRsbNKey4OyOrzrtT ghye5sOpszFtcxAcAbtiS6IvNp4WR/JlkaEbDJE3vSfnP93AXAZZ4GdVs7oegEF8 wMxw+4ThCksvj/qsuB5XTEcuMWxYwswDltDOgReov9DVWCKuC/7wzU7yg+7vKi0n IZZrh+TXpx0aYbgG9toGyGeh6XLRXAnyPFFXUkr+e1G1NiSot07Ul4p1lf9m1GeG XULBL17L3EzhN2w+bMk5g== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786102638; x=1786189038; bh=A JEIXCOdsK1VdJ+vTB/+b5PTVsVX+FRzDaAsHSTP7IA=; b=JmtyROkD4PTSKr6Ym 7sg+dQQSTmcCqEp4Ml+Q8hQb8UgYCm9848RuqaVbWsV7UUih68O//+dnchC1G9zO Gjyd6YDThoDcmz1PmrMtCOC1aTJ7zUQUyjbBM23kvrKKgrIZgVfRFLAhn0Fz1zHP i+dvA9zf3sa+wRbMWwEG0wiQvY1iigRzyidsVG+zmFC346d35R7ZyaFmFrOREJfe KPa3+ug4bZGLhNEAjRHj80VoD4TbcvGRd1TeiqU8ctCD7nbuP/9xC1Pz91ogmMsO PjRGeeg5UQ1UlXemhAIft2rK4WNo5BeH3CuzR8tDDsS2nIpurNcfvSiX4yCNTgZl f3OrQ== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEfW3x2Qrq7ELqxlr/x69vk0aYaa/E/2AWrFAl7yu0Vg+jIGlS0mjOCk5J5e5/NJH T3J+l8M+cQ15gxg/u6FQhwxKudsrf7tN5cWcEThveGDS0KcOU0uMrzGx0YBmylEV4PfT2X SbTqClZY3SINVohR/2NzYRxeOu5tU+SgKqY1sS5oM8U15+ky277S+G1Ww7YAUM5Z6EK46p TQK1COWwyEQl/dhwe51H5lWCLFzTj8K9hkk/JBO/DvOgEgoe5sMihzFP04eKTIrZCF8dte tQwhoNU+zszgTlPncqprVsvTwAgpjyPciGBDZXKQScTz4u1orAuTZg8Fi2VomZsbF2qQLb zIB6vfCNv4v0dcseiUkN3H0nffdrzRAjDXnr/zcE4pHIWwhQSIJcHpNKfPSSatE5vtlSwn DtWUIKhvp1jGXwuxT4icZJ7eqOMoDp9O3eN/sGEFk1QcUkAUefbMoY+f6cBb/eCJOZp5Fv w3A5qOkKNQDIKdjK3fDOPysd5GAuUjiWQZW+5CYr0Pg9TIHi8EHfQKgk6Ip7WUVR5HqKPp DudH9yV/0Ag4i+mp5KiXkaccnRZFhzimORXQyyVo6QHCrlCjus7skssIHkNSCl2guqkiYo UOZv3k9Fp2gNbtDp0Cw80/IB3MeBJb4AXGUJW88X/CQw/EcsEbAANNWBqb5Q X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 7 Aug 2026 07:37:18 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v2 11/17] selftests/mm: cover a shared-source collapse write race Date: Fri, 7 Aug 2026 12:36:41 +0100 Message-ID: <20260807113647.3744609-12-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260807113647.3744609-1-kirill@shutemov.name> References: <20260807113647.3744609-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" collapse_fork checks that a fork-shared range collapses in the process that asks for it while the co-sharer keeps its own page, but the co-sharer sits still while that happens. Add a case where the co-sharer writes to the shared range throughout the collapse. CoW has to keep the two sides apart under those writes: the collapsing child must see the content from before the fork, and the writing parent must see only its own writes. It passes on an unmodified kernel, so it pins down isolation that khugepaged collapse already provides. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) Tested-by: Muhammad Usama Anjum --- tools/testing/selftests/mm/khugepaged.c | 69 +++++++++++++++++++++++++ 1 file changed, 69 insertions(+) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index d3ff1b196e09..a5683f08694f 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -1225,6 +1225,72 @@ static void collapse_max_ptes_shared(struct collapse= _context *c, struct mem_ops ksft_test_result_report(exit_status, "%s\n", __func__); } =20 +/* + * Content stays isolated while a co-sharer writes concurrently. A shared + * source is copied live (not frozen), relying on it being CoW - immutable + * for the duration of the copy; a co-sharer's write goes to a CoW copy. T= he + * collapsing child must see the pre-fork content, the writing parent only + * its own writes. + */ +static void collapse_fork_cow_race(struct collapse_context *c, struct mem_= ops *ops) +{ + const unsigned long shared =3D 64 * page_size; + const int stride =3D page_size / sizeof(int); + int wstatus, child_status, i, n =3D shared / page_size; + /* volatile: the loop below must really store, on every iteration */ + volatile int *ip; + void *p; + + p =3D ops->setup_area(1); + ip =3D p; + ops->fault(p, 0, shared); /* shared prefix, pre-fork pattern */ + + ksft_print_msg("Fork, collapse in the child while the parent rewrites..."= ); + if (!fork()) { + int collapse_status; + + ops->fault(p, shared, hpage_pmd_size); /* private remainder */ + c->collapse("Collapse a range shared with a writing co-sharer", + p, 1, ops, true); + collapse_status =3D exit_status; + for (i =3D 0; i < n; i++) + if (ip[i * stride] !=3D i + 0xdead0000) + break; + if (i =3D=3D n) + success("OK"); + else + fail("Fail: child content"); + /* The content check must not bury a failed collapse. */ + if (exit_status !=3D KSFT_FAIL) + exit_status =3D collapse_status; + ops->cleanup_area(p, hpage_pmd_size); + _exit(exit_status); + } + + /* Hammer the parent's own writes over the shared prefix. */ + for (int it =3D 0; it < 200000; it++) + for (i =3D 0; i < n; i++) + ip[i * stride] =3D i + 0xbeef0000; + + wait(&wstatus); + /* A child that died reading the racing pages is a failure, not a zero. */ + child_status =3D WIFEXITED(wstatus) ? WEXITSTATUS(wstatus) : KSFT_FAIL; + + ksft_print_msg("Check the parent sees only its own writes..."); + for (i =3D 0; i < n; i++) + if (ip[i * stride] !=3D i + 0xbeef0000) + break; + if (i =3D=3D n) + success("OK"); + else + fail("Fail: parent content"); + ops->cleanup_area(p, hpage_pmd_size); + /* Same again: our own check must not bury the child's verdict. */ + if (exit_status !=3D KSFT_FAIL) + exit_status =3D child_status; + ksft_test_result_report(exit_status, "%s\n", __func__); +} + static void madvise_collapse_existing_thps(struct collapse_context *c, struct mem_ops *ops) { @@ -1778,6 +1844,9 @@ int main(int argc, char **argv) TEST(collapse_max_ptes_shared, khugepaged_context, anon_ops); TEST(collapse_max_ptes_shared, madvise_context, anon_ops); =20 + TEST(collapse_fork_cow_race, khugepaged_context, anon_ops); + TEST(collapse_fork_cow_race, madvise_context, anon_ops); + TEST(madvise_collapse_existing_thps, madvise_context, anon_ops); TEST(madvise_collapse_existing_thps, madvise_context, read_only_file_ops); TEST(madvise_collapse_existing_thps, madvise_context, read_write_file_rea= d_ops); --=20 2.54.0 From nobody Tue Sep 29 13:20:38 2026 Received: from fhigh-b6-smtp.messagingengine.com (fhigh-b6-smtp.messagingengine.com [202.12.124.157]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CD2A4472F8C; Fri, 7 Aug 2026 11:37:24 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.157 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102649; cv=none; b=j7LCePl1xBh6XmZEX9T0HS6x+Cau4+ZC+p3K0/6SnORIoon88Cq2WaX0jWIqmiGmZn6SkhtJ90j7ubwe0DuyAOV3G4KiuDO0rF25YOugRcfKken3v4tR3bmAlUiHQM35HmAz76Yj83CwdrdZyW+ClLGlQdfmG28yN0IK+GMx0eQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102649; c=relaxed/simple; bh=jZYp90GmHJmWv/lPXeRlskGLB+yopqApbDcFj2jQR5E=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=d43JyyoZkb9u+SKOdcsS3kaD1RnrhkMxNYZlLH6sqarWglv5JN5gA8VvtIW1aaj5yhUl0ESqdbTpyYPyP9V8uk4/tHA/al0sntzWqjCbYCC5OojwQPC98pIMaXvBpNANyGhJNeBWxkNbThxt8Vy+iBGqmOfFMDc1cReddBWjr0w= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=uxQ/dJq3; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=lDWQSBXl; arc=none smtp.client-ip=202.12.124.157 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="uxQ/dJq3"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="lDWQSBXl" Received: from phl-compute-02.internal (phl-compute-02.internal [10.202.2.42]) by mailfhigh.stl.internal (Postfix) with ESMTP id 617787A00AC; Fri, 7 Aug 2026 07:37:21 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-02.internal (MEProxy); Fri, 07 Aug 2026 07:37:22 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786102641; x= 1786189041; bh=8tAM9g0684oRzmvSRuRjYFDutiQh5p3+9TYSRlUK8Mw=; b=u xQ/dJq3m96Z/ZX7AptbBwSbJbBFxJ609byOlEM2pwXJd1CxA+YRqkzFU+5NMyjpf gyRQItI7A2JW9deYSn8Ij3P+7FqFQBeBR4j9D1IfnE/Etv4tvqOmMhBWo7NBivJB BKVnzMyy/Euv2UHcqssGhRXX8y05vqyO01VVXl3lOTw1WtptB7dyIXdabtWTLqbj 33YWyYUlmTxL8uSHLNPqhoZxPr+EDUmabcLDwR4HxR0JwSWzaDazCp4b1F9or2ub e2+2tsChoQUvEOiETCbW7LYDZGcJjGVWSWTXcwIiapQu5IxqzdWbjMnHQfRA/fAg RXroHxsm6YKiUuj5dvq6w== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786102641; x=1786189041; bh=8 tAM9g0684oRzmvSRuRjYFDutiQh5p3+9TYSRlUK8Mw=; b=lDWQSBXldyOvTowSo o6jgucV1dVhi0gMv74hLWH5exR77eRY65K9uwEA78XGs1C0YTXWpPjJsVJDKIK0d A+VyamLVS/IrMVfAjVb7nGiEggCrX/7TRZCWDfqsHEDqndr7/88KwmDp+ijsWRp1 JC7ijWRQgR56yrmbnXCuFF0mYW2A2KqLrrkFzhByoWx2qUaAmKPty8bbNKvTKlX0 P9OAPuBZH3Jrg/8TFOkJ3toquxM/8+OEy/CUm57w6S/UiH707xctWKX68rDjA1pq lZqepM420O39dDyKrdbA3Ius3U9LnJRWS3rXjyUXihv36zznVYYMCSRu7mxmT/5E xbYAA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFHgVqOo5hmjJljb3IAZGkXyyxzQaD6i5RNmNk424gVmZkPzr1NfNWN42YMHBLBJ7 ouwgSkV7uP+O116AnyQ77X3VN5vAdqq1zvzV33EYJRHUa/cQ0OT9UF9tJbf4RJE1rh86et dgK/XsSmEjRvIZ4/DwnHkulf5IWkeizHGDAErUJF9mnMadsJqDG0SAHuOjgfipI02yAqvF a9Gp7D1P4aM7ojieUjy73reTq6pHjAlslJzhWF2Bayc7rzuN1loFnb3tH1A/9l8/u/BKz2 kEn4hJCLyjrltgNf/ADveISAZdc/G8K4aY+0IiJqAeENTN6qSce54NhN3knOmRl67eS+iD DB/QX6xfKcr5RTHe3Fwx41XsDl2llue/0k0iHG3Lz7tcse7fjrM4kmFD8x/ltDWFR0S/6j g3jExQ3/y710ZEVdf1dY9K2eOWeoheh/bH9sPncAG9GLM5pJu54o3R4VujgiqolMtPX0X6 VAUv9x8ygaHlyqBsODovA5QpLCdyzh6DJdbfrASOOyXwujHMs54A+S+kcowCkvEM/m8zUb XrRjdcnMoWOMC6MRih1xZPflaVhp0PV4JZIZSU3B0R9lN6jFv0whY8PKf9CAAP5ZtbaCtC 3nzgoZtL52vh1X/MauVErS5FJMqE0ozjCuUuESABlSc0ZsELK58+Z1nmeGPg X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 7 Aug 2026 07:37:20 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v2 12/17] selftests/mm: run every supported collapse order by default Date: Fri, 7 Aug 2026 12:36:42 +0100 Message-ID: <20260807113647.3744609-13-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260807113647.3744609-1-kirill@shutemov.name> References: <20260807113647.3744609-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The mTHP collapse cases only run when the caller names both the context and an order, so a plain ./khugepaged covers the PMD contexts on anon and nothing else. run_vmtests.sh pinned order 4 and covered no other. Without -c, run the mTHP cases once per supported anon THP order below the PMD, and include that context in both the no-argument invocation and "all". -c still pins one order, and now says what is wrong when the order is above the PMD or unsupported instead of printing the usage text. An order at or below the -s source order is skipped: the sources would already be the size being asked for. Both orders end up as array indices and shift counts, so -s rejects a negative order and -c anything at or below 0, rather than letting either reach them. The mTHP context has only anon cases, so a run that names a different mem_type -- "all:shmem", say -- drops it again rather than refusing to start. Naming both explicitly still does. A case carries the order it was registered at, so a result names it: # Run test: collapse_order_max_ptes_none (mthp_khugepaged:anon, order 6) On x86-64 with 4K pages that is orders 2 through 8, and ./khugepaged goes from 28 results to 77 in 21 seconds, so run_vmtests.sh can drop its pinned order-4 line. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) Tested-by: Muhammad Usama Anjum --- tools/testing/selftests/mm/khugepaged.c | 111 ++++++++++++++++++---- tools/testing/selftests/mm/run_vmtests.sh | 2 - 2 files changed, 91 insertions(+), 22 deletions(-) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index a5683f08694f..a202fe359cf6 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -31,6 +31,9 @@ static unsigned long page_size; static int hpage_pmd_nr; static int anon_order; static int collapse_order; +static bool collapse_order_given; +static int collapse_orders[NR_ORDERS]; +static int nr_collapse_orders; static int pagemap_fd =3D -1; static int kpageflags_fd =3D -1; =20 @@ -1556,6 +1559,7 @@ static void usage(void) fprintf(stderr, "\t\t-s: mTHP size, expressed as page order.\n"); fprintf(stderr, "\t\t Defaults to 0. Use this size for anon or shmem a= llocations.\n"); fprintf(stderr, "\t\t-c: collapse order for mTHP collapse, expressed as p= age order.\n"); + fprintf(stderr, "\t\t Defaults to every supported order below the PMD.= \n"); fprintf(stderr, "\t\t With -s, -s names the mTHP source order for the\= n"); fprintf(stderr, "\t\t mixed-source case (source order below the target= ).\n"); exit(1); @@ -1563,6 +1567,7 @@ static void usage(void) =20 static void parse_test_type(int argc, char **argv) { + bool mthp_context_implied =3D false; int opt; char *buf; const char *token; @@ -1574,6 +1579,7 @@ static void parse_test_type(int argc, char **argv) break; case 'c': collapse_order =3D atoi(optarg); + collapse_order_given =3D true; break; case 'h': default: @@ -1581,12 +1587,23 @@ static void parse_test_type(int argc, char **argv) } } =20 + /* + * Both orders end up as array indices and shift counts, so neither + * can be negative, and a zero collapse order asks for base pages. + */ + if (anon_order < 0) + ksft_exit_fail_msg("-s takes an order, which cannot be negative\n"); + if (collapse_order_given && collapse_order <=3D 0) + ksft_exit_fail_msg("-c takes an order above 0, not %d\n", + collapse_order); + argv +=3D optind; argc -=3D optind; =20 if (argc =3D=3D 0) { - /* Backwards compatibility */ + /* Everything that needs no argument of its own: anon, every context */ khugepaged_context =3D &__khugepaged_context; + mthp_khugepaged_context =3D &__mthp_khugepaged_context; madvise_context =3D &__madvise_context; anon_ops =3D &__anon_ops; return; @@ -1597,13 +1614,19 @@ static void parse_test_type(int argc, char **argv) =20 if (!strcmp(token, "all")) { khugepaged_context =3D &__khugepaged_context; + mthp_khugepaged_context =3D &__mthp_khugepaged_context; madvise_context =3D &__madvise_context; + + /* + * "all" sweeps the mTHP context in, but it only has anon + * cases: step it aside for the other mem_types rather than + * refusing the whole run. + */ + mthp_context_implied =3D true; } else if (!strcmp(token, "khugepaged")) { khugepaged_context =3D &__khugepaged_context; } else if (!strcmp(token, "mthp_khugepaged")) { mthp_khugepaged_context =3D &__mthp_khugepaged_context; - if (collapse_order <=3D 0 || collapse_order >=3D hpage_pmd_order) - usage(); } else if (!strcmp(token, "madvise")) { madvise_context =3D &__madvise_context; } else { @@ -1619,20 +1642,20 @@ static void parse_test_type(int argc, char **argv) read_write_file_write_ops =3D &__read_write_file_write_ops; anon_ops =3D &__anon_ops; shmem_ops =3D &__shmem_ops; - if (mthp_khugepaged_context) - usage(); } else if (!strcmp(buf, "anon")) { anon_ops =3D &__anon_ops; } else if (!strcmp(buf, "file")) { read_only_file_ops =3D &__read_only_file_ops; read_write_file_read_ops =3D &__read_write_file_read_ops; read_write_file_write_ops =3D &__read_write_file_write_ops; - if (mthp_khugepaged_context) + if (mthp_khugepaged_context && !mthp_context_implied) usage(); + mthp_khugepaged_context =3D NULL; } else if (!strcmp(buf, "shmem")) { shmem_ops =3D &__shmem_ops; - if (mthp_khugepaged_context) + if (mthp_khugepaged_context && !mthp_context_implied) usage(); + mthp_khugepaged_context =3D NULL; } else { usage(); } @@ -1654,9 +1677,14 @@ struct test_case { struct mem_ops *ops; const char *desc; test_fn fn; + int order; /* mTHP contexts: the collapse order */ }; =20 -#define MAX_TEST_CASES 64 +/* + * Enough for every case at every order the kernel offers: the mTHP context + * runs its cases once per supported order below the PMD. + */ +#define MAX_TEST_CASES 256 static struct test_case test_cases[MAX_TEST_CASES]; static int nr_test_cases; =20 @@ -1669,6 +1697,7 @@ static int nr_test_cases; .ops =3D o, \ .desc =3D #t, \ .fn =3D t, \ + .order =3D collapse_order, \ }; \ } \ } while (0) @@ -1709,10 +1738,40 @@ int main(int argc, char **argv) =20 parse_test_type(argc, argv); =20 - if (mthp_khugepaged_context && - !(thp_supported_orders() & (1UL << collapse_order))) - ksft_exit_skip("Order %d is not a supported anon THP order\n", - collapse_order); + if (mthp_khugepaged_context) { + unsigned long orders =3D thp_supported_orders(); + + if (collapse_order_given) { + /* -c pins one order; it has to be one we can build */ + if (collapse_order >=3D hpage_pmd_order) + ksft_exit_fail_msg("-c takes an order below the PMD order (%d)\n", + hpage_pmd_order); + if (!(orders & (1UL << collapse_order))) + ksft_exit_skip("Order %d is not a supported anon THP order\n", + collapse_order); + if (collapse_order <=3D anon_order) + ksft_exit_skip("-c %d needs a source order below it, -s says %d\n", + collapse_order, anon_order); + collapse_orders[nr_collapse_orders++] =3D collapse_order; + } else { + /* + * Otherwise every order a collapse could produce. -s + * makes the fault path hand out folios of that order, + * so a target at or below it has nothing to collapse: + * the sources are already the size being asked for. + */ + int first =3D anon_order ? anon_order + 1 : MIN_MTHP_ORDER; + + if (first < MIN_MTHP_ORDER) + first =3D MIN_MTHP_ORDER; + for (int i =3D first; i < hpage_pmd_order; i++) { + if (orders & (1UL << i)) + collapse_orders[nr_collapse_orders++] =3D i; + } + if (!nr_collapse_orders) + ksft_print_msg("mTHP cases skipped: no order above the source\n"); + } + } =20 if (mthp_khugepaged_context) { pagemap_fd =3D open("/proc/self/pagemap", O_RDONLY); @@ -1773,7 +1832,17 @@ int main(int argc, char **argv) TEST(collapse_full, khugepaged_context, read_write_file_read_ops); TEST(collapse_full, khugepaged_context, read_write_file_write_ops); TEST(collapse_full, khugepaged_context, shmem_ops); - TEST(collapse_full, mthp_khugepaged_context, anon_ops); + for (int i =3D 0; i < nr_collapse_orders; i++) { + collapse_order =3D collapse_orders[i]; + TEST(collapse_full, mthp_khugepaged_context, anon_ops); + TEST(collapse_empty, mthp_khugepaged_context, anon_ops); + TEST(collapse_single_mthp, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_single_window, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_partial_window, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_max_ptes_none, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_mixed_sources, mthp_khugepaged_context, anon_ops); + } + TEST(collapse_full, madvise_context, anon_ops); TEST(collapse_full, madvise_context, read_only_file_ops); TEST(collapse_full, madvise_context, read_write_file_read_ops); @@ -1781,14 +1850,8 @@ int main(int argc, char **argv) TEST(collapse_full, madvise_context, shmem_ops); =20 TEST(collapse_empty, khugepaged_context, anon_ops); - TEST(collapse_empty, mthp_khugepaged_context, anon_ops); TEST(collapse_empty, madvise_context, anon_ops); =20 - TEST(collapse_single_mthp, mthp_khugepaged_context, anon_ops); - TEST(collapse_order_single_window, mthp_khugepaged_context, anon_ops); - TEST(collapse_order_partial_window, mthp_khugepaged_context, anon_ops); - TEST(collapse_order_max_ptes_none, mthp_khugepaged_context, anon_ops); - TEST(collapse_order_mixed_sources, mthp_khugepaged_context, anon_ops); =20 TEST(collapse_single_pte_entry, khugepaged_context, anon_ops); TEST(collapse_single_pte_entry, khugepaged_context, read_only_file_ops); @@ -1862,7 +1925,15 @@ int main(int argc, char **argv) for (int i =3D 0; i < nr_test_cases; i++) { struct test_case *t =3D &test_cases[i]; =20 - ksft_print_msg("\n# Run test: %s (%s:%s)\n", t->desc, t->ctx->name, t->o= ps->name); + if (t->ctx =3D=3D &__mthp_khugepaged_context) { + collapse_order =3D t->order; + ksft_print_msg("\n# Run test: %s (%s:%s, order %d)\n", + t->desc, t->ctx->name, t->ops->name, + t->order); + } else { + ksft_print_msg("\n# Run test: %s (%s:%s)\n", t->desc, + t->ctx->name, t->ops->name); + } t->fn(t->ctx, t->ops); } =20 diff --git a/tools/testing/selftests/mm/run_vmtests.sh b/tools/testing/self= tests/mm/run_vmtests.sh index 2652a7920b80..8bf898b71350 100755 --- a/tools/testing/selftests/mm/run_vmtests.sh +++ b/tools/testing/selftests/mm/run_vmtests.sh @@ -412,8 +412,6 @@ CATEGORY=3D"thp" run_test ./khugepaged all:shmem =20 CATEGORY=3D"thp" run_test ./khugepaged -s 4 all:shmem =20 -CATEGORY=3D"thp" run_test ./khugepaged -c 4 mthp_khugepaged:anon - # Try to create XFS if not provided if [ -z "${SPLIT_HUGE_PAGE_TEST_XFS_PATH}" ]; then if test_selected "thp"; then --=20 2.54.0 From nobody Tue Sep 29 13:20:38 2026 Received: from fhigh-b6-smtp.messagingengine.com (fhigh-b6-smtp.messagingengine.com [202.12.124.157]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C37DE472F6C; Fri, 7 Aug 2026 11:37:27 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.157 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102652; cv=none; b=D3OD464sTnqbeRwXj/uZxgrFgWKWjc+VT4gvtYUHPsKEeIGZCQNKiHR/3AC/d93Pu4A3n2qg8vyhWgYS/VNFeBC1/N0ymVHTUbzHeWRpykA8PX92DIUnIXszmdhLf9dwgumf6c0oBBCThg9PjzCONEh0QxRTjcwoW8ABqLsgiOM= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102652; c=relaxed/simple; bh=yUZcBoiVopGsKNkpsCC5OH4xXwc0qA8AOnln8opJ+8s=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=u7TLtvBeXlaN0EikZJaWWoSgvyG8OoYdfRCTAbzSHLy/FVngFa2wAw7vxJByZ0y+bXZzi558JGbPHqKzLRmdbec4HtiLwHSQwPSln2TPQ/InOilY/jkLuMmOiANbvisZr5+KtRKxgH1WWp0pslWYcPvGjOelvpt20KVKlWLvPc4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=zSX7pb7X; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=YWAKNyIS; arc=none smtp.client-ip=202.12.124.157 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="zSX7pb7X"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="YWAKNyIS" Received: from phl-compute-04.internal (phl-compute-04.internal [10.202.2.44]) by mailfhigh.stl.internal (Postfix) with ESMTP id BB7F77A004F; Fri, 7 Aug 2026 07:37:23 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-04.internal (MEProxy); Fri, 07 Aug 2026 07:37:24 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786102643; x= 1786189043; bh=O73xqgElWTRjAuAEasIQOV2g23H+aaWw3pIMG8sIcFk=; b=z SX7pb7X/Hd4cG6bFGLFdaLKUtrR76UhlUeLI8oFHBcIlRvg7RYT7Y3Zv1fQJmyUk FTAjLBvfO8zlrVcm4sSiblWC6d087F0ShHwkXSpmvT22i1RTw36MLfhaMoc47Dot Cffz5JpQbzZ4IHOagsAbcdJdyGcNH5/iYbgB0lHER88a03SfbeoN9USxMPqEr/1X KntEONQl6/2t6k77AkLTIWchMWn49/s920jZjazrnWrU7n8BUYFPcWk36QmjnwRD ngTsT9FRaMflUBMFw4+Hwbsxwl/B67he0HknP/N3s0tG2hXbfy4D+Oadc/BdIyuU oZuXngyo+83RYtJubcrHw== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786102643; x=1786189043; bh=O 73xqgElWTRjAuAEasIQOV2g23H+aaWw3pIMG8sIcFk=; b=YWAKNyISaDOTeGIM4 z1wrdYyoArya8DnZ2ka2gv/zNmyligxfIdWKS4iZ2HzWGifKmHA+rU+m88d8jWI8 fatHGGn29FRFyOX33oyrheN2pNubJZayrF+9pTI4TuS3yoybgIg5YPMuZ67Yu/Aq 6gzZXx2nTveHMpxVIhS+R9sQXPpbcvBnZhATHRlubpt3RdYph5F2Se0wWF8a0hNO VvjiRVR5jhtUBj5XxLmFd5bQtJZVZ6tEx5tjU42GjZQAltENOWHJ946feeEAb7Zt WcRiR9f0GI9LRgN7foWlhC4G61bXSmUfBTHpv0jmRf8zU//0bMU4tmOAdSJPmY6Q MI+3w== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGUERIwmKnJzNqrmH6VlEmVcWeiClBef7dL76l3JFxgCJIE1wbu/wgjcLR687OyMh BQrMCOBv45wH7h3ooBeneK7mWIj7iIytnpxBlgzGmz5SHufCsdgROAqJKtMeokx5eB2AAp cXT5mMO8imvID7jHzWhqIClk07P/42mlBI32M882inm3UlyT/GsY3qxbman5/VlsuJS9+T F+1+aXRxbMzBOR7RkHQ9dyf5K4OhFzw16ipJCuGwyYR0rlnsi17XpZoH3bdrAT++N0O7xU IF8fSnfJTEWFFODGe4CSmBZThmf/6wAYTC6d1nWOkM6O/FhjItWthf5fql0LkcaRcJIWRy JL+hFZ/gRMyTy77bn7guBmdAZvK76PZDbgYkLybrKwaNUAwJZRgpJcW4Y1Ta6Bs1FhHCfC WNK4ldNl+Sl5ROCrlffbw2Nwk4eNzkvn/GYmoatOXdM5AGssgPqrX4ZRtjnEAJsbpSncI7 3X5mXHHf239KzDMdwhTKRTejE5fnzoDYmOV1h9bjUTiZMVsuOskqqIDZEe21A1CieBok8o pT6kONFMBQYTtNk5qo6dcOQFnVTiSwHUseu1hDWjq9Byeq9hbUTR9dH+XCVi0LEyk41xAU YHzV3i7wUi+HL2QmKm60uJy2rOtcdb4TxqNFJBUYVFvJYoO2rPLNr9pSN9YA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 7 Aug 2026 07:37:23 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v2 13/17] selftests/mm: verify synchronous khugepaged driving is attributable Date: Fri, 7 Aug 2026 12:36:43 +0100 Message-ID: <20260807113647.3744609-14-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260807113647.3744609-1-kirill@shutemov.name> References: <20260807113647.3744609-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The khugepaged tests attribute outcomes through the huge_memory tracepoints. The anon-path events carry no virtual address, but mm_collapse_huge_page_isolate() reports a source folio PFN and order, and the test matches those against the PFNs it read from pagemap before the pass. An attempt counts from whichever signal fires, so the check does not depend on which anon tracepoint a kernel emits. Add small tracefs helpers to vm_util and khugepaged_sync_check: per step, prepare one aligned window, record its source PFNs, run one khugepaged_full_pass() barrier, and require the window collapsed with exactly one attributed attempt. scan_sleep_millisecs is set high, so the test only finishes in time if the sysfs store really wakes the daemon. Passes 5/5 on x86-64 4K and arm64 64K. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) Tested-by: Muhammad Usama Anjum --- tools/testing/selftests/mm/Makefile | 1 + .../selftests/mm/khugepaged_sync_check.c | 217 ++++++++++++++++++ tools/testing/selftests/mm/run_vmtests.sh | 2 + tools/testing/selftests/mm/vm_util.c | 41 ++++ tools/testing/selftests/mm/vm_util.h | 4 + 5 files changed, 265 insertions(+) create mode 100644 tools/testing/selftests/mm/khugepaged_sync_check.c diff --git a/tools/testing/selftests/mm/Makefile b/tools/testing/selftests/= mm/Makefile index 2093fcf6e915..b2d6e5c12934 100644 --- a/tools/testing/selftests/mm/Makefile +++ b/tools/testing/selftests/mm/Makefile @@ -105,6 +105,7 @@ TEST_GEN_FILES +=3D merge TEST_GEN_FILES +=3D rmap TEST_GEN_FILES +=3D folio_split_race_test TEST_GEN_FILES +=3D folio_order_check +TEST_GEN_FILES +=3D khugepaged_sync_check =20 ifneq ($(ARCH),arm64) TEST_GEN_FILES +=3D soft-dirty diff --git a/tools/testing/selftests/mm/khugepaged_sync_check.c b/tools/tes= ting/selftests/mm/khugepaged_sync_check.c new file mode 100644 index 000000000000..30d3fb519fb2 --- /dev/null +++ b/tools/testing/selftests/mm/khugepaged_sync_check.c @@ -0,0 +1,217 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Synchronous khugepaged driving check. + * + * Race tests drive khugepaged through the existing sysfs controls: a + * store to scan_sleep_millisecs wakes the daemon, and full_scans + * advancing by two is a completion barrier for one full pass that + * started after setup (khugepaged_full_pass()). Verify the pair gives + * deterministic, attributable results: one barrier step over one + * prepared window produces exactly one collapse attempt on that + * window's source pages (mm_collapse_huge_page_isolate events filtered + * by source PFN and order) and the window is collapsed + * afterwards, repeatably. + * + * scan_sleep_millisecs is set to 60s to prove the wake path: without + * the wake, one barrier step would sleep multiples of that and blow + * the timeout. It also keeps the daemon from free-running between + * steps, per the khugepaged_full_pass() discipline. + */ +#define _GNU_SOURCE +#include +#include +#include +#include +#include +#include + +#include "kselftest.h" +#include "vm_util.h" +#include "hugepage_settings.h" + +#define BASE_ADDR ((void *)(1UL << 30)) +#define TARGET_ORDER 2 /* smallest order khugepaged considers */ +#define NR_ITERATIONS 5 + +static int pagemap_fd; +static int kpageflags_fd; +static int trace_events_fd =3D -1; +static unsigned long hpage_pmd_size; + +/* + * Each step switches the events off again, but a helper can still give up + * on us in between (a failing sysfs write ends the test from inside + * thp_write_num()), and huge_memory events left on are the whole machine's + * problem, not this test's. + */ +static void trace_events_off(void) +{ + if (trace_events_fd >=3D 0) + tracing_events_enable(trace_events_fd, false); +} + +/* + * Count collapse attempts attributable to our window: legacy-engine + * isolate events whose scan_pfn is one of the window's source PFNs, + * plus batch-engine per-candidate install events at the window's + * address. Either engine reports exactly once per attempt. + */ +static int count_attributed(unsigned long *pfns, int nr_pfns, + unsigned long addr, unsigned int order) +{ + char line[1024]; + int count =3D 0; + FILE *fp; + + fp =3D tracing_open_trace(); + if (!fp) + ksft_exit_fail_msg("Cannot open trace buffer\n"); + + while (fgets(line, sizeof(line), fp)) { + char *s; + unsigned long val; + unsigned int ord; + char *o; + int i; + + s =3D strstr(line, "mm_collapse_huge_page_isolate:"); + if (s) { + if (sscanf(s, "mm_collapse_huge_page_isolate: scan_pfn=3D0x%lx", + &val) !=3D 1) + continue; + o =3D strstr(s, "order=3D"); + if (!o || sscanf(o, "order=3D%u", &ord) !=3D 1 || + ord !=3D order) + continue; + for (i =3D 0; i < nr_pfns; i++) { + if (val =3D=3D pfns[i]) { + count++; + break; + } + } + continue; + } + + s =3D strstr(line, "mm_collapse_candidate:"); + if (s) { + if (!strstr(s, "pass=3Dinstall") || + !strstr(s, "result=3Dsucceeded")) + continue; + o =3D strstr(s, "addr=3D"); + if (!o || sscanf(o, "addr=3D0x%lx", &val) !=3D 1 || + val !=3D addr) + continue; + o =3D strstr(s, "order=3D"); + if (!o || sscanf(o, "order=3D%u", &ord) !=3D 1 || + ord !=3D order) + continue; + count++; + } + } + fclose(fp); + return count; +} + +static void one_step(int iteration) +{ + const size_t window =3D getpagesize() << TARGET_ORDER; + const int nr_pages =3D 1 << TARGET_ORDER; + unsigned long pfns[1 << TARGET_ORDER]; + bool collapsed, passed; + int attributed; + char *p; + int i; + + p =3D mmap(BASE_ADDR, hpage_pmd_size, PROT_READ | PROT_WRITE, + MAP_ANONYMOUS | MAP_PRIVATE | MAP_FIXED_NOREPLACE, -1, 0); + if (p !=3D BASE_ADDR) + ksft_exit_fail_perror("mmap() window"); + + /* Prepare one window; record its source PFNs. */ + for (i =3D 0; i < nr_pages; i++) { + p[i * getpagesize()] =3D i + 1; + pfns[i] =3D pagemap_get_pfn(pagemap_fd, p + i * getpagesize()); + if (pfns[i] =3D=3D -1UL) + ksft_exit_fail_msg("Source page not present\n"); + } + + /* Clear first: with the events still off there is nothing to undo. */ + if (tracing_clear_trace()) + ksft_exit_fail_msg("Cannot clear the trace buffer\n"); + if (tracing_events_enable(trace_events_fd, true)) + ksft_exit_fail_msg("Cannot enable huge_memory events\n"); + + madvise(p, hpage_pmd_size, MADV_HUGEPAGE); + /* Wait up to 120 seconds for the pass to complete. */ + passed =3D khugepaged_full_pass(120); + + /* Off before anything that can give up: the events are system-wide. */ + if (tracing_events_enable(trace_events_fd, false)) + ksft_exit_fail_msg("Cannot disable huge_memory events\n"); + if (!passed) + ksft_exit_fail_msg("khugepaged did not complete a full pass\n"); + + collapsed =3D is_range_backed_by_folio_orders(p, window, TARGET_ORDER, + pagemap_fd, kpageflags_fd); + attributed =3D count_attributed(pfns, nr_pages, (unsigned long)p, + TARGET_ORDER); + + ksft_test_result(collapsed && attributed =3D=3D 1, + "step %d: window collapsed, %d attributed result(s)\n", + iteration, attributed); + + munmap(p, hpage_pmd_size); +} + +int main(void) +{ + struct thp_settings settings; + int i; + + ksft_print_header(); + + if (!thp_available()) + ksft_exit_skip("Transparent Hugepages not available\n"); + if (!(thp_supported_orders() & (1UL << TARGET_ORDER))) + ksft_exit_skip("Order %d is not a supported anon THP order\n", + TARGET_ORDER); + + hpage_pmd_size =3D read_pmd_pagesize(); + if (!hpage_pmd_size) + ksft_exit_fail_msg("Reading PMD pagesize failed\n"); + pagemap_fd =3D open("/proc/self/pagemap", O_RDONLY); + if (pagemap_fd < 0) + ksft_exit_fail_perror("open(/proc/self/pagemap)"); + kpageflags_fd =3D open("/proc/kpageflags", O_RDONLY); + if (kpageflags_fd < 0) + ksft_exit_skip("open(\"/proc/kpageflags\") requires root\n"); + trace_events_fd =3D tracing_events_open("huge_memory"); + if (trace_events_fd < 0) + ksft_exit_skip("huge_memory events require tracefs and root\n"); + atexit(trace_events_off); + + ksft_set_plan(NR_ITERATIONS); + + thp_save_settings(); + thp_read_settings(&settings); + settings.thp_enabled =3D THP_MADVISE; + settings.thp_defrag =3D THP_DEFRAG_ALWAYS; + settings.khugepaged.defrag =3D 1; + settings.khugepaged.scan_sleep_millisecs =3D 60000; + settings.khugepaged.alloc_sleep_millisecs =3D 60000; + settings.khugepaged.max_ptes_none =3D (hpage_pmd_size / getpagesize()) - = 1; + /* One wake must complete one full pass; see khugepaged_full_pass(). */ + settings.khugepaged.pages_to_scan =3D 1UL << 24; + for (i =3D 0; i < NR_ORDERS; i++) + settings.hugepages[i].enabled =3D THP_NEVER; + settings.hugepages[TARGET_ORDER].enabled =3D THP_INHERIT; + /* Base of the settings stack; the bottom entry is never popped. */ + thp_push_settings(&settings); + + for (i =3D 0; i < NR_ITERATIONS; i++) + one_step(i); + + thp_restore_settings(); + + ksft_finished(); +} diff --git a/tools/testing/selftests/mm/run_vmtests.sh b/tools/testing/self= tests/mm/run_vmtests.sh index 8bf898b71350..c0f69da3fd3b 100755 --- a/tools/testing/selftests/mm/run_vmtests.sh +++ b/tools/testing/selftests/mm/run_vmtests.sh @@ -404,6 +404,8 @@ CATEGORY=3D"cow" run_test ./cow =20 CATEGORY=3D"thp" run_test ./folio_order_check =20 +CATEGORY=3D"thp" run_test ./khugepaged_sync_check + CATEGORY=3D"thp" run_test ./khugepaged =20 CATEGORY=3D"thp" run_test ./khugepaged -s 2 diff --git a/tools/testing/selftests/mm/vm_util.c b/tools/testing/selftests= /mm/vm_util.c index 3f586f2c3d33..3b2835279d37 100644 --- a/tools/testing/selftests/mm/vm_util.c +++ b/tools/testing/selftests/mm/vm_util.c @@ -598,6 +598,47 @@ bool is_range_backed_by_folio_orders(char *start, size= _t len, int order, return true; } =20 +#define TRACEFS_ROOT "/sys/kernel/tracing" + +/* + * Open the enable file of one ftrace event subsystem (e.g. "huge_memory"). + * Returns a descriptor for tracing_events_enable(), or -1 if tracefs or t= he + * subsystem is not there. The events are system-wide state: whoever + * switches them on owns them until it switches them off, including on the + * paths where the test gives up. + */ +int tracing_events_open(const char *subsys) +{ + char path[256]; + + snprintf(path, sizeof(path), TRACEFS_ROOT "/events/%s/enable", + subsys); + return open(path, O_WRONLY); +} + +int tracing_events_enable(int fd, bool enable) +{ + if (pwrite(fd, enable ? "1" : "0", 1, 0) !=3D 1) + return -1; + return 0; +} + +/* Drop what the trace buffer holds so far. */ +int tracing_clear_trace(void) +{ + int fd =3D open(TRACEFS_ROOT "/trace", O_WRONLY | O_TRUNC); + + if (fd < 0) + return -1; + close(fd); + return 0; +} + +FILE *tracing_open_trace(void) +{ + return fopen(TRACEFS_ROOT "/trace", "r"); +} + /* If `ioctls' non-NULL, the allowed ioctls will be returned into the var = */ int uffd_register_with_ioctls(int uffd, void *addr, uint64_t len, bool miss, bool wp, bool minor, uint64_t *ioctls) diff --git a/tools/testing/selftests/mm/vm_util.h b/tools/testing/selftests= /mm/vm_util.h index ce05bce4670d..10c7be46e44c 100644 --- a/tools/testing/selftests/mm/vm_util.h +++ b/tools/testing/selftests/mm/vm_util.h @@ -119,6 +119,10 @@ int close_procmap(struct procmap_fd *procmap); int write_sysfs(const char *file_path, unsigned long val); int read_sysfs(const char *file_path, unsigned long *val); bool softdirty_supported(void); +int tracing_events_open(const char *subsys); +int tracing_events_enable(int fd, bool enable); +int tracing_clear_trace(void); +FILE *tracing_open_trace(void); =20 static inline int open_self_procmap(struct procmap_fd *procmap_out) { --=20 2.54.0 From nobody Tue Sep 29 13:20:38 2026 Received: from fhigh-b6-smtp.messagingengine.com (fhigh-b6-smtp.messagingengine.com [202.12.124.157]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 10DF94746C9; Fri, 7 Aug 2026 11:37:28 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.157 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102654; cv=none; b=Ehg5j79HvvqPT/G/1Cw8YP5KUV8Q9yxZMXo/58YwFc1gZjMCwK9Hcat1IJyoUJ49eNWBtAdKTUMg7J5d5LmC3S6G989onl1ERxUttSZkM3avH5BhgFW0CCH1x+IMpJV0D+B6MrhisIPYcVQ9gzzwj/J1F6y4e6R47NmexjXp1W0= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102654; c=relaxed/simple; bh=TpJa6XMY1dtPBZjOCt895S5qE/WDNrIE1LOpAWQx4qY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=A8lKJkhyV2I2aXjSecYMuyuFo21SQ8997/qlFYKpwNJdpKKQ53ZTCYdyUcPqvQ7Ybc6NBKalhEMywD4RmoEQobSQbH+uHBgiPDfckUKib/zZx41C5EMNyn9ec6f1Z45uDl5juICZiK5gx1QIEA8qvwlIjpmRO8vLyq7mmmIY0iI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=ty6vA8VD; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=PWvzBg0/; arc=none smtp.client-ip=202.12.124.157 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="ty6vA8VD"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="PWvzBg0/" Received: from phl-compute-02.internal (phl-compute-02.internal [10.202.2.42]) by mailfhigh.stl.internal (Postfix) with ESMTP id 0FE8D7A012C; Fri, 7 Aug 2026 07:37:26 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-02.internal (MEProxy); Fri, 07 Aug 2026 07:37:26 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786102645; x= 1786189045; bh=m+sHgjFARkaaiaNC+/FgRc830KGjOObydI5mFYZxies=; b=t y6vA8VDFS6QOc+eJrPUhfldWsDpJS91fBHotWJ4l1dak0o3maLzLBsqcbMhMu58r 0IeiARbhw83pD1iJw6iDTZK730MJj42SrMt0oHz4DCEUC5M3HEqWzpPItVeG5Pzv QfrKfvFc9axSq93ng3r0Pc1n7W0zDYYhsrTHBnI8rkP8pvJkg1SU3tavoyM6pj92 pH1rpC7F/wEGJqO399rC1oTC/FOLfgkoj8VeYtGJEaxpyRB9Hwhu7CM/ajII9Xuq v6ai/cWZu0Bywd7Ktrj8k1MB9TQNLT+w0QVYeVbxIh6HSgTw4bgt/knx6JLGkSsQ AboyVDBmQMhB4Sj8ig+Og== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786102645; x=1786189045; bh=m +sHgjFARkaaiaNC+/FgRc830KGjOObydI5mFYZxies=; b=PWvzBg0/gbOkxq++W 3/5sglMVO5ntG6+poVmC1nFckOxxb0DLezopKvB45upcu3ayex4LtxdAFLJ5x2Ka Qhc+KNDlXYt2iN4eVXj/mMpOcjKmaghrT/jOnS16Uanhdz9/ssSQ6cBsqPhiG2PJ le8MA3u43qknNkNfa6rbGFRsfdxHZOJMB5xik7jmyc+wFBzQS5sWDcu+21mW8w79 OvL6hhK0fhfWhvdUUWut4WkTkOLDbTWxW/8Hr6boE1kRj88BtApWeZLx/zBs3SIv IgODALOISyFV1/ygxEdOW06z4h8yKJSqzR9iPbx60lwK6vE4vKw+NFTsRP20n0t1 6H/dw== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGUERIwmKnJzNqrmH6VlEmVcWeiClBef7dL76l3JFxgCJIE1wbu/wgjcLR687OyMh BQrMCOBv45wH7h3ooBeneK7mWIj7iIytnpxBlgzGmz5SHufCsdgROAqJKtMeokx5eB2AAp cXT5mMO8imvID7jHzWhqIClk07P/42mlBI32M882inm3UlyT/GsY3qxbman5/VlsuJS9+T F+1+aXRxbMzBOR7RkHQ9dyf5K4OhFzw16ipJCuGwyYR0rlnsi17XpZoH3bdrAT++N0O7xU IF8fSnfJTEWFFODGe4CSmBZThmf/6wAYTC6d1nWOkM6O/FhjItWthf5fql0LkcaRcJIWN0 zXgxNDWNr+qBBvcNySFIYAtkqpKlAEpsXaVD9K+94Q4eM/H+ZpGLcPJrSNFXt8s98plKxS 6cAxMWEBa0XIWFm1JjsR9Xck7JxcC+fibTtFAhQ2hETnLwDkHSrpTzhvS6jImINCaqeHlt B9e14FmI4jURMwOo0jXXbxpVinMn1sI1jxA9LeeVZU1d/tTACZPjVlmlxit2LU/qKHtSZD Eb9ED/Lkc3oIlLwAWjEu/WE09T06zxKzNEHiu6aZdl/CsMIUXsZHFbvxgLsnRUCsGml7e1 /WeS/n8UR8xs9oYcqZKwP8DGANPbM5NYe4yc11DR6kkzVg8ZPTiq+BNAAFbw X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 7 Aug 2026 07:37:25 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v2 14/17] selftests/mm: add khugepaged race harness Date: Fri, 7 Aug 2026 12:36:44 +0100 Message-ID: <20260807113647.3744609-15-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260807113647.3744609-1-kirill@shutemov.name> References: <20260807113647.3744609-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Collapse serialises against faults, GUP, fork, mremap and zapping through a protocol of locks, TLB flushes and refcount checks. No khugepaged selftest exercises any of it under contention. Add khugepaged_race. Two faulters, an MADV_DONTNEED thread, a transient FOLL_PIN thread (gup_test), a forker and an mremap thread work the same ranges while one of three drivers collapses: stepped khugepaged, one full pass at a time via khugepaged_full_pass(), so each step covers a known extent; free khugepaged left to run (scan_sleep_millisecs=3D0), for soak; madvise an MADV_COLLAPSE and MADV_DONTNEED loop. Every mode runs in turn unless -m names one, five seconds each. All anon THP orders are enabled as inherit and max_ptes_none is 0, so a window collapses only once fully populated and the racing MADV_DONTNEED steers selection across orders. The rule is that a racing page reads as its pattern or as zero, never anything else. The faulters and fork children check it throughout, and a final sweep checks it again. The kernel's own assertions -- DEBUG_VM, page_table_check, KASAN, lockdep -- are the other half of the oracle, so read dmesg too. The threads share three PMD-sized areas plus one for the mremap thread; -a sets the count, since at a 512M PMD the default is several gigabytes. -d sets the soak length. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) Tested-by: Muhammad Usama Anjum --- tools/testing/selftests/mm/Makefile | 1 + tools/testing/selftests/mm/khugepaged_race.c | 410 +++++++++++++++++++ tools/testing/selftests/mm/run_vmtests.sh | 2 + 3 files changed, 413 insertions(+) create mode 100644 tools/testing/selftests/mm/khugepaged_race.c diff --git a/tools/testing/selftests/mm/Makefile b/tools/testing/selftests/= mm/Makefile index b2d6e5c12934..308bbad73c11 100644 --- a/tools/testing/selftests/mm/Makefile +++ b/tools/testing/selftests/mm/Makefile @@ -106,6 +106,7 @@ TEST_GEN_FILES +=3D rmap TEST_GEN_FILES +=3D folio_split_race_test TEST_GEN_FILES +=3D folio_order_check TEST_GEN_FILES +=3D khugepaged_sync_check +TEST_GEN_FILES +=3D khugepaged_race =20 ifneq ($(ARCH),arm64) TEST_GEN_FILES +=3D soft-dirty diff --git a/tools/testing/selftests/mm/khugepaged_race.c b/tools/testing/s= elftests/mm/khugepaged_race.c new file mode 100644 index 000000000000..3dc651dc56b3 --- /dev/null +++ b/tools/testing/selftests/mm/khugepaged_race.c @@ -0,0 +1,410 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * khugepaged race harness. + * + * Runs collapse against concurrent faults, transient GUP pins + * (gup_test), fork, mremap and MADV_DONTNEED over the same ranges, in + * one of three driver modes: + * + * stepped khugepaged, one full pass at a time through + * khugepaged_full_pass(), so a step covers a known extent; + * free khugepaged left to run (scan_sleep_millisecs=3D0), for soak; + * madvise MADV_COLLAPSE in a loop. + * + * All anon THP orders are enabled (inherit) and max_ptes_none is 0, so a + * window has to be fully populated before khugepaged will collapse it, and + * the racing MADV_DONTNEED decides which orders it can still use. + * + * Correctness signals: every racing page must read as its pattern or + * zero (MADV_DONTNEED), never anything else. The faulters and the fork + * children check that continuously, a final sweep checks it once more, pl= us + * whatever DEBUG_VM / page_table_check / KASAN / lockdep report in + * dmesg, which the caller is expected to inspect. + */ +#define _GNU_SOURCE +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "kselftest.h" +#include "vm_util.h" +#include "hugepage_settings.h" +#include "../../../../mm/gup_test.h" + +#define BASE_ADDR ((void *)(1UL << 30)) + +/* + * Shared playground for faults/pins/fork/dontneed: several PMD-sized + * areas the racing threads spread across, plus one area owned by the + * mremap thread. More areas means more independent regions collapsing + * at once; the default suits a normal machine. On a memory-constrained + * host -- or under emulation, where a 512M PMD (arm64/64K) makes the + * default playground multi-gigabyte -- pass -a to shrink it. + */ +#define DEFAULT_SHARED_AREAS 3 +static int nr_shared_areas; +static int nr_areas; + +static unsigned long hpage_pmd_size; +static unsigned long page_size; +static char *region; /* NR_AREAS * hpage_pmd_size */ +static char *mremap_area; /* region + NR_SHARED_AREAS areas */ +static char *mremap_scratch; /* well above the region */ +static int gup_fd =3D -1; +static volatile int stop; +static volatile int corrupted; + +static unsigned int pattern(unsigned long page_idx) +{ + unsigned int val =3D (unsigned int)page_idx * 2654435761U; + + return val ? val : 1; /* never collides with the zero-fill */ +} + +static void check_page(unsigned long page_idx) +{ + unsigned int val =3D *(unsigned int *)(region + page_idx * page_size); + + if (val && val !=3D pattern(page_idx)) { + corrupted =3D 1; + ksft_print_msg("Corruption at page %lu: %#x !=3D %#x\n", + page_idx, val, pattern(page_idx)); + } +} + +static unsigned long rand_page(unsigned int *seed) +{ + return (unsigned long)rand_r(seed) % + (nr_shared_areas * hpage_pmd_size / page_size); +} + +static void *faulter_fn(void *arg) +{ + unsigned int seed =3D (unsigned long)arg; + + while (!stop) { + unsigned long page_idx =3D rand_page(&seed); + + if (rand_r(&seed) & 1) + *(unsigned int *)(region + page_idx * page_size) =3D + pattern(page_idx); + else + check_page(page_idx); + } + return NULL; +} + +static void *dontneed_fn(void *arg) +{ + unsigned int seed =3D (unsigned long)arg; + + while (!stop) { + unsigned long page_idx =3D rand_page(&seed); + unsigned long nr =3D 1UL << (rand_r(&seed) % 6); /* 1..32 pages */ + + madvise(region + page_idx * page_size, nr * page_size, + MADV_DONTNEED); + usleep(rand_r(&seed) % 500); + } + return NULL; +} + +static void *pinner_fn(void *arg) +{ + unsigned int seed =3D (unsigned long)arg; + + while (!stop) { + struct gup_test gup =3D {}; + unsigned long page_idx =3D rand_page(&seed); + + gup.addr =3D (unsigned long)(region + page_idx * page_size); + gup.size =3D 16 * page_size; + gup.nr_pages_per_call =3D 16; + gup.gup_flags =3D 1; /* FOLL_WRITE */ + /* Racing MADV_DONTNEED makes transient failures expected. */ + ioctl(gup_fd, PIN_FAST_BENCHMARK, &gup); + usleep(rand_r(&seed) % 200); + } + return NULL; +} + +static void *forker_fn(void *arg) +{ + unsigned int seed =3D (unsigned long)arg; + + while (!stop) { + pid_t pid =3D fork(); + + if (pid =3D=3D 0) { + for (int i =3D 0; i < 16; i++) + check_page(rand_page(&seed)); + _exit(corrupted); + } + if (pid > 0) { + int wstatus; + + if (waitpid(pid, &wstatus, 0) < 0) + ksft_exit_fail_perror("waitpid()"); + /* A child killed on the read counts too, not just its exit code. */ + if (!WIFEXITED(wstatus) || WEXITSTATUS(wstatus)) + corrupted =3D 1; + } + usleep(rand_r(&seed) % 2000); + } + return NULL; +} + +static void *mremapper_fn(void *arg) +{ + unsigned int seed =3D (unsigned long)arg; + + while (!stop) { + void *p; + + p =3D mremap(mremap_area, hpage_pmd_size, hpage_pmd_size, + MREMAP_MAYMOVE | MREMAP_FIXED, mremap_scratch); + if (p =3D=3D MAP_FAILED) + ksft_exit_fail_perror("mremap() away"); + for (int i =3D 0; i < 8; i++) + mremap_scratch[(rand_r(&seed) % + (hpage_pmd_size / page_size)) * page_size] =3D 1; + p =3D mremap(mremap_scratch, hpage_pmd_size, hpage_pmd_size, + MREMAP_MAYMOVE | MREMAP_FIXED, mremap_area); + if (p =3D=3D MAP_FAILED) + ksft_exit_fail_perror("mremap() back"); + usleep(rand_r(&seed) % 2000); + } + return NULL; +} + +static unsigned long now_ms(void) +{ + struct timeval tv; + + gettimeofday(&tv, NULL); + return tv.tv_sec * 1000UL + tv.tv_usec / 1000; +} + +static void usage(void) +{ + fprintf(stderr, + "Usage: khugepaged_race [-d seconds] [-m stepped|free|madvise] [-a areas= ]\n" + "\tWithout -m, every mode runs in turn.\n" + "\t-d: seconds per mode (default 5)\n" + "\t-a: number of shared PMD-sized playground areas (default 3)\n"); + exit(1); +} + +int main(int argc, char **argv) +{ + static const char * const thread_names[] =3D { + "faulter", "faulter2", "dontneed", "pinner", "forker", + "mremapper", + }; + void *(*const thread_fns[])(void *) =3D { + faulter_fn, faulter_fn, dontneed_fn, pinner_fn, forker_fn, + mremapper_fn, + }; + const int nr_threads =3D ARRAY_SIZE(thread_names); + pthread_t threads[ARRAY_SIZE(thread_names)]; + static const char * const all_modes[] =3D { "stepped", "free", "madvise" = }; + const char *one_mode[1]; + const char * const *modes =3D all_modes; + int nr_modes =3D ARRAY_SIZE(all_modes); + const char *mode_arg =3D NULL; + struct thp_settings settings; + unsigned long end_ms; + int duration_s =3D 5; + unsigned long thread_mask =3D ~0UL; + int nr_areas_arg =3D 0; + unsigned long i; + int steps =3D 0; + int opt; + + while ((opt =3D getopt(argc, argv, "a:d:m:t:h")) !=3D -1) { + switch (opt) { + case 'a': + nr_areas_arg =3D atoi(optarg); + break; + case 'd': + duration_s =3D atoi(optarg); + break; + case 'm': + mode_arg =3D optarg; + break; + case 't': + /* debug: bitmask of racing threads to start */ + thread_mask =3D strtoul(optarg, NULL, 0); + break; + default: + usage(); + } + } + if (mode_arg) { + if (strcmp(mode_arg, "stepped") && strcmp(mode_arg, "free") && + strcmp(mode_arg, "madvise")) + usage(); + one_mode[0] =3D mode_arg; + modes =3D one_mode; + nr_modes =3D 1; + } + + ksft_print_header(); + if (!thp_available()) + ksft_exit_skip("Transparent Hugepages not available\n"); + + page_size =3D getpagesize(); + hpage_pmd_size =3D read_pmd_pagesize(); + if (!hpage_pmd_size) + ksft_exit_fail_msg("Reading PMD pagesize failed\n"); + + gup_fd =3D open("/sys/kernel/debug/gup_test", O_RDWR); + if (gup_fd < 0) + ksft_exit_skip("/sys/kernel/debug/gup_test requires CONFIG_GUP_TEST and = root\n"); + + nr_shared_areas =3D nr_areas_arg > 0 ? nr_areas_arg : DEFAULT_SHARED_AREA= S; + nr_areas =3D nr_shared_areas + 1; + + /* + * The mremap thread moves its area to this address and back, and + * MREMAP_FIXED unmaps whatever is in the way without saying so. Claim + * the address here, so a layout that does not match this assumption + * fails now instead of losing a mapping later. Nothing else in the + * process maps this low: thread stacks and malloc arenas come from the + * top-down mmap area, well above. + */ + mremap_scratch =3D (char *)BASE_ADDR + 2 * nr_areas * hpage_pmd_size; + if (mmap(mremap_scratch, hpage_pmd_size, PROT_NONE, + MAP_ANONYMOUS | MAP_PRIVATE | MAP_FIXED_NOREPLACE, + -1, 0) !=3D (void *)mremap_scratch) + ksft_exit_fail_perror("mmap() mremap scratch"); + + ksft_set_plan(nr_modes); + + thp_save_settings(); + thp_read_settings(&settings); + + /* + * A base entry for the stack, so that the pop at the end of a mode + * always has something to write back: thp_pop_settings() on an empty + * stack has no settings to apply and gives up. + */ + thp_push_settings(&settings); + + for (int m =3D 0; m < nr_modes; m++) { + const char *mode =3D modes[m]; + + thp_read_settings(&settings); + settings.thp_enabled =3D THP_MADVISE; + settings.thp_defrag =3D THP_DEFRAG_ALWAYS; + settings.shmem_enabled =3D SHMEM_NEVER; + settings.khugepaged.defrag =3D 1; + settings.khugepaged.scan_sleep_millisecs =3D + strcmp(mode, "free") ? 1000 : 0; + settings.khugepaged.alloc_sleep_millisecs =3D 10; + /* + * Strict occupancy: mTHP collapse only supports 0 or + * HPAGE_PMD_NR - 1 and coerces anything else to 0 anyway, and 0 + * also keeps khugepaged from burning the whole step in doomed + * PMD-sized allocations on 512M-PMD configs: under racing + * MADV_DONTNEED a fully populated PMD area is rare. + */ + settings.khugepaged.max_ptes_none =3D 0; + settings.khugepaged.pages_to_scan =3D + nr_areas * (hpage_pmd_size / page_size) * 8; + for (i =3D 0; i < NR_ORDERS; i++) { + if (thp_supported_orders() & (1UL << i)) + settings.hugepages[i].enabled =3D THP_INHERIT; + } + /* Popped at the end of this mode, before the next one. */ + thp_push_settings(&settings); + + region =3D mmap(BASE_ADDR, nr_areas * hpage_pmd_size, + PROT_READ | PROT_WRITE, MAP_ANONYMOUS | + MAP_PRIVATE | MAP_FIXED_NOREPLACE, -1, 0); + if (region !=3D BASE_ADDR) + ksft_exit_fail_perror("mmap() playground"); + mremap_area =3D region + nr_shared_areas * hpage_pmd_size; + + /* Populate so the first pass has something to collapse. */ + for (i =3D 0; i < nr_shared_areas * hpage_pmd_size / page_size; i++) + *(unsigned int *)(region + i * page_size) =3D pattern(i); + memset(mremap_area, 1, hpage_pmd_size); + madvise(region, nr_areas * hpage_pmd_size, MADV_HUGEPAGE); + + for (i =3D 0; i < nr_threads; i++) { + if (!(thread_mask & (1UL << i))) { + threads[i] =3D 0; + continue; + } + if (pthread_create(&threads[i], NULL, thread_fns[i], + (void *)(i + 1))) + ksft_exit_fail_perror("pthread_create()"); + } + + end_ms =3D now_ms() + duration_s * 1000UL; + if (!strcmp(mode, "stepped")) { + while (now_ms() < end_ms && !corrupted) { + if (!khugepaged_full_pass(600)) + ksft_exit_fail_msg("khugepaged pass timed out\n"); + steps++; + } + } else if (!strcmp(mode, "free")) { + while (now_ms() < end_ms && !corrupted) + usleep(100 * 1000); + } else { /* madvise */ + while (now_ms() < end_ms && !corrupted) { + for (i =3D 0; i < nr_shared_areas; i++) { + madvise(region + i * hpage_pmd_size, + hpage_pmd_size, MADV_COLLAPSE); + } + madvise(region, nr_shared_areas * hpage_pmd_size, + MADV_DONTNEED); + steps++; + } + } + + stop =3D 1; + for (i =3D 0; i < nr_threads; i++) { + if (threads[i]) + pthread_join(threads[i], NULL); + } + + /* Final integrity sweep. */ + for (i =3D 0; i < nr_shared_areas * hpage_pmd_size / page_size; i++) + check_page(i); + + ksft_test_result(!corrupted, + "%s: %ds, %d steps, no corruption\n", + mode, duration_s, steps); + + /* + * Hand the address space and the settings back before the + * next mode: it maps the region at the same fixed address, + * and its scan cadence differs. + */ + munmap(region, nr_areas * hpage_pmd_size); + thp_pop_settings(); + stop =3D 0; + steps =3D 0; + + if (corrupted) { + /* Memory is suspect; the rest would prove nothing. */ + while (++m < nr_modes) + ksft_test_result_skip("%s: skipped after corruption\n", + modes[m]); + break; + } + } + + thp_restore_settings(); + ksft_finished(); +} diff --git a/tools/testing/selftests/mm/run_vmtests.sh b/tools/testing/self= tests/mm/run_vmtests.sh index c0f69da3fd3b..fc61907aa3b2 100755 --- a/tools/testing/selftests/mm/run_vmtests.sh +++ b/tools/testing/selftests/mm/run_vmtests.sh @@ -406,6 +406,8 @@ CATEGORY=3D"thp" run_test ./folio_order_check =20 CATEGORY=3D"thp" run_test ./khugepaged_sync_check =20 +CATEGORY=3D"thp" run_test ./khugepaged_race + CATEGORY=3D"thp" run_test ./khugepaged =20 CATEGORY=3D"thp" run_test ./khugepaged -s 2 --=20 2.54.0 From nobody Tue Sep 29 13:20:38 2026 Received: from fout-b5-smtp.messagingengine.com (fout-b5-smtp.messagingengine.com [202.12.124.148]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0557847426D; Fri, 7 Aug 2026 11:37:31 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.148 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102656; cv=none; b=DSQs2rwQgWGOrt0MMVF/KY9ZzF2xg11J1+bO389kIX5/UmWJSxzrspdZ7xvmCjDgIIfoVx5ytdCDInV23jCecDPnDFi7oORr/UwfMG34CP+myhRaxsq+JkKVF4SnzrFLQHq+NvT8HmH6rHRtg/lftx6xgmUjtGQFwGxkuWR5Dm4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102656; c=relaxed/simple; bh=DtaUsi/AmaF/PfbZkl7UuFW6bWfuvRjmbzsb/2kyjmM=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=QpSt7jTsjyGhq9PVGeBNrX6D0r4lUXcxIX71w+g/ry+NrZdx1eS5UkFPLugnwnGLFvXRWkFvQAQrO+7U6K0ftLmLtxH5ZvIQatOeyJf/J5nsNisxegdsYCjU/jtBUw+BGTfd2RhrcYnL+9YwzMVWQM4m8kauadH1t53F+CoF7jE= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=Q9hN0SdJ; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=g8Yh/WF6; arc=none smtp.client-ip=202.12.124.148 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="Q9hN0SdJ"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="g8Yh/WF6" Received: from phl-compute-01.internal (phl-compute-01.internal [10.202.2.41]) by mailfout.stl.internal (Postfix) with ESMTP id 78CC51D0008B; Fri, 7 Aug 2026 07:37:28 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-01.internal (MEProxy); Fri, 07 Aug 2026 07:37:29 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786102648; x= 1786189048; bh=pirjW+/uVLVtcgVpdoTceZVcLT7vNoA014N2uc8FC80=; b=Q 9hN0SdJ9Y9uH3hM+MxqgXo0HDBkt/N9uqhuZ6VZBY6lFDAeKwIzjVJ3ZhxC4ugbs m27ai2iux7fEFmVLViKGIras0QSTknlk0lid6Nziv1PTjnS9+xZDLyPBeuZICTKB /kO5DnhliidD2z1hVOTa0+QSfm21dQRkaQS4cm/fZwqIQrJJHZKGdq6vfyakSzaH wVQDLarjIpz9JdDM/WihTv5gUgbEGTjh2RN6OWl9uXCHKyk/zca9tO5xlXqopF3R UDyZmc8J5GBjdj/bTH0raHkmND+neoUS1R8gyTjuHJfWWuTLJ7+iCYjwTurVOAyj 5ZpDPM/R0FAknXCnjvwaA== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786102648; x=1786189048; bh=p irjW+/uVLVtcgVpdoTceZVcLT7vNoA014N2uc8FC80=; b=g8Yh/WF6MX99h9Wlw o85ZrnLBr8qmIxdASmUlXqOHCW2hkCk/qc/i/QzOfmQTk7nzMfl6SwM0+4jown6w slHt1s56OegoN5Ph4WVono6Vntqi4i+jpM2xPhJpQ5g2moOkCtda1Xa8bko2FOS9 h+Y5MrG0Uk0RsDDjL46a/MSVbvzLgHhvUGFOB97bTSTGDlPcsxoA1b2cmalMbhY+ SK+ooA4+SGKis82L/3POJahPeJ5d7Gjg2vMV2fUmF5QZVxAnYwssQjT1hquzxGq4 hmIA1JVjRuEVIJIdrwdFVG2OLr+w6Lc353FNiL8A5+Zxet7RHpBO/kuwrzOsjKVL sG4pA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEcBCtHwgFqhRPkncf0k+1ylINMWkiJotqLKgP2XUmyCnaVLH6jJFUJYUtwVpCetV bzy51nCrHgVQnKUbL6xUc35a2wACIZl+NbR3y/w0ModeWlKP4SKCXfcn2Mi8esZBMBdaWx faB4jzBJOawj2LBnFTgyCBiywrPp7FpNgfXB5cgYmB0wgvQ/HhhWAfKjMf/weJgLhMF/VV yXrMoja9EZiHjuXu8bb94hiDoU2fqI8TuwJJmzzN4TUL6MxSST+GnkpA2ZnKWVewXCphJB Lec89iAB+bDvFyV3zsxjbJ8cA1IEQ3aF+gxi5HDHtN75oTRJtbf/Sd54tDaiXy5fa0cQX5 hoGETtTIfFLGj1cGbXs4/N0rNKNVrZgeu1OU1QhAOZ7bPZc3Aw4JqqhOgEIexTauJZxkHU TbWF4BgZy0M4TbmGMdjoesA5287sX/xP4xHusiyuAp+GBpqhHmkmTvjKuXPRjKfVOqi0I1 RL9MiNNUq75n+PI2g1y2ZZc7rmd3xeZEl1E6coDyGfmfuDk/5ftusYa+ukpqGLMxSVfYAS ijiJYmhr3aMVUvE1AQl64BB8i5VpuNLTyO9VuJ1p/8t2Lnq1L+CNDOKErEKx3n3kTvpvfI 3DwLmLDIO16LxOm4k3YjuSFLS0SjcFaAcQbmVLF7pxpDdfjB3r545/i6foFw X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 7 Aug 2026 07:37:27 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v2 15/17] selftests/mm: race collapse of windows with holes Date: Fri, 7 Aug 2026 12:36:45 +0100 Message-ID: <20260807113647.3744609-16-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260807113647.3744609-1-kirill@shutemov.name> References: <20260807113647.3744609-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The harness pins max_ptes_none to 0, so khugepaged only collapses a window once every PTE in it is present. Collapsing a window that has holes never happens, and that is a different path: a hole is not copied from anywhere but zero-filled into the new folio, and the slot is re-checked under the page table lock at install time in case a racing fault filled it first. Run both ends of the occupancy scale, one after the other, for every driver mode -- mTHP collapse supports only those two, 0 and HPAGE_PMD_NR - 1, and coerces anything between them to 0. Each result says which end it ran: ok 1 stepped/strict: 5s, 231 steps, no corruption ok 2 stepped/holes: 5s, 194 steps, no corruption -z narrows a run to the hole-heavy end, the way -m narrows it to one mode. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) Tested-by: Muhammad Usama Anjum --- tools/testing/selftests/mm/khugepaged_race.c | 57 +++++++++++++------- 1 file changed, 39 insertions(+), 18 deletions(-) diff --git a/tools/testing/selftests/mm/khugepaged_race.c b/tools/testing/s= elftests/mm/khugepaged_race.c index 3dc651dc56b3..a2a91906285b 100644 --- a/tools/testing/selftests/mm/khugepaged_race.c +++ b/tools/testing/selftests/mm/khugepaged_race.c @@ -11,9 +11,12 @@ * free khugepaged left to run (scan_sleep_millisecs=3D0), for soak; * madvise MADV_COLLAPSE in a loop. * - * All anon THP orders are enabled (inherit) and max_ptes_none is 0, so a - * window has to be fully populated before khugepaged will collapse it, and - * the racing MADV_DONTNEED decides which orders it can still use. + * All anon THP orders are enabled (inherit). Occupancy runs at both ends + * of what mTHP collapse supports: max_ptes_none 0, where a window must be + * fully populated, and HPAGE_PMD_NR - 1, where a window full of holes + * collapses too. The holes are not copied from anywhere -- they are + * zero-filled, and re-checked under the page table lock at install time in + * case a racing fault got there first. * * Correctness signals: every racing page must read as its pattern or * zero (MADV_DONTNEED), never anything else. The faulters and the fork @@ -196,9 +199,11 @@ static unsigned long now_ms(void) static void usage(void) { fprintf(stderr, - "Usage: khugepaged_race [-d seconds] [-m stepped|free|madvise] [-a areas= ]\n" + "Usage: khugepaged_race [-d seconds] [-m stepped|free|madvise] [-z] [-a = areas]\n" "\tWithout -m, every mode runs in turn.\n" "\t-d: seconds per mode (default 5)\n" + "\tBoth occupancy limits run unless -z asks for holes only.\n" + "\t-z: only max_ptes_none =3D HPAGE_PMD_NR - 1 (hole-heavy)\n" "\t-a: number of shared PMD-sized playground areas (default 3)\n"); exit(1); } @@ -216,6 +221,9 @@ int main(int argc, char **argv) const int nr_threads =3D ARRAY_SIZE(thread_names); pthread_t threads[ARRAY_SIZE(thread_names)]; static const char * const all_modes[] =3D { "stepped", "free", "madvise" = }; + static const int all_nones[] =3D { 0, 1 }; /* strict, holes */ + const int *nones =3D all_nones; + int nr_nones =3D ARRAY_SIZE(all_nones); const char *one_mode[1]; const char * const *modes =3D all_modes; int nr_modes =3D ARRAY_SIZE(all_modes); @@ -225,11 +233,12 @@ int main(int argc, char **argv) int duration_s =3D 5; unsigned long thread_mask =3D ~0UL; int nr_areas_arg =3D 0; + bool holes_only =3D false; unsigned long i; int steps =3D 0; int opt; =20 - while ((opt =3D getopt(argc, argv, "a:d:m:t:h")) !=3D -1) { + while ((opt =3D getopt(argc, argv, "a:d:m:t:zh")) !=3D -1) { switch (opt) { case 'a': nr_areas_arg =3D atoi(optarg); @@ -244,10 +253,18 @@ int main(int argc, char **argv) /* debug: bitmask of racing threads to start */ thread_mask =3D strtoul(optarg, NULL, 0); break; + case 'z': + holes_only =3D true; + break; default: usage(); } } + if (holes_only) { + nones =3D all_nones + 1; + nr_nones =3D 1; + } + if (mode_arg) { if (strcmp(mode_arg, "stepped") && strcmp(mode_arg, "free") && strcmp(mode_arg, "madvise")) @@ -287,7 +304,7 @@ int main(int argc, char **argv) -1, 0) !=3D (void *)mremap_scratch) ksft_exit_fail_perror("mmap() mremap scratch"); =20 - ksft_set_plan(nr_modes); + ksft_set_plan(nr_modes * nr_nones); =20 thp_save_settings(); thp_read_settings(&settings); @@ -299,8 +316,9 @@ int main(int argc, char **argv) */ thp_push_settings(&settings); =20 - for (int m =3D 0; m < nr_modes; m++) { - const char *mode =3D modes[m]; + for (int mn =3D 0; mn < nr_modes * nr_nones; mn++) { + const char *mode =3D modes[mn / nr_nones]; + bool holes =3D nones[mn % nr_nones]; =20 thp_read_settings(&settings); settings.thp_enabled =3D THP_MADVISE; @@ -310,14 +328,16 @@ int main(int argc, char **argv) settings.khugepaged.scan_sleep_millisecs =3D strcmp(mode, "free") ? 1000 : 0; settings.khugepaged.alloc_sleep_millisecs =3D 10; + /* - * Strict occupancy: mTHP collapse only supports 0 or - * HPAGE_PMD_NR - 1 and coerces anything else to 0 anyway, and 0 - * also keeps khugepaged from burning the whole step in doomed - * PMD-sized allocations on 512M-PMD configs: under racing - * MADV_DONTNEED a fully populated PMD area is rare. + * mTHP collapse only supports the two ends of the occupancy + * scale: 0 or HPAGE_PMD_NR - 1 (anything else coerces to 0). + * Strict needs a fully populated window, which is rare under + * racing MADV_DONTNEED; hole-heavy windows collapse instead, + * so the two ends race different paths. */ - settings.khugepaged.max_ptes_none =3D 0; + settings.khugepaged.max_ptes_none =3D holes ? + (hpage_pmd_size / page_size) - 1 : 0; settings.khugepaged.pages_to_scan =3D nr_areas * (hpage_pmd_size / page_size) * 8; for (i =3D 0; i < NR_ORDERS; i++) { @@ -383,8 +403,9 @@ int main(int argc, char **argv) check_page(i); =20 ksft_test_result(!corrupted, - "%s: %ds, %d steps, no corruption\n", - mode, duration_s, steps); + "%s/%s: %ds, %d steps, no corruption\n", + mode, holes ? "holes" : "strict", + duration_s, steps); =20 /* * Hand the address space and the settings back before the @@ -398,9 +419,9 @@ int main(int argc, char **argv) =20 if (corrupted) { /* Memory is suspect; the rest would prove nothing. */ - while (++m < nr_modes) + while (++mn < nr_modes * nr_nones) ksft_test_result_skip("%s: skipped after corruption\n", - modes[m]); + modes[mn / nr_nones]); break; } } --=20 2.54.0 From nobody Tue Sep 29 13:20:38 2026 Received: from fout-b5-smtp.messagingengine.com (fout-b5-smtp.messagingengine.com [202.12.124.148]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B7FF53FC5AE; Fri, 7 Aug 2026 11:37:34 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.148 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102659; cv=none; b=nxS+dFlrRAtMvhPYGruT0HGZ4xkUqFUotZTaJlMMyXuWaapvptvLXyKwfPj9QRiaSPnP00eL7sq/ckqu6mnhuEPXkMPQ+Z7ODvqaUrYT45CApdJ6HBMSc9+89+GdzzMSLidhyJ3SNThQbCwcjfauOMJ/o92rlP15PtQNLfQ9uE4= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102659; c=relaxed/simple; bh=F/TF2/zns+QnkbIDlaS4JtzV6yU5B+IfMUtnKwD6VCI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=GfVgbO0uDITXGnV7cb165DTaX//9wnlggVDnHFkU1Lg/w105e/GwxJNqHPujfUxPr2M4vaireElEDD/dOlQDqeElq9cRhTXQYBKt1nb+vJrwx1VFD18KgB87PgXiSQyAd7DabW2YgoxBVz5weyu7TxQpXKyCa0focAv5rLxH7jo= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=WxjzI+jD; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=DFMFvzsp; arc=none smtp.client-ip=202.12.124.148 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="WxjzI+jD"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="DFMFvzsp" Received: from phl-compute-05.internal (phl-compute-05.internal [10.202.2.45]) by mailfout.stl.internal (Postfix) with ESMTP id 38B421D000DA; Fri, 7 Aug 2026 07:37:31 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-05.internal (MEProxy); Fri, 07 Aug 2026 07:37:31 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786102651; x= 1786189051; bh=OqABGYaTDBjD+cwWJAm5MV885+wK1zKh6mDVYP63LKI=; b=W xjzI+jD+AKuhzGdFBCgf8zaJLRFQ5VC18xsWsxekiPDCXdCcpZo9GbG/u3x0bJrL WKo+Wp6CbYCxRIYss4kRANOy2S0x5Mvj7hwkVNRmaA1HVVFVEp6jKV8dHHNFBIsO 82ez0QTonL9r+IrQ+mOOaOPB8Aj6Untdu229hlkmSITokttshsd93JQeC8bh1nbO SAxAcx/+PyIgUXgWbri6lEM2aCS/8Kc4ZErqdPPKJNM4afpclsk5vwRiNrUFJcWd S3oY6o8FQXkpBqtioc8XP5f8+owugenmUqip1uAX6TReOVNg29wR1xrqTQP7cxRW khvk6WBnQeMxesduSaiUw== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786102651; x=1786189051; bh=O qABGYaTDBjD+cwWJAm5MV885+wK1zKh6mDVYP63LKI=; b=DFMFvzsp6InJdi/IH B4xHaxSxATDcX9SMie17GDw9dOmjs0Bde1NdPYO/yrfHlWR2njTsFBSW36AtNRCD k1b/uLcOJ+jrZRaZAZVh3dLWS7pKIXJKLGICWsN4ylfLCZzIntZX90B5ic+nLVdC o8K9btkxM0uK9qn6a1EBgClO3K+ntTy/IGuvE4DF7n0KhRSQ5PkLF9ao/wzXAROC a/XPy7M+lf8HGFO27LonZayAxkdIali8U4KZFkHIs40/lAQ1gHbzYgIDxABdJRHl F/PAf9r9iY7eSAQUs1sJ2br26/KQKjWgplCQ5Q97A4CAYs6K9zr2ukV/sJfGdWtJ mmBHQ== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEcBCtHwgFqhRPkncf0k+1ylINMWkiJotqLKgP2XUmyCnaVLH6jJFUJYUtwVpCetV bzy51nCrHgVQnKUbL6xUc35a2wACIZl+NbR3y/w0ModeWlKP4SKCXfcn2Mi8esZBMBdaWx faB4jzBJOawj2LBnFTgyCBiywrPp7FpNgfXB5cgYmB0wgvQ/HhhWAfKjMf/weJgLhMF/VV yXrMoja9EZiHjuXu8bb94hiDoU2fqI8TuwJJmzzN4TUL6MxSST+GnkpA2ZnKWVewXCphJB Lec89iAB+bDvFyV3zsxjbJ8cA1IEQ3aF+gxi5HDHtN75oTRJtbf/Sd54tDaiXy5fa0cQol MW02m0ZGtjHD58JE8hq66IpVl0U+UPJ3nIC2pthVrT/8lSmS/xrVKgAZ9KysURaeJoEhw5 WpFpAmC8vf159gkJAGhW8a/tUyRz26I4LvuoRP+Xa0VlHgTOMmYlw0igpjphktySm6WZ13 46PBg7JdeKJ/GB9sE5rFCoFlTqgY+IR6H61d+UuuYKrpZ00RVuMZsiGazJ3uH1j1nzi7u7 5rP73EdU+fkpAuY1FQ1fYfRVDALBGosYQnjzbDtPx9MVdzp/o2L7ailpK8+VfvossfiVSM c+3BZp/vLcPZJs/IZbkaQEAhNUonJ23UQVWpdakD9zo2ikaCXh83KADXdxeg X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 7 Aug 2026 07:37:30 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v2 16/17] selftests/mm: add memory-pressure threads to the khugepaged race harness Date: Fri, 7 Aug 2026 12:36:46 +0100 Message-ID: <20260807113647.3744609-17-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260807113647.3744609-1-kirill@shutemov.name> References: <20260807113647.3744609-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The harness races collapse against faults, pins, fork, mremap and MADV_DONTNEED, but nothing in it elevates a source folio's refcount from the reclaim or compaction side. Add two more threads, and run every mode and occupancy limit both with and without them: - pageout: cycles MADV_PAGEOUT over a dedicated neighbour region, faults it back in and checks the content each round, since a page's pattern must survive the trip through swap. Idle when the host has no swap, because then there is no anon reclaim to drive. - compactor: writes /proc/sys/vm/compact_memory in a loop. Compaction isolates and migrates folios, so it competes with a collapse for the pages it is gathering, with refcount elevations and migration entries of its own. -p narrows a run to the combinations that have them, the way -z narrows the occupancy and -m the mode. Each result says which it ran: ok 2 stepped/strict/pressure: 5s, 88 steps, no corruption Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) Tested-by: Muhammad Usama Anjum --- tools/testing/selftests/mm/khugepaged_race.c | 147 +++++++++++++++++-- 1 file changed, 136 insertions(+), 11 deletions(-) diff --git a/tools/testing/selftests/mm/khugepaged_race.c b/tools/testing/s= elftests/mm/khugepaged_race.c index a2a91906285b..abd142db7cbe 100644 --- a/tools/testing/selftests/mm/khugepaged_race.c +++ b/tools/testing/selftests/mm/khugepaged_race.c @@ -18,6 +18,12 @@ * zero-filled, and re-checked under the page table lock at install time in * case a racing fault got there first. * + * -p adds memory pressure to any of the above: MADV_PAGEOUT cycling + * on a dedicated neighbor region (swap traffic and LRU churn; skipped + * with a note when the host has no swap) and a compact_memory trigger + * loop (compaction migrates source folios, racing collapse's freeze + * with refcount elevation and migration entries of its own). + * * Correctness signals: every racing page must read as its pattern or * zero (MADV_DONTNEED), never anything else. The faulters and the fork * children check that continuously, a final sweep checks it once more, pl= us @@ -61,6 +67,8 @@ static unsigned long page_size; static char *region; /* NR_AREAS * hpage_pmd_size */ static char *mremap_area; /* region + NR_SHARED_AREAS areas */ static char *mremap_scratch; /* well above the region */ +static char *pageout_area; /* -p: dedicated pressure region */ +static size_t pageout_size; static int gup_fd =3D -1; static volatile int stop; static volatile int corrupted; @@ -188,6 +196,70 @@ static void *mremapper_fn(void *arg) return NULL; } =20 +/* + * -p: swap traffic and LRU churn on a region of our own. The content + * check is exact: a page out and back through swap must preserve the + * pattern, and nothing else ever writes here. + */ +static void *pageout_fn(void *arg) +{ + unsigned int seed =3D (unsigned long)arg; + unsigned long nr =3D pageout_size / page_size; + unsigned long i; + + for (i =3D 0; i < nr; i++) + *(unsigned int *)(pageout_area + i * page_size) =3D pattern(i); + + while (!stop) { + madvise(pageout_area, pageout_size, MADV_PAGEOUT); + for (i =3D 0; i < nr && !stop; i++) { + unsigned int val =3D *(unsigned int *)(pageout_area + + i * page_size); + + if (val !=3D pattern(i)) { + corrupted =3D 1; + ksft_print_msg("Pageout corruption at page %lu: %#x !=3D %#x\n", + i, val, pattern(i)); + } + } + usleep(rand_r(&seed) % 2000); + } + return NULL; +} + +/* -p: compaction migrates the collapse sources out from under us. */ +static void *compactor_fn(void *arg) +{ + unsigned int seed =3D (unsigned long)arg; + int fd =3D open("/proc/sys/vm/compact_memory", O_WRONLY); + + if (fd < 0) { + ksft_print_msg("No compact_memory; compactor idle\n"); + return NULL; + } + while (!stop) { + if (write(fd, "1", 1) < 0) + break; + usleep(10000 + rand_r(&seed) % 100000); + } + close(fd); + return NULL; +} + +static bool swap_available(void) +{ + char line[256]; + int lines =3D 0; + FILE *fp =3D fopen("/proc/swaps", "r"); + + if (!fp) + return false; + while (fgets(line, sizeof(line), fp)) + lines++; + fclose(fp); + return lines > 1; +} + static unsigned long now_ms(void) { struct timeval tv; @@ -199,11 +271,14 @@ static unsigned long now_ms(void) static void usage(void) { fprintf(stderr, - "Usage: khugepaged_race [-d seconds] [-m stepped|free|madvise] [-z] [-a = areas]\n" + "Usage: khugepaged_race [-d seconds] [-m stepped|free|madvise] [-z] [-p]= [-a areas]\n" "\tWithout -m, every mode runs in turn.\n" "\t-d: seconds per mode (default 5)\n" "\tBoth occupancy limits run unless -z asks for holes only.\n" "\t-z: only max_ptes_none =3D HPAGE_PMD_NR - 1 (hole-heavy)\n" + "\tRuns with and without memory pressure unless -p asks for\n" + "\tpressure only.\n" + "\t-p: only with the pageout and compaction threads\n" "\t-a: number of shared PMD-sized playground areas (default 3)\n"); exit(1); } @@ -212,18 +287,22 @@ int main(int argc, char **argv) { static const char * const thread_names[] =3D { "faulter", "faulter2", "dontneed", "pinner", "forker", - "mremapper", + "mremapper", "pageout", "compactor", }; void *(*const thread_fns[])(void *) =3D { faulter_fn, faulter_fn, dontneed_fn, pinner_fn, forker_fn, - mremapper_fn, + mremapper_fn, pageout_fn, compactor_fn, }; + const unsigned long pageout_bit =3D 1UL << 6, compactor_bit =3D 1UL << 7; const int nr_threads =3D ARRAY_SIZE(thread_names); pthread_t threads[ARRAY_SIZE(thread_names)]; static const char * const all_modes[] =3D { "stepped", "free", "madvise" = }; static const int all_nones[] =3D { 0, 1 }; /* strict, holes */ + static const int all_press[] =3D { 0, 1 }; /* quiet, under pressure */ const int *nones =3D all_nones; + const int *press =3D all_press; int nr_nones =3D ARRAY_SIZE(all_nones); + int nr_press =3D ARRAY_SIZE(all_press); const char *one_mode[1]; const char * const *modes =3D all_modes; int nr_modes =3D ARRAY_SIZE(all_modes); @@ -232,13 +311,15 @@ int main(int argc, char **argv) unsigned long end_ms; int duration_s =3D 5; unsigned long thread_mask =3D ~0UL; + unsigned long base_mask; int nr_areas_arg =3D 0; bool holes_only =3D false; + bool pressure_only =3D false; unsigned long i; int steps =3D 0; int opt; =20 - while ((opt =3D getopt(argc, argv, "a:d:m:t:zh")) !=3D -1) { + while ((opt =3D getopt(argc, argv, "a:d:m:t:zph")) !=3D -1) { switch (opt) { case 'a': nr_areas_arg =3D atoi(optarg); @@ -256,6 +337,9 @@ int main(int argc, char **argv) case 'z': holes_only =3D true; break; + case 'p': + pressure_only =3D true; + break; default: usage(); } @@ -265,6 +349,11 @@ int main(int argc, char **argv) nr_nones =3D 1; } =20 + if (pressure_only) { + press =3D all_press + 1; + nr_press =3D 1; + } + if (mode_arg) { if (strcmp(mode_arg, "stepped") && strcmp(mode_arg, "free") && strcmp(mode_arg, "madvise")) @@ -304,7 +393,12 @@ int main(int argc, char **argv) -1, 0) !=3D (void *)mremap_scratch) ksft_exit_fail_perror("mmap() mremap scratch"); =20 - ksft_set_plan(nr_modes * nr_nones); + base_mask =3D thread_mask; + if (!swap_available()) + /* No swap, no anon reclaim: compaction-only pressure. */ + ksft_print_msg("no swap: the pageout thread stays idle\n"); + + ksft_set_plan(nr_modes * nr_nones * nr_press); =20 thp_save_settings(); thp_read_settings(&settings); @@ -316,9 +410,17 @@ int main(int argc, char **argv) */ thp_push_settings(&settings); =20 - for (int mn =3D 0; mn < nr_modes * nr_nones; mn++) { - const char *mode =3D modes[mn / nr_nones]; - bool holes =3D nones[mn % nr_nones]; + for (int run =3D 0; run < nr_modes * nr_nones * nr_press; run++) { + int rem =3D run % (nr_nones * nr_press); + const char *mode =3D modes[run / (nr_nones * nr_press)]; + bool holes =3D nones[rem / nr_press]; + bool pressure =3D press[rem % nr_press]; + + thread_mask =3D base_mask; + if (!pressure) + thread_mask &=3D ~(pageout_bit | compactor_bit); + else if (!swap_available()) + thread_mask &=3D ~pageout_bit; =20 thp_read_settings(&settings); settings.thp_enabled =3D THP_MADVISE; @@ -354,6 +456,24 @@ int main(int argc, char **argv) ksft_exit_fail_perror("mmap() playground"); mremap_area =3D region + nr_shared_areas * hpage_pmd_size; =20 + if (thread_mask & pageout_bit) { + /* + * Big enough to cycle real reclaim, small enough not + * to dominate a TCG guest: 4 PMD areas, clamped to + * [16M, 64M]. + */ + pageout_size =3D 4 * hpage_pmd_size; + pageout_size =3D pageout_size < (16UL << 20) ? + (16UL << 20) : + pageout_size > (64UL << 20) ? + (64UL << 20) : pageout_size; + pageout_area =3D mmap(NULL, pageout_size, + PROT_READ | PROT_WRITE, + MAP_ANONYMOUS | MAP_PRIVATE, -1, 0); + if (pageout_area =3D=3D MAP_FAILED) + ksft_exit_fail_perror("mmap() pageout area"); + } + /* Populate so the first pass has something to collapse. */ for (i =3D 0; i < nr_shared_areas * hpage_pmd_size / page_size; i++) *(unsigned int *)(region + i * page_size) =3D pattern(i); @@ -403,8 +523,9 @@ int main(int argc, char **argv) check_page(i); =20 ksft_test_result(!corrupted, - "%s/%s: %ds, %d steps, no corruption\n", + "%s/%s%s: %ds, %d steps, no corruption\n", mode, holes ? "holes" : "strict", + pressure ? "/pressure" : "", duration_s, steps); =20 /* @@ -413,15 +534,19 @@ int main(int argc, char **argv) * and its scan cadence differs. */ munmap(region, nr_areas * hpage_pmd_size); + if (pageout_area) { + munmap(pageout_area, pageout_size); + pageout_area =3D NULL; + } thp_pop_settings(); stop =3D 0; steps =3D 0; =20 if (corrupted) { /* Memory is suspect; the rest would prove nothing. */ - while (++mn < nr_modes * nr_nones) + while (++run < nr_modes * nr_nones * nr_press) ksft_test_result_skip("%s: skipped after corruption\n", - modes[mn / nr_nones]); + modes[run / (nr_nones * nr_press)]); break; } } --=20 2.54.0 From nobody Tue Sep 29 13:20:38 2026 Received: from fhigh-b6-smtp.messagingengine.com (fhigh-b6-smtp.messagingengine.com [202.12.124.157]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 35B8937C92B; Fri, 7 Aug 2026 11:37:36 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.157 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102658; cv=none; b=V0HnTree7ALkQhexs9tq7iEAM72QRYz9j7HfLd5g7oLSKRNvxzVHO0GdHhb0S+Sns8LseEVtuh0m8G2otaZKlpb/KEylTi6sjP93BA8BcAyfO/EgED/4rbiMdIRcR8d5Qmi7qOt+4l1qKAFNhwmcXivTL19mFT9JHXVnnMj+hEQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786102658; c=relaxed/simple; bh=oNBzqD9jsT2yunyygkm7+etc0sV7BqK7one/jTkqfi0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=R+Mi7vTGYVTIcfcu9xsdUPQX6INu/t+79LF/KEJyb3CysR7EB0wz52I/Mx8SorEFQT2Z54CWEgWHs1yT765EMU2H+lfSx4X0eIWpsGAxf65Q+zvrOo7he8ocTDGiZ5uHGWC90ffmLPDjIhiKglYAhbnSQ43eUbZ85VQXwFSW2vk= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=xbftsFIg; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=mNnlgo5w; arc=none smtp.client-ip=202.12.124.157 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="xbftsFIg"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="mNnlgo5w" Received: from phl-compute-01.internal (phl-compute-01.internal [10.202.2.41]) by mailfhigh.stl.internal (Postfix) with ESMTP id D90FE7A0120; Fri, 7 Aug 2026 07:37:33 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-01.internal (MEProxy); Fri, 07 Aug 2026 07:37:34 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786102653; x= 1786189053; bh=j0pTWFpZyp+VvcRD0WxQqHE3ctvFSewhwoAIj25qxck=; b=x bftsFIgYkMSok8nICxfbAq1vmfsz4Z0KdfsegX/1uf08o1cQaXApfJ5J8j0OTmd5 C3PiD41M+AzixMpV8qI9YmCiXzNyGUd0NdwDVL7izDqRe2IjfthjC+myekxxpZE6 xZkLw3stIVqqyrw66g1jQ6DTxLsmsTQTN/relElWyUQaW6MG+570UBEiaIC/hSc8 rUBIKuXZy3zBQkTNbhXwteZi+tD7y+ewMMG0PZK4knc6JQ1K8vMPqe0a+Cioh61s +zJyBW2lcVo2ZrJ/eaWhAZpOGVYQyB9u6Z95p7x8syxqVzm5V+cfoFWjqorKQWPo OwctHlJA+VDcceIGu6DvQ== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786102653; x=1786189053; bh=j 0pTWFpZyp+VvcRD0WxQqHE3ctvFSewhwoAIj25qxck=; b=mNnlgo5wXwMK2z/P+ oeXSTUdjLJAgU/EBotX2PNmPJa0tc/1GpX1VDT6yzsXNIaddJn8xpQD+p/a9Apmy bwP/pyrqEVLDM9feG7wzBEyWazq0Qd8tozZOg6kDy7IG+LdD6k3Wv6HRYabHpmIm sF6eZBZ8mGE4Ch7U7/LaPZQO9qoqIS9uWN3Ea5A4a5NL0hjc1cym9Kulz2oq9xeN u4Fqe/FTcpOLleg+ds/3/c29oj3/ml/Vqfb1ka5nUSmFAAVzJOaJ+ouTwTdFb8v8 L9/gDxGWDNnfQ9KD0PUSniknK3Is0dqAcxhqqdOenvof00yi/f0jXXquEvznI5Zq iA1Fw== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFHgVqOo5hmjJljb3IAZGkXyyxzQaD6i5RNmNk424gVmZkPzr1NfNWN42YMHBLBJ7 ouwgSkV7uP+O116AnyQ77X3VN5vAdqq1zvzV33EYJRHUa/cQ0OT9UF9tJbf4RJE1rh86et dgK/XsSmEjRvIZ4/DwnHkulf5IWkeizHGDAErUJF9mnMadsJqDG0SAHuOjgfipI02yAqvF a9Gp7D1P4aM7ojieUjy73reTq6pHjAlslJzhWF2Bayc7rzuN1loFnb3tH1A/9l8/u/BKz2 kEn4hJCLyjrltgNf/ADveISAZdc/G8K4aY+0IiJqAeENTN6qSce54NhN3knOmRl67eS+Qh /GDhBAG3wst46ljWIEG6h6ptOpuEu5vMqJ7CiNfirp26y5hXcRosGEtwZNX9q+/Mx5m/Wm 18OvmXPsEUykWNTvUXxAMk9ehrL5JG3kQbKJE6IMEikl2yq2a8p7fPSXxLgRyc+IBLe4eB sHARQUv0hKUKXc+Rk3Cw9iDyfk+gvnaNxGQ9Gr62oeu4PJGH5NWqcHovNLmY9RM7CkdLgE kG3sT4dCq8bDdk36o+UvVus9mH46NhTOVedbykvpE+hKqvsrqn8J3aRu6ncnXEqXGp8lcJ ktWemNxd0IqVDSIhQADkbPFZlfudXSNAeqIjWmB5lgnaRnE6RQ2TpJsIS97Q X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Fri, 7 Aug 2026 07:37:33 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v2 17/17] selftests/mm: zap whole PTE tables in the khugepaged race harness Date: Fri, 7 Aug 2026 12:36:47 +0100 Message-ID: <20260807113647.3744609-18-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260807113647.3744609-1-kirill@shutemov.name> References: <20260807113647.3744609-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The harness's MADV_DONTNEED thread zaps 1 to 32 pages at a time, never a whole PMD-aligned area. Empty-table reclaim (CONFIG_PT_RECLAIM) only engages when a zap spans a full table, so no soak ever ran it against a collapse -- fuzzing had to find that class instead: the page table vanishing between the engine's park and install passes. Make the thread zap a whole PMD-aligned area once every 64 iterations, and keep the fine-grained zaps as the common case. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) Tested-by: Muhammad Usama Anjum --- tools/testing/selftests/mm/khugepaged_race.c | 19 +++++++++++++++++-- 1 file changed, 17 insertions(+), 2 deletions(-) diff --git a/tools/testing/selftests/mm/khugepaged_race.c b/tools/testing/s= elftests/mm/khugepaged_race.c index abd142db7cbe..1fbd20771a5a 100644 --- a/tools/testing/selftests/mm/khugepaged_race.c +++ b/tools/testing/selftests/mm/khugepaged_race.c @@ -121,8 +121,23 @@ static void *dontneed_fn(void *arg) unsigned long page_idx =3D rand_page(&seed); unsigned long nr =3D 1UL << (rand_r(&seed) % 6); /* 1..32 pages */ =20 - madvise(region + page_idx * page_size, nr * page_size, - MADV_DONTNEED); + /* + * Once in a while zap a whole PMD-aligned area: only a zap + * spanning the full table triggers the empty-table reclaim + * (CONFIG_PT_RECLAIM), which can free the table under a + * collapse that is midway through it. Sub-table zaps never + * reach that path. + */ + if (!(rand_r(&seed) % 64)) { + unsigned long area =3D page_idx / + (hpage_pmd_size / page_size); + + madvise(region + area * hpage_pmd_size, + hpage_pmd_size, MADV_DONTNEED); + } else { + madvise(region + page_idx * page_size, + nr * page_size, MADV_DONTNEED); + } usleep(rand_r(&seed) % 500); } return NULL; --=20 2.54.0