From nobody Tue Sep 29 04:44:47 2026 Received: from fhigh-b1-smtp.messagingengine.com (fhigh-b1-smtp.messagingengine.com [202.12.124.152]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A06B844C645; Wed, 12 Aug 2026 13:23:11 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.152 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786540994; cv=none; b=hUr6bSLkZy2EHnFeXbDqnXpQsFZSIG4KCl0XurxnYrWqwac9L3l58X7ZEkfGQJaUdgL9+0+t67NzFDJ9YV2KY1mPVMPdd/4/bG/c91de1KnZ2eRrxU2Hf9qqMPkSS2tlOncbrmX2GtikgOIpXheEfFJmOlmfuYas2avyOEZH0ro= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786540994; c=relaxed/simple; bh=DPU3hl4iFCHFuBCJbEFx058WI14nmACSIi1wqqVRxh0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=KrWSiusldy9Fi3lTrlBH0SXJr+q6pIMxOYaxqUWSipbb6aIH/2TKuIQ6Rqx2ZURAEiPBX40c2npx7HxnlNSC/obu1Az0iiTHE9u+juyuGf2tVRqWuNu24DwBrXrfSCiPmJhoNpfwhnrbqphLjkZia2/MtLich+vu7Fr6CoMlZKs= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=w6sy2mRP; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=WS/CpC5r; arc=none smtp.client-ip=202.12.124.152 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="w6sy2mRP"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="WS/CpC5r" Received: from phl-compute-04.internal (phl-compute-04.internal [10.202.2.44]) by mailfhigh.stl.internal (Postfix) with ESMTP id 2C4E37A00F7; Wed, 12 Aug 2026 09:23:10 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-04.internal (MEProxy); Wed, 12 Aug 2026 09:23:10 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786540990; x= 1786627390; bh=dFVegsV3dbue/WpZaO1ywLiWC7Si3PJqV3I+1pYiJpM=; b=w 6sy2mRPTMtYuyXT7SOYVmXl5NjLm5GBxukeObUH6rzKetbsBFV7ASdEGTVRAj3HI pxOaZD7F7GXV71xARLgxnUQM2FF+CvAJLyi9Mls8EgceOTDHfn+RQSfyZtVXArgO XrXZIHwlTb5L/VMrtPR81C8kpz9PJ4N4LR47yH//mqd4ovCWDGNAMtrfpP6e6zpb LI4ubGbq1dU3tlAkmT5afEWWZKcnqVU4PFR1nRt/lxhykKhAa3+jFPjGYqRwMEEV XDyv+RCUiZX6YJh6EjnQBwKesj+UIZwrPJ+zsSOfv+AlNQlQjEnJZn36rs+w2b3N 4xLpCzLg11Rx/GwWonfwQ== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786540990; x=1786627390; bh=d FVegsV3dbue/WpZaO1ywLiWC7Si3PJqV3I+1pYiJpM=; b=WS/CpC5rJP+IE6ydp +eyJ09FCWELGS4/g2Iuwq7Hsz7G5mNhaG07e25lecZMwxPjjglJUT7r7w8ehrVS6 lwWN/W9LyMoEk1j+xtxLmQFLZhc4Isy0PkRmFDbtORFGwohG25GmEH9MvqK1q0zQ ryllkbLOyc82gx391TeChv3ZApqhb0Kk7RoH0/Vf+mi4ygCI+hXfLPSdcQvohItW vNqF//RSpAKgTcizY6SMGUpYjnJjChg31+cjGtKBkJvC9VtbsjkovwbEw3HkzMnC jV/HW5iwmgzsdD9/l4E6Kbbd5znTM0RpAdBlfVD4HPRy9Q7PGa3O84cdEjbDi62f xBitA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEWZPjKBYDoSh3uCuJ0ogiZvSuDj1RAgMuSmimalRft9LCPvZxyrluChLfTbkjo2r KkCK8Ttm4mqxi+W2/UjRoyI4gRuaMCGjIyoE2sYafZhhTXl0iIGCKEOr9/WVtZfHPc/pjz XsLpzK8U2Xu6F30SrJOujIlp1bmlomKUtpU/iiXSkRy6wwaI7NV8wAm1PQ0lHgRG0nCRKX WKIh8TtevEly/nhHwG2deu/klJhGkXN7h/CDYOOyqjOIkeF7MAFsC6fZtQDVeVkgzM4xU1 Py5NrFUehZdXXi5LHUYxqNjNVNbuxWN2P2YJW6J2IL8s+zlVHHPN2YdKOlFVuMcbnWK9RR he9rwe5OluvMSoZqJuhbcBqFRkamYKpUstN2ZIOi0PqxW8vU9FKgLKQmDf6VUCnKAQqAKK LKfSV1JqAWZZI35Umt11vo+D0VQpW5CMHQewiZHIhNYSwL0WunauJZTW72Si2RCbBxCrI2 BLsRPDosTNNtjZRTQiK/cLEA+xKIOWLPVccONblN75rB4vSsls/CzPnovVsLJhOpzFiS/B UqomOZqCJ0Lu+0+wAtFKYPFd/4sE+/tKt8gMwGA7v8pgFrE+DhcglpAgXlu9z1idk3YvNd nA5AZbt8QYP3wW1iBN//oI/dRAJZiIvYp94BPDbFE28g4ZnWdjWyEuJzas9Q X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Wed, 12 Aug 2026 09:23:09 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v3 01/18] selftests/mm: raise the khugepaged test-case cap Date: Wed, 12 Aug 2026 14:22:47 +0100 Message-ID: <20260812132304.199287-2-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260812132304.199287-1-kirill@shutemov.name> References: <20260812132304.199287-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" TEST() ends the run with "MAX_TEST_CASES is too small" when the table fills, and the table holds 64. A full invocation -- every context crossed with every memory type, which needs a directory for the file and shmem ones -- already registers 63, so the next case added anywhere aborts the whole suite before a single test runs. Raise it to 256. The table is a static array of small structs, so the room costs nothing worth counting, and the cases added after this have somewhere to go. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) Acked-by: Usama Arif Reviewed-by: Mike Rapoport (Microsoft) --- tools/testing/selftests/mm/khugepaged.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index 0d6c71ed2fae..018b0698229c 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -1324,7 +1324,7 @@ struct test_case { test_fn fn; }; =20 -#define MAX_TEST_CASES 64 +#define MAX_TEST_CASES 256 static struct test_case test_cases[MAX_TEST_CASES]; static int nr_test_cases; =20 --=20 2.54.0 From nobody Tue Sep 29 04:44:47 2026 Received: from fhigh-b1-smtp.messagingengine.com (fhigh-b1-smtp.messagingengine.com [202.12.124.152]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 38EEA449B05; Wed, 12 Aug 2026 13:23:14 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.152 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786540997; cv=none; b=gtoHtKv6QbY/zLBIcDd1cdNDRTnv9Mx9Wt3SPYz83HqspBp0fkEIppDLMROa8rt83qBLUCJESUYFAAxZyAQeH0Ksw5maKWvl6jyYiZwJIdf9zC0Dq21trxtVljnYYSYCKBX9TSYCrq8eL1ijVWPnZFYWuKyzbx+Rjk8tvWELX0c= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786540997; c=relaxed/simple; bh=H0TqXrshxFAXwdZIal+W2xcXTqtId6iLEnI8xJeIFlQ=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=jqaCR7iuDZ9jpsemiVQepQdic9M74j2xhMiIVGwoNxu4U32Tk2q5q+6DJmq/8sqTjIlCkIP3QkcwEzU5w5+xh8HjYcZ4T9s1MZ27HmFmBp7s6nrD75BxXW1eiHqhLTgO5ugZjdObHhCZ0yVpYWVl34tGnGrZkun+GcDp8FhayA4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=pRC/52Nz; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=DWl702yu; arc=none smtp.client-ip=202.12.124.152 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="pRC/52Nz"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="DWl702yu" Received: from phl-compute-01.internal (phl-compute-01.internal [10.202.2.41]) by mailfhigh.stl.internal (Postfix) with ESMTP id AF2DC7A0106; Wed, 12 Aug 2026 09:23:13 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-01.internal (MEProxy); Wed, 12 Aug 2026 09:23:14 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786540993; x= 1786627393; bh=QU5cOor2Ab9pLoeGxJ+Fiu9sKQhetOKhAO/atZ7UTrU=; b=p RC/52NzJJsirR+4TDQ7A08DdLIlBhGNLqUZLDkxE7dPfazDfVjp1DZdT9Ev5mShu vY+ADo1LybGGDNx2gTMaztsuRmpJq2ipzy6Mlx4Wy47ib+xRjHniNzlOgSrJY0Xf Cczpk3SlgHB4vjsJv/vMSwRDuH/4FxE6bRVExfluLUZ0FrYtUdaZPiJ7II4O6Kxq 2wJETt5+E5omyL53ncUHwyT7uAVrNMTiOWBO5HDlD7Iygtq7s/7xgKpwh9kMMYfP xJqediQ/D7uRXY4LEJHIWG+9WIoEHq/KNmh8Kmz3MKh6L9yV8X46bJyoU0zIY/OH HKgDwLRPAIpMQZ6X6na0Q== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786540993; x=1786627393; bh=Q U5cOor2Ab9pLoeGxJ+Fiu9sKQhetOKhAO/atZ7UTrU=; b=DWl702yuuHnkiOlUX X+ssNVIK6b19rNaaywVIH+CKKITcBO+w1IllVjvijCCevOg+gzviAqtL3LxMR+Dh e9sT1MgVr9NjCrV3xMFwckk2Wi205eEMHsDfds8zycdBrqCKaTTsYLMLUtZRtYds 6G3git8eoEaKmP/dqMthOjignJOuMb0W4M/wDOwIXwLxOoga9XiEwM/ehwAwqkRo nHsexYmodp7LPUL3uDYsG/T5z3N2NrxbnaRyYBx53FwHRT8RKgN4oRauHcb8Z+rw GhHePx8sjGhQhe36kvNas2rczX0Ok9al/I8l6K2eKYVXxzaI3n9poYqpjt2Yxb4T AHP3g== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEWZPjKBYDoSh3uCuJ0ogiZvSuDj1RAgMuSmimalRft9LCPvZxyrluChLfTbkjo2r KkCK8Ttm4mqxi+W2/UjRoyI4gRuaMCGjIyoE2sYafZhhTXl0iIGCKEOr9/WVtZfHPc/pjz XsLpzK8U2Xu6F30SrJOujIlp1bmlomKUtpU/iiXSkRy6wwaI7NV8wAm1PQ0lHgRG0nCRKX WKIh8TtevEly/nhHwG2deu/klJhGkXN7h/CDYOOyqjOIkeF7MAFsC6fZtQDVeVkgzM4xU1 Py5NrFUehZdXXi5LHUYxqNjNVNbuxWN2P2YJW6J2IL8s+zlVHHPN2YdKOlFVuMcbnWK9Qx kVqpeqVrcf8ZaStEY2VI8+3FWhIr21qDtmYPrSFHLNbwKtjKBMX5CJhId5vlxW1ejp+ZbT UQbEYHl8E0uRICu7ifmawK/TtGPzD9a2v2Ji8WOvraPW5mnOoDCNpqpACSUFLHEzdJY2xz +M5SOZpuG/y8PGt3CPQlLKczqCAHQeBmpGFllRpnfIT7qZqUY3cOJeUbWKswmuJkgk4ylz yXzzt+h9A4hT4cKZxd/tfRdcL9mdMK23tw3+UIvAnIz3tSRfwXHrJd8VaZt8wDmoxfTM8f ba6RfcTYpmZBdecltXKgKCyXs9QlX39+utlhyNGtdi4dYkR0EvMWRVCiudag X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Wed, 12 Aug 2026 09:23:12 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v3 02/18] selftests/mm: skip collapse_compound_extreme where the PMD is too large Date: Wed, 12 Aug 2026 14:22:48 +0100 Message-ID: <20260812132304.199287-3-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260812132304.199287-1-kirill@shutemov.name> References: <20260812132304.199287-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" collapse_compound_extreme builds a PTE table full of distinct PTE-mapped compound pages by cycling hpage_pmd_nr fault-time THPs through mremap. It therefore needs hpage_pmd_nr PMD-order allocations in a row. That is fine at a 2M PMD (4K base pages) or a 32M one (16K), but a 512M PMD -- arm64 with 64K base pages -- makes each of those an order-13 allocation, which the allocator cannot reliably hand out even once, let alone 8192 times. The failure is not a quiet one: the case calls ksft_exit_fail_msg(), so the whole binary stops and every case after it is lost. Skip the case where the PMD is larger than 32M. The MADV_COLLAPSE cases still cover PMD-order collapse on those configurations, and 4K and 16K PMDs are unaffected. Assisted-by: Claude-Code:claude-opus-5 Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) Reviewed-by: Mike Rapoport (Microsoft) --- tools/testing/selftests/mm/khugepaged.c | 10 ++++++++++ 1 file changed, 10 insertions(+) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index 018b0698229c..48eb74c255f6 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -974,6 +974,16 @@ static void collapse_compound_extreme(struct collapse_= context *c, struct mem_ops void *p; int i; =20 + /* + * The test needs hpage_pmd_nr PMD-order allocations, which is likely to + * fail for large PMD sizes. Skip if the PMD size is over 32M. + */ + if (hpage_pmd_size > (32UL << 20)) { + ksft_test_result_skip("%s: PMD too large for fault-time THP construction= \n", + __func__); + return; + } + p =3D ops->setup_area(1); ksft_print_msg("Construct PTE page table full of different PTE-mapped com= pound pages\n"); for (i =3D 0; i < hpage_pmd_nr; i++) { --=20 2.54.0 From nobody Tue Sep 29 04:44:47 2026 Received: from fhigh-b1-smtp.messagingengine.com (fhigh-b1-smtp.messagingengine.com [202.12.124.152]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9026237DAD7; Wed, 12 Aug 2026 13:23:17 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.152 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786540999; cv=none; b=YBNarF5b3IRR+5D5BKUgrrXpQ9hhuorDLC11BXjWyNN0b9xl8hCHVjw3U1bGLde/1D5cBVVrvaLeFuEP7DSk8E1c+pApQIqnAoJxV5K21NLc9Dwxk7K2ctCwetImtP333lpjJviDF6E/2woGcKvz897D3GIuI4XXjG/rR+ypKjQ= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786540999; c=relaxed/simple; bh=tYHjq9AfOnS4weJ6c4sLCMMcu+sjNXqoJxifv8Shr/k=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=WOabeqLS+dcjmuDF5tvnwPUNlpsjRC/t1Z2sF5r9kVwHkSOk70WWJ9fYoXVKu78/rzs4l/lvMT5y77/H9MGi6PHFH9F9EVj8VjTt6JjVFXlhtul3VWFAeGSpC94F9JHPNtvt0uqH9uDqYyUo/N44fYK/cKADfCu3B987Di+8phI= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=aiiosAOa; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=D1/ntZao; arc=none smtp.client-ip=202.12.124.152 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="aiiosAOa"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="D1/ntZao" Received: from phl-compute-01.internal (phl-compute-01.internal [10.202.2.41]) by mailfhigh.stl.internal (Postfix) with ESMTP id 478457A010C; Wed, 12 Aug 2026 09:23:16 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-01.internal (MEProxy); Wed, 12 Aug 2026 09:23:16 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786540996; x= 1786627396; bh=MJbY/SFy75qdRi7w5cKO1mRNu7cIMLlux0wZ8SdfTcU=; b=a iiosAOa4lnIcQmGxw3xyoXiABW5LfDKUemJ8PwlNXaysVYmy1zfqoZSnDyQN+oAi W2V4GVNkMjAzEeNPyAzVb9uUuKKR17KLN8zV6V1BQFPDxNgJHbT6CTM9JNUcjVUQ K9KWplKuoY2m3pfrchow+I9upXOXfVg6kF+uz7KcOXP0H8I7vXtHeCkuDv7SzANi XDZpXth0IYOL7GY+P6qXA+JOT1DhRWS/eOwQTMvvmftcZ8EILT6qmA0w/8Bl0Ocy geaZg01+6Re1R4DUHmsIfiyP8+ufAvQC8ZZVsG24mzocC7AUCCj5qk51PC6uUfbv b779JmXqIqHrSy1ahwO4w== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786540996; x=1786627396; bh=M JbY/SFy75qdRi7w5cKO1mRNu7cIMLlux0wZ8SdfTcU=; b=D1/ntZaoXHDqQY25H Y2A5uM4Zy9Mfrz6K86Xnut4YEMMtHDV+hRovrpBJ4d/KFoVQPNxBogKLXJ4VjOyq kSS7udh0fK23gd41EaTftHSIJH4ze6kCo+bM0uvF5ipVRJbQNzf1GVAD8/N2uYAY 7s3diYhQOxDuUtfQfOx3mmGUHA7YzF6jG6JBKsSb4/66sGJiAOSMgpoSOEdXsFS2 Vuw3td++xfIWBweXiSER/q7hQPcmLZANQKO+tVd+hW290a+pFBaFrwDVag8JAkXX LkZkt3REhbSYkDmacPPVHWa/6ywQoHn9s/7og6CX+VtA3+focCjMxxMta4+DR9NF d7TxA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFzqtrNoSTXX2MRzKyN6p8XXqdjDSX08KYyJpXp82eYC1ae2GegaTOhbDzreYT/K0 OAsc1zsfjWszwuQXvOMxV3FM9x50+HqAsA48u5Kq52QtHo+TJJPCEG+EOeKn5Il8AepNuY amkSO7LHncjzMJKy/SPUs50UUrTlzGP0ayf7b+h/q5NreZMQSrrfXSzcd6tsdBCg4Hhmbc 7COQKDHpqWJD3HpoJ7/Bw0J3tcRV9tZIkmRGQ4nZ/O76iTBc+eN6GYtayu9aLLTJal0DNN hI5LqlugC7KhSFwlZIIFqKKnJB7XY7vZnEBEdRJ8Mvpq9sE5OflKN5J+H3ujbniJsnhlY7 qvofwkhupd3+ldjSFNE07ZJ6AiypvcrGoG9bilNX/w56U6sBddmljfwCtBGcGBcJl6c1dK 6LGhiz6rnHZc9AbDCNDPCZw4HgoNvPHpE0+SgOAZ1eLVGqCVgR/KwLxiC3cdzJfUHp6670 se8/8fH+r0wC8pjS6JwIkHpLmpsNqcTEYrpHAj6zHkZ6FflNj9aLmG9tSF58bwKbKVAHEn O4OccZKzkSFnPJsLnotJOSRo3IRFIx//B5JtVWZLhX8tnh7KHfgOYrELZm7LsxWFEfS9xX AygVIV0KsKuBFx18VKm0EeWgSNMTbkyUC553L7qTZtO5Ehb/HBGxlMszaI/w X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Wed, 12 Aug 2026 09:23:15 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v3 03/18] selftests/mm: scale khugepaged's collapse wait with the PMD size Date: Wed, 12 Aug 2026 14:22:49 +0100 Message-ID: <20260812132304.199287-4-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260812132304.199287-1-kirill@shutemov.name> References: <20260812132304.199287-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" wait_for_scan() gives every case the same three seconds, whatever the huge page costs to build. collapse_full asks for four of them: 8M at a 2M PMD, but 2G at a 512M PMD -- arm64 with 64K base pages. Three seconds is thin at that size rather than generous. Across 80 runs of collapse_full on arm64 with 64K pages the wait was half a second in 73 of them, with a tail to two seconds, and the case has timed out in a full matrix run, reporting a failure for a collapse that was still going. Keep three seconds as the floor and add a second per 128M to collapse. A 2M PMD is unchanged. A 512M PMD gets 19 seconds, which is headroom over the observed tail rather than a measured requirement. The budget bounds how long a real failure takes to report, not how long a passing case waits: wait_for_scan() returns as soon as the collapse turns up. arm64/64K: khugepaged all:anon 21 pass/1 fail -> 22 pass/0 fail. x86-64 is unchanged. Assisted-by: Claude-Code:claude-opus-5 Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) Reviewed-by: Mike Rapoport (Microsoft) --- tools/testing/selftests/mm/khugepaged.c | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index 48eb74c255f6..a5ada78d90ee 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -583,8 +583,10 @@ static bool wait_for_scan(const char *msg, char *p, si= ze_t len, int nr_hpages, int collap_order, struct mem_ops *ops) { unsigned long hpage_size =3D page_size << collap_order; + /* Three seconds as a floor, plus a second per 128M to collapse */ + const unsigned long bytes =3D (unsigned long)nr_hpages * hpage_size; + int timeout =3D 6 + 2 * (bytes / (128UL << 20)); int full_scans; - int timeout =3D 6; /* 3 seconds */ =20 /* Sanity check */ if (!ops->check_huge(p, len, 0, hpage_size)) --=20 2.54.0 From nobody Tue Sep 29 04:44:47 2026 Received: from fhigh-b1-smtp.messagingengine.com (fhigh-b1-smtp.messagingengine.com [202.12.124.152]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3F0064503F7; Wed, 12 Aug 2026 13:23:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.152 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541001; cv=none; b=HiKkVQ6NGQ9il0j8hXQZMCEsaDXS/1JM5yFutM+CV2NzP7KrMccdhI24xKjTv95f2C2GwEI25hhoYSmtY+T5eQi1o5idAWwz95rgryDD1O1iZ85f5N0Dr4QmVw2YQFh9p45478Sm1j/vWeiGkWcsO99d80BZI2q+NeKBmp5Ells= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541001; c=relaxed/simple; bh=MCFYuxTEdi/0ELeKv4opLLuRr3T+vXEvJTsuz8bC+os=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Dvl3BupGgHNhww6oEN5csq9DnLTX+yBW3ROl6Lk1UZjL6RSK244EPFjnybqeB4KP5y7ixciKgcabwv/K/DYnk+by42CHTd9L3iXk6g9qIxbkcTPwrWaTVFYcM9h55e1DlvYs6JBpY9LB7eLLJoPJYnJttGX33h6MjXLeFRV8yaM= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=hWR6F651; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=WBxWWHuT; arc=none smtp.client-ip=202.12.124.152 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="hWR6F651"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="WBxWWHuT" Received: from phl-compute-01.internal (phl-compute-01.internal [10.202.2.41]) by mailfhigh.stl.internal (Postfix) with ESMTP id 1042F7A00D8; Wed, 12 Aug 2026 09:23:19 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-01.internal (MEProxy); Wed, 12 Aug 2026 09:23:19 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786540998; x= 1786627398; bh=DF6XlsE3dfRr0xz9ihNAdyd1y2VRfqNvU66IRArOZEQ=; b=h WR6F651wjr7kxuKTtzxDR0x9diJvksCsJp8s6NJEtgTnAxeA5FCMl0hhZ6kAvvxr gTainIu9jqgofPJKLylovUYHVEQZAdQfXrKCcyo36cWA8dDexeOx4kBRJ6CPwdgh 30G/hgL811VYz97ehy1oxWNjMlhEEgVJvlBs0udJZyPGu88vjJIGfi2sgqS/Ha2z U7WTXrFaOpiUGf6foEV03P0U2mT4IwxSxiZZCdcvVkgjiyLsLfQa3f00rsApdI13 QEIHQ3OLXWsQY6PKquk7ZQDMot7jLHOJ6w++JKee5Sazzbh/AmFwqcq5Zqb8ZE6s ByPDZidU3a8AtZ+mlbfdg== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786540998; x=1786627398; bh=D F6XlsE3dfRr0xz9ihNAdyd1y2VRfqNvU66IRArOZEQ=; b=WBxWWHuT6MDoSl//n MJYbu/wzrDKD0JkzyGIqXAPcL/VtLdp02tydbvPq5iUiU5g0BCR9IJ7EqHkK64hl g3NTH0macTaCkGjfaI3UEkfVOF+OgO4Xs1KYPsl+Cm5uQTTeq3E8Xp7PA8otJQqr 3hg7zsTSa7MkiWpU1cxuEf2fnescopubdtpqkx6RIV7i2T69xp/Jn3/fxZDAN8hL zjWmNWNS0QCM3GLA2/1IM6nENYFLuaZV3jqM3FAgA1UQ0+WaV7HdH3uz/oZ8fEUL N0u6HHQTygdH57XsT5LVvc1h/TAFRkERiRq3ko1Xu13q1zdbq6oCioZstDPTbZV+ dTLzw== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFzqtrNoSTXX2MRzKyN6p8XXqdjDSX08KYyJpXp82eYC1ae2GegaTOhbDzreYT/K0 OAsc1zsfjWszwuQXvOMxV3FM9x50+HqAsA48u5Kq52QtHo+TJJPCEG+EOeKn5Il8AepNuY amkSO7LHncjzMJKy/SPUs50UUrTlzGP0ayf7b+h/q5NreZMQSrrfXSzcd6tsdBCg4Hhmbc 7COQKDHpqWJD3HpoJ7/Bw0J3tcRV9tZIkmRGQ4nZ/O76iTBc+eN6GYtayu9aLLTJal0DNN hI5LqlugC7KhSFwlZIIFqKKnJB7XY7vZnEBEdRJ8Mvpq9sE5OflKN5J+H3ujbniJsnhlYI xji+6SmTFIQHxOuLmkmDVGavg0RuSKJZF7XoXVbUudqM9qjcQTvodLsb6C8iCX1SNL3Pkr qTC/Gfpp7WJRpjVUw/htERWzSW6ctrDxWN5OIVKdPBe6rx8N/aiFz6c3DafXPMeq8VXFAi 6HVJlDqdRrV1tnvNqxl55teaK4UWpOdAfeUZk48tdS1rF6SAsYdUiF+wPN5/GM2xKWXzP2 t7JYdbF+su8Sh0O5pXCYqk8pSorwSfj/04N2pbZTnhDX9EacJQuHhurJ3mQzqThkw/jYu9 y4zbK892IeTWKp8C8x/yjYr8qVLaPz1JQ0vN7l8zE0Ssb5TNjrJeYffjqKuQ X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Wed, 12 Aug 2026 09:23:18 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v3 04/18] selftests/mm: skip khugepaged page cache cases without a PMD folio Date: Wed, 12 Aug 2026 14:22:50 +0100 Message-ID: <20260812132304.199287-5-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260812132304.199287-1-kirill@shutemov.name> References: <20260812132304.199287-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The page cache caps folio order at MAX_PAGECACHE_ORDER, which sits below the PMD order where a PMD is 512M -- arm64 with 64K base pages. A PMD-sized page cache folio is then impossible, so MADV_COLLAPSE answers -EINVAL and khugepaged passes over the range. The shmem cases ask for one anyway, so four of them fail and the run bails out in the middle. The cap is one global, so it rules out every file mapping, not just shmem: shmem_huge_global_enabled() drops the PMD order from what it allows, and file_thp_enabled() refuses a regular file whose mapping cannot hold a PMD folio. Skip both mem types when the PMD order has no per-order shmem_enabled control, which is the readable form of the cap: that control is created for the orders in THP_ORDERS_ALL_FILE_DEFAULT. A run left with nothing to collapse into skips outright. Anonymous collapse is unaffected: its orders are not capped this way. Assisted-by: Claude-Code:claude-opus-5 Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/khugepaged.c | 31 +++++++++++++++++++++++++ 1 file changed, 31 insertions(+) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index a5ada78d90ee..c049ac997def 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -1391,6 +1391,37 @@ int main(int argc, char **argv) =20 setbuf(stdout, NULL); =20 + /* + * The page cache caps folio order at MAX_PAGECACHE_ORDER, which + * xas_split_alloc() puts below the PMD order on arm64 with 64K pages. + * A PMD-sized page cache folio is then impossible, so the kernel + * refuses these collapses by design and there is nothing to test. + * + * The cap is one global, so it rules out every file mapping, not just + * shmem: shmem_huge_global_enabled() drops the PMD order from what it + * allows, and file_thp_enabled() refuses a regular file whose mapping + * cannot hold a PMD folio. + * + * The per-order shmem_enabled control below is what makes the cap + * readable: it is created for the orders in THP_ORDERS_ALL_FILE_DEFAULT, + * which is the cap and nothing else, so whether the PMD order has one + * answers for a regular file as much as for shmem. + */ + if (!(thp_shmem_supported_orders() & (1UL << hpage_pmd_order))) { + if (shmem_ops) { + ksft_print_msg("no PMD-order page cache folio: skipping shmem\n"); + shmem_ops =3D NULL; + } + if (read_only_file_ops) { + ksft_print_msg("no PMD-order page cache folio: skipping file\n"); + read_only_file_ops =3D NULL; + read_write_file_read_ops =3D NULL; + read_write_file_write_ops =3D NULL; + } + if (!anon_ops && !shmem_ops && !read_only_file_ops) + ksft_exit_skip("Nothing left to collapse into\n"); + } + default_settings.khugepaged.max_ptes_none =3D hpage_pmd_nr - 1; default_settings.khugepaged.max_ptes_swap =3D hpage_pmd_nr / 8; default_settings.khugepaged.max_ptes_shared =3D hpage_pmd_nr / 2; --=20 2.54.0 From nobody Tue Sep 29 04:44:47 2026 Received: from fhigh-b1-smtp.messagingengine.com (fhigh-b1-smtp.messagingengine.com [202.12.124.152]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2521443D502; Wed, 12 Aug 2026 13:23:22 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.152 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541004; cv=none; b=gSPNd+gmx59AAJVSZ1kwEC6FgRyvnZA2J+Kvw2EzNtw0kQDhMs1Gn/9W//GWL/E2TylkIwz4zTcy6+Wdre0YcAM2iFSfz2E8yS+3YXJ47+bLvUdWNUSEKMBZQo1LE70cs4lHJ1cnNNMLIPQ4Q90hm4kXBL2mf+dCtfIQ4dwxWuo= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541004; c=relaxed/simple; bh=CyfeleKriuIS1HvCKiihKcZb3I464m9rEOk828X1saY=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=G77cOEra71KSffx4q9Kk20sGHboFjH9d7r9h2q0+v4caTNJmpFh6FGQEakez2D8cJ+DHCZ9hiJoufhYckD9VpDZi4vqQIGXtxcCaCv4yyt0nfXcJmFn17tkVuxhFz2xOVo/8UknxqKDpnjgZBMQ0MlX5384ohgzDbz5NPrkR6E8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=unjqq8I7; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=HFCJV6Mg; arc=none smtp.client-ip=202.12.124.152 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="unjqq8I7"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="HFCJV6Mg" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfhigh.stl.internal (Postfix) with ESMTP id BF1927A00F7; Wed, 12 Aug 2026 09:23:21 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-06.internal (MEProxy); Wed, 12 Aug 2026 09:23:22 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786541001; x= 1786627401; bh=DdpQk+Hx40PJfU17MORiNeZLHwz7nsmGHrsYbIXLhW8=; b=u njqq8I7a/k1zSPjBpZ+XKODmjpe5C35qprcsdSvnx2XIQ3ve82Qx1uQ4M6Laeo/R nWIhP4zft+towIzrF7n67ZwLl+OMiwCSzcDqQI009RMGbe/PatfaQ5HLppREC3QB gWb+c2ntNYXwaJHSHn+vCNSskNbwxG67Pc7p3k2/fHBt8d5zAsRuoepTxNAvnsuU HE/fl6NV69rkqIQVwd2AkcViKJ7IcVDXKV6ZOJ8BMaxjFgDbGmrwMqoxDTy0tuYg 2fm6f3MY5E8ouTM8DAQR3AHDx205hSds10qD32gR1WCSxV7ktnJZuav7FpaHabMH ofmo/Pp1vhFbKB+4XmsHA== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786541001; x=1786627401; bh=D dpQk+Hx40PJfU17MORiNeZLHwz7nsmGHrsYbIXLhW8=; b=HFCJV6Mgjz/2NAcDf T8guo1yAwxL54jEcHRCl1teK01WRTmYb4YCfUShhmNYA4lTcFRrhrgsC34DQpXSJ QtPjkxpj8M1XK/eY9KwT4jZvY5NDm4Hgzjzri/TEkrZMJg283ipBCCsGlc3QOnZf FP7oOss7Ogm2LGRoCwi2xpcyF7AA2D/FdQLwIgIinIkA/YmJtJgXTjTnthlbagR7 fkZXJnDusbLfqTXL18S177AIdmQLQomG6Y5Y739GljVeTNWDuWPZRa3d8pqOrpVN Fs7AiBdRatQkCPbNOHJVzhoJmnlvCJThNLJgj0BhOLfshXCYjRszBgIh67mLjGx1 gEjug== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEWZPjKBYDoSh3uCuJ0ogiZvSuDj1RAgMuSmimalRft9LCPvZxyrluChLfTbkjo2r KkCK8Ttm4mqxi+W2/UjRoyI4gRuaMCGjIyoE2sYafZhhTXl0iIGCKEOr9/WVtZfHPc/pjz XsLpzK8U2Xu6F30SrJOujIlp1bmlomKUtpU/iiXSkRy6wwaI7NV8wAm1PQ0lHgRG0nCRKX WKIh8TtevEly/nhHwG2deu/klJhGkXN7h/CDYOOyqjOIkeF7MAFsC6fZtQDVeVkgzM4xU1 Py5NrFUehZdXXi5LHUYxqNjNVNbuxWN2P2YJW6J2IL8s+zlVHHPN2YdKOlFVuMcbnWK9ia 7TYo9q4qWL1B9/s0ZJakCaEATTtG18RYcIrguVUBbR2bL9tTJafJsGwoPXNytH5YxGRiu6 grTOV7pH/QxrbEPmMc+dcw1Y7/X9Z+NuzOj6vvbRXMRBde1CMd4VR9QSWb29CJBNaP1NOD b3c6jdZiHR5SvgFJNuvNN0AwxRi1gijWos12yTPYEG7xZiHfVS4xTvEgPbYOFwtChmIH92 Ioi7m/ZzGvHNgcBHduYcHSsixX8oCToRzYs5Fa21Um8R9L7Ekq5wd0lvGHFI8CbKDz8Ljs qp2yzUCSeZqnn1hdxCqTSHIW0VPP3E7kwG9QfvDF2aTFTcUZHJ7awAq7rntg X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Wed, 12 Aug 2026 09:23:20 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v3 05/18] selftests/mm: keep khugepaged out of the swapout the swap cases set up Date: Wed, 12 Aug 2026 14:22:51 +0100 Message-ID: <20260812132304.199287-6-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260812132304.199287-1-kirill@shutemov.name> References: <20260812132304.199287-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" collapse_swapin_single_pte() and collapse_max_ptes_swap() page a range out and then require smaps to report exactly the count they asked for. Two things keep it from arriving. MADV_PAGEOUT is best effort, so the count often turns up a moment late. And wait_for_scan() leaves MADV_HUGEPAGE behind, so khugepaged is still working on the range: collapsing one with up to max_ptes_swap pages swapped out means reading them back in, and the daemon empties the swap as fast as the case fills it. On arm64 with 64K pages, where max_ptes_swap is 1024 pages, that is 64M a step and the case loses: # Swapout 1024 of 8192 pages... Fail not ok 10 collapse_max_ptes_swap Ask again for up to two seconds, with the range held out of the daemon's reach while asking. The collapse each case runs next puts MADV_HUGEPAGE back, so only the setup is affected. If the pages still will not go, skip. is_swap_enabled() covers a machine with no swap; what is left -- swap too small, full, capped by a memcg, busy with writeback -- is not the kernel under test refusing. An error from madvise() itself still ends the run. Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: Kiryl Shutsemau (Meta) Reviewed-by: Muhammad Usama Anjum Tested-by: Muhammad Usama Anjum --- tools/testing/selftests/mm/khugepaged.c | 53 +++++++++++++++++++------ 1 file changed, 41 insertions(+), 12 deletions(-) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index c049ac997def..8458cd2ff0df 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -241,6 +241,41 @@ static bool check_swap(void *addr, unsigned long size) return swap; } =20 +/* + * Page the range out and wait for the swap count to say so. + * + * Two things get in the way. MADV_PAGEOUT is best effort: + * shrink_folio_list() leaves a folio alone when it cannot reclaim it right + * away, and one still under writeback from an earlier pageout is the comm= on + * case, so the count the caller asks for arrives a moment later. And a r= ange + * an earlier collapse left MADV_HUGEPAGE is one khugepaged is still worki= ng + * on: collapsing a range with up to max_ptes_swap pages swapped out means + * reading those pages back in, so the daemon undoes the pageout as fast a= s it + * is asked for. Keep the range out of its reach; the collapse the caller= runs + * next puts MADV_HUGEPAGE back. + * + * Failing to get the pages out is the machine's answer, not the kernel's = -- + * swap too small, swap full, a memcg cap, a folio still under writeback -= - so + * callers skip rather than fail. An error from madvise() is different, a= nd + * ends the run here. + */ +static bool swapout_range(void *p, unsigned long size) +{ + int i; + + if (madvise(p, size, MADV_NOHUGEPAGE)) + ksft_exit_fail_perror("madvise(MADV_NOHUGEPAGE)"); + + for (i =3D 0; i < 40; i++) { + if (madvise(p, size, MADV_PAGEOUT)) + ksft_exit_fail_perror("madvise(MADV_PAGEOUT)"); + if (check_swap(p, size)) + return true; + usleep(50 * 1000); + } + return false; +} + static void *alloc_mapping(int nr) { void *p; @@ -855,12 +890,10 @@ static void collapse_swapin_single_pte(struct collaps= e_context *c, struct mem_op p =3D ops->setup_area(1); ops->fault(p, 0, hpage_pmd_size); =20 - if (madvise(p, page_size, MADV_PAGEOUT)) - ksft_exit_fail_perror("madvise(MADV_PAGEOUT)"); - if (check_swap(p, page_size)) { + if (swapout_range(p, page_size)) { success("OK"); } else { - fail("Fail"); + skip("Could not swap out"); goto out; } =20 @@ -887,12 +920,10 @@ static void collapse_max_ptes_swap(struct collapse_co= ntext *c, struct mem_ops *o p =3D ops->setup_area(1); ops->fault(p, 0, hpage_pmd_size); =20 - if (madvise(p, (max_ptes_swap + 1) * page_size, MADV_PAGEOUT)) - ksft_exit_fail_perror("madvise(MADV_PAGEOUT)"); - if (check_swap(p, (max_ptes_swap + 1) * page_size)) { + if (swapout_range(p, (max_ptes_swap + 1) * page_size)) { success("OK"); } else { - fail("Fail"); + skip("Could not swap out"); goto out; } =20 @@ -904,12 +935,10 @@ static void collapse_max_ptes_swap(struct collapse_co= ntext *c, struct mem_ops *o ops->fault(p, 0, hpage_pmd_size); ksft_print_msg("Swapout %d of %d pages...", max_ptes_swap, hpage_pmd_nr); - if (madvise(p, max_ptes_swap * page_size, MADV_PAGEOUT)) - ksft_exit_fail_perror("madvise(MADV_PAGEOUT)"); - if (check_swap(p, max_ptes_swap * page_size)) { + if (swapout_range(p, max_ptes_swap * page_size)) { success("OK"); } else { - fail("Fail"); + skip("Could not swap out"); goto out; } =20 --=20 2.54.0 From nobody Tue Sep 29 04:44:47 2026 Received: from fhigh-b1-smtp.messagingengine.com (fhigh-b1-smtp.messagingengine.com [202.12.124.152]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DC903453A30; Wed, 12 Aug 2026 13:23:25 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.152 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541007; cv=none; b=kspQu0AcfpHtCZWlQAhi5DekKQJWQ3wy233Mpeh5zwW9ReJqKFECtksKPThpmnXSik/0TaE+OTaob4t+qkT0Y/92LMTxrf/FHrkogv6yqUIc8sG8JTWzzJK5uXABAm39ZPRxzx2cOsIJczmgYprFuhKU0rZV+KpbBPZ6w4PiDCk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541007; c=relaxed/simple; bh=W7ao8u1w7MSR5XGEvkcXkyP+JSEW5IKuVPy4aouKQ30=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=qVC3RynO/xZVDKrATsxknv3DzYuC4rxh6lxTj17zQuNt7bEyGVIWlB/DP7nZt/pM6GYPZb6YaYl+A7o1J4sDY2fg+3ajBBi5/mVJ7QuCmvEYsDQVQ2FGh3XadySlNrTrErctQF+EFFqJZ9Ekrp/hAWT238t6P3IHN7DlCdwVQ9o= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=w2m+TTze; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=URahXN19; arc=none smtp.client-ip=202.12.124.152 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="w2m+TTze"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="URahXN19" Received: from phl-compute-05.internal (phl-compute-05.internal [10.202.2.45]) by mailfhigh.stl.internal (Postfix) with ESMTP id 802567A010C; Wed, 12 Aug 2026 09:23:24 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-05.internal (MEProxy); Wed, 12 Aug 2026 09:23:25 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786541004; x= 1786627404; bh=efnlWGTfHMI6UMevBb/oPcFQfcqOFRMF5YmaGx8srVo=; b=w 2m+TTzekYG1h2SgpHz+DmFwZv8QAutAiRxEHITxyhsnhwvJcV0ffY2o81i38szb2 f12Dvx4R9vZXnO68cJo8Bh6fxtb9HY7CycYIdJOLVsFu2sTyPidiEq71XxqyhoSa V0PwymnO1QnfLn35SkvmzLGdm8wVxVuhGJXsYrmgPicrkIPyMu8NUcs47s9L9H/H EA9IBHprvsdhO1B0pd+QnPhQTi4+lljOGc/rbvMBqWjwmz8z+iYMd/iIFzDj/Gcn xtrIxivXSj6l5dkzO2C9qYd784fpvjqC/GqBy0Np+1tCMPcmXZocSg5lrhG+BnxX fwt88UmdnJZLrA04OH3fg== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786541004; x=1786627404; bh=e fnlWGTfHMI6UMevBb/oPcFQfcqOFRMF5YmaGx8srVo=; b=URahXN19dcwX7z748 0lGwTQVtoA+CYQz3cpK7WFrOHv/ag9pN6XtY9hA5Vf0yFXdYu3VR/CvBXWtXzCs1 zR1ieziL7KNn0nczoD8HQVuhUalYGO+Qsvy9roJ/GQhAuulrJ9GKeu8yKR5pHTeI XS7t7xJCzJhSABVyZzftZeBQDcxbNi7Y6xmzWFfkeWCEMLdw3OdOCQDgfGwVHyrn MGX7R9Z8a957xt1MAFyNz0WwWzNi0rs64+gudhyohRcQaAiAgfZwAkyrr+zqi2vS 1B0VMcA9dgtcmEX2/UNlDDubD7MQKgDF1wBAguyVknpaToHupi/MKicwAr9W1SMB as7nw== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEWZPjKBYDoSh3uCuJ0ogiZvSuDj1RAgMuSmimalRft9LCPvZxyrluChLfTbkjo2r KkCK8Ttm4mqxi+W2/UjRoyI4gRuaMCGjIyoE2sYafZhhTXl0iIGCKEOr9/WVtZfHPc/pjz XsLpzK8U2Xu6F30SrJOujIlp1bmlomKUtpU/iiXSkRy6wwaI7NV8wAm1PQ0lHgRG0nCRKX WKIh8TtevEly/nhHwG2deu/klJhGkXN7h/CDYOOyqjOIkeF7MAFsC6fZtQDVeVkgzM4xU1 Py5NrFUehZdXXi5LHUYxqNjNVNbuxWN2P2YJW6J2IL8s+zlVHHPN2YdKOlFVuMcbnWK9MZ cZzAisCav8t92G6y8EDE2O2NeImb6F3Kjul8xOVSg8SqONLZk/6PbrD52wNSqsb8obsWLR iTzUtunyCn4j5KEbFWlb891nKD/xanmWrCVMSvpCzuzJw9KBGCzmwmCvLBz3rjuKGU4szr tKrNL81C+NNfaaoVysR8XxTjI0QZ4FKoleflzByOKnCekcdRFTIUq6K5K3qyzNABWCkKDx WZQ8tJPYsIIFjU54wpPrYmGPFCHo+Zb8CW+gh6EY8aZUe6Vy9ca3raQLeiMaJHnsr/9F1d WIox07JLLn9J2mvuCq20B4HchKOFU0Z36EhhvHHORk2dGFo1TxmzMxwFBbeQ X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Wed, 12 Aug 2026 09:23:23 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v3 06/18] selftests/mm: move is_backed_by_folio() into vm_util Date: Wed, 12 Aug 2026 14:22:52 +0100 Message-ID: <20260812132304.199287-7-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260812132304.199287-1-kirill@shutemov.name> References: <20260812132304.199287-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The khugepaged selftest is about to gain mTHP collapse coverage, which needs to check that an address range is backed by a folio of a given order. split_huge_page_test.c already has the building block: is_backed_by_folio() reads the compound head and tail flags from /proc/kpageflags to classify the folio behind a page. Move it into vm_util so other tests can use it. No functional change. Assisted-by: Claude-Code:claude-opus-5 Acked-by: Mike Rapoport (Microsoft) Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) --- .../selftests/mm/split_huge_page_test.c | 62 ------------------- tools/testing/selftests/mm/vm_util.c | 62 +++++++++++++++++++ tools/testing/selftests/mm/vm_util.h | 2 + 3 files changed, 64 insertions(+), 62 deletions(-) diff --git a/tools/testing/selftests/mm/split_huge_page_test.c b/tools/test= ing/selftests/mm/split_huge_page_test.c index 86a603692826..0adfe7dde7e5 100644 --- a/tools/testing/selftests/mm/split_huge_page_test.c +++ b/tools/testing/selftests/mm/split_huge_page_test.c @@ -42,68 +42,6 @@ const char *kpageflags_proc =3D "/proc/kpageflags"; int pagemap_fd; int kpageflags_fd; =20 -static bool is_backed_by_folio(char *vaddr, int order, int pagemap_fd, - int kpageflags_fd) -{ - const uint64_t folio_head_flags =3D KPF_THP | KPF_COMPOUND_HEAD; - const uint64_t folio_tail_flags =3D KPF_THP | KPF_COMPOUND_TAIL; - const unsigned long nr_pages =3D 1UL << order; - unsigned long pfn_head; - uint64_t pfn_flags; - unsigned long pfn; - unsigned long i; - - pfn =3D pagemap_get_pfn(pagemap_fd, vaddr); - - /* non present page */ - if (pfn =3D=3D -1UL) - return false; - - if (pageflags_get(pfn, kpageflags_fd, &pfn_flags)) - goto fail; - - /* check for order-0 pages */ - if (!order) { - if (pfn_flags & (folio_head_flags | folio_tail_flags)) - return false; - return true; - } - - /* non THP folio */ - if (!(pfn_flags & KPF_THP)) - return false; - - pfn_head =3D pfn & ~(nr_pages - 1); - - if (pageflags_get(pfn_head, kpageflags_fd, &pfn_flags)) - goto fail; - - /* head PFN has no compound_head flag set */ - if ((pfn_flags & folio_head_flags) !=3D folio_head_flags) - return false; - - /* check all tail PFN flags */ - for (i =3D 1; i < nr_pages; i++) { - if (pageflags_get(pfn_head + i, kpageflags_fd, &pfn_flags)) - goto fail; - if ((pfn_flags & folio_tail_flags) !=3D folio_tail_flags) - return false; - } - - /* - * check the PFN after this folio, but if its flags cannot be obtained, - * assume this folio has the expected order - */ - if (pageflags_get(pfn_head + nr_pages, kpageflags_fd, &pfn_flags)) - return true; - - /* If we find another tail page, then the folio is larger. */ - return (pfn_flags & folio_tail_flags) !=3D folio_tail_flags; -fail: - ksft_exit_fail_msg("Failed to get folio info\n"); - return false; -} - static int check_after_split_folio_orders(char *vaddr_start, size_t len, int pagemap_fd, int kpageflags_fd, int orders[], int nr_orders) { diff --git a/tools/testing/selftests/mm/vm_util.c b/tools/testing/selftests= /mm/vm_util.c index 80bc9f597b52..5db1a7774f49 100644 --- a/tools/testing/selftests/mm/vm_util.c +++ b/tools/testing/selftests/mm/vm_util.c @@ -494,6 +494,68 @@ int pageflags_get(unsigned long pfn, int kpageflags_fd= , uint64_t *flags) return 0; } =20 +bool is_backed_by_folio(char *vaddr, int order, int pagemap_fd, + int kpageflags_fd) +{ + const uint64_t folio_head_flags =3D KPF_THP | KPF_COMPOUND_HEAD; + const uint64_t folio_tail_flags =3D KPF_THP | KPF_COMPOUND_TAIL; + const unsigned long nr_pages =3D 1UL << order; + unsigned long pfn_head; + uint64_t pfn_flags; + unsigned long pfn; + unsigned long i; + + pfn =3D pagemap_get_pfn(pagemap_fd, vaddr); + + /* non present page */ + if (pfn =3D=3D -1UL) + return false; + + if (pageflags_get(pfn, kpageflags_fd, &pfn_flags)) + goto fail; + + /* check for order-0 pages */ + if (!order) { + if (pfn_flags & (folio_head_flags | folio_tail_flags)) + return false; + return true; + } + + /* non THP folio */ + if (!(pfn_flags & KPF_THP)) + return false; + + pfn_head =3D pfn & ~(nr_pages - 1); + + if (pageflags_get(pfn_head, kpageflags_fd, &pfn_flags)) + goto fail; + + /* head PFN has no compound_head flag set */ + if ((pfn_flags & folio_head_flags) !=3D folio_head_flags) + return false; + + /* check all tail PFN flags */ + for (i =3D 1; i < nr_pages; i++) { + if (pageflags_get(pfn_head + i, kpageflags_fd, &pfn_flags)) + goto fail; + if ((pfn_flags & folio_tail_flags) !=3D folio_tail_flags) + return false; + } + + /* + * check the PFN after this folio, but if its flags cannot be obtained, + * assume this folio has the expected order + */ + if (pageflags_get(pfn_head + nr_pages, kpageflags_fd, &pfn_flags)) + return true; + + /* If we find another tail page, then the folio is larger. */ + return (pfn_flags & folio_tail_flags) !=3D folio_tail_flags; +fail: + ksft_exit_fail_msg("Failed to get folio info\n"); + return false; +} + /* If `ioctls' non-NULL, the allowed ioctls will be returned into the var = */ int uffd_register_with_ioctls(int uffd, void *addr, uint64_t len, bool miss, bool wp, bool minor, uint64_t *ioctls) diff --git a/tools/testing/selftests/mm/vm_util.h b/tools/testing/selftests= /mm/vm_util.h index 9a49af88702e..56a28ce7d029 100644 --- a/tools/testing/selftests/mm/vm_util.h +++ b/tools/testing/selftests/mm/vm_util.h @@ -97,6 +97,8 @@ int64_t allocate_transhuge(void *ptr, int pagemap_fd); int pageflags_get(unsigned long pfn, int kpageflags_fd, uint64_t *flags); int gather_folio_orders(char *vaddr_start, size_t len, int pagemap_fd, int kpageflags_fd, int orders[], int nr_orders); +bool is_backed_by_folio(char *vaddr, int order, int pagemap_fd, + int kpageflags_fd); =20 int uffd_register(int uffd, void *addr, uint64_t len, bool miss, bool wp, bool minor); --=20 2.54.0 From nobody Tue Sep 29 04:44:47 2026 Received: from fhigh-b1-smtp.messagingengine.com (fhigh-b1-smtp.messagingengine.com [202.12.124.152]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 668A44562BA; Wed, 12 Aug 2026 13:23:28 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.152 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541009; cv=none; b=ZiEo0PfGsEkcdkJsY0wPdsGruYkhlsqgQ6WOAURcb12KlXrcWPXd6UDdpJlEKO/5DGEtwGubBfVHYL+QorBiRD040m+VMoJE/098BxvMT4UA0avp6ZJtDBT5Zy0nClXISof6odpousPM4gU9iiOsmZVKXf43oN+/qdh2qwHuJ8Q= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541009; c=relaxed/simple; bh=JNDbRMrTvWoJBFV0dbie0OAT5FgLL64sP8zbsB7hhgk=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=l42nDwfbeTYcbZy5cj5fF36QrVBAuB+2BpWzTEVBFWqrB5llLtfTEi5pdjMbPdc78GGbv+yodi0ROB8ur7EE1Wc5CDzNlUXcuLi8uezx9E95NbG8DP+hcEk0OFC2efofytNE+iM/puMKjPBeRJavBUVUNmXk32D3Wqodor+KpoY= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=QXUbyLny; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=KPaN37AQ; arc=none smtp.client-ip=202.12.124.152 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="QXUbyLny"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="KPaN37AQ" Received: from phl-compute-01.internal (phl-compute-01.internal [10.202.2.41]) by mailfhigh.stl.internal (Postfix) with ESMTP id 2D21E7A00D8; Wed, 12 Aug 2026 09:23:27 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-01.internal (MEProxy); Wed, 12 Aug 2026 09:23:27 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786541007; x= 1786627407; bh=7/1Od9SJ6IhS+FCj32O5qawyk1eH8bdwgb4C18ahL3g=; b=Q XUbyLnygX6Hg80MGiM5NtW/Et2v4mgvuNkaN/mXkvIweCC6AMxMMBp4XytNuAXaB AqH2R9pIBMg6QfNb/IiphyivuSXTeMRjdrILMEEaQZVYuIom5xT9quL+p6k4eJZD YFoiUeTMdijymZ5tiOTn6DaqNEkdPVD9V2Cf/Bu4e4KWmDN9ez25+9lHkr7smv+G NAeIN5qrS43z1vKzPgxNUlsXI5/T26Uq2eT7nIzVO1EUb1hQE7Hh5ddN7S7k3jDL iNulunUp4VOe9Ha07eDG4eIfoIIopEBWxv/aCTwnIAKgxARe+zzMMm6hV1ZEw+1G G0MErmvqgPDPP38qJQkFw== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786541007; x=1786627407; bh=7 /1Od9SJ6IhS+FCj32O5qawyk1eH8bdwgb4C18ahL3g=; b=KPaN37AQb6Ihsr5EM bzfrZySkcG0An+V8801fZGNmWzjnrYcWn8iabQL/U/1tFZdjX6pwuv41yRZWloVP WG6INcHEWRF0QZjwpuRrohdIgy+E1yJ1nA0SACltYpI6ThddOcQAKq/0QQctv9s+ UXAOjycmwB37tcVa+JpC9VMcjd+zuhYwU7u/mytgV3lKWRuAeO7khEEIR0xP5Z3W 1a1Sf8pecn7bYh7yJSDQDHv5Boagq9i6hs3eW1tn/dqoY18XhK1UyNUAy49HZvRW f5a8vVUpcagvoU7uZqVMC00SEfVH4KzzvK65JpIwykEU3fIncuef5DRxpF4mA/xy PPDCA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEqChZjxU8lqj02ht3p8wYuzmrYEvg5d87zO8AQZiDFbK4iKtuSdYsxN+c2ZoTQsd zQCY87aTj+m9TAVwdwmdnHKtWVRXEr5fYrOyK1UHI3A+jsshF0fdPBAQVaID4Wjh7c0VEi dMMoETa/m0t33XZQU/UKXL4JlwUkkgq4rn6LMFa3pgr2l0Q8gT02Cw2LTLC9qVW3zH4Meh qM5Q9QulYPPXeKlcMtfCTbhFhqEmjDBhZKmzHTdwgvJ4UJe4phkHLyfIFsN1xpYIRuVU2w 753LB2HCYBO3y/gIk+V2c2W66P34Qvl7+NN+EKfE2PSluCBz1GoS4+OMA2hGu38Uv91MGe 4sTC5n92K1OBThe1XxwGhr1GgpYY+WoIkK4tbDPzVa+TSQLKZWYQ/zq5apU+YI+o3tVECK P+PdOrfvJf3ku5k0iSYdGKAgoT4PHJG1I6KcwVNl2YHf0vo513CyZEEe0x+JfEndboES2a 9udf7lXx7WK+fRrtAl3v2jY+cP3lSrTFT68GhvpMYPTHYRgt+TTydDiZL6Y8S7oXvastST Z1Yt6tB4jTmzKMQ9DjoHIquzS15QUsewYVi/NQgWUjfAZe92uTmiDHYNszwvxNeu/uUbeH Ivg6hBCV9ysXNP206u2eN2kMLiEqfSltcj4+UB7B5xMbfeDtHdW3nlbvWbOg X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Wed, 12 Aug 2026 09:23:26 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v3 07/18] selftests/mm: add folio-order check for address ranges Date: Wed, 12 Aug 2026 14:22:53 +0100 Message-ID: <20260812132304.199287-8-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260812132304.199287-1-kirill@shutemov.name> References: <20260812132304.199287-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" An mTHP collapse test has to ask whether a range came out as folios of one particular order, placed where a collapse would place them. Nothing answers that today: is_backed_by_folio() classifies the folio behind a single page, and check_huge_anon() reads smaps AnonHugePages, which only accounts PMD mappings. Add is_range_backed_by_folio_orders(). For every order-aligned window of the range it requires a present head PFN at its natural alignment and a contiguous PFN run across the window. A window stitched together from several folios, or one holding a folio shifted off its natural position, fails -- which is what makes the helper usable for "this window collapsed and that one did not". Assisted-by: Claude-Code:claude-opus-5 Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/vm_util.c | 42 ++++++++++++++++++++++++++++ tools/testing/selftests/mm/vm_util.h | 2 ++ 2 files changed, 44 insertions(+) diff --git a/tools/testing/selftests/mm/vm_util.c b/tools/testing/selftests= /mm/vm_util.c index 5db1a7774f49..c9bd6c92fa41 100644 --- a/tools/testing/selftests/mm/vm_util.c +++ b/tools/testing/selftests/mm/vm_util.c @@ -556,6 +556,48 @@ bool is_backed_by_folio(char *vaddr, int order, int pa= gemap_fd, return false; } =20 +/* + * Check whether every order-@order window of [start, len) maps exactly one + * folio of that order, head to tail. The address range must be naturally + * aligned, each window's PFN run must be contiguous, and a window's first + * PFN must be the folio head. + * + * This is the check "did this range collapse into order-@order folios": a + * window assembled from parts of several folios, or mapping a folio shift= ed + * from its natural position, fails. + */ +bool is_range_backed_by_folio_orders(char *start, size_t len, int order, + int pagemap_fd, int kpageflags_fd) +{ + const unsigned long nr_pages =3D 1UL << order; + const size_t window =3D nr_pages * psize(); + char *vaddr; + + if ((uintptr_t)start % window || len % window) + return false; + + for (vaddr =3D start; vaddr < start + len; vaddr +=3D window) { + unsigned long pfn =3D pagemap_get_pfn(pagemap_fd, vaddr); + unsigned long i; + + /* Not present, or not mapping the folio head. */ + if (pfn =3D=3D -1UL || pfn % nr_pages) + return false; + + for (i =3D 1; i < nr_pages; i++) { + if (pagemap_get_pfn(pagemap_fd, vaddr + i * psize()) !=3D + pfn + i) + return false; + } + + if (!is_backed_by_folio(vaddr, order, pagemap_fd, + kpageflags_fd)) + return false; + } + + return true; +} + /* If `ioctls' non-NULL, the allowed ioctls will be returned into the var = */ int uffd_register_with_ioctls(int uffd, void *addr, uint64_t len, bool miss, bool wp, bool minor, uint64_t *ioctls) diff --git a/tools/testing/selftests/mm/vm_util.h b/tools/testing/selftests= /mm/vm_util.h index 56a28ce7d029..39dfb18dc10c 100644 --- a/tools/testing/selftests/mm/vm_util.h +++ b/tools/testing/selftests/mm/vm_util.h @@ -99,6 +99,8 @@ int gather_folio_orders(char *vaddr_start, size_t len, int pagemap_fd, int kpageflags_fd, int orders[], int nr_orders); bool is_backed_by_folio(char *vaddr, int order, int pagemap_fd, int kpageflags_fd); +bool is_range_backed_by_folio_orders(char *start, size_t len, int order, + int pagemap_fd, int kpageflags_fd); =20 int uffd_register(int uffd, void *addr, uint64_t len, bool miss, bool wp, bool minor); --=20 2.54.0 From nobody Tue Sep 29 04:44:47 2026 Received: from fhigh-b1-smtp.messagingengine.com (fhigh-b1-smtp.messagingengine.com [202.12.124.152]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2B6E34582D9; Wed, 12 Aug 2026 13:23:30 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.152 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541012; cv=none; b=JDmwotTsofXRCOLAzZUcPP4bYXPnumprKLzJ8Cbg/N0E9VyEJ4ZgzndMymUYII6fZx+USe7SOpA8TWzhuNN9io8mefBixcnRachOUR1KK2K/IsS3I/BlCDbqMTeGXnZb1Q0a0uq+g1yevDGdXCb4IgmqQ64lDnYd0ZOb0G4O9zc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541012; c=relaxed/simple; bh=abfrz0+esfYJAG2Kjxv04Ai/3vnXF6siE7sj5P1GiwQ=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=eOjU1SLUjctKypkoq13HY+a1Cxupx16vu9/+rUno9gZ2vCLbz3meHpzsYpsHVgJTvCqe3/FAQGL9jid9HkIo65pEaS2tYQjFy0smgB+miAANXTTARUZkb73X4yAkaVAe2Lim0B6fVx+/lftCo+OXLmFVok00Gu9DSocOQvQ+04o= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=Y2aRuXQh; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=XT0vHhyN; arc=none smtp.client-ip=202.12.124.152 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="Y2aRuXQh"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="XT0vHhyN" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfhigh.stl.internal (Postfix) with ESMTP id CF1877A015F; Wed, 12 Aug 2026 09:23:29 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-06.internal (MEProxy); Wed, 12 Aug 2026 09:23:30 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786541009; x= 1786627409; bh=eV+c3895JCTdH1VJjEuBsFvGHhdtHUWemWq2+damYtQ=; b=Y 2aRuXQhyHjkKd1QYCLfq3PlNdfxXF2j3x3C8qlHoqHlX/rkNfbzNfTx8gLlmCsjt frBXJK9mZscWUaX0CbUwekyFNa/xp+pMpvqKpM003vGeB9E/uxSPt422TUxeI1pu H05cK34WJplIq9l1yl1kXXxgUycd6zmwLcIwyvV0qqJWNaZobfrPho7s4cvumPYQ dXe9/ClGBhEqZBNHqseCcW9nx28MWce4yk/lzHgVdAhsQhKa+JTwz7WCMFVb0G+u 3qA07B/9AYK4o1IHWPzMuGo5RFIapbJ0ueNbHNnC//Pwrbcw6Oe2PWK1VRDdxtc4 Yl+x1SAA4Gm+qu7g4oLAg== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786541009; x=1786627409; bh=e V+c3895JCTdH1VJjEuBsFvGHhdtHUWemWq2+damYtQ=; b=XT0vHhyNr/jesPv6Q VX0YnJt2YxvhOqeI6fyPGxDTW52oUNd7hkLqd+129F8HU5BzxDi7NsA6V+V2HcCV 916NTzORmL6OHOADG7SjRJgC48SxTdirNzoXLLlftorN2bvFzCq78Cp2dAp1sJqA /U/siWtzoissMESH0rcqHdf6OnhQqX2swlOKUigvB518nUQZ0tevzdZUABJZrqmW +/tSfCSnIKVF063oOON3EYKXZn8XN/e6gTCoes4SldVYRDM4H4I/WZzUxcKFy9lP LySje7yMXvDIE0cPP6YhO9BaU7If0mT5Lk7OhrJwPyM5RZhlmD0jt/brQQIyY3kM /vn1Q== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFzqtrNoSTXX2MRzKyN6p8XXqdjDSX08KYyJpXp82eYC1ae2GegaTOhbDzreYT/K0 OAsc1zsfjWszwuQXvOMxV3FM9x50+HqAsA48u5Kq52QtHo+TJJPCEG+EOeKn5Il8AepNuY amkSO7LHncjzMJKy/SPUs50UUrTlzGP0ayf7b+h/q5NreZMQSrrfXSzcd6tsdBCg4Hhmbc 7COQKDHpqWJD3HpoJ7/Bw0J3tcRV9tZIkmRGQ4nZ/O76iTBc+eN6GYtayu9aLLTJal0DNN hI5LqlugC7KhSFwlZIIFqKKnJB7XY7vZnEBEdRJ8Mvpq9sE5OflKN5J+H3ujbniJsnhlPB GEVJb5AijGvAC/NfopsI7czdo2wONzWrV/c4G9br9asrjG5pXZXAqMU7dITwfFGwUqL5H7 Tt9EK0KvS5pS8oQqKEsQj+/gU1mEhUmzFys1sm1kG0Hx3rtMQdzqXrLd0yamhSyBWOgH8B BMA1jaxo1CXtHaGldha0M2yx5byWDwdQM/9pPpxbnBhC0d/2tmPCiF79DZ+mln+zzKedZF z3DHrwJ/hkLHGyYTJyVcRgmqKYrIVYma/O5X8snp0s51EKg0vzdvkGdOOjoLPDiZmw0fXs 0+PIF83+Fu7S1EwOCgEcfFcvW1c+ZFSv6ULiKU+o4EmBlD/rWgXOnST6UGJw X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Wed, 12 Aug 2026 09:23:28 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v3 08/18] selftests/mm: add folio-order detection self-check Date: Wed, 12 Aug 2026 14:22:54 +0100 Message-ID: <20260812132304.199287-9-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260812132304.199287-1-kirill@shutemov.name> References: <20260812132304.199287-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The upcoming khugepaged mTHP tests detect collapse results with the vm_util folio-order helpers instead of smaps AnonHugePages, which only sees PMD mappings. Before any collapse test trusts those helpers, make sure they agree with the kernel about what backs a mapping. For every anon THP order the kernel supports, fault memory in with only that order enabled and require the helpers to classify the backing as exactly that order: not a neighbouring order, and 4K-backed memory as order 0. Runs in the thp category of run_vmtests.sh. Verified on x86-64 4K (orders 0, 2-9) and arm64 64K (orders 0, 2-13). While here: ALIGN() moves into vm_util.h, where the next test that needs a round-up can find it, and hmm-tests.c drops its own copy. Assisted-by: Claude-Code:claude-opus-5 Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/Makefile | 1 + .../testing/selftests/mm/folio_order_check.c | 137 ++++++++++++++++++ tools/testing/selftests/mm/hmm-tests.c | 1 - tools/testing/selftests/mm/migration.c | 1 - tools/testing/selftests/mm/run_vmtests.sh | 2 + tools/testing/selftests/mm/vm_util.h | 2 + 6 files changed, 142 insertions(+), 2 deletions(-) create mode 100644 tools/testing/selftests/mm/folio_order_check.c diff --git a/tools/testing/selftests/mm/Makefile b/tools/testing/selftests/= mm/Makefile index 2d5366196e30..2093fcf6e915 100644 --- a/tools/testing/selftests/mm/Makefile +++ b/tools/testing/selftests/mm/Makefile @@ -104,6 +104,7 @@ TEST_GEN_FILES +=3D guard-regions TEST_GEN_FILES +=3D merge TEST_GEN_FILES +=3D rmap TEST_GEN_FILES +=3D folio_split_race_test +TEST_GEN_FILES +=3D folio_order_check =20 ifneq ($(ARCH),arm64) TEST_GEN_FILES +=3D soft-dirty diff --git a/tools/testing/selftests/mm/folio_order_check.c b/tools/testing= /selftests/mm/folio_order_check.c new file mode 100644 index 000000000000..93030a42c3cc --- /dev/null +++ b/tools/testing/selftests/mm/folio_order_check.c @@ -0,0 +1,137 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Self-check for the vm_util folio-order detection helpers, + * is_backed_by_folio() and is_range_backed_by_folio_orders(). + * + * For every anon THP order the kernel supports, fault memory in with only + * that order enabled and verify the helpers report exactly that order: + * not a neighbouring order, and plain 4K memory as order 0. The helpers + * are what the khugepaged mTHP tests use to detect collapse results, so + * they must agree with the kernel's own idea of the backing before any + * collapse test relies on them. + */ +#define _GNU_SOURCE +#include +#include +#include +#include +#include + +#include "kselftest.h" +#include "vm_util.h" +#include "hugepage_settings.h" + +static int pagemap_fd; +static int kpageflags_fd; + +/* mmap an anon VMA of exactly @size bytes at a @size-aligned address. */ +static char *alloc_aligned(size_t size) +{ + size_t len =3D size * 2; + uintptr_t aligned; + char *p; + + p =3D mmap(NULL, len, PROT_READ | PROT_WRITE, + MAP_ANONYMOUS | MAP_PRIVATE, -1, 0); + if (p =3D=3D MAP_FAILED) + ksft_exit_fail_perror("mmap()"); + + aligned =3D ALIGN((uintptr_t)p, size); + if (aligned !=3D (uintptr_t)p) + munmap(p, aligned - (uintptr_t)p); + if (aligned + size !=3D (uintptr_t)p + len) + munmap((char *)aligned + size, + (uintptr_t)p + len - aligned - size); + + return (char *)aligned; +} + +/* + * Enable only @order (order 0: nothing), fault one aligned window in and + * check the helpers see exactly @order. + */ +static void check_order(int order) +{ + struct thp_settings settings =3D *thp_current_settings(); + size_t size =3D psize() << order; + bool ok =3D true; + char *p; + int i; + + for (i =3D 0; i < NR_ORDERS; i++) + settings.hugepages[i].enabled =3D THP_NEVER; + if (order) + settings.hugepages[order].enabled =3D THP_ALWAYS; + thp_push_settings(&settings); + + p =3D alloc_aligned(size); + *p =3D 1; + + if (!is_range_backed_by_folio_orders(p, size, order, + pagemap_fd, kpageflags_fd)) { + ksft_print_msg("order %d not detected after fault\n", order); + ok =3D false; + } + + /* A lower order must be rejected: the folio is larger. */ + if (order && is_range_backed_by_folio_orders(p, size, order - 1, + pagemap_fd, + kpageflags_fd)) { + ksft_print_msg("order %d also reported as order %d\n", + order, order - 1); + ok =3D false; + } + + /* Order 0 pages must not look like any large folio, and vice versa. */ + if (order && is_range_backed_by_folio_orders(p, size, 0, + pagemap_fd, + kpageflags_fd)) { + ksft_print_msg("order %d also reported as order 0\n", order); + ok =3D false; + } + + munmap(p, size); + thp_pop_settings(); + + ksft_test_result(ok, "order %d classified\n", order); +} + +int main(void) +{ + struct thp_settings settings; + unsigned long orders; + int order; + + ksft_print_header(); + + if (!thp_available()) + ksft_exit_skip("Transparent Hugepages not available\n"); + + pagemap_fd =3D open("/proc/self/pagemap", O_RDONLY); + if (pagemap_fd < 0) + ksft_exit_fail_perror("open(/proc/self/pagemap)"); + kpageflags_fd =3D open("/proc/kpageflags", O_RDONLY); + if (kpageflags_fd < 0) + ksft_exit_skip("open(\"/proc/kpageflags\") requires root\n"); + + orders =3D thp_supported_orders(); + if (!orders) + ksft_exit_skip("No supported THP orders\n"); + + ksft_set_plan(__builtin_popcountl(orders) + 1); + + thp_save_settings(); + thp_read_settings(&settings); + /* Base of the settings stack; the bottom entry is never popped. */ + thp_push_settings(&settings); + + check_order(0); + for (order =3D 1; order < NR_ORDERS; order++) { + if (!(orders & (1UL << order))) + continue; + check_order(order); + } + + + ksft_finished(); +} diff --git a/tools/testing/selftests/mm/hmm-tests.c b/tools/testing/selftes= ts/mm/hmm-tests.c index e2642eca0d02..df426f9218e7 100644 --- a/tools/testing/selftests/mm/hmm-tests.c +++ b/tools/testing/selftests/mm/hmm-tests.c @@ -65,7 +65,6 @@ enum { #define HMM_PATH_MAX 64 #define NTIMES 10 =20 -#define ALIGN(x, a) (((x) + (a - 1)) & (~((a) - 1))) /* Just the flags we need, copied from mm.h: */ =20 #ifndef FOLL_WRITE diff --git a/tools/testing/selftests/mm/migration.c b/tools/testing/selftes= ts/mm/migration.c index f19d53c69576..fd35f8a7b5b8 100644 --- a/tools/testing/selftests/mm/migration.c +++ b/tools/testing/selftests/mm/migration.c @@ -20,7 +20,6 @@ =20 #define TWOMEG (2<<20) #define RUNTIME (20) -#define ALIGN(x, a) (((x) + (a - 1)) & (~((a) - 1))) =20 HUGETLB_SETUP_DEFAULT_PAGES(1) =20 diff --git a/tools/testing/selftests/mm/run_vmtests.sh b/tools/testing/self= tests/mm/run_vmtests.sh index d09f9f6a384e..2652a7920b80 100755 --- a/tools/testing/selftests/mm/run_vmtests.sh +++ b/tools/testing/selftests/mm/run_vmtests.sh @@ -402,6 +402,8 @@ CATEGORY=3D"pfnmap" run_test ./pfnmap # COW tests CATEGORY=3D"cow" run_test ./cow =20 +CATEGORY=3D"thp" run_test ./folio_order_check + CATEGORY=3D"thp" run_test ./khugepaged =20 CATEGORY=3D"thp" run_test ./khugepaged -s 2 diff --git a/tools/testing/selftests/mm/vm_util.h b/tools/testing/selftests= /mm/vm_util.h index 39dfb18dc10c..ce05bce4670d 100644 --- a/tools/testing/selftests/mm/vm_util.h +++ b/tools/testing/selftests/mm/vm_util.h @@ -10,6 +10,8 @@ #include =20 #define BIT_ULL(nr) (1ULL << (nr)) +#define ALIGN(x, a) (((x) + (a) - 1) & ~((a) - 1)) + #define PM_SOFT_DIRTY BIT_ULL(55) #define PM_MMAP_EXCLUSIVE BIT_ULL(56) #define PM_UFFD_WP BIT_ULL(57) --=20 2.54.0 From nobody Tue Sep 29 04:44:47 2026 Received: from fhigh-b1-smtp.messagingengine.com (fhigh-b1-smtp.messagingengine.com [202.12.124.152]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B214E451992; Wed, 12 Aug 2026 13:23:33 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.152 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541015; cv=none; b=ddO2f316ZvohriN2liUskCZdJ3F5WFAhpSeB8KDLhJ2t3BP286ZVBl7Gs9BlI9JOkV7IJzsAN6KZ3Sf/khpZleu9+A4YQumEnR3cwi30q3TVPg5ugHJCI21rrB9E+8VB7mYpDIi8yds41uc/HzIaIstPlh027wDYUHrwh5xZtBc= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541015; c=relaxed/simple; bh=NiZXq6PSkV3W5VwahZWuS9VpZvEGsgLGi73zxA+K2DQ=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=en/kboxS/56oc4dcqunAlAK2gf3S2OAUG7ylXPbFahYMNV7XpxbUtgiK2xg6iR4m9nSGKLGYiAMxoyYMBk08xwJKQg1nQT+0A3SwM1pC+NSI/cNzile+48TAh4FWR7J1HYL9xgz741sb5N3S8QB8rBcg6KSgDVW6HnwRFHAEgh4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=d+bIRM8D; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=kvo68HBQ; arc=none smtp.client-ip=202.12.124.152 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="d+bIRM8D"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="kvo68HBQ" Received: from phl-compute-03.internal (phl-compute-03.internal [10.202.2.43]) by mailfhigh.stl.internal (Postfix) with ESMTP id 586F17A00D8; Wed, 12 Aug 2026 09:23:32 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-03.internal (MEProxy); Wed, 12 Aug 2026 09:23:32 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786541012; x= 1786627412; bh=vEcbFcHSc0RIpAv6Xeg8b8khiwkn4nugTeMN0M/IuN0=; b=d +bIRM8DtQgQ5NhEwnn9zLY5xsSRcpLiMaKCWyYH4Sl09LYB1asLwLL0ed7bTbwfB EHXtzvhrVVnLaabR7/z043cxfbXCE+mVZlZktcVCq0kstHW6fCmJqflBE9hzjTI8 C46BPRFOgu2oSav7g5jJZ8NIhDZFnViC6gK/r5XGl5dDXuMkAKih0eeAtlMFapl5 2BjErd8yBqKa38WIOIS0rnQZ3abtkxcsAukhYkhktQML+uTMwQpcVNWtIJOqygxz VO7eKH3FfqjKbIyBxQcrIeioHfg556xw3XDR7f5SE71+i2L/j4FsvoY94mBAawj6 iRl+7GZJVEfKMOCsSUzvw== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786541012; x=1786627412; bh=v EcbFcHSc0RIpAv6Xeg8b8khiwkn4nugTeMN0M/IuN0=; b=kvo68HBQXjgCRBQN7 6heYYkKOcONSCUBx4Ut0vP3ILxs21c5dRodLxl7kck4B02RT98jHxUUdml5AxheJ zbWoCX6u0BHb0qeOiphpPDASK3qoIVJ2S5rkALrM2akvRKH4EovQYtN9YERxs1/3 iCXzFJeB7+oLEeDIgfl8hxm+9kVf6/Rw1XghYE4Tm45HneIsGtZqD5MAgwHKw9Wt 4bFncveQ27pfCoTx3/43QEC5HZVWQdnvaG1PLbK9TLHB4oqRuz2Ra9fEnF2IL93E BcuO6/Ywx4NOFW6CHn0zTa2o1Hy4mvakkygR9nkEF0WyPOg0Ky6/UL7rGMttBE7j bN7Iw== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGK88F9UMnXZ5YHCfZ+3CEyhDPNn2cwpG7gczg1QtUW3+KH6lRiOTrWgwtmC8SwM8 lS9N2AAOkbX57R9oIGEGX/Y22a2ZgCAGYteUW+bA41LZ08JSUTSnxLP4aLeGJ+l2igRpiI z056Op0LI1NxVlka48nXFKSS8a/uwNZK+39708VFH13YHS1sNHrX5zvuU5299hb11Ao5H/ zQrZEfdZM4G7c0rNyXZ30syjB6Y8P8LTAhLTAHwdc8i3T8CZWzvPOusg5fBb0jumhRSeE4 e5clat+a+U2DXBe0wlr1aBcOwBQzkKo/UFV6TzMBSg5j9cF+ay7ufreZHSHrXPVD1XvoqN nGajMc9LFlqtqdyNBq16ebJCwpL9n0hNsyrSX0E74XPEhuZsPSQdOcHhjksiUT7jfzCQN5 iBSxH59wM/CgQqAHSwNJhfxAEsTymBgLjfOffondkWrzKSe/7P1mH45ya+i7fOvTeG2e3e fEy3DOAGUCoQH+fXGrmdf9ulLu/vDnqrnyQ5qo2cAEtLiHOdRb7M+36TZHrFYXvfiHPGsS k7XZ/bEBQVWi5ABJ/mqzgHtU3UbMZj7i4qYjKCY5tyD55E+dJJMBuh32mh6PN016C75DzM RIVwsXzmayy2BQdFdA4CL2hgOMrG2qzf+dNZzV8Y98BEZ36/pwhrKN0i9c1Q X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Wed, 12 Aug 2026 09:23:31 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v3 09/18] selftests/mm: add khugepaged completion barrier helper Date: Wed, 12 Aug 2026 14:22:55 +0100 Message-ID: <20260812132304.199287-10-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260812132304.199287-1-kirill@shutemov.name> References: <20260812132304.199287-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Race and functional tests need to drive khugepaged in step: set up a layout, let one full scan pass over it, check the result. The khugepaged selftest already waits for full_scans to advance by two, but only makes progress if scan_sleep_millisecs happens to be short. Lift it into khugepaged_full_pass() and drive it through sysfs: any store to scan_sleep_millisecs wakes the daemon, so the barrier completes whatever the scan cadence. It wakes once per missing pass -- over-waking queues a straggler pass that overlaps what the caller sets up next. One wake completes one pass only if the whole mm list fits in a scan batch, so callers need a large pages_to_scan. Settings pushes must not start passes either, so thp_write_settings() now writes a khugepaged knob only when its value changes. That helper, thp_update_num(), is exported for tests wanting the same restraint. Assisted-by: Claude-Code:claude-opus-5 Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) --- .../testing/selftests/mm/hugepage_settings.c | 72 ++++++++++++++++--- .../testing/selftests/mm/hugepage_settings.h | 3 + 2 files changed, 66 insertions(+), 9 deletions(-) diff --git a/tools/testing/selftests/mm/hugepage_settings.c b/tools/testing= /selftests/mm/hugepage_settings.c index d7917dce3aba..8afcdf9793bb 100644 --- a/tools/testing/selftests/mm/hugepage_settings.c +++ b/tools/testing/selftests/mm/hugepage_settings.c @@ -183,6 +183,17 @@ void thp_read_settings(struct thp_settings *settings) } } =20 +/* + * Write only on change: any store to a khugepaged sysfs knob wakes the + * daemon, and settings pushes/pops must not start scan passes nobody + * asked for -- khugepaged_full_pass() is the only sanctioned wake. + */ +void thp_update_num(const char *name, unsigned long num) +{ + if (thp_read_num(name) !=3D num) + thp_write_num(name, num); +} + void thp_write_settings(struct thp_settings *settings) { struct khugepaged_settings *khugepaged =3D &settings->khugepaged; @@ -198,15 +209,15 @@ void thp_write_settings(struct thp_settings *settings) shmem_enabled_strings[settings->shmem_enabled]); thp_write_num("use_zero_page", settings->use_zero_page); =20 - thp_write_num("khugepaged/defrag", khugepaged->defrag); - thp_write_num("khugepaged/alloc_sleep_millisecs", - khugepaged->alloc_sleep_millisecs); - thp_write_num("khugepaged/scan_sleep_millisecs", - khugepaged->scan_sleep_millisecs); - thp_write_num("khugepaged/max_ptes_none", khugepaged->max_ptes_none); - thp_write_num("khugepaged/max_ptes_swap", khugepaged->max_ptes_swap); - thp_write_num("khugepaged/max_ptes_shared", khugepaged->max_ptes_shared); - thp_write_num("khugepaged/pages_to_scan", khugepaged->pages_to_scan); + thp_update_num("khugepaged/defrag", khugepaged->defrag); + thp_update_num("khugepaged/alloc_sleep_millisecs", + khugepaged->alloc_sleep_millisecs); + thp_update_num("khugepaged/scan_sleep_millisecs", + khugepaged->scan_sleep_millisecs); + thp_update_num("khugepaged/max_ptes_none", khugepaged->max_ptes_none); + thp_update_num("khugepaged/max_ptes_swap", khugepaged->max_ptes_swap); + thp_update_num("khugepaged/max_ptes_shared", khugepaged->max_ptes_shared); + thp_update_num("khugepaged/pages_to_scan", khugepaged->pages_to_scan); =20 if (dev_queue_read_ahead_path[0]) write_num(dev_queue_read_ahead_path, settings->read_ahead_kb); @@ -230,6 +241,49 @@ void thp_write_settings(struct thp_settings *settings) } } =20 +/* + * Completion barrier for khugepaged: wait until a full scan pass that + * started after this call has finished. full_scans must advance by two; + * a +1 step may complete a pass that examined this mm before the + * caller's setup was in place. + * + * Any store to scan_sleep_millisecs wakes the daemon, so the barrier works + * whatever the configured scan cadence -- but a store can be lost. + * __sleep_millisecs_store() clears khugepaged_sleep_expire and wakes the + * queue; if the daemon is between scans rather than sleeping, it sets + * khugepaged_sleep_expire itself on the way into khugepaged_wait_work() a= nd + * then sleeps for the full interval, having never seen the store. So keep + * storing until the pass lands; a store while the daemon is awake costs + * nothing and does not queue an extra pass. + * + * One wake completes one full pass only if the whole mm list fits in + * one scan batch, so callers must pair this with a large + * pages_to_scan. + */ +bool khugepaged_full_pass(unsigned int timeout_s) +{ + unsigned long deadline_ms =3D timeout_s * 1000UL; + unsigned long sleep_ms =3D + thp_read_num("khugepaged/scan_sleep_millisecs"); + unsigned long elapsed_ms =3D 0; + int pass; + + for (pass =3D 0; pass < 2; pass++) { + unsigned long target =3D + thp_read_num("khugepaged/full_scans") + 1; + + while (thp_read_num("khugepaged/full_scans") < target) { + if (elapsed_ms >=3D deadline_ms) + return false; + thp_write_num("khugepaged/scan_sleep_millisecs", + sleep_ms); + usleep(10 * 1000); + elapsed_ms +=3D 10; + } + } + return true; +} + struct thp_settings *thp_current_settings(void) { if (!settings_index) { diff --git a/tools/testing/selftests/mm/hugepage_settings.h b/tools/testing= /selftests/mm/hugepage_settings.h index 726c73c43c05..ba7d38370d43 100644 --- a/tools/testing/selftests/mm/hugepage_settings.h +++ b/tools/testing/selftests/mm/hugepage_settings.h @@ -70,6 +70,7 @@ int thp_read_string(const char *name, const char * const = strings[]); void thp_write_string(const char *name, const char *val); unsigned long thp_read_num(const char *name); void thp_write_num(const char *name, unsigned long num); +void thp_update_num(const char *name, unsigned long num); =20 void thp_write_settings(struct thp_settings *settings); void thp_read_settings(struct thp_settings *settings); @@ -83,6 +84,8 @@ static inline void thp_save_settings(void) hugepage_save_settings(/* thp =3D */ true, /* hugetlb =3D */ false); } =20 +bool khugepaged_full_pass(unsigned int timeout_s); + void thp_set_read_ahead_path(char *path); unsigned long thp_supported_orders(void); unsigned long thp_shmem_supported_orders(void); --=20 2.54.0 From nobody Tue Sep 29 04:44:47 2026 Received: from fout-b7-smtp.messagingengine.com (fout-b7-smtp.messagingengine.com [202.12.124.150]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7141A451997; Wed, 12 Aug 2026 13:23:36 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.150 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541018; cv=none; b=AYUkbR357PhSms+A38EEzRUfFkWNwvkmydXUiwfn2W0XP1YLvWA3JhkiyVHsNg2iQlB16GFLIiR1DbwgxTBl3k4eu2irwbNOcXUR8ycHml+UTjTdE58gPC01qPaYhGyAM4kzrp0WgpU/afx50/V51PCtT6Xeli9maGBm0nY/VRY= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541018; c=relaxed/simple; bh=lskO5+k5+lVn83+SxQNXVwsxUEiml217LU9Qer+YkN4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=lkfxQQow1y3GiZEejYjJjJhnRb4V8cNJKDFi/ycXvxXow4h8OhxyRlbMhuvJFwLhE/9sR8Oj7+nsboYrSciRLfn4Ts3aiPbmDeu6hd/WJR+j06//A12X48Ps2HbhJmsc9A73yvL6l9adC8x7Wt1w9WYvt1J6GOg5nNY9IBohfS8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=Lw9IHymN; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=Th715Xjc; arc=none smtp.client-ip=202.12.124.150 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="Lw9IHymN"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="Th715Xjc" Received: from phl-compute-11.internal (phl-compute-11.internal [10.202.2.51]) by mailfout.stl.internal (Postfix) with ESMTP id 2811E1D00165; Wed, 12 Aug 2026 09:23:35 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-11.internal (MEProxy); Wed, 12 Aug 2026 09:23:35 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786541015; x= 1786627415; bh=oeZs2UseJcgjssEuDoTxdhaLLmHeagTkQsJsYI528D8=; b=L w9IHymNYFUdCX1n6pCnEqiO+w5U8ENtkxk1NsACctkKVOn8ROw2eWYZjIw9TOVwG +dfYW0bzbX4DOXEusvKGcvRlhleMhuL+J+EOKTfHZeR7nAxTMqrwJEUh3iHz7Cqr vzB02XkWuy0tWTjw/wsF2H5Szvv8gOfle8oALZICrMVXO0DwACv521Wmh07dYItv D0lxxyl6TC2MlMaTMvCKBkX7OKSLYFVT2V1lGpuYx0rv/6Md5fj9l30eWzYqSAxx FqGeeS/u0o61GEPdCTeF/+f4Tvu7K/kxDH7rDsGRmTbbVjslCB6NC80yo1ceXogC omURnqsB7DfUeIfI9S/bQ== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786541015; x=1786627415; bh=o eZs2UseJcgjssEuDoTxdhaLLmHeagTkQsJsYI528D8=; b=Th715XjcVgT36pmDR c/ZVhfKCOWpXTnv/BnxSATstrdYwEFdAJq1yGEnIFot48XQeKEO/sCHnhk010ml1 2D+d+6mZTqgwL4c3OLXV0vBSDz71OSBYNpfHeCYClROrksn2zImTokxuJ8EZdeg7 t49iuMGFYzNGS6bnQD8qtId2+bMGlrQ+slkRvP5SGCl7Lvdaj7gp/HGpAAglIyc0 Kxw6QUd1Es389lIu+9ygAqguUbPCuqdmkdivxU6RIgBkeuSoaHAcmminYicYo/ct 7pGVl1sy6fdH0hbBSp2NiWPbZqrTho2RN3h1ORytQFaozZUPZzRmzEWcnn7k2QnM LAqNg== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGK88F9UMnXZ5YHCfZ+3CEyhDPNn2cwpG7gczg1QtUW3+KH6lRiOTrWgwtmC8SwM8 lS9N2AAOkbX57R9oIGEGX/Y22a2ZgCAGYteUW+bA41LZ08JSUTSnxLP4aLeGJ+l2igRpiI z056Op0LI1NxVlka48nXFKSS8a/uwNZK+39708VFH13YHS1sNHrX5zvuU5299hb11Ao5H/ zQrZEfdZM4G7c0rNyXZ30syjB6Y8P8LTAhLTAHwdc8i3T8CZWzvPOusg5fBb0jumhRSeE4 e5clat+a+U2DXBe0wlr1aBcOwBQzkKo/UFV6TzMBSg5j9cF+ay7ufreZHSHrXPVD1XvofL E3n9Xy4pBFhY5FAp8tzoG+gD4sN0Aj2N6U89gzjxC+Hl3rn8uf6IbEup57zqPScVTNeItr VxEvT8CC31vyw7k0COC3c6WQDWNraJU38zlqU+WWaMZTjyYgjY7irO9mnzqjB3O39Cpa3X pgZb7Ph8QvjQEphfAJrapAxRDrIyLYuqlIBDlodTDvE7ZbmhIcdkELCsoBCuskCn/fKo/b xo+96fAwhS/eVvb3npy+RFRhX8jcWRVWhKXiQcvf92Wtn3CFIe0imOHRyiAbaV8tDUZ3aa piBPrd8f7E/y1QqP9LO1KYqgb78c+3sRd9aYJdq/llNs9MNerhnjsOW1xKLQ X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Wed, 12 Aug 2026 09:23:34 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v3 10/18] selftests/mm: add order-parameterized khugepaged collapse cases Date: Wed, 12 Aug 2026 14:22:56 +0100 Message-ID: <20260812132304.199287-11-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260812132304.199287-1-kirill@shutemov.name> References: <20260812132304.199287-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The mthp_khugepaged context runs the generic cases at a sub-PMD order, which answers how many folios of that order a range ends up with. It cannot say which window they are in, so "the populated window collapsed and its neighbour did not" and "one window collapsed twice" look alike. Add four cases that check each aligned window on its own, with the folio-order helpers in vm_util: - collapse_order_single_window: only the populated window collapses; - collapse_order_partial_window: the default max_ptes_none lets a window with one present PTE collapse; - collapse_order_max_ptes_none: with max_ptes_none=3D0 a full window collapses and one missing a page does not; - collapse_order_mixed_sources: sources that are already large folios of a smaller order collapse to the target. Each case faults its region before MADV_HUGEPAGE with only the target order enabled, so the sources are order 0 and the result can only come from khugepaged. They wait for a full pass rather than for the result to appear: without a completed pass, "not collapsed" and "not scanned yet" are the same thing. Assisted-by: Claude-Code:claude-opus-5 Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/khugepaged.c | 230 ++++++++++++++++++++++++ 1 file changed, 230 insertions(+) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index 8458cd2ff0df..7889ea1cd322 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -31,6 +31,8 @@ static unsigned long page_size; static int hpage_pmd_nr; static int anon_order; static int collapse_order; +static int pagemap_fd =3D -1; +static int kpageflags_fd =3D -1; =20 #define PID_SMAPS "/proc/self/smaps" #define TEST_FILE "collapse_test_file" @@ -1250,6 +1252,216 @@ static void madvise_retracted_page_tables(struct co= llapse_context *c, ksft_test_result_report(exit_status, "%s\n", __func__); } =20 +/* Smallest order khugepaged will consider for mTHP collapse. */ +#define MIN_MTHP_ORDER 2 + +/* + * Order-parameterized collapse cases for the mthp_khugepaged context. Wh= at + * they add over the generic cases run under that context is per-window + * detection: which aligned window collapsed, and which of its neighbours = did + * not. check_huge() answers how many folios of the order the range holds, + * which cannot tell one window from another. + * + * The region is faulted before MADV_HUGEPAGE, and the target order is only + * enabled for madvise, so the sources are always order 0 and the collapse + * product can only have come from khugepaged. + */ +static size_t mthp_window_size(void) +{ + return page_size << collapse_order; +} + +static void mthp_push_target_order(void) +{ + struct thp_settings settings =3D *thp_current_settings(); + int i; + + /* + * The target order, for madvise only, and nothing else enabled: the + * cases fault their region before MADV_HUGEPAGE, so the sources are + * order 0 whatever -s asked the fault path for. That matters for the + * cases built around a hole -- a large source folio would fill it in + * and the window would collapse after all. + * collapse_order_mixed_sources enables the source order it wants on + * top of this. + */ + settings.thp_enabled =3D THP_NEVER; + for (i =3D 0; i < NR_ORDERS; i++) + settings.hugepages[i].enabled =3D THP_NEVER; + settings.hugepages[collapse_order].enabled =3D THP_MADVISE; + thp_push_settings(&settings); +} + +static bool window_collapsed(void *p, size_t len) +{ + return is_range_backed_by_folio_orders(p, len, collapse_order, + pagemap_fd, kpageflags_fd); +} + +/* No aligned window in [p, p + len) is backed at the target order. */ +static bool window_not_collapsed(void *p, size_t len) +{ + size_t window =3D mthp_window_size(); + char *addr =3D p; + + for (; len >=3D window; addr +=3D window, len -=3D window) { + if (window_collapsed(addr, window)) + return false; + } + return true; +} + +static bool khugepaged_wait_full_pass(void) +{ + /* Wait up to 30 seconds for the pass to complete. */ + return khugepaged_full_pass(30); +} + +static void collapse_order_single_window(struct collapse_context *c, + struct mem_ops *ops) +{ + size_t window =3D mthp_window_size(); + void *p; + + mthp_push_target_order(); + + p =3D ops->setup_area(1); + ops->fault(p, window, 2 * window); + if (!window_not_collapsed(p, hpage_pmd_size)) + ksft_exit_fail_msg("Unexpected large folio after fault\n"); + + madvise(p, hpage_pmd_size, MADV_HUGEPAGE); + ksft_print_msg("Collapse one fully populated window..."); + if (!khugepaged_wait_full_pass()) + fail("Timeout"); + else if (window_collapsed(p + window, window) && + window_not_collapsed(p, window) && + window_not_collapsed(p + 2 * window, + hpage_pmd_size - 2 * window)) + success("OK"); + else + fail("Fail"); + + validate_memory(p, window, 2 * window); + ops->cleanup_area(p, hpage_pmd_size); + thp_pop_settings(); + ksft_test_result_report(exit_status, "%s\n", __func__); +} + +static void collapse_order_partial_window(struct collapse_context *c, + struct mem_ops *ops) +{ + void *p; + + mthp_push_target_order(); + + p =3D ops->setup_area(1); + ops->fault(p, 0, page_size); + if (!window_not_collapsed(p, hpage_pmd_size)) + ksft_exit_fail_msg("Unexpected large folio after fault\n"); + + madvise(p, hpage_pmd_size, MADV_HUGEPAGE); + ksft_print_msg("Collapse window with single PTE entry present..."); + if (!khugepaged_wait_full_pass()) + fail("Timeout"); + else if (window_collapsed(p, mthp_window_size())) + success("OK"); + else + fail("Fail"); + + validate_memory(p, 0, page_size); + ops->cleanup_area(p, hpage_pmd_size); + thp_pop_settings(); + ksft_test_result_report(exit_status, "%s\n", __func__); +} + +static void collapse_order_max_ptes_none(struct collapse_context *c, + struct mem_ops *ops) +{ + struct thp_settings settings; + size_t window =3D mthp_window_size(); + void *p; + + mthp_push_target_order(); + settings =3D *thp_current_settings(); + settings.khugepaged.max_ptes_none =3D 0; + thp_push_settings(&settings); + + p =3D ops->setup_area(1); + ops->fault(p, 0, 2 * window - page_size); + if (!window_not_collapsed(p, hpage_pmd_size)) + ksft_exit_fail_msg("Unexpected large folio after fault\n"); + + madvise(p, hpage_pmd_size, MADV_HUGEPAGE); + ksft_print_msg("Collapse full window, not the one missing a page..."); + if (!khugepaged_wait_full_pass()) + fail("Timeout"); + else if (window_collapsed(p, window) && + window_not_collapsed(p + window, window)) + success("OK"); + else + fail("Fail"); + + validate_memory(p, 0, 2 * window - page_size); + ops->cleanup_area(p, hpage_pmd_size); + thp_pop_settings(); + thp_pop_settings(); + ksft_test_result_report(exit_status, "%s\n", __func__); +} + +static void collapse_order_mixed_sources(struct collapse_context *c, + struct mem_ops *ops) +{ + struct thp_settings settings; + void *p; + + if (collapse_order <=3D MIN_MTHP_ORDER) { + ksft_test_result_skip("%s: no source order below target\n", + __func__); + return; + } + + mthp_push_target_order(); + + /* Fault the whole region as order-MIN_MTHP_ORDER folios. */ + settings =3D *thp_current_settings(); + settings.hugepages[MIN_MTHP_ORDER].enabled =3D THP_ALWAYS; + thp_push_settings(&settings); + p =3D ops->setup_area(1); + ops->fault(p, 0, hpage_pmd_size); + thp_pop_settings(); + + /* + * The order is enabled, but the allocator can still fall back under + * fragmentation. That leaves nothing to collapse from, which is the + * machine's answer rather than a reason to end the run. + */ + if (!is_range_backed_by_folio_orders(p, hpage_pmd_size, MIN_MTHP_ORDER, + pagemap_fd, kpageflags_fd)) { + ksft_print_msg("No order-%d sources to collapse...", + MIN_MTHP_ORDER); + skip("Skip"); + ops->cleanup_area(p, hpage_pmd_size); + thp_pop_settings(); + ksft_test_result_report(exit_status, "%s\n", __func__); + return; + } + + madvise(p, hpage_pmd_size, MADV_HUGEPAGE); + ksft_print_msg("Collapse region backed by smaller large folios..."); + if (!khugepaged_wait_full_pass()) + fail("Timeout"); + else if (window_collapsed(p, hpage_pmd_size)) + success("OK"); + else + fail("Fail"); + + validate_memory(p, 0, hpage_pmd_size); + ops->cleanup_area(p, hpage_pmd_size); + thp_pop_settings(); + ksft_test_result_report(exit_status, "%s\n", __func__); +} + static void usage(void) { fprintf(stderr, "\nUsage: ./khugepaged [OPTIONS] [dir]\n\n"); @@ -1418,6 +1630,20 @@ int main(int argc, char **argv) =20 parse_test_type(argc, argv); =20 + if (mthp_khugepaged_context && + !(thp_supported_orders() & (1UL << collapse_order))) + ksft_exit_skip("Order %d is not a supported anon THP order\n", + collapse_order); + + if (mthp_khugepaged_context) { + pagemap_fd =3D open("/proc/self/pagemap", O_RDONLY); + if (pagemap_fd < 0) + ksft_exit_fail_perror("open(/proc/self/pagemap)"); + kpageflags_fd =3D open("/proc/kpageflags", O_RDONLY); + if (kpageflags_fd < 0) + ksft_exit_fail_perror("open(/proc/kpageflags)"); + } + setbuf(stdout, NULL); =20 /* @@ -1480,6 +1706,10 @@ int main(int argc, char **argv) TEST(collapse_empty, madvise_context, anon_ops); =20 TEST(collapse_single_mthp, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_single_window, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_partial_window, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_max_ptes_none, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_mixed_sources, mthp_khugepaged_context, anon_ops); =20 TEST(collapse_single_pte_entry, khugepaged_context, anon_ops); TEST(collapse_single_pte_entry, khugepaged_context, read_only_file_ops); --=20 2.54.0 From nobody Tue Sep 29 04:44:47 2026 Received: from fhigh-b1-smtp.messagingengine.com (fhigh-b1-smtp.messagingengine.com [202.12.124.152]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CEAD845A299; Wed, 12 Aug 2026 13:23:38 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.152 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541020; cv=none; b=pA15HrM8EQ5gEx1vTClFw34UXuqFeAPpXBlLE/vmFuuO+3mTeKTqLVqU5RQXW4v56eCrYwbjFWnjFGBfRrrWkAcNNtRcybEo/jNUTVLwqWq2Nc3OQmRwrNWydTNaTScq+U8r1sqa/IJEkMhkTw1uc2fF0djeRrOtG2gLjoTaw18= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541020; c=relaxed/simple; bh=Be7S6KcCLp+9jOhsTbJF/KzYvfOV6S6UuaSwquGZ03I=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=HH8XPwctOkyRQ3R8n0TSgrPsklDD8lVzBP5lMkR9+F6Ulkwysu2uNLsvSBqKfRsW1gKyhj61yV6b8lIkMOwadMIuf2g6qSyjzof0HMkQuYlMVRxVLMI1usyLALuvwEmeRREniZe/3ukPd7SjA8Ty/x5KRlEo4ni52Gt8+tGu1a8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=VNboEGhm; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=RIEyJqyS; arc=none smtp.client-ip=202.12.124.152 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="VNboEGhm"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="RIEyJqyS" Received: from phl-compute-01.internal (phl-compute-01.internal [10.202.2.41]) by mailfhigh.stl.internal (Postfix) with ESMTP id A6A447A0173; Wed, 12 Aug 2026 09:23:37 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-01.internal (MEProxy); Wed, 12 Aug 2026 09:23:38 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786541017; x= 1786627417; bh=pS5GXvFi1fnnWCzMO/zMVeDQZR9m016msjFusJIgzZw=; b=V NboEGhmK6Pla17ahIWFTw4i3XELykb3oHfqwW6OgeTb3LJu8CUt8hT0s9XQQTZAE CdC6PAU88mu3k0sU4/5PmnaWZq+w9QWvCOK0UZ+MNRr9tUffrcbpSz07N52Bc0FE SZ8kHK+7kWLXH2YbOdU6O+PBv1QTpsGhT7TQq5mOpMjS+0H3K9H+Nu4xD28OICUy COVk1wcbfKNwmJZql9r9XkDs6ax+xM13Bir+8pzy1lx1mVnaERyeKDJNuVlAvULZ DMvmjjZ7KmNTBu2M9pxyYxjW7FhtKDsqkPqOdyDA+42k4kKTdXiXg07m2DecEqlL Cdq83SEHzYvkvrbNLYAFw== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786541017; x=1786627417; bh=p S5GXvFi1fnnWCzMO/zMVeDQZR9m016msjFusJIgzZw=; b=RIEyJqySDJB3Cbkll BW1kF8PFqZD2EzQEM5neL7nAOFvxYV/E9ilA8AEuZCrRv4TLnkRMWMQtJorifho6 b2hlURWa/H6xpDTXuvgEKtQr2WH5yzdYIXhDAFj/Py+IFOmtggnfYpmAmYPdk0lR qDYhWXQYpTrZfJQvhr/o/wM22vUEhimghAGLB7Y01o7DoPYVNEtigDwYNwYoH3io HYEJNA9b1prHfZ8d/4HmcmHFcXSJQ2teCRPdTa1aZ2P6QMAnrbqhtVKo32v+faxE KV1ygVvb1L5Qhm4TIH53K63XiPhoI1QP6zw/CcQ/BlO5a41E3ttxIwpF+/6LKEC1 rHLVw== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFkdXlhB1yyyGZX3J7dnv16PFzu5W734vHfSlZD/y8c/c/hfyf9JcTqEblNR0Fovw r+VLRtLwlZT2tYuTqMJsjjjBxJMgoAM7D6pmLt36lViB2yT50Uubj7I3KQF8rzeVMBK9wq rm57aKmY5NwWN7q607eE1gnukbthccXm6lhrmbDGqX6XovHG8JRMLfqJj+jBW3X680r35h P9o/w0aRWYBwLgRKYH9oTyb2lbiE5fXNRZyQYMHMmm7yuduJ70EU4YPdENzrqcZn4mtIUi PPdmMfUKqzUyYVrtf6UgPYBrZU3VzLkrgoUPTsKVKEpiHQHdj45X731bFwNLlQpPGy3QFN V7dgSNKYqlBiXdTHQIw/lw5HlTJhpNpUrF9ugBKaUaWKSH5j7b1IYg/YApSXWwvfWBfuhr QAXox8IkePgY6WG5m4Oynz5vJSaOTxCdYVSLz9VnAkaKtkRO/FBY+OQ3RH210u/pRfffHq w5bIr5qXPR2cA9VuUDzpQKZTnAgtC/nBo+NNBNXkcdlA94Bxo1TKm8bb7GlOnfSIqIBOaK J4PtG3ednYkGlLmWVW0qLO319MeCqp/3M3mmL2oqfVzkWFmpPYaL8PhukIYOudoQ9MI0En 8bK9jL1xCgVmT/Wthen8UJ7B7wOIpYEHlydIpTC1Td5KcLGLtrRIiwlFfnFQ X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Wed, 12 Aug 2026 09:23:36 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v3 11/18] selftests/mm: parameterize the mixed-source collapse case by source order Date: Wed, 12 Aug 2026 14:22:57 +0100 Message-ID: <20260812132304.199287-12-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260812132304.199287-1-kirill@shutemov.name> References: <20260812132304.199287-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" collapse_order_mixed_sources faults its region as order-2 folios and collapses them to the -c target. Order 2 sits below the contpte threshold on both arm64 page-size configurations, so nothing in this suite unfolds a contpte source on purpose. Let -s name the source order alongside -c, which was rejected before. The case then faults at that order, keeping order 2 when -s is absent, and the source order has to be a supported mTHP order below the target. Other cases are unaffected: they enable their source order locally. "-s 5 -c 7" on arm64/64K then collapses contpte-mapped sources into a larger mTHP. Assisted-by: Claude-Code:claude-opus-5 Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/khugepaged.c | 25 +++++++++++++++---------- 1 file changed, 15 insertions(+), 10 deletions(-) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index 7889ea1cd322..3099e7b720d9 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -1412,10 +1412,13 @@ static void collapse_order_max_ptes_none(struct col= lapse_context *c, static void collapse_order_mixed_sources(struct collapse_context *c, struct mem_ops *ops) { + int source_order =3D anon_order ? anon_order : MIN_MTHP_ORDER; struct thp_settings settings; void *p; =20 - if (collapse_order <=3D MIN_MTHP_ORDER) { + /* Sources must be a supported mTHP order strictly below the target. */ + if (source_order >=3D collapse_order || + !(thp_supported_orders() & (1UL << source_order))) { ksft_test_result_skip("%s: no source order below target\n", __func__); return; @@ -1423,23 +1426,22 @@ static void collapse_order_mixed_sources(struct col= lapse_context *c, =20 mthp_push_target_order(); =20 - /* Fault the whole region as order-MIN_MTHP_ORDER folios. */ + /* Fault the whole region as order-@source_order folios. */ settings =3D *thp_current_settings(); - settings.hugepages[MIN_MTHP_ORDER].enabled =3D THP_ALWAYS; + settings.hugepages[source_order].enabled =3D THP_ALWAYS; thp_push_settings(&settings); p =3D ops->setup_area(1); ops->fault(p, 0, hpage_pmd_size); thp_pop_settings(); =20 /* - * The order is enabled, but the allocator can still fall back under - * fragmentation. That leaves nothing to collapse from, which is the - * machine's answer rather than a reason to end the run. + * The order is enabled and supported, but the allocator can still fall + * back under fragmentation. That leaves nothing to collapse from, + * which is the machine's answer rather than a reason to end the run. */ - if (!is_range_backed_by_folio_orders(p, hpage_pmd_size, MIN_MTHP_ORDER, + if (!is_range_backed_by_folio_orders(p, hpage_pmd_size, source_order, pagemap_fd, kpageflags_fd)) { - ksft_print_msg("No order-%d sources to collapse...", - MIN_MTHP_ORDER); + ksft_print_msg("No order-%d sources to collapse...", source_order); skip("Skip"); ops->cleanup_area(p, hpage_pmd_size); thp_pop_settings(); @@ -1448,7 +1450,8 @@ static void collapse_order_mixed_sources(struct colla= pse_context *c, } =20 madvise(p, hpage_pmd_size, MADV_HUGEPAGE); - ksft_print_msg("Collapse region backed by smaller large folios..."); + ksft_print_msg("Collapse region backed by order-%d sources...", + source_order); if (!khugepaged_wait_full_pass()) fail("Timeout"); else if (window_collapsed(p, hpage_pmd_size)) @@ -1479,6 +1482,8 @@ static void usage(void) fprintf(stderr, "\t\t-s: mTHP size, expressed as page order.\n"); fprintf(stderr, "\t\t Defaults to 0. Use this size for anon or shmem a= llocations.\n"); fprintf(stderr, "\t\t-c: collapse order for mTHP collapse, expressed as p= age order.\n"); + fprintf(stderr, "\t\t With -s, -s names the mTHP source order for the\= n"); + fprintf(stderr, "\t\t mixed-source case (source order below the target= ).\n"); exit(1); } =20 --=20 2.54.0 From nobody Tue Sep 29 04:44:47 2026 Received: from fhigh-b1-smtp.messagingengine.com (fhigh-b1-smtp.messagingengine.com [202.12.124.152]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 78317453A36; Wed, 12 Aug 2026 13:23:41 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.152 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541022; cv=none; b=tMMaEwUFB7tW/5L6FYJjUkbzg0z/Idh4mu/Gg3901XIRTw27k3Ell5N6oQU8vtd7htNE/ynWBR/Jyig4AcZ+Zw717rRyuSBEWOpmBT0BOpFOqNa0ZzJv1XYJxDk7yvyWxTdmuHbOB8fB0SMNDgRkvdS8Xv0bpxS50116mtVehss= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541022; c=relaxed/simple; bh=l+FhFkhBEBhzqVzwYB7BWQprhxPOlLsgkwz6Fi6bk2g=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=T0rsgyt/KwmQ2J1sLy0SqxVq8WeEjERv2hNAMTNewnNw9Ua3VX5Ifxnb64OnSmONfTA2ys8TzDnXuTfoSxEoRGEKieAXNJkkMccQ7TcQ+LYRDksSxySOuxnGzZ8rpqOcY5e54Pn4u9+2mlDaCiSWM1qhCIrweh1vdgEh1n6Bz3A= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=uhASvvD7; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=BLzgpCa1; arc=none smtp.client-ip=202.12.124.152 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="uhASvvD7"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="BLzgpCa1" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfhigh.stl.internal (Postfix) with ESMTP id 2E3357A0176; Wed, 12 Aug 2026 09:23:40 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-06.internal (MEProxy); Wed, 12 Aug 2026 09:23:40 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786541020; x= 1786627420; bh=yE9UcmZWN8CYeDypZM4ijLjNlZ/odamXd7gxHiGqY8Q=; b=u hASvvD7GOICdfoUSDkwvcr6MX9Vjo+PMCAx2KtGmcxRqgEpZKxMd/eKMLzP+HmC4 hzCd44Fns4gIdUZ+IsBC8SBRewsa36m01xE2pw/QzoIfRLnASyENuLC1zR/82e5B HgtkAcfrxQ+22/4ZWHmal45bbi4M5kiXpAudHDhGhVVKtiQVOC8xM2NNJuK1kiRJ EHIkxrXgMiGF8RpP7Zx+Hc73ADIoKjzqi9ZesA1X+jOjeDu0wVnOkXZnDWGybX7G LDYvUmFIt1mJy36cmrxHK+SmmE4XKxckQEur71fzrqQD8gStL3Ar+BCj5eYZsRgg +7O6wa562NK1IWPzenO4w== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786541020; x=1786627420; bh=y E9UcmZWN8CYeDypZM4ijLjNlZ/odamXd7gxHiGqY8Q=; b=BLzgpCa1S2kn7xZeV 2LeFbPCgkhGmw+Hw0U6ntUMsMrg9Oe1xF5IyYAikp7vZpH6WlJlRjNrFjcEaaE4M P8zAjw0mYsuqXchpvtcDC5CoImsX8lQ7EU0YwhPWZ3bMSJG/zgSPt9rsCu5l3CJW yEZxFzh1wEBoCBXb1FYvqE5/fx8fC1GDNGgIpVf41yZ54p+vNs7VBRNaII+wNxxZ TZrp+39ejrWWNapiL48k5xqaS3aQaDSsbk4ujlDI7ZQLIHJCpJlF4wh6d769v0nl pxinCaIpmwvELQXCeJ5g6/OG5TQc9D+FlLbRS+X0LTpfdGTTRCXlHP1JSX6KIXJL EBxsQ== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEJjeAG5mPidrlxx5WYeorxbJVgJIKvTA5wXegfbw/cux8oxcUciNIpDIuTJE+lqf 6MvKMtIYnfrpE9CE9brIFiUghXXeSu7acy7yTiQzGdGRUaVwSNB0Wtfy8GXhxbblHF9PD0 zPXwLCOAHUDQMRV+OGFoRd51Bfw/8mHZFt6NMjCskGz8KRys5NyppP3tQz3pg+1FIu5GBp PJQAWgZTF13VhomE3d+3g7Xo8LQtMXvrRUSh5YzlZkIRUVCw4Olhx0/o/iEmltBto/jcsB 4mCen6EwKCCQNN9U1Ldmtp3E1EMhPbaX1d9tgEecmRx82uYLegpVLGl6lWthe5rRIcaIdx /lwSL5gwXuq3E4Lw/HBc9fXact2cFi1yiSHQEFhRz3naFma7cY/dEuMeG761JoaF7Uof1w yNYdpA0Vx0SkMA85H6uqkKj9m0fa07KOU7GdORi9GqP018gPW4vRu1GctHkp9pcAKQr9VY V6xLlZDxU4h1joU2gZbKO9mvfcFNQ3VDN5H8Abq4XpZHcYQlGaJC3eev6f8qvLelEul0qw MWLochdmzlZ5pPWIk+kn8fAx8gPJfBacUdqwsMyz2EcS75QJjh7GVkvlKwhnSVI9gZPKn5 B0l+LcBLnsBrLvRbM/n1z5H6gofPVqbnt06qPATKf9PLHA1qbXOIxeQ6u2OA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Wed, 12 Aug 2026 09:23:39 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v3 12/18] selftests/mm: cover a shared-source collapse write race Date: Wed, 12 Aug 2026 14:22:58 +0100 Message-ID: <20260812132304.199287-13-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260812132304.199287-1-kirill@shutemov.name> References: <20260812132304.199287-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" collapse_fork checks that a fork-shared range collapses in the process that asks for it while the co-sharer keeps its own page, but the co-sharer sits still while that happens. Add a case where the co-sharer writes to the shared range throughout the collapse. CoW has to keep the two sides apart under those writes: the collapsing child must see the content from before the fork, and the writing parent must see only its own writes. It passes on an unmodified kernel, so it pins down isolation that khugepaged collapse already provides. Assisted-by: Claude-Code:claude-opus-5 Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/khugepaged.c | 69 +++++++++++++++++++++++++ 1 file changed, 69 insertions(+) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index 3099e7b720d9..9c8830fed0a2 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -1207,6 +1207,72 @@ static void collapse_max_ptes_shared(struct collapse= _context *c, struct mem_ops ksft_test_result_report(exit_status, "%s\n", __func__); } =20 +/* + * Content stays isolated while a co-sharer writes concurrently. A shared + * source is copied live (not frozen), relying on it being CoW - immutable + * for the duration of the copy; a co-sharer's write goes to a CoW copy. T= he + * collapsing child must see the pre-fork content, the writing parent only + * its own writes. + */ +static void collapse_fork_cow_race(struct collapse_context *c, struct mem_= ops *ops) +{ + const unsigned long shared =3D 64 * page_size; + const int stride =3D page_size / sizeof(int); + int wstatus, child_status, i, n =3D shared / page_size; + /* volatile: the loop below must really store, on every iteration */ + volatile int *ip; + void *p; + + p =3D ops->setup_area(1); + ip =3D p; + ops->fault(p, 0, shared); /* shared prefix, pre-fork pattern */ + + ksft_print_msg("Fork, collapse in the child while the parent rewrites..."= ); + if (!fork()) { + int collapse_status; + + ops->fault(p, shared, hpage_pmd_size); /* private remainder */ + c->collapse("Collapse a range shared with a writing co-sharer", + p, 1, ops, true); + collapse_status =3D exit_status; + for (i =3D 0; i < n; i++) + if (ip[i * stride] !=3D i + 0xdead0000) + break; + if (i =3D=3D n) + success("OK"); + else + fail("Fail: child content"); + /* The content check must not bury a failed collapse. */ + if (exit_status !=3D KSFT_FAIL) + exit_status =3D collapse_status; + ops->cleanup_area(p, hpage_pmd_size); + _exit(exit_status); + } + + /* Hammer the parent's own writes over the shared prefix. */ + for (int it =3D 0; it < 200000; it++) + for (i =3D 0; i < n; i++) + ip[i * stride] =3D i + 0xbeef0000; + + wait(&wstatus); + /* A child that died reading the racing pages is a failure, not a zero. */ + child_status =3D WIFEXITED(wstatus) ? WEXITSTATUS(wstatus) : KSFT_FAIL; + + ksft_print_msg("Check the parent sees only its own writes..."); + for (i =3D 0; i < n; i++) + if (ip[i * stride] !=3D i + 0xbeef0000) + break; + if (i =3D=3D n) + success("OK"); + else + fail("Fail: parent content"); + ops->cleanup_area(p, hpage_pmd_size); + /* Same again: our own check must not bury the child's verdict. */ + if (exit_status !=3D KSFT_FAIL) + exit_status =3D child_status; + ksft_test_result_report(exit_status, "%s\n", __func__); +} + static void madvise_collapse_existing_thps(struct collapse_context *c, struct mem_ops *ops) { @@ -1770,6 +1836,9 @@ int main(int argc, char **argv) TEST(collapse_max_ptes_shared, khugepaged_context, anon_ops); TEST(collapse_max_ptes_shared, madvise_context, anon_ops); =20 + TEST(collapse_fork_cow_race, khugepaged_context, anon_ops); + TEST(collapse_fork_cow_race, madvise_context, anon_ops); + TEST(madvise_collapse_existing_thps, madvise_context, anon_ops); TEST(madvise_collapse_existing_thps, madvise_context, read_only_file_ops); TEST(madvise_collapse_existing_thps, madvise_context, read_write_file_rea= d_ops); --=20 2.54.0 From nobody Tue Sep 29 04:44:47 2026 Received: from fout-b7-smtp.messagingengine.com (fout-b7-smtp.messagingengine.com [202.12.124.150]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0F45C45C704; Wed, 12 Aug 2026 13:23:43 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.150 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541025; cv=none; b=LL/wFdg1YcDfqTNMAQdu8NIq4uziX5a7xD2wyWKzuKYzJ1Md1264xGs5BexByyAvtsxNDuoEl5xYXoHw/PVCKuATA7GN9g0xeE3ndpDzqMQEHaKiKRUSxrOvxzWvjkQ+OjOemlC5iWFJ48GtsxGcwE5MfrFOyAWe1zsAeZPDRWk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541025; c=relaxed/simple; bh=MVPtYn0BbjGeUCqhsSUxODxC73aKZw2N13CpUSkBo2E=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Y67ykArqXn9Tmoqxc5Gf1cxZlK3K4IN2MQNFD/j32Q4urQ0sMP4cfM69dcUq22E4mSgrmDBHN/B5KVFO5yFZiGftPL4DEhQCAc7OicoLfcXT20OcWhczmN/y3zK/rmJWXWvKWkRNHNtAliVagJ0lewGfjapq1s8z3pVN254wTHU= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=DrRJ4skz; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=ZD3siZzM; arc=none smtp.client-ip=202.12.124.150 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="DrRJ4skz"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="ZD3siZzM" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfout.stl.internal (Postfix) with ESMTP id B90F21D00170; Wed, 12 Aug 2026 09:23:42 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-06.internal (MEProxy); Wed, 12 Aug 2026 09:23:43 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786541022; x= 1786627422; bh=8jKw/Rh2BzJjfFUeMIccyUKsF4XLAd9SgrDy5I5e8ZA=; b=D rRJ4skzYiRYr5OL37Xe+9qufDx7Lzxdk3MB63J/PCpZJQq26e9oKhgY0J7LUGluW zrR78xjvubiDIB0mPkcCAhQQlRzM0kSHWfzEB3cTm4OEYOl65q9fZlCxAWYeWU36 klqvcIq3uNSnMv+jLhe7icgx/C9QBU6o+VlgbqIM15KgWQvielbEuT11em58hy4r HNytnpEHYiaIwhuKPfzQ6yn57Xu2YzzWwZl+dm8pp3hUuho5lxvY6sBCMJiLjgxK /r07KtP5bfR4VRoEfTDIUoAY1kDpN6cErtfpV5dkZVgZljVuyIILVxEEL2yTMMKL LM02s4nKn/iPd6RziFrjw== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786541022; x=1786627422; bh=8 jKw/Rh2BzJjfFUeMIccyUKsF4XLAd9SgrDy5I5e8ZA=; b=ZD3siZzMuhvbgQ6mn j4siAbf+JaheR5KMndotrjoJibIwh8Rx17OL27Jwp++VgM54k1kNDyEQ0DXJBhm3 nXX0AyG75ychXRA2F3eImG6mdL8YRqj3Z7DzAnDXXE9Fz7Jsg40GmdU/FrNXTJ1o d1j5Ep/H+1+s3OFclXEN7g5GDzLWtr+8ilVWg8stCr9fXheaZUHljI9gUE8PoGHV rwytThgfd3kn6aBJSennUVoIV85lHcZKzX3Xep6xPrTFVG5c7L0FiD7Rp9upNIwW v1MbIxVJVIAC5ECOGCycIMURl+Bsf7HcJoWBtXKbDkDo/+ZF8vkwLmE94HvcFyLP 473YA== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEJjeAG5mPidrlxx5WYeorxbJVgJIKvTA5wXegfbw/cux8oxcUciNIpDIuTJE+lqf 6MvKMtIYnfrpE9CE9brIFiUghXXeSu7acy7yTiQzGdGRUaVwSNB0Wtfy8GXhxbblHF9PD0 zPXwLCOAHUDQMRV+OGFoRd51Bfw/8mHZFt6NMjCskGz8KRys5NyppP3tQz3pg+1FIu5GBp PJQAWgZTF13VhomE3d+3g7Xo8LQtMXvrRUSh5YzlZkIRUVCw4Olhx0/o/iEmltBto/jcsB 4mCen6EwKCCQNN9U1Ldmtp3E1EMhPbaX1d9tgEecmRx82uYLegpVLGl6lWthe5rRIcaITh sTJxhOaQa0UDlgJNE1X0o3KXKkJIR7vSM7Z6eevJ8msNKRqMwVvCiUidN0tRNeaYtAC8DY EarZSmGvSZnVQ0bq7TbqkGVcxX3ZqE6DlyK2cROhVkzNx0GJmFzhBdan28pI7NjPtnDz8t l9O1/0Hsa+Y+ihi+BscZkdEqyXCey8BGiviv1JhVAfFKQ1w9WaIj2nU4MXnll8XNIIU4Aw UGeFg/1EvbbF/XcY8WJlj+c9w+luTI0H5paeUrYh+Ig+Yp9Jbx6Ka+rfpwqH8mX+HjYMVR bf5kYEdmxvoC5T89HKOazFgyrr/R5WJY0h/fhQO1xiYc4tnRoqQQBhTPK6Iw X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Wed, 12 Aug 2026 09:23:41 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v3 13/18] selftests/mm: run every supported collapse order by default Date: Wed, 12 Aug 2026 14:22:59 +0100 Message-ID: <20260812132304.199287-14-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260812132304.199287-1-kirill@shutemov.name> References: <20260812132304.199287-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The mTHP collapse cases only run when the caller names both the context and an order, so a plain ./khugepaged covers the PMD contexts on anon and nothing else. run_vmtests.sh pinned order 4 and covered no other. Without -c, run the mTHP cases once per supported anon THP order below the PMD, and include that context in both the no-argument invocation and "all". -c still pins one order, and now says what is wrong when the order is above the PMD or unsupported instead of printing the usage text. An order at or below the -s source order is skipped: the sources would already be the size being asked for. Both orders end up as array indices and shift counts, so -s rejects a negative order and -c anything at or below 0, rather than letting either reach them. The mTHP context has only anon cases, so a run that names a different mem_type -- "all:shmem", say -- drops it again rather than refusing to start. Naming both explicitly still does. A case carries the order it was registered at, so a result names it: # Run test: collapse_order_max_ptes_none (mthp_khugepaged:anon, order 6) On x86-64 with 4K pages that is orders 2 through 8, and ./khugepaged goes from 28 results to 77 in 21 seconds, so run_vmtests.sh can drop its pinned order-4 line. Assisted-by: Claude-Code:claude-opus-5 Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/khugepaged.c | 111 ++++++++++++++++++---- tools/testing/selftests/mm/run_vmtests.sh | 2 - 2 files changed, 92 insertions(+), 21 deletions(-) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selfte= sts/mm/khugepaged.c index 9c8830fed0a2..f5773d447542 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -31,6 +31,9 @@ static unsigned long page_size; static int hpage_pmd_nr; static int anon_order; static int collapse_order; +static bool collapse_order_given; +static int collapse_orders[NR_ORDERS]; +static int nr_collapse_orders; static int pagemap_fd =3D -1; static int kpageflags_fd =3D -1; =20 @@ -1548,6 +1551,7 @@ static void usage(void) fprintf(stderr, "\t\t-s: mTHP size, expressed as page order.\n"); fprintf(stderr, "\t\t Defaults to 0. Use this size for anon or shmem a= llocations.\n"); fprintf(stderr, "\t\t-c: collapse order for mTHP collapse, expressed as p= age order.\n"); + fprintf(stderr, "\t\t Defaults to every supported order below the PMD.= \n"); fprintf(stderr, "\t\t With -s, -s names the mTHP source order for the\= n"); fprintf(stderr, "\t\t mixed-source case (source order below the target= ).\n"); exit(1); @@ -1555,6 +1559,7 @@ static void usage(void) =20 static void parse_test_type(int argc, char **argv) { + bool mthp_context_implied =3D false; int opt; char *buf; const char *token; @@ -1566,6 +1571,7 @@ static void parse_test_type(int argc, char **argv) break; case 'c': collapse_order =3D atoi(optarg); + collapse_order_given =3D true; break; case 'h': default: @@ -1573,12 +1579,25 @@ static void parse_test_type(int argc, char **argv) } } =20 + /* + * Both orders end up as array indices and shift counts, so neither + * can be negative, and a zero collapse order asks for base pages. + */ + if (anon_order < 0 || anon_order > hpage_pmd_order) + ksft_exit_fail_msg("-s takes an order in 0..%d, not %d\n", + hpage_pmd_order, anon_order); + if (collapse_order_given && + (collapse_order <=3D 0 || collapse_order > hpage_pmd_order)) + ksft_exit_fail_msg("-c takes an order in 1..%d, not %d\n", + hpage_pmd_order, collapse_order); + argv +=3D optind; argc -=3D optind; =20 if (argc =3D=3D 0) { - /* Backwards compatibility */ + /* Everything that needs no argument of its own: anon, every context */ khugepaged_context =3D &__khugepaged_context; + mthp_khugepaged_context =3D &__mthp_khugepaged_context; madvise_context =3D &__madvise_context; anon_ops =3D &__anon_ops; return; @@ -1589,13 +1608,19 @@ static void parse_test_type(int argc, char **argv) =20 if (!strcmp(token, "all")) { khugepaged_context =3D &__khugepaged_context; + mthp_khugepaged_context =3D &__mthp_khugepaged_context; madvise_context =3D &__madvise_context; + + /* + * "all" sweeps the mTHP context in, but it only has anon + * cases: step it aside for the other mem_types rather than + * refusing the whole run. + */ + mthp_context_implied =3D true; } else if (!strcmp(token, "khugepaged")) { khugepaged_context =3D &__khugepaged_context; } else if (!strcmp(token, "mthp_khugepaged")) { mthp_khugepaged_context =3D &__mthp_khugepaged_context; - if (collapse_order <=3D 0 || collapse_order >=3D hpage_pmd_order) - usage(); } else if (!strcmp(token, "madvise")) { madvise_context =3D &__madvise_context; } else { @@ -1611,20 +1636,20 @@ static void parse_test_type(int argc, char **argv) read_write_file_write_ops =3D &__read_write_file_write_ops; anon_ops =3D &__anon_ops; shmem_ops =3D &__shmem_ops; - if (mthp_khugepaged_context) - usage(); } else if (!strcmp(buf, "anon")) { anon_ops =3D &__anon_ops; } else if (!strcmp(buf, "file")) { read_only_file_ops =3D &__read_only_file_ops; read_write_file_read_ops =3D &__read_write_file_read_ops; read_write_file_write_ops =3D &__read_write_file_write_ops; - if (mthp_khugepaged_context) + if (mthp_khugepaged_context && !mthp_context_implied) usage(); + mthp_khugepaged_context =3D NULL; } else if (!strcmp(buf, "shmem")) { shmem_ops =3D &__shmem_ops; - if (mthp_khugepaged_context) + if (mthp_khugepaged_context && !mthp_context_implied) usage(); + mthp_khugepaged_context =3D NULL; } else { usage(); } @@ -1646,8 +1671,13 @@ struct test_case { struct mem_ops *ops; const char *desc; test_fn fn; + int order; /* mTHP contexts: the collapse order */ }; =20 +/* + * Enough for every case at every order the kernel offers: the mTHP context + * runs its cases once per supported order below the PMD. + */ #define MAX_TEST_CASES 256 static struct test_case test_cases[MAX_TEST_CASES]; static int nr_test_cases; @@ -1661,6 +1691,7 @@ static int nr_test_cases; .ops =3D o, \ .desc =3D #t, \ .fn =3D t, \ + .order =3D collapse_order, \ }; \ } \ } while (0) @@ -1701,10 +1732,40 @@ int main(int argc, char **argv) =20 parse_test_type(argc, argv); =20 - if (mthp_khugepaged_context && - !(thp_supported_orders() & (1UL << collapse_order))) - ksft_exit_skip("Order %d is not a supported anon THP order\n", - collapse_order); + if (mthp_khugepaged_context) { + unsigned long orders =3D thp_supported_orders(); + + if (collapse_order_given) { + /* -c pins one order; it has to be one we can build */ + if (collapse_order >=3D hpage_pmd_order) + ksft_exit_fail_msg("-c takes an order below the PMD order (%d)\n", + hpage_pmd_order); + if (!(orders & (1UL << collapse_order))) + ksft_exit_skip("Order %d is not a supported anon THP order\n", + collapse_order); + if (collapse_order <=3D anon_order) + ksft_exit_skip("-c %d needs a source order below it, -s says %d\n", + collapse_order, anon_order); + collapse_orders[nr_collapse_orders++] =3D collapse_order; + } else { + /* + * Otherwise every order a collapse could produce. -s + * makes the fault path hand out folios of that order, + * so a target at or below it has nothing to collapse: + * the sources are already the size being asked for. + */ + int first =3D anon_order ? anon_order + 1 : MIN_MTHP_ORDER; + + if (first < MIN_MTHP_ORDER) + first =3D MIN_MTHP_ORDER; + for (int i =3D first; i < hpage_pmd_order; i++) { + if (orders & (1UL << i)) + collapse_orders[nr_collapse_orders++] =3D i; + } + if (!nr_collapse_orders) + ksft_print_msg("mTHP cases skipped: no order above the source\n"); + } + } =20 if (mthp_khugepaged_context) { pagemap_fd =3D open("/proc/self/pagemap", O_RDONLY); @@ -1765,7 +1826,17 @@ int main(int argc, char **argv) TEST(collapse_full, khugepaged_context, read_write_file_read_ops); TEST(collapse_full, khugepaged_context, read_write_file_write_ops); TEST(collapse_full, khugepaged_context, shmem_ops); - TEST(collapse_full, mthp_khugepaged_context, anon_ops); + for (int i =3D 0; i < nr_collapse_orders; i++) { + collapse_order =3D collapse_orders[i]; + TEST(collapse_full, mthp_khugepaged_context, anon_ops); + TEST(collapse_empty, mthp_khugepaged_context, anon_ops); + TEST(collapse_single_mthp, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_single_window, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_partial_window, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_max_ptes_none, mthp_khugepaged_context, anon_ops); + TEST(collapse_order_mixed_sources, mthp_khugepaged_context, anon_ops); + } + TEST(collapse_full, madvise_context, anon_ops); TEST(collapse_full, madvise_context, read_only_file_ops); TEST(collapse_full, madvise_context, read_write_file_read_ops); @@ -1773,14 +1844,8 @@ int main(int argc, char **argv) TEST(collapse_full, madvise_context, shmem_ops); =20 TEST(collapse_empty, khugepaged_context, anon_ops); - TEST(collapse_empty, mthp_khugepaged_context, anon_ops); TEST(collapse_empty, madvise_context, anon_ops); =20 - TEST(collapse_single_mthp, mthp_khugepaged_context, anon_ops); - TEST(collapse_order_single_window, mthp_khugepaged_context, anon_ops); - TEST(collapse_order_partial_window, mthp_khugepaged_context, anon_ops); - TEST(collapse_order_max_ptes_none, mthp_khugepaged_context, anon_ops); - TEST(collapse_order_mixed_sources, mthp_khugepaged_context, anon_ops); =20 TEST(collapse_single_pte_entry, khugepaged_context, anon_ops); TEST(collapse_single_pte_entry, khugepaged_context, read_only_file_ops); @@ -1854,7 +1919,15 @@ int main(int argc, char **argv) for (int i =3D 0; i < nr_test_cases; i++) { struct test_case *t =3D &test_cases[i]; =20 - ksft_print_msg("\n# Run test: %s (%s:%s)\n", t->desc, t->ctx->name, t->o= ps->name); + if (t->ctx =3D=3D &__mthp_khugepaged_context) { + collapse_order =3D t->order; + ksft_print_msg("\n# Run test: %s (%s:%s, order %d)\n", + t->desc, t->ctx->name, t->ops->name, + t->order); + } else { + ksft_print_msg("\n# Run test: %s (%s:%s)\n", t->desc, + t->ctx->name, t->ops->name); + } t->fn(t->ctx, t->ops); } =20 diff --git a/tools/testing/selftests/mm/run_vmtests.sh b/tools/testing/self= tests/mm/run_vmtests.sh index 2652a7920b80..8bf898b71350 100755 --- a/tools/testing/selftests/mm/run_vmtests.sh +++ b/tools/testing/selftests/mm/run_vmtests.sh @@ -412,8 +412,6 @@ CATEGORY=3D"thp" run_test ./khugepaged all:shmem =20 CATEGORY=3D"thp" run_test ./khugepaged -s 4 all:shmem =20 -CATEGORY=3D"thp" run_test ./khugepaged -c 4 mthp_khugepaged:anon - # Try to create XFS if not provided if [ -z "${SPLIT_HUGE_PAGE_TEST_XFS_PATH}" ]; then if test_selected "thp"; then --=20 2.54.0 From nobody Tue Sep 29 04:44:47 2026 Received: from fhigh-b1-smtp.messagingengine.com (fhigh-b1-smtp.messagingengine.com [202.12.124.152]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5E4ED45C71B; Wed, 12 Aug 2026 13:23:46 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.152 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541028; cv=none; b=ZfN3d0ft1TpXQiPWErsM6VZGJ98bUxmzqCJApYkbCK4SnHdTnKCZ/Lop3swPFn3bC2kE1yAs60ww1gcI43/sUOD7kYQWFX5JbzLka7Lur8zx4g2jitiTqGJD7uPcygIOD0cJJw8OA3drjozOLiKQpAbIQnMVCDdtN7qm/0weeMg= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541028; c=relaxed/simple; bh=0eJN8oWrcQrFzP+pVgHgg80N5Lq8LdRUnTOswfFc+iA=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=TzyVZCF1LSHVMN9ubks1I1Y9ETpUfw5JiFV9y8mv7IXrC3aQla2YC/BWW1MemEVN3jGf1Aiw9UH+KJ0oETUSY2xHeGLLWDcAcINJf2ocL5JAS1lnK3Tw3pT/CVvA97ezujDFtBX0Nkq1/phKhuviIu8eJX03GV7yNRpT9COQFDA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=xMiUF7Tn; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=LNMkxvW6; arc=none smtp.client-ip=202.12.124.152 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="xMiUF7Tn"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="LNMkxvW6" Received: from phl-compute-02.internal (phl-compute-02.internal [10.202.2.42]) by mailfhigh.stl.internal (Postfix) with ESMTP id 552437A00F7; Wed, 12 Aug 2026 09:23:45 -0400 (EDT) Received: from phl-frontend-03 ([10.202.2.162]) by phl-compute-02.internal (MEProxy); Wed, 12 Aug 2026 09:23:45 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786541025; x= 1786627425; bh=6PELDQDgdL1ofsn+a6Ycmvq65G/VhZ4TTdp7hsL+eOo=; b=x MiUF7TnyTe9UL11b243rk2E1RA2R4fDL4ah/nlAVyRWWi4qf8nPckK5tgdtcld4c amqEg70rD+o91Rlh4Wk8d5xNQo9+hQyVgMoG+feDe4DkXS2Ax7J49mFjEvTPskJz 6Z1W9E9FffqH31qf55x0Ko2/bbyuduzfsw7/3giWt59+XOD9+AMdEVK2oxgkdwCy POTeWi5tkwXCZl5BNi5YSE/7kgb8OlbzIWtxtDL27jCNf7+wdBkOQSilu1BwkH3L Q6IqvDJSMqb+u+DsqiVMgkCDLP3RG9C+OpjL9Oj5KS1E9Yq33a8TAniOskOcONp/ R+uho20OGMllHEjAFfmxA== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786541025; x=1786627425; bh=6 PELDQDgdL1ofsn+a6Ycmvq65G/VhZ4TTdp7hsL+eOo=; b=LNMkxvW6ZyEdq3GZK ZC8Lv5PvbN5woCHPVrn9wPHkWcqh1qxCR9IsEryNGe4HgcDt6J7hNQ5KD2trBVoM bNuS8wMYo9xSGramz+QZMFhBK95tu9RaSPr36vdCP+cVTvRjr+GDLDL+vuPm1Vyg OnhMgszKsvVZAYa0EjiKgXx41iCpZ58CRnLAMI007l2KX/zNp3SQVL0Dpv9j6P+1 Tdxrnm5YnyjO4nHpUdU1R6KCCRbYF3zo5z3izz1XE025uI1ZJzm4f4l7C/x0VjgB 0Yi5U1QRnI52IwqyvCxlxgYVsETKdrLmOY/46QWcX3SvLdP2/8Wn4McHOK8aQ/I0 wwKaQ== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEWZPjKBYDoSh3uCuJ0ogiZvSuDj1RAgMuSmimalRft9LCPvZxyrluChLfTbkjo2r KkCK8Ttm4mqxi+W2/UjRoyI4gRuaMCGjIyoE2sYafZhhTXl0iIGCKEOr9/WVtZfHPc/pjz XsLpzK8U2Xu6F30SrJOujIlp1bmlomKUtpU/iiXSkRy6wwaI7NV8wAm1PQ0lHgRG0nCRKX WKIh8TtevEly/nhHwG2deu/klJhGkXN7h/CDYOOyqjOIkeF7MAFsC6fZtQDVeVkgzM4xU1 Py5NrFUehZdXXi5LHUYxqNjNVNbuxWN2P2YJW6J2IL8s+zlVHHPN2YdKOlFVuMcbnWK9sh Zh/iqUff1BlKwk2ETqlAUo3CDRrLY2hXcGdUqRTk7FGYeW6lpDxfJ9e3an+QrSOGW0dAzi nKdzgBpwSozJafjPi1DxKzjrwPtjRSD4wUHlSxYEOg7uHmWIvOfvgY2BsteYEKaL7mqlP3 PUIfbZB3mezZe/jUfjy8YTeXCWIJJ6Je2o0f+HB90Sg/vBo1Lp4YohG8oKeX+z0/NnVuEQ P0RBZGkV9qsKtMuEozJ0hGIxEUvuGO7K8wUpLvvSLL6lv1v6kcOWDU68CIp1xk+5BG7l1v vJrQ66VNbBl050PX/fFLJDzCXCECXV0lhzNUCbgKlnPw4GFyWE3wvIw5Kh7w X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Wed, 12 Aug 2026 09:23:44 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v3 14/18] selftests/mm: verify synchronous khugepaged driving is attributable Date: Wed, 12 Aug 2026 14:23:00 +0100 Message-ID: <20260812132304.199287-15-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260812132304.199287-1-kirill@shutemov.name> References: <20260812132304.199287-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The khugepaged tests attribute outcomes through the huge_memory tracepoints. The anon-path events carry no virtual address, but mm_collapse_huge_page_isolate() reports a source folio PFN and order, and the test matches those against the PFNs it read from pagemap before the pass. An attempt counts from whichever signal fires, so the check does not depend on which anon tracepoint a kernel emits. Add small tracefs helpers to vm_util and khugepaged_sync_check: per step, prepare one aligned window, record its source PFNs, run one khugepaged_full_pass() barrier, and require the window collapsed with exactly one attributed attempt. scan_sleep_millisecs is set high, so the test only finishes in time if the sysfs store really wakes the daemon. Passes 5/5 on x86-64 4K and arm64 64K. Assisted-by: Claude-Code:claude-opus-5 Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/Makefile | 1 + .../selftests/mm/khugepaged_sync_check.c | 217 ++++++++++++++++++ tools/testing/selftests/mm/run_vmtests.sh | 2 + tools/testing/selftests/mm/vm_util.c | 41 ++++ tools/testing/selftests/mm/vm_util.h | 4 + 5 files changed, 265 insertions(+) create mode 100644 tools/testing/selftests/mm/khugepaged_sync_check.c diff --git a/tools/testing/selftests/mm/Makefile b/tools/testing/selftests/= mm/Makefile index 2093fcf6e915..b2d6e5c12934 100644 --- a/tools/testing/selftests/mm/Makefile +++ b/tools/testing/selftests/mm/Makefile @@ -105,6 +105,7 @@ TEST_GEN_FILES +=3D merge TEST_GEN_FILES +=3D rmap TEST_GEN_FILES +=3D folio_split_race_test TEST_GEN_FILES +=3D folio_order_check +TEST_GEN_FILES +=3D khugepaged_sync_check =20 ifneq ($(ARCH),arm64) TEST_GEN_FILES +=3D soft-dirty diff --git a/tools/testing/selftests/mm/khugepaged_sync_check.c b/tools/tes= ting/selftests/mm/khugepaged_sync_check.c new file mode 100644 index 000000000000..30d3fb519fb2 --- /dev/null +++ b/tools/testing/selftests/mm/khugepaged_sync_check.c @@ -0,0 +1,217 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Synchronous khugepaged driving check. + * + * Race tests drive khugepaged through the existing sysfs controls: a + * store to scan_sleep_millisecs wakes the daemon, and full_scans + * advancing by two is a completion barrier for one full pass that + * started after setup (khugepaged_full_pass()). Verify the pair gives + * deterministic, attributable results: one barrier step over one + * prepared window produces exactly one collapse attempt on that + * window's source pages (mm_collapse_huge_page_isolate events filtered + * by source PFN and order) and the window is collapsed + * afterwards, repeatably. + * + * scan_sleep_millisecs is set to 60s to prove the wake path: without + * the wake, one barrier step would sleep multiples of that and blow + * the timeout. It also keeps the daemon from free-running between + * steps, per the khugepaged_full_pass() discipline. + */ +#define _GNU_SOURCE +#include +#include +#include +#include +#include +#include + +#include "kselftest.h" +#include "vm_util.h" +#include "hugepage_settings.h" + +#define BASE_ADDR ((void *)(1UL << 30)) +#define TARGET_ORDER 2 /* smallest order khugepaged considers */ +#define NR_ITERATIONS 5 + +static int pagemap_fd; +static int kpageflags_fd; +static int trace_events_fd =3D -1; +static unsigned long hpage_pmd_size; + +/* + * Each step switches the events off again, but a helper can still give up + * on us in between (a failing sysfs write ends the test from inside + * thp_write_num()), and huge_memory events left on are the whole machine's + * problem, not this test's. + */ +static void trace_events_off(void) +{ + if (trace_events_fd >=3D 0) + tracing_events_enable(trace_events_fd, false); +} + +/* + * Count collapse attempts attributable to our window: legacy-engine + * isolate events whose scan_pfn is one of the window's source PFNs, + * plus batch-engine per-candidate install events at the window's + * address. Either engine reports exactly once per attempt. + */ +static int count_attributed(unsigned long *pfns, int nr_pfns, + unsigned long addr, unsigned int order) +{ + char line[1024]; + int count =3D 0; + FILE *fp; + + fp =3D tracing_open_trace(); + if (!fp) + ksft_exit_fail_msg("Cannot open trace buffer\n"); + + while (fgets(line, sizeof(line), fp)) { + char *s; + unsigned long val; + unsigned int ord; + char *o; + int i; + + s =3D strstr(line, "mm_collapse_huge_page_isolate:"); + if (s) { + if (sscanf(s, "mm_collapse_huge_page_isolate: scan_pfn=3D0x%lx", + &val) !=3D 1) + continue; + o =3D strstr(s, "order=3D"); + if (!o || sscanf(o, "order=3D%u", &ord) !=3D 1 || + ord !=3D order) + continue; + for (i =3D 0; i < nr_pfns; i++) { + if (val =3D=3D pfns[i]) { + count++; + break; + } + } + continue; + } + + s =3D strstr(line, "mm_collapse_candidate:"); + if (s) { + if (!strstr(s, "pass=3Dinstall") || + !strstr(s, "result=3Dsucceeded")) + continue; + o =3D strstr(s, "addr=3D"); + if (!o || sscanf(o, "addr=3D0x%lx", &val) !=3D 1 || + val !=3D addr) + continue; + o =3D strstr(s, "order=3D"); + if (!o || sscanf(o, "order=3D%u", &ord) !=3D 1 || + ord !=3D order) + continue; + count++; + } + } + fclose(fp); + return count; +} + +static void one_step(int iteration) +{ + const size_t window =3D getpagesize() << TARGET_ORDER; + const int nr_pages =3D 1 << TARGET_ORDER; + unsigned long pfns[1 << TARGET_ORDER]; + bool collapsed, passed; + int attributed; + char *p; + int i; + + p =3D mmap(BASE_ADDR, hpage_pmd_size, PROT_READ | PROT_WRITE, + MAP_ANONYMOUS | MAP_PRIVATE | MAP_FIXED_NOREPLACE, -1, 0); + if (p !=3D BASE_ADDR) + ksft_exit_fail_perror("mmap() window"); + + /* Prepare one window; record its source PFNs. */ + for (i =3D 0; i < nr_pages; i++) { + p[i * getpagesize()] =3D i + 1; + pfns[i] =3D pagemap_get_pfn(pagemap_fd, p + i * getpagesize()); + if (pfns[i] =3D=3D -1UL) + ksft_exit_fail_msg("Source page not present\n"); + } + + /* Clear first: with the events still off there is nothing to undo. */ + if (tracing_clear_trace()) + ksft_exit_fail_msg("Cannot clear the trace buffer\n"); + if (tracing_events_enable(trace_events_fd, true)) + ksft_exit_fail_msg("Cannot enable huge_memory events\n"); + + madvise(p, hpage_pmd_size, MADV_HUGEPAGE); + /* Wait up to 120 seconds for the pass to complete. */ + passed =3D khugepaged_full_pass(120); + + /* Off before anything that can give up: the events are system-wide. */ + if (tracing_events_enable(trace_events_fd, false)) + ksft_exit_fail_msg("Cannot disable huge_memory events\n"); + if (!passed) + ksft_exit_fail_msg("khugepaged did not complete a full pass\n"); + + collapsed =3D is_range_backed_by_folio_orders(p, window, TARGET_ORDER, + pagemap_fd, kpageflags_fd); + attributed =3D count_attributed(pfns, nr_pages, (unsigned long)p, + TARGET_ORDER); + + ksft_test_result(collapsed && attributed =3D=3D 1, + "step %d: window collapsed, %d attributed result(s)\n", + iteration, attributed); + + munmap(p, hpage_pmd_size); +} + +int main(void) +{ + struct thp_settings settings; + int i; + + ksft_print_header(); + + if (!thp_available()) + ksft_exit_skip("Transparent Hugepages not available\n"); + if (!(thp_supported_orders() & (1UL << TARGET_ORDER))) + ksft_exit_skip("Order %d is not a supported anon THP order\n", + TARGET_ORDER); + + hpage_pmd_size =3D read_pmd_pagesize(); + if (!hpage_pmd_size) + ksft_exit_fail_msg("Reading PMD pagesize failed\n"); + pagemap_fd =3D open("/proc/self/pagemap", O_RDONLY); + if (pagemap_fd < 0) + ksft_exit_fail_perror("open(/proc/self/pagemap)"); + kpageflags_fd =3D open("/proc/kpageflags", O_RDONLY); + if (kpageflags_fd < 0) + ksft_exit_skip("open(\"/proc/kpageflags\") requires root\n"); + trace_events_fd =3D tracing_events_open("huge_memory"); + if (trace_events_fd < 0) + ksft_exit_skip("huge_memory events require tracefs and root\n"); + atexit(trace_events_off); + + ksft_set_plan(NR_ITERATIONS); + + thp_save_settings(); + thp_read_settings(&settings); + settings.thp_enabled =3D THP_MADVISE; + settings.thp_defrag =3D THP_DEFRAG_ALWAYS; + settings.khugepaged.defrag =3D 1; + settings.khugepaged.scan_sleep_millisecs =3D 60000; + settings.khugepaged.alloc_sleep_millisecs =3D 60000; + settings.khugepaged.max_ptes_none =3D (hpage_pmd_size / getpagesize()) - = 1; + /* One wake must complete one full pass; see khugepaged_full_pass(). */ + settings.khugepaged.pages_to_scan =3D 1UL << 24; + for (i =3D 0; i < NR_ORDERS; i++) + settings.hugepages[i].enabled =3D THP_NEVER; + settings.hugepages[TARGET_ORDER].enabled =3D THP_INHERIT; + /* Base of the settings stack; the bottom entry is never popped. */ + thp_push_settings(&settings); + + for (i =3D 0; i < NR_ITERATIONS; i++) + one_step(i); + + thp_restore_settings(); + + ksft_finished(); +} diff --git a/tools/testing/selftests/mm/run_vmtests.sh b/tools/testing/self= tests/mm/run_vmtests.sh index 8bf898b71350..c0f69da3fd3b 100755 --- a/tools/testing/selftests/mm/run_vmtests.sh +++ b/tools/testing/selftests/mm/run_vmtests.sh @@ -404,6 +404,8 @@ CATEGORY=3D"cow" run_test ./cow =20 CATEGORY=3D"thp" run_test ./folio_order_check =20 +CATEGORY=3D"thp" run_test ./khugepaged_sync_check + CATEGORY=3D"thp" run_test ./khugepaged =20 CATEGORY=3D"thp" run_test ./khugepaged -s 2 diff --git a/tools/testing/selftests/mm/vm_util.c b/tools/testing/selftests= /mm/vm_util.c index c9bd6c92fa41..ee1334778391 100644 --- a/tools/testing/selftests/mm/vm_util.c +++ b/tools/testing/selftests/mm/vm_util.c @@ -598,6 +598,47 @@ bool is_range_backed_by_folio_orders(char *start, size= _t len, int order, return true; } =20 +#define TRACEFS_ROOT "/sys/kernel/tracing" + +/* + * Open the enable file of one ftrace event subsystem (e.g. "huge_memory"). + * Returns a descriptor for tracing_events_enable(), or -1 if tracefs or t= he + * subsystem is not there. The events are system-wide state: whoever + * switches them on owns them until it switches them off, including on the + * paths where the test gives up. + */ +int tracing_events_open(const char *subsys) +{ + char path[256]; + + snprintf(path, sizeof(path), TRACEFS_ROOT "/events/%s/enable", + subsys); + return open(path, O_WRONLY); +} + +int tracing_events_enable(int fd, bool enable) +{ + if (pwrite(fd, enable ? "1" : "0", 1, 0) !=3D 1) + return -1; + return 0; +} + +/* Drop what the trace buffer holds so far. */ +int tracing_clear_trace(void) +{ + int fd =3D open(TRACEFS_ROOT "/trace", O_WRONLY | O_TRUNC); + + if (fd < 0) + return -1; + close(fd); + return 0; +} + +FILE *tracing_open_trace(void) +{ + return fopen(TRACEFS_ROOT "/trace", "r"); +} + /* If `ioctls' non-NULL, the allowed ioctls will be returned into the var = */ int uffd_register_with_ioctls(int uffd, void *addr, uint64_t len, bool miss, bool wp, bool minor, uint64_t *ioctls) diff --git a/tools/testing/selftests/mm/vm_util.h b/tools/testing/selftests= /mm/vm_util.h index ce05bce4670d..10c7be46e44c 100644 --- a/tools/testing/selftests/mm/vm_util.h +++ b/tools/testing/selftests/mm/vm_util.h @@ -119,6 +119,10 @@ int close_procmap(struct procmap_fd *procmap); int write_sysfs(const char *file_path, unsigned long val); int read_sysfs(const char *file_path, unsigned long *val); bool softdirty_supported(void); +int tracing_events_open(const char *subsys); +int tracing_events_enable(int fd, bool enable); +int tracing_clear_trace(void); +FILE *tracing_open_trace(void); =20 static inline int open_self_procmap(struct procmap_fd *procmap_out) { --=20 2.54.0 From nobody Tue Sep 29 04:44:47 2026 Received: from fhigh-b1-smtp.messagingengine.com (fhigh-b1-smtp.messagingengine.com [202.12.124.152]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id ECB6D45517F; Wed, 12 Aug 2026 13:23:48 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.152 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541030; cv=none; b=uKN3uFtAdu9hqVCwbIPYUfraQs9a+KRLvuAFmtRbVoKwPS7HM5RRm2eODgVvyGsfDlEI2iPQ+ighvAA0q1g34PKelkDGHPN8QBaALMVZghWX7HALwIHmpvmAUqS+aGAm/3lkj/eLDx5BgD+YYOSFNWqxZs9+opIuYY6y9yjMUJE= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541030; c=relaxed/simple; bh=IxIYmZill7P9IK3f0Umjv0+TTCBiiNwTmbnhLIldafo=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=WfHgCvmzKBmyPr5sEuQzwtkWDbOHYd+W55LKLOHG5WHrdloa1zNIPDn1ejfdcIJKB25l5pqa4z817ZsdUQ6b6SM2nQjxary4r66QzWy9tyZcvnRX1nRqhyh/N3vkv4nm4qan+01hEzwt2M2H2DnSPFfijc0g9loemVphi214Y+8= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=RKFRcnUi; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=aI1phcga; arc=none smtp.client-ip=202.12.124.152 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="RKFRcnUi"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="aI1phcga" Received: from phl-compute-12.internal (phl-compute-12.internal [10.202.2.52]) by mailfhigh.stl.internal (Postfix) with ESMTP id D900A7A0160; Wed, 12 Aug 2026 09:23:47 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-12.internal (MEProxy); Wed, 12 Aug 2026 09:23:48 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786541027; x= 1786627427; bh=9oFfQCqKMm/SRuwfBKl+h3wjprs8HFm3Wg7AwMGv1us=; b=R KFRcnUi44U2l6boPRdaVVvH6gb2KV+xf0LjuR6A1ffFKLu8nw/fxYxTtSWJSYlie FiZ7If8zPxzqQpfr2DUACe8eQpwUTdNtDc7XiN2PzVnc+uw8xgeN8sjcoeYE4N8L Y7Dgoj/Iamgob2Kc0F2zXEzypV7aXumc9AWdDcwNBuKLmv1/KwAh/E4yfJZsh9lL q0tBT1drLH96YfJDtxqqJ5zc4/gcslYKLibKSlXI6jOSUL3NQEttZIqo60Tb0c1f Lo/WedLO7e7YrnJQ6zH/uFTukJ64p8uL20AZEel5b6paepn3o1TsCDfqnZs5kIQV YLB7j+eKLRfYZNkGleruA== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786541027; x=1786627427; bh=9 oFfQCqKMm/SRuwfBKl+h3wjprs8HFm3Wg7AwMGv1us=; b=aI1phcgaXMW/AnGfn JbQgufLsPj3DeZh6L4SZPJR7c4gDi9Yx4G4/QO1xaGKXM9TkYnWUlnEEjjk/3m87 Yw0TO6A890J1wRu6KFOhEeTcOhdxRF2h7pcbqE+4s/a4TZ+PYC4KtkrfLqhuWiSw PU7ZlPxCXtRk8je2xcdAEtkCGtiy+U8Q2kjmpXPlzNhJKQDZ6UzF+qjwWi+mEbTB TVlnjx2peDfn9nLlHicKj8bT3P6RqJcKlNDjcrfJeNwGoJaMyTBFFyTdDseZeezb eQRaAP+U4fKNRy8oT5eSSc8eYtDT06nZLUgu1lActaxI8D8yRA+WTW5YxRNHbnZ/ 54I4A== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTEWZPjKBYDoSh3uCuJ0ogiZvSuDj1RAgMuSmimalRft9LCPvZxyrluChLfTbkjo2r KkCK8Ttm4mqxi+W2/UjRoyI4gRuaMCGjIyoE2sYafZhhTXl0iIGCKEOr9/WVtZfHPc/pjz XsLpzK8U2Xu6F30SrJOujIlp1bmlomKUtpU/iiXSkRy6wwaI7NV8wAm1PQ0lHgRG0nCRKX WKIh8TtevEly/nhHwG2deu/klJhGkXN7h/CDYOOyqjOIkeF7MAFsC6fZtQDVeVkgzM4xU1 Py5NrFUehZdXXi5LHUYxqNjNVNbuxWN2P2YJW6J2IL8s+zlVHHPN2YdKOlFVuMcbnWK9n+ gMTk4WgLkBkenumpDQFUN9Fx+TeJoxCVG8CZRQgGqQpM7yxSCigP0CZHQrFni2NmMVhFzF JwcWfDUaC35gcJW/Ysr7QnWcKi3pdMoQYao4BZjPZ50t0KQOHbOEPsfMSSYHKYo0Fk1MfR lE4JRLi0km31O2Lv0PQShhLHynfRCsVV2OLxchoXEc/di5G1kF+MFl4nHBgvX1SwA+k9s+ GAKiI8swXWbzhXBnl94tJXohsuStdSy8btM1QyPy8yC+AuNWjaN83+OmzMu5d7wlgWnyLs tXR/l8tpjQJGRgl1ywR13NuLVtkirnksonjoac6p7DmdB5podVjwJUMMb7tg X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Wed, 12 Aug 2026 09:23:46 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v3 15/18] selftests/mm: add khugepaged race harness Date: Wed, 12 Aug 2026 14:23:01 +0100 Message-ID: <20260812132304.199287-16-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260812132304.199287-1-kirill@shutemov.name> References: <20260812132304.199287-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" Collapse serialises against faults, GUP, fork, mremap and zapping through a protocol of locks, TLB flushes and refcount checks. No khugepaged selftest exercises any of it under contention. Add khugepaged_race. Two faulters, an MADV_DONTNEED thread, a transient FOLL_PIN thread (gup_test), a forker and an mremap thread work the same ranges while one of three drivers collapses: stepped khugepaged, one full pass at a time via khugepaged_full_pass(), so each step covers a known extent; free khugepaged left to run (scan_sleep_millisecs=3D0), for soak; madvise an MADV_COLLAPSE and MADV_DONTNEED loop. Every mode runs in turn unless -m names one, five seconds each. All anon THP orders are enabled as inherit and max_ptes_none is 0, so a window collapses only once fully populated and the racing MADV_DONTNEED steers selection across orders. The rule is that a racing page reads as its pattern or as zero, never anything else. The faulters and fork children check it throughout, and a final sweep checks it again. The kernel's own assertions -- DEBUG_VM, page_table_check, KASAN, lockdep -- are the other half of the oracle, so read dmesg too. The threads share three PMD-sized areas plus one for the mremap thread; -a sets the count, since at a 512M PMD the default is several gigabytes. -d sets the soak length. Assisted-by: Claude-Code:claude-opus-5 Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/Makefile | 1 + tools/testing/selftests/mm/khugepaged_race.c | 441 +++++++++++++++++++ tools/testing/selftests/mm/run_vmtests.sh | 2 + 3 files changed, 444 insertions(+) create mode 100644 tools/testing/selftests/mm/khugepaged_race.c diff --git a/tools/testing/selftests/mm/Makefile b/tools/testing/selftests/= mm/Makefile index b2d6e5c12934..308bbad73c11 100644 --- a/tools/testing/selftests/mm/Makefile +++ b/tools/testing/selftests/mm/Makefile @@ -106,6 +106,7 @@ TEST_GEN_FILES +=3D rmap TEST_GEN_FILES +=3D folio_split_race_test TEST_GEN_FILES +=3D folio_order_check TEST_GEN_FILES +=3D khugepaged_sync_check +TEST_GEN_FILES +=3D khugepaged_race =20 ifneq ($(ARCH),arm64) TEST_GEN_FILES +=3D soft-dirty diff --git a/tools/testing/selftests/mm/khugepaged_race.c b/tools/testing/s= elftests/mm/khugepaged_race.c new file mode 100644 index 000000000000..d9e421ede6de --- /dev/null +++ b/tools/testing/selftests/mm/khugepaged_race.c @@ -0,0 +1,441 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * khugepaged race harness. + * + * Runs collapse against concurrent faults, transient GUP pins + * (gup_test), fork, mremap and MADV_DONTNEED over the same ranges, in + * one of three driver modes: + * + * stepped khugepaged, one full pass at a time through + * khugepaged_full_pass(), so a step covers a known extent; + * free khugepaged left to run (scan_sleep_millisecs=3D0), for soak; + * madvise MADV_COLLAPSE in a loop. + * + * All anon THP orders are enabled (inherit) and max_ptes_none is 0, so a + * window has to be fully populated before khugepaged will collapse it, and + * the racing MADV_DONTNEED decides which orders it can still use. + * + * Correctness signals: every racing page must read as its pattern or + * zero (MADV_DONTNEED), never anything else. The faulters and the fork + * children check that continuously, a final sweep checks it once more, pl= us + * whatever DEBUG_VM / page_table_check / KASAN / lockdep report in + * dmesg, which the caller is expected to inspect. + */ +#define _GNU_SOURCE +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include "kselftest.h" +#include "vm_util.h" +#include "hugepage_settings.h" +#include "../../../../mm/gup_test.h" + +#define BASE_ADDR ((void *)(1UL << 30)) + +/* + * Shared playground for faults/pins/fork/dontneed: several PMD-sized + * areas the racing threads spread across, plus one area owned by the + * mremap thread. More areas means more independent regions collapsing + * at once; the default suits a normal machine. On a memory-constrained + * host -- or under emulation, where a 512M PMD (arm64/64K) makes the + * default playground multi-gigabyte -- pass -a to shrink it. + */ +#define DEFAULT_SHARED_AREAS 3 +static int nr_shared_areas; +static int nr_areas; + +static unsigned long hpage_pmd_size; +static unsigned long page_size; +static char *region; /* NR_AREAS * hpage_pmd_size */ +static char *mremap_area; /* region + NR_SHARED_AREAS areas */ +static char *mremap_scratch; /* well above the region */ +static int gup_fd =3D -1; +static volatile int stop; +static volatile int corrupted; + +static unsigned int pattern(unsigned long page_idx) +{ + unsigned int val =3D (unsigned int)page_idx * 2654435761U; + + return val ? val : 1; /* never collides with the zero-fill */ +} + +/* Zero means never written; anything else must be this page's pattern */ +static bool page_is_corrupt(unsigned long page_idx, unsigned int *val) +{ + *val =3D *(unsigned int *)(region + page_idx * page_size); + + return *val && *val !=3D pattern(page_idx); +} + +static void check_page(unsigned long page_idx) +{ + unsigned int val; + + if (page_is_corrupt(page_idx, &val)) { + corrupted =3D 1; + ksft_print_msg("Corruption at page %lu: %#x !=3D %#x\n", + page_idx, val, pattern(page_idx)); + } +} + +static unsigned long shared_pages(void) +{ + return nr_shared_areas * hpage_pmd_size / page_size; +} + +static unsigned long rand_page(unsigned int *seed) +{ + return (unsigned long)rand_r(seed) % shared_pages(); +} + +/* Pages left from @page_idx, so a range never reaches the mremap thread's= area */ +static unsigned long room_from(unsigned long page_idx, unsigned long want) +{ + unsigned long left =3D shared_pages() - page_idx; + + return want < left ? want : left; +} + +static void *faulter_fn(void *arg) +{ + unsigned int seed =3D (unsigned long)arg; + + while (!stop) { + unsigned long page_idx =3D rand_page(&seed); + + if (rand_r(&seed) & 1) + *(unsigned int *)(region + page_idx * page_size) =3D + pattern(page_idx); + else + check_page(page_idx); + } + return NULL; +} + +static void *dontneed_fn(void *arg) +{ + unsigned int seed =3D (unsigned long)arg; + + while (!stop) { + unsigned long page_idx =3D rand_page(&seed); + unsigned long nr =3D 1UL << (rand_r(&seed) % 6); /* 1..32 pages */ + + madvise(region + page_idx * page_size, + room_from(page_idx, nr) * page_size, MADV_DONTNEED); + usleep(rand_r(&seed) % 500); + } + return NULL; +} + +static void *pinner_fn(void *arg) +{ + unsigned int seed =3D (unsigned long)arg; + + while (!stop) { + struct gup_test gup =3D {}; + unsigned long page_idx =3D rand_page(&seed); + + unsigned long nr =3D room_from(page_idx, 16); + + gup.addr =3D (unsigned long)(region + page_idx * page_size); + gup.size =3D nr * page_size; + gup.nr_pages_per_call =3D nr; + gup.gup_flags =3D 1; /* FOLL_WRITE */ + /* Racing MADV_DONTNEED makes transient failures expected. */ + ioctl(gup_fd, PIN_FAST_BENCHMARK, &gup); + usleep(rand_r(&seed) % 200); + } + return NULL; +} + +static void *forker_fn(void *arg) +{ + unsigned int seed =3D (unsigned long)arg; + + while (!stop) { + pid_t pid =3D fork(); + + if (pid =3D=3D 0) { + unsigned int val; + int bad =3D 0; + + /* + * No stdio in the child: a thread may have held + * stdout's lock when we forked, and printing under an + * inherited lock hangs. The parent reports what the + * exit status says. + */ + for (int i =3D 0; i < 16; i++) + bad |=3D page_is_corrupt(rand_page(&seed), &val); + _exit(bad); + } + if (pid > 0) { + int wstatus; + + if (waitpid(pid, &wstatus, 0) < 0) + ksft_exit_fail_perror("waitpid()"); + /* A child killed on the read counts too, not just its exit code. */ + if (!WIFEXITED(wstatus) || WEXITSTATUS(wstatus)) + corrupted =3D 1; + } + usleep(rand_r(&seed) % 2000); + } + return NULL; +} + +static void *mremapper_fn(void *arg) +{ + unsigned int seed =3D (unsigned long)arg; + + while (!stop) { + void *p; + + p =3D mremap(mremap_area, hpage_pmd_size, hpage_pmd_size, + MREMAP_MAYMOVE | MREMAP_FIXED, mremap_scratch); + if (p =3D=3D MAP_FAILED) + ksft_exit_fail_perror("mremap() away"); + for (int i =3D 0; i < 8; i++) + mremap_scratch[(rand_r(&seed) % + (hpage_pmd_size / page_size)) * page_size] =3D 1; + p =3D mremap(mremap_scratch, hpage_pmd_size, hpage_pmd_size, + MREMAP_MAYMOVE | MREMAP_FIXED, mremap_area); + if (p =3D=3D MAP_FAILED) + ksft_exit_fail_perror("mremap() back"); + usleep(rand_r(&seed) % 2000); + } + return NULL; +} + +static unsigned long now_ms(void) +{ + struct timeval tv; + + gettimeofday(&tv, NULL); + return tv.tv_sec * 1000UL + tv.tv_usec / 1000; +} + +static void usage(void) +{ + fprintf(stderr, + "Usage: khugepaged_race [-d seconds] [-m stepped|free|madvise] [-a areas= ]\n" + "\tWithout -m, every mode runs in turn.\n" + "\t-d: seconds per mode (default 5)\n" + "\t-a: number of shared PMD-sized playground areas (default 3)\n"); + exit(1); +} + +int main(int argc, char **argv) +{ + static const char * const thread_names[] =3D { + "faulter", "faulter2", "dontneed", "pinner", "forker", + "mremapper", + }; + void *(*const thread_fns[])(void *) =3D { + faulter_fn, faulter_fn, dontneed_fn, pinner_fn, forker_fn, + mremapper_fn, + }; + const int nr_threads =3D ARRAY_SIZE(thread_names); + pthread_t threads[ARRAY_SIZE(thread_names)]; + static const char * const all_modes[] =3D { "stepped", "free", "madvise" = }; + const char *one_mode[1]; + const char * const *modes =3D all_modes; + int nr_modes =3D ARRAY_SIZE(all_modes); + const char *mode_arg =3D NULL; + struct thp_settings settings; + unsigned long end_ms; + int duration_s =3D 5; + unsigned long thread_mask =3D ~0UL; + int nr_areas_arg =3D 0; + unsigned long i; + int steps =3D 0; + int opt; + + while ((opt =3D getopt(argc, argv, "a:d:m:t:h")) !=3D -1) { + switch (opt) { + case 'a': + nr_areas_arg =3D atoi(optarg); + break; + case 'd': + duration_s =3D atoi(optarg); + break; + case 'm': + mode_arg =3D optarg; + break; + case 't': + /* debug: bitmask of racing threads to start */ + thread_mask =3D strtoul(optarg, NULL, 0); + break; + default: + usage(); + } + } + if (mode_arg) { + if (strcmp(mode_arg, "stepped") && strcmp(mode_arg, "free") && + strcmp(mode_arg, "madvise")) + usage(); + one_mode[0] =3D mode_arg; + modes =3D one_mode; + nr_modes =3D 1; + } + + ksft_print_header(); + if (!thp_available()) + ksft_exit_skip("Transparent Hugepages not available\n"); + + page_size =3D getpagesize(); + hpage_pmd_size =3D read_pmd_pagesize(); + if (!hpage_pmd_size) + ksft_exit_fail_msg("Reading PMD pagesize failed\n"); + + gup_fd =3D open("/sys/kernel/debug/gup_test", O_RDWR); + if (gup_fd < 0) + ksft_exit_skip("/sys/kernel/debug/gup_test requires CONFIG_GUP_TEST and = root\n"); + + nr_shared_areas =3D nr_areas_arg > 0 ? nr_areas_arg : DEFAULT_SHARED_AREA= S; + nr_areas =3D nr_shared_areas + 1; + + /* + * The mremap thread moves its area to this address and back, and + * MREMAP_FIXED unmaps whatever is in the way without saying so. Claim + * the address here, so a layout that does not match this assumption + * fails now instead of losing a mapping later. Nothing else in the + * process maps this low: thread stacks and malloc arenas come from the + * top-down mmap area, well above. + */ + mremap_scratch =3D (char *)BASE_ADDR + 2 * nr_areas * hpage_pmd_size; + if (mmap(mremap_scratch, hpage_pmd_size, PROT_NONE, + MAP_ANONYMOUS | MAP_PRIVATE | MAP_FIXED_NOREPLACE, + -1, 0) !=3D (void *)mremap_scratch) + ksft_exit_fail_perror("mmap() mremap scratch"); + + ksft_set_plan(nr_modes); + + thp_save_settings(); + thp_read_settings(&settings); + + /* + * A base entry for the stack, so that the pop at the end of a mode + * always has something to write back: thp_pop_settings() on an empty + * stack has no settings to apply and gives up. + */ + thp_push_settings(&settings); + + for (int m =3D 0; m < nr_modes; m++) { + const char *mode =3D modes[m]; + + thp_read_settings(&settings); + settings.thp_enabled =3D THP_MADVISE; + settings.thp_defrag =3D THP_DEFRAG_ALWAYS; + settings.shmem_enabled =3D SHMEM_NEVER; + settings.khugepaged.defrag =3D 1; + settings.khugepaged.scan_sleep_millisecs =3D + strcmp(mode, "free") ? 1000 : 0; + settings.khugepaged.alloc_sleep_millisecs =3D 10; + /* + * Strict occupancy: mTHP collapse only supports 0 or + * HPAGE_PMD_NR - 1 and coerces anything else to 0 anyway, and 0 + * also keeps khugepaged from burning the whole step in doomed + * PMD-sized allocations on 512M-PMD configs: under racing + * MADV_DONTNEED a fully populated PMD area is rare. + */ + settings.khugepaged.max_ptes_none =3D 0; + settings.khugepaged.pages_to_scan =3D + nr_areas * (hpage_pmd_size / page_size) * 8; + for (i =3D 0; i < NR_ORDERS; i++) { + if (thp_supported_orders() & (1UL << i)) + settings.hugepages[i].enabled =3D THP_INHERIT; + } + /* Popped at the end of this mode, before the next one. */ + thp_push_settings(&settings); + + region =3D mmap(BASE_ADDR, nr_areas * hpage_pmd_size, + PROT_READ | PROT_WRITE, MAP_ANONYMOUS | + MAP_PRIVATE | MAP_FIXED_NOREPLACE, -1, 0); + if (region !=3D BASE_ADDR) + ksft_exit_fail_perror("mmap() playground"); + mremap_area =3D region + nr_shared_areas * hpage_pmd_size; + + /* Populate so the first pass has something to collapse. */ + for (i =3D 0; i < nr_shared_areas * hpage_pmd_size / page_size; i++) + *(unsigned int *)(region + i * page_size) =3D pattern(i); + memset(mremap_area, 1, hpage_pmd_size); + madvise(region, nr_areas * hpage_pmd_size, MADV_HUGEPAGE); + + for (i =3D 0; i < nr_threads; i++) { + if (!(thread_mask & (1UL << i))) { + threads[i] =3D 0; + continue; + } + if (pthread_create(&threads[i], NULL, thread_fns[i], + (void *)(i + 1))) + ksft_exit_fail_perror("pthread_create()"); + } + + end_ms =3D now_ms() + duration_s * 1000UL; + if (!strcmp(mode, "stepped")) { + while (now_ms() < end_ms && !corrupted) { + if (!khugepaged_full_pass(600)) + ksft_exit_fail_msg("khugepaged pass timed out\n"); + steps++; + } + } else if (!strcmp(mode, "free")) { + while (now_ms() < end_ms && !corrupted) + usleep(100 * 1000); + } else { /* madvise */ + while (now_ms() < end_ms && !corrupted) { + for (i =3D 0; i < nr_shared_areas; i++) { + madvise(region + i * hpage_pmd_size, + hpage_pmd_size, MADV_COLLAPSE); + } + madvise(region, nr_shared_areas * hpage_pmd_size, + MADV_DONTNEED); + steps++; + } + } + + stop =3D 1; + for (i =3D 0; i < nr_threads; i++) { + if (threads[i]) + pthread_join(threads[i], NULL); + } + + /* Final integrity sweep. */ + for (i =3D 0; i < nr_shared_areas * hpage_pmd_size / page_size; i++) + check_page(i); + + ksft_test_result(!corrupted, + "%s: %ds, %d steps, no corruption\n", + mode, duration_s, steps); + + /* + * Hand the address space and the settings back before the + * next mode: it maps the region at the same fixed address, + * and its scan cadence differs. + */ + munmap(region, nr_areas * hpage_pmd_size); + thp_pop_settings(); + stop =3D 0; + steps =3D 0; + + if (corrupted) { + /* Memory is suspect; the rest would prove nothing. */ + while (++m < nr_modes) + ksft_test_result_skip("%s: skipped after corruption\n", + modes[m]); + break; + } + } + + thp_restore_settings(); + ksft_finished(); +} diff --git a/tools/testing/selftests/mm/run_vmtests.sh b/tools/testing/self= tests/mm/run_vmtests.sh index c0f69da3fd3b..fc61907aa3b2 100755 --- a/tools/testing/selftests/mm/run_vmtests.sh +++ b/tools/testing/selftests/mm/run_vmtests.sh @@ -406,6 +406,8 @@ CATEGORY=3D"thp" run_test ./folio_order_check =20 CATEGORY=3D"thp" run_test ./khugepaged_sync_check =20 +CATEGORY=3D"thp" run_test ./khugepaged_race + CATEGORY=3D"thp" run_test ./khugepaged =20 CATEGORY=3D"thp" run_test ./khugepaged -s 2 --=20 2.54.0 From nobody Tue Sep 29 04:44:47 2026 Received: from fout-b7-smtp.messagingengine.com (fout-b7-smtp.messagingengine.com [202.12.124.150]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id F2FC045DF6C; Wed, 12 Aug 2026 13:23:51 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.150 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541033; cv=none; b=ZSkwhC3F1WcoZHtbppJWv7zwOBNG7gFI8r1OwfaRwCcdnlNdwGQJgzow/46L3p2U1dFXrK2qvbp30B+s7o4k8ou8CQ8loxz+BxPTKqdbC4YBxeUM0L2k3YM75g7DFmrVi3kt9XHwoT0xtZRMsCQaEe+N4uIFnswUcfiFmyAmIUk= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541033; c=relaxed/simple; bh=M0jQLbu3ZUjWnCXL3Xv44CUXusQeA4EgoDT5i+Id1gw=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=G96XaMuOqJt9e+rWPJTRBSvLLStyBZKHgN9LXI/SCQAWYHLQiuq/LOtHEWlCU3ESm4VDr2twjyU34LRo8mPmF8r0kgIPdAfzSdlU/TuORqrDDKQRxmlRds2iBSnziiBSWjeEJZwrKfN8d20prAqpDDr2j5jErnl4AqrcTxet3M4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=S9H1VZRq; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=WtsbVfUF; arc=none smtp.client-ip=202.12.124.150 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="S9H1VZRq"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="WtsbVfUF" Received: from phl-compute-09.internal (phl-compute-09.internal [10.202.2.49]) by mailfout.stl.internal (Postfix) with ESMTP id ACA201D0016F; Wed, 12 Aug 2026 09:23:50 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-09.internal (MEProxy); Wed, 12 Aug 2026 09:23:51 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786541030; x= 1786627430; bh=ymEJqkwUinB+T+KKBklAPXcC2k5sBcIeWFiKz2tJ1JM=; b=S 9H1VZRqLFa3pXg6pF3HLc29eLjfAHk3GKk60lm5wFJqrDUCvvedrIFbZ59UPmLae 87oOL5Vf9c8+IbyITmpBQPTfLW14RFaqI9HPcWrNwsbm6DcWd4iZNk3BOdl3ugr7 G89EertJStnBDXm5fmGDfzahPV5aBuV5wT9+KD2MoMGtT85rJpoptRp5DP+Fsbvt /SslAXJdk4AZgSOJ39X1Nzmn5I3xMaOv8B56uU4ikUSx3SRM3f3fa4tdaYdN1X2s /Y7VhaSM1n0g8JYjLZgTxKmNPrf0kNUYJzsnbgBm82ZBYpBss3UD1FaDlLDj88LX Pqd+hcwQU8Qg+5UVeXO6A== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786541030; x=1786627430; bh=y mEJqkwUinB+T+KKBklAPXcC2k5sBcIeWFiKz2tJ1JM=; b=WtsbVfUFoSwajttct pMFE3KBG6bt45uwtQsVtE8TBordu7tf5RXfr/pm4YBBrJfNLtXha71jlu2/Kt1Y5 2zAY0V+fkQtp2A3CG715rykCYdvnOPLG8fd98Q4lxnvZv2bXU4pkx6VB+gY/BvVy x3FqBjeZrL1VsdLsHK34vReQk6vkGPniZ5u6CktJabhN889JdpUSZNg68AEnKnIB zIDvI/XzBo7fOYWXiac9HvOsLFl42TgUKdxH9Y22iJBcL1y0X0gM6ib2yoJY018J y9PU+W+GlQ6XkIeclt2b9w9NXrZCsUztWc735CII5X3YZyENbHBIRrWpSO2Cfobp YouTw== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTGK88F9UMnXZ5YHCfZ+3CEyhDPNn2cwpG7gczg1QtUW3+KH6lRiOTrWgwtmC8SwM8 lS9N2AAOkbX57R9oIGEGX/Y22a2ZgCAGYteUW+bA41LZ08JSUTSnxLP4aLeGJ+l2igRpiI z056Op0LI1NxVlka48nXFKSS8a/uwNZK+39708VFH13YHS1sNHrX5zvuU5299hb11Ao5H/ zQrZEfdZM4G7c0rNyXZ30syjB6Y8P8LTAhLTAHwdc8i3T8CZWzvPOusg5fBb0jumhRSeE4 e5clat+a+U2DXBe0wlr1aBcOwBQzkKo/UFV6TzMBSg5j9cF+ay7ufreZHSHrXPVD1XvoiK 0gdJKI8U7fbIGj0HZXI1c4pGX6rHFYyG2kcmUon0bnMEum6bTcKdn+AS3ZN6sQ9l8HpyGR Jhf86xrJtorxgBIdnEOe8djhD4Akbuh/xMBZMkPCnoqkfkdePxuB66oPTKp9nmhn19rafD teQichvbgkO21ko9xW+rsH+2LOyLHer3IyunhZ+cNF72Q4tlbZciZOu9NCGFXA7jKAMPiF G88p6TycyVjRW11TsMM8hRfdyIYSNw/Sb5Lpcs9wgmcoi/OmzTDrV+D/6li9grvIytJpXJ 8saFUo2wylxtQeciGVI7lp33+rMUZ0kMCXnkkjS5RDMGLx+T+ptToq8ZTwnQ X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Wed, 12 Aug 2026 09:23:49 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v3 16/18] selftests/mm: race collapse of windows with holes Date: Wed, 12 Aug 2026 14:23:02 +0100 Message-ID: <20260812132304.199287-17-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260812132304.199287-1-kirill@shutemov.name> References: <20260812132304.199287-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The harness pins max_ptes_none to 0, so khugepaged only collapses a window once every PTE in it is present. Collapsing a window that has holes never happens, and that is a different path: a hole is not copied from anywhere but zero-filled into the new folio, and the slot is re-checked under the page table lock at install time in case a racing fault filled it first. Run both ends of the occupancy scale, one after the other, for every driver mode -- mTHP collapse supports only those two, 0 and HPAGE_PMD_NR - 1, and coerces anything between them to 0. Each result says which end it ran: ok 1 stepped/strict: 5s, 231 steps, no corruption ok 2 stepped/holes: 5s, 194 steps, no corruption -z narrows a run to the hole-heavy end, the way -m narrows it to one mode. Assisted-by: Claude-Code:claude-opus-5 Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/khugepaged_race.c | 61 ++++++++++++++------ 1 file changed, 42 insertions(+), 19 deletions(-) diff --git a/tools/testing/selftests/mm/khugepaged_race.c b/tools/testing/s= elftests/mm/khugepaged_race.c index d9e421ede6de..ba4f5c36bf88 100644 --- a/tools/testing/selftests/mm/khugepaged_race.c +++ b/tools/testing/selftests/mm/khugepaged_race.c @@ -11,9 +11,12 @@ * free khugepaged left to run (scan_sleep_millisecs=3D0), for soak; * madvise MADV_COLLAPSE in a loop. * - * All anon THP orders are enabled (inherit) and max_ptes_none is 0, so a - * window has to be fully populated before khugepaged will collapse it, and - * the racing MADV_DONTNEED decides which orders it can still use. + * All anon THP orders are enabled (inherit). Occupancy runs at both ends + * of what mTHP collapse supports: max_ptes_none 0, where a window must be + * fully populated, and HPAGE_PMD_NR - 1, where a window full of holes + * collapses too. The holes are not copied from anywhere -- they are + * zero-filled, and re-checked under the page table lock at install time in + * case a racing fault got there first. * * Correctness signals: every racing page must read as its pattern or * zero (MADV_DONTNEED), never anything else. The faulters and the fork @@ -227,9 +230,11 @@ static unsigned long now_ms(void) static void usage(void) { fprintf(stderr, - "Usage: khugepaged_race [-d seconds] [-m stepped|free|madvise] [-a areas= ]\n" + "Usage: khugepaged_race [-d seconds] [-m stepped|free|madvise] [-z] [-a = areas]\n" "\tWithout -m, every mode runs in turn.\n" "\t-d: seconds per mode (default 5)\n" + "\tBoth occupancy limits run unless -z asks for holes only.\n" + "\t-z: only max_ptes_none =3D HPAGE_PMD_NR - 1 (hole-heavy)\n" "\t-a: number of shared PMD-sized playground areas (default 3)\n"); exit(1); } @@ -247,6 +252,9 @@ int main(int argc, char **argv) const int nr_threads =3D ARRAY_SIZE(thread_names); pthread_t threads[ARRAY_SIZE(thread_names)]; static const char * const all_modes[] =3D { "stepped", "free", "madvise" = }; + static const int all_nones[] =3D { 0, 1 }; /* strict, holes */ + const int *nones =3D all_nones; + int nr_nones =3D ARRAY_SIZE(all_nones); const char *one_mode[1]; const char * const *modes =3D all_modes; int nr_modes =3D ARRAY_SIZE(all_modes); @@ -256,11 +264,12 @@ int main(int argc, char **argv) int duration_s =3D 5; unsigned long thread_mask =3D ~0UL; int nr_areas_arg =3D 0; + bool holes_only =3D false; unsigned long i; int steps =3D 0; int opt; =20 - while ((opt =3D getopt(argc, argv, "a:d:m:t:h")) !=3D -1) { + while ((opt =3D getopt(argc, argv, "a:d:m:t:zh")) !=3D -1) { switch (opt) { case 'a': nr_areas_arg =3D atoi(optarg); @@ -275,10 +284,18 @@ int main(int argc, char **argv) /* debug: bitmask of racing threads to start */ thread_mask =3D strtoul(optarg, NULL, 0); break; + case 'z': + holes_only =3D true; + break; default: usage(); } } + if (holes_only) { + nones =3D all_nones + 1; + nr_nones =3D 1; + } + if (mode_arg) { if (strcmp(mode_arg, "stepped") && strcmp(mode_arg, "free") && strcmp(mode_arg, "madvise")) @@ -318,7 +335,7 @@ int main(int argc, char **argv) -1, 0) !=3D (void *)mremap_scratch) ksft_exit_fail_perror("mmap() mremap scratch"); =20 - ksft_set_plan(nr_modes); + ksft_set_plan(nr_modes * nr_nones); =20 thp_save_settings(); thp_read_settings(&settings); @@ -330,8 +347,9 @@ int main(int argc, char **argv) */ thp_push_settings(&settings); =20 - for (int m =3D 0; m < nr_modes; m++) { - const char *mode =3D modes[m]; + for (int mn =3D 0; mn < nr_modes * nr_nones; mn++) { + const char *mode =3D modes[mn / nr_nones]; + bool holes =3D nones[mn % nr_nones]; =20 thp_read_settings(&settings); settings.thp_enabled =3D THP_MADVISE; @@ -341,14 +359,16 @@ int main(int argc, char **argv) settings.khugepaged.scan_sleep_millisecs =3D strcmp(mode, "free") ? 1000 : 0; settings.khugepaged.alloc_sleep_millisecs =3D 10; + /* - * Strict occupancy: mTHP collapse only supports 0 or - * HPAGE_PMD_NR - 1 and coerces anything else to 0 anyway, and 0 - * also keeps khugepaged from burning the whole step in doomed - * PMD-sized allocations on 512M-PMD configs: under racing - * MADV_DONTNEED a fully populated PMD area is rare. + * mTHP collapse only supports the two ends of the occupancy + * scale: 0 or HPAGE_PMD_NR - 1 (anything else coerces to 0). + * Strict needs a fully populated window, which is rare under + * racing MADV_DONTNEED; hole-heavy windows collapse instead, + * so the two ends race different paths. */ - settings.khugepaged.max_ptes_none =3D 0; + settings.khugepaged.max_ptes_none =3D holes ? + (hpage_pmd_size / page_size) - 1 : 0; settings.khugepaged.pages_to_scan =3D nr_areas * (hpage_pmd_size / page_size) * 8; for (i =3D 0; i < NR_ORDERS; i++) { @@ -414,8 +434,9 @@ int main(int argc, char **argv) check_page(i); =20 ksft_test_result(!corrupted, - "%s: %ds, %d steps, no corruption\n", - mode, duration_s, steps); + "%s/%s: %ds, %d steps, no corruption\n", + mode, holes ? "holes" : "strict", + duration_s, steps); =20 /* * Hand the address space and the settings back before the @@ -429,9 +450,11 @@ int main(int argc, char **argv) =20 if (corrupted) { /* Memory is suspect; the rest would prove nothing. */ - while (++m < nr_modes) - ksft_test_result_skip("%s: skipped after corruption\n", - modes[m]); + while (++mn < nr_modes * nr_nones) + ksft_test_result_skip("%s/%s: skipped after corruption\n", + modes[mn / nr_nones], + nones[mn % nr_nones] ? + "holes" : "strict"); break; } } --=20 2.54.0 From nobody Tue Sep 29 04:44:47 2026 Received: from fhigh-b1-smtp.messagingengine.com (fhigh-b1-smtp.messagingengine.com [202.12.124.152]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 2545545FFC2; Wed, 12 Aug 2026 13:23:54 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.152 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541036; cv=none; b=KZbpvi/ybpS+AokxL13on4U9HwnN4kEiTVvY2AGazb4GZ1REq5XE9CwE6Fm/4MiyqGD4Hr72apNYNNzmw2Rx9D9GgsQsrftVvlrVebiBbhFVEWN9UFNVPKyuA5R8qDQoiSnfyvCeLJb8ErAB4V9qD03QzHuxLG+sS7DxMUubJ8o= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541036; c=relaxed/simple; bh=oLj6va39kYAJ9dSmxcl2rZDHnUkr22chuoGH1tAyi6s=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=AvMIOXXg2e/jffVcOrTV2PoTKgPLbMZFW9b+1+TY5YvML1ACSDi8IDr2SH55lJd3/eJJf5i7cawptWZAUtNeS9NFMAf2AocEM5WHNUADX7zSFPij+97Yrsd+i9BREjdkNcVpBijdM6kTrxFX52CtvuwSXH0JcFL9652JnDtNd+4= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=QQGr3i+Y; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=XAdy8p4P; arc=none smtp.client-ip=202.12.124.152 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="QQGr3i+Y"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="XAdy8p4P" Received: from phl-compute-06.internal (phl-compute-06.internal [10.202.2.46]) by mailfhigh.stl.internal (Postfix) with ESMTP id 1B1807A0173; Wed, 12 Aug 2026 09:23:53 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-06.internal (MEProxy); Wed, 12 Aug 2026 09:23:53 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786541032; x= 1786627432; bh=N7LCReNWni5AyzqrKHBrxiaeJCnTTFdJVGlYhEXGLCk=; b=Q QGr3i+YcqmaXLYc1+qBt5+N9VHD/fm8L+kkEalfLzOVAkuOJuPtww7QclBkQ3o98 k6N9h+JaQH+fdqWjVQrLFA/xDdeF+H5khW1qnUrhUzHuIj7mJdXyl111b4zloiOu Fkfa9/UTSgOGd50i1H6k1LsgNUYyGwU+Lt64ZofWUa3JfjEOnkZvE7Qonle6CDHe 1hqu7edgD5Fe9lSMB6POY2cDDCDzBT0/fKrb7bq3oFE4zlxhddegz3u4MTR+ZeEg /iAwY4A4IifPM0GGa70qZb8WD5T0tKPsuG8r97BSInL6uG0D6gxSDB1WzdqH/ugr dO2JF3M6raqcE7fpX29tA== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786541032; x=1786627432; bh=N 7LCReNWni5AyzqrKHBrxiaeJCnTTFdJVGlYhEXGLCk=; b=XAdy8p4PcD9p7T995 LyY9WetgpWpoaAuifi/iSEG8knfe6Ey7IH4oIOl3JbLHlNNsLJYlSyUtl/zCHtIC ae8wbn7O6Llqq5V5JDwKu0xfoKG0z5f38ZdbKD+D7IAmDv90NxZ7znPaDPEN8KBN WncYC8vqiFEd46eQKArCjY7wE9DLNATuh1lTnGfLB6p09oybXgj7k9W6sojbdvnt TNLx9Nq56SMuaz0MNAaHJzpj3HY9ziTFGWo2Bj+NcFoA4OeZsCwR6I/CB6jW8LRN 8a+7E9qDMZJHpZRa/zFjaRNW91pcbooJtzwptBlBdUD468xagEObyfBYtLdKJE1U 3GQJw== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFkdXlhB1yyyGZX3J7dnv16PFzu5W734vHfSlZD/y8c/c/hfyf9JcTqEblNR0Fovw r+VLRtLwlZT2tYuTqMJsjjjBxJMgoAM7D6pmLt36lViB2yT50Uubj7I3KQF8rzeVMBK9wq rm57aKmY5NwWN7q607eE1gnukbthccXm6lhrmbDGqX6XovHG8JRMLfqJj+jBW3X680r35h P9o/w0aRWYBwLgRKYH9oTyb2lbiE5fXNRZyQYMHMmm7yuduJ70EU4YPdENzrqcZn4mtIUi PPdmMfUKqzUyYVrtf6UgPYBrZU3VzLkrgoUPTsKVKEpiHQHdj45X731bFwNLlQpPGy3Qck 3rH4D3RYGs85EQdsgI+nO0kGwXnp1+YPBPQEZ674qFSOTpHnJMdRTrEnca3M3nqPQDMI8/ ot33MuJyjdjtqfgPRq9ROijZZ1wIUlw7rPGF0hcNfCu2dViVV4NYep85twdkb01CCSV7wr 5LzGBvkUx8M0vZPNJavpmA1lEY8+Xmei4gP0gtjC4GUILlPtElB/7LhUvA8qzAb7tdsg5E 7biggFc6E1CeiLviqvxjkxXH+5n2eXOrJn0eeU3BOoJ1aKI9U9fZnq2ihJeIUH3ydyPOzS 4VHbf0mHgOLh5uk1YDcHD+S5UsCKHaTQT7LiHh6qZtmj8vI+p2Ve1aGavVgw X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Wed, 12 Aug 2026 09:23:52 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v3 17/18] selftests/mm: add memory-pressure threads to the khugepaged race harness Date: Wed, 12 Aug 2026 14:23:03 +0100 Message-ID: <20260812132304.199287-18-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260812132304.199287-1-kirill@shutemov.name> References: <20260812132304.199287-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The harness races collapse against faults, pins, fork, mremap and MADV_DONTNEED, but nothing in it elevates a source folio's refcount from the reclaim or compaction side. Add two more threads, and run every mode and occupancy limit both with and without them: - pageout: cycles MADV_PAGEOUT over a dedicated neighbour region, faults it back in and checks the content each round, since a page's pattern must survive the trip through swap. Idle when the host has no swap, because then there is no anon reclaim to drive. - compactor: writes /proc/sys/vm/compact_memory in a loop. Compaction isolates and migrates folios, so it competes with a collapse for the pages it is gathering, with refcount elevations and migration entries of its own. -p narrows a run to the combinations that have them, the way -z narrows the occupancy and -m the mode. Each result says which it ran: ok 2 stepped/strict/pressure: 5s, 88 steps, no corruption Assisted-by: Claude-Code:claude-opus-5 Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/khugepaged_race.c | 158 +++++++++++++++++-- 1 file changed, 144 insertions(+), 14 deletions(-) diff --git a/tools/testing/selftests/mm/khugepaged_race.c b/tools/testing/s= elftests/mm/khugepaged_race.c index ba4f5c36bf88..bde36af03e49 100644 --- a/tools/testing/selftests/mm/khugepaged_race.c +++ b/tools/testing/selftests/mm/khugepaged_race.c @@ -18,6 +18,12 @@ * zero-filled, and re-checked under the page table lock at install time in * case a racing fault got there first. * + * -p adds memory pressure to any of the above: MADV_PAGEOUT cycling + * on a dedicated neighbor region (swap traffic and LRU churn; skipped + * with a note when the host has no swap) and a compact_memory trigger + * loop (compaction migrates source folios, racing collapse's freeze + * with refcount elevation and migration entries of its own). + * * Correctness signals: every racing page must read as its pattern or * zero (MADV_DONTNEED), never anything else. The faulters and the fork * children check that continuously, a final sweep checks it once more, pl= us @@ -61,6 +67,8 @@ static unsigned long page_size; static char *region; /* NR_AREAS * hpage_pmd_size */ static char *mremap_area; /* region + NR_SHARED_AREAS areas */ static char *mremap_scratch; /* well above the region */ +static char *pageout_area; /* -p: dedicated pressure region */ +static size_t pageout_size; static int gup_fd =3D -1; static volatile int stop; static volatile int corrupted; @@ -219,6 +227,70 @@ static void *mremapper_fn(void *arg) return NULL; } =20 +/* + * -p: swap traffic and LRU churn on a region of our own. The content + * check is exact: a page out and back through swap must preserve the + * pattern, and nothing else ever writes here. + */ +static void *pageout_fn(void *arg) +{ + unsigned int seed =3D (unsigned long)arg; + unsigned long nr =3D pageout_size / page_size; + unsigned long i; + + for (i =3D 0; i < nr; i++) + *(unsigned int *)(pageout_area + i * page_size) =3D pattern(i); + + while (!stop) { + madvise(pageout_area, pageout_size, MADV_PAGEOUT); + for (i =3D 0; i < nr && !stop; i++) { + unsigned int val =3D *(unsigned int *)(pageout_area + + i * page_size); + + if (val !=3D pattern(i)) { + corrupted =3D 1; + ksft_print_msg("Pageout corruption at page %lu: %#x !=3D %#x\n", + i, val, pattern(i)); + } + } + usleep(rand_r(&seed) % 2000); + } + return NULL; +} + +/* -p: compaction migrates the collapse sources out from under us. */ +static void *compactor_fn(void *arg) +{ + unsigned int seed =3D (unsigned long)arg; + int fd =3D open("/proc/sys/vm/compact_memory", O_WRONLY); + + if (fd < 0) { + ksft_print_msg("No compact_memory; compactor idle\n"); + return NULL; + } + while (!stop) { + if (write(fd, "1", 1) < 0) + break; + usleep(10000 + rand_r(&seed) % 100000); + } + close(fd); + return NULL; +} + +static bool swap_available(void) +{ + char line[256]; + int lines =3D 0; + FILE *fp =3D fopen("/proc/swaps", "r"); + + if (!fp) + return false; + while (fgets(line, sizeof(line), fp)) + lines++; + fclose(fp); + return lines > 1; +} + static unsigned long now_ms(void) { struct timeval tv; @@ -230,11 +302,14 @@ static unsigned long now_ms(void) static void usage(void) { fprintf(stderr, - "Usage: khugepaged_race [-d seconds] [-m stepped|free|madvise] [-z] [-a = areas]\n" + "Usage: khugepaged_race [-d seconds] [-m stepped|free|madvise] [-z] [-p]= [-a areas]\n" "\tWithout -m, every mode runs in turn.\n" "\t-d: seconds per mode (default 5)\n" "\tBoth occupancy limits run unless -z asks for holes only.\n" "\t-z: only max_ptes_none =3D HPAGE_PMD_NR - 1 (hole-heavy)\n" + "\tRuns with and without memory pressure unless -p asks for\n" + "\tpressure only.\n" + "\t-p: only with the pageout and compaction threads\n" "\t-a: number of shared PMD-sized playground areas (default 3)\n"); exit(1); } @@ -243,18 +318,22 @@ int main(int argc, char **argv) { static const char * const thread_names[] =3D { "faulter", "faulter2", "dontneed", "pinner", "forker", - "mremapper", + "mremapper", "pageout", "compactor", }; void *(*const thread_fns[])(void *) =3D { faulter_fn, faulter_fn, dontneed_fn, pinner_fn, forker_fn, - mremapper_fn, + mremapper_fn, pageout_fn, compactor_fn, }; + const unsigned long pageout_bit =3D 1UL << 6, compactor_bit =3D 1UL << 7; const int nr_threads =3D ARRAY_SIZE(thread_names); pthread_t threads[ARRAY_SIZE(thread_names)]; static const char * const all_modes[] =3D { "stepped", "free", "madvise" = }; static const int all_nones[] =3D { 0, 1 }; /* strict, holes */ + static const int all_press[] =3D { 0, 1 }; /* quiet, under pressure */ const int *nones =3D all_nones; + const int *press =3D all_press; int nr_nones =3D ARRAY_SIZE(all_nones); + int nr_press =3D ARRAY_SIZE(all_press); const char *one_mode[1]; const char * const *modes =3D all_modes; int nr_modes =3D ARRAY_SIZE(all_modes); @@ -263,13 +342,15 @@ int main(int argc, char **argv) unsigned long end_ms; int duration_s =3D 5; unsigned long thread_mask =3D ~0UL; + unsigned long base_mask; int nr_areas_arg =3D 0; bool holes_only =3D false; + bool pressure_only =3D false; unsigned long i; int steps =3D 0; int opt; =20 - while ((opt =3D getopt(argc, argv, "a:d:m:t:zh")) !=3D -1) { + while ((opt =3D getopt(argc, argv, "a:d:m:t:zph")) !=3D -1) { switch (opt) { case 'a': nr_areas_arg =3D atoi(optarg); @@ -287,6 +368,9 @@ int main(int argc, char **argv) case 'z': holes_only =3D true; break; + case 'p': + pressure_only =3D true; + break; default: usage(); } @@ -296,6 +380,11 @@ int main(int argc, char **argv) nr_nones =3D 1; } =20 + if (pressure_only) { + press =3D all_press + 1; + nr_press =3D 1; + } + if (mode_arg) { if (strcmp(mode_arg, "stepped") && strcmp(mode_arg, "free") && strcmp(mode_arg, "madvise")) @@ -335,7 +424,12 @@ int main(int argc, char **argv) -1, 0) !=3D (void *)mremap_scratch) ksft_exit_fail_perror("mmap() mremap scratch"); =20 - ksft_set_plan(nr_modes * nr_nones); + base_mask =3D thread_mask; + if (!swap_available()) + /* No swap, no anon reclaim: compaction-only pressure. */ + ksft_print_msg("no swap: the pageout thread stays idle\n"); + + ksft_set_plan(nr_modes * nr_nones * nr_press); =20 thp_save_settings(); thp_read_settings(&settings); @@ -347,9 +441,17 @@ int main(int argc, char **argv) */ thp_push_settings(&settings); =20 - for (int mn =3D 0; mn < nr_modes * nr_nones; mn++) { - const char *mode =3D modes[mn / nr_nones]; - bool holes =3D nones[mn % nr_nones]; + for (int run =3D 0; run < nr_modes * nr_nones * nr_press; run++) { + int rem =3D run % (nr_nones * nr_press); + const char *mode =3D modes[run / (nr_nones * nr_press)]; + bool holes =3D nones[rem / nr_press]; + bool pressure =3D press[rem % nr_press]; + + thread_mask =3D base_mask; + if (!pressure) + thread_mask &=3D ~(pageout_bit | compactor_bit); + else if (!swap_available()) + thread_mask &=3D ~pageout_bit; =20 thp_read_settings(&settings); settings.thp_enabled =3D THP_MADVISE; @@ -385,6 +487,24 @@ int main(int argc, char **argv) ksft_exit_fail_perror("mmap() playground"); mremap_area =3D region + nr_shared_areas * hpage_pmd_size; =20 + if (thread_mask & pageout_bit) { + /* + * Big enough to cycle real reclaim, small enough not + * to dominate a TCG guest: 4 PMD areas, clamped to + * [16M, 64M]. + */ + pageout_size =3D 4 * hpage_pmd_size; + pageout_size =3D pageout_size < (16UL << 20) ? + (16UL << 20) : + pageout_size > (64UL << 20) ? + (64UL << 20) : pageout_size; + pageout_area =3D mmap(NULL, pageout_size, + PROT_READ | PROT_WRITE, + MAP_ANONYMOUS | MAP_PRIVATE, -1, 0); + if (pageout_area =3D=3D MAP_FAILED) + ksft_exit_fail_perror("mmap() pageout area"); + } + /* Populate so the first pass has something to collapse. */ for (i =3D 0; i < nr_shared_areas * hpage_pmd_size / page_size; i++) *(unsigned int *)(region + i * page_size) =3D pattern(i); @@ -434,8 +554,9 @@ int main(int argc, char **argv) check_page(i); =20 ksft_test_result(!corrupted, - "%s/%s: %ds, %d steps, no corruption\n", + "%s/%s%s: %ds, %d steps, no corruption\n", mode, holes ? "holes" : "strict", + pressure ? "/pressure" : "", duration_s, steps); =20 /* @@ -444,17 +565,26 @@ int main(int argc, char **argv) * and its scan cadence differs. */ munmap(region, nr_areas * hpage_pmd_size); + if (pageout_area) { + munmap(pageout_area, pageout_size); + pageout_area =3D NULL; + } thp_pop_settings(); stop =3D 0; steps =3D 0; =20 if (corrupted) { /* Memory is suspect; the rest would prove nothing. */ - while (++mn < nr_modes * nr_nones) - ksft_test_result_skip("%s/%s: skipped after corruption\n", - modes[mn / nr_nones], - nones[mn % nr_nones] ? - "holes" : "strict"); + while (++run < nr_modes * nr_nones * nr_press) { + rem =3D run % (nr_nones * nr_press); + + ksft_test_result_skip("%s/%s%s: skipped after corruption\n", + modes[run / (nr_nones * nr_press)], + nones[rem / nr_press] ? + "holes" : "strict", + press[rem % nr_press] ? + "/pressure" : ""); + } break; } } --=20 2.54.0 From nobody Tue Sep 29 04:44:48 2026 Received: from fhigh-b1-smtp.messagingengine.com (fhigh-b1-smtp.messagingengine.com [202.12.124.152]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5761E45040B; Wed, 12 Aug 2026 13:23:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=202.12.124.152 ARC-Seal: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541039; cv=none; b=CqZI7Q5Y48g/4WKCvLNKVppWC9VYELT1Da2LvLPaK8lKviePH8dwYpjTNe06bsmYxRe1OvVIYlNw1li0SZPjnqTAsfPI57KzHjdKI/Q0NjE5dUQrghtAr2kRX5lX7GPWL2am1OdQs6aChACnwvw8r61X4YuoY6fn4Dd1Gu/I7/k= ARC-Message-Signature: i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786541039; c=relaxed/simple; bh=Syn9bBoQ4/V3bvTB3ND/k09AkQOy6hRO4+fSt1PnXy0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=MrVP1av/gVoSOrbRdLtYLlQmZH1ExnvDVcE8FnYyklj0phFvZnzhnhflF73lr/bhcmZLxyQMVzb8UyrhtdwBacVIuDGF7mxnp1BJe9ZJ0GTmCrCxHbNh+6S7JN4kDxZvk2RoqjPKaYCe1SEXfiRkbpgsZDRk7P1pVw8podzV0QA= ARC-Authentication-Results: i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name; spf=pass smtp.mailfrom=shutemov.name; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b=WAI67wcs; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=XUxiESsR; arc=none smtp.client-ip=202.12.124.152 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=shutemov.name Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=shutemov.name Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=shutemov.name header.i=@shutemov.name header.b="WAI67wcs"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="XUxiESsR" Received: from phl-compute-12.internal (phl-compute-12.internal [10.202.2.52]) by mailfhigh.stl.internal (Postfix) with ESMTP id 9E37A7A015F; Wed, 12 Aug 2026 09:23:55 -0400 (EDT) Received: from phl-frontend-04 ([10.202.2.163]) by phl-compute-12.internal (MEProxy); Wed, 12 Aug 2026 09:23:56 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=shutemov.name; h=cc:cc:content-transfer-encoding:content-type:date:date:from :from:in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to; s=fm1; t=1786541035; x= 1786627435; bh=kgwf8/Ke7VGpiXTGkaQbuCrVb2lFqUgVIPPrp0NabqI=; b=W AI67wcsCbhXxN5ySnYvIchAVFeBeAgIPa1rrr+fIiwwm+gg7sjzHHuwpOoYBxQ8n ChQjU+C0cnT+H5uB9aBaf6Y4/7ocb/pfwlyDUvgleidJmL7Az8Ml12zWteVeWwN9 De7yN6Vhk9lBCN9RBLhjfE+o1idleXplTIf52+qyKYJd6wYTt8mOlMyohVvuubzW i1ivBsbSmi3JIEnLqho2FDrD65Y/tVYg8kRb1ZFbT91oDCo3Ie0hg8chK3cs3+bd w0M4Euzgcyb4sK4q1H84DJarbF65dG0h8Eye7Ramzxi/r102sF44kvKmjrDixUvv ZkA2O25Xto2tDTbvWkd0A== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:date:date:feedback-id:feedback-id:from:from :in-reply-to:in-reply-to:message-id:mime-version:references :reply-to:subject:subject:to:to:x-me-proxy:x-me-sender :x-me-sender:x-sasl-enc; s=fm3; t=1786541035; x=1786627435; bh=k gwf8/Ke7VGpiXTGkaQbuCrVb2lFqUgVIPPrp0NabqI=; b=XUxiESsRJ2NdKdLAJ zyc0sb7bj50PGVkpd3+Posb5XFYM8rOQA3LbqIHz620YgpUPsVaeuSPK4xKAuVO6 UkJFV2EBGsTdi9noQoWqnaMDoj9ssfDdp4zHCNADMMR87GR3Pi1kVQmrPHQe/GXb tKNEidYXA51rlxQI8dtdyqhS768ADDYnggIf57Ug1hJJcHYWbYXcv3anlaKBhYBC Z5agvmpiLMPiwVeoky7L/FxcoLmBseA3SCBNAsXM1y8u25hWlwwUKD/uDrE0zxCH Fb+NkHqVgGpSgdxL0BZ9O7n3NwSwqFzKmhv24TfL6RPho1OxpoTiWKG1VpECxvqQ oxe5g== X-ME-Sender: X-ME-Received: X-ME-Proxy-Cause: dmFkZTFzqtrNoSTXX2MRzKyN6p8XXqdjDSX08KYyJpXp82eYC1ae2GegaTOhbDzreYT/K0 OAsc1zsfjWszwuQXvOMxV3FM9x50+HqAsA48u5Kq52QtHo+TJJPCEG+EOeKn5Il8AepNuY amkSO7LHncjzMJKy/SPUs50UUrTlzGP0ayf7b+h/q5NreZMQSrrfXSzcd6tsdBCg4Hhmbc 7COQKDHpqWJD3HpoJ7/Bw0J3tcRV9tZIkmRGQ4nZ/O76iTBc+eN6GYtayu9aLLTJal0DNN hI5LqlugC7KhSFwlZIIFqKKnJB7XY7vZnEBEdRJ8Mvpq9sE5OflKN5J+H3ujbniJsnhlZt AO8s3IAdX3RECZ65RDz6vZeKF9y7e5/y26Zz/FaKwgKRPbu3b56Zt0D9mkA5eY/On1BDUN I9ma2NHB/D+64EcN3N7LxO3fy2AgeqQ31XBqVh/XaUCXD1q9B5JOs7C72aocdlKdm/FV5i Ns9weHmhUOW4IdfeIeNN/rOuwSH9MM0OsOa+KOf0Vn5HvZG9aU1ZOYDUy0RLkBhiG/nxOo MIblt/EzJD+PYQjYcMDVjPPBq+vd4tb14ejXbM5KGhtJdmIwOX2nHCfpMA7zZxGXjpQQyX 0ON8a6SdRxRGhFE1GKNFxde/DmqcwOVMa4FV5LNzOODt1ffD7KWFqyklEEUA X-ME-Proxy: Feedback-ID: ie3994620:Fastmail Received: by mail.messagingengine.com (Postfix) with ESMTPA; Wed, 12 Aug 2026 09:23:54 -0400 (EDT) From: Kiryl Shutsemau To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, nico.pache@linux.dev Cc: baolin.wang@linux.alibaba.com, baohua@kernel.org, dev.jain@arm.com, hughd@google.com, lance.yang@linux.dev, liam@infradead.org, mhocko@suse.com, rppt@kernel.org, ryan.roberts@arm.com, shuah@kernel.org, surenb@google.com, usama.arif@linux.dev, vbabka@kernel.org, ziy@nvidia.com, usama.anjum@arm.com, agordeev@linux.ibm.com, linux-mm@kvack.org, linux-kselftest@vger.kernel.org, linux-kernel@vger.kernel.org, kas@kernel.org Subject: [PATCH v3 18/18] selftests/mm: zap whole PTE tables in the khugepaged race harness Date: Wed, 12 Aug 2026 14:23:04 +0100 Message-ID: <20260812132304.199287-19-kirill@shutemov.name> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260812132304.199287-1-kirill@shutemov.name> References: <20260812132304.199287-1-kirill@shutemov.name> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="utf-8" From: "Kiryl Shutsemau (Meta)" The harness's MADV_DONTNEED thread zaps 1 to 32 pages at a time, never a whole PMD-aligned area. Empty-table reclaim (CONFIG_PT_RECLAIM) only engages when a zap spans a full table, so no soak ever ran it against a collapse -- fuzzing had to find that class instead: the page table vanishing between the engine's park and install passes. Make the thread zap a whole PMD-aligned area once every 64 iterations, and keep the fine-grained zaps as the common case. Assisted-by: Claude-Code:claude-opus-5 Tested-by: Muhammad Usama Anjum Signed-off-by: Kiryl Shutsemau (Meta) --- tools/testing/selftests/mm/khugepaged_race.c | 20 ++++++++++++++++++-- 1 file changed, 18 insertions(+), 2 deletions(-) diff --git a/tools/testing/selftests/mm/khugepaged_race.c b/tools/testing/s= elftests/mm/khugepaged_race.c index bde36af03e49..3601c030d43e 100644 --- a/tools/testing/selftests/mm/khugepaged_race.c +++ b/tools/testing/selftests/mm/khugepaged_race.c @@ -141,8 +141,24 @@ static void *dontneed_fn(void *arg) unsigned long page_idx =3D rand_page(&seed); unsigned long nr =3D 1UL << (rand_r(&seed) % 6); /* 1..32 pages */ =20 - madvise(region + page_idx * page_size, - room_from(page_idx, nr) * page_size, MADV_DONTNEED); + /* + * Once in a while zap a whole PMD-aligned area: only a zap + * spanning the full table triggers the empty-table reclaim + * (CONFIG_PT_RECLAIM), which can free the table under a + * collapse that is midway through it. Sub-table zaps never + * reach that path. + */ + if (!(rand_r(&seed) % 64)) { + unsigned long area =3D page_idx / + (hpage_pmd_size / page_size); + + madvise(region + area * hpage_pmd_size, + hpage_pmd_size, MADV_DONTNEED); + } else { + madvise(region + page_idx * page_size, + room_from(page_idx, nr) * page_size, + MADV_DONTNEED); + } usleep(rand_r(&seed) % 500); } return NULL; --=20 2.54.0